VLDB 2026 Research / reviewers in the wild / expert
Shenghua Fan
dblp:345/6824
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-3062-3028ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse Tuning Enhances Plasticity in PTM-based Continual LearningabstractContinual Learning with Pre-trained Models holds great promise for efficient adaptation across sequential tasks. However, most existing approaches freeze PTMs and rely on auxiliary modules like prompts or adapters, limiting model plasticity and leading to suboptimal generalization when facing significant distribution shifts. While full fine-tuning can improve adaptability, it risks disrupting crucial pre-trained knowledge. In this paper, we propose Mutual Information-guided Sparse Tuning (MIST), a plug-and-play method that selectively updates a small subset of PTM parameters, less than 5%, based on sensitivity to mutual information objectives. MIST enables effective task-specific adaptation while preserving generalization. To further reduce interference, we introduce strong sparsity regularization by randomly dropping gradients during tuning, resulting in fewer than 0.5% of parameters being updated per step. Applied before standard freeze-based methods, MIST consistently boosts performance across diverse continual learning benchmarks. Experiments show that integrating our method into multiple baselines yields significant performance gains. Shenghua Fan, Shuyu Dong, Yujin Zheng, Dingwen Wang, Fan Lyu |
AAAI | 2 |
| 2026 | Constructing Enhanced Mutual Information for Online Class-Incremental Learning
Fan Lyu, Shenghua Fan, Yujin Zheng, Dingwen Wang |
IEEE Trans. Multim. | 3 |
| 2025 | Wavelet-based Feature Representation Framework for Event Stream RecognitionabstractEvent streams generated by event cameras exhibit low data redundancy and preserve precise temporal information through Address Event Representation (AER), which differs significantly from the outputs of traditional frame-based cameras. However, traditional Artificial Neural Networks (ANNs) and Spiking Neural Networks (SNNs) often convert these streams into static frames, which can lead to some loss of temporal features. In this work, we propose a novel wavelet-based framework, Tempo-Spatial Wavelet Transform (TSWT) framework, designed to optimize both ANNs and SNNs. The Temporal Wavelet Transform (TWT) Module within this framework produces a wavelet-based event representation that captures more precise temporal correlations from the event streams. Additionally, the Spatial Wavelet Transform (SWT) Module provides a wavelet-based method for selecting frequency features. By integrating these two modules with a tailored loss function, TSWT enhances the network’s ability to learn tempo-spatial information effectively. Experimental results demonstrate that both ANNs and SNNs equipped with the proposed TSWT framework achieve superior performance across various datasets. Xingyu Pan, Xi Chen 0078, Shenghua Fan |
ICME | 4 |
| 2025 | RDD-DETR: An End-to-End Real-Time Road Damage Detection Algorithm
Ziang He, Zhuoxuan Zhao, Shenghua Fan |
PRCV (16) | 3 |
| 2025 | NOT-156: Night Object Tracking Using Low-Light and Thermal Infrared: From Multimodal Common-Aperture Camera to Benchmark DatasetsabstractNight object tracking (NOT) is aimed at tracking objects under low-illumination conditions at night. Existing works concentrate on thermal infrared modality, while some RGB and thermal infrared (RGB-T) data also contain night scenes. However, night scenes in these datasets are mostly well-lit, making it challenging to fully cover low-illumination scenarios. In this article, we focus on the NOT task and build up a novel low-light visible and thermal infrared (LOL-T) multimodal benchmark dataset for NOT-156. To achieve night vision, we design a common-aperture LOL-T camera by integrating a highly dynamic low-light visible imaging sensor with a thermal infrared sensor in a common aperture optical system. The proposed dataset consists of 156 video sequences and a total of 170k annotated frames, including various low-illumination night scenes such as dark rooms, streets, corridors, and so on. Compared with existing datasets, NOT-156 has more comprehensive and distinctive attributes (thermal variation, noise, high illumination overexposure, etc.). Comprehensive experiments are carried out to evaluate the performance of the advanced visible, infrared, and visible-thermal trackers on the proposed NOT-156 dataset. The authors believe that NOT-156 has great potential in the application and development of night vision. The dataset will be made available athttp://rsidea.whu.edu.cn/NOT156_dataset.htm. Xinyu Wang 0003, Shenghua Fan, Xiaobing Dai, Yuting Wan, Zengliang Zhu, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Cross-scale content-based full Transformer network with Bayesian inference for object tracking
Shenghua Fan, Xi Chen 0078, Chu He |
Multim. Tools Appl. | 1 |
| 2023 | Multiple frequency-spatial network for RGBT tracking in the presence of motion blur
Shenghua Fan, Xi Chen 0078, Chu He, Lei Yu 0006, Zhongjie Mao, Yujin Zheng |
Neural Comput. Appl. | 1 |
| 2023 | Bayesian Dumbbell Diffusion Model for RGBT Object Tracking With Enriched PriorsabstractRGBT tracking can be accomplished by constructing Bayesian estimators that incorporate fusion prior distributions for the visible (RGB) and thermal (T) modalities. Such estimators enable the computation of a posterior distribution for the variables of interest to locate the target. Incorporating rich prior information can improve the performance of predictors. However, current RGBT trackers face limited fusion prior data. To mitigate this issue, we propose a novel tracker, BD$^{2}$Track, which employs a diffusion model. Firstly, this paper introduces a dumbbell diffusion model, and employ convolution networks and the dumbbell model to derive the fusion feature prior information from various index frames in the same tracking video sequence. Secondly, we propose a plug-and-play channel augmented joint learning strategy to derive the images prior distribution. This strategy not only homogeneously generates modality-relevant prior information but also increases the distance between positive and negative samples within the modality, while reducing the distance between modalities during fusion. Results demonstrate promising performance in the GTOT, RGBT234, LasHeR, and VTUAV-ST datasets, surpassing other state-of-the-art trackers. Shenghua Fan, Chu He, Chenxia Wei, Yujin Zheng, Xi Chen 0078 |
IEEE Signal Process. Lett. | 1 |