VLDB 2026 Research / reviewers in the wild / expert
Shangquan Sun
dblp:346/0940
· DBLP profile ↗
10ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0002-6292-2495ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 6 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Disentangle to Fuse: Toward Content Preservation and Cross-Modality Consistency for Multi-Modality Image FusionabstractMulti-modal image fusion (MMIF) aims to integrate complementary information from heterogeneous sensor modalities. However, substantial cross-modality discrepancies hinder joint scene representation and lead to semantic degradation in the fused output. To address this limitation, we propose C2MFuse, a novel framework designed to preserve content while ensuring cross-modality consistency. To the best of our knowledge, this is the first MMIF approach to explicitly disentangle style and content representations across modalities for image fusion. C2MFuse introduces a content-preserving style normalization mechanism that suppresses modality-specific variations while maintaining the underlying scene structure. The normalized features are then progressively aggregated to enhance fine-grained details and improve content completeness. In light of the lack of ground truth and the inherent ambiguity of the fused distribution, we further align the fused representation with a well-defined source modality, thereby enhancing semantic consistency and reducing distributional uncertainty. Additionally, we introduce an adaptive consistency loss with learnable transformation, which provides dynamic, modality-aware supervision by enforcing global consistency across heterogeneous inputs. Extensive experiments on five datasets across three representative MMIF tasks demonstrate that C2MFuse achieves efficient and high-quality fusion, surpasses existing methods, and generalizes effectively to downstream visual applications. Xinran Qin, Yuning Cui 0001, Shangquan Sun, Ruoyu Chen 0001, Wenqi Ren, Alois C. Knoll, Xiaochun Cao |
IEEE Trans. Image Process. | 3 |
| 2025 | Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video DerainingabstractSignificant progress has been made in video restoration under rainy conditions over the past decade, largely propelled by advancements in deep learning. Nevertheless, existing methods that depend on paired data struggle to generalize effectively to real-world scenarios, primarily due to the disparity between synthetic and authentic rain effects. To address these limitations, we propose a dual-branch spatiotemporal state-space model to enhance rain streak removal in video sequences. Specifically, we design spatial and temporal state-space model layers to extract spatial features and incorporate temporal dependencies across frames, respectively. To improve multi-frame feature fusion, we derive a dynamic stacking filter, which adaptively approximates statistical filters for superior pixel-wise feature refinement. Moreover, we develop a median stacking loss to enable semi-supervised learning by generating pseudo-clean patches based on the sparsity prior of rain. To further explore the capacity of deraining models in supporting other vision-based tasks in rainy environments, we introduce a novel real-world benchmark focused on object detection and tracking in rainy conditions. Our method is extensively evaluated across multiple benchmarks containing numerous synthetic and real-world rainy videos, consistently demonstrating its superiority in quantitative metrics, visual quality, efficiency, and its utility for downstream tasks. Shangquan Sun, Wenqi Ren, Juxiang Zhou, Jianhou Gan, Xiaochun Cao |
CVPR | 1 |
| 2025 | NDFormer: A Mixed-Scale Transformer with Enhanced Nonlinearity for Nighttime Image Deraining
Zhirui Liu, Shangquan Sun, Yuning Cui 0001, Dehong Kong, Wenqi Ren, Kin-Man Lam 0001 |
PRCV (9) | 2 |
| 2025 | DI-Retinex: Digital-Imaging Retinex Model for Low-Light Image Enhancement
Shangquan Sun, Wenqi Ren, Jingyang Peng, Fenglong Song, Xiaochun Cao |
Int. J. Comput. Vis. | 1 |
| 2025 | NDMamba: Dual-Prior State-Space Model for Nighttime DerainingabstractRecent advancements in deep learning, particularly through Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), have led to significant progress in nighttime image deraining. However, current architectures still struggle to strike an optimal balance between computational efficiency and restoration performance. Moreover, existing methods often fail to fully exploit the intrinsic characteristics of low-light conditions and inadequately model the interaction between rain and illumination. To overcome these challenges, we propose NDMamba, a dual-prior-guided state-space model that addresses nighttime deraining by incorporating degradation cues related to both lighting and rain distribution. Inspired by the Retinex theory, which suggests that rain streak distribution is influenced by the reflectance component of a scene, we propose a Prior Extraction Module (PEM) to jointly model lighting conditions and rain degradation. Furthermore, we design a Prior-Guided Mamba Block (PGMB), which comprises a Lighting-Adaptive Vision State-Space Module (LVSSM) that incorporates illumination priors, and a Rain Distribution Guidance Module (RDGM) to enhance local features in a more refined manner. Extensive experiments demonstrate that NDMamba outperforms state-of-the-art methods on both synthetic and real-world benchmark datasets. Our code is publicly available at https://github.com/tandaily/NDMamba. Zhirui Liu, Shangquan Sun, Chaopeng Li, Wenqi Ren |
IEEE Trans. Image Process. | 2 |
| 2024 | Logit Standardization in Knowledge DistillationabstractKnowledge distillation involves transferring soft labels from a teacher to a student using a shared temperature-based softmax function. However, the assumption of a shared temperature between teacher and student implies a mandatory exact match between their logits in terms of logit range and variance. This side-effect limits the performance of student, considering the capacity discrepancy between them and the finding that the innate logit relations of teacher are sufficient for student to learn. To address this issue, we propose setting the temperature as the weighted standard deviation of logit and performing a plug-and-play Z-score pre-process of logit standardization before applying softmax and Kullback-Leibler divergence. Our pre-process enables student to focus on essential logit relationsfrom teacher rather than requiring a magnitude match, and can improve the performance of existing logit-based distillation methods. We also show a typical case where the conventional setting of sharing temperature between teacher and student cannot reliably yield the authentic dis-tillation evaluation; nonetheless, this challenge is success-fully alleviated by our Z-score. We extensively evaluate our method for various student and teacher models on CIFAR-100 and ImageNet, showing its significant superiority. The vanilla knowledge distillation powered by our pre-process can achieve favorable performance against state-of-the-art methods, and other distillation variants can obtain considerable gain with the assistance of our pre-process. The codes, pre-trained models and logs are released on Github. Shangquan Sun, Wenqi Ren, Jingzhi Li 0002, Rui Wang 0032, Xiaochun Cao |
CVPR | 1 |
| 2024 | Restoring Images in Adverse Weather Conditions via Histogram Transformer
Shangquan Sun, Wenqi Ren, Xinwei Gao, Rui Wang 0032, Xiaochun Cao |
ECCV (22) | 1 |
| 2024 | EnsIR: An Ensemble Algorithm for Image Restoration via Gaussian Mixture ModelsabstractImage restoration has experienced significant advancements due to the development of deep learning. Nevertheless, it encounters challenges related to ill-posed problems, resulting in deviations between single model predictions and ground-truths. Ensemble learning, as a powerful machine learning technique, aims to address these deviations by combining the predictions of multiple base models. Most existing works adopt ensemble learning during the design of restoration models, while only limited research focuses on the inference-stage ensemble of pre-trained restoration models. Regression-based methods fail to enable efficient inference, leading researchers in academia and industry to prefer averaging as their choice for post-training ensemble. To address this, we reformulate the ensemble problem of image restoration into Gaussian mixture models (GMMs) and employ an expectation maximization (EM)-based algorithm to estimate ensemble weights for aggregating prediction candidates. We estimate the range-wise ensemble weights on a reference set and store them in a lookup table (LUT) for efficient ensemble inference on the test set. Our algorithm is model-agnostic and training-free, allowing seamless integration and enhancement of various pre-trained image restoration models. It consistently outperforms regression-based methods and averaging ensemble approaches on 14 benchmarks across 3 image restoration tasks, including super-resolution, deblurring and deraining. The codes and all estimated weights have been released in Github. Shangquan Sun, Wenqi Ren, Zikun Liu 0001, Hyunhee Park, Rui Wang 0032, Xiaochun Cao |
NeurIPS | 1 |
| 2023 | Event-Aware Video Deraining via Multi-Patch Progressive LearningabstractIn this paper, we address the problem of video-based rain streak removal by developing an event-aware multi-patch progressive neural network. Rain streaks in video exhibit correlations in both temporal and spatial dimensions. Existing methods have difficulties in modeling the characteristics. Based on the observation, we propose to develop a module encoding events from neuromorphic cameras to facilitate deraining. Events are captured asynchronously at pixel-level only when intensity changes by a margin exceeding a certain threshold. Due to this property, events contain considerable information about moving objects including rain streaks passing though the camera across adjacent frames. Thus we suggest that utilizing it properly facilitates deraining performance non-trivially. In addition, we develop a multi-patch progressive neural network. The multi-patch manner enables various receptive fields by partitioning patches and the progressive learning in different patch levels makes the model emphasize each patch level to a different extent. Extensive experiments show that our method guided by events outperforms the state-of-the-art methods by a large margin in synthetic and real-world datasets. Shangquan Sun, Wenqi Ren, Jingzhi Li 0002, Kaihao Zhang, Meiyu Liang, Xiaochun Cao |
IEEE Trans. Image Process. | 1 |
| 2022 | Rethinking Image Restoration for Object DetectionabstractAlthough image restoration has achieved significant progress, its potential to assist object detectors in adverse imaging conditions lacks enough attention. It is reported that the existing image restoration methods cannot improve the object detector performance and sometimes even reduce the detection performance. To address the issue, we propose a targeted adversarial attack in the restoration procedure to boost object detection performance after restoration. Specifically, we present an ADAM-like adversarial attack to generate pseudo ground truth for restoration training. Resultant restored images are close to original sharp images, and at the same time, lead to better results of object detection. We conduct extensive experiments in image dehazing and low light enhancement and show the superiority of our method over conventional training and other domain adaptation and multi-task methods. The proposed pipeline can be applied to all restoration methods and detectors in both one- and two-stage. Shangquan Sun, Wenqi Ren, Tao Wang 0053, Xiaochun Cao |
NeurIPS | 1 |