VLDB 2026 Research / reviewers in the wild / expert
Rui Sun 0004
dblp:01/3595-4
· DBLP profile ↗
15ranked-venue papers
10as first author
12since 2021 · last 2026
0000-0002-1547-161XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalized adversarial feature aggregation and KAN-enhanced network for semi-supervised visible-infrared person re-identification
Rui Sun 0004, Jicheng Shen, Guoxi Huang, Jingjing Wu 0001 |
Image Vis. Comput. | 1 |
| 2026 | Semantic-Aware and Semi-Fragile Diffusion Watermarking for Proactive Deepfake DetectionabstractThe rapid progress of deepfake technology, which primarily manipulates facial identity and image semantics, has made detection and defense critically important. Conventional global watermarking methods offer limited capacity for protecting key semantic content, as they typically rely on uniformly distributed watermarks across the entire image. This letter presents a method that weave watermarks as intrinsic components into the semantic content of images (facial regions) in the latent space. By aligning watermark embedding regions with facial content, we establish an inherent fragility mechanism wherein any deepfake manipulation that modifies facial semantics inevitably disrupts the watermark, enabling precise detection. Simultaneously, adversarial training of the extractor ensures robustness against conventional signal processing operations. A local entropy perception module dynamically adjusts embedding intensity based on regional texture complexity, maintaining high perceptual fidelity. Extensive experiments indicate that compared to advanced methods, the proposed approach maintains robustness against conventional benign operations while achieving reliable detection of deepfake forgeries, thereby enabling precise protection of image semantic content. Rui Sun 0004, Xiaolu Yu, Yuwei Dai, Yaofei Wang |
IEEE Signal Process. Lett. | 1 |
| 2026 | Learning Corruption-Invariant Components and Cross-Modal Correspondence for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised Visible-Infrared Person Re-Identification (US-VI-ReID) has great potential prospects because it does not require label information. However, corrupted pedestrian images collected due to corruption factors in real-world scenarios (e.g., noise, blur, and weather changes) largely limit the scalability of US-VI-ReID. In this paper, we explore the robustness of US-VI-ReID for the first time and propose a Multi-Granularity Spatial-Frequency Prototype Learning (MSPL) framework. The framework mainly consists of Multi-Channel Soft Augmentation (MSA), Robust Frequency Domain Feature Learning (RFL) module and Cross-modal Spatial-Frequency Prototype Matching (CSPM). Specifically, the MSA alleviates the sensitivity of model to color and abnormal samples through rich channel combinations and soft erasing. Subsequently, the RFL performs deep global filtering and amplitude attention compensated InstanceNorm to complete frequency and style modulation, concentrating on degradation-robust frequency content. Finally, the CSPM is designed to achieve multi-granularity prototype contrastive learning on cluster level and view level, then conduct cross-modal matching of multi-granularity spatial-frequency prototypes, thus establishing robust label association. With the above modules, our proposed framework can learn corruption-invariant feature components and generate robust cross-modal correspondence from unlabeled cross-modal images. Extensive experiments demonstrate that our MSPL outperforms other state-of-the-art methods by a large margin on the challenging SYSU-MM01-C and RegDB-C, while maintaining competitive on the SYSU-MM01 and RegDB. Rui Sun 0004, Guoxi Huang, Jingjing Wu 0001, Wei Jia 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Self-supervised polarization image dehazing method via frequency domain generative adversarial networks
Rui Sun 0004, Tanbin Liao, Zhiguo Fan |
Pattern Recognit. | 1 |
| 2025 | Robust multimodal face anti-spoofing via frequency-domain feature refinement and aggregation
Rui Sun 0004, Xiaolu Yu, Xinjian Gao |
Pattern Recognit. Lett. | 1 |
| 2025 | Implicit Alignment-Based Cross-Modal Symbiotic Network for Text-to-Image Person Re-IdentificationabstractText-to-image person re-identification aims to utilize textual descriptions to retrieve specific person images from large image databases. The core challenge of this task lies in the significant feature differences between the abstract nature of text and the intuitiveness of images. Existing solutions primarily rely on explicit alignment of global or fine-grained local features, which lack flexibility and struggle to effectively capture and leverage subtle features and relationship information in multimodal data. Particularly, for different images of the same person, the emphasis in feature extraction should be adjusted according to the differences in text descriptions. To address these issues, this paper proposes a Cross-Modal Symbiotic Network (CMSN) based on implicit alignment. First, CMSN employs an Implicit Multi-scale Feature Integration (IMFI) module to implicitly extract and fuse multiscale features from images and text, thereby adaptively capturing the feature relationships between the two modalities. Second, a Combined Representation Learning (CRL) module is used to produce a combined representation of the text and image features, utilizing a Combined-Representation Identity Alignment (CRIA) loss to align and constrain the identity centers of the three feature vectors. Finally, we design a Semi-Positive Triplet (SPT) loss function, which defines semi-positive samples using other images and texts of the same identity, providing additional supervisory information to the model and further reducing modality heterogeneity. Extensive experiments on the CUHK-PEDES dataset demonstrate that CMSN achieves an impressive Rank-1 and mAP accuracy of 76.46% and 70.28%, respectively, significantly outperforming existing SOTA methods. Rui Sun 0004, Guoxi Huang, Jingjing Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Visible thermal person re-identification via multi-branch modality residual complementary learning
Rui Sun 0004, Yiheng Yu |
Image Vis. Comput. | 2 |
| 2024 | Text-augmented Multi-Modality contrastive learning for unsupervised visible-infrared person re-identification
Rui Sun 0004, Guoxi Huang |
Image Vis. Comput. | 1 |
| 2024 | Diffusion Augmentation and Pose Generation Based Pre-Training Method for Robust Visible-Infrared Person Re-IdentificationabstractCross-Modal Visible-Infrared Person Re-identification (VI-REID) constitutes a vital application for constructing all-time surveillance systems. However, the current VI-REID model exhibits significant performance deterioration in noisy environments. Existing algorithms endeavor to mitigate this challenge through fine-tuning stages. We contend that, in contrast to fine-tuning stages, the pre-training phase can effectively exploit the attributes of extensive unlabeled data, thereby facilitating the development of a robust VI-REID model. Therefore, in this paper, we propose a pre-training method for VI-REID based on Diffusion Augmentation and Pose Generation (DAPG), aiming to enhance the robustness and recognition rate of VI-REID models in the presence of damaged scenes. Multiple transfer experiments on the SYSU-MM01 and RegDB datasets demonstrate that our method outperforms existing self-supervised methods, as evidenced by the results. Rui Sun 0004, Guoxi Huang, Ruirui Xie |
IEEE Signal Process. Lett. | 1 |
| 2024 | Robust Visible-Infrared Person Re-Identification Based on Polymorphic Mask and Wavelet Graph Convolutional NetworkabstractWhen deploying re-identification (ReID) models in the field of public safety, understanding the robustness of models to various types of corrupted images is crucial. Unfortunately, in the real world, images are always contaminated (e.g., noise, blur, and weather changes), which is ignored by existing visible-infrared person re-identification (VI-ReID) models. The performance of existing models tested in corrupted scenes is severely degraded. Therefore, learning corruption-invariant representations for corrupted images in VI-ReID is valuable and deserves further investigation. We design a polymorphic masked wavelet graph convolutional network for VI-ReID under corrupted scenes. Firstly, a cross-modality data augmentation algorithm is designed to construct a mixed image set that merges multi-modality attributes to improve robustness against interference. Secondly, a dual-branch network consisting of a global branch and a graph structure branch is designed. The global branch extracts overall information. While the graph structure branch is a wavelet-based graph convolutional module that utilizes the robustness of human structural information to corruptions and modalities, it can filter noise and extract discriminative features specifically targeted for cross-modality scenes. Finally, the global branch and the graph structure branch are integrated, and modality consistency loss is designed to match the branches with hetero-center triplet loss. Experiments show that our method can effectively alleviate degradation problems under corrupted environments such as noise, blur, digitization, and weather changes, and achieve state-of-the-art on corrupted datasets. Besides, it still maintains good performance on clean datasets, facilitating the reliable deployment of VI-ReID in real-world scenarios. Rui Sun 0004, Ruirui Xie, Jun Gao 0006 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Multi-cues underwater image restoration algorithm combined with light field technologyabstractAbstract Underwater images suffer from color distortion, low clarity and halos problems due to light absorption, particle scattering and non‐uniform illumination. To address these degradation issues, a multi‐cues underwater image restoration algorithm combined with light field technology is proposed. First, based on Epipolar Plane Image, the light field cue transmittance containing depth information is calculated. Then, according to the turbidity of the underwater image, the light field cue transmittance and the polarization cue transmittance are fused to obtain the multi‐cues transmittance, which can effectively reduce the effect of particle scattering and color bias. Finally, the background light is estimated through the all‐focus operation, which can effectively overcome the distortion of an underwater single image and simultaneously reduce the halo phenomenon. Experimental results show that the method achieves the best results evaluated by UCIQE, UIQM, PSNR, and SSIM, and the restored color under the method is closer to the actual image than other underwater restoration methods. Liwen Cui, Zhiguo Fan, Rui Sun 0004 |
IET Image Process. | 4 |
| 2022 | Data gap decomposed by auxiliary modality for NIR-VIS heterogeneous face recognitionabstractAbstract In the dark scene at night, the face images captured by ordinary visible light (VIS) are generally poor quality and very dim, while the near‐infrared (NIR) can capture high definition and recognizable face images at night. The NIR‐VIS Heterogeneous face recognition has become a hot research field, which helps to build an all‐weather face recognition system. NIR‐VIS HFR is sophisticated because of the large visual difference between NIR images and VIS images. In order to reduce the difficulty of such cross‐modality invariant feature learning, this paper proposes a cross‐modality data gap decomposed by auxiliary modality method (DGD) for NIR‐VIS HFR. First, the brightness component (Y component) of VIS image YCbCr space is used as the auxiliary modality to decompose the cross‐modality data gap. The lightness component retained the structural information of VIS image and was similar to the colour information of NIR modality; in this way, the huge gap between the NIR data and the VIS data is decomposed into two smaller gaps, thus reducing the difficulty of network learning. Second, the data of the three modalities are input into the weight sharing network and training under the combined guidance of cross‐modality gap decomposition loss and intra‐modality gap loss; in this way, the modality invariant features can be learned faster and better. Extensive experiments were conducted on two commonly used datasets CASIA NIR‐VIS 2.0 and Oulu‐CASIA NIR‐VIS to evaluate DGD method. Experimental results indicate DGD method has competitive performance compared with the latest methods. Rui Sun 0004, Xiaoquan Shan, Jun Gao 0006 |
IET Image Process. | 1 |
| 2020 | EPI-based Oriented Relation Networks for Light Field Depth Estimation
Kunyuan Li, Jun Zhang 0017, Rui Sun 0004, Jun Gao 0006 |
BMVC | 3 |
| 2020 | Unsupervised video summarization via clustering validity index
Ye Zhao 0001, Yanrong Guo, Rui Sun 0004, Zhengqiong Liu, Dan Guo 0001 |
Multim. Tools Appl. | 3 |
| 2019 | Robust visual tracking based on convolutional neural network with extreme learning machine
Rui Sun 0004, Xiaoxing Yan |
Multim. Tools Appl. | 1 |