Shicai Wei

dblp:300/6440 · DBLP profile ↗
← Back
11ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0001-5744-2035ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FreSCo: Joint Frequency-Aware and Spatial Control for Image Zero-Shot Style Transfer
abstract
Latent Diffusion Models (LDMs) have become a cornerstone for zero-shot style transfer in multimedia content creation, but they frequently struggle with a critical trade-off between artistic stylization fidelity and semantic structural preservation. A key oversight in existing methods is the neglect of frequency-domain distinctions in visual signals, which leads to prevalent issues like content drift and style leakage. To address these limitations, we propose FreSCo, a novel training-free framework that explicitly decouples content and style through dual-domain control mechanisms. First, the Dynamic Wavelet Latent Fusion (DWLF) module decomposes latent features via Discrete Wavelet Transform (DWT), injecting style exclusively into high-frequency texture sub-bands while boosting spectral energy to counteract VAE-induced smoothing. Second, the VAE-Compressed Masking strategy encodes edge maps directly into the latent space, resolving pixel-latent misalignment for precise spatial control. We construct a comprehensive benchmark with 1,280 content-style pairs to rigorously evaluate performance. Extensive experiments demonstrate that FreSCo achieves state-of-the-art results, generating high-fidelity artistic textures while maintaining superior structural consistency across diverse multimedia content creation scenarios compared to existing baselines.
Tingrun Chen, Xudong Ling, Shicai Wei, Guiduo Duan, Yue Zhang 0042
ICMR3
2025 Improving Multimodal Learning via Imbalanced Learning
Shicai Wei, Chunbo Luo, Yang Luo 0001
ICCV1
2025 Boosting Multimodal Learning via Disentangled Gradient Learning
abstract
Multimodal learning often encounters the under-optimized problem and may have worse performance than unimodal learning. Existing methods attribute this problem to the imbalanced learning between modalities and rebalance them through gradient modulation. However, they fail to explain why the dominant modality in multimodal models also underperforms that in unimodal learning. In this work, we reveal the optimization conflict between the modality encoder and modality fusion module in multimodal models. Specifically, we prove that the cross-modal fusion in multimodal models decreases the gradient passed back to each modality encoder compared with unimodal models. Consequently, the performance of each modality in the multimodal model is inferior to that in the unimodal model. To this end, we propose a disentangled gradient learning (DGL) framework to decouple the optimization of the modality encoder and modality fusion module in the multimodal model. DGL truncates the gradient back-propagated from the multimodal loss to the modality encoder and replaces it with the gradient from unimodal loss. Besides, DGL removes the gradient back-propagated from the unimodal loss to the modality fusion module. This helps eliminate the gradient interference between the modality encoder and modality fusion module while ensuring their respective optimization processes. Finally, extensive experiments on multiple types of modalities, tasks, and frameworks with dense cross-modal interaction demonstrate the effectiveness and versatility of the proposed DGL. Code is available at \href{https://github.com/shicaiwei123/ICCV2025-GDL}{https://github.com/shicaiwei123/ICCV2025-GDL}
Shicai Wei, Chunbo Luo, Yang Luo 0001
ICCV1
2024 Scale Decoupled Distillation
abstract
Logit knowledge distillation attracts increasing attention due to its practicality in recent studies. However, it of-ten suffers inferior performance compared to the feature knowledge distillation. In this paper, we argue that existing log it-based methods may be sub-optimal since they only leverage the global logit output that couples multiple se-mantic knowledge. This may transfer ambiguous knowl-edge to the student and mislead its learning. To this end, we propose a simple but effective method, i.e., Scale De-coupled Distillation (SDD), for logit knowledge distillation. SDD decouples the global logit output into multi-ple local logit outputs and establishes distillation pipelines for them. This helps the student to mine and inherit fine-grained and unambiguous logit knowledge. Moreover, the decoupled knowledge can be further divided into consis-tent and complementary logit knowledge that transfers the semantic information and sample ambiguity, respectively. By increasing the weight of complementary parts, SDD can guide the student to focus more on ambiguous samples, im-proving its discrimination ability. Extensive experiments on several benchmark datasets demonstrate the effective-ness of SDD for wide teacher-student pairs, especially in the fine-grained classification task. Code is available at: https://github.comishicaiwei123/SDD-CVPR2024
Shicai Wei, Chunbo Luo, Yang Luo 0001
CVPR1
2024 Robust Multimodal Learning via Representation Decoupling
Shicai Wei, Yang Luo 0001, Yuji Wang, Chunbo Luo
ECCV (42)1
2024 Convolution Meets Transformer: Efficient Hybrid Transformer for Semantic Segmentation with Very High Resolution Imagery
abstract
In this paper, we introduce an efficient and lightweight hybrid Transformer architecture, ingeniously integrating convolutions within Transformer blocks for semantic segmentation of remote sensing Very High Resolution (VHR) imagery. To simultaneously avoid the high computational complexity in the shallow layers and capture the local representations of the VHR images, we propose the Group-Team Convolution Modulation (GTCM) module that uses convolutions to approximate the effect of attention mechanisms and modulates features in channel dimension. Additionally, to enlarge the effective receptive field (ERF) in the decoder, based on the grouping philosophy, we adopt dilated convolutions with multiple dilated rates to further enhance the performance. The superiority and efficiency of our proposed hybrid structure are demonstrated by outperforming state-of-the-art methods on the Vaihingen and Potsdam datasets with relatively lower complexity and fewer parameters.
Yuji Wang, Ruojun Zhao, Shicai Wei, Jingchen Ni, Yang Luo 0001, Chunbo Luo
IGARSS3
2024 Gradient Decoupled Learning With Unimodal Regularization for Multimodal Remote Sensing Classification
abstract
The joint use of multisource remote-sensing data for Earth observation has drawn much attention due to its robust performance. Although many methods have been proposed to fuse multimodal data, they tend to improve the interaction of different modality data while ignoring the optimization of each modality. Existing studies show that high-performance modalities will suppress the learning of weak ones, leading to under-optimized multimodal learning. To this end, we propose a general framework called gradient decoupled network (GDNet) to assist the multimodal remote sensing (RS) classification. GDNet guides each modality encoder in the multimodal model to learn probabilistic representations instead of deterministic ones. This helps decouple their gradient, reducing their influence on each other and encouraging them to learn the modality-specific information. Then, we further introduce the unimodal regularization for each modality encoder to align their logit output with the multimodal one and label distribution simultaneously. This helps introduce independent gradient paths for each morality encoder to accelerate their optimization when preserving the modality-share information. Finally, extensive experiments conducted on three benchmark datasets demonstrate that the proposed GDNet can effectively address the under-optimized problem in multimodal RS image classification. Code is available athttps://github.com/shicaiwei123/TGRS-GDNet.
Shicai Wei, Chunbo Luo, Xiaoguang Ma, Yang Luo 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Privileged Modality Learning via Multimodal Hallucination
abstract
Learning based on multimodal data has attracted increasing interest recently. While a variety of sensory modalities can be collected for training, not all of them are always available in practical scenarios, which raises the challenge to infer with incomplete modality. This article presents a general framework termed multimodal hallucination (MMH) to bridge the gap between ideal training scenarios and real-world deployment scenarios with incomplete modality data by transferring the complete multimodal knowledge to the hallucination network with incomplete modality input. Compared with the modality hallucination methods that restore privileged modalities information for late fusion, the proposed framework not only helps to preserve the crucial cross-modal cues but relates the study in complete modalities and in incomplete modalities. Then, we introduce two strategies called region-aware distillation and discrepancy-aware distillation to transfer the response-based and joint-representation-based knowledge of pre-trained multimodal networks, respectively. Region-aware distillation establishes and weights knowledge transferring pipelines between the response of multimodal and hallucination networks at multiple regions, which guides the hallucination network to focus on discriminative regions and avoid wasted gradients. Discrepancy-aware distillation guides the hallucination network to mimic the local inter-sample distance of multimodal representations, which enables the hallucination network to acquire the inter-class discrimination refined by multimodal cues. Extensive experiments on multimodal action recognition and face anti-spoofing demonstrate the proposed multimodal hallucination framework can overcome the problem of incomplete modality input in various scenes and achieve state-of-the-art performance.https://github.com/shicaiwei123/TMM-MMH
Shicai Wei, Chunbo Luo, Yang Luo 0001, Jialang Xu
IEEE Trans. Multim.1
2023 MMANet: Margin-Aware Distillation and Modality-Aware Regularization for Incomplete Multimodal Learning
abstract
Multimodal learning has shown great potentials in numerous scenes and attracts increasing interest recently. However, it often encounters the problem of missing modality data and thus suffers severe performance degradation in practice. To this end, we propose a general framework called MMANet to assist incomplete multimodal learning. It consists of three components: the deployment network used for inference, the teacher network transferring comprehensive multimodal information to the deployment network, and the regularization network guiding the deployment network to balance weak modality combinations. Specifically, we propose a novel margin-aware distillation (MAD) to assist the information transfer by weighing the sample contribution with the classification uncertainty. This encourages the deployment network to focus on the samples near decision boundaries and acquire the refined inter-class margin. Besides, we design a modality-aware regularization (MAR) algorithm to mine the weak modality combinations and guide the regularization network to calculate prediction loss for them. This forces the deployment network to improve its representation ability for the weak modality combinations adaptively. Finally, extensive experiments on multimodal classification and segmentation tasks demonstrate that our MMANet outperforms the state-of-the-art significantly. Code is available at: https://github.com/shicaiwei123/MMANet
Shicai Wei, Chunbo Luo, Yang Luo 0001
CVPR1
2023 Diversity-Guided Distillation With Modality-Center Regularization for Robust Multimodal Remote Sensing Image Classification
abstract
Multimodal learning has shown great potential in remote sensing image classification and attracted increasing interest in the community. Although it is preferable to collect multiple modalities for training, not all of them are available in practical scenarios. To this end, we propose a general diversity-guided distillation network (DGDNet) with modality-center regularization to facilitate accurate model inference when modalities are missing. Compared with existing modality reconstruction methods, DGDNet does not need prior knowledge of the missing modality and can handle various missing modalities via only one model. Specifically, DGDNet consists of two components: the deployment network extracting the modality-invariant representation for robust inference and the teacher network transferring comprehensive multimodal information to the deployment network. This enables the deployment network to learn the modality invariant and specific information simultaneously while maintaining robustness for incomplete modality input. In particular, we design a novel diversity-guided distillation method that transfers knowledge by matching the feature diversity. This helps overcome the representation heterogeneity when encouraging the deployment network to learn modality-specific information. Besides, a modality-center regularization strategy is proposed to address the unbalanced training of teacher and deployment networks by constraining the intra-class inter-modality variations. This helps alleviate the underfitting for weak modality, improving the model performance. Finally, extensive experiments demonstrate that the proposed DGDNet can address the problem of missing modalities effectively and achieves state-of-the-art performance. The code is available at https://github.com/shicaiwei123/TGRS-DGDNet.
Shicai Wei, Yang Luo 0001, Chunbo Luo
IEEE Trans. Geosci. Remote. Sens.1
2023 MSH-Net: Modality-Shared Hallucination With Joint Adaptation Distillation for Remote Sensing Image Classification Using Missing Modalities
abstract
Learning based multimodal data has attracted increasing interest in the remote sensing community owing to its robust performance. Although it is preferable to collect multiple modalities for training, not all of them are available in practical scenarios due to the restriction of imaging conditions. Therefore, how to assist the model inference with missing modalities is significant for multimodal remote sensing image processing. In this work, we propose a general framework called modality-shared hallucination network (MSH-Net) to address this issue by reconstructing complete modality-shared features from the incomplete inference modalities. Compared to conventional privilege modality hallucination methods, MSH-Net does not only help preserve the cross-modal interactions for model inference, but also scales well with the increasing number of missing modalities. We further develop a novel joint adaptation distillation (JAD) method that guides the hallucination model to learn the modality-shared knowledge from the multimodal model by matching the joint probability distributions between representation and groundtruth. This overcomes the representation heterogeneity caused by the discrepancy between inputs and structures of multimodal and hallucination model, while preserving the decision boundaries refined by multimodal cues. Finally, extensive experiments conducted on four common modality combinations demonstrate that the proposed MSH-Net can effectively address the problem of missing modalities and achieve state-of-the-art performance. Code is available at: https://github.com/shicaiwei123/MSHNet.
Shicai Wei, Yang Luo 0001, Xiaoguang Ma, Peng Ren 0001, Chunbo Luo
IEEE Trans. Geosci. Remote. Sens.1