Zihan Cao

dblp:235/8988 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-7532-2122ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 A bio-inspired photopic-scotopic duplex network for low-light image enhancement
Zihan Cao, Manli Wang, Changsen Zhang, Yannan Shi, Dekui Li, Yingxu Qiao
Neurocomputing1
2026 Zero-shot unsupervised learning with unfolded equilibrium network for remote sensing pansharpening
Jieyi Zhu, Zihan Cao, Liang-Jian Deng
Pattern Recognit.2
2025 Hyperspectral Pansharpening via Diffusion Models with Iteratively Zero-Shot Guidance
abstract
Hyperspectral pansharpening refers to fusing a panchromatic image (PAN) and a low-resolution hyperspectral image (LR-HSI) to obtain a high-resolution hyperspectral image (HR-HSI). Recently, guiding pre-trained diffusion models (DMs) has demonstrated significant potential in this area, leveraging their powerful representational abilities while avoiding complex training processes. However, these DMs are often trained on RGB images, not well-suited for pansharpening tasks, limited in adapting to the hyperspectral images. In this work, we propose a novel guided diffusion scheme with zero-shot guidance and neural spatialspectral decomposition (NSSD) to iteratively generate the RGB detail image and map the RGB detail image to target HR-HSI. Specifically, zero-shot guidance employs an auxiliary neural network that trained only with a PAN and LR-HSI to guide pre-trained DMs in generating the RGB detail image, informed by specific prior knowledge. Then, NSSD establishes a spectral mapping from the generated RGB detail image to the final HR-HSI. Extensive experiments are conducted on Pavia, Washington DC, Chukusei, and FR1 datasets to demonstrate that the proposed method significantly enhances the performance of DMs for hyperspectral pansharpening tasks, outperforming existing methods across multiple metrics and achieving improvements in visualization results. The code is available at https://github.com/Jin-liangXiao/DM-zs.
Jin-Liang Xiao, Ting-Zhu Huang, Liang-Jian Deng, Guang Lin 0002, Zihan Cao, Chao Li 0013, Qibin Zhao
CVPR5
2025 Taming Flow Matching With Unbalanced Optimal Transport Into Fast Pansharpening
abstract
Pansharpening, a pivotal task in remote sensing for fusing high-resolution panchromatic and multispectral imagery, has garnered significant research interest. Recent advancements employing diffusion models based on stochastic differential equations (SDEs) have demonstrated state-of-the-art performance. However, the inherent multi-step sampling process of SDEs imposes substantial computational overhead, hindering practical deployment. While existing methods adopt efficient samplers, knowledge distillation, or retraining to reduce sampling steps (e.g., from 1,000 to fewer steps), such approaches often compromise fusion quality. In this work, we propose the Optimal Transport Flow Matching (OTFM) framework, which integrates the dual formulation of unbalanced optimal transport (UOT) to achieve one-step, high-quality pansharpening. Unlike conventional OT formulations that enforce rigid distribution alignment, UOT relaxes marginal constraints to enhance modeling flexibility, accommodating the intrinsic spectral and spatial disparities in remote sensing data. Furthermore, we incorporate task-specific regularization into the UOT objective, enhancing the robustness of the flow model. The OTFM framework enables simulation-free training and single-step inference while maintaining strict adherence to pansharpening constraints. Experimental evaluations across multiple datasets demonstrate that OTFM matches or exceeds the performance of previous regression-based models and leading diffusion-based methods while only needing one sampling step. Codes are available at https://github.com/294coder/PAN-OTFM.
Zihan Cao, Liang-Jian Deng
ICCV1
2025 MMAIF: Multi-Task and Multi-Degradation All-in-One for Image Fusion with Language Guidance
abstract
Image fusion, a fundamental low-level vision task, aims to integrate multiple image sequences into a single output while preserving as much information as possible from the input. However, existing methods face several significant limitations: 1) requiring task- or dataset-specific models; 2) neglecting real-world image degradations (\textit{e.g.}, noise), which causes failure when processing degraded inputs; 3) operating in pixel space, where attention mechanisms are computationally expensive; and 4) lacking user interaction capabilities. To address these challenges, we propose a unified framework for multi-task, multi-degradation, and language-guided image fusion. Our framework includes two key components: 1) a practical degradation pipeline that simulates real-world image degradations and generates interactive prompts to guide the model; 2) an all-in-one Diffusion Transformer (DiT) operating in latent space, which fuses a clean image conditioned on both the degraded inputs and the generated prompts. Furthermore, we introduce principled modifications to the original DiT architecture to better suit the fusion task. Based on this framework, we develop two versions of the model: Regression-based and Flow Matching-based variants. Extensive qualitative and quantitative experiments demonstrate that our approach effectively addresses the aforementioned limitations and outperforms previous restoration+fusion and all-in-one pipelines. Codes are available at https://github.com/294coder/MMAIF.
Zihan Cao, Liang-Jian Deng
ICCV1
2025 An Efficient Image Fusion Network Exploiting Unifying Language and Mask Guidance
abstract
Image fusion aims to merge image pairs collected by different sensors over the same scene, preserving their distinct features. Recent works have often focused on designing various image fusion losses, developing different network architectures, and leveraging downstream tasks (e.g., object detection) for image fusion. However, a few studies have explored how language and semantic masks can serve as guidance to aid image fusion. In this paper, we investigate how the combination of language and masks can guide image fusion tasks, discarding the previously complex frameworks, which rely on downstream tasks, GAN-based cycle training, diffusion models, or deep image priors. Additionally, we exploit a recurrent neural network-like architecture to build a lightweight network that avoids the quadratic-cost of traditional attention mechanisms. To adapt the receptance weighted key value (RWKV) model to an image modality, we modify it into a bidirectional version using an efficient scanning strategy (ESS). To guide image fusion by language and mask features, we introduce a multi-modal fusion module (MFM) to facilitate information exchange. Comprehensive experiments show that the proposed framework achieved state-of-the-art results in various image fusion tasks (i.e., visible-infrared image fusion, multi-focus image fusion, multi-exposure image fusion, medical image fusion, hyperspectral and multispectral image fusion, and pansharpening).
Zihan Cao, Yu-Jie Liang, Liang-Jian Deng, Gemine Vivone
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Fully-Connected Transformer for Multi-Source Image Fusion
abstract
Multi-source image fusion combines the information coming from multiple images into one data, thus improving imaging quality. This topic has aroused great interest in the community. How to integrate information from different sources is still a big challenge, although the existing self-attention based transformer methods can capture spatial and channel similarities. In this paper, we first discuss the mathematical concepts behind the proposed generalized self-attention mechanism, where the existing self-attentions are considered basic forms. The proposed mechanism employs multilinear algebra to drive the development of a novel fully-connected self-attention (FCSA) method to fully exploit local and non-local domain-specific correlations among multi-source images. Moreover, we propose a multi-source image representation embedding it into the FCSA framework as a non-local prior within an optimization problem. Some different fusion problems are unfolded into the proposed fully-connected transformer fusion network (FC-Former). More specifically, the concept of generalized self-attention can promote the potential development of self-attention. Hence, the FC-Former can be viewed as a network model unifying different fusion tasks. Compared with state-of-the-art methods, the proposed FC-Former method exhibits robust and superior performance, showing its capability of faithfully preserving information.
Zihan Cao, Ting-Zhu Huang, Liang-Jian Deng, Jocelyn Chanussot, Gemine Vivone
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 A Novel State Space Model with Local Enhancement and State Sharing for Image Fusion
abstract
In image fusion tasks, images from different sources possess distinct characteristics.This has driven the development of numerous methods to explore better ways of fusing them while preserving their respective characteristics.Mamba, as a state space model, has emerged in the field of natural language processing.Recently, many studies have attempted to extend Mamba to vision tasks.However, due to the nature of images different from causal language sequences, the limited state capacity of Mamba weakens its ability to model image information.Additionally, the sequence modeling ability of Mamba is only capable of spatial information and cannot effectively capture the rich spectral information in images.Motivated by these challenges, we customize and improve the vision Mamba network designed for the image fusion task.Specifically, we propose the local-enhanced vision Mamba block, dubbed as LEVM.The LEVM block can improve local information perception of the network and simultaneously learn local and global spatial information.Furthermore, we propose the state sharing technique to enhance spatial details and integrate spatial and spectral information.Finally, the overall network is a multi-scale structure based on vision Mamba, called LE-Mamba.Extensive experiments show the proposed methods achieve state-of-the-art results on multispectral pansharpening and multispectral and hyperspectral image fusion datasets, and demonstrate the effectiveness of the proposed approach.Code can be accessed at https://github.com/294coder/Efficient-MIF.
Zihan Cao, Liang-Jian Deng
ACM Multimedia1
2024 Linearly-evolved Transformer for Pan-sharpening
Junming Hou, Zihan Cao, Naishan Zheng, Xuan Li 0012, Xiaofeng Cong, Danfeng Hong, Man Zhou 0003
ACM Multimedia2
2024 Fourier-enhanced Implicit Neural Fusion Network for Multispectral and Hyperspectral Image Fusion
abstract
Recently, implicit neural representations (INR) have made significant strides in various vision-related domains, providing a novel solution for Multispectral and Hyperspectral Image Fusion (MHIF) tasks. However, INR is prone to losing high-frequency information and is confined to the lack of global perceptual capabilities. To address these issues, this paper introduces a Fourier-enhanced Implicit Neural Fusion Network (FeINFN) specifically designed for MHIF task, targeting the following phenomena: The Fourier amplitudes of the HR-HSI latent code and LR-HSI are remarkably similar; however, their phases exhibit different patterns. In FeINFN, we innovatively propose a spatial and frequency implicit fusion function (Spa-Fre IFF), helping INR capture high-frequency information and expanding the receptive field. Besides, a new decoder employing a complex Gabor wavelet activation function, called Spatial-Frequency Interactive Decoder (SFID), is invented to enhance the interaction of INR features. Especially, we further theoretically prove that the Gabor wavelet activation possesses a time-frequency tightness property that favors learning the optimal bandwidths in the decoder. Experiments on two benchmark MHIF datasets verify the state-of-the-art (SOTA) performance of the proposed method, both visually and quantitatively. Also, ablation studies demonstrate the mentioned contributions. The code can be available at https://github.com/294coder/Efficient-MIF.
Yu-Jie Liang, Zihan Cao, Shangqi Deng, Hong-Xia Dou, Liang-Jian Deng
NeurIPS2
2024 SSDiff: Spatial-spectral Integrated Diffusion Model for Remote Sensing Pansharpening
abstract
Pansharpening is a significant image fusion technique that merges the spatial content and spectral characteristics of remote sensing images to generate high-resolution multispectral images. Recently, denoising diffusion probabilistic models have been gradually applied to visual tasks, enhancing controllable image generation through low-rank adaptation (LoRA). In this paper, we introduce a spatial-spectral integrated diffusion model for the remote sensing pansharpening task, called SSDiff, which considers the pansharpening process as the fusion process of spatial and spectral components from the perspective of subspace decomposition. Specifically, SSDiff utilizes spatial and spectral branches to learn spatial details and spectral features separately, then employs a designed alternating projection fusion module (APFM) to accomplish the fusion. Furthermore, we propose a frequency modulation inter-branch module (FMIM) to modulate the frequency distribution between branches. The two components of SSDiff can perform favorably against the APFM when utilizing a LoRA-like branch-wise alternative fine-tuning method. It refines SSDiff to capture component-discriminating features more sufficiently. Finally, extensive experiments on four commonly used datasets, i.e., WorldView-3, WorldView-2, GaoFen-2, and QuickBird, demonstrate the superiority of SSDiff both visually and quantitatively. The code is available at https://github.com/Z-ypnos/SSdiff_main.
Liang-Jian Deng, Zihan Cao, Hong-Xia Dou
NeurIPS4
2024 RF-Based Drone Detection Enhancement via a Generalized Denoising and Interference-Removal Framework
abstract
Radio frequency-based (RF-based) detection methods are currently the main means of countering drones. However, these prevalent approaches frequently exhibit deficiencies in effectively addressing noise and interference, making them potentially unsuitable for application in realistic urban environments. This paper proposes a generalized RF signal-enhanced framework that explicitly addresses noise and interference. We decompose the RF signal into three components and uniformly integrate them into the proposed framework for decomposition. To accomplish this, three innovative loss functions and two appropriate neural networks are devised. To validate our framework, we create a real-world drone RF dataset sampled from urban surroundings, faithfully representing drone RF signals in real-world scenarios. Experimental results demonstrate that our framework exhibits satisfactory denoising and interferenceremoval performance, significantly improving the accuracy of multiple detection methods.
Zihan Cao, Julan Xie, Wei Zhang 0099, Zishu He
IEEE Signal Process. Lett.2
2023 DBL-MPE: Deep Broad Learning for Prediction of Response to Neo-adjuvant Chemotherapy Using MRI-Based Multi-angle Maximal Enhancement Projection in Breast Cancer
Zihan Cao, Zhenwei Shi 0002, Xiaomei Huang, Chu Han, Peng Xu 0004, Zaiyi Liu
ICIC (3)1
2023 Fed-CSA: Channel Spatial Attention and Adaptive Weights Aggregation-Based Federated Learning for Breast Tumor Segmentation on MRI
Zhenwei Shi 0002, Xiaomei Huang, Chu Han, Zihan Cao, Peng Xu 0004, Zaiyi Liu
ICIC (3)5
2022 High Accuracy Pressure Sensing With Sagnac Interferometry Based On Deep Learning Approach
abstract
In this paper, we proposed a pressure sensor using a Sagnac interferometer based on a side-hole fiber (SHF) with the assistance of deep learning. A convolutional neural network (CNN) was built to identify the spectra of different pressures since the traditional tracing method will face spectral overlap problems when the shift of the spectrum exceeds the free spectral range (FSR). The spectra of pressures ranging from 0 Mpa to 5 MPa with a step of 0.1 MPa will be normalized firstly and then sent to the CNN model for training. The precited result shows that the coefficient of determination$R^{2}$is 99.99987% with the root mean square error (RMSE) equal to$1.6537\times 10^{-3}$MPa. Additionally, a similar structure was constructed to demonstrate the universality of the proposed CNN model. The model can also get a good performance, although the receiving device has a low resolution, showing its great potential for developing a low-cost sensing system.
Yongchang Mei, Shengqi Zhang, Zihan Cao, Titi Xia, Xingwen Yi, Zhengyong Liu
MMSP3
2021 Analysis on Teeth Occlusion Distribution Based on Segmentation and Registration Algorithm
abstract
Occlusal contact status of teeth is a key indicator for orthodontic and periodontal disease treatment. Digital analysis includes tooth position recognition and occlusal contact distribution estimation. In this study, we propose a cascade two-stage point-wise network named Teeth Segmentation Network (TSegNet) based on self-attention mechanism to address teeth segmentation task. And a template-based registration method is proposed to analyze the status of teeth occlusion. In TSegNet, spatial and channel attention are used to improve the performance feature extraction. Template-registration-based occlusal distribution analysis method reduced the labeled number of training samples. To the best of our knowledge, it is the first study on occlusal contact analyzing by using computer-aided-diagnosis technique. Experiment results illustrate the effectiveness and robustness of our proposed method.
Zihan Cao, Xinwu Sun, Gangyuan Chen, Yan Liu 0054, Xinggang Liu, Dongxiang Zheng, Ling Wang 0013
BIBM1