Litao Qu

dblp:343/6800 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0009-0179-5515ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 70% Representation and self-supervised learning · 22% Vision and language · 8%
Computer graphics and multimedia
2 papers
Image and video processing · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.622025
Dual-Granularity Semantic Guided Sparse Routing Diffusion Model for General Pansharpening · CVPR 2025
CrossDiff: Exploring Self-SupervisedRepresentation of Pansharpening via Cross-Predictive Diffusion Model · IEEE Trans. Image Process. 2024
Image and video processing › image fusion › remote sensing image fusion
pansharpening
1.622025
Dual-Granularity Semantic Guided Sparse Routing Diffusion Model for General Pansharpening · CVPR 2025
CrossDiff: Exploring Self-SupervisedRepresentation of Pansharpening via Cross-Predictive Diffusion Model · IEEE Trans. Image Process. 2024
Machine learning › Generative modeling › diffusion model › score-based generative model
denoising diffusion probabilistic model
0.812024
CrossDiff: Exploring Self-SupervisedRepresentation of Pansharpening via Cross-Predictive Diffusion Model · IEEE Trans. Image Process. 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
pretext task
0.812024
CrossDiff: Exploring Self-SupervisedRepresentation of Pansharpening via Cross-Predictive Diffusion Model · IEEE Trans. Image Process. 2024
Image and video processing
image fusion
0.812024
CrossDiff: Exploring Self-SupervisedRepresentation of Pansharpening via Cross-Predictive Diffusion Model · IEEE Trans. Image Process. 2024
Computer vision › Vision and language
vision-language model
0.312025
Dual-Granularity Semantic Guided Sparse Routing Diffusion Model for General Pansharpening · CVPR 2025
Image and video processing › image fusion
remote sensing image fusion
0.212024
CrossDiff: Exploring Self-SupervisedRepresentation of Pansharpening via Cross-Predictive Diffusion Model · IEEE Trans. Image Process. 2024

Methods — techniques the papers use, named apart from their topics

sparse routing · 1.7mixture of experts · 1.7geochat · 1.7cross-predictive diffusion · 1.5self-supervised pretraining · 0.8self-supervised pre-training · 0.8
YearPublicationVenuePosition
2025 Dual-Granularity Semantic Guided Sparse Routing Diffusion Model for General Pansharpening
abstract
Pansharpening aims at integrating complementary information from panchromatic and multispectral images. Available deep-learning based pansharpening methods typically perform exceptionally with particular satellite datasets. At the same time, it has been observed that these models also exhibit scene dependence, for example, if the majority of the training samples come from the urban scenes, the model’s performance may decline in the river scene. To address the domain gap produced by varying satellite sensors and distinct scenes, we propose a dual-granularity semantic guided sparse routing diffusion model for general pansharpening. By utilizing the large Vision-Language Models (VLMs) in the field of geoscience, e.g, GeoChat, we introduce the dual granularity semantics to generate dynamic sparse routing scores for adaptation of different satellite sensors and scenes. This scene-level and region-level dual-granularity semantic information serves as guidance for dynamically activating specialized experts within the diffusion model. Extensive experiments on WorldView-3, QuickBird, and GaoFen-2 datasets show the effectiveness of our proposed method. Notably, the proposed method outperforms the comparison approaches in adapting to new satellite sensors and scenes. The codes are available at https://github.com/codgodtao/SGDiff.
Yinghui Xing, Litao Qu, Shizhou Zhang, Di Xu 0010, Yingkun Yang, Yanning Zhang 0001
CVPR2
2025 Temperature-Aware Adaptive Federated Distillation for Energy-Constrained AIoT with Non-IID Data
abstract
Federated Learning (FL) can help multiple Internet of Things (IoT) devices to collaboratively train a machine learning model to provide intelligent services and applications (FL-AIoT). Due to IoT devices' limited storage capacity and energy, fresh data collected by devices often overwrites outdated data and establishes a heterogeneous data distribution. This causes the global model to forget outdated data's characteristics (i.e., catastrophic forgetting). Existing methods incorporate knowledge distillation into FL (i.e., federated distillation, FD) to extract and integrate characteristics from both fresh and outdated data, but they use fixed distillation temperatures for different devices, which overlooks that fixed distillation temperatures cannot match the heterogeneous data distribution on different devices and degrades global model accuracy. To this end, we propose a Federated Dynamic Decoupled Distillation method based on Logits distribution (Fed3DL). Specifically, Fed3DL utilizes decoupled distillation to mitigate catastrophic forgetting. To alleviate the impact of heterogeneous data distributions, Fed3DL novelly builds an adaptive temperature-aware mechanism to dynamically adjust the distillation temperature of each device based on the distribution of Logits. Additionally, Fed3DL introduces a regularization term into the local distillation loss to reduce inter-class characteristics disparity and improve model accuracy. Experiments on two datasets show that compared with the best of the 5 baselines, Fed3DL can improve the global model accuracy by an average of 3.40 %, reduce the forgetting rate by an average of 4.98 %, and achieve the lowest inter-class accuracy disparity.
Yingchi Mao, Jiakai Zhang, Litao Qu, Benteng Zhang, Xiaoming He 0004
VTC2025-Spring3
2024 Complementary Fusion Network Based on Frequency Hybrid Attention for Pansharpening
abstract
Pansharpening is a feasible way to obtain the high-resolution (HR) multispectral (MS) images by using panchromatic (PAN) images to sharpen low-resolution MS images. Despite its great advances, most existing pansharpening methods neglect the importance of integrating local and non-local characteristics of images, resulting in the imbalance of spatial and spectral distribution. In this paper, we propose a complementary fusion network (CFNet) based on frequency hybrid attention mechanism for pansharpening. By introducing the frequency transformation and the deformable cross-attention, our model takes image-wide receptive field into consideration to explore global feature learning. Combined with the convolutional layers with local receptive field, CFNet can well capture local and non-local features. Experimental results demonstrate that the proposed method outperforms the comparison methods in terms of visual and quantitative qualities.
Yinghui Xing, Litao Qu, Kai Zhang 0010, Yan Zhang 0127, Xiuwei Zhang 0001, Yanning Zhang 0001
ICASSP2
2024 Empower Generalizability for Pansharpening Through Text-Modulated Diffusion Model
abstract
Pansharpening is crucial to remote sensing applications by fusing high-resolution (HR) panchromatic (PAN) images with low-resolution multispectral (LRMS) images to generate HR multispectral (HRMS) images. Recently, diffusion probabilistic models (DPMs) have provided high-quality results than regression-based methods when trained on specific pairwise data for their specific purpose. However, their performance degrades when applied to a new satellite dataset, which represents different imaging properties and spectral ranges, limiting the generalization ability of them. For better generalizability of pansharpening, in this article, we propose a text-modulated diffusion model (TMDiff) for unified pansharpening of different satellites. TMDiff takes a text-modulated 3-D UNet (TM3DU) as denoising network to gradually recover HRMS through iterative refinement over multiple time steps. By introducing satellite’s physical properties as text prompts, TM3DU is able to learn meta-knowledge across different satellites and thus can sharpen LRMS images with diverse spatial and spectral attributes. Extensive experiments on various satellite datasets demonstrate the state-of-the-art performance of our model in both qualitative and quantitative metrics. Furthermore, our model exhibits superior generalization ability to unseen datasets, highlighting its practical significance. Code is available athttps://github.com/codgodtao/TMDiff.
Yinghui Xing, Litao Qu, Shizhou Zhang, Jiapeng Feng, Xiuwei Zhang 0001, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 CrossDiff: Exploring Self-SupervisedRepresentation of Pansharpening via Cross-Predictive Diffusion Model
abstract
Fusion of a panchromatic (PAN) image and corresponding multispectral (MS) image is also known as pansharpening, which aims to combine abundant spatial details of PAN and spectral information of MS images. Due to the absence of high-resolution MS images, available deep-learning-based methods usually follow the paradigm of training at reduced resolution and testing at both reduced and full resolution. When taking original MS and PAN images as inputs, they always obtain sub-optimal results due to the scale variation. In this paper, we propose to explore the self-supervised representation for pansharpening by designing a cross-predictive diffusion model, named CrossDiff. It has two-stage training. In the first stage, we introduce a cross-predictive pretext task to pre-train the UNet structure based on conditional Denoising Diffusion Probabilistic Model (DDPM). While in the second stage, the encoders of the UNets are frozen to directly extract spatial and spectral features from PAN and MS images, and only the fusion head is trained to adapt for pansharpening task. Extensive experiments show the effectiveness and superiority of the proposed model compared with state-of-the-art supervised and unsupervised methods. Besides, the cross-sensor experiments also verify the generalization ability of proposed self-supervised representation learners for other satellite datasets. Code is available at https://github.com/codgodtao/CrossDiff.
Yinghui Xing, Litao Qu, Shizhou Zhang, Kai Zhang 0010, Yanning Zhang 0001, Lorenzo Bruzzone
IEEE Trans. Image Process.2