VLDB 2026 Research / reviewers in the wild / expert
Qinglong Cao
dblp:313/4952
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0003-2742-026XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Latent Knowledge-Guided Video Diffusion for Scientific Phenomena Generation from a Single Initial FrameabstractVideo diffusion models have achieved impressive results in natural scene generation, yet they struggle to generalize to scientific phenomena such as fluid simulations and meteorological processes, where underlying dynamics are governed by scientific laws. These tasks pose unique challenges, including severe domain gaps, limited training data, and the lack of descriptive language annotations. To handle this dilemma, we extracted the latent scientific phenomena knowledge and further proposed a fresh framework that teaches video diffusion models to generate scientific phenomena from a single initial frame. Particularly, static knowledge is extracted via pre-trained masked autoencoders, while dynamic knowledge is derived from pre-trained optical flow prediction. Subsequently, based on the aligned spatial relations between the CLIP vision and language encoders, the visual embeddings of scientific phenomena, guided by latent scientific phenomena knowledge, are projected to generate the pseudo-language prompt embeddings in both spatial and frequency domains. By incorporating these prompts and fine-tuning the video diffusion model, we enable the generation of videos that better adhere to scientific laws. Extensive experiments on both computational fluid dynamics simulations and real-world typhoon observations demonstrate the effectiveness of our approach, achieving superior fidelity and consistency across diverse scientific scenarios. Qinglong Cao, Chao Ma 0004, Yuntian Chen, Xiaokang Yang 0001 |
AAAI | 1 |
| 2026 | Dual-perspective filter pruning via diversity and independence collaboration
Chenyang Gao, Qinglong Cao, Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003 |
Pattern Recognit. | 2 |
| 2025 | Auto-Regressive Moving Diffusion Models for Time Series ForecastingabstractTime series forecasting (TSF) is essential in various domains, and recent advancements in diffusion-based TSF models have shown considerable promise. However, these models typically adopt traditional diffusion patterns, treating TSF as a noise-based conditional generation task. This approach neglects the inherent continuous sequential nature of time series, leading to a fundamental misalignment between diffusion mechanisms and the TSF objective, thereby severely impairing performance. To bridge this misalignment, and inspired by the classic Auto-Regressive Moving Average (ARMA) theory, which views time series as continuous sequential progressions evolving from previous data points, we propose a novel Auto-Regressive Moving Diffusion (ARMD) model to first achieve the continuous sequential diffusion-based TSF. Unlike previous methods that start from white Gaussian noise, our model employs chain-based diffusion with priors, accurately modeling the evolution of time series and leveraging intermediate state information to improve forecasting accuracy and stability. Specifically, our approach reinterprets the diffusion process by considering future series as the initial state and historical series as the final state, with intermediate series generated using a sliding-based technique during the forward process. This design aligns the diffusion model's sampling procedure with the forecasting objective, resulting in an unconditional, continuous sequential diffusion TSF model. Extensive experiments conducted on seven widely used datasets demonstrate that our model achieves state-of-the-art performance, significantly outperforming existing diffusion-based TSF models. Qinglong Cao, Yuntian Chen |
AAAI | 2 |
| 2025 | Domain Prompt Learning with Quaternion Networks (Extended Abstract)abstractFoundational vision-language models (VLMs) like CLIP have revolutionized image recognition, but adapting them to specialized domains with limited data remains challenging. We propose Domain Prompt Learning with Quaternion Networks (DPLQ), which leverages domain-specific foundation models and quaternion-based prompt tuning to effectively transfer recognition capabilities. Our method achieves state-of-the-art results in remote sensing and medical imaging tasks. This extended abstract highlights the key contributions and performance of DPLQ. Qinglong Cao, Zhengqin Xu, Yuntian Chen, Chao Ma 0004, Xiaokang Yang 0001 |
IJCAI | 1 |
| 2025 | Open-Vocabulary High-Resolution Remote Sensing Image Semantic SegmentationabstractOpen-vocabulary image semantic segmentation (OVS) seeks to segment images into semantic regions across an open set of categories. Existing OVS methods commonly depend on foundational vision-language models and utilize similarity computation to tackle OVS tasks. However, these approaches are predominantly tailored to natural images and struggle with the unique characteristics of high-resolution remote sensing images, such as rapidly changing orientations and significant scale variations. To tackle this dilemma, we propose the first OVS framework specifically designed for high-resolution remote sensing imagery, introducing a rotation-aggregative similarity computation module to enhance segmentation across varying orientations and a multi-scale feature integration strategy to generate scale-aware semantic masks. Additionally, we establish the first open-sourced OVS benchmark for remote sensing, comprising four public datasets. Experiments demonstrate that our framework effectively addresses orientation and scale challenges, achieving state-of-the-art performance. All codes and datasets are available at https://github.com/caoql98/OVRS. Qinglong Cao, Yuntian Chen, Chao Ma 0004, Xiaokang Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | HSFormer: Multiscale Hybrid Sparse Transformer for Uncertainty-Aware Cloud and Shadow RemovalabstractClouds and their shadows hinder accurate analysis of optical remote sensing imagery, making cloud removal an indispensable preprocessing step in remote sensing. However, existing methods lack efficient long-range modeling capabilities and overlook the impact of cloud irregularities and uncertainties on cloud removal. Additionally, resolving spectral confusion between cloud shadows and surface information remains a significant challenge. To tackle this issue, this study introduces an innovative cloud removal algorithm termed the multiscale hybrid sparse transformer (HSFormer), which adaptively removes clouds and shadows while reconstructing land surface semantics. HSFormer leverages pixel correlation explicit sparsity and uncertainty-driven implicit sparsity to maximize attention gains, enabling efficient cloud recognition and removal. The global pixel correlation based on attention relations enhances the semantic integrity of reconstructed images and avoids information loss across frequency domains. Furthermore, the uncertainty-guided adaptive receptive field enhances the model’s ability to resolve complex cloud-covered spatial relationships and reduces the spatial uncertainty of the reconstructed image. Experiments on simulated cloud shadow, real RICE, and full-band WHUS2-CRv datasets demonstrate HSFormer’s superiority over existing methods, improving PSNR and SSIM by 0.49% and 0.62%, respectively, in average evaluations across all bands of the WHUS2-CRv dataset and effectively addresses spectral aliasing between cloud shadows and dark surfaces. Changqi Sun, Yuntian Chen, Qinglong Cao, Longfeng Nie, Zhenzhong Zeng, Dongxiao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Domain-Controlled Prompt LearningabstractLarge pre-trained vision-language models, such as CLIP, have shown remarkable generalization capabilities across various tasks when appropriate text prompts are provided. However, adapting these models to specific domains, like remote sensing images (RSIs), medical images, etc, remains unexplored and challenging. Existing prompt learning methods often lack domain-awareness or domain-transfer mechanisms, leading to suboptimal performance due to the misinterpretation of specific images in natural image patterns. To tackle this dilemma, we proposed a Domain-Controlled Prompt Learning for the specific domains. Specifically, the large-scale specific domain foundation model (LSDM) is first introduced to provide essential specific domain knowledge. Using lightweight neural networks, we transfer this knowledge into domain biases, which control both the visual and language branches to obtain domain-adaptive prompts in a directly incorporating manner. Simultaneously, to overcome the existing overfitting challenge, we propose a novel noisy-adding strategy, without extra trainable parameters, to help the model escape the suboptimal solution in a global domain oscillation manner. Experimental results show our method achieves state-of-the-art performance in specific domain image recognition datasets. Our code is available at https://github.com/caoql98/DCPL. Qinglong Cao, Zhengqin Xu, Yuntian Chen, Chao Ma 0004, Xiaokang Yang 0001 |
AAAI | 1 |
| 2024 | Domain Prompt Learning with Quaternion NetworksabstractPrompt learning has emerged as a potent and resource-efficient technique in large Vision-Language Models (VLMs). However, its application in adapting VLMs to specialized domains like remote sensing and medical imaging, termed domain prompt learning, remains relatively unexplored. Although large-scale domain-specific foundation models offer a potential solution, their focus on a singular vision level presents challenges in prompting both vision and language modalities. To address this limitation, we propose leveraging domain-specific knowledge from these foundation models to transfer the robust recognition abilities of VLMs from generalized to specialized domains, employing quaternion networks. Our method entails utilizing domain-specific vision features from domain-specific foundation models to guide the transformation of generalized contextual embeddings from the language branch into a specialized space within quaternion networks. Furthermore, we introduce a hierarchical approach that derives vision prompt features by analyzing intermodal relationships between hierarchical language prompt features and domain-specific vision features. Through this mechanism, quaternion networks can effectively explore intermodal relationships in specific domains, facilitating domain-specific vision-language contrastive learning. Extensive experiments conducted on domain-specific datasets demonstrate that our proposed method achieves new state-of-the-art results in prompt learning. Codes are available at https://github.com/caoq198/DPLQ. Qinglong Cao, Zhengqin Xu, Yuntian Chen, M. Chao, Xiaokang Yang 0001 |
CVPR | 1 |
| 2024 | Break the Bias: Delving Semantic Transform Invariance for Few-Shot SegmentationabstractFew-shot semantic segmentation (FSS) aims to segment objects of unseen classes in query images with only a few annotated support images. Existing FSS algorithms typically focus on mining category representations from the single-view support to match semantic objects of the single-view query. However, the limited annotated samples render the single-view matching struggle to perceive the varying characteristics of novel objects, which results in a restricted learning space for novel categories and further induces a biased segmentation with demoted parsing performance. To address this challenge, inspired by the semantic transform invariance, this paper proposes a fresh few-shot segmentation framework to break the bias and perform invariant segmentation in a multi-view matching manner. Specifically, original and transform support features from different perspectives with the same semantics are learnable fused to obtain the transform invariance prototype with a stronger category representation ability. Simultaneously, aiming at providing better parsing guidance, the Transform Invariance Guidance Mask Generation (TIGM) module is proposed to integrate prior knowledge from different perspectives. Finally, segmentation predictions from varying views are complementarily merged in the Transform Invariance Semantic Prediction (TISP) module to decide the uncertain area and yield precise segmentation predictions. Extensive experiments on both PASCAL-5i and COCO-20i datasets demonstrate the effectiveness of our approach and show that our method could achieve state-of-the-art performance. Code is available at https://github.com/caoql98/BBD. Qinglong Cao, Yuntian Chen, Chao Ma 0004, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Few-Shot Rotation-Invariant Aerial Image Semantic SegmentationabstractFew-shot aerial image semantic segmentation is a challenging task that requires precisely parsing unseen-category objects in query aerial images with limited annotated support aerial images. Formally, category prototypes would be extracted from support samples to segment query images in a pixel-to–pixel matching manner. However, aerial objects in aerial images are often distributed with arbitrary orientations, and varying orientations could cause a dramatic feature change. This unique property of aerial images renders conventional matching manner without consideration of orientations fails to activate same-category objects with different orientations. Furthermore, the oscillation of the confidence scores in existing rotation-insensitive algorithms, engendered by the striking changes of object orientations, often leads to false recognition of lower-scored rotated semantic objects. To tackle these challenges, inspired by the intrinsic rotation invariance in aerial images, we propose a novel few-shot rotation-invariant aerial semantic segmentation network (FRINet) to efficiently segment aerial semantic objects with diverse orientations. Specifically, through extracting orientation-varying yet category-consistent support information, FRINet provides rotation-adaptive matching for each query feature in a feature-aggregation manner. Meanwhile, to encourage consistent predictions for aerial objects with arbitrary orientations, segmentation predictions from different orientations are supervised by the same label and further fused to obtain the final rotation-invariant prediction in a complementary manner. Moreover, aiming at providing a better solution searching space, the backbones are newly pre-trained in the base category to basically boost the segmentation performance. Extensive experiments on the few-shot aerial image semantic segmentation benchmark demonstrate that the proposed FRINet achieves a new state-of-the-art performance. The code is available at https://github.com/caoql98/FRINet. Qinglong Cao, Yuntian Chen, Chao Ma 0004, Xiaokang Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Learning to Assess Image Quality Like an ObserverabstractHuman observers are the ultimate receivers and evaluators of the image visual information and have powerful perception ability of visual quality with short-term global perception and long-term regional observation. Thus, it is natural to design an image quality assessment (IQA) computational model to act like an observer for accurately predicting the human perception of image quality. Inspired by this, here, we propose a novel observer-like network (OLN) to perform IQA by jointly considering the global glimpsing information and local scanning information. Specifically, the OLN consists of a global distortion perception (GDP) module and a local distortion observation (LDO) module. The GDP module is designed to mimic the observer's global perception of image quality through performing classification of images' distortion categories and levels. Simultaneously, to simulate the human local observation behavior, the LDO module attempts to gather the long-term regional observation information of the distorted images by continuously tracing the human scanpath in the observer-like scanning manner. By leveraging the bilinear pooling layer to collaborate the short-term global perception with the long-term regional observation, our network precisely predicts the quality scores of distorted images, such as human observers. Comprehensive experiments on the public datasets powerfully demonstrate that the proposed OLN achieves state-of-the-art performance. Xiwen Yao, Qinglong Cao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Learning Non-target Knowledge for Few-shot Semantic SegmentationabstractExisting studies in few-shot semantic segmentation only focus on mining the target object information, however, often are hard to tell ambiguous regions, especially in non-target regions, which include background (BG) and Distracting Objects (DOs). To alleviate this problem, we propose a novel framework, namely Non-Target Region Eliminating (NTRE) network, to explicitly mine and eliminate BG and DO regions in the query. First, a BG Mining Module (BGMM) is proposed to extract the BG region via learning a general BG prototype. To this end, we design a BG loss to supervise the learning of BGMM only using the known target object segmentation ground truth. Then, a BG Eliminating Module and a DO Eliminating Module are proposed to successively filter out the BG and DO information from the query feature, based on which we can obtain a BG and DO-free target object segmentation result. Furthermore, we propose a prototypical contrastive learning algorithm to improve the model ability of distinguishing the target object from DOs. Extensive experiments on both PASCAL-5iand COCO-20idatasets show that our approach is effective despite its simplicity. Code is available at https://github.com/LIUYUANWEI98/NERTNet Yuanwei Liu, Nian Liu 0002, Qinglong Cao, Xiwen Yao, Junwei Han 0001, Ling Shao 0001 |
CVPR | 3 |
| 2022 | Scale-Aware Detailed Matching for Few-Shot Aerial Image Semantic SegmentationabstractFew-shot semantic segmentation, aiming to segment query images with a few annotated support samples, has drawn increasing attention. Most existing few-shot methods leverage the single prototype obtained from global average pooling to represent all support information and further use the extracted prototype to segment the query images in a matching manner. Although promising results for natural images have been reported, these methods cannot be directly applied on aerial images. The main reason comes from that the extracted single support prototype can only provide a coarse guidance for matching between query and support images and could not handle the large variance of objects’ appearances and scales. To deal with these challenges on aerial images, we propose a scale-aware few-shot semantic segmentation network to perform detailed matching with multiple prototypes. More specifically, the detailed matching module is first constructed to compute the pixel-level similarity between the query features and the extracted multiple support prototypes for providing more accurate parsing guidance. Subsequently, to address the problem of scale imbalance, the scale-aware focal loss is designed to dynamically down-weight the loss assigned to large well-parsed objects and focus training on tiny hard-parsed objects. To facilitate the reproducible research on the task of few-shot semantic segmentation in aerial images, we further provide a few-shot segmentation benchmark iSAID-$5^{\mathrm {i}}$constructed from the large-scale iSAID dataset[1]. Comprehensive experiments and comparisons with the state-of-the-art few-shot segmentation methods on the iSAID-$5^{\mathrm {i}}$dataset clearly demonstrate the superiority of our proposed method. The code and dataset are available athttps://github.com/caoql98/SDM. Xiwen Yao, Qinglong Cao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |