VLDB 2026 Research / reviewers in the wild / expert
Kailang Cao
dblp:248/4693
· DBLP profile ↗
10ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-9093-2711ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hyperspectral Target Detection Based on Generative Self-Supervised Learning With Wavelet TransformabstractRecently, generative self-supervised learning (GSSL) has gained extensive attention in hyperspectral remote sensing. For the hyperspectral target detection (HTD) task, traditional GSSL-based algorithms usually require hyperspectral images (HSIs) as additional datasets for pretraining, which are relatively resource-intensive and time-consuming. To better interpret the spectral-spatial information of HSIs while alleviating the dependence on large-scale hyperspectral datasets, we develop a novel two-stage framework for HTD based on GSSL in this article. In the preprocessing for the input HSI, a dimensional transformation (DT) module and a coarse detection reference (CDR) module are constructed to produce feature patches as training samples for subsequent pretraining and fine-tuning. In the pretraining stage for spectral-spatial reconstruction, we construct an asymmetric autoencoder (AE) architecture which leverages the transformer blocks with long-range perception to extract generalized features and explore discriminative feature representations of the input HSI. Specifically, a dual-stream wavelet patch embedding (DWPE) module is proposed to integrate the wavelet transform (WT) mechanism with the convolutional neural networks (CNNs), which extracts robust spectral-spatial features by performing convolutional operations with different frequency components of WT. In the fine-tuning stage, a novel signature-constrained cross-entropy (SC-CE) loss function is proposed to constrain the network optimization. For the final detection, a pixel-level fusion based on coarse detection based pixel-level fusion (CDPF) module is employed after inference to further suppress the interference from background. Experimental results on six real HSIs demonstrate that the proposed method achieves superior detection performance while maintaining the generalization of the pretrained model. Shuai Wang 0057, Yunsong Li 0001, Weiying Xie, Kai Jiang 0001, Kailang Cao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | A Signature-Constrained Two-Stage Framework for Hyperspectral Target Detection Based on Generative Self-Supervised LearningabstractRecently, hyperspectral target detection (HTD) technique based on deep learning (DL) has been developed rapidly. However, existing algorithms show poor generalization across different hyperspectral images (HSIs), where repeated training and inference based on limited prior information are necessary. To liberate HTD from dependence on the quantity and the quality of training samples, this article proposes a signature-constrained two-stage framework for HTD (HTD-STF) based on generative self-supervised learning (GSSL). In the first stage for pre-training, to realize spectral-spatial reconstruction, we build an asymmetric autoencoder (AE) employing transformer blocks with long-range perception for generalized feature extraction. During pre-training, the spectral-spatial similarity loss is designed to improve the effect of reconstruction. In the second stage for fine-tuning and detection, the signature is utilized in preprocessing, training and inference, respectively. Specifically, the coarse sample mining and tiling strategy in preprocessing not only facilitates the framework in flexible input dimension, but also provides pseudo labels for end-to-end training. During training, we adopt the signature as guidance for feature-level fusion, which alleviates the impact of sample imbalance. After training, the final inference based on pixel-level fusion refines the original output. For ideal GSSL, the HyperMix-10K, a new large-scale hyperspectral dataset, has been constructed in this work, which contains numerous unlabeled HSIs captured in various scenes. Experimental results and analysis on real HSIs verify the effectiveness and generalization ability of the HTD-STF. Shuai Wang 0057, Yunsong Li 0001, Weiying Xie, Kai Jiang 0001, Kailang Cao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Model-Driven Deep Pipeline With Uncertainty-Aware Bundle Adjustment for Satellite PhotogrammetryabstractBundle adjustment (BA), a vital technology in satellite photogrammetry, directly determines the quality of geographic information mapping. However, the existing BA methods suffer from bottlenecks in the cases of limited stereo views caused by input guidance inadequacy and biased modeling. To conquer these issues, a model-driven satellite photogrammetry deep pipeline (SPDP) is proposed in this article. Specifically, for the triplets of remote sensing images (RSIs), the fusion feature maps are extracted by our attention-driven multiscale feature extractor (AMFE), which emphasizes the image information and provides guidance for the subsequent multiview geometric processing. Following that, with the feature error volume as input, a dedicated feature-metric error perceptron module (FEPM) is built to infer the observation uncertainty and predict the pixel-wise compensations. Furthermore, a novel uncertainty-aware BA (UBA) is implemented to derive accurate and robust 3-D point clouds, which introduces the BA model transformation and the specialized iterative refinement to enhance the observation error elimination capability. The detailed experimental results demonstrate the feasibility and effectiveness of the proposed pipeline, which is significant for remote sensing surveys and mapping. Kailang Cao, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Structure-Aware Graph Convolution Network for Point Cloud ParsingabstractPoint clouds are becoming a popular medium to describe 3D scenes, benefitting from their accuracy and completeness in expressing the spatial and geometrical information of objects. However, due to the disorder and uneven distribution nature, merely selecting neighbors for point clouds in Euclidean space is inefficient and position-ignoring. To fill this gap, we propose a structure-aware graph convolution network (SA-GCN), which consists of an adaptive dilated KNN module (ADKNN), a learnable graph filter (LGF), and a structure-aware feature transformation module (SFT). Specially, the ADKNN module can dynamically adjust the range of grouping neighbor points, while being universal to improve the performance of arbitrary KNN-based methods. Moreover, with the localized auxiliary information provided by LGF, our SFT module disentangles the spatial details as a sort of coding guidance for better deep feature representations. Extensive experimental results on point cloud classification and segmentation tasks demonstrate the superiority of our proposed network. Fengda Hao, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Kailang Cao |
IEEE Trans. Multim. | 5 |
| 2022 | Cascaded geometric feature modulation network for point cloud processing
Fengda Hao, Rui Song 0003, Jiaojiao Li 0001, Kailang Cao, Yunsong Li 0001 |
Neurocomputing | 4 |
| 2021 | An Extreme Learning Machine Correction Network for High Precision Satellite Attitude DeterminationabstractThe fusion framework of star sensor and gyro based on adaptive Kalman filter is widely used in satellite pose estimation. However, the discretization and linearization inevitably introduce system errors, which degrades of the filtering accuracy. To address this problem, we propose a high-precision satellite attitude determination algorithm based on extreme learning machine network correction. We design a dedicated network for error compensation and trained the parameters effectively. In attitude calculation procedure, the forward fusion filtering of star sensor and gyro data is performed firstly by using the adaptive Kalman filter. Then the filtering estimation results are compensated by the extreme learning machine network proposed in this paper. After that, backward smoothing is performed to solve the high-precision attitude. Simulation results show that armed with the compensation procedure of the proposed extreme learning machine network, the accuracy of estimated pose is significantly improved. Kailang Cao, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Weijiao Jiang |
IGARSS | 1 |
| 2021 | Pansharpening of Hyperspectral Images with Detail Guided Feature ModulationabstractPansharpening of hyperspectral image (HSI), which makes use of the detail information contained in the high-resolution panchromatic (HR-PAN) image to sharpen the low-resolution HSI (LR-HSI), is an essential technology to enhance the spatial resolution of HSI. In this paper, we propose a detail guided feature modulation residual network (DGFM-Net) to address the HS pansharpening problem, which is able to effectively integrate details extracted from the PAN image into the pansharpened result. Specifically, we elaborately design a novel feature modulation (FM) module with the guidance of PAN detail information to modulate HSI features flexibly and incorporate PAN details adaptively. The modulated features are then fed to the residual reconstruction (RR) block to recover the difference between the upsampled HSI and the HR-HSI by efficient residual learning. Finally, the upsampled HSI is combined with the estimated residual HSI to produce the desired HR-HSI. Experiments on the Pavia Center data set confirm that the proposed DGFM-Net outperforms several state-of-the-art HS pansharpening methods. Yuxuan Zheng, Jiaojiao Li 0001, Yunsong Li 0001, Kailang Cao |
IGARSS | 4 |
| 2021 | ASFM-Net: Asymmetrical Siamese Feature Matching Network for Point CompletionabstractWe tackle the problem of object completion from point clouds and propose a novel point cloud completion network employing an Asymmetrical Siamese Feature Matching strategy, termed as ASFM-Net. Specifically, the Siamese auto-encoder neural network is adopted to map the partial and complete input point cloud into a shared latent space, which can capture detailed shape prior. Then we design an iterative refinement unit to generate complete shapes with fine-grained details by integrating prior information. Experiments are conducted on the PCN dataset and the Completion3D benchmark, demonstrating the state-of-the-art performance of the proposed ASFM-Net. Our method achieves the 1st place in the leaderboard of Completion3D and outperforms existing methods with a large margin, about 12%. The codes and trained models are released publicly at https://github.com/Yan-Xia/ASFM-Net. Yaqi Xia, Yan Xia 0003, Wei Li 0111, Rui Song 0003, Kailang Cao, Uwe Stilla |
ACM Multimedia | 5 |
| 2020 | Deep Residual Learning for Boosting the Accuracy of Hyperspectral PansharpeningabstractRecently, deep learning (DL) has gained impressive achievements in the field of remote sensing image fusion. However, most of the previous DL-based fusion methods are originally designed for multispectral pansharpening, which cannot be readily employed to hyperspectral pansharpening due to the much wider spectral range and lower spatial resolution of a hyperspectral image (HSI). In this letter, a novel framework based on deep residual learning is proposed for hyperspectral pansharpening. The proposed framework consists mainly of two parts. First, the initialized HSI with the enhanced spatial resolution is generated through contrast limited adaptive histogram equalization (CLAHE) and guided filter. Then, a deep residual convolutional neural network (DRCNN) is introduced to map the residuals between the initialized HSI and the reference HSI for further boosting the fusion accuracy. Experimental results demonstrate that the proposed framework can achieve superior performance compared with the existing state-of-the-art pansharpening methods, especially in terms of edge details enhancement. Yuxuan Zheng, Jiaojiao Li 0001, Yunsong Li 0001, Kailang Cao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2019 | Discriminative Feature Learning With Distance Constrained Stacked Sparse Autoencoder for Hyperspectral Target DetectionabstractTarget detection (TD) is one of the major tasks in hyperspectral image (HSI) processing, and its performance is greatly affected by the background. Feature extraction (FE) has been an effective way to mine discriminative information, especially FE based on deep learning, which can learn the intrinsic properties of data to further improve the detection performance. Unlike supervised networks, unsupervised stacked sparse autoencoders (SSAEs) can learn deep and nonlinear features without any labeled data. However, SSAEs usually require a supervised fine-tuned model to obtain better discrimination, which is not feasible for TD, since the prior information is generally insufficient. In this letter, we introduce a distance constraint that is added to the SSAE to form a new distance constrained SSAE (DCSSAE) network. Specifically, the distance constraint maximizes the distinction between the target pixels and other background pixels in the feature space. Then, using the discriminative features learned from the DCSSAE, a simple detector using radial basis function kernel is derived for background suppression. Experiments on two HSIs demonstrate that the deep spectral features learned from the DCSSAE are more distinguishable, and our proposed detector, namely, the DCSSAE detector, outperforms several popular detectors, especially in background suppression. Yanzi Shi, Jie Lei 0001, Yaping Yin, Kailang Cao, Yunsong Li 0001, Chein-I Chang |
IEEE Geosci. Remote. Sens. Lett. | 4 |