EDBT 2026 Demo / reviewers in the wild / expert
Kyle Gao
dblp:303/9577 · also Kyle Yilin Gao
· DBLP profile ↗
14ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-8320-6308ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Multi-View Images to Learn Domain-Invariant Discriminative Embeddings for Cross-View Geo-LocalizationabstractCross-view geo-localization (CVGL) aims to match images of the same location captured from different viewpoints, such as those captured by Unmanned Aerial Vehicles (UAVs) and satellite platforms. The task is particularly challenging due to significant variations in scale, viewpoint, and illumination. Most existing methods employ symmetric sampling strategy to construct drone–satellite image pairs for deep metric learning, but neglect the potential of incorporating multi-view drone images to enhance the viewpoint robustness of features. To address this, we propose leveraging multi-view images to learn Domain-Invariant Discriminative Embeddings (DIDE) for CVGL. DIDE introduces an Inter-view Feature Aggregation Module (IFAM), which dynamically integrates multi-view drone information into robust embeddings. These are used in contrastive learning with satellite embeddings within batches to learn view-invariant discriminative features, while representation learning further improves scene discrimination across batches. To reduce the domain gap, DIDE constructs and aligns drone and satellite prototypes for effective cross-domain feature alignment. Furthermore, we adopt a parameter-efficient transfer learning strategy that leverages the capabilities of pre-trained foundation models while fine-tuning only dual adapters, significantly reducing the trainable parameters. DIDE achieves the state-of-the-art on University-1652 and University-160k, competitive results on SUES-200, and demonstrates strong cross-dataset transferability, with fewer training parameters and lower computational cost. Ziyi Chen 0001, Dilong Li, Jin Gou, Cheng Wang 0003, Kyle Gao, Jonathan Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | L4DR: LiDAR-4DRadar Fusion for Weather-Robust 3D Object DetectionabstractLiDAR-based 3D object detection is crucial for autonomous driving. However, due to the quality deterioration of LiDAR point clouds, it suffers from performance degradation in adverse weather conditions. Fusing LiDAR with the weatherrobust 4D radar sensor is expected to solve this problem; however, it faces challenges of significant differences in terms of data quality and the degree of degradation in adverse weather. To address these issues, we introduce L4DR, a weather-robust 3D object detection method that effectively achieves LiDAR and 4D Radar fusion. Our L4DR proposes Multi-Modal Encoding (MME) and Foreground-Aware Denoising (FAD) modules to reconcile sensor gaps, which is the first exploration of the complementarity of early fusion between LiDAR and 4D radar. Additionally, we design an Inter-Modal and IntraModal ({IM}2) parallel feature extraction backbone coupled with a Multi-Scale Gated Fusion (MSGF) module to counteract the varying degrees of sensor degradation under adverse weather conditions. Experimental evaluation on a VoD dataset with simulated fog proves that L4DR is more adaptable to changing weather conditions. It delivers a significant performance increase under different fog levels, improving the 3D mAP by up to 20.0% over the traditional LiDAR-only approach. Moreover, the results on the K-Radar dataset validate the consistent performance improvement of L4DR in realworld adverse weather conditions. Xun Huang 0003, Ziyu Xu 0002, Qiming Xia, Yan Xia 0003, Jonathan Li 0001, Kyle Gao, Chenglu Wen, Cheng Wang 0003 |
AAAI | 8 |
| 2025 | Digital Buildings Analysis: 3-D Modeling, GIS Integration, and Visual Descriptions Using Gaussian Splatting, ChatGPT/Deepseek, and Google Maps PlatformabstractWe propose a Digital Building Analysis (DBA), a digital system for building-scale cloud-based data integration and data analytics. By connecting to cloud mapping platforms such as Google Map Platforms APIs, by leveraging state-of-the-art multi-agent Large Language Models data analysis using ChatGPT(4o) and Deepseek-V3/R1, and by using our Gaussian Splatting-based mesh extraction pipeline, our framework can retrieve a building’s 3D model, visual descriptions, and achieve cloud-based mapping integration with large language model-based data analytics using a building’s address, postal code, or geographic coordinates, and be easily extended to perform data analysis on other cloud-based data streams. Kyle Gao, Dening Lu, Liangzhi Li 0002, Hongjie He 0003, Linlin Xu, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | Enhanced 3-D Urban Scene Reconstruction and Point Cloud Densification Using Gaussian Splatting and Google Earth ImageryabstractThree-dimensional urban scene reconstruction and modeling is a crucial research area. From a technical perspective, it is an interdisciplinary research area spanning computer vision, computer graphics, and photogrammetry. Its applications span across multiple disciplines including autonomous navigation with 3-D scene understanding, remote sensing/photogrammetry for the creation of 3-D maps from aerial/drone/satellite images, geographic information systems with urban digital twins, augmented and virtual reality with photorealistic scene reconstructions. Using Google Earth imagery, we create a 3-D Gaussian splatting (3DGS) model of the Waterloo region centered on the University of Waterloo, and are able to achieve view-synthesis results far exceeding previous 3-D view-synthesis results based on neural radiance fields (NeRFs)which we demonstrate in our benchmark. We also retrieve the 3-D geometry of the scene using the 3-D point cloud extracted from the 3DGS model, thereby reconstructing both the 3-D geometry and photorealistic lighting of the large-scale urban scene. Kyle Gao, Dening Lu, Hongjie He 0003, Linlin Xu, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Exploring Token Serialization for Mamba-Based LiDAR Point Cloud SegmentationabstractLiDAR point cloud segmentation has increasingly benefited from the application of Mamba-based models. However, unordered and irregular natures of point clouds necessitates serialization, which significantly impacts the performance of Mamba-based methods. This paper explores the critical role of token serialization in Mamba-based point cloud processing, using the pure Mamba network, PointMamba, as the baseline. We systematically investigated existing point cloud serialization methods, evaluating their performance on two challenging LiDAR datasets: the airborne MultiSpectral LiDAR (MS-LiDAR) dataset and the aerial DALES dataset. To explore the inherent factors of serialization contributing to Mamba’s performance, we design novel indicators for serialization quality, focusing on spatial and semantic proximity. These indicators are validated across all datasets, offering a valuable reference and guidance for advancing token serialization in Mamba-based point cloud processing. Guided by these indicators, we proposed a new point cloud serialization method that integrates spatial and semantic features through a weighted comprehensive distance matrix. The proposed method achieves superior accuracy on both LiDAR datasets, surpassing existing approaches, and establishes a strong foundation for advancing Mamba-based point cloud processing. Dening Lu, Kyle Gao, Jonathan Li 0001, Dedong Zhang, Linlin Xu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Synchronizing Spatiotemporal Reflectance Fusion via Dual Bayesian Nonparametric Inference and Explicit DownsamplingabstractRemote sensing images from a single sensor restrict in terms of spatial resolution and temporal resolution and they cannot simultaneously achieve high-precision and high-frequency synchronous observations. This paper proposes a spatiotemporal reflectance fusion method using double Bayesian nonparametric inferences for coupled feature space. Overcoming the limitations of existing fusion models, our approach incorporates downsampling, establishes a coupled feature space fusion model, and introduces a Beta-Bernoulli process to learn specific dictionaries. A dictionary atomic indicator ensures precise sparse projection relationships between coupled feature spaces. To address resolution disparities, a two-step Bayesian nonparametric fusion framework is employed. Experimental results demonstrate superior fusion performance, particularly in capturing surface reflectance changes in images depicting phenology or land-cover type changes. This method offers a promising solution to balance spatial and temporal resolutions, enhancing dynamic ecosystem monitoring and phenological parameter inversion. Biao Zhang 0001, Hongjie He 0003, Zhouzhou Liu, Kyle Gao |
IGARSS | 5 |
| 2024 | 3DGTN: 3-D Dual-Attention GLocal Transformer Network for Point Cloud Classification and SegmentationabstractAlthough the application of Transformers to 3-D point cloud processing has achieved significant progress and success, it is still challenging for existing 3-D Transformer methods to efficiently and accurately learn both valuable global and local features for improved applications. This article presents a novel point cloud representational learning network, called 3-D Dual Self-attention global local (GLocal) Transformer Network (3DGTN), for improved feature learning in both classification and segmentation tasks, with the following key contributions. First, a GLocal feature learning (GFL) block with the dual self-attention mechanism [i.e., a novel point-patch self-attention, called PPSA, and a channel-wise self-attention (CSA)] is designed to efficiently learn the global and local context information. Second, the GFL block is integrated with a multiscale Graph Convolution-based local feature aggregation (LFA) block, leading to a GLocal information extraction module that can efficiently capture critical information. Third, a series of GLocal modules are used to construct a new hierarchical encoder–decoder structure to enable the learning of information in different scales in a hierarchical manner. The proposed framework is evaluated on both classification and segmentation datasets, demonstrating that the proposed method is capable of outperforming many state-of-the-art methods on both synthetic and LiDAR data. Our code has been released athttps://github.com/d62lu/3DGTN. Dening Lu, Kyle Gao, Qian Xie 0001, Linlin Xu, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Research on Fast Detection Method of Wind Turbine in Remote Sensing Image Land Area Based on YoloabstractWith the development of the social economy, wind turbines are taking up a larger and larger share of new energy sources. The detection of the number and spatial distribution of wind turbines in remotely sensed images holds great scientific significance. Wind turbines are difficult to identify in remote sensing images therefore, a fast detection method based on deep learning is proposed. First, we extract potential wind turbine candidate regions from wind speed, slope, and land use data. Second, the YOLO v5 model was trained using our labeled wind turbine detection dataset. Finally, the images of the candidate regions were used for wind turbine detection using the trained optimal model. The proposed method was demonstrated to have a recall of 94.87% and an accuracy of 82.04% through the experimental results. The proposed method for wind turbine detection is not only reasonable and effective but also offers a heightened level of efficiency. Deliang Chen, Taotao Cheng, Kyle Gao, Sarah Narges Fatholahi, Jonathan Li 0001 |
IGARSS | 4 |
| 2023 | Natural Language Aided Remote Sensing Image Few-Shot ClassificationabstractWe aim to improve the efficiency of traditional deep learning methods for remote sensing by reducing the reliance on annotated data and minimizing training time. Instead of using large-scale unimodal remote sensing image datasets for pre-training, we propose the use of multimodal data (text-image pairs), which we believe to be more effective. To enhance the model's generalization performance in the remote sensing domain and achieve accurate remote sensing image scene classification, we employ the Feature Adaptive Embedding Module. For this purpose, we introduce a cross-modal comparison learning network that is based on openly accessible generalized datasets. This network is capable of recognizing specific photo scenarios from remote sensing photographs, maximizing the accuracy of classification. Deliang Chen, Jianbo Xiao, Kyle Gao, Sarah Narges Fatholahi, Jonathan Li 0001 |
IGARSS | 3 |
| 2023 | Nighttime Light Missing Data Retrieval Using Modis Version 6 Satellite Data and Mask Dilated Partial Convolutional Neural NetworkabstractNighttime Lights (NTLs) remote sensing imagery contains tremendous information and has been shown to accurately predict a region’s human dynamics, economic health and energy consumption. Despite its usefulness, NTLs imagery is less widely available than other remote sensing data modalities. Several challenges appear when attempting to reconstruct NTLs data, either from other data modalities or existing NTLs data. These include complex non-linear relationships between NTLs and multispectral bands, non-matching spatial and temporal coverage, and different atmospheric and cloud conditions. This study attempts to create an out-of-the-box model that compensates for missing NTLs data using widely available daytime data in a broadly generalizable manner. The proposed project has two objectives: the construction of an image-to-image dataset mapping daytime multispectral images (MODIS V6 Land Surface Reflectance, MODIS V6 Land Cover, MODIS V6 Vegetation Indices) to NTLs images, and the reconstruction of NTLs data using deep learning techniques by researching, creating, and employing the state-of-the-art architecture of the Mask Partial Convolutional Neural Network in conjunction with dilated convolutions. The project will facilitate the training of new models for predicting missing NTLs and make NTLs data more accessible for future remote sensing research. Xuanchen Liu, Shuxin Qiao, Kyle Gao, Hongjie He 0003, Lingfei Ma, Jonathan Li 0001 |
IGARSS | 3 |
| 2023 | BrGAN: Blur Resist Generative Adversarial Network With Multiple Joint Dilated Residual Convolutions for Chlorophyll Color Image RestorationabstractThis paper presents a Blur Resist Generative Adversarial Network (GAN) (BrGAN) with multiple joint dilated residual convolutions for chlorophyll image restoration of the Geostationary Ocean Color Imager (GOCI). First, a publicly available dataset was built to support this study. Second, a multiple attention perception mechanism and a multiple joint dilated residual convolution module was proposed to cope with the challenge of large missing areas in GOCI chlorophyll images. Third, a patch GAN based discrimination module was proposed to avoid the restored areas with generating mosaic and shadows. Our experimental results demonstrate that the BrGAN can reach 37.06 in the peak signal-to-noise ratio (PSNR) and 0.0485 in the Learned Perceptual Image Patch Similarity (LPIPS), respectively. The comparative study shows that the BrGAN achieves the highest effectiveness and advancement among other seven state-of-the-art methods. Ziyi Chen 0001, Yuhua Luo, Yiping Chen 0002, Jing Wang 0049, Dilong Li, Kyle Gao, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | SAR-Optical Image Matching With Semantic Position Probability DistributionabstractWe propose a deep learning framework of Semantic Position Probability Distribution for SAR-optical image matching, termed as SPPD. Unlike the pixel-by-pixel searching matching method, a correspondence is directly obtained by an outputted matching position probability distribution. First, multiscale pyramidal features are created for each pixel in the SAR and optical images by using two weight-sharing ResNet-50 + Feature Pyramid Network (FPN) networks. The features containing high-level semantic information are then embedded into the proposed image Position Attention Module to obtain the spatial position dependencies between two images. Then, we present a loss function for semantic position matching to optimize the network from both semantic information and pixel alignment perspectives, converting the probability distribution of semantic matching positions into a point-to-point matching problem. In this paper, the SAR and optical images are set as the sensed and reference images. The effects of different image sizes, training label types, and loss function weights on matching accuracy are explored to obtain the optimal parameter settings for matching. The experimental results show that the proposed method is insensitive to image deformation and achieves cross-modal matching for SAR-optical images with high accuracy compared with the best matching method on different scene images, with several orders of magnitude faster inferences time. Liangzhi Li 0002, Kyle Gao, Hongjie He 0003, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | 3DCTN: 3D Convolution-Transformer Network for Point Cloud ClassificationabstractPoint cloud classification is a fundamental task in 3D applications. However, it is challenging to achieve effective feature learning due to the irregularity and unordered nature of point clouds. Lately, 3D Transformers have been adopted to improve point cloud processing. Nevertheless, massive Transformer layers tend to incur huge computational and memory costs. This paper presented a novel hierarchical framework that incorporated convolutions with Transformers for point cloud classification, named 3D Convolution-Transformer Network (3DCTN). It combined the strong local feature learning ability of convolutions with the remarkable global context modeling capability of Transformers. Our method had two main modules operating on the downsampling point sets. Each module consisted of a multi-scale local feature aggregating (LFA) block and a global feature learning (GFL) block, which were implemented by using the Graph Convolution and Transformer respectively. We also conducted a detailed investigation on a series of self-attention variants to explore better performance for our network. Various experiments on ModelNet40 and ScanObjectNN datasets demonstrated that our method achieves state-of-the-art classification performance with a lightweight design. The code is publicly available athttps://github.com/d62lu/3DCTN. Dening Lu, Qian Xie 0001, Kyle Gao, Linlin Xu, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | The Impact of Data Volume on Performance of Depp Learning Based Building Rooftop Extraction Using Very High Spatial Resolution Aerial ImagesabstractBuilding rooftop data are of importance in several urban applications and in natural disaster management. In contrast to traditional surveying and mapping, by using high spatial resolution aerial images, deep learning-based building rooftops extraction methods are efficient and accurate. Although more training data is preferred in deep learning-based tasks, the effect of data volume on building extraction models is underexplored. Therefore, the paper explores the impact of data volume on the performance of building rooftop extraction from very-high-spatial-resolution (VHSR) images using deep learning-based methods. To do so, we manually labelled 0.12m spatial resolution aerial images and perform a comparative analysis of models trained on datasets of different sizes using popular deep learning architectures for segmentation tasks, including Fully Convolutional Networks (FCN)-8s, U-Net and DeepLabv3+. The experiments showed that with more training data, algorithms converged faster and achieved higher accuracy, while better algorithms were able to better mitigate the lack of training data. Hongjie He 0003, Yuwei Cai, Zijian Jiang, Qiutong Yu, Sarah Narges Fatholahi, Yan Liu 0043, Hasti Andon Petrosians, Bingxu Hu, Liyuan Qing, Zhehan Zhang, Hongzhang Xu, Kyle Gao, Linlin Xu, Jonathan Li 0001 |
IGARSS | 16 |