VLDB 2026 Research / reviewers in the wild / expert
Hongjie He 0003
dblp:34/5183-3
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0003-3839-5821ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BCWildfire: A Long-term Multi-factor Dataset and Deep Learning Benchmark for Boreal Wildfire Risk PredictionabstractWildfire risk prediction remains a critical yet challenging task due to the complex interactions among fuel conditions, meteorology, topography, and human activity. Despite growing interest in data-driven approaches, publicly available benchmark datasets that support long-term temporal modeling, large-scale spatial coverage, and multimodal drivers remain scarce. To address this gap, we present a 25-year, daily-resolution wildfire dataset covering 240 million hectares across British Columbia and surrounding regions. The dataset includes 38 covariates, encompassing active fire detections, weather variables, fuel conditions, terrain features, and anthropogenic factors. Using this benchmark, we evaluate a diverse set of time-series forecasting models, including CNN-based, linear-based, Transformer-based, and Mamba-based architectures. We also investigate effectiveness of position embedding and the relative importance of different fire-driving factors. Zhengsen Xu, Sibo Cheng, Hongjie He 0003, Wentao Sun, Jonathan Li 0001, Lincoln Linlin Xu |
AAAI | 4 |
| 2025 | Digital Buildings Analysis: 3-D Modeling, GIS Integration, and Visual Descriptions Using Gaussian Splatting, ChatGPT/Deepseek, and Google Maps PlatformabstractWe propose a Digital Building Analysis (DBA), a digital system for building-scale cloud-based data integration and data analytics. By connecting to cloud mapping platforms such as Google Map Platforms APIs, by leveraging state-of-the-art multi-agent Large Language Models data analysis using ChatGPT(4o) and Deepseek-V3/R1, and by using our Gaussian Splatting-based mesh extraction pipeline, our framework can retrieve a building’s 3D model, visual descriptions, and achieve cloud-based mapping integration with large language model-based data analytics using a building’s address, postal code, or geographic coordinates, and be easily extended to perform data analysis on other cloud-based data streams. Kyle Gao, Dening Lu, Liangzhi Li 0002, Hongjie He 0003, Linlin Xu, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Enhanced 3-D Urban Scene Reconstruction and Point Cloud Densification Using Gaussian Splatting and Google Earth ImageryabstractThree-dimensional urban scene reconstruction and modeling is a crucial research area. From a technical perspective, it is an interdisciplinary research area spanning computer vision, computer graphics, and photogrammetry. Its applications span across multiple disciplines including autonomous navigation with 3-D scene understanding, remote sensing/photogrammetry for the creation of 3-D maps from aerial/drone/satellite images, geographic information systems with urban digital twins, augmented and virtual reality with photorealistic scene reconstructions. Using Google Earth imagery, we create a 3-D Gaussian splatting (3DGS) model of the Waterloo region centered on the University of Waterloo, and are able to achieve view-synthesis results far exceeding previous 3-D view-synthesis results based on neural radiance fields (NeRFs)which we demonstrate in our benchmark. We also retrieve the 3-D geometry of the scene using the 3-D point cloud extracted from the 3DGS model, thereby reconstructing both the 3-D geometry and photorealistic lighting of the large-scale urban scene. Kyle Gao, Dening Lu, Hongjie He 0003, Linlin Xu, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Loss Functions Analysis of Performance Improvements in Single-Image Super-ResolutionabstractThe building rooftop delineation plays a critical role in urban planning and management. Current methods on improving the delineation accuracy of the building rooftops mostly focused on the data improvements and methodology improvement. However, as the current deep learning networks have already reached high accuracy, the data quality has then become the important topic. Although loss functions serve the purpose of quantifying reconstruction errors and directing the optimization process of the model in super-resolution (SR), limited research related to their impact on SR were implemented. In this study, we focused on improving the spatial resolution of the building datasets by investigating numerous loss functions, including the Mean Absolute Error (MAE) loss function, the Mean-Squared Error (MSE) loss function, the SmoothL1 loss function and the Charbonnier loss function. With our proposed Single-Image Super-Resolution (SISR) network, Residual Feature Aggregation- Pyramid Vision Transformer- Involution Network (RFA-PVTInvoNet) and other typical SISR networks, the loss functions were applied to further enhance the performance. By using Peak Signal-Noise Ratio (PSNR) and Similarity Structure Index Measurement (SSIM) as the evaluation metrics, the RFA-PVTInvoNet with Charbonnier loss function showed the highest performance compared with other approaches, with PSNR of 22.046 dB and SSIM of 0.502 on the WHU Building Dataset, which demonstrate the superior performance of the Charbonnier loss in SISR. Yuwei Cai, Hongjie He 0003, Zhimeng He |
IGARSS | 2 |
| 2024 | Synchronizing Spatiotemporal Reflectance Fusion via Dual Bayesian Nonparametric Inference and Explicit DownsamplingabstractRemote sensing images from a single sensor restrict in terms of spatial resolution and temporal resolution and they cannot simultaneously achieve high-precision and high-frequency synchronous observations. This paper proposes a spatiotemporal reflectance fusion method using double Bayesian nonparametric inferences for coupled feature space. Overcoming the limitations of existing fusion models, our approach incorporates downsampling, establishes a coupled feature space fusion model, and introduces a Beta-Bernoulli process to learn specific dictionaries. A dictionary atomic indicator ensures precise sparse projection relationships between coupled feature spaces. To address resolution disparities, a two-step Bayesian nonparametric fusion framework is employed. Experimental results demonstrate superior fusion performance, particularly in capturing surface reflectance changes in images depicting phenology or land-cover type changes. This method offers a promising solution to balance spatial and temporal resolutions, enhancing dynamic ecosystem monitoring and phenological parameter inversion. Biao Zhang 0001, Hongjie He 0003, Zhouzhou Liu, Kyle Gao |
IGARSS | 3 |
| 2024 | An Enhanced Trans-Involution Network for Building Footprint Extraction from High Resolution OrthoimageryabstractThe amalgamation of data visualization and geospatial insights has driven significantly advancements in remote sensing for applications like damage detection and urban planning, particularly building rooftop extraction from high spatial satellite imagery. However, building rooftop extraction using deep learning methods often results in outputs with unclear margin delineation. In this study, we propose a novel approach that combines Transformer architectures, involution, and an enhanced U-net (E-Unet) [1] to improve building footprint extraction performance. Our method demonstrates remarkable accuracy in complex urban environments in the Waterloo Building Dataset. Transformers, renowned for their success in natural language processing, have excelled in adeptness at analyzing sequential data. By using the embedding and multi-head attention blocks, this method is becoming increasingly valuable for building extraction. Involution, in turn, augments neural networks by providing spatial-specific adaptability, effectively extracting inter-band features and surpassing convolutional limitations. Through comprehensive comparative model experiments on the Waterloo building dataset, the optimal architecture was identified. The model significantly enhances accuracy when the Transformer architecture is integrated at the output of the E-Unet. Our proposed network achieves outstanding at the crucial metric values of IoU, mIoU, Precision, F1-score in 81.2, 91.9, 92.9, 89.8 (%), surpassing established frameworks such as FCN-8s, U-Net, DeepLab v3+, Fast Statistical Convolutional Neural Network (SCNN), High-Resolution Net (HRNet) v2, Mask R-CNN, as well as E-Unet. Zhimeng He, Yuwei Cai, Hongjie He 0003, Xinyan Xian, Brian W. Barrett |
IGARSS | 3 |
| 2024 | Detection of Small Objects from UAV Imagery via an Improved Swin TransformerabstractAutomated detection of small objects such as vehicles in images of complex urban environments taken by unmanned aerial vehicles (UAVs) is one of the most challenging tasks in computer vision and remote sensing communities. Convolutional neural networks (CNNs)-based deep learning models have been widely used to automatically detect objects in UAV images given their high performance. However, their detection accuracy is still unsatisfactory, particularly when it comes to small objects, due to the shortcomings of CNNs. Therefore, in this study, we propose a Swin Transformer-based model that incorporates convolutions with the Swin Transformer to extract more local information, mitigating the problem of small object detection from complex backgrounds in UAV images and further improving the detection accuracy. By using the Swin Transformer, our model leverages both the local feature extraction of convolutions and the global feature modeling of transformers. The framework comprises two primary modules: a Local Context Enhancement (LCE) module and a Residual U-Feature Pyramid Network (RSU-FPN) module. Additionally, it incorporates a loss function that combines L1 loss with Normalized Gaussian Wasserstein Distance. Our experimental results obtained on the UAV Detection and Tracking (UAVDT) dataset indicated that our proposed method increased the average precision (AP) by 21.6%, 22.3% and 25.5% over Cascade Region-based CNN (R-CNN), Faster Region based CNN (R-CNN) with ResNet-50 and with Pyramid Vision Transformer (PVT) B0, and Dynamic R-CNN detectors, respectively, indicating its effectiveness and reliability on small object detection from UAV images. Weidong Liang, Jingtian Tan, Hongjie He 0003, Hongzhang Xu, Jonathan Li 0001 |
IGARSS | 3 |
| 2024 | UnderstAnding Bag of Tricks of Deep Learning-Based Semantic Segmentation in Pavement Crack DetectionabstractThe rapid development of deep learning has significantly enhanced the performance of models in the detection of pavement cracks, thereby facilitating the deployment of deep learning-based approaches into real-world applications. Nevertheless, it is worth noting that deep learning-based crack detection models represent complex amalgamations of deep learning networks and model training strategies, with the latter frequently being overlooked. Therefore, in this paper, we focus on various techniques in data augmentation, and model deployment stages that are commonly employed in deep learning-based semantic segmentation models. Through extensive experiments, the effectiveness of these techniques in crack detection is evaluated, aiming to provide guidance for subsequent crack detection experiments and project implementations. Consequently, the experiments demonstrate data augmentation methods such as color jittering and CutMix can effectively improve model performance by altering the distribution of the training dataset. Additionally, in case of crack datasets with limited samples and severe class imbalance, loss function selection and pre-training weights can be crucial in model deployment. Zhengsen Xu, Hiayan Guan, Hongjie He 0003, Jonathan Li 0001 |
IGARSS | 3 |
| 2024 | HigherNet-DST: Higher-Resolution Network With Dynamic Scale Training for Rooftop DelineationabstractHigh-definition (HD) maps of building rooftops or footprints are important for urban application and disaster management. Rapid creation of such HD maps through rooftop delineation at the city scale using high-resolution satellite and aerial images with deep learning methods has become feasible and drawn much attention. However, the scale variance issue in rooftop delineation limited the overall performance. Existing methods exhibit considerably poor performance in rooftop delineation of small buildings. In this paper, we propose a new method, namely the Higher Resolution Network with Dynamic Scale Training (HigherNet-DST) to overcome the scale variance problem in rooftop delineation. Specifically, the DST is applied in the model training phase to reduce the negative impact of scale variance. Then, a scale-aware backbone, namely the Higher Resolution Network, is adopted to enhance the feature representation. Finally, the high-resolution supervision targets are used to further boost the delineation performance. Our method was tested on four publicly accessible building datasets and the results demonstrated that our method achieved the highest performance in rooftop delineation among the existing methods. Extensive experiments showed the superior performance of our method with an AP of 68.5% on the AICrowd Building Dataset and an IoU of 82.6% On the Inria Building Dataset, respectively, which surpassed many state-of-the-art (SOTA) methods. On the WHU Building Dataset and the Waterloo Building Dataset, our method also achieved the highest performance among the benchmarked methods, showing the high performance of our method for building boundary delineation. Hongjie He 0003, Lingfei Ma, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Enhancing Spatial Resolution of Building Datasets Using Transformer-Based Single-Image Super-ResolutionabstractThe spatial resolution of Earth Observation (EO) images plays a key role in building footprint extraction. For the spatial resolution enhancement, deep learning-based image super-resolution methods have been widely used due to their remarkable performance. Transformer-based networks are effective and has drawn much attention in computer vision but underutilized in remote sensing, especially for super-resolving building datasets. Therefore, in this paper, we developed a novel transformer-based Single-Image Super-Resolution (SISR) method, named Pyramid Vision Transformer-Residual Feature Aggregation Network (PVT_RFANet), to improve the spatial resolution of building datasets. Specifically, the PVT v2 network was embedded into our Momentum Spatial-Channel Attention Residual Feature Aggregation Network (MSCA-RFANet). Moreover we conducted a comparative study to compare our method with Bicubic interpolation (BI), Super-resolution Convolutional Neural Network (SRCNN), Deep Recursive Residual Network (DRRN), SRResNet, and MSCA-RFANet. Using Peak Signal-Noise Ratio (PSNR) and Similarity Structure Index Measurement (SSIM) as the evaluation metrics, our method showed highest performance with the PSNR of 22.01 dB and the SSIM of 0.50 on the WHU Building Dataset, which demonstrated the superior performance of the proposed method. Yuwei Cai, Hongjie He 0003, Zhimeng He, Michael A. Chapman, Jing Li 0040, Lingfei Ma, Jonathan Li 0001 |
IGARSS | 2 |
| 2023 | Nighttime Light Missing Data Retrieval Using Modis Version 6 Satellite Data and Mask Dilated Partial Convolutional Neural NetworkabstractNighttime Lights (NTLs) remote sensing imagery contains tremendous information and has been shown to accurately predict a region’s human dynamics, economic health and energy consumption. Despite its usefulness, NTLs imagery is less widely available than other remote sensing data modalities. Several challenges appear when attempting to reconstruct NTLs data, either from other data modalities or existing NTLs data. These include complex non-linear relationships between NTLs and multispectral bands, non-matching spatial and temporal coverage, and different atmospheric and cloud conditions. This study attempts to create an out-of-the-box model that compensates for missing NTLs data using widely available daytime data in a broadly generalizable manner. The proposed project has two objectives: the construction of an image-to-image dataset mapping daytime multispectral images (MODIS V6 Land Surface Reflectance, MODIS V6 Land Cover, MODIS V6 Vegetation Indices) to NTLs images, and the reconstruction of NTLs data using deep learning techniques by researching, creating, and employing the state-of-the-art architecture of the Mask Partial Convolutional Neural Network in conjunction with dilated convolutions. The project will facilitate the training of new models for predicting missing NTLs and make NTLs data more accessible for future remote sensing research. Xuanchen Liu, Shuxin Qiao, Kyle Gao, Hongjie He 0003, Lingfei Ma, Jonathan Li 0001 |
IGARSS | 4 |
| 2023 | SAR-Optical Image Matching With Semantic Position Probability DistributionabstractWe propose a deep learning framework of Semantic Position Probability Distribution for SAR-optical image matching, termed as SPPD. Unlike the pixel-by-pixel searching matching method, a correspondence is directly obtained by an outputted matching position probability distribution. First, multiscale pyramidal features are created for each pixel in the SAR and optical images by using two weight-sharing ResNet-50 + Feature Pyramid Network (FPN) networks. The features containing high-level semantic information are then embedded into the proposed image Position Attention Module to obtain the spatial position dependencies between two images. Then, we present a loss function for semantic position matching to optimize the network from both semantic information and pixel alignment perspectives, converting the probability distribution of semantic matching positions into a point-to-point matching problem. In this paper, the SAR and optical images are set as the sensed and reference images. The effects of different image sizes, training label types, and loss function weights on matching accuracy are explored to obtain the optimal parameter settings for matching. The experimental results show that the proposed method is insensitive to image deformation and achieves cross-modal matching for SAR-optical images with high accuracy compared with the best matching method on different scene images, with several orders of magnitude faster inferences time. Liangzhi Li 0002, Kyle Gao, Hongjie He 0003, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Automated Detection of Oil/Gas Well Sites Detection from Multi-Source High Spatial Resolution ImagesabstractWith the development of oil/gas production, its adverse impact has drawn much attention. Therefore, automated detecting oil/gas well sites become important. Current research has been focused on detection using sole source RGB images, which was not the common case in remote sensing. In this study, we explored the use of a combination of the Residual Channel Attention Network (RCAN) and the state-of-the-art object detection method, You Only Look Once (YOLO) v4, to detect oil/gas well sites from multi-sensor images. To testify the feasibility of the combination, we selected 18 RapidEye and 15 WorldView images which cover the oil sands area in Alberta, Canada. We applied a pre-trained RCAN to unify the spatial resolution of different images to 2m/pixel. To maximize the feature space, we preserved 5 bands which were available in both images. YOLO v4 was applied on cropped image patches to detect oil/gas well sites. The experiment results showed that using the framework proposed in this study, oil/gas well sites can be localized accurately although the bounding boxes of the sites may not perfectly align with the objects. Hongjie He 0003, Hongzhang Xu, Michael A. Chapman, Yiping Chen 0002, Jonathan Li 0001 |
IGARSS | 1 |
| 2021 | Monitoring Surface Deformation Over Oilfield Using MT-Insar and Production Well DataabstractSurface displacements associated with the average subsidence due to hydrocarbon exploitation in southwest of Iran which has a long history in oil production, can lead to significant damages to surface and subsurface structures, and requires serious consideration. In this study, the Small BAseline Subset (SBAS) approach, which is a multitemporal Interferometric Synthetic Aperture Radar (InSAR) algorithm was employed to resolve ground deformation in the Marun region, Iran. A total of 22 interferograms were generated using 10 Envisat ASAR images. The mean velocity map obtained in the Line-Of-Sight (LOS) direction of satellite to the ground reveals the maximum subsidence on order of 13.5 mm per year over the field due to both tectonic and non-tectonic features. In order to assess the effect of non-tectonic features such as petroleum extraction on ground surface displacement, the results of InSAR have been compared with the oil production rate, which have shown a good agreement. Sarah Narges Fatholahi, Hongjie He 0003, Awase Syed, Jonathan Li 0001 |
IGARSS | 2 |
| 2021 | The Impact of Data Volume on Performance of Depp Learning Based Building Rooftop Extraction Using Very High Spatial Resolution Aerial ImagesabstractBuilding rooftop data are of importance in several urban applications and in natural disaster management. In contrast to traditional surveying and mapping, by using high spatial resolution aerial images, deep learning-based building rooftops extraction methods are efficient and accurate. Although more training data is preferred in deep learning-based tasks, the effect of data volume on building extraction models is underexplored. Therefore, the paper explores the impact of data volume on the performance of building rooftop extraction from very-high-spatial-resolution (VHSR) images using deep learning-based methods. To do so, we manually labelled 0.12m spatial resolution aerial images and perform a comparative analysis of models trained on datasets of different sizes using popular deep learning architectures for segmentation tasks, including Fully Convolutional Networks (FCN)-8s, U-Net and DeepLabv3+. The experiments showed that with more training data, algorithms converged faster and achieved higher accuracy, while better algorithms were able to better mitigate the lack of training data. Hongjie He 0003, Yuwei Cai, Zijian Jiang, Qiutong Yu, Sarah Narges Fatholahi, Yan Liu 0043, Hasti Andon Petrosians, Bingxu Hu, Liyuan Qing, Zhehan Zhang, Hongzhang Xu, Kyle Gao, Linlin Xu, Jonathan Li 0001 |
IGARSS | 1 |