Wufan Zhao

dblp:285/7595 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-0265-3465ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Sat2City: 3D City Generation from a Single Satellite Image with Cascaded Latent Diffusion
abstract
Recent advancements in generative models have enabled 3D urban scene generation from satellite imagery, unlocking promising applications in gaming, digital twins, and beyond. However, most existing methods rely heavily on neural rendering techniques, which hinder their ability to produce detailed 3D structures on a broader scale, largely due to the inherent structural ambiguity derived from relatively limited 2D observations. To address this challenge, we propose Sat2City, a novel framework that synergizes the representational capacity of sparse voxel grids with latent diffusion models, tailored specifically for our novel 3D city dataset. Our approach is enabled by three key components: (1) A cascaded latent diffusion framework that progressively recovers 3D city structures from satellite imagery, (2) a Re-Hash operation at its Variational Autoencoder (VAE) bottleneck to compute multi-scale feature grids for stable appearance optimization and (3) an inverse sampling strategy enabling implicit supervision for smooth appearance transitioning.To overcome the challenge of collecting real-world city-scale 3D models with high-quality geometry and appearance, we introduce a dataset of synthesized large-scale 3D cities paired with satellite-view height maps. Validated on this dataset, our framework generates detailed 3D structures from a single satellite image, achieving superior fidelity compared to existing city generation models.
Tongyan Hua, Lutao Jiang, Ying-Cong Chen, Wufan Zhao
ICCV4
2025 Depth2Elevation: Scale Modulation With Depth Anything Model for Single-View Remote Sensing Image Height Estimation
abstract
Accurate terrain elevation estimation from remote sensing data is essential for a multitude of geographic applications. Specifically, image-based elevation estimation has garnered growing attention due to advancements in optical sensor development and automated analysis algorithms, such as machine learning. In this context, deep learning methods, including Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), have recently enhanced the feature extraction ability and estimation accuracy of this task. Despite the distinct advantages afforded by each architectural paradigm, current methods are frequently impeded in their ability to discern subtle height variations within complex scenes and are ill-equipped to effectively tackle the extraction of features across both large and small scales. Although vision foundation models have shown significant advances in remote sensing analysis, their effectiveness for height estimation remains unexplored. In this study, we introduce the foundation model in the field of elevation estimation and propose a novel Depth to Elevation (Depth2Elevation) model, marking the first application of the Depth Anything Model (DAM) to height estimation in remote sensing images. First, we introduce the scale modulator for modulating partial encoders in the original DAM, which enables DAM to capture subtle representations of localized objects at different scales. Secondly, we further enhance the model’s representational capability by using a resolution-agnostic decoder architecture, which enables DAM to learn features at different spatial scales efficiently. We conducted comprehensive experiments on several benchmark datasets. Compared to strong baselines, our method achieves an average relative improvement of at most 42% on the latest large-scale benchmark dataset GAMUS and shows the best generalization ability across different scenarios.
Zhongcheng Hong, Wufan Zhao
IEEE Trans. Geosci. Remote. Sens.4
2023 Vectorizing Planar Roof Structure From Very High Resolution Remote Sensing Images Using Transformers
abstract
Grasping the roof structure of a building is a key part of building reconstruction. Directly predicting the geometric structure of the roof from a raster image to a vectorized representation, however, remains challenging. This paper introduces an efficient and accurate parsing method based upon a vision Transformer we dubbed Roof-Former. Our method consists of three steps: 1) Image encoder and edge node initialization, 2) Image feature fusion with an enhanced segmentation refinement branch, and 3) Edge filtering and structural reasoning. The vertex and edge heat map F1-scores have increased by 2.0% and 1.9% on the VWB dataset when compared to HEAT. Additionally, qualitative evaluations suggest that our method is superior to the current state-of-the-art. It indicates effectiveness for extracting global image information and maintaining the consistency and topological validity of the roof structure.
Wufan Zhao, Claudio Persello, Xianwei Lv 0002, Alfred Stein
IGARSS1
2022 Multiscale Context Aggregation Network for Building Change Detection Using High Resolution Remote Sensing Images
abstract
The existing methods of building change detection (CD) using remote sensing (RS) images are still deficient in handling scale variation and class imbalance problems, indicating a decrease in the robustness of small-object detection and pseudo-change information. Thus, a novel building CD framework called the multiscale context aggregation network (MSCANet) is proposed. The high-resolution network is integrated into the feature extracting stage to maintain high-resolution representations throughout the whole process. Then, multiscale context information is aggregated using a scale-aware feature pyramid module (FPM). Recognition performance can be improved from discriminant feature representation learning by using a channel–spatial attention module. Furthermore, a class-balanced loss is proposed to reduce the impact of class imbalance in long-tail datasets. Experimental results from using the LEVIR-CD and SZTAKI AirChange benchmark datasets prove the superiority of the MSCANet over the other baseline methods, with improved maximum F1 scores of 5.28 and 8.47, respectively.
Wufan Zhao
IEEE Geosci. Remote. Sens. Lett.2
2022 Rotation-Aware Building Instance Segmentation From High-Resolution Remote Sensing Images
abstract
While extracting buildings from high-resolution remote sensing imagery has been widely conducted in automatic surveying and mapping, challenges remain in building extraction over complex scenes, particularly for densely rotated objects with fuzzy boundaries. This study proposes a rotation-aware building instance segmentation network (RotSegNet) that integrates a refined rotated detector to extract rotation equivariant and invariant features. A boundary refinement module is added to the segmentation network to extract fine-grained boundary features. We evaluated our method using the WHU building dataset and AFCities dataset. Our RotSegNet generated a minimum of 2.5% and 0.8% mAP on the two datasets compared with other state-of-the-art methods, which shows the superiority of our method. Results also show that the proposed method can produce regularized buildings with high geometric accuracy.
Wufan Zhao, Jiaming Na, Mengmeng Li 0002, Hu Ding 0002
IEEE Geosci. Remote. Sens. Lett.1
2021 End-to-End Roofline Extraction from Very-High-Resolution Remote Sensing Images
abstract
Roof shape information is essential for creating 3D building models. However, the automated extracting of roof structures from Earth observation data is a difficult task involving significant uncertainties caused by scene complexity and limited multi-source data coverage. This paper introduces the integrally-attracted wireframe parsing (IAWP) framework to reconstruct building rooflines as a planar graph from remotely sensed images with a single forward pass. We add global geometric line priors through the Hough transform into deep networks to better extract the linear geometric features. We perform experiments on the vectorizing world building (VWB) dataset. The investigated method improves the F-score metrics of corner points/edges by 0.1%/7.7% and 0.6%/1.1%, respectively. Visual comparison results also indicate that the HT-IHT block gives consistent improvements in terms of geometric regularity.
Wufan Zhao, Claudio Persello, Alfred Stein
IGARSS1
2020 Building Instance Segmentation and Boundary Regularization from High-Resolution Remote Sensing Images
abstract
Building extraction from remote sensing images using convolutional neural networks (CNNs) has been an active research topic in recent years. Most results obtained by CNN-based algorithms, however, still have common issues with the precision of the delineation of building outlines and the separation of different buildings. Recently, efforts have been made towards the automation of building outline regularization. This paper employs a new instance segmentation framework named Hybrid Task Cascade (HTC) as baseline model, integrating detection and segmentation as a joint multi-stage processing. We further integrate regularization methods such as convex hull and Douglas-Peucker algorithm to obtain accurately segmented edges. The method is tested on the crowdAI benchmark dataset by comparing with alternative state-of-the-art models (i.e., Mask R-CNN). The results show that our method achieves better instance segmentation results and improves the results in terms of geometric regularity of building segments.
Wufan Zhao, Claudio Persello, Alfred Stein
IGARSS1