Wenmiao Hu

dblp:305/5239 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-5902-2306ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Traffic Sign Localization and Orientation Classification for Automated Map Updating
abstract
High Definition (HD) maps, containing detailed road information, are essential for autonomous driving and many geo-related tasks. Recent developments in computer vision make it possible to automate the labor-intensive HD map maintenance work, such as localizing traffic signs within a road network. However, updating traffic signs to HD maps is non-trivial, as it not only requires precise geo-location but also requires confirming whether a sign belongs to a specific road. In our work, we develop an end-to-end automated traffic sign update system, termed AutoTS, which is capable of using an image sequence collected during vehicle operation to extract the geo-location of a traffic sign and determine whether it belongs to the road driven on, from its orientation. In AutoTS, we design a noise and sparsity adaptive localization module, which can filter noisy location points and derive a geo-location from sparse location points. To identify the orientation of traffic signs, we devise a position-aware orientation classification module, which uses the ROI feature and the position-aware SIFT feature to explore the orientation characteristic and understand the road context. To facilitate the evaluation of the proposed method, we construct a traffic sign localization and orientation classification benchmark, KITTI-TS. Our AutoTS achieves an MAE of 2.38 meters in traffic sign localization, while the accuracy in orientation classification reaches 88.89%.
Xianjing Han, Wenmiao Hu, Xuemeng Song, Hannes Kruppa, See-Kiong Ng, Roger Zimmermann
IEEE Trans. Circuits Syst. Video Technol.2
2025 GAN-Assisted Road Segmentation from Satellite Imagery
abstract
Geo-information extraction from satellite imagery has become crucial to carry out large-scale ground surveys in a short amount of time. With the increasing number of commercial satellites launched into orbit in recent years, high-resolution RGB color remote sensing imagery has attracted a lot of attention. However, because of the high cost of image acquisition and even more complicated annotation procedures, there are limited high-resolution satellite datasets available. Compared to close-range imagery datasets, existing satellite datasets have a much lower number of images and cover only a few scenarios (cities, background environments, etc.). They may not be sufficient for training robust learning models that fit all environmental conditions or be representative enough for training regional models that optimize for local scenarios. Instead of collecting and annotating more data, using synthetic images could be another solution to boost the performance of a model. This article proposes a GAN-assisted training scheme for road segmentation from high-resolution RGB color satellite images, which includes three critical components: (a) synthetic training sample generation, (b) synthetic training sample selection, and (c) assisted training strategy. Apart from the GeoPalette and cSinGAN image generators introduced in our prior work, this article explains in detail how to generate new training pairs using OpenStreetMap (OSM) and introduces a new set of evaluation metrics for selecting synthetic training pairs from a pool of generated samples. We conduct extensive quantitative and qualitative experiments to compare different image generators and training strategies. Our experiments on the downstream road segmentation task show that (1) our proposed metrics are more aligned with the trained model performance compared to commonly used GAN evaluation metrics such as the Fréchet inception distance (FID); and (2) by using synthetic data with the best training strategy, the model performance, mean Intersection over Union (mean IoU), is improved from 60.92% to 64.44%, when 1,000 real training pairs are available for learning, which reaches a similar level of performance as a model that is standard-trained with 4,000 real images (64.59%), i.e., enabling a 4-fold reduction in real dataset size.
Wenmiao Hu, Yifang Yin, Ying Kiat Tan, An Tran, Hannes Kruppa, Roger Zimmermann
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Utilizing Very High-resolution Optical RGB Satellite Imagery in Geo-information Extraction for Fine-scale Map-making
abstract
Maps are the fundamental elements of any navigation and localization system. With the fast expansion of urban areas and the increasing complexity of modern cities, traditional mapping techniques cannot meet the need for frequent map updates with enriched map details. This work investigates how geo-referenced very high-resolution (VHR) RGB satellite imagery can be used to extract geo-information to support the creation of fine-scale maps and facilitate map updates, focusing on two sub-directions. First, to alleviate the data scarcity issue of large-scale geo-information extraction from satellite imagery-based datasets, (1) GAN-assisted road segmentation proposes a new assisted training scheme to improve the model performance when the training dataset is limited and (2) a context-enhanced satellite-imagery dataset is created for large-scale parking lot detection to improve the type diversity of target objects. Second, to support rich map attribute geo-information extraction and vision-based navigation using street-view imagery, new methods are proposed to improve location and orientation extraction of the street-view imagery via cross-view matching with satellite imagery.
Wenmiao Hu
ACM Multimedia1
2023 CrossMatch: Source-Free Domain Adaptive Semantic Segmentation via Cross-Modal Consistency Training
abstract
Source-free domain adaptive semantic segmentation has gained increasing attention recently. It eases the requirement of full access to the source domain by transferring knowledge only from a well-trained source model. However, reducing the uncertainty of the target pseudo labels becomes inevitably more challenging without the supervision of the labeled source data. In this work, we propose a novel asymmetric two-stream architecture that learns more robustly from noisy pseudo labels. Our approach simultaneously conducts dual-head pseudo label denoising and cross-modal consistency regularization. Towards the former, we introduce a multimodal auxiliary network during training (and discard it during inference), which effectively enhances the pseudo labels' correctness by leveraging the guidance from the depth information. Towards the latter, we enforce a new cross-modal pixel-wise consistency between the predictions of the two streams, encouraging our model to behave smoothly for both modality variance and image perturbations. It serves as an effective regularization to further reduce the impact of the inaccurate pseudo labels in source-free unsupervised domain adaptation. Experiments on GTA5 → Cityscapes and SYNTHIA → Cityscapes benchmarks demonstrate the superiority of our proposed method, obtaining the new state-of-the-art mIoU of 57.7% and 57.5%, respectively.
Yifang Yin, Wenmiao Hu, Zhenguang Liu, Guanfeng Wang, Shili Xiang, Roger Zimmermann
ICCV2
2023 PetalView: Fine-grained Location and Orientation Extraction of Street-view Images via Cross-view Local Search
abstract
Satellite-based street-view information extraction by cross-view matching refers to a task that extracts the location and orientation information of a given street-view image query by using one or multiple geo-referenced satellite images. Recent work has initiated a new research direction to find accurate information within a local area covered by one satellite image centered at a location prior (e.g., from GPS). It can be used as a standalone solution or complementary step following a large-scale search with multiple satellite candidates. However, these existing works require an accurate initial orientation (angle) prior (e.g., from IMU) and/or do not efficiently search through all possible poses. To allow efficient search and to give accurate prediction regardless of the existence or the accuracy of the angle prior, we present PetalView extractors with multi-scale search. The PetalView extractors give semantically meaningful features that are equivalent across two drastically different views, and the multi-scale search strategy efficiently inspects the satellite image from coarse to fine granularity to provide sub-meter and sub-degree precision extraction. Moreover, when an angle prior is given, we propose a learnable prior angle mixer to utilize this information. Our method obtains the best performance on the VIGOR dataset and successfully improves the performance on KITTI dataset test~1 set with the recall within 1 meter (r@1m) for location estimation to 68.88% and recall within 1 degree (r@1d) 21.10% when no angle prior is available, and with angle prior achieves stable estimations at r@1m and r@1d above 70% and 21%, up to a 40-degree noise level.
Wenmiao Hu, Yichen Zhang 0002, Yuxuan Liang 0002, Xianjing Han, Yifang Yin, Hannes Kruppa, See-Kiong Ng, Roger Zimmermann
ACM Multimedia1
2022 Beyond Geo-localization: Fine-grained Orientation of Street-view Images by Cross-view Matching with Satellite Imagery
abstract
Street-view imagery provides us with novel experiences to explore different places remotely. Carefully calibrated street-view images (e.g., Google Street View) can be used for different downstream tasks, e.g., navigation, map features extraction. As personal high-quality cameras have become much more affordable and portable, an enormous amount of crowdsourced street-view images are uploaded to the internet, but commonly with missing or noisy sensor information. To prepare this hidden treasure for "ready-to-use" status, determining missing location information and camera orientation angles are two equally important tasks. Recent methods have achieved high performance on geo-localization of street-view images by cross-view matching with a pool of geo-referenced satellite imagery. However, most of the existing works focus more on geo-localization than estimating the image orientation. In this work, we re-state the importance of finding fine-grained orientation for street-view images, formally define the problem and provide a set of evaluation metrics to assess the quality of the orientation estimation. We propose two methods to improve the granularity of the orientation estimation, achieving 82.4% and 72.3% accuracy for images with estimated angle errors below 2 degrees for CVUSA and CVACT datasets, corresponding to 34.9% and 28.2% absolute improvement compared to previous works. Integrating fine-grained orientation estimation in training also improves the performance on geo-localization, giving top 1 recall 95.5%/85.5% and 86.8%/80.4% for orientation known/unknown tests on the two datasets.
Wenmiao Hu, Yichen Zhang 0002, Yuxuan Liang 0002, Yifang Yin, Andrei Georgescu, An Tran, Hannes Kruppa, See-Kiong Ng, Roger Zimmermann
ACM Multimedia1
2022 A Context-enriched Satellite Imagery Dataset and an Approach for Parking Lot Detection
abstract
Automatic detection of geoinformation from satellite images has been a fundamental yet challenging problem, which aims to reduce the manual effort of human annotators in maintaining an up-to-date digital map. There are currently several high-resolution satellite imagery datasets that are publicly available. However, the associated ground-truth annotations are limited to road, building, and land use, while the annotations of other geographic objects or attributes are mostly not available. To bridge the gap, we present Grab-Pklot, the first high-resolution and context-enriched satellite imagery dataset for parking lot detection. Our dataset consists of 1344 satellite images with the ground-truth annotations of carparks in Singapore. Motivated by the observation that carparks are mostly co-appear with other geographic objects, we associate each satellite image in our dataset with the surrounding contextual information of road and building, given in the format of multi-channel images. As a side contribution, we present a fusion-based segmentation approach to demonstrate that the parking lot detection accuracy can be improved by modeling the correlations between parking lots and other geographic objects. Experiments on our dataset provide baseline results as well as new insights into the challenges and opportunities in parking lot detection from satellite images.
Yifang Yin, Wenmiao Hu, An Tran, Hannes Kruppa, Roger Zimmermann, See-Kiong Ng
WACV2
2021 GeoPalette: Road Segmentation with Limited Satellite Imagery
abstract
In recent years, Geo-information extraction from high-resolution satellite imagery has attracted a lot of attention. However, because of the high cost of image acquisition and annotation, there are limited datasets available. Compared to close-range imagery datasets, existing satellite datasets have a much lower number of images and cover only a few scenarios (cities, background environments, etc.). They may not be sufficient for training robust learning models that fit all environmental conditions or be representative enough for training regional models that optimize for local scenarios. In this study, we propose GeoPalette, a Generative Adversarial Network (GAN) based tool to generate additional synthetic training samples for boosting model performance when the training dataset is limited. Our experiments on road segmentation show that using additional synthetic data can improves the model performance mean Intersection over Union (mIoU) from 60.92% to 64.44%, when 1,000 real training pairs are available for learning, which reaches a similar level of performance as a model is standard-trained on 4,000 real pairs (64.59%), i.e., a 4-fold reduction in real dataset size.
Wenmiao Hu, Yifang Yin, Ying Kiat Tan, An Tran, Hannes Kruppa, Roger Zimmermann
SIGSPATIAL/GIS1
2021 Multimodal Fusion of Satellite Images and Crowdsourced GPS Traces for Robust Road Attribute Detection
abstract
Automatic inference of missing road attributes (e.g., road type and speed limit) for enriching digital maps has attracted significant research attention in recent years. A number of machine learning based approaches have been proposed to detect road attributes from GPS traces, dash-cam videos, or satellite images. However, existing solutions mostly focus on a single modality without modeling the correlations among multiple data sources. To bridge the gap, we present a multimodal road attribute detection method, which improves the robustness by performing pixel-level fusion of crowdsourced GPS traces and satellite images. A GPS trace is usually given by a sequence of location, bearing, and speed. To align it with satellite imagery in the spatial domain, we render GPS traces into a sequence of multi-channel images that simultaneously capture the global distribution of the GPS points, the local distribution of vehicles' moving directions and speeds, and their temporal changes over time, at each pixel. Unlike previous GPS based road feature extraction methods, our proposed GPS rendering does not require map matching in the data preprocessing step. Moreover, our multimodal solution addresses single-modal challenges such as occlusions in satellite images and data sparsity in GPS traces by learning the pixel-wise correspondences among different data sources. Extensive experiments have been conducted on two real-world datasets in Singapore and Jakarta. Compared with previous work, our method is able to improve the detection accuracy on road attributes by a large margin.
Yifang Yin, An Tran, Ying Zhang 0047, Wenmiao Hu, Guanfeng Wang, Jagannadan Varadarajan, Roger Zimmermann, See-Kiong Ng
SIGSPATIAL/GIS4