Muhammad Shahzad 0002

dblp:14/5514-2 · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0002-8278-9118ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 12 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Systems, architecture and hardware · 1Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MPTSNet: Integrating Multiscale Periodic Local Patterns and Global Dependencies for Multivariate Time Series Classification
abstract
Multivariate Time Series Classification (MTSC) is crucial in extensive practical applications, such as environmental monitoring, medical EEG analysis, and action recognition. Real-world time series datasets typically exhibit complex dynamics. To capture this complexity, RNN-based, CNN-based, Transformer-based, and hybrid models have been proposed. Unfortunately, current deep learning-based methods often neglect the simultaneous construction of local features and global dependencies at different time scales, lacking sufficient feature extraction capabilities to achieve satisfactory classification accuracy. To address these challenges, we propose a novel Multiscale Periodic Time Series Network (MPTSNet), which integrates multiscale local patterns and global correlations to fully exploit the inherent information in time series. Recognizing the multi-periodicity and complex variable correlations in time series, we use the Fourier transform to extract primary periods, enabling us to decompose data into multiscale periodic segments. Leveraging the inherent strengths of CNN and attention mechanism, we introduce the PeriodicBlock, which adaptively captures local patterns and global dependencies while offering enhanced interpretability through attention integration across different periodic scales. The experiments on UEA benchmark datasets demonstrate that the proposed MPTSNet outperforms 21 existing advanced baselines in the MTSC tasks.
Yang Mu, Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
AAAI2
2025 TransPose Re-ID: transformers for pose invariant person Re-identification
abstract
Person re-identification (Re-ID) is a computer vision task that involves recognizing and tracking individuals across multiple non-overlapping cameras or over time within the same camera view. It is particularly important in surveillance systems, where it can help in identifying potential threats or tracking suspects. Convolutional neural networks (CNNs) have been used to extract invariant person representation for this challenging task. However, CNNs do not consider global dependencies in their initial layers, causing some vital information to be lost during the convolution process. The development of vision-based transformers has opened up new research avenues for person re-identification. This work proposes a purely transformer-based solution, called TansPose Re-ID, that learns pose-invariant person representations. The proposed system uses a vision transformer baseline and enhances its architecture by introducing multiple streams to learn global and local dependencies as well as pose invariance in person images. The architecture includes a Global Self-Attention Module (GSM) and a Local Self-Attention Module (LSM) that jointly learn global and local patch-based person embeddings. The LSM is further improved by stochastically grouping local patches and aligning them. Additionally, an attention feature learning module (AFLM) is introduced in the LSM to handle pose and viewpoint variations. The proposed method is evaluated on two public Re-ID benchmarks, Market1501 and DukeMTMC-ReID, and demonstrates superior performance compared to existing transformer baselines.
Nazia Perwaiz, Muhammad Shahzad 0002, Muhammad Moazam Fraz
J. Exp. Theor. Artif. Intell.2
2025 TVFace: towards large-scale unsupervised face recognition in video streams
Atif Khurshid, Bostan Khan, Muhammad Shahzad 0002, Muhammad Moazam Fraz
Pattern Anal. Appl.3
2024 Smart surveillance with simultaneous person detection and re-identification
Nazia Perwaiz, Muhammad Moazam Fraz, Muhammad Shahzad 0002
Multim. Tools Appl.3
2024 Beyond local patches: Preserving global-local interactions by enhancing self-attention via 3D point cloud tokenization
abstract
Transformer-based architectures have recently shown impressive performance on various point cloud understanding tasks such as 3D object shape classification and semantic segmentation. Particularly, this can be attributed to their self-attention mechanism, which has the ability to capture long-range dependencies. However, current methods have constrained it to operate in local patches due to its quadratic memory constraints. This hinders their generalization ability and scaling capacity due to the loss of non-locality in early layers. To tackle this issue, we propose a window-based transformer architecture that captures long-range dependencies while aggregating information in the local patches. We do this by interacting each window with a set of global point cloud tokens — a representative subset of the entire scene — and augmenting the local geometry through a 3D Histogram of Oriented Gradients (HOG) descriptor. Through a series of experiments on segmentation and classification tasks, we show that our model exceeds the state-of-the-art on S3DIS semantic segmentation (+1.67% mIoU), ShapeNetPart part segmentation (+1.03% instance mIoU) and performs competitively on ScanObjectNN 3D object classification.1
Muhammad Shahzad 0002, Saqib Ali Khan, Muhammad Moazam Fraz, Xiao Xiang Zhu 0001
Pattern Recognit.2
2023 Time Series-based Active Labeling Framework for Curating a Multispectral Sentinel 2 Imagery Dataset for Crop Type Mapping
abstract
Acquiring ground-truth data for crop mapping is a challenging task in developing countries. Limited resources, inconsistent agricultural practices across the country, and inadequate infrastructure at the administrative level pose a significant challenge in collecting dataset information. This study proposes an active labelling framework for automated ground truth data generation in areas with limited or no ground truth information available. Sentinel 2 images and vegetative indices are visually interpreted for the identification of Wheat and Rice fields. This hand-labelled data is incorporated into the framework as training data to generate crop maps at 10m resolution using light weight, ConvLSTM model. Testing is performed for two spatially distinct districts with different field sizes and crop distribution in Pakistan. The results are consistent with an accuracy of 77.8%, F1 score of 87.3% and IOU of 70.7% for Gujranwala and 76%, 82% and 70% for Sargodha, promising the applicability of the proposed approach to generate large-scale crop labels, thereby enhancing the efficiency of crop mapping efforts in developing countries
Vaneeza Mehmood, Ramesha Murtaza, Zuhair Zafar, Muhammad Shahzad 0002, Karsten Berns, Muhammad Moazam Fraz
IGARSS4
2023 Towards a Benchmark EO Semantic Segmentation Dataset for Uncertainty Quantification
abstract
In order to achieve the objective of accurate and reliable use of deep neural networks for Earth Observation in large-scale scene understanding and interpretation, a large and diverse dataset with proper quantification of uncertainty is required. In this work, we exemplify the lack of a benchmark dataset and present the progress of a novel benchmark dataset for uncertainty quantification of deep learning models in the classic problem of building segmentation from overhead imagery. We present a synthetic dataset where synthetic UAV images were rendered from 3D mesh models of Berlin, Germany. The building masks were extracted from precise LoD-2 building models of the same area. We compare and contrast the performances of baseline methods for semantic segmentation and various uncertainty quantification techniques on this dataset. The experiments show that U-Net is the most accurate model with mIoU of 0.812. Moreover, the Bayesian model is found to be the most reliable uncertainty quantification method on our dataset, with the least ECE.
Dawood Wasif, Yuanyuan Wang 0002, Muhammad Shahzad 0002, Rudolph Triebel, Xiao Xiang Zhu 0001
IGARSS3
2023 Residual learning with annularly convolutional neural networks for classification and segmentation of 3D point clouds
Rabbia Hassan, Muhammad Moazam Fraz, A. Rajput, Muhammad Shahzad 0002
Neurocomputing4
2023 Ubiquitous vision of transformers for person re-identification
Nazia Perwaiz, Muhammad Shahzad 0002, Muhammad Moazam Fraz
Mach. Vis. Appl.2
2023 DCARN: Deep Context Aware Recurrent Neural Network for Semantic Segmentation of Large Scale Unstructured 3D Point Cloud
Saba Mehmood, Muhammad Shahzad 0002, Muhammad Moazam Fraz
Neural Process. Lett.2
2023 Person re-identification: A retrospective on domain specific open challenges and future trends
Asmat Zahra, Nazia Perwaiz, Muhammad Shahzad 0002, Muhammad Moazam Fraz
Pattern Recognit.3
2023 Per-former: rethinking person re-identification using transformer augmented with self-attention and contextual mapping
N. Pervaiz, Muhammad Moazam Fraz, Muhammad Shahzad 0002
Vis. Comput.3
2022 Mitigating Distribution Shift for Multi-Sensor Classification
abstract
Distribution shift may pose significant challenges in Earth observation, especially when dealing with significantly differ-ent sensors like multispectral optical and Synthetic Aperture Radar (SAR). Deep learning models trained for optical image classification generally do not generalize well for SAR images. This is due to very marked differences between them. Though there is a considerable amount of works on domain adaptation, only few deal with such strong differences. Towards this, we propose a co-teaching based domain adaptation method using dual classifier head, a Multi-layer Perceptron (MLP) classi-fier and a Graph Neural Network (GNN) classifier. The two classifier heads teach each other in an iterative manner, thus gradually adapting both of them for target classification. We experimentally demonstrate the efficacy of the proposed approach on Sentinel 2 (optical) as source and Sentinel 1 (SAR) images as target - both product of Copernicus program of European Space Agency.
Sudipan Saha, Shan Zhao 0007, Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
IGARSS3
2022 A multiapproach generalized framework for automated solution suggestion of support tickets
abstract
Nowadays, customer support systems are one of the key factors in maintaining any big company's reputation and success. These systems are capable of handling a large number of tickets systemically and provides a mechanism to track/logs the communication between customer and support agents. Companies invest huge amounts of money in training support agents and deploying customer care services for their products and services. Support agents are responsible for handling different customer queries and implementing required actions to solve a particular issue or problem raised by the service/product user. In a bigger picture, customer support systems could receive a large amount of ticket raised depending upon the number of users and services being offered. Customer care service gets directly affected due to the high volume of tickets and a limited number of support agents. Therefore, providing support agents with the recommendations about the possible resolution actions for a new ticket would be helpful and can save a lot of time. This study is focused on the development of an end-to-end framework for suggesting resolution actions rather than recommending free form resolution text against a newly raised ticket. To develop such a system, the pipeline is broadly divided into four components that are data preprocessing, actions extractor, resolution predictor, and evaluation. In actions extractor module, we have proposed a technique to identify and extract actionable phrases from resolution text. For resolution predictor, we have proposed two different pipelines that are referred as “Similarity Search Model” and “End-to-End Model.” The similarity search method is based on a ticket similarity search to find the most relevant historical tickets which then leads to corresponding resolution actions. On the other hand, end-to-end model make use of actions extractor module directly and implemented in a way to directly predict resolution actions. To compare and evaluate the mentioned methods on the same ground, we also proposed an actions evaluation criterion which uses BertScore and METEOR score jointly to compute the score against actual and predicted actions for a particular test ticket. The analysis and experiments are performed on the real-world IBM ticket data set. Overall, we observed that end-to-end model outperformed similarity search-based methods and achieved better performance and scores comparatively. The trained models and code are available at https://bit.ly/2GbUBVk.
Syed S. Ali Zaidi, Muhammad Moazam Fraz, Muhammad Shahzad 0002, Sharifullah Khan
Int. J. Intell. Syst.3
2022 Unsupervised Single-Scene Semantic Segmentation for Earth Observation
abstract
Earth observation data has huge potential to enrich our knowledge about our planet. An important step in many Earth observation tasks is semantic segmentation. Generally, a large number of pixelwise labeled images are required to train deep models for supervised semantic segmentation. On the contrary, strong inter-sensor and geographic variations impede the availability of annotated training data in Earth observation. In practice, most Earth observation tasks use only the target scene without assuming availability of any additional scene, labeled or unlabeled. Keeping in mind such constraints, we propose a semantic segmentation method that learns to segment from a single scene, without using any annotation. Earth observation scenes are generally larger than those encountered in typical computer vision datasets. Exploiting this, the proposed method samples smaller unlabeled patches from the scene. For each patch an alternate view is generated by simple transformations, e.g., addition of noise. Both views are then processed through a two-stream network and weights are iteratively refined using deep clustering, spatial consistency, and contrastive learning in the pixel space. The proposed model automatically segregates the major classes present in the scene and produces the segmentation map. Extensive experiments on four Earth observation datasets collected by different sensors show the effectiveness of the proposed method. Implementation is available at https://gitlab.lrz.de/ai4eo/cd/-/tree/main/unsupContrastiveSemanticSeg.
Sudipan Saha, Muhammad Shahzad 0002, Lichao Mou, Qian Song, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 Sustainable Vehicle-Assisted Edge Computing for Big Data Migration in Smart Cities
abstract
Smart cities are based on connected devices generating large quantities of data every instant. These data can be stored at a nearby edge location for initial processing but later sending the data to the backend data centers for storage and further analysis consumes considerable network bandwidth. In this article, we propose a large-scale data migration framework using vehicles. The framework uses a neural network to identify suitable vehicles as data mules, ones moving toward the data destination, potentially reducing the load from backend networks in terms of bandwidth usage and overall energy consumption. We compare the framework with data transfers using the traditional Internet and an approach without machine intelligence. The proposed framework performs well in terms of data loss, transfer time, energy, and CO2 emissions. From experiments, we demonstrate that the approach achieves a 67% success rate with data transfers $193\times $ faster than the average Internet bandwidth of 21.28 Mb/s. Moreover, the resulting CO2 emissions for 30-TB data transfers stood at 6.403 kg, which is significantly lower compared to 1172.8 kg for the Internet.
Maria Kanwal, Asad Waqar Malik, Anis Ur Rahman 0001, Imran Mahmood, Muhammad Shahzad 0002
IEEE Internet Things J.5
2019 VR-PROUD: Vehicle Re-identification using PROgressive Unsupervised Deep architecture
Raja Muhammad Saad Bashir, Muhammad Shahzad 0002, Muhammad Moazam Fraz
Pattern Recognit.2
2019 Buildings Detection in VHR SAR Images Using Fully Convolution Neural Networks
abstract
This paper addresses the highly challenging problem of automatically detecting man-made structures especially buildings in very high-resolution (VHR) synthetic aperture radar (SAR) images. In this context, this paper has two major contributions. First, it presents a novel and generic workflow that initially classifies the spaceborne SAR tomography (TomoSAR) point clouds-generated by processing VHR SAR image stacks using advanced interferometric techniques known as TomoSAR-into buildings and nonbuildings with the aid of auxiliary information (i.e., either using openly available 2-D building footprints or adopting an optical image classification scheme) and later back project the extracted building points onto the SAR imaging coordinates to produce automatic large-scale benchmark labeled (buildings/nonbuildings) SAR data sets. Second, these labeled data sets (i.e., building masks) have been utilized to construct and train the state-of-the-art deep fully convolution neural networks with an additional conditional random field represented as a recurrent neural network to detect building regions in a single VHR SAR image. Such a cascaded formation has been successfully employed in computer vision and remote sensing fields for optical image classification but, to our knowledge, has not been applied to SAR images. The results of the building detection are illustrated and validated over a TerraSAR-X VHR spotlight SAR image covering approximately 39 km2-almost the whole city of Berlin- with the mean pixel accuracies of around 93.84%.
Muhammad Shahzad 0002, Michael Maurer, Friedrich Fraundorfer, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.1
2018 Two Stream Deep CNN-RNN Attentive Pooling Architecture for Video-Based Person Re-identification
Wajeeha Ansar, Muhammad Moazam Fraz, Muhammad Shahzad 0002, Imad Gohar, Sajid Javed, Soon Ki Jung
CIARP3
2018 Extraction of Buildings in VHR SAR Images Using Fully Convolution Neural Networks
abstract
Modern spaceborne synthetic aperture radar (SAR) sensors, such as TerraSAR-X/TanDEM-X and COSMO-SkyMed, can deliver very high resolution (VHR) data beyond the inherent spatial scales (on the order of 1m) of buildings, constituting invaluable data source for large-scale urban mapping. Processing this VHR data with advanced interferometric techniques, such as SAR tomography (TomoSAR), enables the generation of 3-D (or even 4-D) TomoSAR point clouds from space. In this paper, we present a novel and generic workflow that exploits these TomoSAR point clouds in a way that is capable to automatically produce benchmark annotated (buildings/non-buildings) SAR datasets. These annotated datasets (building masks) have been utilized to construct and train the state-of-the-art deep Fully Convolution Neural Networks with an additional Conditional Random Field represented as a Recurrent Neural Network to detect building regions in a single VHR SAR image. The results of building detection are illustrated and validated over TerraSAR-X VHR spotlight SAR image covering approximately 39 km2- almost the whole city of Berlin - with mean pixel accuracies of around 93.84%.
Muhammad Shahzad 0002, Michael Maurer, Friedrich Fraundorfer, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS1
2018 End to End Person Re-Identification for Automated Visual Surveillance
abstract
Applications of Deep learning based methods are enormously growing in order to help the blind to see the world and/or enable the deaf to hear and speak. Re-Identification of a person among a set of cameras, is one of the latest challenges, in Computer Vision. Person Re-Identification deals with matching images of the same person over multiple non-overlapping camera views. Commonly, the task of Re-Id is broken down into three sub-modules, which are detection, tracking, and matching. Most of the techniques use manually annotated bounding boxes and only focus on matching between probes and cropped candidate images. This is not desirable in a real-time environment where the localization of object boundaries is not available. The target person needs to be identified from the complete image which may contain many distractors. To address the issue we investigated how the localization and matching of the target person can be done without using any prior annotation of bounding boxes. Our proposed method is based on an end to end deep learning technique, which not only targets matching but localization of objects as well. It handles detection and Re-Identification together. The research provides an end-to-end implementation of person tracking across multiple cameras in Surveillance context. The model is tested under diverse situations and resulted in a higher retrieval accuracy. The whole network is jointly optimized, using CIFM loss and fine-tuned to get better accuracy. The proposed approach outperforms state of the art methods on the PRW datasets, which demonstrates the effectiveness and generalization ability of our proposed approach.
Saadia Batool, Muhammad Zeeshan Ali, Muhammad Shahzad 0002, Muhammad Moazam Fraz
IPAS3
2016 Forest remote sensing on the individual tree level by airborne millimeterwave SAR
abstract
This paper presented experimental results discussing the potential of millimeterwave SAR for forest remote sensing on the individual tree level. As can be seen from the experimental results, although there is a certain amount of canopy penetration, a significant part of the signal response is received from the tree crowns. This provides both interesting perspectives for an analysis of forest volumes by continuous TomoSAR models as well as the reconstruction of individual tree models by utilization of discrete TomoSAR models.
Michael Schmitt 0003, Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
IGARSS2
2016 Automatic Detection and Reconstruction of 2-D/3-D Building Shapes From Spaceborne TomoSAR Point Clouds
abstract
Modern spaceborne synthetic aperture radar (SAR) sensors, such as TerraSAR-X/TanDEM-X and COSMO-SkyMed, can deliver very high resolution (VHR) data beyond the inherent spatial scales of buildings. Processing these VHR data with advanced interferometric techniques, such as SAR tomography (TomoSAR), allows for the generation of four-dimensional point clouds, containing not only the 3-D positions of the scatterer location but also the estimates of seasonal/temporal deformation on the scale of centimeters or even millimeters, making them very attractive for generating dynamic city models from space. Motivated by these chances, the authors have earlier proposed approaches that demonstrated first attempts toward reconstruction of building facades from this class of data. The approaches work well when high density of facade points exists, and the full shape of the building could be reconstructed if data are available from multiple views, e.g., from both ascending and descending orbits. However, there are cases when no or only few facade points are available. This usually happens for lower height buildings and renders the detection of facade points/regions very challenging. Moreover, problems related to the visibility of facades mainly facing toward the azimuth direction (i.e., facades orthogonally oriented to the flight direction) can also cause difficulties in deriving the complete structure of individual buildings. These problems motivated us to reconstruct full 2-D/3-D shapes of buildings via exploitation of roof points. In this paper, we present a novel and complete data-driven framework for the automatic (parametric) reconstruction of 2-D/3-D building shapes (or footprints) using unstructured TomoSAR point clouds particularly generated from one viewing angle only. The proposed approach is illustrated and validated by examples using TomoSAR point clouds generated using TerraSAR-X high-resolution spotlight data stacks acquired from ascending orbit covering two different test areas, with one containing simple moderate-sized buildings in Las Vegas, USA and the other containing relatively complex building structures in Berlin, Germany.
Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.1
2015 Robust Reconstruction of Building Facades for Large Areas Using Spaceborne TomoSAR Point Clouds
abstract
With data provided by modern meter-resolution synthetic aperture radar (SAR) sensors and advanced multipass interferometric techniques such as tomographic SAR inversion (TomoSAR), it is now possible to reconstruct the shape and monitor the undergoing motion of urban infrastructures on the scale of centimeters or even millimeters from space in very high level of details. The retrieval of rich information allows us to take a step further toward generation of 4-D (or even higher dimensional) dynamic city models, i.e., city models that can incorporate temporal (motion) behavior along with the 3-D information. Motivated by these opportunities, the authors proposed an approach that first attempts to reconstruct facades from this class of data. The approach works well for small areas containing only a couple of buildings. However, towards automatic reconstruction for the whole city area, a more robust and fully automatic approach is needed. In this paper, we present a complete extended approach for automatic (parametric) reconstruction of building facades from 4-D TomoSAR point cloud data and put particular focus on robust reconstruction of large areas. The proposed approach is illustrated and validated by examples using TomoSAR point clouds generated from a stack of TerraSAR-X high-resolution spotlight images from ascending orbit covering an approximately 2- km2high-rise area in the city of Las Vegas.
Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.1
2014 Automatic large area reconstruction of building façades from spaceborne TomoSAR point clouds
abstract
Improved resolution of SAR sensors and advanced multipass interferometric techniques, such as tomographic SAR inversion (TomoSAR), opens up new possibilities of 4D (or even higher dimensional) imaging that can be potentially used to reconstruct dynamic models of entire cities, i.e., city models that can incorporate temporal (motion) behaviour along with the 3D information. Motivated by these chances, this paper presents an approach that systematically allows automatic reconstruction of building façades from 4D point cloud generated from tomographic SAR processing and put particular focus on robust reconstruction of large areas. The approach is modular and is illustrated/validated by examples using TomoSAR point clouds generated from a stack of TerraSAR-X high-resolution spotlight images from ascending orbit covering approx. 2 km2high rise area in the city of Las Vegas.
Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
IGARSS1
2014 Facade Reconstruction Using Multiview Spaceborne TomoSAR Point Clouds
abstract
Recent advances in very high resolution tomographic synthetic aperture radar inversion (TomoSAR) using multiple data stacks from different viewing angles enables us to generate 4-D (space-time) point clouds of the illuminated area from space with a point density comparable to LiDAR. They can be potentially used for facade reconstruction and deformation monitoring in urban environment. In this paper, we present the first attempt to reconstruct facades from this class of data: First, the facade region is extracted using the density estimates of the points projected to the ground plane, the extracted facade points are then clustered into individual facades by means of orientation analysis, surface (flat or curved) model parameters of the segmented building facades are further estimated, and the geometric primitives such as intersection points of the adjacent facades are determined to complete the reconstruction process. The proposed approach is illustrated and validated by examples using TomoSAR point clouds generated from stacks of TerraSAR-X high-resolution spotlight images from two viewing angles, i.e., both ascending and descending orbits. The performance of the proposed approach is systematically analyzed. To explore the possible applications, we refine the elevation estimate of each raw TomoSAR point by using its more accurate azimuth and range coordinates and the corresponding reconstructed building facade model. Compared to the raw TomoSAR point clouds, significantly improved elevation positioning accuracy is achieved. Finally, a first example of the reconstructed 4-D city model is presented.
Xiao Xiang Zhu 0001, Muhammad Shahzad 0002
IEEE Trans. Geosci. Remote. Sens.2
2013 Reconstruction of building façades using spaceborne multiview TomoSAR point clouds
abstract
In this paper we present an approach that allows automatic reconstruction of building façades from 4D point cloud generated from tomographic SAR processing. The approach is modular and works by extracting façade points from the point density projected onto the ground plane. Individual façades are segmented using an unsupervised clustering procedure. Surface (flat or curved) model parameters of the segmented building façades are further estimated and finally the geometric primitives such as intersection points of the adjacent façades are determined to complete the reconstruction process. The proposed approach is illustrated and validated by examples using TomoSAR point clouds generated from TerraSAR-X high resolution spotlight images.
Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
IGARSS1
2012 Façade structure reconstruction using spaceborne TomoSAR point clouds
abstract
Very high resolution SAR tomography using multiple data stacks from different viewing angles enables us for the first time to generate 4D point clouds of the illuminated area from space with a point density comparable to LiDAR. They can be potentially used for façade reconstruction and monitoring in urban environment. In this paper, we propose an approach for façade detection and reconstruction from such point clouds. Firstly, the façade region is extracted by thresholding the point density on the ground plane. The extracted façades points are then clustered into segments corresponding to individual façades by means of slope analysis. Surface (flat or curved) model parameters of the segmented building façades are further estimated. Finally, the elevation estimates of each raw TomoSAR point is refined by using its more accurate azimuth and range coordinates, and the corresponding reconstructed surface model of the façade. The proposed approach is illustrated and validated by examples using TomoSAR point clouds generated from a stack of 25 TerraSAR-X high spotlight images.
Muhammad Shahzad 0002, Xiao Xiang Zhu 0001, Richard Bamler
IGARSS1