Jie Shan

dblp:33/4970 · DBLP profile ↗
← Back
31ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 10 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 A WebGIS-based digital twin platform for intelligent operation and maintenance of rail transit infrastructure
Wei Huang 0014, Junhua Xiao, Jie Shan, Mengbo Liu, Weian Guo, Yushan Zhu, Jing Zhang 0124
Expert Syst. Appl.4
2025 Hierarchical DRL-based Service Placement, UAV Placement, and Resource Allocation in MEC-enabled AGINs with Fairness Guarantee
abstract
This paper addresses hierarchical service placement, UAV trajectory design, access control, and resource scheduling issues in multi-access edge computing (MEC)-enabled time division multiple access (TDMA)-based air-ground integrated networks (AGINs) comprising multiple unmanned aerial vehicles (UAVs) and a ground base station. UAVs perform environmental sensing and data collection, but their limited onboard computing and energy resources pose challenges for handling heterogeneous, delay-sensitive tasks. To tackle this, we propose a two-tier actor-based deep reinforcement learning (DRL) algorithm for multi-timescale decision making. A high-level deep Q-Network (DQN) agent determines frame-level service placement, while a low-level improved deep deterministic policy gradient (IDDPG) agent manages time-slot-level UAV trajectory, access control, bandwidth allocation, and transmission power. A constraint-aware reward mechanism ensures learning stability and feasibility. Simulations show that the proposed method outperforms baseline schemes in task success rate, energy efficiency, and scheduling fairness.
Jianbo Du, Zhixiang Deng, Jing Jiang 0026, Jie Shan, Defeng Ren
VTC2025-Fall5
2025 Learning-Based Spatial Interpolation for Sparse Space Altimetry Measurements
abstract
Generating continuous models from sparse measurements remains a key challenge in remote sensing due to the limited availability and high cost of dense data collection. Traditional interpolation methods such as ordinary Kriging and natural neighbor interpolation often degrade significantly in accuracy under sparse measurements. To address this, we introduce T-GMSI, a Transformer-based Generative Model for Spatial Interpolation approach, which leverages the Vision Transformer (ViT) architecture for high-quality spatial interpolation from sparse inputs. We then apply T-GMSI to the task of digital elevation model (DEM) generation using sparse spaceborne laser altimetry data. Results show that T-GMSI maintains high accuracy even with over 70% data sparsity and generalizes well across diverse landscapes without requiring fine-tuning. Compared to baseline methods, T-GMSI reduces root mean square error (RMSE) by 40% and 25% over ordinary Kriging and natural neighbor interpolation, respectively, on airborne lidar datasets, and by 23% and 10% on spaceborne laser altimeter data. It also outperforms a state-of-the-art conditional generative adversarial network (CEDGAN), improving RMSE by 35% and 20% for airborne and spaceborne data, respectively. The study highlights the potential of learning-based interpolation methods for improving Earth observation modeling with sparse and incomplete remote sensing data.
Xiangxi Tian, Jie Shan
IEEE Trans. Geosci. Remote. Sens.2
2024 Drone Image Pose Estimation With Simultaneous Road Correspondence
abstract
Estimation of image or camera pose, including three position and three rotation parameters, is a fundamental task in remote sensing. When global navigation satellite system (GNSS) signal is unavailable or interferes with accurate pose, estimation is still challenging since its initials can be rather rough and the search space is large. Conventional solutions to such problems rely on ground control points that are identifiable in the images; however, this is an unrealistic or inconvenient requirement for many applications. We report a novel approach that uses a 3-D road database to determine the image pose without known image-road correspondence. After the road database is projected to images using the current estimated pose, we calculate a set of spectral and geometric road attributes for these projections. Taking advantage of the geometry of an image triplet, road projections from two neighboring images are transferred to the third one. For each image triplet, the final estimated pose minimizes the differences among the attributes of the road projections, extracted image roads, and transferred roads. Tests using two drone image datasets demonstrate that this approach can obtain an accurate pose at an uncertainty of less than 2.2 m and 2.6° compared to the bundle adjustment (BA) results. The use of an image triplet considerably improves the robustness and accuracy of independent single-image pose estimation. The developed visual hull quality metric can intrinsically determine the reliability of the estimated pose. This work will benefit image-based positioning and mapping applications where images have no ground control points or GNSS signals are degraded.
Jie Shan
IEEE Trans. Geosci. Remote. Sens.2
2024 ICESat-2 Controlled Integration of GEDI and SRTM Data for Large-Scale Digital Elevation Model Generation
abstract
Recent developments in spaceborne laser altimetry have revolutionized the way DEMs can be created with unprecedented productivity and accuracy. This paper aims to develop a holistic mathematic framework that is able to combine current multiple space-based terrain data sources for large-scale DEM generation. To achieve this goal, the developed framework can accommodate the heterogeneity of the involved multi-sensor data. Under the context model estimation or regression, the ICESat-2 ATL08 terrain dataset is treated as the observations for the target variable, while the predictor variables are derived from the GEDI Level 2A terrain data and the SRTM DEM. Three different models are then applied to determine their performance and identify the most accurate and robust one. Comparative evaluation for areas of over 11,000 square kilometers demonstrates that the support vector regression approach consistently yields superior and satisfactory results, surpassing the quality of current SRTM 30 m and 90 m DEMs. Using the 3DEP DEM as independent reference, the corrected SRTM DEMs exhibit a substantial reduction in the mean of the DEM error by one order of magnitude (~10 times), and a 27% to 36% significant improvement in RMSE. As for the corrected GEDI L2A terrain data, it achieves an exceptional accuracy of -0.4±1.4 m for plain suburban area and -0.6±9.3 m for mountain area. This study underscores the benefits of integrating multiple spaceborne altimeter data and the necessity of adopting holistic data integration models for such purpose.
Xiangxi Tian, Jie Shan
IEEE Trans. Geosci. Remote. Sens.2
2023 A Block-Wise Image-Based Mobile Localization Using Google Streetview
abstract
Accurate localization is a prerequisite for autonomous navigation and intelligent driving. However, this can be challenging due to urban canyons and no line-of-sight propagation of GNSS signals. In addition, remote areas with limited and unstable GNSS signals also exhibit difficulties in reliable and accurate positioning. Recent advances in foundation geospatial data and geotagged photographs, such as OpenStreetMap and Google Street View (GSV), provide a valuable reference for the localization of moving vehicles by using images taken aboard [1] .
Zhixin Li 0004, Jie Shan
IGARSS2
2023 Rigorous Oversampling Correction for Landsat 8 Lunar Observations
abstract
Motivation.Landsat 8 Operational Land Imager (OLI) has the capability to acquire Moon images in their orbits. These lunar images provide an additional option for on-orbit radiometric calibration using the Moon as a reference source. However, due to the slow scanning (attitude) rate of the satellite, the lunar images are usually elongated in the along-track direction making the shape of the Moon an ellipse instead of a circle. It is essential to correct such oversampling effect to obtain accurate digital numbers (DNs) or irradiance and pixel counts of the Moon disk.Existing Work.In general, when using lunar images for radiometric calibration, oversampling was accounted for by simply dividing the total irradiance by a certain ratio or oversampling factor [1] - [3]. However, this oversampling factor varies with the observers (sensors) and their attitude maneuver, such as 4.57 for Terra [4] and 8 for Landsat 8 [5]. Lunar edge points could be detected through edge detection [3], derivatives of the distribution of brightness [6], or gray-level based segmentation. Ellipse fitting was then applied to the extracted Moon edge to determine its semi-major and semi-minor. The semi-major axis was considered the elongated Moon radius in the image and the semi-minor axis was the actual radius. The oversampling factor was the ratio of these two axes [7] - [8]. Moreover, the oversampling factor can also be calculated using the average pitch rate [4] or roll angle [9] of the satellite when scanning across the Moon. However, the oversampling effect is different for each frame or among push-broom image lines since the viewing geometry and the oversampling factor are also changing during the Moon observation. All reported work produces a single oversampling factor and uses it for the oversampling correction of all lines in the image. Therefore, rigorous oversampling correction is needed.Fig. 1.(a) Flowchart of the proposed rigorous oversampling correction approach, including two main parts, i.e., line of sight intersection calculation and weighted interpolation-based oversampling correction, (b) illustration of the footprints of corrected and oversampled pixels in one column, and the calculation of the interpolation weights at a pixel based on a triangular function.Method.We propose a rigorous oversampling correction method for lunar images collected by Landsat 8 OLI. As shown in Fig. 1 (a), the variation of oversampling effect for each image line (row) is considered and corrected to produce a lunar image that would ideally be collected with non-overlapped and non-gapped footprints on the lunar surface. Due to the oversampling effect, the information from a single ground instantaneous field of view (GIFOV) is captured by the same detector in multiple adjacent image rows whose footprints overlap with the current row. In addition to oversampling between the same detector in different image rows, there are offsets between even and odd detectors in one image. Fig. 1 (b) shows an example of the GIFOVs of pixels collected by the same detector in three adjacent rows in one column of the corrected image (the red, green, and blue squares), denoted as the target footprints, and the footprints of pixels in the oversampled image (the gray windows). The cross signs in Fig. 1 (b) represent the centers of the gray windows. In the original image, the information collected by each pixel footprint is proportional to the overlapped area with the target footprints. Hence, the pixel value in the corrected image can be derived from the weighted interpolation of pixels in the oversampled image, where the weights are determined by the overlapped areas of their footprints on the lunar surface using a triangular function. An illustration for the calculation of the n-th pixel in a column of the corrected image is provided in Fig. 1 (b). The calculation of overlapped areas between pixel footprints is achieved by calculating the intersections of the lines-of-sight vector [10] with the lunar surface. The imaging time for each pixel is determined and used for the interpolation of the attitude and ephemeris data. The LOS vectors are constructed given the sensor imaging parameters, i.e., the coefficients of two groups of the Legendre polynomials for the along-track and across-track direction. The LOS vectors in the sensor coordinate system are then converted to the Earth-centered inertial (ECI) coordinate system. Based on the relationships between the satellite positions, LOS vectors, and Moon positions at specific observation time, the LOS intersection on lunar surface can be determined by solving a system of quadratic equations.Findings.The lunar images we used for this study were collected by Landsat 8 OLI on Jan. 29, 2021. The OLI consists of 14 sensor chip assemblies (SCAs), each of which has 9 bands. They observed the Moon in two orbits, producing 15 multi-spectral lunar images in total (one of the SCAs observed the Moon twice). The oversampling correction was conducted on each SCA and each band. Edge detection and ellipse fitting were conducted for evaluation. Since the ideal correction should result in a nearly circular Moon, the flattening and azimuth of the corrected Moon disk are calculated for assessment. The corrected Moon images showed that our method can achieve a near-circular Moon disk with ideally non-overlapped and non-gapped pixel footprints for all SCAs and bands. For example, Fig. 2 provides the oversample correction results for the images of SCA 5. The flattening of the Moon disk is less than 0.02 except for SCA 13. As for the consistency, the results of 9 bands in the same SCA are similar while larger difference is found for different SCAs. Overall, our method provides a rigorous solution to oversampling correction of Landsat 8 OLI lunar images. Variations of oversampling with different viewing geometries are considered, resulting in corrected images with desirable non-overlapped and non-gapped pixel footprints and more accurate DNs.Fig. 2.An example of Landsat 8 lunar images from band 1-9 (in the order of their placement on the SCA along track) of SCA 5 before and after correction.
Jie Shan
IGARSS2
2023 Validation of the Icesat-2 Lake Water Level Product
abstract
The ICESat-2 ATL13 product provides measurements of inland water surface height or water level. However, it has been shown that such water level observations are subject to noticeable uncertainties and need to be carefully handled for reliable water level determination. This paper presents an approach to detect outliers in the ATL13 product so that accurate outcome can be achieved. Lake Huron and Lake Superior of the Great Lakes are selected as the study areas. The extracted water levels from ATL13 over a period of four years are validated by using field observations at the closest NOAA hydrological stations. This work demonstrates the critical need on outlier removal and the capability of the ATL13 data. A bias of 9-10 cm is found in the ATL13 product. Such an uncertainty is positively related to the frozen precipitation, but mostly independent from the laser beam intensity and data acquisition time.
Renfei Li, Xiangxi Tian, Jie Shan
IGARSS3
2023 Towards Global DEM Generation by Combining GEDI and Icesat-2 Data
abstract
Widely used in geoscience, global or large-scale digital elevation model (DEM) is an essential depiction of the 3-dimenional information of the bare earth [1] . Among others, the most popular DEMs include the Shuttle Radar Topography Mission (SRTM) DEM (1" for USA and 3" for global), ASTER GDEM (30m), and several other similar ones [1] . The most significant problem for these existing global DEMs is the temporal latency (more than 20-year-old for SRTM [2] , more than 10-year-old for ASTER [3] ). Furthermore, some of these foundation data lack consistencies due to the inclusion of multiple data sources collected over a long period of time [1] . The successful launches of the Global Ecosystem Dynamics Investigation (GEDI) mission in December 2018 [4] and the Ice, Cloud, and land Elevation Satellite-2 (ICESat-2) mission [5] in September 2018 provide current, complementary, and dense on-orbit global elevation with unprecedented accuracy and coverage in the history of space laser altimetry. However, many researchers have found that GEDI and ICESat-2 data have inconsistent quality, e.g., a root mean square error (RMSE) of 4.48 m for GEDI terrain height [6] and an uncertainty ranging from 0.2 m to 2 m for ICESat-2 ATL08 terrain height [7] . Beyond the abundant research focusing on quality assessment, there are a significant amount of work on wall-to-wall mapping by integrating GEDI and other Earth observations to overcome the spatial heterogeneity of spaceborne lidar data [8] - [11] . However, there are no recent studies to generate terrain height using GEDI since it is originally designed for forestry studies; and the capability and accuracy of its terrain measurements are less promising than canopy measurements [6] , [7] . Subsequently, few research was reported to combine the terrain measurements from GEDI and ICESat-2 to achieve an even denser coverage than using GEDI or ICESat-2 alone.
Xiangxi Tian, Jie Shan
IGARSS2
2023 Generative Building Feature Estimation From Satellite Images
abstract
Urban and environmental researchers seek to obtain building features (e.g., building shapes, counts, and areas) at large scales. However, blurriness, occlusions, and noise from prevailing satellite images severely hinder the performance of image segmentation, super-resolution, or deep-learning-based translation networks. In this article, we combine globally available satellite images and spatial geometric feature datasets to create a generative modeling framework that enables obtaining significantly improved accuracy in per-building feature estimation and the generation of visually plausible building footprints. Our approach is a novel design that compensates for the degradation present in satellite images by using a novel deep network setup that includes segmentation, generative modeling, and adversarial learning for instance-level building features. Our method has proven its robustness through large-scale prototypical experiments covering heterogeneous scenarios from dense urban to sparse rural. Results show better quality over advanced segmentation networks for urban and environmental planning, and show promise for future continental-scale urban applications.
Jie Shan, Daniel G. Aliaga
IEEE Trans. Geosci. Remote. Sens.2
2023 Detection of Signal and Ground Photons From ICESat-2 ATL03 Data
abstract
The Advanced Topographic Laser Altimeter System (ATLAS) laser altimeter aboard the Ice, Cloud, and Land Elevation Satellite (ICESat-2) can measure the elevation of the Earth’s surface with unprecedented spatial detail. However, the quality of the derived signal and ground photons depends on the signal-to-noise ratio and canopy coverage. Current algorithms underperform for data collected during daytime over mountain areas with dense canopy. We demonstrate a novel procedure for signal photon detection and subsequent ground photon detection from ICESat-2 ATL03 data. We first introduce a gravity-based density model to characterize the anisotropic properties of photon distribution. Through jointly using the photon densities from the weak–strong beam pair, we are able to find key photons that have high probability being signals. A directional regional growing approach then takes these key photons as seeds to label all remaining signal photons. Finally, we introduce a weighted iterative median filter (WIMF) algorithm to identify ground photons whose height is closest to the estimated ground surface. A total of 36 ATL03 beams of two entire counties in USA are used for test and evaluation. Compared to the ATL03 and ATL08 algorithms, our signal photon finding method is more robust to the variation of topography, canopy coverage, and data collection time. Remarkably, the mislabeling caused by the after-pulsing effect does not present in our detected signal photons. Comparing current ATL03 and ATL08 products, the detected ground photons from our method are more consistent with reference to the 3DEP DEM, especially for strong beam data collected during daytime in dense canopy, high relief areas.
Xiangxi Tian, Jie Shan
IEEE Trans. Geosci. Remote. Sens.2
2022 A parallel compact firefly algorithm for the control of variable pitch wind turbine
Jie Shan, Shu-Chuan Chu 0001, ShaoWei Weng, Jeng-Shyang Pan 0001, Shi-Jie Jiang, Shiguang Zheng
Eng. Appl. Artif. Intell.1
2021 PSA-Net: Deep learning-based physician style-aware segmentation network for postoperative prostate cancer clinical target volumes
Anjali Balagopal, Howard E. Morgan, Michael Dohopolski, Ramsey Timmerman, Jie Shan, Daniel F. Heitjan, Dan Nguyen, Raquibul Hannan, Aurelie Garant, Neil Desai, Steve B. Jiang
Artif. Intell. Medicine5
2021 Comprehensive Evaluation of the ICESat-2 ATL08 Terrain Product
abstract
Current spaceborne lidar Ice, Cloud, and Land Elevation Satellite (ICESat)-2 provides ATL08 product for global terrain height, whose quality properties are yet to be fully understood. This article performs a comprehensive evaluation on its quality by using 3-D elevation program (3DEP) digital elevation model (DEM) and hundreds of survey marks of two counties in the USA. The evaluation is carried out in terms of data specification, survey marks, land cover, season and time (day or night) of acquisition, incidence angle, and terrain slope. The ATL08 height errors are further modeled as a function of laser incidence angle and canopy coverage. It is found out the height from ATL08 product lies between the 3DEP DEM and true ground surface. The uncertainty of ATL08 height is 0.2 m for plain terrain, and 2 m for mountainous terrain where the majority of ATL08 segments are not useful for terrain extraction. The terrain height gets more underestimated by ATL08 products in regions with large terrain slope or incidence angle and more overestimated where the terrain is covered by dense canopy. Furthermore, seasonal variation of terrain height error can be as high as 0.8 m, while the impact of acquisition time is less than 0.3 m. We expect these findings to be informative for the utilization of ATL08 terrain product by worldwide users, and for the improvement of future ATL08 product.
Xiangxi Tian, Jie Shan
IEEE Trans. Geosci. Remote. Sens.2
2020 Semantic Labeling of ALS Point Cloud via Learning Voxel and Pixel Representations
abstract
Semantic labeling is a fundamental task that can provide useful semantics for many other 3-D processing tasks. To tackle the challenge of airborne laser scanning (ALS) point cloud classification, current state-of-the-art methods leverage the capabilities of deep learning. However, they are limited due to the weaknesses of the isolated use of individual representations of point clouds. To address this issue, this letter presents a novel network, VPNet, which ensembles voxel and pixel representation-based networks, to predict class probabilities for each light detection and ranging (LiDAR) point. A fully connected conditional random field-based global refinement is then performed over each point in the point cloud to produce a fine-grained classification result. On the ISPRS 3-D Semantic Labeling Contest, our solution sets a new state of the art by improving the highest average F1-score and the highest average per-class accuracy from 69.3% to 73.9%, and 69.0% to 74.9%, respectively. The overall accuracy of our approach is 84.0%.
Nannan Qin, Xiangyun Hu, Puzuo Wang, Jie Shan
IEEE Geosci. Remote. Sens. Lett.4
2019 Graph Attention Convolution for Point Cloud Semantic Segmentation
abstract
Standard convolution is inherently limited for semantic segmentation of point cloud due to its isotropy about features. It neglects the structure of an object, results in poor object delineation and small spurious regions in the segmentation result. This paper proposes a novel graph attention convolution (GAC), whose kernels can be dynamically carved into specific shapes to adapt to the structure of an object. Specifically, by assigning proper attentional weights to different neighboring points, GAC is designed to selectively focus on the most relevant part of them according to their dynamically learned features. The shape of the convolution kernel is then determined by the learned distribution of the attentional weights. Though simple, GAC can capture the structured features of point clouds for fine-grained segmentation and avoid feature contamination between objects. Theoretically, we provided a thorough analysis on the expressive capabilities of GAC to show how it can learn about the features of point clouds. Empirically, we evaluated the proposed GAC on challenging indoor and outdoor datasets and achieved the state-of-the-art results in both scenarios.
Lei Wang 0084, Yuchun Huang, Yaolin Hou, Shenman Zhang, Jie Shan
CVPR5
2019 Fast and Robust Matching for Multimodal Remote Sensing Image Registration
abstract
While image matching has been studied in remote sensing community for decades, matching multimodal data [e.g., optical, light detection and ranging (LiDAR), synthetic aperture radar (SAR), and map] remains a challenging problem because of significant nonlinear intensity differences between such data. To address this problem, we present a novel fast and robust template matching framework integrating local descriptors for multimodal images. First, a local descriptor [such as histogram of oriented gradient (HOG) and local self-similarity (LSS) or speeded-up robust feature (SURF)] is extracted at each pixel to form a pixelwise feature representation of an image. Then, we define a fast similarity measure based on the feature representation using the fast Fourier transform (FFT) in the frequency domain. A template matching strategy is employed to detect correspondences between images. In this procedure, we also propose a novel pixelwise feature representation using orientated gradients of images, which is named channel features of orientated gradients (CFOG). This novel feature is an extension of the pixelwise HOG descriptor with superior performance in image matching and computational efficiency. The major advantages of the proposed matching framework include: 1) structural similarity representation using the pixelwise feature description and 2) high computational efficiency due to the use of FFT. The proposed matching framework has been evaluated using many different types of multimodal images, and the results demonstrate its superior matching performance with respect to the state-of-the-art methods.
Yuanxin Ye, Lorenzo Bruzzone, Jie Shan, Francesca Bovolo, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.3
2017 Fast and robust structure-based multimodal geospatial image matching
abstract
This paper presents a fast and robust framework integrating local features for the matching of multimodal geospatial data (e.g., optical, LiDAR, SAR and map). In the proposed framework, local feature descriptors, such as Histogram of Oriented Gradient (HOG) and Local Self Similarity (LSS), are first extracted for every pixel to form a pixel-wise structural feature representation of an image. Then we define a similarity metric based on the feature representation in frequency domain using the 3 Dimensional Fast Fourier Transform (3DFFT) technique, followed by a template matching scheme to detect control points between multimodal data. The proposed framework is based on the hypothesis that structural similarity between images is preserved across different modalities. The major advantages of this framework include (1) structural similarity representation using pixel-wise feature description and (2) high computational efficiency due to the use of 3DFFT. Experimental results on different types of multimodal geospatial data show more accurate matching performance of the proposed framework than the state-of-the-art methods.
Yuanxin Ye, Lorenzo Bruzzone, Jie Shan, Li Shen 0004
IGARSS3
2017 Robust Registration of Multimodal Remote Sensing Images Based on Structural Similarity
abstract
Automatic registration of multimodal remote sensing data [e.g., optical, light detection and ranging (LiDAR), and synthetic aperture radar (SAR)] is a challenging task due to the significant nonlinear radiometric differences between these data. To address this problem, this paper proposes a novel feature descriptor named the histogram of orientated phase congruency (HOPC), which is based on the structural properties of images. Furthermore, a similarity metric named HOPCncc is defined, which uses the normalized correlation coefficient (NCC) of the HOPC descriptors for multimodal registration. In the definition of the proposed similarity metric, we first extend the phase congruency model to generate its orientation representation and use the extended model to build HOPCncc. Then, a fast template matching scheme for this metric is designed to detect the control points between images. The proposed HOPCncc aims to capture the structural similarity between images and has been tested with a variety of optical, LiDAR, SAR, and map data. The results show that HOPCncc is robust against complex nonlinear radiometric differences and outperforms the state-of-the-art similarities metrics (i.e., NCC and mutual information) in matching performance. Moreover, a robust registration method is also proposed in this paper based on HOPCncc, which is evaluated using six pairs of multimodal remote sensing images. The experimental results demonstrate the effectiveness of the proposed method for multimodal image registration.
Yuanxin Ye, Jie Shan, Lorenzo Bruzzone, Li Shen 0004
IEEE Trans. Geosci. Remote. Sens.2
2016 Decomposing LiDAR waveforms with nonparametric classification methods
abstract
Waveform decomposition is an important step in full-waveform LiDAR remote sensing. Under the Gaussian Mixture Model, the conventional parametric classification algorithm of Expectation-Maximization (EM) is among the most widely applied ones to decompose the waveforms. This paper introduces nonparametric classification methods, such as K-means and mean-shift to decompose the LiDAR waveforms. The experiments demonstrate that a properly selected nonparametric method can model the asymmetry of a waveform, which is ignored in the conventional parametric model based method. Furthermore, the skewness of the decomposed waveform is conspicuous to be utilized for separating bare ground and forest.
Serkan Ural, Jie Shan
IGARSS3
2016 A Fuzzy Mean-Shift Approach to Lidar Waveform Decomposition
abstract
Waveform decomposition is a common step for exploitation of full-waveform lidar data. Much effort has been focused on designing algorithms based on the assumption that the returned waveforms follow a Gaussian mixture model where each component is a Gaussian. However, many real examples show that the waveform components can be neither Gaussian nor symmetric even when the emitted signal is Gaussian or symmetric. This paper proposes a nonparametric mixture model to represent lidar waveforms without any constraints on the shape of the waveform components. A fuzzy mean-shift algorithm is then developed to decompose the waveforms. This approach has the following properties: 1) It does not assume that the waveforms follow any parametric or functional distributions; 2) the waveform decomposition is treated as a fuzzy data clustering problem and the number of components is determined during the time of decomposition; and 3) neither peak selection nor noise floor filtering prior to the decomposition is needed. Experiments are conducted on a dataset collected over a dense forest area where significant skewed waveforms are demonstrated. As the result of the waveform decomposition, a highly dense point cloud is generated, followed by a subsequent filtering step to create a fine digital elevation model. Compared with the conventional expectation-maximization method, the fuzzy mean-shift approach yielded practically comparable and similar results. However, it is about three times faster and tends to lead to slightly fewer artifacts in the resultant digital elevation model.
Serkan Ural, John Anderson 0002, Jie Shan
IEEE Trans. Geosci. Remote. Sens.4
2015 Automatic Recognition of Cloud Images by Using Visual Saliency Features
abstract
Automatic cloud detection from satellite imagery is a necessary preprocessing step in remote sensing. Given that humans can easily “see” clouds in an image because of salient region features, we adopt a visual attention technique in computer vision to automatically identify images with a significant cloud cover. The proposed method generates a rough cloud mask by using a top-down visual saliency model to qualitatively distinguish cloud images from noncloud images. First, an image is downsized for rapid processing. Some basic saliency maps of clouds are then generated by multilevel segmentation, the computation of cloud visual saliency features, and feature classification. Thereafter, we fuse the basic saliency maps by using a most-votes-win strategy to generate the cloud mask. With the cloud mask, a threshold is used to classify the images as cloud or noncloud images. A total of 200 RapidEye images are tested by using the algorithm. Of the cloud images, 92% are correctly identified. The average processing time is 1.8 s per image.
Xiangyun Hu, Jie Shan
IEEE Geosci. Remote. Sens. Lett.3
2014 Minimum description length constrained LiDAR waveform decomposition
abstract
Waveform decomposition is a necessary step for the exploitation of full waveform LiDAR data. Much effort has been focused on designing algorithms to decompose the waveform into a fixed number of components. However, the determination of the appropriate number of components in a waveform, though crucial, is rarely studied. This paper introduces an order identification method, Minimum Description Length (MDL) to estimate the number of components. MDL requires the addition of a penalty term in the model fitness to account for the over-fitting of high order models. The convexity of MDL in terms of the number of components makes it possible to find out optimal model with 2-4 iterations in most cases. We applied the MDL-based estimation method to analyze a dataset collected by a Riegl Q680i LiDAR system. The procedure is demonstrated in this paper.
Serkan Ural, John Anderson 0002, Jie Shan
IGARSS4
2014 Road Centerline Extraction in Complex Urban Scenes From LiDAR Data Based on Multiple Features
abstract
Automatic extraction of roads from images of complex urban areas is a very difficult task due to the occlusions and shadows of contextual objects, and complicated road structures. As light detection and ranging (LiDAR) data explicitly contain direct 3-D information of the urban scene and are less affected by occlusions and shadows, they are a good data source for road detection. This paper proposes to use multiple features to detect road centerlines from the remaining ground points after filtering. The main idea of our method is to effectively detect smooth geometric primitives of potential road centerlines and to separate the connected nonroad features (parking lots and bare grounds) from the roads. The method consists of three major steps, i.e., spatial clustering based on multiple features using an adaptive mean shift to detect the center points of roads, stick tensor voting to enhance the salient linear features, and a weighted Hough transform to extract the arc primitives of the road centerlines. In short, we denote our method as Mean shift, Tensor voting, Hough transform (MTH). We evaluated the method using the Vaihingen and Toronto data sets from the International Society for Photogrammetry and Remote Sensing Test Project on Urban Classification and 3-D Building Reconstruction. The completeness of the extracted road network on the Vaihingen data and the Toronto data are 81.7% and 72.3%, respectively, and the correctness are 88.4% and 89.2%, respectively, yielding the best performance compared with template matching and phase-coded disk methods.
Xiangyun Hu, Jie Shan, Jianqing Zhang, Yongjun Zhang 0002
IEEE Trans. Geosci. Remote. Sens.3
2013 Local Edge Distributions for Detection of Salient Structure Textures and Objects
abstract
Automatic detection of regions of salient texture and objects is useful for analysis of remotely sensed imagery, such as for land cover classification, object detection, and change detection. Intuitively, the local edges on an image indicate spectral discontinuity and the existence of structure texture or objects. This letter explores a simple method for measuring the saliency of texture and objects based on the edge density and spatial evenness of the edge distribution in the local window of each pixel. This method generates a saliency map by computing the saliency index of each pixel. By segmenting the saliency map, the salient structure texture regions and the locations of objects can be extracted. The algorithm requires only the window size as the input parameter and is relatively simple to implement. Experiments using high-resolution images show its effectiveness and accuracy in the detection of salient structure texture regions, such as crops and residential areas, and man-made objects, such as airplanes, cars, etc.
Xiangyun Hu, Jiajie Shen, Jie Shan
IEEE Geosci. Remote. Sens. Lett.3
2013 Attraction-Repulsion Model-Based Subpixel Mapping of Multi-/Hyperspectral Imagery
abstract
This paper presents a new subpixel mapping method based on subpixel attraction-repulsion. The proposed method is formulated as an optimization problem with respect to attraction-repulsion among subpixels and is used to reconstruct a finer spatial resolution image from a lower resolution one. A comprehensive experiment is conducted to demonstrate the performance of the proposed method, by comparing it with the other three existing subpixel mapping methods, i.e., linear optimization, pixel swapping and spatial attraction model methods. In the experiment, both a synthetic image with known fractional abundances and an EO-1 Hyperion hyperspectral image of Shanghai were used to evaluate performances of the subpixel mapping methods. The experimental result shows that by using spatial dependence with attraction between the same types of ground objects and repulsion between different types of these objects, the proposed subpixel mapping method achieves a better performance on subpixel mapping than the other three methods.
Xiaohua Tong, Jie Shan, Huan Xie 0001, Miaolong Liu
IEEE Trans. Geosci. Remote. Sens.3
2011 Adaptive morphological filtering for DEM generation
abstract
This paper presents a novel approach to LiDAR filtering of ALS (Airborne Laser Scanning) data. The main effort is devoted to simplify and overcome the shortcomings of the existing morphological filtering algorithms. The proposed approach is based on the morphological erosion operation. The filtering is applied only to the points of discontinuity, which are identified from their residuals. The threshold for identifying the points of discontinuity is adaptively determined during the iteration. Our experiments show the proposed approach produces satisfying results in different terrain conditions with minimal change of parameters.
KyoHyouk Kim, Jie Shan
IGARSS2
2010 Segmentation and Reconstruction of Polyhedral Building Roofs From Aerial Lidar Point Clouds
abstract
This paper presents a solution framework for the segmentation and reconstruction of polyhedral building roofs from aerial LIght Detection And Ranging (lidar) point clouds. The eigenanalysis is first carried out for each roof point of a building within its Voronoi neighborhood. Such analysis not only yields the surface normal for each lidar point but also separates the lidar points into planar and nonplanar ones. In the second step, the surface normals of all planar points are clustered with the fuzzyk-means method. To optimize this clustering process, a potential-based approach is used to estimate the number of clusters, while considering both geometry and topology for the cluster similarity. The final step of segmentation separates the parallel and coplanar segments based on their distances and connectivity, respectively. Building reconstruction starts with forming an adjacency matrix that represents the connectivity of the segmented planar segments. A roof interior vertex is determined by intersecting all planar segments that meet at one point, whereas constraints in the form of vertical walls or boundary are applied to determine the vertices on the building outline. Finally, an extended boundary regularization approach is developed based on multiple parallel and perpendicular line pairs to achieve topologically consistent and geometrically correct building models. This paper describes the detail principles and implementation steps for the aforementioned solution framework. Results of a number of buildings with diverse roof complexities are presented and evaluated.
Aparajithan Sampath, Jie Shan
IEEE Trans. Geosci. Remote. Sens.2
2008 Fuzzy inference guided cellular automata urban-growth modelling using multi-temporal satellite images
abstract
This paper presents a fuzzy inference guided cellular automata approach. Semantic or linguistic knowledge on urban development is expressed as fuzzy rules, based on which fuzzy inference is applied to determine the urban development potential for each pixel. A defuzzification process converts the development potential to the required neighbourhood development level, which is taken by cellular automata as initial approximation for its transition rules. Such approximations are updated through spatial calibration over townships and temporal calibration with multi‐temporal satellite images. Assessment of the modelling results is based on three evaluation measures: fitness and Type I and Type II errors. The approach is applied to model the growth of the city of Indianapolis, Indiana over a period of 30 years from 1973 to 2003. A fitness level of 100 ±20% with 30% average errors can be achieved for 80% of the townships in urban‐growth prediction.
S. Al-kheder, Jun Wang 0139, Jie Shan
Int. J. Geogr. Inf. Sci.3
2005 Performance evaluation for pan-sharpening techniques
Qian Du 0001, Oguz Gungor, Jie Shan
IGARSS3
1998 Visualizing 3-D Geographical Data with VRML
abstract
The article discusses visualizing and interacting with 3D geographical data via VRML in the Web environment. For this purpose, the Web based desktop VR for geographical information is conceptualized. After VRML is briefly introduced, issues in modeling geographical data such as modeling geometry, topology and appearance are addressed. Sample visualization results for terrain and buildings are given, followed by discussions on the application of VRML. Cooperation within computer graphics and GIS is prospected in automatic model reconstruction and in exploring VRML application potential in visualizing geographical information.
Jie Shan
Computer Graphics International1