Zhipeng Gui

dblp:13/8410 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0001-9467-9680ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-authorSecurity and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Extraction of geoprocessing modeling knowledge from crowdsourced Google Earth Engine scripts by coordinating large and small language models
abstract
The widespread use of online geoinformation platforms, such as Google Earth Engine (GEE), has produced numerous scripts. Extracting domain knowledge from these crowdsourced scripts supports understanding of geoprocessing workflows. Small Language Models (SLMs) are effective for semantic embedding but struggle with complex code; Large Language Models (LLMs) can summarize scripts, yet lack consistent geoscience terminology to express knowledge. In this paper, we propose Geo-CLASS, a knowledge extraction framework for geospatial analysis scripts that coordinates large and small language models. Specifically, we designed domain-specific schemas and a schema-aware prompt strategy to guide LLMs to generate and associate entity descriptions, and employed SLMs to standardize the outputs by mapping these descriptions to a constructed geoscience knowledge base. Experiments on 237 GEE scripts, selected from 295,943 scripts in total, demonstrated that our framework outperformed LLM baselines, including Llama-3, GPT-3.5 and GPT-4o. In comparison, the proposed framework improved accuracy in recognizing entities and relations by up to 31.9% and 12.0%, respectively. Ablation studies and performance analysis further confirmed the effectiveness of key components and the robustness of the framework. Geo-CLASS has the potential to enable the construction of geoprocessing modeling knowledge graphs, facilitate domain-specific reasoning and advance script generation via Retrieval-Augmented Generation (RAG).
Zhipeng Gui, Jianyuan Liang, Dehua Peng, Wenzhang Wei, Shuyang Hou, Huayi Wu
Int. J. Geogr. Inf. Sci.2
2026 Topological and semantic contrastive graph clustering by Ricci curvature augmentation and hypergraph fusion
abstract
Contrastive graph clustering is an advanced technology in the field of cluster analysis. By leveraging graph neural networks and contrastive learning paradigm, it enables the coupling of topological structure and node semantic information for attributed graph networks. Graph augmentation and positive sample selection are two essentials of contrastive graph clustering. However, existing graph augmentation methods tend to disrupt the cluster structures, and most positive sample selectors suffer from the false negative sample problem. In this paper, we propose a Topological and Semantic Contrastive Graph Clustering (TSCGC) model consisting of three learning components. The representation learning component augments original graph using Ricci curvature to preserve the cluster structure, and introduces hypergraph view to capture high-order relationships. Graph and hypergraph convolutional networks are used to encode the triple-view embeddings. Meanwhile, we develop a dual contrastive learning component to extract the topological and semantic information. To reduce the number of false negatives, it utilizes K-means to generate pseudo cluster labels to guide the selection of positive samples. The self-supervised learning component is leveraged to align the three graph views. The final clustering results are obtained by performing K-means on the aligned embeddings. We demonstrated the effectiveness by comparing the performance of TSCGC with 13 clustering baselines on six real-world networks. Ablations verified the validity of key components and the impact of parameter settings were also analyzed. We further applied TSCGC to identify the function types of 10,370 buildings in ShenZhen City, China based on multi-source geospatial data. It achieved the highest accuracy and exhibit significant potential in handling complex network structures and high-dimensional node features. The code is available at: https://github.com/ZPGuiGroupWhu/TSCGC .
Dehua Peng, Guangyao Fang, Zhipeng Gui, Huayi Wu
Knowl. Based Syst.3
2026 Efficient Diffusion-Based 3D Human Pose Estimation With Hierarchical Temporal Pruning
abstract
Diffusion models have demonstrated strong capabilities in generating high-fidelity 3D human poses, yet their iterative nature and multi-hypothesis requirements incur substantial computational cost. In this paper, we propose an efficient diffusion-based 3D human pose estimation framework with a Hierarchical Temporal Pruning (HTP) strategy, which dynamically prunes redundant pose tokens across both frame and semantic levels while preserving critical motion dynamics. HTP operates in a staged, top-down manner: (1) Temporal Correlation-Enhanced Pruning (TCEP) identifies essential frames by analyzing inter-frame motion correlations through adaptive temporal graph construction; (2) Sparse-Focused Temporal MHSA (SFT MHSA) leverages the resulting frame-level sparsity to reduce attention computation, focusing on motion-relevant tokens; and (3) Mask-Guided Pose Token Pruner (MGPTP) performs fine-grained semantic pruning via clustering, retaining only the most informative pose tokens. Experiments on Human3.6M and MPI-INF-3DHP show that HTP reduces training MACs by 38.5%, inference MACs by 56.8%, and improves inference speed by an average of 81.1% compared to prior diffusion-based methods, while achieving state-of-the-art performance.
Yuquan Bi, Hongsong Wang 0001, Xinli Shi, Zhipeng Gui, Jie Gui, Yuan Yan Tang
IEEE Trans. Circuits Syst. Video Technol.4
2026 MeanCut: Greedy Graph Clustering by Fast Maximum Spanning Tree and Degree Descent Criterion
abstract
As the most typical graph clustering method, spectral clustering is popular and attractive due to its remarkable performance, easy implementation, and strong adaptability. Classical spectral clustering measures the edge weights using pairwise Euclidean similarity and resolves the optimal graph partitioning by relaxing the constraints of indicator matrix and decomposing the Laplacian matrix. However, Euclidean similarity might cause skew graph cuts when handling non-spherical clusters, and the relaxation strategy introduces information loss. Meanwhile, spectral clustering requires specifying the number of clusters, which is difficult to determine without enough prior knowledge. In this work, we propose a greedy-optimized scheme for resolving the indicator matrix using path-based similarity and degree descent criterion. It yields an indicator matrix with strictly binary entries without destructive relaxation and discretization steps. Path-based similarity can enhance the intra-cluster associations of arbitrary-shaped clusters, while degree descending is theoretically proven to be the best order to minimize our proposed objective function MeanCut. Moreover, we define a density gradient factor to separate clusters with fuzzy boundaries, and develop a fast maximum spanning tree algorithm to improve the scalability of similarity calculation. The effectiveness of MeanCut has been demonstrated on synthetic datasets and real-world benchmarks. By fusing multi-view image features, MeanCut outperforms cutting-edge subspace clustering methods in face recognition. The code is available at:https://github.com/ZPGuiGroupWhu/MeanCut.
Dehua Peng, Zhipeng Gui, Jie Gui, Huayi Wu
IEEE Trans. Fuzzy Syst.2
2026 Axial-View-Oriented Contrastive Adversarial Training for Robust Point Cloud Recognition
abstract
Contrastive adversarial training emerges as an effective approach to enhancing model robustness in safety-critical applications, particularly point cloud recognition for autonomous driving and medical imaging. However, existing point cloud adversarial training methods mainly emphasize global contrastive learning while overlooking local geometric variations induced by adversarial perturbations. Motivated by the spatial and intensity variations of perturbations across axial views, we propose AVOC, a novel local-global adversarial training framework that utilizes axial-view-oriented contrastive learning. This framework leverages the smallest axial view for local contrastive learning, as it exhibits the highest perturbation differences, and utilizes the largest axial view for global contrastive learning, as it preserves global structural consistency. We conduct comprehensive experiments across four representative architectures, demonstrating significant robustness improvements on widely-adopted recognition benchmarks, including ModelNet40, ShapeNetPart, ModelNet40-C, and ScanObjectNN-C, and further validate its effectiveness on the large-scale KITTI benchmark for 3D object detection. Our results across diverse perturbation scenarios, encompassing white-box attacks, black-box attacks, and natural perturbations, demonstrate the consistent and significant model robustness enhancement of our proposed method.
Jie Gui, Yu-Xin Zhang 0004, Xiaofeng Cong, Baosheng Yu, Zhipeng Gui, Yuan Yan Tang, James T. Kwok
IEEE Trans. Inf. Forensics Secur.5
2026 PANDA: Diffusion-Guided Purification and Adaptation for Robust Point Cloud Classification Against Adversarial Attack
Yu-Xin Zhang 0004, Xiaofeng Cong, Minjing Dong, Zhipeng Gui, Jie Gui, Yuan Yan Tang, James T. Kwok
IEEE Trans. Inf. Forensics Secur.4
2025 PEACE: Empowering Geologic Map Holistic Understanding with MLLMs
abstract
Geologic map, as a fundamental diagram in geology science, provides critical insights into the structure and composition of Earth’s subsurface and surface. These maps are indispensable in various fields, including disaster assessment, resource exploration, and civil engineering. Despite their significance, current Multimodal Large Language Models (MLLMs) often fall short in geologic map understanding. This gap is primarily due to the challenging nature of cartographic generalization, which involves handling high-resolution map, managing multiple associated components, and requiring domain-specific knowledge. To quantify this gap, we construct GeoMap-Bench, the first-ever benchmark for evaluating MLLMs in geologic map understanding, which assesses the full-scale abilities in extracting, referring, grounding, reasoning, and analyzing. To bridge this gap, we introduce GeoMap-Agent, the inaugural agent designed for geologic map understanding, which features three modules: Hierarchical Information Extraction (HIE), Domain Knowledge Injection (DKI), and Prompt-enhanced Question Answering (PEQA). Inspired by the interdisciplinary collaboration among human scientists, an AI expert group acts as consultants, utilizing a diverse tool pool to comprehensively analyze questions. Through comprehensive experiments, GeoMap-Agent achieves an overall score of 0.811 on GeoMap-Bench, significantly outperforming 0.369 of GPT-4o. Our work, emPowering gEologic mAp holistiC undErstanding (PEACE) with MLLMs, paves the way for advanced AI applications in geology, enhancing the efficiency and accuracy of geological investigations. The code and data are available at https://github.com/microsoft/PEACE.
Yangyu Huang, Qihao Zhao, Zhipeng Gui, Tengchao Lv, Lei Cui 0001, Scarlett Li, Furu Wei
CVPR6
2025 Striking a balance between diversity and regularity: a preference-guided transformer for individual mobility prediction
abstract
Human mobility modeling and prediction are central research topics in GIScience. Although deep learning has led to significant advances in these fields, existing trajectory prediction models still face challenges in capturing the complexity of individual mobility behavior. Regression-based models often overestimate the diversity of human mobility, whereas classification models tend to underestimate it. This study attributes these biases to the models’ limitations in recognizing the spatial relationships among activity locations and mobility heterogeneity across individuals. To address these challenges, we propose the Spatial Preference Map-based Transformer (SPM-Former), explicitly integrating spatial proximity and mobility heterogeneity to enhance trajectory sequence prediction. To capture individual mobility characteristics, SPM-Former utilizes the Spatial Preference Map (SPM) to represent individuals’ spatial visitation preferences and adjacency relationships between locations. Then, we introduce two encoding modules to decode the information hidden within the SPM: one for encoding trajectory-level spatial-temporal information and another for embedding individual-level overall mobility features. Furthermore, we propose a novel optimization method, SPM-Loss, to assess prediction accuracy from the global spatial distribution perspective. Experimental results on a large-scale dataset from Japan demonstrate that SPM-Former outperforms state-of-the-art classification-based models, achieving approximately 3% and 20% improvements in trajectory sequence similarity and overall spatial feature similarity, respectively.
Guangyue Li, Yang Xu 0002, Zhipeng Gui, Luliang Tang
Int. J. Geogr. Inf. Sci.3
2025 A Robust and Efficient Boundary Point Detection Method by Measuring Local Direction Dispersion
abstract
Boundary point detection aims to outline the external contour structure of clusters and enhance the inter-cluster discrimination, thus bolstering the performance of the downstream classification and clustering tasks. However, existing boundary point detectors are sensitive to density heterogeneity or cannot identify boundary points in concave structures and high-dimensional manifolds. In this work, we propose a robust and efficient boundary point detection method based on Local Direction Dispersion (LoDD). The core of boundary point detection lies in measuring the difference between boundary points and internal points. It is a common observation that an internal point is surrounded by its neighbors in all directions, while the neighbors of a boundary point tend to be distributed only in a certain directional range. By considering this observation, we adopt density-independent K-Nearest Neighbors (KNN) method to determine neighboring points and design a centrality metric LoDD using the eigenvalues of the covariance matrix to depict the distribution uniformity of KNN. We also develop a grid-structure assumption of data distribution to determine the parameters adaptively. The effectiveness of LoDD is demonstrated on synthetic datasets, real-world benchmarks, and application of training set split for deep learning model and hole detection on point cloud data. The datasets and toolkit are available at:https://github.com/ZPGuiGroupWhu/lodd.
Dehua Peng, Zhipeng Gui, Jie Gui, Huayi Wu
IEEE Trans. Circuits Syst. Video Technol.2
2025 Dynamic Visual Semantic Sub-Embeddings and Fast Re-Ranking for Image-Text Retrieval
abstract
The core of image-text retrieval is to accurately measure the similarity between different modalities in a unified representation space. However, compared to textual descriptions of a certain perspective, the visual modality has more semantic variations. Therefore, images are usually associated with multiple textual captions in databases. Although popular symmetric embedding methods have explored numerous modal interaction approaches, they often learn toward outputting the average representation of multiple semantic variations within image embeddings. Consequently, information entropy in embeddings is increased, resulting in redundancy and decreased accuracy. In this work, we propose a Dynamic Visual Semantic Sub-Embeddings framework (DVSE) to reduce the information entropy. Specifically, we obtain a set of heterogeneous visual sub-embeddings through dynamic orthogonal constraint loss. To encourage the generated candidate image embeddings to capture various semantic variations, we construct a mixed distribution and employ a variance-aware weighting loss to assign different weights to the optimization process. In addition, we develop a Fast Re-ranking strategy (FR) to efficiently evaluate the retrieval results and enhance the performance. We compare the performance with existing set-based method using five image feature encoders and three text feature encoders on three benchmark datasets: MSCOCO, Flickr30K and CUB Captions. We also show the role of different components by ablation studies and perform a sensitivity analysis of the hyperparameters. The qualitative analysis of visualized bidirectional retrieval and attention maps further demonstrates the ability of our method to encode semantic variations.
Wenzhang Wei, Zhipeng Gui, Changguang Wu, Dehua Peng, Huayi Wu
IEEE Trans. Multim.2
2024 Map retrieval intention recognition based on relevance feedback and geographic semantic guidance: For better understanding user retrieval demands
abstract
Effective retrieval is essential for finding resources in demand handily amidst extensive data records in data warehouse. Mainstream map retrieval methods suffer from intention gap problem and are incapable to describe sophisticated user demands precisely due to the limits of low- and middle-level text or visual feature matching, resulting in unsatisfactory retrieval results. Such limitations are more marked when map retrieval demands were characterized with joint constraints of geographic concepts. To address this issue, we propose a map retrieval intention recognition method to perceive user demands with relevance feedback samples and geographic semantics guidance. Specifically, we construct a hierarchical intention expression model to describe retrieval goals and their multi-dimensional attribute constrains; incorporate geographic ontologies to provide semantic guidance and facilitate recognition; utilize the frequent itemset mining (FIM) algorithm Apriori to generate intention candidates from relevance feedback samples, and search for the optimal intention set by adopting the minimum description length (MDL) principle. The experiments verify the effectiveness of Apriori algorithm and MDL principle on intention recognition. The proposed method outperforms the FIM algorithm Gene Ontology (RuleGO) and the Decision Tree algorithm with Hierarchical Features (DTHF) with higher recognition accuracy and noise tolerance. Furthermore, through our sample augmentation strategy, the method yields promising recognition accuracy even when the feedback sample size is as low as ten, substantially reducing the feedback burden in human-computer interactions. We envision that the application of our method in spatial data infrastructures (SDIs), such as geoportals and catalogue services, could enhance the quality of service and user experience in geospatial data discovery.
Zhipeng Gui, Xinjie Liu, Zhipeng Ling, Fa Li, Zelong Yang 0001, Huayi Wu, Shuangming Zhao
Inf. Process. Manag.1
2022 Enriching the metadata of map images: a deep learning approach with GIS-based data augmentation
abstract
Maps in the form of digital images are widely available in geoportals, Web pages, and other data sources. The metadata of map images, such as spatial extents and place names, are critical for their indexing and searching. However, many map images have either mismatched metadata or no metadata at all. Recent developments in deep learning offer new possibilities for enriching the metadata of map images via image-based information extraction. One major challenge of using deep learning models is that they often require large amounts of training data that have to be manually labeled. To address this challenge, this paper presents a deep learning approach with GIS-based data augmentation that can automatically generate labeled training map images from shapefiles using GIS operations. We utilize such an approach to enrich the metadata of map images by adding spatial extents and place names extracted from map images. We evaluate this GIS-based data augmentation approach by using it to train multiple deep learning models and testing them on two different datasets: a Web Map Service image dataset at the continental scale and an online map image dataset at the state scale. We then discuss the advantages and limitations of the proposed approach.
Yingjie Hu 0001, Zhipeng Gui, Jimin Wang, Muxian Li
Int. J. Geogr. Inf. Sci.2
2021 LSI-LSTM: An attention-aware LSTM for real-time driving destination prediction by considering location semantics and location importance of trajectory points
abstract
Individual driving final destination prediction supports location-based services such as personalized service recommendations, traffic navigation, and public transport dispatching. However, real-time destination prediction is challenging due to the complexity of temporal dependencies, and the strong influence of travel spatiotemporal semantics and spatial correlations. Besides temporal context, the nearby urban functionalities of traveling zones and departure regions, and the crucial positions on the road network where trajectory points located would reflect the travel intentions of drivers. However, these spatial factors are rarely considered in existing studies. To fill this gap, we propose a real-time individual driving destination prediction model LSI-LSTM based on an attention-aware Long Short-Term Memory (LSTM) by taking Location Semantics and Location Importance of trajectory points into account. More specifically, a trajectory location semantics extraction method (t-LSE) enriches feature description with prior knowledge for implicit travel intentions learning. t-LSE represents urban functionality through Points of Interest (POIs) using Term Frequency-Inverse Document Frequency (TF-IDF). Meanwhile, a novel trajectory spatial attention mechanism (t-SAM) captures the trajectory points that strongly correlate to candidate destinations based on the location importance inferred from the driving status, i.e., turning angle, driving speed, and traveled distance. Comparative experiments with three baseline methods, i.e., Hidden Markov Model, Random Forest, and LSTM, demonstrate significant prediction accuracy improvements of LSI-LSTM on four individual trajectory datasets. Further analyses validate the effectiveness of the proposed semantic extraction method and attention mechanism, and also discuss the factors that may affect the prediction results.
Zhipeng Gui, Yunzeng Sun, Dehua Peng, Fa Li, Huayi Wu, Chi Guo, Wenfei Guo, Jianya Gong
Neurocomputing1
2020 MSGC: Multi-scale grid clustering by fusing analytical granularity and visual cognition for detecting hierarchical spatial patterns
abstract
Spatial clustering is a widely used data mining method for discovery of spatial aggregation pattern. However, existing methods often neglect scale dependence, impeding the full recognition of point patterns and the detection of hierarchical spatial structures. Spatial clustering is scale dependent and linked to the size of analysis unit as well as the hierarchy of visual cognition. Therefore, this paper proposes a novel multi-scale grid clustering (MSGC) algorithm, which fuses dual scale factors, i.e., analytical scale and visual scale that sequentially integrates multi-analytical-scale clustering (MASC) and multi-visual-scale clustering (MVSC). MASC generates multi-granularity grids to transform the analytical scales, and MVSC extracts multi-level clusters to express the hierarchy of visual cognition. Comparative experiments validated the proposed algorithm against the classical Density-based Spatial Clustering of Applications with Noise (DBSCAN) and WaveCluster algorithms on both synthetic and real-world geographic datasets. The results demonstrate that MSGC can generate multi-scale clusters for increased understanding of the spatial aggregation patterns and hierarchical structures of geographic entities. Moreover, it can eliminate noise adaptively and effectively identify clusters with arbitrary shapes. Due to the nature of grid clustering, the low computational complexity enables near real-time visual analytics and efficient point pattern mining on large spatial datasets.
Zhipeng Gui, Dehua Peng, Huayi Wu
Future Gener. Comput. Syst.1
2020 Optimizing and accelerating space-time Ripley 's K function based on Apache Spark for distributed spatiotemporal point pattern analysis
abstract
With increasing point of interest (POI) datasets available with fine-grained spatial and temporal attributes, space–time Ripley’s K function has been regarded as a powerful approach to analyze spatiotemporal point process. However, space–time Ripley’s K function is computationally intensive for point-wise distance comparisons, edge correction and simulations for significance testing. Parallel computing technologies like OpenMP, MPI and CUDA have been leveraged to accelerate the K function, and related experiments have demonstrated the substantial acceleration. Nevertheless, previous works have not extended optimization of Ripley’s K function from space dimension to space–time dimension. Without sophisticated spatiotemporal query and partitioning mechanisms, extra computational overhead can be problematic. Meanwhile, these researches were limited by the restricted scalability and relative expensive programming cost of parallel frameworks and impeded their applications for large POI dataset and Ripley’s K function variations. This paper presents a distributed computing method to accelerate space–time Ripley’s K function upon state-of-the-art distributed computing framework Apache Spark, and four strategies are adopted to simplify calculation procedures and accelerate distributed computing respectively: (1) spatiotemporal index based on R-tree is utilized to retrieve potential spatiotemporally neighboring points with less distance comparison; (2) spatiotemporal edge correction weights are reused by 2-tier cache to reduce repetitive computation in L value estimation and simulations; (3) spatiotemporal partitioning using KDB-tree is adopted to decrease ghost buffer redundancy in partitions and support near-balanced distributed processing; (4) customized serialization with compact representations of spatiotemporal objects and indexes is developed to lower the cost of data transmission. Based on the optimized method, a web-based visual analytics framework prototype has been developed. Experiments prove the feasibility and time efficiency of the proposed method, and also demonstrate its value on promoting applications of space–time Ripley’s K function in ecology, geography, sociology, economics, urban transportation and other fields.
Zhipeng Gui, Huayi Wu, Dehua Peng, Jinghang Wu, Zousen Cui
Future Gener. Comput. Syst.2
2020 A hierarchical temporal attention-based LSTM encoder-decoder model for individual mobility prediction
Fa Li, Zhipeng Gui, Zhao-Yu Zhang 0003, Dehua Peng, Kunxiaojia Yuan, Yunzeng Sun, Huayi Wu, Jianya Gong, Yichen Lei
Neurocomputing2
2019 A quad-tree-based fast and adaptive Kernel Density Estimation algorithm for heat-map generation
abstract
Kernel Density Estimation (KDE) is a classic algorithm for analyzing the spatial distribution of point data, and widely applied in spatial humanities analysis. A heat-map permits intuitive visualization of spatial point patterns estimated by KDE without any overlapping. To achieve a suitable heat-map, KDE bandwidth parameter selection is critical. However, most generally applicable bandwidth selectors of KDE with relatively high accuracy encounter intensive computation issues that impede or limit the applications of KDE in big data era. To solve the complex computation problems, as well as make the bandwidths adaptively suitable for spatially heterogenous distributions, we propose a new Quad-tree-based Fast and Adaptive KDE (QFA-KDE) algorithm for heat-map generation. QFA-KDE captures the aggregation patterns of input point data through a quad-tree-based spatial segmentation function. Different bandwidths are adaptively calculated for locations in different grids calculated by the segmentation function; and density is estimated using the calculated adaptive bandwidths. In experiments, through comparisons with three mostly used KDE methods, we quantitatively evaluate the performance of the proposed method in terms of correctness, computation efficiency and visual effects. Experimental results demonstrate the power of the proposed method in computation efficiency and heat-map visual effects while guaranteeing a relatively high accuracy.
Kunxiaojia Yuan, Xiaoqiang Cheng, Zhipeng Gui, Fa Li, Huayi Wu
Int. J. Geogr. Inf. Sci.3
2014 Optimizing an index with spatiotemporal patterns to support GEOSS Clearinghouse
abstract
A variety of Earth observation systems monitor the Earth and provide petabytes of geospatial data to decision-makers and scientists on a daily basis. However, few studies utilize spatiotemporal patterns to optimize the management of the Big Data. This article reports a new indexing mechanism with spatiotemporal patterns integrated to support Big Earth Observation (EO) metadata indexing for global user access. Specifically, the predefined multiple indices mechanism (PMIM) categorizes heterogeneous user queries based on spatiotemporal patterns, and multiple indices are predefined for various user categories. A new indexing structure, the Access Possibility R-tree (APR-tree), is proposed to build an R-tree-based index using spatiotemporal query patterns. The proposed indexing mechanism was compared with the classic R*-tree index in a number of scenarios. The experimental result shows that the proposed indexing mechanism generally outperforms a regular R*-tree and supports better operation of Global Earth Observation System of Systems (GEOSS) Clearinghouse.
Jizhe Xia, Chaowei Phil Yang, Zhipeng Gui, Kai Liu 0017, Zhenglong Li 0002
Int. J. Geogr. Inf. Sci.3
2013 A performance, semantic and service quality-enhanced distributed search engine for improving geospatial resource discovery
abstract
Geospatial resource discovery is a critical step for developing geographic science applications. With the increasing number of geospatial resources available online, many Spatial Data Infrastructure (SDI) components (e.g. catalogues and portals) have been developed to help manage and discover geospatial resources. However, efficient and accurate geospatial resource discovery is still a big challenge because of the heterogeneity and complexity of decentralized network environments and interdisciplinary semantics. In this article, we report a search engine framework for efficient geospatial resource discovery, which reduces integration costs by leveraging existing Geospatial Cyberinfrastructure (GCI) components. Specifically, (1) the framework provides integration capability and flexibility by adopting the brokering approach, implementing a ‘plug-in’-based framework for metadata processing and proposing a dynamically configurable search workflow; (2) the asynchronous messaging and batch processing-based metadata record retrieval mode enhances the search performance and user interactivity; (3) an embedded semantic support system improves the discovery recall level and precision by providing semantic-based search rule creation and result similarity evaluation functions and (4) the engine assists user decision-making by integrating a service quality monitoring and evaluation system, data/service visualization tools, multiple views and additional information. Experiments and a search example show that the proposed engine helps both scientists and general users search for more accurate results with enhanced performance and user experience through a user-friendly interface.
Zhipeng Gui, Chaowei Phil Yang, Jizhe Xia, Kai Liu 0017, Jing Li 0029, Peter Lostritto
Int. J. Geogr. Inf. Sci.1
2012 An experimental study of open-source cloud platforms for dust storm forecasting
abstract
Cloud computing is becoming a viable computing solution for scientific research and several open-source cloud solutions are available to support scientific studies. However, little has been done to systematically investigate the performance of these solutions in supporting scientific pursuits. Taking dust storm forecasting as an example, we test three popular open-source cloud solutions, namely Eucalyptus, OpenNebula, and CloudStack, on the same hardware and compare against a bare cluster. We find that: (1) compared to the bare cluster, a cloud has about 10% virtualization and management overhead when one virtual machine is used. Overhead increases when more virtual machines are used. Leveraging more virtual resources would not necessarily yield better performance. (2) For computing- and communication-intensive dust storm forecasting, the performance overhead is mainly due to virtualized network rather than virtualized computing resources when more than one virtual machine is involved. (3) Compared to Eucalyptus and CloudStack, OpenNebula provides better support for dust storm forecasting with relatively better performance. The results can provide some insights for scientific community in adopting these open-source cloud solutions.
Qunying Huang, Jizhe Xia, Chaowei Phil Yang, Kai Liu 0017, Jing Li 0029, Zhipeng Gui, Mohammed Anowarul Hassan, Songqing Chen
SIGSPATIAL/GIS6