VLDB 2026 Research / reviewers in the wild / expert
Huayi Wu
dblp:16/3208
· DBLP profile ↗
42ranked-venue papers
7as first author
21since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 13 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Systems, architecture and hardware · 5 · 1 since 2021Computer networks · 3 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | PostGISer: The first end-to-end fine-tuned large language model for PostGIS GeoSQL query generation
Shuyang Hou, Lutong Xie, Haoyue Jiao, Shaowen Wu, Xuefeng Guan, Huayi Wu |
Inf. Process. Manag. | 8 |
| 2026 | GeoSQL-Eval: first evaluation of LLMs on PostGIS-based NL2GeoSQL queries
Shuyang Hou, Haoyue Jiao, Lutong Xie, Shaowen Wu, Xuefeng Guan, Huayi Wu |
Expert Syst. Appl. | 8 |
| 2026 | GeoCogent: an LLM-based agent for geospatial code generationabstractAutomated geospatial code generation is increasingly essential for complex spatial analysis and interdisciplinary GIS applications. However, existing large language models (LLMs) often fail to meet demands for interpretation, syntax adaptation, path retrieval, and validation, frequently producing ‘code hallucinations’—unreadable or non-functional code. To address this, we propose GeoCogent, the first intelligent framework for geospatial code generation powered by LLMs. GeoCogent integrates planning, tool-augmented reasoning, and memory mechanisms to support demand interpretation, dynamic knowledge retrieval, and consistent context maintenance, tackling core challenges in geospatial coding. We also developed and open-sourced GeoCodes, a benchmark dataset with evaluation metrics, and conducted comparative experiments, ablation studies, and case demonstrations. Results show GeoCogent achieves high efficiency across explicit, incomplete, and open-ended requirements, with each mechanism contributing significantly. Moreover, the prototype system supports local deployment and LLM integration, offering a low-barrier, efficient tool for geospatial development. By lowering technical thresholds and improving code reliability, GeoCogent advances intelligent geospatial analysis and enables interdisciplinary users to address complex analytical challenges. Shuyang Hou, Haoyue Jiao, Jianyuan Liang, Zhangxiao Shen, Huayi Wu |
Int. J. Geogr. Inf. Sci. | 6 |
| 2026 | Addressing the challenge of spatiotemporal data sparsity in disasters: predicting flood risks by aggregating distant neighbors on the basis of the principle of geographic similarityabstractWith floods being among the most frequent disasters of the 21st century, timely prediction of flood risks is crucial for early alerts to governments and residents, thereby reducing potential losses. Owing to the uneven spatiotemporal distribution of disaster data and limited sharing, assessing risks in blind spots without monitoring is challenging. In this paper, we propose a flood risk prediction method that integrates spatial neighbors on the basis of geographic similarity. Considering the correlation between surrounding environments and flood occurrence, this method introduces high-order neighborhoods to model internal and external environments of spatial units. Given the sparse spatiotemporal coverage of monitoring data, we used a semi-supervised approach and graph attention to aggregate many unlabeled spatial units to achieve comprehensive representations of these environments. Subsequently, we constructed a semantic space association graph and established local and global features. Through multistage label propagation, we accounted for changes in spatial unit attributes when predicting flood risks. Applied to the “7.20” Zhengzhou extreme rainstorm event, this approach achieved a micro ROC-AUC of 0.7374 for spatial risk prediction and an accuracy of 0.7888 for temporal risk prediction, proving its effectiveness in identifying risks in data-sparse regions and addressing gaps in flood monitoring. Shunli Wang 0005, Rui Li 0046, Huayi Wu |
Int. J. Geogr. Inf. Sci. | 3 |
| 2026 | Extraction of geoprocessing modeling knowledge from crowdsourced Google Earth Engine scripts by coordinating large and small language modelsabstractThe widespread use of online geoinformation platforms, such as Google Earth Engine (GEE), has produced numerous scripts. Extracting domain knowledge from these crowdsourced scripts supports understanding of geoprocessing workflows. Small Language Models (SLMs) are effective for semantic embedding but struggle with complex code; Large Language Models (LLMs) can summarize scripts, yet lack consistent geoscience terminology to express knowledge. In this paper, we propose Geo-CLASS, a knowledge extraction framework for geospatial analysis scripts that coordinates large and small language models. Specifically, we designed domain-specific schemas and a schema-aware prompt strategy to guide LLMs to generate and associate entity descriptions, and employed SLMs to standardize the outputs by mapping these descriptions to a constructed geoscience knowledge base. Experiments on 237 GEE scripts, selected from 295,943 scripts in total, demonstrated that our framework outperformed LLM baselines, including Llama-3, GPT-3.5 and GPT-4o. In comparison, the proposed framework improved accuracy in recognizing entities and relations by up to 31.9% and 12.0%, respectively. Ablation studies and performance analysis further confirmed the effectiveness of key components and the robustness of the framework. Geo-CLASS has the potential to enable the construction of geoprocessing modeling knowledge graphs, facilitate domain-specific reasoning and advance script generation via Retrieval-Augmented Generation (RAG). Zhipeng Gui, Jianyuan Liang, Dehua Peng, Wenzhang Wei, Shuyang Hou, Huayi Wu |
Int. J. Geogr. Inf. Sci. | 8 |
| 2026 | Hybrid Quantum-Transformer Network-Based Probabilistic Multi-Energy Flow Calculation in Integrated Energy SystemabstractProbabilistic energy flow (PEF) is a fundamental task for the operation of conventional model-based integrated energy systems (IES), electricity–gas–heating coupled systems, while challenged by heavy computational burdens and uncertainties introduced by renewable energy sources (RES). To overcome this bottleneck, a novel data-driven hybrid quantum-transformer network (HQTN) is proposed for fast PEF calculation. Specifically, the classical Transformer network captures the dynamic spatial uncertainty features via a novel graph attention mechanism to model the uncertain fluctuating RES features. Additionally, a quantum neural network is proposed to extract the high-dimensional, complex relationships of IES operational data, thereby improving calculational accuracy. Case studies are conducted on both IEEE 33-node electric/20-node gas/17-node heating and 118-node electric/48-node gas/14-node heating systems. Experimental evaluation results verify that the proposed approach achieves high accuracy and maintains strong computational performance in PEF calculation. Huayi Wu, Zhao Xu 0002 |
IEEE Internet Things J. | 1 |
| 2026 | Topological and semantic contrastive graph clustering by Ricci curvature augmentation and hypergraph fusionabstractContrastive graph clustering is an advanced technology in the field of cluster analysis. By leveraging graph neural networks and contrastive learning paradigm, it enables the coupling of topological structure and node semantic information for attributed graph networks. Graph augmentation and positive sample selection are two essentials of contrastive graph clustering. However, existing graph augmentation methods tend to disrupt the cluster structures, and most positive sample selectors suffer from the false negative sample problem. In this paper, we propose a Topological and Semantic Contrastive Graph Clustering (TSCGC) model consisting of three learning components. The representation learning component augments original graph using Ricci curvature to preserve the cluster structure, and introduces hypergraph view to capture high-order relationships. Graph and hypergraph convolutional networks are used to encode the triple-view embeddings. Meanwhile, we develop a dual contrastive learning component to extract the topological and semantic information. To reduce the number of false negatives, it utilizes K-means to generate pseudo cluster labels to guide the selection of positive samples. The self-supervised learning component is leveraged to align the three graph views. The final clustering results are obtained by performing K-means on the aligned embeddings. We demonstrated the effectiveness by comparing the performance of TSCGC with 13 clustering baselines on six real-world networks. Ablations verified the validity of key components and the impact of parameter settings were also analyzed. We further applied TSCGC to identify the function types of 10,370 buildings in ShenZhen City, China based on multi-source geospatial data. It achieved the highest accuracy and exhibit significant potential in handling complex network structures and high-dimensional node features. The code is available at: https://github.com/ZPGuiGroupWhu/TSCGC . Dehua Peng, Guangyao Fang, Zhipeng Gui, Huayi Wu |
Knowl. Based Syst. | 5 |
| 2026 | MeanCut: Greedy Graph Clustering by Fast Maximum Spanning Tree and Degree Descent CriterionabstractAs the most typical graph clustering method, spectral clustering is popular and attractive due to its remarkable performance, easy implementation, and strong adaptability. Classical spectral clustering measures the edge weights using pairwise Euclidean similarity and resolves the optimal graph partitioning by relaxing the constraints of indicator matrix and decomposing the Laplacian matrix. However, Euclidean similarity might cause skew graph cuts when handling non-spherical clusters, and the relaxation strategy introduces information loss. Meanwhile, spectral clustering requires specifying the number of clusters, which is difficult to determine without enough prior knowledge. In this work, we propose a greedy-optimized scheme for resolving the indicator matrix using path-based similarity and degree descent criterion. It yields an indicator matrix with strictly binary entries without destructive relaxation and discretization steps. Path-based similarity can enhance the intra-cluster associations of arbitrary-shaped clusters, while degree descending is theoretically proven to be the best order to minimize our proposed objective function MeanCut. Moreover, we define a density gradient factor to separate clusters with fuzzy boundaries, and develop a fast maximum spanning tree algorithm to improve the scalability of similarity calculation. The effectiveness of MeanCut has been demonstrated on synthetic datasets and real-world benchmarks. By fusing multi-view image features, MeanCut outperforms cutting-edge subspace clustering methods in face recognition. The code is available at:https://github.com/ZPGuiGroupWhu/MeanCut. Dehua Peng, Zhipeng Gui, Jie Gui, Huayi Wu |
IEEE Trans. Fuzzy Syst. | 6 |
| 2025 | A Local Moran's I guided transformer cellular automata for simulating heterogeneous urban growthabstractThe rapid advancement of urbanization in recent decades has attracted extensive application of cellular automata (CA)-based models to simulate urban growth for planning and decision-making. However, the inaccurate representation of heterogeneous spatial interactions between urban units and the neglect of autocorrelated growth patterns in urbanization lead to unreliable simulation results of CA-based models. To address these two limitations, this study proposes a novel CA-based model integrated with Transformer network and Local Moran’s I, namely TL-CA. The Transformer network is built to quantify heterogeneous interaction between neighbors using the self-attention mechanism. Subsequently, Local Moran’s I is employed to implicitly guide the network in learning spatially autocorrelated patterns of urban growth through auxiliary learning. Finally, the development potential estimated from driving factors, i.e. the network output, is incorporated into CA to simulate urban growth. Land use data from Wuhan (2000–2020) are selected to verify TL-CA’s performance. The results demonstrate that TL-CA achieves the highest simulation accuracy, with an average increase in the figure of merit (FoM) of 9.97%. Attention visualization and residual analysis explain the model’s effectiveness in modeling heterogeneous interactions and autocorrelated growth. Additionally, TL-CA exhibits high computational efficiency and low resource consumption, with sufficient potential to support larger-scale research. Qingyang Xu, Xuefeng Guan, Changlan Yang, Huayi Wu |
Int. J. Geogr. Inf. Sci. | 6 |
| 2025 | Geo-FuB: A method for constructing an Operator-Function knowledge base for geospatial code generation with large language modelsabstractThe rapid growth of spatiotemporal data and the increasing demand for geospatial modeling have driven the automation of these tasks with large language models (LLMs) to enhance research efficiency. However, general LLMs often encounter hallucinations when generating geospatial code due to a lack of domain-specific knowledge on geospatial functions and related operators. The retrieval-augmented generation (RAG) technique, integrated with an external operator-function knowledge base, provides an effective solution to this challenge. To date, no widely recognized framework exists for building such a knowledge base. This study presents a comprehensive framework for constructing the operator-function knowledge base, leveraging semantic and structural knowledge embedded in geospatial scripts. The framework consists of three core components: Function Semantic Framework Construction (Geo-FuSE), Frequent Operator Combination Statistics (Geo-FuST), and Combination and Semantic Framework Mapping (Geo-FuM). Geo-FuSE employs techniques like Chain-of-Thought (CoT), TF-IDF, t-SNE, and Gaussian Mixture Models (GMM) to extract semantic features from scripts; Geo-FuST uses Abstract Syntax Trees (AST) and the Apriori algorithm to identify frequent operator combinations; Geo-FuM combines LLMs with a fuzzy matching algorithm to align these combinations with the semantic framework, forming the Geo-FuB knowledge base. The instance of Geo-FuB, named GEE-FuB, has been developed using 154,075 Google Earth Engine scripts and is available at https://github.com/whuhsy/GEE-FuB . Based on a set of well-defined evaluation metrics introduced in this study, the GEE-FuB construction achieved an overall accuracy of 88.89 %, demonstrating a 31 % to 34 % reduction in hallucinations compared to mainstream LLMs without external knowledge integration. This research introduces a novel approach to knowledge mining and knowledge base construction specifically tailored for geospatial code generation tasks, broadening the applications of knowledge base construction and providing valuable theoretical insights, practical examples, and data resources for related research fields. Shuyang Hou, Jianyuan Liang, Zhangxiao Shen, Huayi Wu |
Knowl. Based Syst. | 5 |
| 2025 | A Robust and Efficient Boundary Point Detection Method by Measuring Local Direction DispersionabstractBoundary point detection aims to outline the external contour structure of clusters and enhance the inter-cluster discrimination, thus bolstering the performance of the downstream classification and clustering tasks. However, existing boundary point detectors are sensitive to density heterogeneity or cannot identify boundary points in concave structures and high-dimensional manifolds. In this work, we propose a robust and efficient boundary point detection method based on Local Direction Dispersion (LoDD). The core of boundary point detection lies in measuring the difference between boundary points and internal points. It is a common observation that an internal point is surrounded by its neighbors in all directions, while the neighbors of a boundary point tend to be distributed only in a certain directional range. By considering this observation, we adopt density-independent K-Nearest Neighbors (KNN) method to determine neighboring points and design a centrality metric LoDD using the eigenvalues of the covariance matrix to depict the distribution uniformity of KNN. We also develop a grid-structure assumption of data distribution to determine the parameters adaptively. The effectiveness of LoDD is demonstrated on synthetic datasets, real-world benchmarks, and application of training set split for deep learning model and hole detection on point cloud data. The datasets and toolkit are available at:https://github.com/ZPGuiGroupWhu/lodd. Dehua Peng, Zhipeng Gui, Jie Gui, Huayi Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Dynamic Visual Semantic Sub-Embeddings and Fast Re-Ranking for Image-Text RetrievalabstractThe core of image-text retrieval is to accurately measure the similarity between different modalities in a unified representation space. However, compared to textual descriptions of a certain perspective, the visual modality has more semantic variations. Therefore, images are usually associated with multiple textual captions in databases. Although popular symmetric embedding methods have explored numerous modal interaction approaches, they often learn toward outputting the average representation of multiple semantic variations within image embeddings. Consequently, information entropy in embeddings is increased, resulting in redundancy and decreased accuracy. In this work, we propose a Dynamic Visual Semantic Sub-Embeddings framework (DVSE) to reduce the information entropy. Specifically, we obtain a set of heterogeneous visual sub-embeddings through dynamic orthogonal constraint loss. To encourage the generated candidate image embeddings to capture various semantic variations, we construct a mixed distribution and employ a variance-aware weighting loss to assign different weights to the optimization process. In addition, we develop a Fast Re-ranking strategy (FR) to efficiently evaluate the retrieval results and enhance the performance. We compare the performance with existing set-based method using five image feature encoders and three text feature encoders on three benchmark datasets: MSCOCO, Flickr30K and CUB Captions. We also show the role of different components by ablation studies and perform a sensitivity analysis of the hyperparameters. The qualitative analysis of visualized bidirectional retrieval and attention maps further demonstrates the ability of our method to encode semantic variations. Wenzhang Wei, Zhipeng Gui, Changguang Wu, Dehua Peng, Huayi Wu |
IEEE Trans. Multim. | 6 |
| 2024 | GridMesa: A NoSQL-based big spatial data management system with an adaptive grid approximation model
Xuefeng Guan, Zhaoxing Pang, Xing Kui, Huayi Wu |
Future Gener. Comput. Syst. | 5 |
| 2024 | Map retrieval intention recognition based on relevance feedback and geographic semantic guidance: For better understanding user retrieval demandsabstractEffective retrieval is essential for finding resources in demand handily amidst extensive data records in data warehouse. Mainstream map retrieval methods suffer from intention gap problem and are incapable to describe sophisticated user demands precisely due to the limits of low- and middle-level text or visual feature matching, resulting in unsatisfactory retrieval results. Such limitations are more marked when map retrieval demands were characterized with joint constraints of geographic concepts. To address this issue, we propose a map retrieval intention recognition method to perceive user demands with relevance feedback samples and geographic semantics guidance. Specifically, we construct a hierarchical intention expression model to describe retrieval goals and their multi-dimensional attribute constrains; incorporate geographic ontologies to provide semantic guidance and facilitate recognition; utilize the frequent itemset mining (FIM) algorithm Apriori to generate intention candidates from relevance feedback samples, and search for the optimal intention set by adopting the minimum description length (MDL) principle. The experiments verify the effectiveness of Apriori algorithm and MDL principle on intention recognition. The proposed method outperforms the FIM algorithm Gene Ontology (RuleGO) and the Decision Tree algorithm with Hierarchical Features (DTHF) with higher recognition accuracy and noise tolerance. Furthermore, through our sample augmentation strategy, the method yields promising recognition accuracy even when the feedback sample size is as low as ten, substantially reducing the feedback burden in human-computer interactions. We envision that the application of our method in spatial data infrastructures (SDIs), such as geoportals and catalogue services, could enhance the quality of service and user experience in geospatial data discovery. Zhipeng Gui, Xinjie Liu, Zhipeng Ling, Fa Li, Zelong Yang 0001, Huayi Wu, Shuangming Zhao |
Inf. Process. Manag. | 9 |
| 2024 | Robust scientific text classification using prompt tuning based on data augmentation with L2 regularization
Shijun Shi, Kai Hu 0005, Jie Xie 0001, Ya Guo 0001, Huayi Wu |
Inf. Process. Manag. | 5 |
| 2024 | HSeq2Seq: Hierarchical graph neural network for accurate mobile traffic forecasting
Rihui Xie, Xuefeng Guan, Xinglei Wang, Huayi Wu |
Inf. Sci. | 5 |
| 2024 | Multi-Energy Load Forecasting in Integrated Energy Systems: A Spatial-Temporal Adaptive Personalized Federated Learning ApproachabstractShort-term forecasting of multienergy loads is of paramount significance for integrated energy systems operation. The central forecasting framework is confronted with the privacy disclosure issue. Besides, the intricate interdependencies among diverse energy loads present an opportunity to improve prediction accuracy. To this end, a privacy-preserving spatial-temporal adaptive personalized federated learning model is proposed in this article. Specifically, the proposed federated learning-based decentralized framework enables the sharing of local model weights while ensuring the confidentiality of raw measurement data. Besides, the spatial-temporal transformer leverages the self-attention mechanism to synchronously capture the complex dynamic dependencies among different types of energy load demand. Furthermore, the adaptive local aggregation mechanism is proposed to personalize the local model to address the data heterogeneity and subsequently improve forecasting accuracy. The proposed model is applied to a publicly available dataset. The results show that the proposed model can achieve highly efficient and effective forecasting accuracy. Huayi Wu, Zhao Xu 0002 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Gridtopo-GAN for Distribution System Topology IdentificationabstractDue to the limited presence of monitoring and measurement devices, timely identification of distribution grid topology has been challenging. Therefore, this article proposes a power grid topological generative adversarial network (Gridtopo-GAN) model to identify the distribution grid topology of either meshed or radial structure with limited measurements. By leveraging the topology preserved node embedding architecture, this model can efficiently handle large-scale systems with different topological configurations. Because of the generative capability of GAN, the model is robust enough when fed with bad measurement data, including missing data, commonly encountered in practical applications. Numerical simulations are carried out on the IEEE 33-node system, 118-node, 415-node, and real 76-node distribution systems to demonstrate the effectiveness and efficiency of the proposed topology identification model. Huayi Wu, Zhao Xu 0002, Jian Zhao 0023, Songjian Chai |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Browsing behavior modeling and browsing interest extraction in the trajectories on web map service platforms
Guangsheng Dong, Rui Li 0046, Huayi Wu, Wei Huang 0045, Hongping Zhang |
Expert Syst. Appl. | 3 |
| 2022 | MPR-GAN: A Novel Neural Rendering Framework for MLS Point Cloud With Deep Generative LearningabstractEfficient point cloud visualization is indispensable for practical applications. In the context of point cloud visualization, 3-D rendering can be viewed as the kernel that transforms 3-D points into a 2-D scene image. Compared with traditional point-based rendering (PBR), neural image-based rendering (NIBR) has gradually emerged as a feasible solution for point cloud rendering. To efficiently render sparse and colorless mobile laser scanning (MLS) point cloud, we propose a novel neural rendering framework based on deep generative learning, named MLS point cloud rendering with generative adversarial network (MPR-GAN). In this framework, perspective projection with intrinsic parameter scaling and cumulative distribution normalization is first utilized to transform the 3-D point cloud into a compact 2-D image; a conditional generative adversarial network (CGAN)-based rendering model is then proposed to generate a photorealistic scene image from the projected 2-D image. In this CGAN model, the asymmetric encoder–decoder generator can implement inpainting and true colorization using context feature capturing and edge information perception; a multiscale discriminator is built to guarantee the model output with global consistency and local details. Moreover, a hybrid loss function is designed to improve the visual quality of the generated images with similarity constraints from both the content and the structure. Two public MLS point cloud datasets are selected and employed to carry out extensive evaluation using MPR-GAN and other baseline frameworks. The experimental results demonstrate that MPR-GAN achieves the state-of-the-art rendering performance in terms of peak signal-to-noise ratio (PSNR) and structural similarity (SSIM). Furthermore, the efficiency analysis shows that MPR-GAN can support real-time rendering, achieving end-to-end rendering from raw points. Qingyang Xu, Xuefeng Guan, Huayi Wu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | LSI-LSTM: An attention-aware LSTM for real-time driving destination prediction by considering location semantics and location importance of trajectory pointsabstractIndividual driving final destination prediction supports location-based services such as personalized service recommendations, traffic navigation, and public transport dispatching. However, real-time destination prediction is challenging due to the complexity of temporal dependencies, and the strong influence of travel spatiotemporal semantics and spatial correlations. Besides temporal context, the nearby urban functionalities of traveling zones and departure regions, and the crucial positions on the road network where trajectory points located would reflect the travel intentions of drivers. However, these spatial factors are rarely considered in existing studies. To fill this gap, we propose a real-time individual driving destination prediction model LSI-LSTM based on an attention-aware Long Short-Term Memory (LSTM) by taking Location Semantics and Location Importance of trajectory points into account. More specifically, a trajectory location semantics extraction method (t-LSE) enriches feature description with prior knowledge for implicit travel intentions learning. t-LSE represents urban functionality through Points of Interest (POIs) using Term Frequency-Inverse Document Frequency (TF-IDF). Meanwhile, a novel trajectory spatial attention mechanism (t-SAM) captures the trajectory points that strongly correlate to candidate destinations based on the location importance inferred from the driving status, i.e., turning angle, driving speed, and traveled distance. Comparative experiments with three baseline methods, i.e., Hidden Markov Model, Random Forest, and LSTM, demonstrate significant prediction accuracy improvements of LSI-LSTM on four individual trajectory datasets. Further analyses validate the effectiveness of the proposed semantic extraction method and attention mechanism, and also discuss the factors that may affect the prediction results. Zhipeng Gui, Yunzeng Sun, Dehua Peng, Fa Li, Huayi Wu, Chi Guo, Wenfei Guo, Jianya Gong |
Neurocomputing | 6 |
| 2020 | MSGC: Multi-scale grid clustering by fusing analytical granularity and visual cognition for detecting hierarchical spatial patternsabstractSpatial clustering is a widely used data mining method for discovery of spatial aggregation pattern. However, existing methods often neglect scale dependence, impeding the full recognition of point patterns and the detection of hierarchical spatial structures. Spatial clustering is scale dependent and linked to the size of analysis unit as well as the hierarchy of visual cognition. Therefore, this paper proposes a novel multi-scale grid clustering (MSGC) algorithm, which fuses dual scale factors, i.e., analytical scale and visual scale that sequentially integrates multi-analytical-scale clustering (MASC) and multi-visual-scale clustering (MVSC). MASC generates multi-granularity grids to transform the analytical scales, and MVSC extracts multi-level clusters to express the hierarchy of visual cognition. Comparative experiments validated the proposed algorithm against the classical Density-based Spatial Clustering of Applications with Noise (DBSCAN) and WaveCluster algorithms on both synthetic and real-world geographic datasets. The results demonstrate that MSGC can generate multi-scale clusters for increased understanding of the spatial aggregation patterns and hierarchical structures of geographic entities. Moreover, it can eliminate noise adaptively and effectively identify clusters with arbitrary shapes. Due to the nature of grid clustering, the low computational complexity enables near real-time visual analytics and efficient point pattern mining on large spatial datasets. Zhipeng Gui, Dehua Peng, Huayi Wu |
Future Gener. Comput. Syst. | 3 |
| 2020 | Optimizing and accelerating space-time Ripley 's K function based on Apache Spark for distributed spatiotemporal point pattern analysisabstractWith increasing point of interest (POI) datasets available with fine-grained spatial and temporal attributes, space–time Ripley’s K function has been regarded as a powerful approach to analyze spatiotemporal point process. However, space–time Ripley’s K function is computationally intensive for point-wise distance comparisons, edge correction and simulations for significance testing. Parallel computing technologies like OpenMP, MPI and CUDA have been leveraged to accelerate the K function, and related experiments have demonstrated the substantial acceleration. Nevertheless, previous works have not extended optimization of Ripley’s K function from space dimension to space–time dimension. Without sophisticated spatiotemporal query and partitioning mechanisms, extra computational overhead can be problematic. Meanwhile, these researches were limited by the restricted scalability and relative expensive programming cost of parallel frameworks and impeded their applications for large POI dataset and Ripley’s K function variations. This paper presents a distributed computing method to accelerate space–time Ripley’s K function upon state-of-the-art distributed computing framework Apache Spark, and four strategies are adopted to simplify calculation procedures and accelerate distributed computing respectively: (1) spatiotemporal index based on R-tree is utilized to retrieve potential spatiotemporally neighboring points with less distance comparison; (2) spatiotemporal edge correction weights are reused by 2-tier cache to reduce repetitive computation in L value estimation and simulations; (3) spatiotemporal partitioning using KDB-tree is adopted to decrease ghost buffer redundancy in partitions and support near-balanced distributed processing; (4) customized serialization with compact representations of spatiotemporal objects and indexes is developed to lower the cost of data transmission. Based on the optimized method, a web-based visual analytics framework prototype has been developed. Experiments prove the feasibility and time efficiency of the proposed method, and also demonstrate its value on promoting applications of space–time Ripley’s K function in ecology, geography, sociology, economics, urban transportation and other fields. Zhipeng Gui, Huayi Wu, Dehua Peng, Jinghang Wu, Zousen Cui |
Future Gener. Comput. Syst. | 3 |
| 2020 | A hierarchical temporal attention-based LSTM encoder-decoder model for individual mobility prediction
Fa Li, Zhipeng Gui, Zhao-Yu Zhang 0003, Dehua Peng, Kunxiaojia Yuan, Yunzeng Sun, Huayi Wu, Jianya Gong, Yichen Lei |
Neurocomputing | 8 |
| 2020 | Distribution Network Reconfiguration for Loss Reduction and Voltage Stability With Random Fuzzy Uncertainties of Renewable Energy Generation and LoadabstractThis paper presents a distribution reconfiguration framework to adapt to the distribution network considering the twofold random and fuzzy uncertainties of the wind, photovoltaic power generation, and load demand and considering the power loss reduction and voltage stability. First, the random fuzzy power output models of distributed generation and load are built based on the random fuzzy theory. Second, the corresponding objective functions are established, which are the random fuzzy expected value of active power loss and maximum probability of voltage limit. Third, the modified particle swarm optimization (MPSO) algorithm based on Kruskal algorithm is introduced for the first time to determine the optimal network topology. The Kruskal algorithm is employed to generate a radial network topology directly without checking the loops and islands. Lastly, the proposed method is applied to the IEEE33, PG&E 69-bus distribution systems, 25-bus unbalanced distribution system, as well as a real 109-bus distribution system. The results show that the proposed method has a good performance to solve distribution network reconfiguration problem. Huayi Wu, Mingbo Liu |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | A quad-tree-based fast and adaptive Kernel Density Estimation algorithm for heat-map generationabstractKernel Density Estimation (KDE) is a classic algorithm for analyzing the spatial distribution of point data, and widely applied in spatial humanities analysis. A heat-map permits intuitive visualization of spatial point patterns estimated by KDE without any overlapping. To achieve a suitable heat-map, KDE bandwidth parameter selection is critical. However, most generally applicable bandwidth selectors of KDE with relatively high accuracy encounter intensive computation issues that impede or limit the applications of KDE in big data era. To solve the complex computation problems, as well as make the bandwidths adaptively suitable for spatially heterogenous distributions, we propose a new Quad-tree-based Fast and Adaptive KDE (QFA-KDE) algorithm for heat-map generation. QFA-KDE captures the aggregation patterns of input point data through a quad-tree-based spatial segmentation function. Different bandwidths are adaptively calculated for locations in different grids calculated by the segmentation function; and density is estimated using the calculated adaptive bandwidths. In experiments, through comparisons with three mostly used KDE methods, we quantitatively evaluate the performance of the proposed method in terms of correctness, computation efficiency and visual effects. Experimental results demonstrate the power of the proposed method in computation efficiency and heat-map visual effects while guaranteeing a relatively high accuracy. Kunxiaojia Yuan, Xiaoqiang Cheng, Zhipeng Gui, Fa Li, Huayi Wu |
Int. J. Geogr. Inf. Sci. | 5 |
| 2019 | Understanding the topic evolution of scientific literatures like an evolving city: Using Google Word2Vec model and spatial autocorrelation analysis
Kai Hu 0005, Kunlun Qi, Siluo Yang, Xiaokang Fu, Jie Zheng 0007, Huayi Wu, Ya Guo 0001, Qibing Zhu |
Inf. Process. Manag. | 8 |
| 2018 | Hierarchical decomposition method and combination forecasting scheme for access load on public map service platforms
Rui Li 0046, Huayi Wu, Guangsheng Dong, Jie Jiang 0014 |
Future Gener. Comput. Syst. | 3 |
| 2017 | Spatiotemporal correlation in WebGIS group-user intensive access patternsabstractGroup-user intensive access to WebGIS exhibits spatiotemporal behaviour patterns with aggregation features and regularity distributions when geospatial data are accessed repeatedly over time and aggregated in certain spatial areas. We argue that these observable group-user access patterns provide a foundation for improved optimization of WebGIS so that it can respond to volume intensive requests with a higher quality of service and improve performance. Subsequently, a measure of access popularity distribution must precisely reflect the access aggregation and regularity features found in group-user intensive access. In our research, we considered both the temporal distribution characteristics and spatial correlation in the access popularity of tiled geospatial data (tiles). Based on the observation that group-user access follows a Zipf-like law, we built a tile-access popularity distribution based on time-sequence, to express the access aggregation of group-users with heavy-tailed characteristics. Considering the spatial locality of user-browsed tiles, we built a quantitative expression for the correlation between tile-access popularities and the distances to hotspot tiles, reflecting the attenuation of tile-access popularity to distance. Moreover, given the geographical spatial dependency and scale attribute of tiles, and the time-sequence of tile-access popularity, we built a Poisson regression model to express the degree of correlation among the accesses to adjacent tiles at different scales, reflecting the spatiotemporal correlation in tile access patterns. Experiments verify the accuracy of our Poisson regression model, which we then applied to a cluster-based cache-prefetching scenario. The results show that our model successfully reflects the spatiotemporal aggregation features of group-user intensive access and group-user behaviour patterns in WebGIS. The refined mathematical method in our model represents a time-sequence distribution of intensive access to tiles and the spatial aggregation and correlation in access to tiles at different scales, quantitatively expressing group-user spatiotemporal behaviour patterns with aggregation features and a regular distribution. Our proposed model provides a precise and empirical basis for performance-optimization strategies in WebGIS services, such as planning computing resource allocation and utilization, distributed storage of geospatial data, and providing distributed services so as to respond rapidly to geospatial data requests, thus addressing the challenges of volume-intensive user access. Rui Li 0046, Jiapei Fan, Jie Jiang 0014, Huayi Wu |
Int. J. Geogr. Inf. Sci. | 4 |
| 2015 | Land-Use Scene Classification in High-Resolution Remote Sensing Images Using Improved CorrelatonsabstractExisting methods that incorporate spatial information into a traditional Bag-of-Visual-Words (BoVW) model consider the spatial arrangement of an image but ignore pixel homogeneity in land-use remote sensing images. In this letter, we present an improved correlaton model to jointly integrate appearance, spatial correlation, and pixel homogeneity using multiscale segmentation. The effectiveness of the proposed method was tested on a ground truth image data set of 21 land-use classes manually extracted from high-resolution remote sensing images. The experimental results demonstrate that our improved correlaton model can promote classification and outperforms existing methods such as the traditional BoVW model, spatial pyramid matching model, and the traditional correlaton model. Kunlun Qi, Huayi Wu, Jianya Gong |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2014 | ReSDaP: A Real-Time Data Provision System Architecture for Sensor Webs
Huayi Wu |
W2GIS | 3 |
| 2013 | A Load-balancing method for network GISs in a heterogeneous cluster-based system using access density
Rui Li 0046, Yinfeng Zhang, Zhengquan Xu, Huayi Wu |
Future Gener. Comput. Syst. | 4 |
| 2012 | Spatial data quality and beyondabstractIssues of accuracy, uncertainty, and spatial data quality have been on the top of most GIScience research agendas around the world from the late 1980s. Ever since then, growing research efforts have been directed toward uncertainty characterization in spatial information, analysis, and applications, aiming for better understanding of spatial uncertainty and thus improved methods and techniques for assessing and managing data quality. Impressive progress has been made in various issues concerning data quality. In addition, growing research on extensions to the conventional norms of data quality, such as the quality aspects of geospatial information services, has been observed. Chinese researchers have contributed to this great cause by keeping abreast with the developments abroad and striving for their own innovative work. This paper reviews the past research on data quality-related issues and provides a perspective on future developments. These will be seen not only in continued research on theoretical and technical issues concerning data quality, but also in developments of tools for quality assessment and decision-making under uncertainty through geospatial information processing and applications. DeRen Li, Jingxiong Zhang, Huayi Wu |
Int. J. Geogr. Inf. Sci. | 3 |
| 2012 | Dense Corresponding Pixel Matching Between Aerial Epipolar Images Using an RGB-Belief Propagation AlgorithmabstractA new algorithm, RGB-belief propagation (RGB-BP), for dense corresponding pixel matching between aerial epipolar images is proposed in this letter. Evolved from the traditional belief-propagation algorithm, RGB-BP makes full use of all the three color components, R, G, and B, and simplifies the parameter settings. In order to reduce the impact of the obvious color differences between corresponding pixels, RGB-BP reduces the sensitivity of central pixels and increases the contribution of neighboring pixels. Three pairs of aerial epipolar images are tested using RGB-BP. The experimental results demonstrate the effectiveness of RGB-BP. Bingxuan Guo, Huayi Wu, Jianya Gong, Tong Zhang 0011 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2011 | Sensitivity analysis of the sunshine duration parameters in SEBAL model for estimating winter wheat evapotranspirationabstractIt is an efficient approach to estimate winter wheat evapotranspiration (ET) from remote sensing images, but it remains a difficult problem to generate a series of daily or seasonal ETs through temporal interpolation from discrete ET estimations. The SEBAL (surface energy balance algorithm for land) model works well only for areas where sunny days dominate and with a relatively pollution-free atmosphere. This paper presents a modification of the SEBAL model for temporal ET interpolation which considers sunshine duration. This method is demonstrably more suitable for areas such as the Haihe basin in east China. A sensitivity analysis of the sunshine duration parameters was conducted for winter wheat ET in the Shijin irrigation area in Haihe basin for the period starting Oct. 2007 and ending Jun. 2008. Seasonal ET interpolations were based on MODIS products and validated by ground measurements from the Luancheng Agro-Ecosystem Experimental Station. The results show that the seasonal ET interpolations taking account of sunshine duration in the SEBAL model are better than results without considering the sunshine duration. A small change in the sunshine duration parameters results in significant change in the seasonal ET interpolation. Therefore, the SEBAL model is highly sensitive to the sunshine duration parameters and such a modification of SEBAL model improves the results of temporal interpolation. Huayi Wu |
IGARSS | 2 |
| 2011 | An optimized framework for seamlessly integrating OGC Web Services to support geospatial sciencesabstractOGC Web Services (OWS) are essential building blocks for the national and global spatial data infrastructure (NSDI and GSDI) and the geospatial cyberinfrastructure (GCI). Web Map Service (WMS), Web Feature Service (WFS), Web Coverage Service (WCS), and Catalogue Service for Web (CSW) have been increasingly adopted to serve scientific data. Interoperable services can facilitate the integration of different scientific applications by searching, finding, and utilizing the large number of scientific data and Web services. However, these services are widely dispersed and hard to be found and utilized with acceptable performance. This is especially true when developing a science application to seamlessly integrate multiple geographically dispersed services. Focusing on the integration of distributed OWS resources, we propose a layer-based service-oriented integration framework and relevant optimization technologies to search and utilize relevant resources. Specifically, (1) an AJAX (Asynchronous JAvaScript and eXtensible Markup Language)-based synchronous multi-catalogue search is proposed and utilized to enhance the multi-catalogue searching performance; (2) a layer-based search engine with spatial, temporal, and performance criteria is proposed and used for identifying better services; (3) a service capabilities clearinghouse (SCCH) is proposed and developed to address the service issues identified by a statistical experiment. A science application of data correlation analysis is used as an example to demonstrate the performance enhancement of the proposed framework. Zhenglong Li 0002, Chaowei Phil Yang, Huayi Wu, Wenwen Li 0001 |
Int. J. Geogr. Inf. Sci. | 3 |
| 2010 | Earth science data records sharing supported by the Spatial Web PortalabstractThis paper reports our research on utilizing OGC (Open Geospatial Consortium) or de facto standards to support the ESDRs (Earth Science Data Records) sharing within the SWP (Spatial Web Portal). In MEaSUREs (Making Earth Science Data Records for Use in Research Environments) project, we leverage our extensive experience with emerging technologies to develop a SWP to provide distributed data access, distributed data analysis and integrated data display. In this study, three parts are investigated on how to share ESDRs by MEaSUREs SWP including a) establishing a publicly-accessible ESDR data portal, b) developing an on-line meta-database to enable quick discovery of relevant ESDR resources, c) using intelligent, Web-enabled services that simplifies for users data access, processing and exchange to distribute ESDRs. Paul R. Houser, Chaowei Phil Yang, Huayi Wu |
IGARSS | 4 |
| 2009 | The integration of wireless sensor networks remote sensing and geographic information systems for autonomous environmental and animal monitoring
Huayi Wu |
IADIS AC (2) | 2 |
| 2007 | QoS multicast routing by using multiple paths/trees in wireless ad hoc networks
Huayi Wu, Xiaohua Jia |
Ad Hoc Networks | 1 |
| 2005 | New paradigm for compressed image quality metric: exploring band similarity with CSF and mutual informationabstractA new paradigm for designing compressed image quality metric is proposed in this paper, of which the most significant characteristic is the use of mutual information, a key concept in information theory which measures statistical dependence between two random variables, to exploring the degree of similarity of spatial visual information distribution across different frequency bands in image. Visual information used in the calculation of mutual information is extracted by contrast sensitivity function (CSF) and local band-limited contrast definition proposed by E. Peli. Our paradigm is more consistent with human perceptual mechanism comparing with the traditional error-summation based ones. The effectiveness of our image quality assessment paradigm is validated by JPEG & JPEG2000 compressed images at different bit rates and images with various types of noises. Haijun Zhu, Huayi Wu |
IGARSS | 2 |
| 2004 | Bandwidth-Guaranteed QoS Multicast Routing by Multiple Paths in AD Hoc Wireless NetworksabstractIn this paper, we investigate the issues of QoS multicast routing in ad hoc wireless networks. Due to limited bandwidth of a wireless node, a QoS multicast call could often be blocked if there does not exist a single multicast tree that has the requested bandwidth, even though there is enough bandwidth in the system to support the call. In this paper we propose a multicast routing scheme by using multiple paths or multiple trees to meet the bandwidth requirement of a call. Three multicast routing strategies are studied, SPT (shortest path tree) based multiple-paths (SPTM), least cost tree based multiple-paths (LCTM) and multiple least cost trees (MLCT). The final routing tree(s) can meet the user's QoS requirements such that the delay from the source node to the furthest destination node shall not exceed the bound and the aggregate bandwidth of the paths or trees shall meet the bandwidth requirement of the call. Extensive simulations have been conducted to evaluate the performance. The simulation results show that the new scheme has three major advantages: 1) it greatly reduces the system blockings; 2) multicast routing is in a fully distributed fashion; 3) the proposed routing protocol follows the format of existing on-demand multicast routing protocols for ad hoc networks, which makes it easy to be incorporated into the existing on-demand routing protocols Huayi Wu, Xiaohua Jia, Yanxiang He, Chuanhe Huang |
ICCCN | 1 |
| 2000 | An algebraic algorithm for point inclusion query
Huayi Wu, Jianya Gong, DeRen Li, Wenzhong Shi |
Comput. Graph. | 1 |