Jinfeng Wang 0001

dblp:25/1601-1 · DBLP profile ↗
← Back
18ranked-venue papers in the field
2as first author
5since 2021 · last 2026
0000-0002-6687-9420ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 16 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 2
YearPublicationVenuePosition
2026 Poisson means of stratified nonhomogeneity: a new method to predict spatial counts
Fengbei Shen, Chengdong Xu, Jinfeng Wang 0001, Maogui Hu, Yuehua Hu
Int. J. Geogr. Inf. Sci.3
2026 A model to identify causality for geographic patterns
abstract
Identifying causal relationships is essential for understanding the mechanisms through which natural and anthropogenic factors interact within Earth systems. However, in spatial cross-sectional data, the absence of temporal ordering poses significant challenges to traditional causal inference methods. This study proposes a novel Geographical Pattern Causality (GPC) model to detect positive, negative, dark causality and its strength between variables in spatial data. Grounded in dynamical systems theory and generalized embedding principles, the method transforms spatial neighbourhoods into lagged sequences, reconstructs the phase space, and compares symbolic trajectories to assess predictability and consistency in pattern changes—thereby inferring both the direction and type of causality. Case studies demonstrated that, compared to correlation analysis and Linear Non-Gaussian Acyclic Model (LiNGAM), the GPC model could reveal latent causal relationships among weakly correlated variables in geographical systems and capture diverse causal patterns. Despite limitations, such as sensitivity to noise and potential biases from proxy variables, the GPC model provides a novel framework for causal inference based on spatial observations, and it advances both the methodological and theoretical development of causality analysis in complex geographical systems.
Zuopei Zhang, Jinfeng Wang 0001
Int. J. Geogr. Inf. Sci.2
2023 Understanding and extending the geographical detector model under a linear regression framework
abstract
The Geographical Detector Model (GDM) is a popular statistical toolkit for geographical attribution analysis. Despite the striking resemblance of the q-statistic in GDM to the R-squared in linear regression models, their explicit connection has not yet been established. This study proves that the q-statistic reduces into the R-squared under a linear regression framework. Under linear regression and moderate-to-strong spatial autocorrelation, Monte Carlo simulation results show that the GDM tends to underestimate the importance of variables. In addition, an almost perfect power law relationship is present between the percentage bias and the degree of the spatial autocorrelations, indicating the presence of fast uplifting bias in response to increasing levels of spatial autocorrelations. We propose an integrated approach for variable importance quantification by bringing together the spatial econometrics model and the game theory based-Shapley value method. By applying our proposed methodology to a case study of land desertification in African, it is found human activity tends to affect land desertification both directly and indirectly. However, such effects appear to be underestimated or undistinguished in the classic GDM.
Guanpeng Dong, Jinfeng Wang 0001, Tonglin Zhang, Xiaoyu Meng, Dongyang Yang, Binbin Lu
Int. J. Geogr. Inf. Sci.3
2022 Spatial rough set-based geographical detectors for nominal target variables
Hexiang Bai, Deyu Li 0001, Jinfeng Wang 0001
Inf. Sci.4
2021 Space-time disease mapping by combining Bayesian maximum entropy and Kalman filter: the BME-Kalman approach
abstract
In this work, a synthesis of the Bayesian maximum entropy (BME) and the Kalman filter (KF) methods, which enhances their individual strengths and overcomes certain of their weaknesses for spatiotemporal mapping purposes, is proposed in a spatiotemporal disease mapping context. The proposed BME-Kalman synthesis allows BME to use information from both parametric regression modeling and KF estimation leading to enhanced knowledge bases. The BME-Kalman synthetic approach is used to study the space-time incidence mapping of the hand, foot and mouth disease (HFMD) in Shandong province (China) during the period May 1st, 2008 to March 19th, 2009. The results showed that the BME-Kalman approach exhibited very good regressive and predictive accuracies, maintained a very good performance even during low-incidence and extremely low-incidence periods, offered an improved description of hierarchical disease characteristics compared to traditional mapping techniques, and provided a clear explanation of the spatial stratified incidence heterogeneity at unsampled locations. The BME-Kalman approach is versatile and flexible so that it can be modified and adjusted according to the needs of the application.
Bisong Hu, Pan Ning, Yi Li 0024, Chengdong Xu, George Christakos, Jinfeng Wang 0001
Int. J. Geogr. Inf. Sci.6
2020 Incorporating spatial association into statistical classifiers: local pattern-based prior tuning
abstract
This paper proposes a new classification method for spatial data by adjusting prior class probabilities according to local spatial patterns. First, the proposed method uses a classical statistical classifier to model training data. Second, the prior class probabilities are estimated according to the local spatial pattern and the classifier for each unseen object is adapted using the estimated prior probability. Finally, each unseen object is classified using its adapted classifier. Because the new method can be coupled with both generative and discriminant statistical classifiers, it performs generally more accurately than other methods for a variety of different spatial datasets. Experimental results show that this method has a lower prediction error than statistical classifiers that take no spatial information into account. Moreover, in the experiments, the new method also outperforms spatial auto-logistic regression and Markov random field-based methods when an appropriate estimate of local prior class distribution is used.
Hexiang Bai, Peter M. Atkinson, Qian Chen 0023, Jinfeng Wang 0001
Int. J. Geogr. Inf. Sci.5
2020 Spatial interpolation of marine environment data using P-MSN
abstract
When a marine study area is large, the environmental variables often present spatially stratified non-homogeneity, violating the spatial second-order stationary assumption. The stratified non-homogeneous surface can be divided into several stationary strata with different means or variances, but still with close relationships between neighboring strata. To give the best linear-unbiased estimator for those environmental variables, an interpolated version of the mean of the surface with stratified non-homogeneity (MSN) method called point mean of the surface with stratified non-homogeneity (P-MSN) was derived. P-MSN distinguishes the spatial mean and variogram in different strata and borrows information from neighboring strata to improve the interpolation precision near the strata boundary. This paper also introduces the implementation of this method, and its performance is demonstrated in two case studies, one using ocean color remote sensing data, and the other using marine environment monitoring data. The predictions of P-MSN were compared with ordinary kriging, stratified kriging, kriging with an external drift, and empirical Bayesian kriging, the most frequently used methods that can handle some extent of spatial non-homogeneity. The results illustrated that for spatially stratified non-homogeneous environmental variables, P-MSN outperforms other methods by simultaneously improving interpolation precision and avoiding artificially abrupt changes along the strata boundaries.
Bingbo Gao, Mao-Gui Hu, Jinfeng Wang 0001, Chengdong Xu, Hai-Mei Fan, Haiyuan Ding
Int. J. Geogr. Inf. Sci.3
2019 A spatial heterogeneity-based rough set extension for spatial data
abstract
When classical rough set (CRS) theory is used to analyze spatial data, there is an underlying assumption that objects in the universe are completely randomly distributed over space. However, this assumption conflicts with the actual situation of spatial data. Generally, spatial heterogeneity and spatial autocorrelation are two important characteristics of spatial data. These two characteristics are important information sources for improving the modeling accuracy of spatial data. This paper extends CRS theory by introducing spatial heterogeneity and spatial autocorrelation. This new extension adds spatial adjacency information into the information table. Many fundamental concepts in CRS theory, such as the indiscernibility relation, equivalent classes, and lower and upper approximations, are improved by adding spatial adjacency information into these concepts. Based on these fundamental concepts, a new reduct and an improved rule matching method are proposed. The new reduct incorporates spatial heterogeneity in selecting the feature subset which can preserve the local discriminant power of all features, and the new rule matching method uses spatial autocorrelation to improve the classification ability of rough set-based classifiers. Experimental results show that the proposed extension significantly increased classification or segmentation accuracy, and the spatial reduct required much less time than classical reduct.
Hexiang Bai, Deyu Li 0001, Jinfeng Wang 0001
Int. J. Geogr. Inf. Sci.4
2016 Driving forces and their interactions of built-up land expansion based on the geographical detector - a case study of Beijing, China
abstract
Scientific interpretation of the driving forces of built-up land expansion is essential to urban planning and policy-making. In general, built-up land expansion results from the interactions of different factors, and thus, understanding the combined impacts of built-up land expansion is beneficial. However, previous studies have primarily been concerned with the separate effect of each driver, rather than the interactions between the drivers. Using the built-up land expansion in Beijing from 2000 to 2010 as a study case, this research aims to fill this gap. A spatial statistical method, named the geographical detector, was used to investigate the effects of physical and socioeconomic factors. The effects of policy factors were also explored using physical and socioeconomic factors as proxies. The results showed that the modifiable areal unit problem existed in the geographical detector, and 4000 m might be the optimal scale for the classification performed in this study. At this scale, the interactions between most factors enhanced each other, which indicated that the interactions had greater effects on the built-up land expansion than any single factor. In addition, two pairs of nonlinear enhancement, the greatest enhancement type, were found between the distance to rivers and two socioeconomic factors: the total investment in fixed assets and GDP. Moreover, it was found that the urban plans, environmental protection policies and major events had a great impact on built-up land expansion. The findings of this study verify that the geographical detector is applicable in analysing the driving forces of built-up land expansion. This study also offers a new perspective in researching the interactions between different drivers.
Hongrun Ju, Zengxiang Zhang, Lijun Zuo, Jinfeng Wang 0001, Shengrui Zhang, Xiao Wang 0022
Int. J. Geogr. Inf. Sci.4
2016 Detecting nominal variables' spatial associations using conditional probabilities of neighboring surface objects' categories
Hexiang Bai, Deyu Li 0001, Jinfeng Wang 0001
Inf. Sci.4
2015 A stratified optimization method for a multivariate marine environmental monitoring network in the Yangtze River estuary and its adjacent sea
abstract
An efficient monitoring network is very important in accessing the marine environmental quality and its protection and management. In an estuary, there are fronts that separate distinctly different water masses and affect material transport, nutrient distribution, pollutant aggregation, and diffusion. This stratified heterogeneous surface neither satisfies the stationary requirements of kriging, nor can be handled adequately by removing a spatially continuous trend. This article presents a stratified optimization method for a multivariate monitoring network. In this method, principal component analysis (PCA) was used to reduce the dimensionality of the correlated targets, and the mean of surface with nonhomogeneity (MSN) method was adopted to produce the best linear unbiased estimator for a spatially stratified heterogeneous surface that failed to satisfy the requirements for a kriging estimate. The existing monitoring network in the Yangtze River estuary and its adjacent sea, which was designed by purposive sampling year ago was optimized as an illustration. The optimization consisted of two steps: reduce the redundant monitoring sites and then optimally add new sites to the remaining sites. After optimization, the inclusion of 51 sites in the monitoring network was found to produce a smaller total estimated error than that of the current network, which has 70 sites; moreover, the use of 55 sites can produce a higher precision of estimation for all three principal components (PCs) than that of the current 70 sites. The results demonstrated that the proposed method is suitable for optimizing environmental monitoring sites that have dominant stratified nonhomogeneity and that involve multiple factors.
Bingbo Gao, Jinfeng Wang 0001, Hai-Mei Fan, Kan Xu, Mao-Gui Hu
Int. J. Geogr. Inf. Sci.2
2015 Sampling design optimization of a wireless sensor network for monitoring ecohydrological processes in the Babao River basin, China
abstract
Optimal selection of observation locations is an essential task in designing an effective ecohydrological process monitoring network, which provides information on ecohydrological variables by capturing their spatial variation and distribution. This article presents a geostatistical method for multivariate sampling design optimization, using a universal cokriging (UCK) model. The approach is illustrated by the design of a wireless sensor network (WSN) for monitoring three ecohydrological variables (land surface temperature, precipitation and soil moisture) in the Babao River basin of China. After removal of spatial trends in the target variables by multiple linear regression, variograms and cross-variograms of regression residuals are fit with the linear model of coregionalization. Using weighted mean UCK variance as the objective function, the optimal sampling design is obtained using a spatially simulated annealing algorithm. The results demonstrate that the UCK model-based sampling method can consider the relationship of target variables and environmental covariates, and spatial auto- and cross-correlation of regression residuals, to obtain the optimal design in geographic space and attribute space simultaneously. Compared with a sampling design without consideration of the multivariate (cross-)correlation and spatial trend, the proposed sampling method reduces prediction error variance. The optimized WSN design is efficient in capturing spatial variation of the target variables and for monitoring ecohydrological processes in the Babao River basin.
J. H. Wang, Gerard B. M. Heuvelink, Xin Li 0029, Jinfeng Wang 0001
Int. J. Geogr. Inf. Sci.6
2010 Using rough set theory to identify villages affected by birth defects: the example of Heshun, Shanxi, China
abstract
This article uses rough set theory to explore spatial decision rules in neural-tube birth defects and searches for novel spatial factors related to the disease. The whole rule induction process includes data transformation, searching for attribute reducts, rule generation, prediction or classification, and accuracy assessment. We use Heshun as an example, where neural-tube birth defects are prevalent, to validate the approach. About 50% of the villages in Heshun are used as the sample data, from which all of the rules are extracted. Meanwhile, the other villages are used as reference data. The rules extracted from the training data are then applied to the reference data. The result shows that the rules' generalization is reasonably good. Moreover, a novel relationship between the spatial attributes and the neural-tube birth defects was discovered. That is, the villages that lie in Watershed 9 of this district and that are also associated with a gradient of between 16° and 25° are vulnerable to neural-tube birth defects. This result paves the road for predicting where high rates of neural-tube birth defects will occur and can be used as a preliminary step in finding a direct cause for the disease.
Hexiang Bai, Jinfeng Wang 0001, Yi-Lan Liao
Int. J. Geogr. Inf. Sci.3
2010 Using spatial analysis and Bayesian network to model the vulnerability and make insurance pricing of catastrophic risk
abstract
Vulnerability refers to the degree of an individual subject to the damage arising from a catastrophic disaster. It is affected by multiple indicators that include hazard intensity, environment, and individual characteristics. The traditional area aggregate approach does not differentiate the individuals exposed to the disaster. In this article, we propose a new solution of modeling vulnerability. Our strategy is to use spatial analysis and Bayesian network (BN) to model vulnerability and make insurance pricing in a spatially explicit manner. Spatial analysis is employed to preprocess the data, for example kernel density analysis (KDA) is employed to quantify the influence of geo-features on catastrophic risk and relate such influence to spatial distance. BN provides a consistent platform to integrate a variety of indicators including those extracted by spatial analysis techniques to model uncertainty of vulnerability. Our approach can differentiate attributes of different individuals at a finer scale, integrate quantitative indicators from multiple-sources, and evaluate the vulnerability even with missing data. In the pilot study case of seismic risk, our approach obtains a spatially located result of vulnerability and makes an insurance price at a finer scale for the insured buildings. The result obtained with our method is informative for decision-makers to make a spatially located planning of buildings and allocation of resources before, during, and after the disasters.
Lianfa Li, Jinfeng Wang 0001, Hareton K. N. Leung
Int. J. Geogr. Inf. Sci.2
2010 Integration of GP and GA for mapping population distribution
abstract
Mapping population distribution is an important field of geographical and related research because of the frequent need to combine spatial data representing socio‐demographic information across various incompatible spatial units. However, the research may become very complex and difficult when a population in multiple places is estimated by various factors. Previous efforts in the field have contributed to the selection of appropriate independent variables and the creation of different population models. However, the level of accuracy obtainable with these studies is limited by the spatial heterogeneity of population distribution within the individual census districts, particularly in large rural areas. A high‐accuracy modelling method for population estimation based on integration of Genetic Programming (GP) and Genetic Algorithms (GA) with Geographic Information Systems (GIS) is presented in this paper. GIS was applied to identify and quantify a set of natural and socioeconomic factors which contributed to population distribution, and then GP and GA were used to build and optimise the population model to automatically transform census population data to regular grids. The study indicated that the proposed method performed much better than the stepwise regression analysis and adapted gravity model methods in estimating the population of both urban and rural areas. More importantly, this proposed method could provide a single, unified approach to mapping population distribution in various areas because the paradigms of these algorithms are general.
Yilan Liao, Jinfeng Wang 0001, Xinhu Li
Int. J. Geogr. Inf. Sci.2
2010 Sample surveying to estimate the mean of a heterogeneous surface: reducing the error variance through zoning
abstract
One of the major sources of uncertainty associated with geographical data in GIS arises when they are the outcome of a sampling process. It is well known that when sampling from a spatially autocorrelated homogeneous surface, stratification reduces the error variance of the estimator of the population mean. In this study, we evaluate the efficiency of different spatial sampling strategies when the surface is not homogeneous. When the surface is first-order heterogeneous (the mean of the surface varies across the map), we examine the effects of stratifying it into first-order homogeneous zones prior to the usual stratification for a systematic or stratified random sample. We investigate the effect of this form of spatial heterogeneity on the performance of different methods for estimating the population mean and its error variance. We do so by distinguishing between the real surface to be surveyed (ℜ), the sampling frame (ℑ) including the choice of zoning, and the statistical estimators (Ψ). The study shows that zoning improves estimator efficiency when sampling a heterogeneous surface. Systematic comparison provides rules of thumb for choice of sample design, sample statistics and uncertainty estimation, based on considering different spatial heterogeneities on real surfaces.
Jinfeng Wang 0001, Robert Haining, Zhidong Cao
Int. J. Geogr. Inf. Sci.1
2010 Geographical Detectors-Based Health Risk Assessment and its Application in the Neural Tube Defects Study of the Heshun Region, China
abstract
Physical environment, man‐made pollution, nutrition and their mutual interactions can be major causes of human diseases. These disease determinants have distinct spatial distributions across geographical units, so that their adequate study involves the investigation of the associated geographical strata. We propose four geographical detectors based on spatial variation analysis of the geographical strata to assess the environmental risks of health: the risk detector indicates where the risk areas are; the factor detector identifies factors that are responsible for the risk; the ecological detector discloses relative importance between the factors; and the interaction detector reveals whether the risk factors interact or lead to disease independently. In a real‐world study, the primary physical environment (watershed, lithozone and soil) was found to strongly control the neural tube defects (NTD) occurrences in the Heshun region (China). Basic nutrition (food) was found to be more important than man‐made pollution (chemical fertilizer) in the control of the spatial NTD pattern. Ancient materials released from geological faults and subsequently spread along slopes dramatically increase the NTD risk. These findings constitute valuable input to disease intervention strategies in the region of interest.
Jinfeng Wang 0001, Xinhu Li, George Christakos, Yi-Lan Liao, Tin Zhang, Xue Gu
Int. J. Geogr. Inf. Sci.1
2005 Typhoon insurance pricing with spatial decision support tools
abstract
In disaster insurance and reinsurance, GIS has been used to visualize and manage geospatial data and to help vulnerability and risk analysis for years. However, hazard insurance is a multidisciplinary issue that involves complex factors and uncertainty. GIS, if used alone, has limited functionality due to poor incorporation of intelligence and spatial statistics. The Spatial Decision Support System (SDSS) presented in this paper, addresses some of the deficiencies of traditional GIS, by providing powerful tools to support disaster insurance pricing that involves procedural and declarative knowledge. In the SDSS, the knowledge‐based system shell, using the open‐source CLIPS and supporting fuzziness and uncertainty, can be applied in at least three phases: hazard simulation, fuzzy comprehensive evaluation of risk, and query for insurance pricing. The libraries of statistics and spatial statistics provide a robust support for analysis of spatial factors, including spatial correlation between zones vulnerable to hazard and spatial variation of exposures. The GIS components provide sophisticated visualization and database management support for geospatial data, helping easily locate the insured points and risk zones as well as exploratory analysis of spatial data. Standard database management interfaces are used to manage other aspatial data. COM, an industry‐wide interface protocol, tightly integrates these technologies (the expert shell, GIS, spatial statistics and DBM within an integral system), and can be used to develop mixed complex algorithms in support of other COM objects. An application of typhoon insurance pricing is demonstrated with a case study in Guangdong, China. Developed as a suite of generic tools with abilities to deal with the complex problem of disaster insurance involving spatial factors and field knowledge, this prototype SDSS can also be applied to other disaster insurance and fields that involve similar spatial decision making.
Lianfa Li, Jinfeng Wang 0001, Chengyi Wang 0001
Int. J. Geogr. Inf. Sci.2