Feng Lu 0004

dblp:97/6483-4 · DBLP profile ↗
← Back
21ranked-venue papers in the field
0as first author
17since 2021 · last 2026
0000-0001-6573-2550ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 16Information Retrieval & Web Search · 5
YearPublicationVenuePosition
2026 A geographic evolutionary framework with multi-task optimization of automatic hyperparameter tuning for spatially stratified machine learning models
abstract
The rapid growth of spatial data across various fields has made data-driven machine learning (ML) methods increasingly essential for spatial analysis. The performance of ML methods heavily depends on hyperparameter tuning (HPT), which becomes particularly challenging when dealing with spatial stratified heterogeneity (SSH) that causes significant variability in statistical characteristics across sub-regions and demands separate local models. Traditional HPT methods either apply uniform hyperparameters or treat each model independently, ignoring spatial associations between adjacent models. To address this gap, we propose a geographic evolutionary (GeoEvo) framework with multi-task optimization to account for spatial associations in collaborative HPT of local ML models under SSH. GeoEvo formulates the problem of a single objective optimization with multi-task constraints to jointly optimize hyperparameters across multiple local models. The framework introduces an evolution operator for geographic proximity differential (Geo-DE) to enable collaboration among spatially adjacent models and a geographic selection operator (Geo-SL) to promote hyperparameter sharing, improve resource utilization and accelerate convergence. Extensive experiments based on soil organic carbon stocks and PM2.5 concentration datasets demonstrated that GeoEvo enhanced accuracy and stability while maintaining computational efficiency. Subsequent visual analytics revealed that considering the diversity and spatial correlations within the hyperparameter space enhanced the optimization process and improved prediction accuracy.
Shifen Cheng, Feng Lu 0004
Int. J. Geogr. Inf. Sci.3
2026 Predicting human-activity intensity in urban areas with a prior-enhanced probabilistic-deterministic model
abstract
Although numerous models have been proposed to predict the intensity of human activities in urban areas, two major issues hamper the performance of existing models: (1) fail to incorporate appropriate prior knowledge instrumental for improving accuracy and interpretability; (2) fail to integrate probabilistic and deterministic predictions to achieve complementary strengths, namely uncertainty quantification and high predictive accuracy. To address these challenges, we proposed a prior-enhanced dual-mode spatiotemporal graph neural network (PED-STGNN) to support both probabilistic and deterministic predictions. Specifically, we introduced a hypergraph node-to-vector (hypernode2vec) method to capture the multivariate functional similarity prior derived from complex and multivariate relations between urban regions. This functional similarity characterizes urban systems more precisely than existing methods relying on first-order pairwise relations. It improves accuracy and interpretability while enabling spatial modeling of higher-order multivariate relations beyond first-order pairwise relations. We also designed a plug-and-play probabilistic prediction module that enables switches between probabilistic and deterministic modes. Experiments based on the human activity intensity in Fuzhou, China, demonstrated the advantages in accuracy, interpretability and multi-scenario applicability.
Sheng Wu 0004, Peixiao Wang, Hengcai Zhang, Shifen Cheng, Feng Lu 0004
Int. J. Geogr. Inf. Sci.6
2026 Unveiling heterogeneity in tourist-generated content quality: A large language model approach with multi-dimensional analysis
abstract
The rapid growth of tourist-generated content demands scalable and reliable quality assessment methods. This study introduces an LLM-driven framework that combines parameter-efficient fine-tuning and prompt engineering to evaluate content quality accurately and interpretably. Applied to 484,930 reviews from MaFengWo, TripAdvisor, and Ctrip, the approach achieves superior performance (RMSE=0.3040, NDCG@100=0.500, BERTScore=77.95%) with 9 × higher efficiency. Spatial-temporal-semantic analyses reveal platform-specific quality patterns: MaFengWo exhibits prominent spatial centrality and stable temporal cointegration; TripAdvisor demonstrates simplified core-periphery structures with high volatility; Ctrip presents dynamic multicentricity particularly in Shanghai. Two domestic platforms, MaFengWo and Ctrip, expose systematic deficiency on the theme ‘ Decision-making Plan ’ (92.9∼96.4% lacking operational suggestions), while international TripAdvisor emphasizes ‘ Practical Information ’ and ‘ Consumption Activity ’ but 40.76% neglects original viewpoints. Heterogeneous network analysis identifies the behavioral signatures of high-reliability user—preference attachment, quality stability, and profile homogeneity. This work bridges theoretical rigor with operational scalability, demonstrating the potential of LLMs in content governance for digital tourism.
Jialiang Gao, Lizhu Chen, Peng Peng 0007, Yang Xu 0054, Feng Lu 0004, Christophe Claramunt
Inf. Process. Manag.7
2026 Adaptive model selection and ensemble via spatiotemporal graph-guided expert routing
Lizeng Wang, Shifen Cheng, Feng Lu 0004
Inf. Process. Manag.3
2026 Structure-aware multi-view urban representation learning with coordinated fusion and alignment
Jinghui Wei, Sheng Wu 0004, Shifen Cheng, Peixiao Wang, Feng Lu 0004
Inf. Process. Manag.5
2025 STAGE: a spatiotemporal-knowledge enhanced multi-task generative adversarial network (GAN) for trajectory generation
abstract
Individual trajectory data play a pivotal role in various application fields, such as urban planning, traffic control, and epidemic simulation. Despite the diverse means for data collection in current times, the real-world trajectory data in practical application remains severely limited due to concerns over personal privacy. In this study, we designed a Spatiotemporal-knowledge enhanced multi-TAsk GEnerative adversarial network (GAN), named STAGE, to generate synthetic trajectories that statistically resemble the real data without recycling personal information. In STAGE, we designed a multi-task generator with three stages of spatio-temporal generation tasks, i.e. activity-sequence generation task, township-level trajectory generation task, and neighborhood-level trajectory generation task, with the last one as the main task while the other two as auxiliary tasks. Meanwhile, we designed a spatial consistency loss in the adversarial training process to assess the spatial consistency of generated trajectories at different spatial scales. Experiment results show that compared to the baselines, trajectories generated by our method have closer data distributions to the real ones. We argued that the designs of spatiotemporal-knowledge enhanced generation tasks and training loss benefit the spatiotemporal generation processes, which help reproduce the temporal patterns of human daily activities and spatial distribution of human movements.
Zhongcai Cao, Kang Liu 0010, Li Ning 0001, Ling Yin 0001, Feng Lu 0004
Int. J. Geogr. Inf. Sci.6
2025 An explainable spatial interpolation method considering spatial stratified heterogeneity
abstract
Spatial interpolation is essential for handling sparsity and missing spatial data. Current machine learning-based spatial interpolation methods are subject to the statistical constraints of spatial stratified heterogeneity (SSH), normally involving separate modeling of each stratum and simple weighted averaging to integrate intra-stratum and inter-strata features. However, these models overlook the different contributions of inter-strata features to different locations within a stratum (heterogeneous inter-strata associations, HIA) and the explanation of spatial effects on the interpolation process, leading to suboptimal and unreliable interpolation outcomes. This article proposes a novel explainable spatial interpolation method considering SSH (X-SSHM). Spatial and environmental features are utilized to describe intra-stratum and inter-strata information, which are fed into random forest-based learners to achieve high-level semantic feature mapping. Geographically weighted regression is employed to integrate intra-stratum and inter-strata features to achieve a unified expression of SSH and HIA, obtaining the final interpolation result. Geographically weighted Shapley (GSHAP) is proposed to decompose the marginal contributions of intra-stratum and inter-strata features. Model performance is evaluated on simulated and soil organic matter datasets. X-SSHM outperformed five baselines regarding interpolation accuracy. Moreover, statistical methods validated X-SSHM’s ability to elucidate the mechanisms by which SSH, spatial autocorrelation and HIA affect the model interpolation process.
Shifen Cheng, Lizeng Wang, Feng Lu 0004
Int. J. Geogr. Inf. Sci.5
2025 A tensor decomposition method based on embedded geographic meta-knowledge for urban traffic flow imputation
abstract
Accurate and reliable traffic flow data are essential for intelligent transportation systems; however, limitations arising from hardware and communication costs often lead to missing data. Tensor decomposition is widely used to address these issues. However, existing imputation methods employ a fixed geographic feature similarity matrix to constrain the tensor decomposition process, which fails to accurately capture the spatial heterogeneity of traffic flows, thus limiting the imputation accuracy and robustness. This study proposes a tensor decomposition method embedded with geographic meta-knowledge (Meta-TD) to accurately determine the spatial heterogeneity of traffic flows. The key innovation is establishing a dynamic relationship between the geographic meta-knowledge and spatial heterogeneity of traffic flows, and then using the spatial heterogeneity of the traffic flows to constrain the tensor decomposition process. Experimental results based on real urban traffic flows demonstrated the superiority of Meta-TD over fifteen baseline models under random, block, and long time-series missing patterns, achieving reductions in MAE, RMSE, and MAPE of 6.97–97.05%, 3.33–94.68%, and 0.72–90.89%, respectively. Notably, Meta-TD maintained high accuracy for sudden changes in traffic flow states, evidencing its robustness to varying missing data rates and distribution patterns. This adaptability makes it highly suitable for complex and dynamic urban traffic environments.
Xiaoyue Luo, Shifen Cheng, Lizeng Wang, Yuxuan Liang 0002, Feng Lu 0004
Int. J. Geogr. Inf. Sci.5
2025 Efficient inference of large-scale air quality using a lightweight ensemble predictor
abstract
Accurate and efficient air quality prediction is crucial for public health protection and environmental sustainability. While numerous grid-based and graph-based prediction models have been developed, they encounter challenges in large-scale scenarios: (1) Grid-based models, though computationally efficient, have limited prediction accuracy in large-scale sparse scenarios; (2) Graph-based models, despite higher prediction accuracy, suffer from significant computational inefficiencies when dealing with a large number of sensors, i.e. graph nodes. To address these issues, we propose a Lightweight Ensemble Predictor (LiEnPred) for efficient air quality prediction in large-scale sparse scenarios. First, we present a data structure transformation algorithm that converts sparse monitoring sensors from graph structures to compact grid structures, preserving the connections between graph nodes. Next, we present a lightweight parameter-shared spatio-temporal dilation convolution network that efficiently captures spatio-temporal dependencies in air quality data without significantly increasing computation time or parameter scale. In our experiments, we collected air quality data from over 2000 sensors across China over the past three years and evaluated LiEnPred’s prediction performance in large-scale scenarios using PM2.5 and NO2 concentration data. The experimental results demonstrate that the proposed LiEnPred model matches or exceeds the predictive accuracy of eight baselines with faster time efficiency and fewer model parameters.
Peixiao Wang, Hengcai Zhang, Feng Lu 0004, Tong Zhang 0009
Int. J. Geogr. Inf. Sci.4
2024 An ensemble spatial prediction method considering geospatial heterogeneity
abstract
Ensemble learning synthesizes the advantages of different models and has been widely applied in the field of spatial prediction. However, the nonlinear constraints of spatial heterogeneity on the model ensemble process make it difficult to adaptively determine the ensemble weights, greatly limiting the predictive ability of the ensemble learning model. This paper therefore proposes a novel geographical spatial heterogeneous ensemble learning method (GSH-EL). Firstly, the geographically weighted regression model, geographically optimal similarity model, and random forest model are used as three base learners to express local spatial heterogeneity, global feature correlation, and nonlinear relationship of geographic elements, respectively. Then, a spatially weighted ensemble neural network module (SWENN) of GSH-EL is proposed to express spatial heterogeneity by exploring the complex nonlinear relationship between the spatial proximity and ensemble weights. Finally, the outputs of the three base learners are combined with the spatial heterogeneous ensemble weights from SWENN to obtain the spatial prediction results. The proposed method is validated on the PM2.5 air quality and landslide dataset in China, both of which obtain more accurate prediction results than the existing ensemble learning strategies. The results confirm the need to accurately express spatial heterogeneity in the model ensemble process.
Shifen Cheng, Lizeng Wang, Peixiao Wang, Feng Lu 0004
Int. J. Geogr. Inf. Sci.4
2024 Simulating human mobility with a trajectory generation framework based on diffusion model
abstract
Most mobility modeling methods are designed to solve specific tasks, leading to questions regarding their deficiency in generalizability. Inspired by the bloom of foundation models, we proposed a Trajectory Generation framework based on the Diffusion Model (TrajGDM) to capture the universal mobility pattern in a trajectory dataset by learning the trajectory generation process. The process is modeled as a step-by-step uncertainty-reducing process, in which a deep learning network with a novel training method is proposed to learn from the process. We compared the proposed trajectory generation method with six baselines on two public trajectory datasets. The results showed that the similarity between the generated and real trajectory movements measured by the Jensen-Shannon Divergence improved significantly on both datasets. Moreover, we applied zero-shot inferences on two basic trajectory tasks: trajectory prediction and trajectory reconstruction. The accuracy improved by a maximum of 25.6% on two tasks. The universal mobility pattern that is suitable for solving multiple trajectory tasks is verified, inferring the strong generalizability of our model. Finally, the study provides insights into artificial intelligence’s understanding of human mobility by exploring the way the model maps the trajectory in the latent space into reality.
Chen Chu, Hengcai Zhang, Peixiao Wang, Feng Lu 0004
Int. J. Geogr. Inf. Sci.4
2024 Act2Loc: a synthetic trajectory generation method by combining machine learning and mechanistic models
abstract
Human mobility data play a crucial role in many fields such as infectious diseases, transportation, and public safety. Although the development of Information and Communication Technologies (ICTs) has made it easy to collect individual-level positioning records, raw individual trajectory data are still limited in availability and usability due to privacy issues. Developing models to generate synthetic trajectories that are statistically close to the real data is a promising solution. This study proposed a novel trajectory generation method called Act2Loc (Activity to Location), which combined machine learning and mechanistic models. First, an activity-sequence generation model was constructed based on machine learning models (i.e. K-medoids and Transformer) to generate individual activity sequences aligning with human activity patterns. Then, a spatial-location selection model was proposed based on mechanistic models (e.g. Universal Opportunity model) to explicitly determine the specific locations of the activities in each generated sequence. Experimental results showed that compared to baselines based on purely machine learning or mechanistic models, Act2Loc can better reproduce the spatio-temporal characteristics of the real data, with additional advantage of low data requirements for training, proving its potential for generating synthetic trajectories in practice. This research offers new insights on knowledge-guided GeoAI models for human mobility.
Kang Liu 0010, Shifen Cheng, Song Gao 0001, Ling Yin 0001, Feng Lu 0004
Int. J. Geogr. Inf. Sci.6
2024 Identifying the cargo types of road freight with semi-supervised trajectory semantic enhancement
abstract
Identifying road freight cargo types is crucial for regional economic interaction and transportation optimization. Existing methods primarily rely on manual labeling and the rule, neither of which can achieve automated semantic enhancement of large-scale road freight trajectories. Consequently, this study proposes a semi-supervised trajectory semantic enhancement method for identifying cargo types based on trajectory feature extraction and point-of-interest (POI) association. The raw trajectories are segmented and enriched with the closest POIs. The sample labeling method with POI semantic enhancement is then proposed using company registration information. Finally, the spatiotemporal and sequential features of labeled freight trips are extracted to build a self-training semi-supervised model for identifying the cargo type of road freight. Experimental studies on real trajectory data demonstrate superior accuracy and robustness compared to existing methods, with accuracy and F1 values reaching 81.4 and 0.77%, respectively. The proposed sample labeling method improves representativeness and universality, increasing accuracy by 7.8–14.4% and F1 value by 8.5–34.5% compared to the rule-based method. The semi-supervised model improves accuracy by 8.9% and F1 value by 29.1% compared to the supervised model when only 10.0% of samples were labeled. This method enables automatic and full-sample cargo type identification in real-world large-scale transportation systems.
Shifen Cheng, Beibei Zhang 0002, Feng Lu 0004
Int. J. Geogr. Inf. Sci.4
2024 Mining tourist preferences and decision support via tourism-oriented knowledge graph
Jialiang Gao, Peng Peng 0007, Feng Lu 0004, Christophe Claramunt, Peiyuan Qiu, Yang Xu 0054
Inf. Process. Manag.3
2023 TrajGDM: A New Trajectory Foundation Model for Simulating Human Mobility
abstract
Capturing the universal movement pattern and simulating human mobility is one of the most important trajectory data-mining tasks. Most of the current mobility modeling methods are specially designed to solve a specific task, which leads to questions regarding generalizability. Aiming to construct a general trajectory foundation model to overcome this weakness, we proposed a generative Trajectory Generation framework based on Diffusion Model (TrajGDM) to capture the universal mobility pattern and simulate human mobility. It is capable of solving multiple trajectory tasks through learning the generation of the trajectory. The generation process of a trajectory is modeled as a step-by-step uncertainty reducing process. A trajectory generator network is proposed to estimate the uncertainty in each step, and a trajectory diffusion and generation process is defined to train the model to simulate the real dataset. Finally, we compared the proposed method with 6 baselines on 2 public trajectory datasets: T-Drive and Geo-life. By comparing 5 different evaluation metrics, the result showed that the similarity between generated and real trajectories' movement character measured by Jensen-Shannon Divergence (JSD) improved by at least 50.3% in both datasets. It also addresses the problem of generating diverse trajectories, which is ignored by most previous models. Moreover, we applied zero-shot inferences on two basic trajectory tasks: trajectory prediction and trajectory reconstruction. The zero-shot prediction accuracy of our model is up to 23.4% higher than the benchmark, and the reconstruction accuracy improves by a maximum of 25.6%.
Chen Chu, Hengcai Zhang, Feng Lu 0004
SIGSPATIAL/GIS3
2023 Towards travel recommendation interpretability: Disentangling tourist decision-making process via knowledge graph
Jialiang Gao, Peng Peng 0007, Feng Lu 0004, Christophe Claramunt, Yang Xu 0054
Inf. Process. Manag.3
2021 Prediction of human activity intensity using the interactions in physical and social spaces through graph convolutional networks
abstract
Dynamic human activity intensity information is of great importance in many location-based applications. However, two limitations remain in the prediction of human activity intensity. First, it is hard to learn the spatial interaction patterns across scales for predicting human activities. Second, social interaction can help model the activity intensity variation but is rarely considered in the existing literature. To mitigate these limitations, we proposed a novel dynamic activity intensity prediction method with deep learning on graphs using the interactions in both physical and social spaces. In this method, the physical interactions and social interactions between spatial units were integrated into a fused graph convolutional network to model multi-type spatial interaction patterns. The future activity intensity variation was predicted by combining the spatial interaction pattern and the temporal pattern of activity intensity series. The method was verified with a country-scale anonymized mobile phone dataset. The results demonstrated that our proposed deep learning method with combining graph convolutional networks and recurrent neural networks outperformed other baseline approaches. This method enables dynamic human activity intensity prediction from a more spatially and socially integrated perspective, which helps improve the performance of modeling human dynamics.
Mingxiao Li 0001, Song Gao 0001, Feng Lu 0004, Kang Liu 0010, Hengcai Zhang, Wei Tu 0001
Int. J. Geogr. Inf. Sci.3
2020 A lightweight ensemble spatiotemporal interpolation model for geospatial data
abstract
Missing data is a common problem in the analysis of geospatial information. Existing methods introduce spatiotemporal dependencies to reduce imputing errors yet ignore ease of use in practice. Classical interpolation models are easy to build and apply; however, their imputation accuracy is limited due to their inability to capture spatiotemporal characteristics of geospatial data. Consequently, a lightweight ensemble model was constructed by modelling the spatiotemporal dependencies in a classical interpolation model. Temporally, the average correlation coefficients were introduced into a simple exponential smoothing model to automatically select the time window which ensured that the sample data had the strongest correlation to missing data. Spatially, the Gaussian equivalent and correlation distances were introduced in an inverse distance-weighting model, to assign weights to each spatial neighbor and sufficiently reflect changes in the spatiotemporal pattern. Finally, estimations of the missing values from temporal and spatial were aggregated into the final results with an extreme learning machine. Compared to existing models, the proposed model achieves higher imputation accuracy by lowering the mean absolute error by 10.93 to 52.48% in the road network dataset and by 23.35 to 72.18% in the air quality station dataset and exhibits robust performance in spatiotemporal mutations.
Shifen Cheng, Peng Peng 0007, Feng Lu 0004
Int. J. Geogr. Inf. Sci.3
2018 Fine-grained prediction of urban population using mobile phone location data
abstract
Fine-grained prediction of urban population is of great practical significance in many domains that require temporally and spatially detailed population information. However, fine-grained population modeling has been challenging because the urban population is highly dynamic and its mobility pattern is complex in space and time. In this study, we propose a method to predict the population at a large spatiotemporal scale in a city. This method models the temporal dependency of population by estimating the future inflow population with the current inflow pattern and models the spatial correlation of population using an artificial neural network. With a large dataset of mobile phone locations, the model’s prediction error is low and only increases gradually as the temporal prediction granularity increases, and this model is adaptive to sudden changes in population caused by special events.
Jie Chen 0077, Tao Pei, Shih-Lung Shaw, Feng Lu 0004, Mingxiao Li 0001, Shifen Cheng, Xiliang Liu, Hengcai Zhang
Int. J. Geogr. Inf. Sci.4
2016 Understanding the bias of call detail records in human mobility research
abstract
In recent years, call detail records (CDRs) have been widely used in human mobility research. Although CDRs are originally collected for billing purposes, the vast amount of digital footprints generated by calling and texting activities provide useful insights into population movement. However, can we fully trust CDRs given the uneven distribution of people’s phone communication activities in space and time? In this article, we investigate this issue using a mobile phone location dataset collected from over one million subscribers in Shanghai, China. It includes CDRs (~27%) plus other cellphone-related logs (e.g., tower pings, cellular handovers) generated in a workday. We extract all CDRs into a separate dataset in order to compare human mobility patterns derived from CDRs vs. from the complete dataset. From an individual perspective, the effectiveness of CDRs in estimating three frequently used mobility indicators is evaluated. We find that CDRs tend to underestimate the total travel distance and the movement entropy, while they can provide a good estimate to the radius of gyration. In addition, we observe that the level of deviation is related to the ratio of CDRs in an individual’s trajectory. From a collective perspective, we compare the outcomes of these two datasets in terms of the distance decay effect and urban community detection. The major differences are closely related to the habit of mobile phone usage in space and time. We believe that the event-triggered nature of CDRs does introduce a certain degree of bias in human mobility research and we suggest that researchers use caution to interpret results derived from CDR data.
Shih-Lung Shaw, Yang Xu 0002, Feng Lu 0004, Jie Chen 0077, Ling Yin 0001
Int. J. Geogr. Inf. Sci.4
2006 A Forced Transplant Algorithm for Dynamic R-tree Implementation
Mingbo Zhang, Feng Lu 0004, Changxiu Cheng
DEXA2