Hangli Ge

dblp:180/3320 · DBLP profile ↗
← Back
7ranked-venue papers in the field
2as first author
7since 2021 · last 2025
0000-0001-8025-3013ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 5 (2 first)Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2025 XCKAN: Federated Catalog for Data Discovery in Dataspaces
Hangli Ge, Hideaki Takeda 0001, Takeshi Sagara, Naho Kitano, Noboru Koshizuka
IEEE Big Data1
2025 Leveraging Visitor Mobility and IoT Sensor Networks for Sustainable Waste Management
Slamet Kristanto Tirto Utomo, Hangli Ge, Noboru Koshizuka
IEEE Big Data2
2025 Place with Intention: An Empirical Attendance Predictive Study of Expo 2025 Osaka, Kansai, Japan
Dizhi Huang, Hangli Ge, Masahiro Sano, Takeaki Ohdake, Kazuma Hatano, Noboru Koshizuka
IEEE Big Data3
2025 Robust and Efficient Human Mobility Data Processing through the Lens of Topological Persistence
abstract
Large-scale human mobility data (e.g. GPS data) encodes valuable information interested by various fields. Extracting stay and movement behaviors from noisy positioning record sequences is a critical preliminary step to utilize human mobility data. For the past two decades, this processing has been founded on a simple intuition proposed by Hariharan and Zheng et al.[18, 48], which uses manually selected parameters to make recursive, rule-based classification as to whether data points in a positioning record sequence constitute noise, move, or stay. This de facto processing approach, despite its simplicity, is inherently sensitive to parameter choice and suffers from the low efficiency of sequential processing. These inherent limitations make it practically infeasible, when confronted with the large-scale, fine-grained human mobility datasets in industry. To address this fundamental problem in human mobility data utilization, we innotatively rethink the distinction in representation patterns of noise/stay/move within the positioning record sequence from the lens of topological persistence, culminating in a novel pipeline for robust and efficient human mobility data processing. This is grounded in our empirical observation that topological persistence features of stay/move/noise exhibit robust and generalizable discriminability across variations in parameter choice, individuals, and geographical regions. By introducing the Laplacian to simplify the computation of topological persistence features, our processing pipeline is capable to exploit GPUs' parallel capacity for efficient processing. Experiments on real-world GPS datasets totaling up to thousand billion data points demonstrate that our method produces processing results comparable to those of human annotators, while requiring only 10% of the time consumed by previous approaches. We further show that our method is scalable for cumulative data volume and remains effective in identifying stay/move behaviors that traditional techniques consistently fail to handle, even under conditions of severe positioning errors and diverse behavioral patterns.
Lifeng Lin 0003, Hangli Ge, Takashi Michikata, Kazuma Hatano, Ryosuke Shibasaki, Noboru Koshizuka
SIGSPATIAL/GIS2
2025 CausalMob: Causal Human Mobility Prediction with LLMs-derived Human Intentions toward Public Events
abstract
Large-scale human mobility exhibits spatial and temporal patterns that can assist policymakers in decision making. Although traditional prediction models attempt to capture these patterns, they are often affected by nonperiodic public events, such as disasters and occasional celebrations. Since regular human mobility patterns are affected by these events, estimating their causal effects is critical to accurate mobility predictions. News articles provide unique perspectives on these events, though processing them is a challenge. In this study, we propose a causality based prediction model, CausalMob, to analyze the causal effects of public events. We first utilize large language models (LLMs) to extract human intentions from news and transform them into features that act as causal treatments. Next, the model learns representations of spatio-temporal regional covariates from multiple data sources to serve as confounders for causal inference. Finally, we present a causal effect estimation framework to ensure that event features remain independent of confounders during prediction. Based on large-scale real-world data, the experimental results show that the proposed model excels in human mobility prediction, outperforming state-of-the-art models.
Hangli Ge, Jiawei Wang 0005, Zipei Fan, Renhe Jiang, Ryosuke Shibasaki, Noboru Koshizuka
KDD (1)2
2024 FRTP: Federating Route Search Records to Enhance Long-term Traffic Prediction
abstract
Accurate traffic prediction, especially predicting traffic conditions several days in advance is essential for intelligent transportation systems (ITS). Such predictions enable mid- and long-term traffic optimization, which is crucial for efficient transportation planning. However, the inclusion of diverse external features, alongside the complexities of spatial relationships and temporal uncertainties, significantly increases the complexity of forecasting models. Additionally, traditional approaches have handled data preprocessing separately from the learning model, leading to inefficiencies caused by repeated trials of preprocessing and training. In this study, we propose a federated architecture capable of learning directly from raw data with varying features and time granularities or lengths. The model adopts a unified design that accommodates different feature types, time scales, and temporal periods. Our experiments focus on federating route search records and begin by processing raw data within the model framework. Unlike traditional models, this approach integrates the data federation phase into the learning process, enabling compatibility with various time frequencies and input/output configurations. The accuracy of the proposed model is demonstrated through evaluations using diverse learning patterns and parameter settings. The results show that online search log data is useful for forecasting long-term traffic, highlighting the model’s adaptability and efficiency.
Hangli Ge, Itsuki Matsunaga, Dizhi Huang, Noboru Koshizuka
IEEE Big Data1
2022 Traffic Congestion Prediction Using Toll and Route Search Log Data
abstract
Predicting future people’s behavior can significantly impact various industries. Intelligent transportation system (ITS) advancement, in particular, depends on the ability to predict traffic congestion. If we can do so, we can encourage people to alter their behavior, which reduces traffic congestion, traffic accidents, travel times, and CO2emissions while also promoting the development of applications like dynamic pricing. However, predicting traffic congestion a few days ahead is challenging owing to its spatial and temporal dependence and its nature of being susceptible to external factors, such as weather, local events, and the pandemic of infectious diseases. For these reasons, previous studies have been limited to predicting the next few minutes to a few hours. To address this limitation, we propose using search log data of the toll route search service owned by East Nippon Expressway Co., Ltd. (NEXCO East), which operates expressway services in Japan, as these data are available several days before the prediction and comprehensively explain multiple external factors. We show that search log data can contribute to predicting people’s behavior by verifying the improvement in the accuracy of traffic congestion prediction.
Yuto Kosugi, Itsuki Matsunaga, Hangli Ge, Takashi Michikata, Noboru Koshizuka
IEEE Big Data3