Xiaobai Angela Yao

dblp:224/0395 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-2719-2017ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 46% Representation and self-supervised learning · 23% Deep learning architectures and training · 23%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
0.812024
TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning · NeurIPS 2024
Machine learning › Trustworthy machine learning › dataset bias
geographic bias
0.812024
TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning · NeurIPS 2024
Machine learning › Deep learning architectures and training
positional encoding
0.812024
TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning · NeurIPS 2024
Machine learning › Representation and self-supervised learning
spatial representation learning
0.812024
TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 0.8diffusion transformer · 0.8
YearPublicationVenuePosition
2025 Exploring human mobility: a time-informed approach to pattern mining and sequence similarity
abstract
The surge in the availability of spatial big data has sparked increased interest in researching human mobility patterns. Despite this, discovering human mobility patterns from such spatial big data and assessing the similarity between patterns remains a formidable challenge. This study introduces two novel methods: the Time-Informed pattern mining (TiPam) method for frequent pattern mining and a Time-Aware Longest Common Subsequence (T-LCS) algorithm for assessing similarity between time-conscious sequences. Leveraging these innovative algorithms, our research introduces an analytical framework for analyzing human mobility patterns at both individual and aggregated levels. As a case study, this proposed workflow is applied to examine the daily mobility patterns of voluntary mobile phone users in Kampala, Uganda. The 135 participants are found in four distinct groups labeled with distinct mobility properties for users in each group: "stay-at-home," "unoccupied," "education-oriented," and "work-oriented." The results effectively showcase the efficiency of the framework and the novel techniques employed. The framework's versatility extends to human mobility studies with other forms of data and across various research fields.
Xiaobai Angela Yao, Christopher C. Whalen, Noah Kiwanuka
Int. J. Geogr. Inf. Sci.2
2024 SRL: Towards a General-Purpose Framework for Spatial Representation Learning
abstract
Representation learning (RL) techniques are widely adopted in areas such as natural language processing and computer vision, with prominent examples such as attention and ConvNet architectures. In comparison, many GeoAI works still rely on feature engineering or data conversion to represent spatial data (e.g., points, polylines, polygons, 3D building models, etc.) as features in formats that are easier for neural networks to handle. The neural network architectures remain unchanged, and the need for feature engineering has become a bottleneck for applying deep learning to new tasks in the age of big data. In this paper, we advocate the idea of developing learnable spatial representation modules, which not only enable spatial reasoning but also enable neural nets to directly consume (i.e., encoding) or generate (i.e., decoding) spatial data. We propose Spatial Representation Learning (SRL), a new general-purpose representation learning framework for spatial reasoning. We discuss the key challenges of spatial representation learning including multi-scale RL, continuous RL, shape-centric RL, noise-robust RL, heterogeneity-aware RL, and fairness-aware RL. We also discuss the critical role and potential of SRL in various geospatial subdomains and how this technique can lead to a new generation of GeoAI.
Gengchen Mai, Xiaobai Angela Yao, Yiqun Xie, Jinmeng Rao, Hao Li 0019, Qing Zhu 0011, Ni Lao
SIGSPATIAL/GIS2
2024 TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning
abstract
Spatial representation learning (SRL) aims at learning general-purpose neural network representations from various types of spatial data (e.g., points, polylines, polygons, networks, images, etc.) in their native formats. Learning good spatial representations is a fundamental problem for various downstream applications such as species distribution modeling, weather forecasting, trajectory generation, geographic question answering, etc. Even though SRL has become the foundation of almost all geospatial artificial intelligence (GeoAI) research, we have not yet seen significant efforts to develop an extensive deep learning framework and benchmark to support SRL model development and evaluation. To fill this gap, we propose TorchSpatial, a learning framework and benchmark for location (point) encoding,which is one of the most fundamental data types of spatial representation learning. TorchSpatial contains three key components: 1) a unified location encoding framework that consolidates 15 commonly recognized location encoders, ensuring scalability and reproducibility of the implementations; 2) the LocBench benchmark tasks encompassing 7 geo-aware image classification and 10 geo-aware imageregression datasets; 3) a comprehensive suite of evaluation metrics to quantify geo-aware models’ overall performance as well as their geographic bias, with a novel Geo-Bias Score metric. Finally, we provide a detailed analysis and insights into the model performance and geographic bias of different location encoders. We believe TorchSpatial will foster future advancement of spatial representationlearning and spatial fairness in GeoAI research. The TorchSpatial model framework and LocBench benchmark are available at https://github.com/seai-lab/TorchSpatial, and the Geo-Bias Score evaluation framework is available at https://github.com/seai-lab/PyGBS.
Nemin Wu, Zeping Liu, Yanlin Qi, Jielu Zhang, Joshua Ni, Xiaobai Angela Yao, Lan Mu, Stefano Ermon, Tanuja Ganu, Akshay Uttama Nambi, Ni Lao, Gengchen Mai
NeurIPS8
2019 Representation and analytical models for location-based big data
abstract
The last decade has seen an exponential growth in location-based big data research. Indeed, the availability of fine-grained location-based big data has created unprecedented opportunities for rese...
Xiaobai Angela Yao, Haosheng Huang, Bin Jiang 0004, Jukka Matthias Krisp
Int. J. Geogr. Inf. Sci.1