Yunfan Kang

dblp:244/1338 · DBLP profile ↗
← Back
7ranked-venue papers in the field
4as first author
5since 2021 · last 2026
0000-0001-8488-201XORCID · corroborated

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 7 (4 first)
YearPublicationVenuePosition
2026 IMS: Incremental Max-P Regionalization With Statistical Constraints
Yunfan Kang, Yiyang Bian, Qinma Kang, Amr Magdy 0001
IEEE Trans. Knowl. Data Eng.1
2024 Pyneapple-L: Scalable Expressive Learning-based Spatial Analysis
abstract
This paper demonstrates Pyneapple-L, an open-source library designed to enhance scalable spatial analysis through learning-based techniques. Through collaboration with social scientists and domain experts, we identify scalability challenges inherent in conventional spatial analysis methods, particularly as the data size increases. Pyneapple-L addresses these challenges by leveraging learning-based models to offer scalable solutions. We demonstrate two modules: scalable learning of spatial hotspots along spatial networks and augmented geographically weighted regression. To showcase Pyneapple-L, we have developed a user-friendly frontend web application to interact with different datasets, algorithms, model configurations, and visualize outcomes on interactive maps that support both broad and analytical views.
Yongyi Liu, Nicolas Lee, Yunfan Kang, Mohammad Reza Shahneh, Ahmed R. Mahmood, Vishal Rohith Chinnam, Aparna Vivek Sarawadekar, Samet Oymak, Ibrahim Sabek, Amr Magdy 0001
SIGSPATIAL/GIS3
2024 Pyneapple-R: Scalable and Expressive Spatial Regionalization
abstract
This paper demonstrates Pyneapple-R, an open-source library for scalable and expressive regionalization. Re-gionalization algorithms, also known as the ‘spatially-constrained clustering algorithms', have been widely adopted in spatial analysis tasks and now evolving towards a more large-scale and fine-scale direction. Through collaborations with social scientists and domain experts, we have identified emerging challenges in existing regionalization techniques, particularly regarding scalability and expressiveness. As data volumes continue to grow and regionalization algorithms become increasingly crucial to decision-making across various fields, enhancing these aspects can significantly impact the quality and effectiveness of re-search and applications. To address these challenges, Pyneapple-R provides novel algorithms for regionalization queries including the expressive p-regions algorithm, the scalable max-p regions algorithm, and the expressive max-p regions problem. To show-case Pyneapple-R, we have developed frontend web applications that enable users to interact with the algorithms by selecting constraints or simply engaging in conversation with the system to issue queries with the help of popular AI models. Interactive notebooks, designed to demonstrate the superiority and simplicity of Pyneapple-R, provide varying levels of detail to help social scientists and developers explore its full potential.
Yunfan Kang, Yongyi Liu, Hussah Alrashid, Akash Bilgi, Siddhant Purohit, Ahmed Mahmood, Sergio J. Rey, Amr Magdy 0001
ICDE1
2023 Scalable Evaluation of Local K-Function for Radius-Accurate Hotspot Detection in Spatial Networks
abstract
The widespread of geotagged data combined with modern map services allows for the accurate attachment of data to spatial networks. Applying statistical analysis, such as hotspot detection, over spatial networks is very important for precise quantification and patterns analysis, which empowers effective decision-making in various important applications. Existing hotspot detection algorithms on spatial networks either lack statistical evidence on detected hotspots, such as clustering, or they provide statistical evidence at a prohibitive computational overhead. In this paper, we propose efficient algorithms for detecting hotspots based on the network local K-function for predefined and unknown hotspot radii. The network local K-function is a widely adopted statistical approach for network pattern analysis that enables the understanding of the density and distribution of activities and events in the spatial network. However, its practical application has been limited due to the inefficiency of existing algorithms, particularly for large-sized networks. Extensive experimental evaluation using real and synthetic datasets shows that our algorithms are up to 28 times faster than the state-of-the-art algorithms in computing hotspots with a predefined radius and up to more than four orders of magnitude faster in identifying hotspots without a predefined radius.
Yongyi Liu, Yunfan Kang, Ahmed R. Mahmood, Amr Magdy 0001
SIGSPATIAL/GIS2
2022 EMP: Max-P Regionalization with Enriched Constraints
abstract
Spatial regionalization is the process of grouping a set of spatial areas into spatially contiguous and homogeneous regions. This paper introduces an enriched max-p-regions (EMP) problem; a regionalization process that allows enriched user-defined constraints based on SQL aggregate functions. In addition to enabling richer constraints, it enables users to employ multiple constraints simultaneously to significantly push the expressiveness and effectiveness of the existing regionalization literature. The EMP problem is NP-hard and significantly enriches the existing regionalization problems. Such a major enrichment introduces several challenges in both feasibility and scalability. To address these challenges, we propose the FaCT algorithm, a three-phase greedy approach that finds a feasible set of spatial regions that satisfy EMP constraints while supporting large datasets compared to the existing literature. Our extensive experimental evaluation has demonstrated the effectiveness and scalability of our techniques on several real datasets.
Yunfan Kang, Amr Magdy 0001
ICDE1
2020 Microblogs data management: a survey
Amr Magdy 0001, Laila Abdelhafeez, Yunfan Kang, Eric Ong, Mohamed F. Mokbel
VLDB J.3
2019 Scalable Multi-resolution Spatial Visualization for Anthropogenic Litter Data
abstract
This paper demonstrates CleanUpOurWorld; a research spatial database that is designed and deployed to collect, process, query, and visualize anthropogenic litter data. Such data has a significant importance in the field of environmental sciences due to its important use cases. We make a major on-going effort to collect and maintain such data worldwide from different sources through a community of environmental scientists and partner organizations. With the increasing volume of data, existing software packages, such as GIS software, do not scale to process, query, and visualize such data. To overcome this, CleanUpOurWorld digests datasets from diferent sources, with different formats, in a scalable backend that cleans, integrates, and unifies them in a structured form in a relational spatial database. Frontend applications are built to visualize litter data at multiple spatial resolutions.
Yunfan Kang, Ziang Zhao, Amr Magdy 0001, Win Cowger, Andrew B. Gray
SIGSPATIAL/GIS1