Yuhao Kang

dblp:241/5156 · DBLP profile ↗
← Back
8ranked-venue papers in the field
1as first author
7since 2021 · last 2025
0000-0003-3810-9450ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 8 (1 first)
YearPublicationVenuePosition
2025 Leveraging Reinforcement Learning for Maternity Care Resource Reallocation: A Case Study in Florida
abstract
Persistent disparities in access to maternal healthcare across the United States, particularly in rural and underserved communities, have resulted in poor maternal outcomes. Traditional statistical methods, such as Quadratic Programming (QP), have been utilized for healthcare resource reallocation, but they struggle with dynamic, multi-objective geographic optimization problems. This study presents a Reinforcement Learning (RL)-based framework to optimize maternity care resource distribution. We follow the Maximal Accessible Equality Problem (MAEP), aiming to enhance spatial equality by reducing the weighted accessibility variance. We integrate the Two-Step Floating Catchment Area (2SFCA) method for measuring maternity care accessibility and leverage Proximal Policy Optimization (PPO) to dynamically reallocate obstetric facilities across counties in Florida. To account for real-world complexities, we tested our RL framework under three scenario objectives: minimizing distance, obstetric bed supply preservation, and prioritizing underserved counties. Results show the proposed framework effectively reduces accessibility variance by 31.7–49.3%. These findings highlight the potential of RL in offering scalable, data-driven solutions to support equitable maternity health care, benefiting practical applications for implementing actionable health policy decision making.
Andy Qin, Yuhao Kang, Shiqi Wang 0029, Fahui Wang, Peiyin Hung
SIGSPATIAL/GIS2
2025 CartoAgent: a multimodal large language model-powered multi-agent cartographic framework for map style transfer and evaluation
abstract
The rapid development of generative artificial intelligence (GenAI) presents new opportunities to advance the cartographic process. Previous studies have either overlooked the artistic aspects of maps or faced challenges in creating both accurate and informative maps. In this study, we propose CartoAgent, a novel multi-agent cartographic framework powered by multimodal large language models (MLLMs). This framework simulates three key stages in cartographic practice: preparation, map design, and evaluation. At each stage, different MLLMs act as agents with distinct roles to collaborate, discuss, and utilize tools for specific purposes. In particular, CartoAgent leverages MLLMs’ visual aesthetic capability and world knowledge to generate maps that are both visually appealing and informative. By separating style from geographic data, it can focus on designing stylesheets without modifying the vector-based data, thereby ensuring geographic accuracy. As a result, the proposed CartoAgent could effectively produce maps that are not only visually appealing but also accurate and informative. We applied it to a specific task centered on map restyling, namely, map style transfer and evaluation. The effectiveness of this framework was validated through extensive experiments and a human evaluation study. CartoAgent can be extended to support a variety of cartographic design decisions and inform future integrations of GenAI in cartography.
Chenglong Wang 0004, Yuhao Kang, Zhaoya Gong, Yu Feng 0006
Int. J. Geogr. Inf. Sci.2
2024 Estimating urban noise along road network from street view imagery
abstract
Estimating road traffic noise is essential for examining the quality of sounding environment and mitigating such a non-negligible pollutant in urban areas. However, existing estimated models often have limited applicability to specific traffic conditions, while the required parameters may not be readily available for city-wide collection. This paper proposes a data-driven approach for measuring road-level acoustic information of traffic with street view imagery. Specifically, we utilize portable vehicle-equipped hardware for in-situ noise acquisition and employ a deep learning model ResNet to learn high-level visual features from street view images that are closely associated with road traffic noise. The ResNet captures meaningful patterns from the input data, and the output probability vectors are then fed into a Random-Forest regression algorithm to quantitatively estimate the noise in decibels for different road segments. The MAE and RMSE of the DCNN-RF model are 2.01 and 2.71, respectively. Additionally, we employ a gradient-weighted Class Active Mapping approach to visually interpret our deep learning model and explore the significant elements in streetscapes that contribute to the model's estimations. Our proposed framework facilitates low-cost and fine-scale road traffic noise estimations and sheds light on how auditory information could be inferred from street imagery, which may benefit practices in geography and urban planning.
Teng Fei 0001, Yuhao Kang, Guofeng Wu
Int. J. Geogr. Inf. Sci.3
2022 STICC: a multivariate spatial clustering method for repeated geographic pattern discovery with consideration of spatial contiguity
abstract
Spatial clustering has been widely used for spatial data mining and knowledge discovery. An ideal multivariate spatial clustering should consider both spatial contiguity and aspatial attributes. Existing spatial clustering approaches may face challenges for discovering repeated geographic patterns with spatial contiguity maintained. In this paper, we propose a Spatial Toeplitz Inverse Covariance-Based Clustering (STICC) method that considers both attributes and spatial relationships of geographic objects for multivariate spatial clustering. A subregion is created for each geographic object serving as the basic unit when performing clustering. A Markov random field is then constructed to characterize the attribute dependencies of subregions. Using a spatial consistency strategy, nearby objects are encouraged to belong to the same cluster. To test the performance of the proposed STICC algorithm, we apply it in two use cases. The comparison results with several baseline methods show that the STICC outperforms others significantly in terms of adjusted rand index and macro-F1 score. Join count statistics is also calculated and shows that the spatial contiguity is well preserved by STICC. Such a spatial clustering method may benefit various applications in the fields of geography, remote sensing, transportation, and urban planning, etc.
Yuhao Kang, Kunlin Wu, Song Gao 0001, Ignavier Ng, Jinmeng Rao, Shan Ye, Fan Zhang 0011, Teng Fei 0001
Int. J. Geogr. Inf. Sci.1
2022 Theseus: Navigating the Labyrinth of Time-Series Anomaly Detection
abstract
The detection of anomalies in time series has gained ample academic and industrial attention, yet, no comprehensive benchmark exists to evaluate time-series anomaly detection methods. Therefore, there is no final verdict on which method performs the best (and under what conditions). Consequently, we often observe methods performing exceptionally well on one dataset but surprisingly poorly on another, creating an illusion of progress. To address these issues, we thoroughly studied over one hundred papers, and summarized our effort in TSB-UAD, a new benchmark to evaluate univariate time series anomaly detection methods. In this paper, we describe Theseus, a modular and extensible web application that helps users navigate through the benchmark, and reason about the merits and drawbacks of both anomaly detection methods and accuracy measures, under different conditions. Overall, our system enables users to compare 12 anomaly detection methods on 1980 time series, using 13 accuracy measures, and decide on the most suitable method and measure for some application.
Paul Boniol, John Paparrizos, Yuhao Kang, Themis Palpanas, Ruey S. Tsay, Aaron J. Elmore, Michael J. Franklin
Proc. VLDB Endow.3
2022 TSB-UAD: An End-to-End Benchmark Suite for Univariate Time-Series Anomaly Detection
abstract
The detection of anomalies in time series has gained ample academic and industrial attention. However, no comprehensive benchmark exists to evaluate time-series anomaly detection methods. It is common to use (i) proprietary or synthetic data, often biased to support particular claims; or (ii) a limited collection of publicly available datasets. Consequently, we often observe methods performing exceptionally well in one dataset but surprisingly poorly in another, creating an illusion of progress. To address the issues above, we thoroughly studied over one hundred papers to identify, collect, process, and systematically format datasets proposed in the past decades. We summarize our effort in TSB-UAD, a new benchmark to ease the evaluation of univariate time-series anomaly detection methods. Overall, TSB-UAD contains 13766 time series with labeled anomalies spanning different domains with high variability of anomaly types, ratios, and sizes. TSB-UAD includes 18 previously proposed datasets containing 1980 time series and we contribute two collections of datasets. Specifically, we generate 958 time series using a principled methodology for transforming 126 time-series classification datasets into time series with labeled anomalies. In addition, we present data transformations with which we introduce new anomalies, resulting in 10828 time series with varying complexity for anomaly detection. Finally, we evaluate 12 representative methods demonstrating that TSB-UAD is a robust resource for assessing anomaly detection methods. We make our data and code available at www.timeseries.org/TSB-UAD. TSB-UAD provides a valuable, reproducible, and frequently updated resource to establish a leaderboard of univariate time-series anomaly detection methods.
John Paparrizos, Yuhao Kang, Paul Boniol, Ruey S. Tsay, Themis Palpanas, Michael J. Franklin
Proc. VLDB Endow.2
2021 Emotional habitat: mapping the global geographic distribution of human emotion with physical environmental factors using a species distribution model
abstract
Human emotion is an intrinsic psychological state that is influenced by human thoughts and behaviours. Human emotion distribution has been regarded as an important part of emotional geography research. However, it is difficult to form a global scaled map reflecting human emotions at the same sampling density because various emotional sampling data are usually positive occurrences without absence data. In this study, a methodological framework for mapping the global geographic distribution of human emotion is proposed and applied, combining a species distribution model with physical environment factors. State-of-the-art affective computing technology is used to extract human emotions from facial expressions in Flickr photos. Various human emotions are considered as different species to form their ‘habitats’ and predict the suitability, termed as ‘Emotional Habitat’. To our knowledge, this framework is the first method to predict emotional distribution from an ecological perspective. Different geographic distributions of seven dimensional emotions are explored and depicted, and emotional diversity and abnormality are detected at the global scale. These results confirm the effectiveness of our framework and offer new insights to understand the relationship between human emotions and the physical environment. Moreover, our method facilitates further rigorous exploration in emotional geography and enriches its content.
Yizhuo Li 0002, Teng Fei 0001, Yingjing Huang, Xiang Li 0086, Fan Zhang 0011, Yuhao Kang, Guofeng Wu
Int. J. Geogr. Inf. Sci.7
2019 A Data-Driven Approach to Understanding and Predicting the Spatiotemporal Availability of Street Parking
abstract
Searching for a parking spot in metropolitan areas is a great challenge comparable to the Hunger Games, especially in highly populated areas such as downtown districts and job centers. On-street parking is often a cost-effective choice compared to parking facilities such as garages and parking lots. However, limited space and complex parking regulation rules make the search process of on-street parking very difficult. To this end, we propose a data-driven framework for understanding and predicting the spatiotemporal availability of on-street parking using the NYC parking tickets open data, points of interest (POI) data and human mobility data. Four popular types of spatial analysis units (i.e., point, street, census tract, and grid) are used to examine the effects of spatial scale in machine learning predictive models. The results show that random forest works the best with the highest accuracy scores for the spatiotemporal availability classification across all four spatial analysis scales.
Mingxiao Li 0001, Song Gao 0001, Yunlei Liang, Joseph Marks, Yuhao Kang, Moyin Li
SIGSPATIAL/GIS5