Viswadeep Lebakula

dblp:309/2169 · DBLP profile ↗
← Back
3ranked-venue papers in the field
2as first author
3since 2021 · last 2024
0000-0001-5293-5914ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3 (2 first)
YearPublicationVenuePosition
2024 Geographical Insights into Suicide Mortality Through Spatial Machine Learning
abstract
Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2improved from 0.59 to 0.67) and local (R2improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.
Viswadeep Lebakula, Swapna S. Gokhale, Anuj J. Kapadia, Jodie Trafton, Alina Peluso
IEEE Big Data1
2024 At Risk Population Estimates for Belarus, Poland and Slovakia with Machine Learning
abstract
High-resolution gridded population modeling is crucial for various applications, including disaster response planning, infectious disease spread modeling, climate change impact estimation, policy development, and more. Multiple gridded population datasets have been developed, each tailored to meet specific objectives. Among them, LandScan Global dataset is designed to represent ambient and unwarned population distributions. However, this dataset relies on a statistical approach that requires manual adjustments, making it time consuming and labour intensive. Existing machine learning (ML) methods often train and test at different spatial resolutions, potentially leading to inflated results, and they rely on Census population totals for disaggregation. To address these limitations, in this study we developed population estimates using ML models trained and tested at a consistent 30 arc-second resolution (≈1 square kilometer), specifically using Random Forest (RF) and XGBoost. These models were trained on 2020 datum to predict for 2021 for three countries: Belarus, Poland, and Slovakia. Our findings show that both RF (MAE varies from 5.75 to 13.25) and XGBoost (MAE varies from 8.15 to 23.44) model performance is close to LandScan Global estimates. Furthermore, neither of the models performed the best across all grid cells: the RF model was more effective in areas with lower populations, while XGBoost excelled in more densely populated regions. The proposed approach can be used for countries where the Census data is not available.
Viswadeep Lebakula, Clinton Stipek, Daniel S. Adams, Justin Epting, Marie L. Urban
IEEE Big Data1
2024 Empirically Categorizing the Built Environment in Relation to Height
abstract
Buildings are a core component of the urban environment and affect human populations, energy usage, city development, city planning, and urban heat islands. Buildings span an enormous range of sizes, from a 2m tall shelter to the Burj Khalifa; and at the same time there are widely recognized categories of similar buildings, with homes, office buildings, or skyscrapers as some examples. Currently, there is no consistent method to quantitatively determine how a building should be categorized by its height, or how many categories there should be within the built environment. Additionally, these categories vary spatially, leading to multiple definitions at local scales of what it means to be a tall, medium, or short building. Here, we find across 17.59 million buildings in the United States, Germany, and Japan, that applying a K-nearest neighbor approach to quantitatively bin the built environment outperforms the current state-of-the-art, subjective domain knowledge. This was evidenced as our method of leveraging a K-nearest neighbor improved upon the existing approach of using domain knowledge by 10% with respect to precision, recall, F1-score and accuracy. Our results showcase the finding that it is possible to generate a global and consistent approach to categorizing the built environment in relation to height. This is significant in that there is now a quantitative way to categorize the built environment based on building height at a global scale, allowing researchers a consistent platform for comparison and collaboration across various applications.
Clinton Stipek, Justin Epting, Daniel S. Adams, Viswadeep Lebakula, Taylor Hauser, Christa Brelsford, Allan Ross
IEEE Big Data4