Thilanka Munasinghe

dblp:123/3128 · DBLP profile ↗
← Back
28ranked-venue papers in the field
5as first author
21since 2021 · last 2025
0000-0002-0911-750XORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 28 (5 first)
YearPublicationVenuePosition
2025 Early Detection of Harmful Algal Blooms using Machine Learning
Anirban Acharya, Thilanka Munasinghe, Kevin Rose
IEEE Big Data2
2025 Data Driven Dengue Dynamics via Satellite Data
Rafi Magdon-Ismail, Thilanka Munasinghe, Jennifer C. Wei
IEEE Big Data2
2025 A Knowledge-Based System for Managing Hardware Dependency and Reproducibility in Quantum Machine Learning Workflows
Thilanka Munasinghe, Kimberly A. Cornell, James A. Hendler, George Berg, Jennifer C. Wei
IEEE Big Data1
2025 From Black Box to Insight: Explainable AI for Extreme Event Preparedness
Kiana Vu, Ismet Selçuk Özer, Phung Lai, Thilanka Munasinghe, Jennifer C. Wei
IEEE Big Data5
2024 A Knowledge Graph Framework for Organizing Heterogeneous Datasets for Utilization in Classical and Quantum Computing: Current Challenges and Future Directions
abstract
The lack of representation in interaction within environmental variables found in literature led to the development of a novel framework that reflects the true nature of the inter-connectedness in our environment. We propose an Environmental Interaction Knowledge Graph (EIKG) framework. This general EIKG framework works as the basis for interconnected environ-mental events by knitting interrelated events such as hurricanes leading to storm surges, which lead to flood events that could cause events such as mudslides and landslides. The cascading nature of one event leading to another related event in the environment requires an adequate understanding of each event using contextual information before conducting any data-driven analytics. This vision paper showcases how the EIKG:floods, EIKG:wildfire EIKG:landslides, etc., can be derived from a base case framework of EIKG as those individual events are interconnected with some common denominator variables. As an example, the precipitation variable is used in the flood case study as well as in the wildfire or drought case study, as excessive precipitation levels lead to floods, and lack of precipitation leads to droughts and wildfires. We identify the precipitation variable as a "common-denominator-variable" in extreme weather events that play a key role in modeling the environment leading to different extreme weather events based on the variability of that variable (varying values where low precipitation leads to drought, and high values lead to floods). Insights from EIKG facilitate data analysis using both classical and Quantum Machine Learning (QML) techniques. The EIKG organizes heterogeneous datasets and integrates relationships to address extreme weather events. This study incorporates various datasets, including mobility data, socioeconomic data from the US Census Bureau, climate data from NASA, and critical infrastructure data.
Thilanka Munasinghe, Kimberly A. Cornell, Jennifer C. Wei, George Berg, James A. Hendler
IEEE Big Data1
2024 Assessment of Quantum ML Applicability for Climate Actions: Comparison of the Variational Quantum Classifier and the Quantum Support Vector Classifier with Classical ML Models
abstract
Climate change refers to significant and long-term alterations in the Earth’s climate patterns, typically resulting from human activities that increase greenhouse gas emissions. Addressing climate change is not merely an option but a necessity, demanding creative solutions and efforts from individuals, researchers, communities, and governments. Despite the capabilities of machine learning (ML) with data-driven solutions promising to combat climate change-related problems, they face challenges stemming from traditional computational methods and prolonged training times, impeding their practical utility. Recent strides in quantum computing have permeated diverse domains, spanning from manufacturing engineering and pharmaceutical discovery to the latest frontier of detecting climate anomalies. With the potential to substantially reduce time and computational complexity, quantum computing shows promise in addressing climate change impacts. Its distinctive features will enable the concurrent exploration of expansive solution spaces, making it well-suited for analyzing extensive climate datasets, simulating intricate climate models, optimizing resource allocation, and discerning patterns in climate data for mitigation and adaptation endeavors. This study explores the potential of using Quantum machine learning (QML) techniques on climate and weather data obtained from NASA Giovannis. We used two QML algorithms, the Quantum Support Vector Classifier (QSVC) and the Variational Quantum Classifier (VQC) models, using the IBM Qiskit ML 0.7.2 ecosystem. We used an actual 127-Qubit IBM Quantum Computer (IBM 127-qubit Eagle) in this study. The methodology and results sections describe the experiences gained from applying and evaluating quantum ML results on climate and weather data obtained from NASA satellites as a novel practical application of quantum computing.
Thilanka Munasinghe, Phung Lai, Jennifer C. Wei, James A. Hendler, Kimberly A. Cornell
IEEE Big Data1
2024 Natural Language Processing for Extracting Rich Disease Data Aligned To Satellite Meteorological Data
abstract
Global climate change is redefining our understanding of how diseases spread. In Sri Lanka, vector-borne diseases such as dengue fever historically surged during the monsoon seasons when temperatures were high enough for mosquito eggs to hatch. Unfortunately, due to rising temperatures and more erratic rainfall patterns, mosquito eggs can now hatch year-round making outbreaks increasingly unpredictable, leading to an alarming rise in hospitalizations and deaths. More data is needed to adapt our response to these diseases in an increasingly warmer world. In the contemporary landscape, a wealth of disease information is available, yet accessibility remains limited due to unstructured data formats such as PDFs. Therefore, converting unstructured disease reports into structured formats is necessary for effectively leveraging data. This paper introduces a comprehensive framework for collecting unstructured disease reports and transforming them into analyzable formats. By creating separate models tailored to each data format, we can ensure accuracy compared to general models. These straightforward models enhance accessibility and empower other researchers to use our tools. The returned structured data can then be harnessed for analysis, statistical purposes, and informing evidence-based public health interventions, thus facilitating more informed decision-making in healthcare. We deploy this framework to produce geospatial data for Sri Lanka and Brazil for many different conditions and align these data with satellite environmental data, providing for the first time a structured, aligned powerful dataset for disease modeling.
Mahi Pasarkar, Junseob Kim, Eoin O'Gara, Alan Zhang, Malik Magdon-Ismail, Thilanka Munasinghe, Jiaqi Weng, David Qiu, Ethan Cruz, Jennifer C. Wei, Ashan Pathirana
IEEE Big Data6
2024 Energy Infrastructure Risk Modeling using Quantum and Classical Machine Learning
abstract
The integrity of energy infrastructure is critical to societal stability, yet it is increasingly vulnerable to diverse natural and man-made hazards. These hazards create interdependent risks that challenge conventional risk management models. This study explores hazard modeling approaches for energy infrastructure using classical machine learning (ML) and quantum machine learning (QML) techniques. We develop Deep Neural Networks (DNN), Naïve Bayes (NB), Decision Tree (DT), and Quantum Neural Networks (QNN) deployed in a 127-qubit quantum computer to classify hazard levels. Further, we propose a Bayesian network to model vulnerability propagation within the infrastructure. Comparing QML models against classical ML benchmarks provides insights into QML’s potential benefits and limitations for complex, high-dimensional data in critical infrastructure applications. This paper presents data analysis, model development, and comparison of classical and quantum methods to advance energy infrastructure risk modeling.
Neelanga Thelasingha, Thilanka Munasinghe
IEEE Big Data2
2024 Graph Representation Learning for Dengue Forecasting
abstract
The global expansion of the dengue belt, driven by climate change and increased urbanization, has led to a significant rise in dengue cases worldwide (1). Early warning systems (EWS) coupled with prompt public health response mechanisms are crucial in mitigating dengue-related morbidity and mortality globally. In Sri Lanka, dengue transmission occurs year-round with two peaks correlating to the southwest monsoon from May to September and the northeast monsoon from October to January (2). The presence of multiple dengue virus serotypes (DENV1–4) complicates epidemiological patterns, as sequential infections with different serotypes can increase the risk of severe disease manifestations detected by surveillance systems (3). Understanding and integrating these virological dynamics, vector dynamics, and real-time surveillance data are essential for developing effective EWS and targeted public health interventions. We propose the use of Graph Neural Networks (GNNs) as an EWS. Using Earth observational data from NASA’s global satellites and dengue incidence data from Sri Lanka’s Ministry of Health, we developed traditional and graph-based EWS to forecast dengue cases across Sri Lanka’s 25 districts between 2013 and 2022. We demonstrate empirically that GNNs incorporating spatiotemporal relations significantly outperform traditional EWS models such as Autoregressive Integrated Moving Average (ARIMA), Random Forest, and Long Short-Term Memory (LSTM). Our source code is available on GitHub.
Jiaqi Weng, David Qiu, Ethan Cruz, Malik Magdon-Ismail, Thilanka Munasinghe, Jennifer C. Wei, Ashan Pathirana, Mahi Pasarkar
IEEE Big Data5
2023 Exploring Power Outage Prediction Using Weather and Socioeconomic Data in the Southeastern Part of the United States
abstract
Electricity has become an indispensable part of our society. As natural disasters strike communities, it is crucial to prepare the power grid for the incoming storm. This exploratory paper delves into the methods used to predict power outages in the United States’ southeastern counties after hurricanes. The prediction is made using a range of machine learning techniques. Four methods, namely Multivariable Linear Regression, Random Forest, Extreme Gradient Boosting (XGBoost), and K-Nearest Neighbors Regression (KNN), are employed to predict the percentage of electric customers in a county that will face a power outage the next day. The accuracy of these models is assessed using regression metrics such as Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and Mean Absolute Error (MAE). To assess a county’s vulnerability to power outages, meteorological, electrical, and socioeconomic data were taken into account. Rolling average data was also included for day-to-day features. A correlation matrix was used to select variables relevant to this analysis. The hyperparameters of each model were chosen based on the parameters that resulted in the lowest MSE in 5-fold cross-validation. After comparing different models, it was observed that the Random Forest Regression method had the lowest MSE of 5.93E-05, indicating sufficiency. These model implementations helped to better understand the nature of the problem and plan future research work toward predicting weather-based power outages, particularly those caused by hurricanes as they are related to precipitation (rainfall) which we included in this study. This analysis suggests that creating a model specifically for predicting power outages is necessary to understand what are the most influential weather-related variables that can be used in data-driven analysis.
Ethan Cruz, Thilanka Munasinghe
IEEE Big Data2
2023 From Satellites to Fields: Machine Learning Applications for Prediction of Corn Production Using NDVI, Precipitation and Land Surface Temperature for Large Producer Countries
abstract
In this work-in-progress study, we aim to determine the predictors of corn production in Iowa (United States), Heilongjiang (China), Mato Grosso (Brazil), Cordoba (Argentina), and Poltava Oblast (Ukraine). The effectiveness of predicting annual corn production using precipitation and land surface temperature (LST) values obtained from the Earth Engine tool, and Moderate Resolution Imaging Spectroradiometer (MODIS) Normalised Difference Vegetation Index (NDVI) values obtained from the National Aeronautics and Space Administration (NASA) Global Inventory Monitoring and Modelling Studies (GIMMS) Global Agricultural Monitoring System, is examined. A comparison is conducted between multiple linear regression, ridge regression, and lasso regression models. The highest adjusted R2values are found in order to identify the optimal model in which the corn yield variance is explained by a combination of variables including the year, NDVI sample and anomaly values, precipitation, and LST. The results show that corn production values in Iowa, Heilongjiang, Mato Grosso, Cordoba, and Poltava Oblast can best be predicted using lasso regression models with adjusted R2values of 0.841, 0.933, 0.847, 0.854, and 0.860 respectively.
Saniya Nangia, Thilanka Munasinghe, Heidi Tubbs, Assaf Anyamba
IEEE Big Data2
2023 Integrating Climate Variable Data in Machine Learning Models for Predictive Analytics of Tomato Yields in California
abstract
Traditionally, agricultural forecasting has relied on empirical methods and basic statistical analysis, such as applying average values from previous years’ yields or using a simple linear fit for next year’s predictions. However, the emergence of data-driven approaches, particularly machine learning algorithms, has revolutionized yield prediction in agriculture. Machine learning techniques have demonstrated their potential to provide accurate predictions. However, existing models often rely on a limited number of input variables for crop yield predictions, which makes them only suitable for specific scenarios. In this study, we have developed four distinct machine learning-based predictors, incorporating various climate factors, including daytime temperature, nighttime temperature, precipitation (rainfall), vegetation index, and evapotranspiration as input variables to predict tomato acreage yields in counties of California, USA. Our results show that regression models constructed using neural networks and linear regression exhibited better performance than other predictors, achieving an average accuracy rate of 70% to 80%. Compared to most of the existing crop yield predictors, our models offer versatility while maintaining a desirable level of predictive accuracy. Expanding the number of input variables, such as nitrogen fertilizer usage etc, and introducing larger spatial and temporal high-resolution datasets for model training can improve our model performance, enabling us to obtain better results in tomato yield prediction.
Tianze Zhu, Tingyi Tan, Shuheng Wang, Thilanka Munasinghe, Heidi Tubbs, Assaf Anyamba
IEEE Big Data5
2022 Landslide Likelihood Prediction using Machine Learning Algorithms
abstract
The supply of electricity via power plants is critical to the operation of many critical infrastructure systems in modern society. Natural hazards can disrupt the power supply, cause power outages that can halt economic growth, and impede emergency response until power is restored. The proposed work aims to predict the landslides likelihood in these critical infrastructure locations in the Northeastern USA using integrated databases of explanatory variables and machine learning algorithms. First, data related to landslides are obtained and merged, including topographic, soil moisture, and precipitation-related data. Five regression algorithms, namely: Random Forest, Extreme Gradient Boosting (XGBoost), K-Nearest Neighbor regression (KNN), Linear Support Vector Regressor (SVR), and Linear regression, are utilized to predict the landslide probability and evaluated on the dataset. The accuracy of the models is assessed by using statistical metrics such as mean absolute error (MAE), mean squared error (MSE), and root mean squared error (RMSE). The study results show that Random Forest outperformed other models with the mutual information feature selection method. It achieved an MSE of 0.0011 with mutual information-based feature selection and an MSE of 0.00157 without feature selection. KNN regressor outperformed the other models with an MSE of 0.00139 with correlation-based information selection. The proposed landslide identification model with Random Forest algorithm shows outstanding robustness and great potential in tackling the landslide likelihood prediction by employing ML algorithms.
Vasundhara Acharya, Anindita Ghosh, Inwon Kang, Thilanka Munasinghe, K. C. Binita
IEEE Big Data4
2022 Using Graph Neural Networks to Investigate the Relationship Between the Socioeconomic Factors and Emergency Medical Service (EMS) Median Response Time in New York City
abstract
In this working-in-progress study, we plan to use Graph Neural Networks (GNNs) to investigate whether and how socio-demographic factors influence p redicted r esponse t imes of EMS services in New York City (NYC). Currently, the application of GNNs to predict the EMS response time is relatively novel. Operational models, which prioritize task-specific features (such as call priority level and day/time of incident) have been deployed in several contexts around the world. Leveraging unique capabilities of different emergent GNN architectures, we will evaluate whether the predictive accuracy of operational models of EMS response time are improved when neighborhood-level socio-economic factors are included. Focusing on different neighborhoods of NYC, we plan to use EMS Incident Dispatch Data obtained from the Fire Department of New York (FDNY) and accessed through the NYC Open Data Portal. We couple incident data with socio-economic factors at the neighborhood level, such as average income, population density, poverty level, and percentage of the population speaking a second language, obtained from the U.S Census Bureau and other sources. Our starting assumption is that response time is associated with different operational features (such as Call Priority Level) and that these features are consistent across neighborhoods. We then represent the call Priority Level Set and Neighborhood Set in a graph G = (U,V,E). The U and V are the node sets for Priority Levels (PLs) and Neighborhoods (NBs), respectively. The relationships or the links between the nodes U and V are represented by the edge set denoted as E, which is a subset of U × V. If relationships between call Priority Level and median response time fluctuate a cross n eighborhoods, i t s uggests that extra-operational factors influence EMS response times, and that there are inequities in response shaped by the socio-demographic features of where the call originated. We explore how we can leverage the novel techniques of Graph Neural Networks in this field of study compared to previous work done using traditional machine learning techniques by the research community.
Thilanka Munasinghe, Brandon Behlendorf
IEEE Big Data1
2021 Binary Classification Edge Cases in Cyclones Using AlexNet and ResNet Neural Networks
abstract
Two convolutional neural network binary classifiers were developed to classify images with and without hurricanes. These binary classifiers were modelled after ResNet, and AlexNet architectures. The cyclone data was gathered with the assistance of the National Aeronautics and Space Administration’s Interagency Implementation and Advanced Concepts Team (NASA IMPACT). The IMPACT Team archived Tag Image File Format (TIFF) images from the NASA Worldview application. This data store containing over 250 TIFF files of cyclone data, with examples of true and false positives of this cyclone phenomena, was then augmented through shear transforms, random rotations, and additional techniques to generate the model data. With this augmented data set, batch normalization was implemented for model training to develop a model that could detect nuanced differences. Ultimately, the ResNet model overfit toward the classification of hurricanes and was unable to correctly classify the false hurricane labelled images.The AlexNet model correctly classified these same images with 75 percent accuracy. This final result proved more so an experiment on a classifier network’s ability to generate improved feature detection when given prior false positive, and true positive data. Additionally, the results proved that when remote sensing rare phenomena use a representative, and stratified sample.
Brendan Donnelly, Thilanka Munasinghe
IEEE BigData2
2021 Chronic Respiratory Disease: Risk Modeling Potential and Limitations
abstract
Chronic Respiratory Diseases (CRDs), including Chronic Obstructive Pulmonary Disease (COPD) and asthma, are among the leading causes of mortality worldwide, with 545 million prevalent cases in 2017. Symptoms of noninfectious CRDs are often exacerbated by ambient air pollution and changes in temperature and humidity. This study explores a novel application of machine learning to forecast CRD risk and discuss its merits and limitations. We developed, trained, and tested a Random Forest regressor using datasets over the United States during 2000-2016, with mortality rate as the target variable. The final regressor produced R-squared values of 0.7526 and 0.7528 for cross-validation and test dataset prediction, respectively, implying that our model generalizes well out-of-sample. The selected features comprise location and temporal encoders, population density, and net primary production. The study reveals significant potential for modeling CRD risk but highlights setbacks due to the primarily noninfectious nature of CRDs, phenomena only identifiable on finer spatiotemporal scales, and data limitations. We also identify methods that may refine our approach and describe future developments that may improve CRD risk modeling.
Alexander He 0002, Thilanka Munasinghe
IEEE BigData2
2021 Hidden Workforce: Analysis on Recognizing Unpaid Domestic Work Through Social Determinants
abstract
The United Nations collects data at gigabyte scale from the member countries for sustainable development goals, and recognizes the impact big data collection has on opportunities and risks for sustainable development. The United Nations (UN) Sustainable Development Goals (SDG) open data hub provided a dataset that details the "proportion of time spent on unpaid domestic chores and care work" world wide. This primary dataset is only a portion of the data available related to the research area. This data contains many related variables, such as the population's age, sex, and location metrics. In an effort to better understand the impact of unpaid domestic work, the dataset was analyzed in conjunction with another dataset from the UN Statistics Division that details the rate of divorce/separation in the world population. Our analysis included methods such as principal component analysis, the k-means clustering algorithm, the random forest clustering algorithm, and implementing a neural network. The principal components were clustered using the k-means algorithm to cluster the variables in a manner that explains the variance present in the dataset. The analysis found that age, sex, and location demographics are key variables that explain the diverse variation between countries and the percentage of time spent on unpaid domestic work.Machine learning algorithms enabled the confirmation of this relationship. Using the key variables identified, a random forest of decision trees and a neural network were generated to classify the percentage of time spent on unpaid domestic work. Similarly, the random forest algorithm and a neural network were also implemented to classify geographical regions. These models were compared to determine the strength of the relationship between age, sex, location metrics, the percentage of time spent, and geographical regions.The analysis detailed in this work strives to identify the social factors that classify the percentage of time spent on unpaid domestic work in accordance with the UN SDG 5.4, which is to "recognize and value unpaid care and domestic work through the provision of public services, infrastructure, and social protection policies and the promotion of shared responsibility within the household and the family as nationally appropriate."
James Hicks, Thilanka Munasinghe
IEEE BigData2
2021 Scraping Unstructured Data to Explore the Relationship between Rainfall Anomalies and Vector-Borne Disease Outbreaks
abstract
According to the World Health Organization (WHO), vector-borne diseases such as malaria and dengue account for 17% of all infectious disease cases and lead to more than 700,000 deaths per year. Tracking and predicting the spread of vector-borne diseases is a vital task that could save hundreds of thousands of lives annually. Oftentimes, the first reports of vector-borne disease outbreaks occur through emails and online reporting systems long before they are officially documented. Tracking and predicting the emergence and spread of vector-borne disease outbreaks requires extracting data from these unstructured sources in combination with historical weather and climate data to understand the underlying background triggers and disease dynamics. In this work, we develop a data extraction pipeline for the online outbreak reporting website ProMED-mail that utilizes a web scraper, transformer neural network summarizer, and named entity recognizer to obtain a dataset of malaria, dengue, zika, and chikungunya outbreaks over the last 30 years. This scraped dataset was further analyzed in association with global rainfall anomalies derived from NASA’s Integrated Multi-satellitE Retrievals for GPM [Global Precipitation Mission] (IMERG) dataset. This preliminary analysis was to understand the effect of global rainfall patterns on the spread of vector-borne diseases. Analysis of the ProMED-mail and GPM data shows that vector-borne disease outbreaks are clustered towards the tropics and outbreaks are often amplified during the rainy seasons. Our scraped dataset can be a valuable tool in creating comprehensive georeferenced disease records for modeling and predicting future outbreaks.
Ethan Joseph, Thilanka Munasinghe, Heidi Tubbs, Bhaskar Bishnoi, Assaf Anyamba
IEEE BigData2
2021 Analysis of Preprocessing Techniques, Keras Tuner, and Transfer Learning on Cloud Street image data
abstract
A Convolutional Neural Network is a powerful tool that has been extensively used for image classification. One specific area of application is remotely sensed images of meteorological phenomena such as cyclones and high latitude dust events. Such images are complicated in nature and hence may require special techniques for feature identification. Cloud streets are another such phenomenon that occurs in nature and is mainly captured in images taken by artificial satellites. In this work, deep learning models are implemented on NASA-IMPACT teams’ cloud street dataset. Three preprocessing techniques were tested to address the drawbacks of the dataset. Gaussian blur, census transformation to extract textural features, data augmentation, and removal of noise were implemented. Then techniques such as Keras tuner are also utilized for hyperparameter tuning to help achieve maximum accuracy. The results show the efficiency with which Keras tuner attempts to direct towards optimal hyperparameters restricting the number of iterations to a low value and obtaining a test dataset accuracy as high as 80.96%. Lastly, binary classification of Cloud Street satellite images is performed by leveraging the benefits of Transfer Learning and pre-trained models. The various architectures that were tested in this work were namely, EfficientNetB7, AlexNet, VGG19 and, InceptionNetV3. Transfer learning provides a quick approach to build deep learning models with good accuracy scores achieving high accuracies of 80.89% for the VGG19 architecture.
Sharmad Joshi, Jessie Ann Owens, Shlok Shah, Thilanka Munasinghe
IEEE BigData4
2021 Image Classification to identify Transverse Cirrus Band Clouds using Convolutional Neural Networks
abstract
The NASA IMPACT team is investigating a certain type of cloud called Transverse Cirrus Bands (aka. TCB), that can cause dangerous levels of aviation turbulence for nearby planes. We wanted to collaborate with the NASA IMPACT team by exploring how machine learning models that can detect Transverse Cirrus Band clouds in satellite images. To accomplish this task we decided to use various CNN models and data preprocessing methods to determine which would produce the most accurate detection. For our CNN models, we chose Simple Sequential, AlexNet, ResNet50, LetNet5, Google MobileNetV2, and VGG-16. For our data preprocessing methods we shrank our training images, converted images into grayscale, and turned images into edges by using edge detection. Our results showed that training CNNs on the originally colored satellite images produced the best results for most models. We believe there is still potential for images converted into edges to improve model accuracy, but we would need further work to calculate an appropriate edge detection threshold for satellite images of different sizes.
Gregory Saini, Thilanka Munasinghe, Kevin Pan
IEEE BigData2
2021 Covid Vaccine Risk Stratification
abstract
In response to the pandemic caused by the rapidly spreading COVID-19 virus, several highly effective vaccines have been developed by Pfizer, Moderna, and Janssen. Despite the promising efficacy of those vaccines, there remains the challenge of properly distributing vaccines to those who need it most in the US. Of particular concern are individuals who are at higher risk due to underlying medical conditions which have been shown to exacerbate COVID-19 symptoms and at times lead to fatal illnesses. In addition to this, a variety of socioeconomic factors have been linked to increased COVID-19 rates and increased mortality, such as race, age, income, mobility, and education level.This project aims to develop an information system to help advise vaccine distributors and state governments on how to effectively distribute vaccines to prioritize high risk individuals. The information system incorporates state-level data of the population with underlying medical conditions, demographics, overall state income, education level, and state mobility to formulate a mortality index. State-level data on the number of vaccines available and doses already administered are also incorporated into the information system to generate a vaccine index. The mortality and vaccine indices for each state are coupled to generate a vaccine priority ranking which can be used to advise vaccine distribution.The prototype can successfully link the data described above to a map of the US and then color code states according to the vaccine priority ranking. Implementation of this prototype will enable optimal vaccine distribution and reduce instances of severe or fatal COVID-19 illnesses as well as reduce costs associated with oversupply of vaccines in a single region. Future work will focus on improving the granularity of data down to the county-level, as well as increasing the scope of the system to the global scale. Additionally, the team plans to expand the application space of this information system to other diseases.
Lucas Standaert, Thilanka Munasinghe, Dongyoung Jang, Catherine Tate, Sean Lossef, Wilson Wong
IEEE BigData2
2020 Exploratory Data Analysis to Understand Social Determinants Important to Global Neonatal Mortality Rate
abstract
The Sustainable Development Goals (SDGs) are a set of targets that the UN hopes all countries will reach by 2030 broadly spanning the range of health, education, racial inequalities, environmental protections, and several other fields. Among these goals includes (Goal 3.2) an aim for all countries to reduce Neonatal Mortality Rates (NMR) to 12 per 1,000 live births. Without properly allocating resources to see the most dramatic shifts in NMR, many countries may be at risk of not meeting these ambitious goals. However, there are many factors which may influence national NMR, and while much previous work has been done to identify factors that influence NMR usually on a nation by nation basis, these factors can tend to vary. The goal of this study is to find factors that consistently lead, by changing them, to a change in NMR for many countries, in order to better inform health policy and resource allocations to the medical sector. This study will serve as an exploratory data analysis step for future studies regarding the impact of several health indicators on NMR per country. Cross-sectional data from the year 2014 were used for this Exploratory Data Analysis (EDA). To identify indicators that showed significant differences between the countries with high NMR and countries with low NMR, Mann-Whitney U Tests were performed. The p-value for each mean comparison was less than the 0.01 significance level. We have built a K-means clustering model to observe the variables' contribution to NMR, as well as a K-means clustering model to observe the same data's contributions to Gross Domestic Product (GDP), to see if both NMR and GDP follow similar trends across our target countries. The clustering for NMR groups of countries showed mostly separate clusters, while the clustering for the same data for the GDP classes showed very little separation, as the most points from each class all occupied the same cluster. To determine the actual amount that each indicator contributed to the data, Principle Component Analysis (PCA) was performed to understand the strongest contributions to the total data variance. The results of this study will serve to highlight the most important areas which must be improved in order to fulfill the Sustainable Development Goals (SDG) by the end of the next decade and to contribute to future studies that utilize longitudinal or more recent data.
Joshua Chuah, Thilanka Munasinghe
IEEE BigData2
2020 Exploring Various Applicable Techniques to Detect Smoke on the Satellite Images
abstract
Every year, the wildfires ravage broad areas of natural forest and nearby regions, causing substantial financial and life losses and deteriorating the air quality. More air hazards are emitted to the atmosphere reaching as high as the stratosphere propagating through the air currents. With the aggravation of climate change, wildfires of either human or natural cause could become more ferocious and devastating. A feasible solution is to detect the wildfire and respond early before the fire spread becomes irreversible. Satellite imagery serves as a cost-effective means to update near-real-time holistic landscape views of land and sea over extended periods. Such an advantage makes early fire detection and warning even in remote areas possible. The rendered images provided by the satellites’ various instruments incorporate various channels to provide real and artificial colors to reveal landscape details imperceptible to the naked eyes. This imagery dataset discussed in this paper derives from NASA’s Aqua and Terra satellites and make available at the NASA-IMPACT data share repository [1]. The dataset totals 704 images of cropped frames and their labeled images taken during both satellites’ extensive flyby observations. The images also contain spatial-temporal information serving as relevant metadata for analysis. This paper provides a survey of the recent advances in neural network-based object detection techniques followed by machine learning and deep learning-based methods to detect and localize smoke. A comprehensive elaboration of the datasets follows the method overview.
Chau-Lin Charly Huang, Thilanka Munasinghe
IEEE BigData2
2020 Investigating the Relationship Between High-Latitude Dust and Precipitation
abstract
The effect of dust on precipitation at high latitudes has not been adequately explored as most of this research has been restricted to mid- and low-latitude regions. Even so, at these lower latitudes the relationship between dust and precipitation still has not been distinctly established. The purpose of this study is to act as a starting point for understanding the role dust plays in precipitation processes at high latitudes. Knowledge of the effects dust can have on precipitation will allow for better prediction and forecasting through an improvement of regional and global climate models. In order to investigate this relationship, we looked at instances of high-latitude dust and precipitation amounts specifically in regions of Alaska, Iceland and Patagonia. The precipitation data were separated in terms of days where dust was present, versus absent (binary basis). Most of the precipitation averages were low with the highest precipitation amounts seen only on days without dust. The data were also separated into four specific days surrounding the dust event (4-day basis); the day before the dust event, the day of, one day after and two days after. The day of and one day after were limited to having lower levels of precipitation whereas the day before and two days after exhibited a broader range of precipitation levels, which included the highest amounts. These initial observations indicated more of a negative relationship between the high-latitude dust events and precipitation. Two classifications of precipitation intensity were created and tested for the purpose of this study. After which, correlation analyses (Chi-Square Test, Fisher's Exact Test and Rank Correlations) and classification modeling (Decision Tree modeling) were applied. The p-values from the correlation analyses were less than 0.05 suggesting that there is likely a statistically significant relationship between the high-latitude dust events and precipitation, and this relationship again appears to be slightly negative. The classification models were able to predict the precipitation amounts for the days surrounding the dust events using the two classifications. However, they were not specifically able to distinguish between each of the categories for the classifications as most of the data were in the lower categories.
Ariane Maharaj, Thilanka Munasinghe
IEEE BigData2
2020 Exploring Image Segmentation Techniques to Detect High Latitude Dust on Satellite Images
abstract
This paper presents a proposal to detect the High-Latitude Dust (HLD) events on the images captured by NASA satellites. We plan to use semantic segmentation techniques [1] to identify the HLD regions at the pixel level. The proposed model would distribute labels to specific pixels that we need to identify. After analyzing and learning those attributes, the trained model would identify the regions in new images with HLD events. However, in model training, the imbalanced data leads to a gap in the model's performance in coastal areas and sea areas, which means that even though the performance of dust detection over the ocean is good, the model cannot correctly detect the dust near or over the coast. Moreover, a gap exists between the different areas with different dust density. Thus, in this paper, we summarize the shortcomings of basic semantic segmentation models for HLD detection and propose a couple of solutions to improve such semantic segmentation models' performance.
Thilanka Munasinghe
IEEE BigData2
2019 IoT Application Development Using MIT App Inventor to Collect and Analyze Sensor Data
abstract
The rapid development of low-cost sensors, smart devices, communication networks, and learning algorithms has enabled data-driven decision making in large-scale systems. However, the development platforms for such Internet of Things (IoT) applications and collecting the data in a cohesive, yet simple manner is not very well understood. MIT App Inventor [1] is an open-source, user-friendly interface to develop mobile applications and has been used by over ten million users worldwide. We have added IoT capability to the MIT App Inventor platform where people can build applications using various sensors and IoT platforms such as Raspberry Pi [2] and Android Things [3] to collect the data through mobile applications. This poster paper presents a realization of easy to program IoT applications in the mobile application space with a focus on collecting data from the connected sensors. We have integrated the Android Things platform with the MIT App Inventor and introduced the existing MIT App Inventor community to the IoT based App development. This integration allows not only the seasoned developers, but also the novice developers to make interesting IoT applications with minimal programming knowledge.
Thilanka Munasinghe, Evan W. Patton, Oshani Seneviratne
IEEE BigData1
2019 Identifying the Relationship between Precipitation and Zika Outbreaks in Argentina
abstract
Dengue, Malaria and Zika are vector-borne diseases caused by mosquitoes that carry the parasites that lead to illnesses. According to the World Health Organization (WHO) hundreds of thousands of people around the world die every year due to disease-transmitting mosquitoes [6]. Mosquito outbreaks occur most commonly in warm climates, in areas close to the equator and tropical regions. Female mosquitoes lay their eggs in the ponds and puddles where water accumulates due to rainfall. In 2016, there was a significant spike in the number of dengue cases in Argentina; there were 79,455 cases of dengue reported in 2016 compared to 3250 cases in 2014 and 4774 in 2015 and went back down to the hundreds in 2017 and rose to thousands again in 2018 [1]. Since Dengue and Zika are spread by the same species of mosquito, Aedes aegypti, [2], the goal of this project is to examine whether the spike in dengue cases in 2016 in Argentina also led to a spike in Zika cases. During this study we will determine if there is a correlation between the number of Zika cases and rainfall precipitation levels. For the analysis, we use available Zika data for Argentina obtained from the Centers for Disease Control and Prevention (CDC) database [8]. More specifically, we are looking at the number of cases per month at a county level in Argentina. The precipitation data is obtained from the National Aeronautics and Space Administration (NASA) through the Global Precipitation Measurement Mission (GPM) [7]. Just like the Zika data, precipitation data is also on a monthly basis and data is available at a county level. There are existing systems for forecasting the outbreak of vector-borne diseases in various countries and each system looks at a particular factor. For example, the Dengue forecasting MOdel Satellite-based System (D-MOSS) is a system that issues warnings of dengue outbreaks eight months before outbreaks are likely to occur in Vietnam [3]. Another existing early warning system is the Predictive fLUshing Mosquito (PLUM) model. PLUM was developed by Draper Scientists in collaboration with scientists from Boston University and the Massachusetts Institute of Technology (MIT) to help predict and decrease outbreaks of dengue fever using observations collected in Singapore, Peru, and Puerto Rico [4]. Our work is inspired by D-MOSS, but instead of including water availability as a component in the prediction and focusing on Vietnam, we are interested in finding out the relationship between Zika outbreaks and precipitation levels in Argentina. The Sustainable Development Goals of the United Nations aim to address global problems of peace, justice, gender equality, good health and many others [5]. Similar to how D-MOSS targets these UN Sustainability goals [3], our project aspires to bring a similar approach when dealing with the Zika Virus. We will produce a report of our analysis results and visualizations that can be used by beneficiaries in Argentina to help in the efforts of combating and controlling the spread of Zika virus.
Lilian Ngweta, Karan Bhanot, Ariane Maharaj, Ian Bogle, Thilanka Munasinghe
IEEE BigData5
2019 Airline Miles Redemption
abstract
The business of Airline firms has deviated from their main business of flying passengers over the past decade. Now they have diversified into other lines of business as well. The revenue model has therefore changed over the past decade. Airline miles is one of the main revenue generating venture for airlines currently. It has been mentioned that airline miles business has turned cash cow for these firms. Their normal way of business, flying passengers, is not an attractive method of running business for them. Selling airline miles allow them to generate a higher revenue. We look at how the redemption of airline miles affect the bottom line of the company using data from publicly available data sources.
Joseph Sebastian, Thilanka Munasinghe
IEEE BigData2