Assaf Anyamba

dblp:33/6064 · DBLP profile ↗
← Back
3ranked-venue papers in the field
0as first author
3since 2021 · last 2023
0000-0003-0932-9585ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3
YearPublicationVenuePosition
2023 From Satellites to Fields: Machine Learning Applications for Prediction of Corn Production Using NDVI, Precipitation and Land Surface Temperature for Large Producer Countries
abstract
In this work-in-progress study, we aim to determine the predictors of corn production in Iowa (United States), Heilongjiang (China), Mato Grosso (Brazil), Cordoba (Argentina), and Poltava Oblast (Ukraine). The effectiveness of predicting annual corn production using precipitation and land surface temperature (LST) values obtained from the Earth Engine tool, and Moderate Resolution Imaging Spectroradiometer (MODIS) Normalised Difference Vegetation Index (NDVI) values obtained from the National Aeronautics and Space Administration (NASA) Global Inventory Monitoring and Modelling Studies (GIMMS) Global Agricultural Monitoring System, is examined. A comparison is conducted between multiple linear regression, ridge regression, and lasso regression models. The highest adjusted R2values are found in order to identify the optimal model in which the corn yield variance is explained by a combination of variables including the year, NDVI sample and anomaly values, precipitation, and LST. The results show that corn production values in Iowa, Heilongjiang, Mato Grosso, Cordoba, and Poltava Oblast can best be predicted using lasso regression models with adjusted R2values of 0.841, 0.933, 0.847, 0.854, and 0.860 respectively.
Saniya Nangia, Thilanka Munasinghe, Heidi Tubbs, Assaf Anyamba
IEEE Big Data4
2023 Integrating Climate Variable Data in Machine Learning Models for Predictive Analytics of Tomato Yields in California
abstract
Traditionally, agricultural forecasting has relied on empirical methods and basic statistical analysis, such as applying average values from previous years’ yields or using a simple linear fit for next year’s predictions. However, the emergence of data-driven approaches, particularly machine learning algorithms, has revolutionized yield prediction in agriculture. Machine learning techniques have demonstrated their potential to provide accurate predictions. However, existing models often rely on a limited number of input variables for crop yield predictions, which makes them only suitable for specific scenarios. In this study, we have developed four distinct machine learning-based predictors, incorporating various climate factors, including daytime temperature, nighttime temperature, precipitation (rainfall), vegetation index, and evapotranspiration as input variables to predict tomato acreage yields in counties of California, USA. Our results show that regression models constructed using neural networks and linear regression exhibited better performance than other predictors, achieving an average accuracy rate of 70% to 80%. Compared to most of the existing crop yield predictors, our models offer versatility while maintaining a desirable level of predictive accuracy. Expanding the number of input variables, such as nitrogen fertilizer usage etc, and introducing larger spatial and temporal high-resolution datasets for model training can improve our model performance, enabling us to obtain better results in tomato yield prediction.
Tianze Zhu, Tingyi Tan, Shuheng Wang, Thilanka Munasinghe, Heidi Tubbs, Assaf Anyamba
IEEE Big Data7
2021 Scraping Unstructured Data to Explore the Relationship between Rainfall Anomalies and Vector-Borne Disease Outbreaks
abstract
According to the World Health Organization (WHO), vector-borne diseases such as malaria and dengue account for 17% of all infectious disease cases and lead to more than 700,000 deaths per year. Tracking and predicting the spread of vector-borne diseases is a vital task that could save hundreds of thousands of lives annually. Oftentimes, the first reports of vector-borne disease outbreaks occur through emails and online reporting systems long before they are officially documented. Tracking and predicting the emergence and spread of vector-borne disease outbreaks requires extracting data from these unstructured sources in combination with historical weather and climate data to understand the underlying background triggers and disease dynamics. In this work, we develop a data extraction pipeline for the online outbreak reporting website ProMED-mail that utilizes a web scraper, transformer neural network summarizer, and named entity recognizer to obtain a dataset of malaria, dengue, zika, and chikungunya outbreaks over the last 30 years. This scraped dataset was further analyzed in association with global rainfall anomalies derived from NASA’s Integrated Multi-satellitE Retrievals for GPM [Global Precipitation Mission] (IMERG) dataset. This preliminary analysis was to understand the effect of global rainfall patterns on the spread of vector-borne diseases. Analysis of the ProMED-mail and GPM data shows that vector-borne disease outbreaks are clustered towards the tropics and outbreaks are often amplified during the rainy seasons. Our scraped dataset can be a valuable tool in creating comprehensive georeferenced disease records for modeling and predicting future outbreaks.
Ethan Joseph, Thilanka Munasinghe, Heidi Tubbs, Bhaskar Bishnoi, Assaf Anyamba
IEEE BigData5