Diego Arenas

dblp:154/6535 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0001-7829-6102ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Missing data as augmentation in the Earth Observation domain: A multi-view learning approach
abstract
Multi-view learning (MVL) leverages multiple sources or views of data to enhance machine learning model performance and robustness. This approach has been successfully used in the Earth Observation (EO) domain, where views have a heterogeneous nature and can be affected by missing data. Despite the negative effect that missing data has on model predictions, the ML literature has used it as an augmentation technique to improve model generalization, like masking the input data. Inspired by this, we introduce novel methods for EO applications tailored to MVL with missing views. Our methods integrate the combination of a set to simulate all combinations of missing views as different training samples. Instead of replacing missing data with a numerical value, we use dynamic merge functions, like average, and more complex ones like Transformer. This allows the MVL model to entirely ignore the missing views, enhancing its predictive robustness. We experiment on four EO datasets with temporal and static views, including state-of-the-art methods from the EO domain. The results indicate that our methods improve model robustness under conditions of moderate missingness, and improve the predictive performance when all views are present. The proposed methods offer a single adaptive solution to operate effectively with any combination of available views.
Francisco Alejandro Mena, Diego Arenas, Andreas Dengel 0001
Neurocomputing2
2024 Impact Assessment of Missing Data in Model Predictions for Earth Observation Applications
abstract
Earth observation (EO) applications involving complex and heterogeneous data sources are commonly approached with machine learning models. However, there is a common assumption that data sources will be persistently available. Different situations could affect the availability of EO sources, like noise, clouds, or satellite mission failures. In this work, we assess the impact of missing temporal and static EO sources in trained models across four datasets with classification and regression tasks. We compare the predictive quality of different methods and find that some are naturally more robust to missing data. The Ensemble strategy, in particular, achieves a prediction robustness up to 100%. We evidence that missing scenarios are significantly more challenging in regression than classification tasks. Finally, we find that the optical view is the most critical view when it is missing individually.
Francisco Alejandro Mena, Diego Arenas, Marcela Charfuelan, Marlon Nuske, Andreas Dengel 0001
IGARSS2
2023 Crop Yield Prediction: An Operational Approach to Crop Yield Modeling on Field and Subfield Level with Machine Learning Models
abstract
Accurate and reliable crop yield prediction is a complex task. The yield of a crop depends on a variety of factors whose accurate measurement and modeling is challenging. At the same time, reliable yield prediction is highly desirable for farmers to optimize crop production. In this paper, we introduce a modeling based on remote sensing data and Machine Learning models evaluated on a large-scale dataset to address the challenge of an operational crop yield estimation and forecasting on field and subfield level. With our approach, we aim towards a global yield modeling based on Machine Learning models which operates across crop types without the need for crop-specific modeling. We demonstrate that our approach learns to map in-field variability for all studied crop types. Overall, the predictions have an error (RRMSE) of around 15% and an R2value of 0.77 at field level.
Patrick Helber, Benjamin Bischke, Peter Habelitz, Cristhian Sanchez, Deepak Pathak, Miro Miranda, Hiba Najjar, Francisco Alejandro Mena, Jayanth Siddamsetty, Diego Arenas, Michaela Vollmer, Marcela Charfuelan, Marlon Nuske, Andreas Dengel 0001
IGARSS10
2023 A Comparative Assessment of Multi-View Fusion Learning For Crop Classification
abstract
With a rapidly increasing amount and diversity of remote sensing (RS) data sources, there is a strong need for multi-view learning modeling. This is a complex task when considering the differences in resolution, magnitude, and noise of RS data. The typical approach for merging multiple RS sources has been input-level fusion, but other - more advanced - fusion strategies may outperform this traditional approach. This work assesses different fusion strategies for crop classification in the CropHarvest dataset. The fusion methods proposed in this work outperform models based on individual views and previous fusion methods. We do not find one single fusion method that consistently outperforms all other approaches. Instead, we present a comparison of multi-view fusion methods for three different datasets and show that, depending on the test region, different methods obtain the best performance. Despite this, we suggest a preliminary criterion for the selection of fusion methods.
Francisco Alejandro Mena, Diego Arenas, Marlon Nuske, Andreas Dengel 0001
IGARSS2
2023 Feature Attribution Methods for Multivariate Time-Series Explainability in Remote Sensing
abstract
Numerous remote sensing applications rely on temporal satellite data, and Deep learning models are increasingly being used for such tasks. Nevertheless, these models operate as black boxes, lacking transparency and understandability. We address this gap by using explainable AI on an agricultural task. Specifically, we trained a recurrent neural network on individual pixels from multispectral time-series of Sentinel-2 satellite images to predict crop yield. We then applied nine feature attribution methods on a sample of the dataset and computed the spectral and temporal contributions to the final individual predictions. The aggregated results were evaluated qualitatively and quantitatively. Results suggest that LIME and Shapley sampling value methods performed best on the quantitative scores, followed by GradientShap. Most backpropagation-based techniques had highly inconsistent scores across the explained data points. Finally, to guide remote sensing practitioners in using Explainable AI on similar datasets, we further discuss some selection criteria to be considered.
Hiba Najjar, Patrick Helber, Benjamin Bischke, Peter Habelitz, Cristhian Sanchez, Francisco Alejandro Mena, Miro Miranda, Deepak Pathak, Jayanth Siddamsetty, Diego Arenas, Michaela Vollmer, Marcela Charfuelan, Marlon Nuske, Andreas Dengel 0001
IGARSS10
2023 Predicting Crop Yield with Machine Learning: An Extensive Analysis of Input Modalities and Models on a Field and Sub-Field Level
abstract
We introduce a simple yet effective early fusion method for crop yield prediction that handles multiple input modalities with different temporal and spatial resolutions. We use high-resolution crop yield maps as ground truth data to train crop and machine learning model agnostic methods at the sub-field level. We use Sentinel-2 satellite imagery as the primary modality for input data with other complementary modalities, including weather, soil, and DEM data. The proposed method uses input modalities available with global coverage, making the framework globally scalable. We explicitly highlight the importance of input modalities for crop yield prediction and emphasize that the best-performing combination of input modalities depends on region, crop, and chosen model.
Deepak Pathak, Miro Miranda, Francisco Alejandro Mena, Cristhian Sanchez, Patrick Helber, Benjamin Bischke, Peter Habelitz, Hiba Najjar, Jayanth Siddamsetty, Diego Arenas, Michaela Vollmer, Marcela Charfuelan, Marlon Nuske, Andreas Dengel 0001
IGARSS10
2023 Influence of Data Cleaning Techniques on Sub-Field Yield Predictions
abstract
Modern combine harvesters can collect geo-located real-time yield measurement while harvesting. This data can be used to train Machine Learning models that predict the yield at sub-field level based on remote sensing input data. The performance of these models is, however, highly dependent on the quality of the yield data. It is therefore important to develop automatic cleaning techniques to correct for common errors in combine harvester yield maps. In this work, we compare different combinations of data cleaning techniques by evaluating their impact on the yield-prediction model performance at field and sub-field level. Our findings indicate that basic cleaning techniques such as absolute thresholds are sufficient at the field level, whereas the performance at the sub-field level is enhanced through the utilization of more intricate statistical cleaning methods.
Cristhian Sanchez, Deepak Pathak, Miro Miranda, Marcela Charfuelan, Patrick Helber, Marlon Nuske, Benjamin Bischke, Peter Habelitz, Nafisur Rahman, Francisco Alejandro Mena, Hiba Najjar, Jayanth Siddamsetty, Diego Arenas, Michaela Vollmer, Andreas Dengel 0001
IGARSS13