VLDB 2026 Research / reviewers in the wild / expert
Rodrigo Vargas
dblp:170/0542
· DBLP profile ↗
11ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0001-6829-5333ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GEOtiled-SG: A Scalable Framework for High-Resolution Terrain Parameter ComputationabstractGEOtiled enables efficient computation of high-resolution terrain parameters from digital elevation models (DEMs) by decomposing large regions into smaller, parallelizable tiles. Originally developed as GEOtiled-G with support for three parameters (Slope, Aspect, Hillshade) via GDAL, we present GEOtiled-SG, an enhanced version that integrates the SAGA GIS library to compute over 15 parameters, expanding its utility in Earth science. To offset SAGA’s computational overhead, GEOtiled-SG introduces three optimizations: concurrent DEM cropping, buffer-aware mosaicking, and unified concurrency across workflow stages. Evaluations show that GEOtiled-SG maintains GEOtiled-G’s performance on the original parameters and offers consistent speedups across the expanded set. The framework is open source on GitHub, with data hosted on Dataverse, supporting reproducible, scalable terrain analysis. Gabriel Laboy, Ian Lumsden, Paula Olaya, Jack D. Marquez, Kin Wai Ng, Rodrigo Vargas, Michela Taufer |
eScience | 6 |
| 2025 | Advancing the GEOtiled Framework Through Scalable Terrain Parameter ComputationabstractThe GEOtiled framework facilitates the scalable and efficient computation of high-resolution terrain parameters using Digital Elevation Models (DEMs) across the Continental United States (CONUS). These parameters are essential in Earth Science applications. This paper presents significant advances to optimizing GEOtiled's performance. In GEOtiled 2.0, we introduce three key optimizations to speed up GEOtiled's overall runtime: concurrent cropping, efficient mosaic operation, and unified parallel processing. These optimizations improve data distribution and handling, reducing computation times while maintaining accuracy. We evaluate performance using DEM data at 30-meter resolution over a region covering the state of Tennessee. Our results demonstrate that the optimized framework provides substantial performance improvements for generating terrain parameters. GEOtiled 2.0 software, openly available on GitHub, advances reproducible, data-driven scientific discovery in Earth Science. Gabriel Laboy, Paula Olaya, Jack D. Marquez, Michael Sutherlin, Rodrigo Vargas, Michela Taufer |
HPDC | 5 |
| 2024 | Multi-Modal Transformer for Compressive LiDARs Using Hyperspectral Imaging Side-InformationabstractCompressive satellite LiDAR (CS-LiDAR) has been recently introduced as a radically different computational sensing and reconstruction approach for LiDAR sensing of Earth. It is based on NASA’s adaptive wavelength scanning LiDAR (AWSL) system. Unlike conventional 1D LiDAR methods, CS-LiDAR utilizes sparse coded laser illumination across a 2D field-of-view. The aim is to compressively capture Earth from hundreds of kilometers above, enabling computational 3D imagery reconstruction with resolution that is comparable to that attained with data collected from just hundreds of meters. The forward imaging model captures the light propagation phenomena affecting the photon pulses transmitted from the sensor to the Earth’s surface and back. This work enhances CS-LiDAR by integrating imaging spectroscopy into a multimodal system and employing a transformer network for the inverse imaging problem, driven by multimodal attention mechanisms. Emulations enabled by enormous observational LiDAR data of Earth, available from NASA’s G-LiHT imaging observatory, highlight the efficacy of methods developed. Nestor Porras-Diaz, Andres Ramirez-Jaime, Gonzalo R. Arce, Rodrigo Vargas, David J. Harding, Mark Stephen, James MacKinnon |
IGARSS | 4 |
| 2024 | Transformer End-to-End Optimization of Compressive LiDARs Using Imaging Spectroscopy Side InformationabstractCompressive satellite LiDAR (CS-LiDAR) has been recently introduced as a radically different computational sensing and reconstruction approach for LiDAR sensing of Earth. It is based on NASA’s adaptive wavelength scanning LiDAR (AWSL) system. Rather than measuring 1D line footprints over a satellite’s swath path as is the norm today, CS-LiDAR adopts sparse coded laser illumination over a 2D wide field-of-view. The objective is to compressively sense Earth from hundreds of km above Earth to then computationally reconstruct the 3D imagery with resolution and coverage as if the data was collected from just hundreds of meters in height. The forward imaging model captures the light propagation phenomena affecting the photon pulses transmitted from the sensor to the Earth’s surface and back. This paper advances CS-LiDAR on many fronts. First, imaging spectroscopy side-information, often jointly available with LiDARs, is integrated into a multimodal imaging system. Secondly, the inverse imaging problem is cast under a transformer network architecture driven by multimodal attention mechanisms. Finally, by directing the snapshot spectral cameras in front of the LiDAR, the transformer mechanisms autonomously adjust the LiDAR’s beam scanning to focus on specific target locations, thus attaining end-to-end optimal adaptive sampling that can respond to varying observational conditions, surface events, and scientific priorities. Emulations enabled by enormous observational LiDAR data of Earth, available from NASA’s G-LiHT imaging observatory, show the advantages attained by methods developed in this work. Nestor Porras-Diaz, Andres Ramirez-Jaime, Gonzalo R. Arce, Karelia Pena-Pena, David J. Harding, Mark Stephen, James MacKinnon, Rodrigo Vargas |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2023 | Enabling Scalability in the Cloud for Scientific Workflows: An Earth Science Use CaseabstractScientific discovery increasingly relies on interoperable, multimodular workflows generating intermediate data. The complexity of managing intermediate data may cause performance losses or unexpected costs. This paper defines an approach to composing these scientific workflows on cloud services, focusing on workflow data orchestration, management, and scalability. We demonstrate the effectiveness of our approach with the SOMOSPIE scientific workflow that deploys machine learning (ML) models to predict high-resolution soil moisture using an HPC service (LSF) and an open-source cloud-native service (K8s) and object storage. Our approach enables scientists to scale from coarse-grained to fine-grained resolution and from a small to a larger region of interest. Using our empirical observations, we generate a cost model for the execution of workflows with hidden intermediate data on cloud services. Paula Olaya, Jakob Lüttgau, Camila Roa, Ricardo M. Llamas, Rodrigo Vargas, Sophia Wen, I-Hsin Chung, Seetharami R. Seelam, Yoonho Park, Jay F. Lofstead, Michela Taufer |
CLOUD | 5 |
| 2023 | GEOtiled: A Scalable Workflow for Generating Large Datasets of High-Resolution Terrain ParametersabstractTerrain parameters such as slope, aspect, and hillshading are essential in various applications, including agriculture, forestry, and hydrology. However, generating high-resolution terrain parameters is computationally intensive, making it challenging to provide these value-added products to communities in need. We present a scalable workflow called GEOtiled that leverages data partitioning to accelerate the computation of terrain parameters from digital elevation models, while preserving accuracy. We assess our workflow in terms of its accuracy and wall time by comparing it to SAGA, which is highly accurate but slow to generate results, and to GDAL, which supports memory optimizations but not data parallelism. We obtain a coefficient of determination (R^2) between GEOtiled and SAGA of 0.794, ensuring accuracy in our terrain parameters. We achieve an X6 speedup compared to GDAL when generating the terrain parameters at a high-resolution (10 m) for the Contiguous United States (CONUS). Camila Roa, Paula Olaya, Ricardo M. Llamas, Rodrigo Vargas, Michela Taufer |
HPDC | 4 |
| 2023 | Building Trust in Earth Science Findings through Data Traceability and Results ExplainabilityabstractTo trust findings in computational science, scientists need workflows that trace the data provenance and support results explainability. As workflows become more complex, tracing data provenance and explaining results become harder to achieve. In this paper, we propose a computational environment that automatically creates a workflow execution's record trail and invisibly attaches it to the workflow's output, enabling data traceability and results explainability. Our solution transforms existing container technology, includes tools for automatically annotating provenance metadata, and allows effective movement of data and metadata across the workflow execution. We demonstrate the capabilities of our environment with the study of SOMOSPIE, an earth science workflow. Through a suite of machine learning modeling techniques, this workflow predicts soil moisture values from the 27 km resolution satellite data down to higher resolutions necessary for policy making and precision agriculture. By running the workflow in our environment, we can identify the causes of different accuracy measurements for predicted soil moisture values in different resolutions of the input data and link different results to different machine learning methods used during the soil moisture downscaling, all without requiring scientists to know aspects of workflow design and implementation. Paula Olaya, Dominic Kennedy, Ricardo M. Llamas, Leobardo Valera, Rodrigo Vargas, Jay F. Lofstead, Michela Taufer |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | Augmenting Singularity to Generate Fine-grained Workflows, Record Trails, and Data ProvenanceabstractThe use of containerization technology in high performance computing (HPC) workflows has substantially increased recently because it makes workflows much easier to develop and deploy. Although many HPC workflows include multiple data and multiple applications, they have traditionally all been bundled together into one monolithic container. This hinders the ability to trace the thread of execution, thus preventing scientists from establishing data provenance, or having workflow reproducibility. To provide a solution to this problem we extend the functionality of a popular HPC container runtime, Singularity. We implement both the ability to compose fine-grained containerized workflows and execute these workflows within the Singularity runtime with automatic metadata collection. Specifically, the new functionality collects a record trail of execution and creates data provenance. The use of our augmented Singularity is demonstrated with an earth science workflow, SOMOSPIE. The workflow is composed via our augmented Singularity which creates fine-grained containers and collects the metadata to trace, explain, and reproduce the prediction of soil moisture at a fine resolution. Dominic Kennedy, Paula Olaya, Jay F. Lofstead, Rodrigo Vargas, Michela Taufer |
e-Science | 4 |
| 2019 | SOMOSPIE: A Modular SOil MOisture SPatial Inference Engine Based on Data-Driven DecisionsabstractThe current availability of soil moisture data over large areas comes from satellite remote sensing technologies (i.e., radar-based systems), but these data have coarse resolution and often exhibit large spatial information gaps. Where data are too coarse or sparse for a given need (e.g., precision farming), one can leverage machine-learning techniques coupled with other sources of environmental information (e.g., topography) to generate gap-free information at a finer spatial resolution (i.e., increased granularity). To this end, we develop a spatial inference engine consisting of modular stages for processing spatial environmental data, generating predictions with machine-learning techniques, and analyzing these predictions. We demonstrate the functionality of this approach and the effects of data processing choices via multiple prediction maps over a United States ecological region with a highly diverse soil moisture profile (i.e., the Middle Atlantic Coastal Plains). The relevance of our work derives from a pressing need to improve the spatial representation of soil moisture for applications in environmental sciences (e.g., ecological niche modeling, carbon monitoring systems, and other Earth system models) and precision farming (e.g., optimizing irrigation practices and other land management decisions). Danny Rorabaugh, Mario Guevara, Ricardo M. Llamas, Joy Kitson, Rodrigo Vargas, Michela Taufer |
eScience | 5 |
| 2017 | Data analytics for modeling soil moisture patterns across united states ecoclimatic domainsabstractOur poster presents a data analytics strategy to enable scientists to model patterns of soil moisture data at different resolutions across the United States. We build upon previous work of Guevara and co-authors with three contributions. First, we introduce divisions of soil moisture into the climatic regions proposed by the National Ecology Observatory Network. Second, we reduce the topological parameters used in modeling soil moisture using Principal Component Analysis. Third, we present an efficient workflow for modeling and visualizing soil moisture data. Thomas Kitson, Paula Olaya, Elizabeth Racca, Michael R. Wyatt II, Mario Guevara, Rodrigo Vargas, Michela Taufer |
IEEE BigData | 6 |
| 2015 | From HPC Performance to Climate Modeling: Transforming Methods for HPC Predictions into Models of Extreme Climate ConditionsabstractIn the past forty years, the high-performance computing (HPC) community has been developing powerful and rigorous tools for predicting the performance of supercomputers from log traces. In this paper, we transform one of these approaches previously used for predicting idle resources in high-end clusters into a method for capturing extreme climate events in geographical locations of interest. Our method uses an analysis based on empirical cumulative distribution functions (ECDFs) to benchmark and model occurrences of climate events including extreme temperature and precipitation. The method comprises two phases: a learning phase and a prediction phase. The learning phase applies the ECDF-based empirical analysis to historical climate data in order to identify suitable modeling and forecasting windows, both given in years. The prediction phase applies the modeling window to the most recent climate data in order to estimate the likelihood that given portions of the region of interest can experience extreme climate events in the forecasting window. The research is the first of its kind to extend HPC performance modeling techniques to study extreme climate events. Ryan McKinney, Vivek K. Pallipuram, Rodrigo Vargas, Michela Taufer |
e-Science | 3 |