Paula Olaya

dblp:213/1518 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-0258-6861ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 GEOtiled-SG: A Scalable Framework for High-Resolution Terrain Parameter Computation
abstract
GEOtiled enables efficient computation of high-resolution terrain parameters from digital elevation models (DEMs) by decomposing large regions into smaller, parallelizable tiles. Originally developed as GEOtiled-G with support for three parameters (Slope, Aspect, Hillshade) via GDAL, we present GEOtiled-SG, an enhanced version that integrates the SAGA GIS library to compute over 15 parameters, expanding its utility in Earth science. To offset SAGA’s computational overhead, GEOtiled-SG introduces three optimizations: concurrent DEM cropping, buffer-aware mosaicking, and unified concurrency across workflow stages. Evaluations show that GEOtiled-SG maintains GEOtiled-G’s performance on the original parameters and offers consistent speedups across the expanded set. The framework is open source on GitHub, with data hosted on Dataverse, supporting reproducible, scalable terrain analysis.
Gabriel Laboy, Ian Lumsden, Paula Olaya, Jack D. Marquez, Kin Wai Ng, Rodrigo Vargas, Michela Taufer
eScience3
2025 Advancing the GEOtiled Framework Through Scalable Terrain Parameter Computation
abstract
The GEOtiled framework facilitates the scalable and efficient computation of high-resolution terrain parameters using Digital Elevation Models (DEMs) across the Continental United States (CONUS). These parameters are essential in Earth Science applications. This paper presents significant advances to optimizing GEOtiled's performance. In GEOtiled 2.0, we introduce three key optimizations to speed up GEOtiled's overall runtime: concurrent cropping, efficient mosaic operation, and unified parallel processing. These optimizations improve data distribution and handling, reducing computation times while maintaining accuracy. We evaluate performance using DEM data at 30-meter resolution over a region covering the state of Tennessee. Our results demonstrate that the optimized framework provides substantial performance improvements for generating terrain parameters. GEOtiled 2.0 software, openly available on GitHub, advances reproducible, data-driven scientific discovery in Earth Science.
Gabriel Laboy, Paula Olaya, Jack D. Marquez, Michael Sutherlin, Rodrigo Vargas, Michela Taufer
HPDC2
2024 PerSSD: Persistent, Shared, and Scalable Data with Node-Local Storage for Scientific Workflows in Cloud Infrastructure
abstract
Computational workflows need to retain data from both intermediate stages and final results to ensure the reproducibility and trustworthiness of scientific discoveries. While cloud infrastructure offers advantages like elasticity and automation, it compromises the persistence of intermediate data to ensure performance and reduce costs. Utilizing node-local storage can enhance performance but requires manual data transfers to persistent storage, making the technique impractical. To address these challenges, we propose a software architecture called Persistent, Shared, and Scalable Data (PerSSD) that integrates cloud operators and a Network File System (NFS) to make node-local data persistent and shareable across cloud nodes while ensuring performance. PerSSD outperforms traditional cloud object storage, achieving 35% reduction in the overall execution time of an earth science workflow, all while ensuring data persistence and shareability.
Paula Olaya, Sophia Wen, Jay F. Lofstead, Michela Taufer
IEEE Big Data1
2023 Enabling Scalability in the Cloud for Scientific Workflows: An Earth Science Use Case
abstract
Scientific discovery increasingly relies on interoperable, multimodular workflows generating intermediate data. The complexity of managing intermediate data may cause performance losses or unexpected costs. This paper defines an approach to composing these scientific workflows on cloud services, focusing on workflow data orchestration, management, and scalability. We demonstrate the effectiveness of our approach with the SOMOSPIE scientific workflow that deploys machine learning (ML) models to predict high-resolution soil moisture using an HPC service (LSF) and an open-source cloud-native service (K8s) and object storage. Our approach enables scientists to scale from coarse-grained to fine-grained resolution and from a small to a larger region of interest. Using our empirical observations, we generate a cost model for the execution of workflows with hidden intermediate data on cloud services.
Paula Olaya, Jakob Lüttgau, Camila Roa, Ricardo M. Llamas, Rodrigo Vargas, Sophia Wen, I-Hsin Chung, Seetharami R. Seelam, Yoonho Park, Jay F. Lofstead, Michela Taufer
CLOUD1
2023 GEOtiled: A Scalable Workflow for Generating Large Datasets of High-Resolution Terrain Parameters
abstract
Terrain parameters such as slope, aspect, and hillshading are essential in various applications, including agriculture, forestry, and hydrology. However, generating high-resolution terrain parameters is computationally intensive, making it challenging to provide these value-added products to communities in need. We present a scalable workflow called GEOtiled that leverages data partitioning to accelerate the computation of terrain parameters from digital elevation models, while preserving accuracy. We assess our workflow in terms of its accuracy and wall time by comparing it to SAGA, which is highly accurate but slow to generate results, and to GDAL, which supports memory optimizations but not data parallelism. We obtain a coefficient of determination (R^2) between GEOtiled and SAGA of 0.794, ensuring accuracy in our terrain parameters. We achieve an X6 speedup compared to GDAL when generating the terrain parameters at a high-resolution (10 m) for the Contiguous United States (CONUS).
Camila Roa, Paula Olaya, Ricardo M. Llamas, Rodrigo Vargas, Michela Taufer
HPDC2
2023 Composable Workflow for Accelerating Neural Architecture Search Using In Situ Analytics for Protein Classification
abstract
Neural architecture search (NAS), which automates the design of neural network (NN) architectures for scientific datasets, requires significant computational resources and time — often on the order of days or weeks of GPU hours and training time. We design the Analytics for Neural Network (A4NN) workflow, a composable workflow that significantly reduces the time and resources required to design accurate and efficient NN architectures. We introduce a parametric fitness prediction strategy and distribute training across multiple accelerators to decrease the aggregated NN training time. A4NN rigorously record neural architecture histories, model states, and metadata to reproduce the search for near-optimal NNs. We demonstrate A4NN’s ability to reduce training time and resource consumption on a dataset generated by an X-ray Free Electron Laser (XFEL) experiment simulation. When deploying A4NN, we decrease training time by up to 37% and epochs required by up to 38%.
Georgia Channing, Ria Patel, Paula Olaya, Ariel Keller Rorabaugh, Osamu Miyashita, Silvina Caíno-Lores, Catherine D. Schuman, Florence Tama, Michela Taufer
ICPP3
2023 Building Trust in Earth Science Findings through Data Traceability and Results Explainability
abstract
To trust findings in computational science, scientists need workflows that trace the data provenance and support results explainability. As workflows become more complex, tracing data provenance and explaining results become harder to achieve. In this paper, we propose a computational environment that automatically creates a workflow execution's record trail and invisibly attaches it to the workflow's output, enabling data traceability and results explainability. Our solution transforms existing container technology, includes tools for automatically annotating provenance metadata, and allows effective movement of data and metadata across the workflow execution. We demonstrate the capabilities of our environment with the study of SOMOSPIE, an earth science workflow. Through a suite of machine learning modeling techniques, this workflow predicts soil moisture values from the 27 km resolution satellite data down to higher resolutions necessary for policy making and precision agriculture. By running the workflow in our environment, we can identify the causes of different accuracy measurements for predicted soil moisture values in different resolutions of the input data and link different results to different machine learning methods used during the soil moisture downscaling, all without requiring scientists to know aspects of workflow design and implementation.
Paula Olaya, Dominic Kennedy, Ricardo M. Llamas, Leobardo Valera, Rodrigo Vargas, Jay F. Lofstead, Michela Taufer
IEEE Trans. Parallel Distributed Syst.1
2022 Augmenting Singularity to Generate Fine-grained Workflows, Record Trails, and Data Provenance
abstract
The use of containerization technology in high performance computing (HPC) workflows has substantially increased recently because it makes workflows much easier to develop and deploy. Although many HPC workflows include multiple data and multiple applications, they have traditionally all been bundled together into one monolithic container. This hinders the ability to trace the thread of execution, thus preventing scientists from establishing data provenance, or having workflow reproducibility. To provide a solution to this problem we extend the functionality of a popular HPC container runtime, Singularity. We implement both the ability to compose fine-grained containerized workflows and execute these workflows within the Singularity runtime with automatic metadata collection. Specifically, the new functionality collects a record trail of execution and creates data provenance. The use of our augmented Singularity is demonstrated with an earth science workflow, SOMOSPIE. The workflow is composed via our augmented Singularity which creates fine-grained containers and collects the metadata to trace, explain, and reproduce the prediction of soil moisture at a fine resolution.
Dominic Kennedy, Paula Olaya, Jay F. Lofstead, Rodrigo Vargas, Michela Taufer
e-Science2
2022 Identifying Structural Properties of Proteins from X-ray Free Electron Laser Diffraction Patterns
abstract
Capturing structural information of a biological molecule is crucial to determine its function and understand its mechanics. X-ray Free Electron Lasers (XFEL) are an experimental method used to create diffraction patterns (images) that can reveal structural information. In this work we design, implement, and evaluate XPSI (X-ray Free Electron Laser-based Protein Structure Identifier), a framework capable of predicting three structural properties in molecules (i.e., orientation, conformation, and protein type) from their diffraction patterns. XPSI predicts these properties with high accuracy in challenging scenarios, such as recognizing orientations despite symmetries in diffraction patterns, distinguishing conformations even when they have similar structures, and identifying protein types under different noise conditions. Our framework shows low computational cost and high prediction accuracy compared to other machine learning methods such as random forest and neural networks.
Paula Olaya, Silvina Caíno-Lores, Vanessa Lama, Ria Patel, Ariel Keller Rorabaugh, Osamu Miyashita, Florence Tama, Michela Taufer
e-Science1
2022 A Methodology to Generate Efficient Neural Networks for Classification of Scientific Datasets
abstract
Neural networks (NNs) are increasingly utilized in high-throughput scientific workflows. In this context, NN efficiency is essential for successful workflow management. We use a multi-objective Neural Architecture Search (NAS), NSGA-Net, to search for highly accurate NNs while optimizing for efficient use of computational resources by minimizing FLoating-point Operations Per Second (FLOPS). We define a domain-agnostic methodology to generate NNs with the support of NSGA-Net, select promising NNs that balance accuracy and FLOPS usage, and refine a subset of NNs in order to curate networks suitable for efficient data analysis. We apply this methodology to a protein diffraction use case. Preliminary results show NNs that efficiently classify conformation of proteins with a final accuracy of 97.7% or higher and using only 187 FLOPS.
Ria Patel, Ariel Keller Rorabaugh, Paula Olaya, Silvina Caíno-Lores, Georgia Channing, Catherine D. Schuman, Osamu Miyashita, Florence Tama, Michela Taufer
e-Science3
2022 NSDF-Cloud: Enabling Ad-Hoc Compute Clusters Across Academic and Commercial Clouds
abstract
Computational resources are increasingly provisioned to users through cloud-like interfaces. Both academic and commercial cloud offerings exist, but no single standardized interface for common actions such as configuration, launching, and termination of virtual resources exists. This imposes huge technical burden on domain scientist that attempt to take advantage of these resources; even expert users spend considerable time to port their applications from one cloud platform to another.
Jakob Lüttgau, Paula Olaya, Naweiluo Zhou, Giorgio Scorzelli, Valerio Pascucci, Michela Taufer
HPDC2
2022 NSDF-FUSE: A Testbed for Studying Object Storage via FUSE File Systems
abstract
This work presents NSDF-FUSE, a testbed for evaluating settings and performance of FUSE-based file systems on top of S3-compatible object storage; the testbed is part of a suite of services from the National Science Data Fabric (NSDF) project (an NSF-funded project that is delivering cyberinfrastructures for data scientists). We demonstrate how NSDF-FUSE can be deployed to evaluate eight different mapping packages that mount S3-compatible object storage to a file system, as well as six data patterns representing different I/O operations on two cloud platforms. NSDF-FUSE is open-source and can be easily extended to run with other software mapping packages and different cloud platforms.
Paula Olaya, Jakob Lüttgau, Naweiluo Zhou, Jay F. Lofstead, Giorgio Scorzelli, Valerio Pascucci, Michela Taufer
HPDC1
2017 Data analytics for modeling soil moisture patterns across united states ecoclimatic domains
abstract
Our poster presents a data analytics strategy to enable scientists to model patterns of soil moisture data at different resolutions across the United States. We build upon previous work of Guevara and co-authors with three contributions. First, we introduce divisions of soil moisture into the climatic regions proposed by the National Ecology Observatory Network. Second, we reduce the topological parameters used in modeling soil moisture using Principal Component Analysis. Third, we present an efficient workflow for modeling and visualizing soil moisture data.
Thomas Kitson, Paula Olaya, Elizabeth Racca, Michael R. Wyatt II, Mario Guevara, Rodrigo Vargas, Michela Taufer
IEEE BigData2