EDBT 2026 Demo / reviewers in the wild / expert
Paula Olaya
dblp:213/1518
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-0258-6861ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GEOtiled-SG: A Scalable Framework for High-Resolution Terrain Parameter ComputationabstractGEOtiled enables efficient computation of high-resolution terrain parameters from digital elevation models (DEMs) by decomposing large regions into smaller, parallelizable tiles. Originally developed as GEOtiled-G with support for three parameters (Slope, Aspect, Hillshade) via GDAL, we present GEOtiled-SG, an enhanced version that integrates the SAGA GIS library to compute over 15 parameters, expanding its utility in Earth science. To offset SAGA’s computational overhead, GEOtiled-SG introduces three optimizations: concurrent DEM cropping, buffer-aware mosaicking, and unified concurrency across workflow stages. Evaluations show that GEOtiled-SG maintains GEOtiled-G’s performance on the original parameters and offers consistent speedups across the expanded set. The framework is open source on GitHub, with data hosted on Dataverse, supporting reproducible, scalable terrain analysis. Gabriel Laboy, Ian Lumsden, Paula Olaya, Jack D. Marquez, Kin Wai Ng, Rodrigo Vargas, Michela Taufer |
eScience | 3 |
| 2025 | Advancing the GEOtiled Framework Through Scalable Terrain Parameter ComputationabstractThe GEOtiled framework facilitates the scalable and efficient computation of high-resolution terrain parameters using Digital Elevation Models (DEMs) across the Continental United States (CONUS). These parameters are essential in Earth Science applications. This paper presents significant advances to optimizing GEOtiled's performance. In GEOtiled 2.0, we introduce three key optimizations to speed up GEOtiled's overall runtime: concurrent cropping, efficient mosaic operation, and unified parallel processing. These optimizations improve data distribution and handling, reducing computation times while maintaining accuracy. We evaluate performance using DEM data at 30-meter resolution over a region covering the state of Tennessee. Our results demonstrate that the optimized framework provides substantial performance improvements for generating terrain parameters. GEOtiled 2.0 software, openly available on GitHub, advances reproducible, data-driven scientific discovery in Earth Science. Gabriel Laboy, Paula Olaya, Jack D. Marquez, Michael Sutherlin, Rodrigo Vargas, Michela Taufer |
HPDC | 2 |
| 2024 | PerSSD: Persistent, Shared, and Scalable Data with Node-Local Storage for Scientific Workflows in Cloud InfrastructureabstractComputational workflows need to retain data from both intermediate stages and final results to ensure the reproducibility and trustworthiness of scientific discoveries. While cloud infrastructure offers advantages like elasticity and automation, it compromises the persistence of intermediate data to ensure performance and reduce costs. Utilizing node-local storage can enhance performance but requires manual data transfers to persistent storage, making the technique impractical. To address these challenges, we propose a software architecture called Persistent, Shared, and Scalable Data (PerSSD) that integrates cloud operators and a Network File System (NFS) to make node-local data persistent and shareable across cloud nodes while ensuring performance. PerSSD outperforms traditional cloud object storage, achieving 35% reduction in the overall execution time of an earth science workflow, all while ensuring data persistence and shareability. Paula Olaya, Sophia Wen, Jay F. Lofstead, Michela Taufer |
IEEE Big Data | 1 |
| 2023 | Enabling Scalability in the Cloud for Scientific Workflows: An Earth Science Use CaseabstractScientific discovery increasingly relies on interoperable, multimodular workflows generating intermediate data. The complexity of managing intermediate data may cause performance losses or unexpected costs. This paper defines an approach to composing these scientific workflows on cloud services, focusing on workflow data orchestration, management, and scalability. We demonstrate the effectiveness of our approach with the SOMOSPIE scientific workflow that deploys machine learning (ML) models to predict high-resolution soil moisture using an HPC service (LSF) and an open-source cloud-native service (K8s) and object storage. Our approach enables scientists to scale from coarse-grained to fine-grained resolution and from a small to a larger region of interest. Using our empirical observations, we generate a cost model for the execution of workflows with hidden intermediate data on cloud services. Paula Olaya, Jakob Lüttgau, Camila Roa, Ricardo M. Llamas, Rodrigo Vargas, Sophia Wen, I-Hsin Chung, Seetharami R. Seelam, Yoonho Park, Jay F. Lofstead, Michela Taufer |
CLOUD | 1 |
| 2023 | GEOtiled: A Scalable Workflow for Generating Large Datasets of High-Resolution Terrain ParametersabstractTerrain parameters such as slope, aspect, and hillshading are essential in various applications, including agriculture, forestry, and hydrology. However, generating high-resolution terrain parameters is computationally intensive, making it challenging to provide these value-added products to communities in need. We present a scalable workflow called GEOtiled that leverages data partitioning to accelerate the computation of terrain parameters from digital elevation models, while preserving accuracy. We assess our workflow in terms of its accuracy and wall time by comparing it to SAGA, which is highly accurate but slow to generate results, and to GDAL, which supports memory optimizations but not data parallelism. We obtain a coefficient of determination (R^2) between GEOtiled and SAGA of 0.794, ensuring accuracy in our terrain parameters. We achieve an X6 speedup compared to GDAL when generating the terrain parameters at a high-resolution (10 m) for the Contiguous United States (CONUS). Camila Roa, Paula Olaya, Ricardo M. Llamas, Rodrigo Vargas, Michela Taufer |
HPDC | 2 |
| 2023 | Composable Workflow for Accelerating Neural Architecture Search Using In Situ Analytics for Protein ClassificationabstractNeural architecture search (NAS), which automates the design of neural network (NN) architectures for scientific datasets, requires significant computational resources and time — often on the order of days or weeks of GPU hours and training time. We design the Analytics for Neural Network (A4NN) workflow, a composable workflow that significantly reduces the time and resources required to design accurate and efficient NN architectures. We introduce a parametric fitness prediction strategy and distribute training across multiple accelerators to decrease the aggregated NN training time. A4NN rigorously record neural architecture histories, model states, and metadata to reproduce the search for near-optimal NNs. We demonstrate A4NN’s ability to reduce training time and resource consumption on a dataset generated by an X-ray Free Electron Laser (XFEL) experiment simulation. When deploying A4NN, we decrease training time by up to 37% and epochs required by up to 38%. Georgia Channing, Ria Patel, Paula Olaya, Ariel Keller Rorabaugh, Osamu Miyashita, Silvina Caíno-Lores, Catherine D. Schuman, Florence Tama, Michela Taufer |
ICPP | 3 |
| 2023 | Building Trust in Earth Science Findings through Data Traceability and Results ExplainabilityabstractTo trust findings in computational science, scientists need workflows that trace the data provenance and support results explainability. As workflows become more complex, tracing data provenance and explaining results become harder to achieve. In this paper, we propose a computational environment that automatically creates a workflow execution's record trail and invisibly attaches it to the workflow's output, enabling data traceability and results explainability. Our solution transforms existing container technology, includes tools for automatically annotating provenance metadata, and allows effective movement of data and metadata across the workflow execution. We demonstrate the capabilities of our environment with the study of SOMOSPIE, an earth science workflow. Through a suite of machine learning modeling techniques, this workflow predicts soil moisture values from the 27 km resolution satellite data down to higher resolutions necessary for policy making and precision agriculture. By running the workflow in our environment, we can identify the causes of different accuracy measurements for predicted soil moisture values in different resolutions of the input data and link different results to different machine learning methods used during the soil moisture downscaling, all without requiring scientists to know aspects of workflow design and implementation. Paula Olaya, Dominic Kennedy, Ricardo M. Llamas, Leobardo Valera, Rodrigo Vargas, Jay F. Lofstead, Michela Taufer |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Augmenting Singularity to Generate Fine-grained Workflows, Record Trails, and Data ProvenanceabstractThe use of containerization technology in high performance computing (HPC) workflows has substantially increased recently because it makes workflows much easier to develop and deploy. Although many HPC workflows include multiple data and multiple applications, they have traditionally all been bundled together into one monolithic container. This hinders the ability to trace the thread of execution, thus preventing scientists from establishing data provenance, or having workflow reproducibility. To provide a solution to this problem we extend the functionality of a popular HPC container runtime, Singularity. We implement both the ability to compose fine-grained containerized workflows and execute these workflows within the Singularity runtime with automatic metadata collection. Specifically, the new functionality collects a record trail of execution and creates data provenance. The use of our augmented Singularity is demonstrated with an earth science workflow, SOMOSPIE. The workflow is composed via our augmented Singularity which creates fine-grained containers and collects the metadata to trace, explain, and reproduce the prediction of soil moisture at a fine resolution. Dominic Kennedy, Paula Olaya, Jay F. Lofstead, Rodrigo Vargas, Michela Taufer |
e-Science | 2 |
| 2022 | Identifying Structural Properties of Proteins from X-ray Free Electron Laser Diffraction PatternsabstractCapturing structural information of a biological molecule is crucial to determine its function and understand its mechanics. X-ray Free Electron Lasers (XFEL) are an experimental method used to create diffraction patterns (images) that can reveal structural information. In this work we design, implement, and evaluate XPSI (X-ray Free Electron Laser-based Protein Structure Identifier), a framework capable of predicting three structural properties in molecules (i.e., orientation, conformation, and protein type) from their diffraction patterns. XPSI predicts these properties with high accuracy in challenging scenarios, such as recognizing orientations despite symmetries in diffraction patterns, distinguishing conformations even when they have similar structures, and identifying protein types under different noise conditions. Our framework shows low computational cost and high prediction accuracy compared to other machine learning methods such as random forest and neural networks. Paula Olaya, Silvina Caíno-Lores, Vanessa Lama, Ria Patel, Ariel Keller Rorabaugh, Osamu Miyashita, Florence Tama, Michela Taufer |
e-Science | 1 |
| 2022 | A Methodology to Generate Efficient Neural Networks for Classification of Scientific DatasetsabstractNeural networks (NNs) are increasingly utilized in high-throughput scientific workflows. In this context, NN efficiency is essential for successful workflow management. We use a multi-objective Neural Architecture Search (NAS), NSGA-Net, to search for highly accurate NNs while optimizing for efficient use of computational resources by minimizing FLoating-point Operations Per Second (FLOPS). We define a domain-agnostic methodology to generate NNs with the support of NSGA-Net, select promising NNs that balance accuracy and FLOPS usage, and refine a subset of NNs in order to curate networks suitable for efficient data analysis. We apply this methodology to a protein diffraction use case. Preliminary results show NNs that efficiently classify conformation of proteins with a final accuracy of 97.7% or higher and using only 187 FLOPS. Ria Patel, Ariel Keller Rorabaugh, Paula Olaya, Silvina Caíno-Lores, Georgia Channing, Catherine D. Schuman, Osamu Miyashita, Florence Tama, Michela Taufer |
e-Science | 3 |
| 2022 | NSDF-Cloud: Enabling Ad-Hoc Compute Clusters Across Academic and Commercial CloudsabstractComputational resources are increasingly provisioned to users through cloud-like interfaces. Both academic and commercial cloud offerings exist, but no single standardized interface for common actions such as configuration, launching, and termination of virtual resources exists. This imposes huge technical burden on domain scientist that attempt to take advantage of these resources; even expert users spend considerable time to port their applications from one cloud platform to another. Jakob Lüttgau, Paula Olaya, Naweiluo Zhou, Giorgio Scorzelli, Valerio Pascucci, Michela Taufer |
HPDC | 2 |
| 2022 | NSDF-FUSE: A Testbed for Studying Object Storage via FUSE File SystemsabstractThis work presents NSDF-FUSE, a testbed for evaluating settings and performance of FUSE-based file systems on top of S3-compatible object storage; the testbed is part of a suite of services from the National Science Data Fabric (NSDF) project (an NSF-funded project that is delivering cyberinfrastructures for data scientists). We demonstrate how NSDF-FUSE can be deployed to evaluate eight different mapping packages that mount S3-compatible object storage to a file system, as well as six data patterns representing different I/O operations on two cloud platforms. NSDF-FUSE is open-source and can be easily extended to run with other software mapping packages and different cloud platforms. Paula Olaya, Jakob Lüttgau, Naweiluo Zhou, Jay F. Lofstead, Giorgio Scorzelli, Valerio Pascucci, Michela Taufer |
HPDC | 1 |
| 2017 | Data analytics for modeling soil moisture patterns across united states ecoclimatic domainsabstractOur poster presents a data analytics strategy to enable scientists to model patterns of soil moisture data at different resolutions across the United States. We build upon previous work of Guevara and co-authors with three contributions. First, we introduce divisions of soil moisture into the climatic regions proposed by the National Ecology Observatory Network. Second, we reduce the topological parameters used in modeling soil moisture using Principal Component Analysis. Third, we present an efficient workflow for modeling and visualizing soil moisture data. Thomas Kitson, Paula Olaya, Elizabeth Racca, Michael R. Wyatt II, Mario Guevara, Rodrigo Vargas, Michela Taufer |
IEEE BigData | 2 |