VLDB 2026 Research / reviewers in the wild / expert
Silvina Caíno-Lores
dblp:155/4941
· DBLP profile ↗
15ranked-venue papers
4as first author
10since 2021 · last 2024
0000-0002-6922-0138ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 7 · 7 since 2021Systems, architecture and hardware · 6 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Workflow Provenance in the Computing Continuum for Responsible, Trustworthy, and Energy-Efficient AIabstractAs Artificial Intelligence (AI) becomes more pervasive in our society, it is crucial to develop, deploy, and assess Responsible and Trustworthy AI (RTAI) models, i.e., those that consider not only accuracy but also other aspects, such as explainability, fairness, and energy efficiency. Workflow provenance data have historically enabled critical capabilities towards RTAI. Provenance data derivation paths contribute to responsible workflows through transparency in tracking artifacts and resource consumption. Provenance data are well-known for their trustworthiness helping explainability, reproducibility, and accountability. However, there are complex challenges to achieve RTAI, which are further complicated by the heterogeneous infrastructure in the computing continuum (Edge-Cloud-HPC) used to develop and deploy models. As a result, a significant research and development gap remains between workflow provenance data management and RTAI. In this paper, we present a vision of the pivotal role of workflow provenance in supporting RTAI and discuss related challenges. We present a schematic view between RTAI and provenance, and highlight open research directions. Renan Souza 0001, Silvina Caíno-Lores, Mark Coletti, Tyler J. Skluzacek, Alexandru Costan, Frédéric Suter, Marta Mattoso, Rafael Ferreira da Silva |
e-Science | 2 |
| 2023 | Online Boosted Gaussian Learners for In-Situ Detection and Characterization of Protein Folding States in Molecular Dynamics SimulationsabstractMolecular Dynamics (MD) simulations are a crucial tool for understanding how proteins fold. In its easiest form, MD simulations can be scaled through data parallelism, this means that multiple folding trajectories can be spawned and executed in parallel, facilitating a more efficient exploration of the protein folding space. However, due to data dependencies, the analysis of MD simulations remains largely as a centralized process. In this work, we propose a data parallel, lightweight technique to learn the characteristics of protein folding states in MD simulations. Contrary to other methods, ours can differentiate relevant states in a single protein folding trajectory without requiring centralized global knowledge of the protein dynamics. As its processing and memory overheads are negligible (in the order of milliseconds per window of frames, and kilo bytes respectively) this technique can be coupled with the simulation for in-situ analysis. Harshita Sahni, Hector Carrillo-Cabada, Ekaterina D. Kots, Silvina Caíno-Lores, Jack D. Marquez, Ewa Deelman, Michel A. Cuendet, Harel Weinstein, Michela Taufer, Trilce Estrada |
e-Science | 4 |
| 2023 | Composable Workflow for Accelerating Neural Architecture Search Using In Situ Analytics for Protein ClassificationabstractNeural architecture search (NAS), which automates the design of neural network (NN) architectures for scientific datasets, requires significant computational resources and time — often on the order of days or weeks of GPU hours and training time. We design the Analytics for Neural Network (A4NN) workflow, a composable workflow that significantly reduces the time and resources required to design accurate and efficient NN architectures. We introduce a parametric fitness prediction strategy and distribute training across multiple accelerators to decrease the aggregated NN training time. A4NN rigorously record neural architecture histories, model states, and metadata to reproduce the search for near-optimal NNs. We demonstrate A4NN’s ability to reduce training time and resource consumption on a dataset generated by an X-ray Free Electron Laser (XFEL) experiment simulation. When deploying A4NN, we decrease training time by up to 37% and epochs required by up to 38%. Georgia Channing, Ria Patel, Paula Olaya, Ariel Keller Rorabaugh, Osamu Miyashita, Silvina Caíno-Lores, Catherine D. Schuman, Florence Tama, Michela Taufer |
ICPP | 6 |
| 2023 | Performance assessment of ensembles of in situ workflows under resource constraintsabstractSummary Scientific breakthroughs in biomolecular methods and improvements in hardware technology have shifted from a long‐running simulation to a large set of shorter simulations running simultaneously, called an ensemble. In an ensemble, simulations are usually coupled with analyses of data produced by the simulations. In situ methods can be used to analyze large volumes of data generated by scientific simulations at runtime (i.e., simulations and analyses are performed concurrently). In this work, we study the execution of ensemble‐based simulations paired with in situ analyses using in‐memory staging methods. Using an ensemble of molecular dynamics in situ workflows with multiple simulations and analyses, we first show that collecting traditional metrics such as makespan, instructions per cycle, memory usage, or cache miss ratio is not sufficient to characterize complex behaviors of ensembles. We propose a method to evaluate the performance of ensembles of workflows that captures multiple resource usage aspects: resource efficiency, resource allocation, and resource provisioning. Experimental results demonstrate that the proposed method can effectively distinguish the performance of different component placements in an ensemble with up to 32 ensemble members. By evaluating different co‐location scenarios, our proposed performance indicators demonstrate benefits of co‐locating simulation and coupled analyses within a compute node. Tu Mai Anh Do, Loïc Pottier, Rafael Ferreira da Silva, Silvina Caíno-Lores, Michela Taufer, Ewa Deelman |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | Reproducing and Extending Analytical Performance Models of Generalized Hierarchical SchedulingabstractWorkflows in High-Performance Computing (HPC) are rapidly changing towards more complex and large-scale workflows. In particular, high-throughput and ensemble workflows are becoming increasingly common. These workflows impose significant burden on current HPC scheduling systems which typically use slow, centralized schedulers. Generalized hierarchical scheduling (GHS) is a potential solution to face modern workflows but is not widely adopted in HPC yet. One difficulty hindering widespread adoption is the lack of performance models to configure and fit application requirements. The few existing models are often built on stick assumptions that can substantially reduce the analysis realism. In this paper, we reproduce the analysis and improve the realism of a state-of-the-art model presented in “An Analytical Performance Model of Generalized Hierarchical Scheduling” [1] by Herbein and co-authors. Specifically, we first reproduce four key analysis studies in the original paper and then expand the model by removing key assumptions, one at a time. In doing so, we extend the realism of the original model. We empirically validate our extended model using three different scenarios and discuss the observed accuracy. Jakob Lüttgau, Silvina Caíno-Lores, Kae Suarez, Dong H. Ahn, Stephen Herbein, Michela Taufer |
e-Science | 2 |
| 2022 | Identifying Structural Properties of Proteins from X-ray Free Electron Laser Diffraction PatternsabstractCapturing structural information of a biological molecule is crucial to determine its function and understand its mechanics. X-ray Free Electron Lasers (XFEL) are an experimental method used to create diffraction patterns (images) that can reveal structural information. In this work we design, implement, and evaluate XPSI (X-ray Free Electron Laser-based Protein Structure Identifier), a framework capable of predicting three structural properties in molecules (i.e., orientation, conformation, and protein type) from their diffraction patterns. XPSI predicts these properties with high accuracy in challenging scenarios, such as recognizing orientations despite symmetries in diffraction patterns, distinguishing conformations even when they have similar structures, and identifying protein types under different noise conditions. Our framework shows low computational cost and high prediction accuracy compared to other machine learning methods such as random forest and neural networks. Paula Olaya, Silvina Caíno-Lores, Vanessa Lama, Ria Patel, Ariel Keller Rorabaugh, Osamu Miyashita, Florence Tama, Michela Taufer |
e-Science | 2 |
| 2022 | A Methodology to Generate Efficient Neural Networks for Classification of Scientific DatasetsabstractNeural networks (NNs) are increasingly utilized in high-throughput scientific workflows. In this context, NN efficiency is essential for successful workflow management. We use a multi-objective Neural Architecture Search (NAS), NSGA-Net, to search for highly accurate NNs while optimizing for efficient use of computational resources by minimizing FLoating-point Operations Per Second (FLOPS). We define a domain-agnostic methodology to generate NNs with the support of NSGA-Net, select promising NNs that balance accuracy and FLOPS usage, and refine a subset of NNs in order to curate networks suitable for efficient data analysis. We apply this methodology to a protein diffraction use case. Preliminary results show NNs that efficiently classify conformation of proteins with a final accuracy of 97.7% or higher and using only 187 FLOPS. Ria Patel, Ariel Keller Rorabaugh, Paula Olaya, Silvina Caíno-Lores, Georgia Channing, Catherine D. Schuman, Osamu Miyashita, Florence Tama, Michela Taufer |
e-Science | 4 |
| 2022 | Ubique: A New Model for Untangling Inter-task Data Dependence in Complex HPC WorkflowsabstractExploiting task parallelism is getting increasingly difficult for diverse and complex scientific workflows running on High Performance Computing (HPC) systems. In this paper, we argue that the difficulty rises from a void in the spectrum of existing data-transfer models for resolving inter-task data dependence within a workflow and propose a novel model to fill that gap: Ubique. The Ubique model combines the best from in-transit and in situ models in order for loosely coupled producer and consumer tasks to run concurrently and to resolve their data dependencies efficiently with little or no modifications to their codes, striking a balance between transparent optimization, productivity, and performance. Our preliminary evaluation suggests that Ubique can significantly outperform the parallel file system (PFS)-based model while offering automatic data transfer and synchronization which are the features lacking in many traditional models. It also identifies the performance characteristics of its key depending subsystems, which must be understood for further broadening its benefits. Jae-Seung Yeom, Dong H. Ahn, Ian Lumsden, Jakob Lüttgau, Silvina Caíno-Lores, Michela Taufer |
e-Science | 5 |
| 2022 | Building High-Throughput Neural Architecture Search Workflows via a Decoupled Fitness Prediction EngineabstractNeural networks (NN) are used in high-performance computing and high-throughput analysis to extract knowledge from datasets. Neural architecture search (NAS) automates NN design by generating, training, and analyzing thousands of NNs. However, NAS requires massive computational power for NN training. To address challenges of efficiency and scalability, we proposePENGUIN, a decoupled fitness prediction engine that informs the search without interfering in it.PENGUINuses parametric modeling to predict fitness of NNs. Existing NAS methods and parametric modeling functions can be plugged intoPENGUINto build flexible NAS workflows. Through this decoupling and flexible parametric modeling,PENGUINreduces training costs: it predicts the fitness of NNs, enabling NAS to terminate training NNs early. Early termination increases the number of NNs that fixed compute resources can evaluate, thus giving NAS additional opportunity to find better NNs. We assess the effectiveness of our engine on 6,000 NNs across three diverse benchmark datasets and three state of the art NAS implementations using the Summit supercomputer. Augmenting these NAS implementations withPENGUINcan increase throughput by a factor of 1.6 to 7.1. Furthermore, walltime tests indicate thatPENGUINcan reduce training time by a factor of 2.5 to 5.3. Ariel Keller Rorabaugh, Silvina Caíno-Lores, J. Travis Johnston, Michela Taufer |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | A Case Study in Scientific Reproducibility from the Event Horizon Telescope (EHT)abstractThis poster presents the first results of an interdisciplinary project aiming to develop and share sustainable knowledge necessary to analyze, understand, and use published scientific results to advance reproducibility in multi-messenger astrophysics. Specifically, the project targets breakthrough work associated with the First M87 Event Horizon Telescope (EHT) and delivers recommendations on how the published results of the first black hole can be effectively reproduced. The project has the potential to advance new discovery in multi-messenger astrophysics by providing guidance for generalizing methods and findings from use cases. Ross Ketron, Jacob Leonard, Brandan Roachell, Ria Patel, R. White, Silvina Caíno-Lores, Nigel Tan, Patrick R. Miles, Karan Vahi, Ewa Deelman, Duncan A. Brown, Michela Taufer |
e-Science | 6 |
| 2020 | Applying big data paradigms to a large scale scientific workflow: Lessons learned and future directions
Silvina Caíno-Lores, Andrei Lapin, Jesús Carretero 0001, Peter G. Kropf |
Future Gener. Comput. Syst. | 1 |
| 2018 | Spark-DIY: A Framework for Interoperable Spark Operations with High Performance Block-Based Data ModelsabstractToday's scientific applications are increasingly relying on a variety of data sources, storage facilities, and computing infrastructures, and there is a growing demand for data analysis and visualization for these applications. In this context, exploiting Big Data frameworks for scientific computing is an opportunity to incorporate high-level libraries, platforms, and algorithms for machine learning, graph processing, and streaming; inherit their data awareness and fault-tolerance; and increase productivity. Nevertheless, limitations exist when Big Data platforms are integrated with an HPC environment, namely poor scalability, severe memory overhead, and huge development effort. This paper focuses on a popular Big Data framework -Apache Spark- and proposes an architecture to support the integration of highly scalable MPI block-based data models and communication patterns with a map-reduce-based programming model. The resulting platform preserves the data abstraction and programming interface of Spark, without conducting any changes in the framework, but allows the user to delegate operations to the MPI layer. The evaluation of our prototype shows that our approach integrates Spark and MPI efficiently at scale, so end users can take advantage of the productivity facilitated by the rich ecosystem of high-level Big Data tools and libraries based on Spark, without compromising efficiency and scalability. Silvina Caíno-Lores, Jesús Carretero 0001, Bogdan Nicolae, Orcun Yildiz, Tom Peterka |
BDCAT | 1 |
| 2017 | Data-Aware Support for Hybrid HPC and Big Data ApplicationsabstractNowadays there is a raising interest in bridging the gap between Big Data application models and data-intensive HPC. This work explores the effects that Big Data-inspired paradigms could have in current scientific applications through the evaluation of a real-world application from the hydrology domain. This evaluation led to experience that portrayed the key aspects of the HPC and Big Data paradigms that made them successful in their respective worlds. With this information, we established a research roadmap to build a platform suitable for HPC hybrid applications, with a focus on efficient data management and fault-tolerance. Silvina Caíno-Lores, Florin Isaila, Jesús Carretero 0001 |
CCGrid | 1 |
| 2016 | Methodological Approach to Data-Centric Cloudification of Scientific Iterative Workflows
Silvina Caíno-Lores, Andrei Lapin, Peter G. Kropf, Jesús Carretero 0001 |
ICA3PP | 1 |
| 2015 | A Multi-Objective Simulator for Optimal Power Dimensioning on Electric Railways using Cloud ComputingabstractPower dimensioning and energy saving have been traditionally two main issues regarding the deployment of
electric grids. Electric railways are also concerned about these issues, and simulators have been traditionally
used to test such infrastructure deployments. The main goal of this paper is to present the Railway electric
Power Consumption Simulator, a simulation model and tool for the railway energy provisioning problem. This
simulator aims to propose electric railway infrastructure deployments, optimizing the quality of the electric
flow supplied to train, as well as saving as much energy as possible. The paper describes the simulator
structure, as well as the ontology used to translate railway infrastructure elements into an electric circuit.
Because these two objectives are conflicting, a multi-objective optimization problem is formulated and solved.
Finally, a standard railway scenario is used to illustrate the capabilities of the tool, trying to find the best
electric substation placements in order to optimize such objectives. The evaluation shows how the tool can
handle hundreds of simulated scenarios using Cloud Computing techniques. Jesús Carretero 0001, Silvina Caíno-Lores, Félix García Carballeira, Alberto García Fernández |
SIMULTECH | 2 |