Germán T. Eizaguirre

dblp:284/8749 · also Germán Telmo Eizaguirre · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0002-2865-9873ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Serverless Data Analytics (Finally) Bridging the Gap: Introducing the Ortzi DataFrame
abstract
Serverless technologies have simplified distributed computing by streamlining resource management and offering out-of-the-box usability of cloud resources. However, serverless computing has yet to fully permeate the broader data analytics community. One of the major reasons causing this slow adoption is the lack of a handy serverless interface for seamlessly running recurrent workloads in the cloud. To fill this gap, we introduce in this work the missing piece in serverless analytics: the Ortzi Dataframe, a practical and intuitive programming abstraction that mirrors pandas DataFrames, so that users can effortlessly run their local, single-threaded Python code at scale in the cloud. Needless to say, such a powerful abstraction is certainly useless if not backed by a serverless analytics system that can operate over it in parallel. For this reason, another major contribution of this paper is a fully-fledged system that can run jobs in parallel across the cloud continuum using the novel Ortzi Dataframes. The new system leverages the specific capabilities of each serverless backend without user intervention. Our evaluation demonstrates that Ortzi enables exploration of the nuanced trade-offs of heterogeneous backends with min-imal programming changes and overhead. By harnessing the seamless nature of Ortzi, we optimize jobs through strategic backend selection, still delivering a user-friendly open source framework for programmers without cloud expertise.
Germán T. Eizaguirre, Marc Hostau, Marc Sánchez Artigas
CLOUD1
2025 Quantifying Serverless Elasticity: The gumeter Benchmark Suite
Germán T. Eizaguirre, Enrique Molina-Giménez, Gerard Finol, Carlos Molina 0004, Pedro García López
ICSOC (1)1
2024 A Seer knows best: Auto-tuned object storage shuffling for serverless analytics
Germán T. Eizaguirre, Marc Sánchez Artigas
J. Parallel Distributed Comput.1
2023 Is Performance of Object Storage Predictable for Serverless I/O Workloads? A Comparative Study
abstract
Serverless architectures abstract resource provisioning away from the user. However, this property may be at odds with performance. One example of this is Function as a Service (FaaS), where the lack of network addressability compels developers to resort to serverless storage services such as AWS S3 to share (intermediate) data between the functions. For IO-bound workflows, the literature has shown that the performance of parallel reads and writes highly depends on the level of parallelism. Simply put, both an excess or a deficiency in the number of functions may lead to longer IO times. The good news is that the provisioning of functions is fast. Consequently, it is feasible to auto-provision the serverless functions to the optimal number to minimize IO latency. For this, the performance of object storage must be predictable and consistent. We confirmed this in the past for IBM COS. And in this paper, we show that the same occurs to AWS S3. Concretely, we prove that the optimal level of parallelism for parallel reads and writes can be approximated analytically for AWS S3.
Germán T. Eizaguirre, Marc Sánchez Artigas
ICNP1
2022 A seer knows best: optimized object storage shuffling for serverless analytics
abstract
Serverless platforms offer high resource elasticity and pay-as-you-go billing, making them a compelling choice for data analytics. To craft a "pure" serverless solution, the common practice is to transfer intermediate data between serverless functions via serverless object storage (IBM COS; AWS S3). However, prior works have led to inconclusive results about the performance of object storage, since they have left large margin for optimization. To verify that object storage has been underrated, we design a novel shuffle manager for serverless data analytics termed Seer. Specifically, Seer dynamically chooses between two shuffle algorithms to maximize performance. The algorithm choice is based on some predictive models, and very importantly, without users having to specify intermediate data sizes at the time of the job submission. We integrate Seer with PyWren-IBM [31], a serverless analytics framework, and evaluate it against both serverful (e.g., Spark) and serverless systems (e.g., Google BigQuery). Our results certify that our new shuffle manager can deliver performance improvements over them.
Marc Sánchez Artigas, Germán T. Eizaguirre
Middleware2