Ivo Jimenez

dblp:17/4858 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
3since 2021 · last 2022
0000-0002-2222-1985ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Storage systems · 50% Performance modeling and evaluation · 36% High-performance computing · 7%
Databases, data mining, and information retrieval
3 papers
Database system architecture and tuning · 78% Query processing and optimization · 22%
Computer networks
1 paper
Network measurement and analytics · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 18 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
0.612022
P-RECS'22: 5th International Workshop on Practical Reproducible Evaluation of Systems · HPDC 2022
Performance modeling and evaluation
performance variability
0.312018
Taming Performance Variability · OSDI 2018
Performance modeling and evaluation
workload characterization
0.312018
Taming Performance Variability · OSDI 2018
Storage systems › distributed storage
distributed shared log
0.312017
Malacology: A Programmable Storage System · EuroSys 2017
Storage systems
file systems
0.312017
Malacology: A Programmable Storage System · EuroSys 2017
Storage systems › file systems › distributed file system
metadata load balancing
0.312017
Malacology: A Programmable Storage System · EuroSys 2017
Storage systems › computational storage
programmable storage
0.312017
Malacology: A Programmable Storage System · EuroSys 2017
Storage systems
distributed storage
0.212016
DAOS and friends: a proposal for an exascale storage system · SC 2016
High-performance computing › supercomputing
exascale computing
0.212016
DAOS and friends: a proposal for an exascale storage system · SC 2016
Storage systems › distributed storage
exascale storage system
0.212016
DAOS and friends: a proposal for an exascale storage system · SC 2016
Computational science and engineering
computational reproducibility
0.212022
P-RECS'22: 5th International Workshop on Practical Reproducible Evaluation of Systems · HPDC 2022
Database system architecture and tuning
index tuning
0.112012
Kaizen: a semi-automatic index advisor · SIGMOD Conference 2012
Software maintenance and evolution
devops
0.112019
Creating repeatable, reusable experimentation pipelines with popper: tutorial · PPoPP 2019
Query processing and optimization
query optimization
0.112010
Data desensitization of customer data for use in optimizer performance experiments · ICDE 2010
Database system architecture and tuning › database design
physical database design
0.112009
QuickStart: An Upfront Client-Based Design Advisor for Parallel Data Warehouses · ICDE 2009
Cloud and datacenter computing
datacenter storage
0.112017
Malacology: A Programmable Storage System · EuroSys 2017
Storage systems
metadata management
0.112016
DAOS and friends: a proposal for an exascale storage system · SC 2016
Database system architecture and tuning
index recommendation
0.012012
Kaizen: a semi-automatic index advisor · SIGMOD Conference 2012

Methods — techniques the papers use, named apart from their topics

benchmarking · 0.9performance evaluation · 0.2string desensitization · 0.2numeric desensitization · 0.2heuristics · 0.1cost calculation · 0.1
YearPublicationVenuePosition
2022 Skyhook: Towards an Arrow-Native Storage System
abstract
With the ever-increasing dataset sizes, several file formats such as Parquet, ORC, and Avro have been developed to store data efficiently, save the network, and interconnect bandwidth at the price of additional CPU utilization. However, with the advent of networks supporting 25–100 Gb/s and storage devices delivering 1,000,000 reqs/sec, the CPU has become the bottleneck trying to keep up feeding data in and out of these fast devices. The result is that data access libraries executed on single clients are often CPU-bound and cannot utilize the scale-out benefits of distributed storage systems. One attractive solution to this problem is to offload data-reducing processing and filtering tasks to the storage layer. However, modifying legacy storage systems to support compute offloading is often tedious and requires an extensive understanding of the system internals. Previous approaches re-implemented functionality of data processing frameworks and access libraries for a particular storage system, a duplication of effort that might have to be repeated for different storage systems. This paper introduces a new design paradigm that allows ex-tending programmable object storage systems to embed existing, widely used data processing frameworks and access libraries into the storage layer with no modifications. In this approach, data processing frameworks and access libraries can evolve independently from storage systems while leveraging distributed storage systems' scale-out and availability properties. We present Skyhook, an example implementation of our design paradigm using Ceph, Apache Arrow, and Parquet. We provide a brief performance evaluation of Skyhook and discuss key results.
Jayjeet Chakraborty, Ivo Jimenez, Sebastiaan Alvarez Rodriguez, Alexandru Uta, Jeff LeFevre, Carlos Maltzahn
CCGRID2
2022 P-RECS'22: 5th International Workshop on Practical Reproducible Evaluation of Systems
abstract
The P-RECS workshop focuses heavily on practical, actionable aspects of reproducibility in broad areas of computational science and data exploration, with special emphasis on issues in which community collaboration can be essential for adopting novel methodologies, techniques and frameworks aimed at addressing some of the challenges we face today. The workshop brings together researchers and experts to share experiences and advance the state of the art in the reproducible evaluation of computer systems, featuring contributed papers and invited talks.
Jay F. Lofstead, Carlos Maltzahn, Ivo Jimenez
HPDC3
2021 Zero-Cost, Arrow-Enabled Data Interface for Apache Spark
abstract
Distributed data processing ecosystems are widespread and their components are highly specialized, such that efficient interoperability is urgent. Recently, Apache Arrow was chosen by the community to serve as a format mediator, providing efficient in-memory data representation. Arrow enables efficient data movement between data processing and storage engines, significantly improving interoperability and overall performance. In this work, we design a new zero-cost data interoperability layer between Apache Spark and Arrow-based data sources through the Arrow Dataset API. Our novel data interface helps separate the computation (Spark) and data (Arrow) layers. This enables practitioners to seamlessly use Spark to access data from all Arrow Dataset API-enabled data sources and frameworks. To benefit our community, we open-source our work and show that consuming data through Apache Arrow is zero-cost: our novel data interface is either on-par or more performant than native Spark.
Sebastiaan Alvarez Rodriguez, Jayjeet Chakraborty, Aaron Chu, Ivo Jimenez, Jeff LeFevre, Carlos Maltzahn, Alexandru Uta
IEEE BigData4
2020 Is Big Data Performance Reproducible in Modern Cloud Networks?
Alexandru Uta, Alexandru Custura, Dmitry Duplyakin, Ivo Jimenez, Jan S. Rellermeyer, Carlos Maltzahn, Robert Ricci, Alexandru Iosup
NSDI4
2019 Creating repeatable, reusable experimentation pipelines with popper: tutorial
abstract
Popper is an experimentation protocol for conducting scientific explorations and writing academic articles following a DevOps approach. The Popper CLI tool helps researchers automate the execution and validation of an experimentation pipeline. In this tutorial we give an introduction to the concepts and CLI tool, and go over hands-on exercises that help.
Ivo Jimenez, Jay F. Lofstead, Carlos Maltzahn
PPoPP1
2018 Cudele: An API and Framework for Programmable Consistency and Durability in a Global Namespace
abstract
HPC and data center scale application developers are abandoning POSIX IO because file system metadata synchronization and serialization overheads of providing strong consistency and durability are too costly - and often unnecessary - for their applications. Unfortunately, designing file systems with weaker consistency or durability semantics excludes applications that rely on stronger guarantees, forcing developers to re-write their applications or deploy them on a different system. We present a framework and API that lets administrators specify their consistency/durability requirements and dynamically assign them to subtrees in the same namespace, allowing administrators to optimize subtrees over time and space for different workloads. We show similar speedups to related work but more importantly, we show performance improvements when we custom fit subtree semantics to applications such as checkpoint-restart (91.7x speedup), user home directories (0.03 standard deviation from optimal), and users checking for partial results (2% overhead).
Michael Sevilla, Ivo Jimenez, Noah Watkins, Jeff LeFevre, Peter Alvaro, Shel Finkelstein, Patrick Donnelly, Carlos Maltzahn
IPDPS2
2018 Taming Performance Variability
Aleksander Maricq, Dmitry Duplyakin, Ivo Jimenez, Carlos Maltzahn, Ryan Stutsman, Robert Ricci, Ana Klimovic
OSDI3
2018 quiho: Automated Performance Regression Testing Using Inferred Resource Utilization Profiles
abstract
We introduce quiho, a framework for profiling application performance that can be used in automated performance regression tests. quiho profiles an application by applying sensitivity analysis, in particular statistical regression analysis (SRA), using application-independent performance feature vectors that characterize the performance of machines. The result of the SRA, feature importance specifically, is used as a proxy to identify hardware and low-level system software behavior. The relative importance of these features serve as a performance profile of an application (termed inferred resource utilization profile or IRUP), which is used to automatically validate performance behavior across multiple revisions of an application»s code base without having to instrument code or obtain performance counters. We demonstrate that quiho can successfully discover performance regressions by showing its effectiveness in profiling application performance for synthetically introduced regressions as well as those found in real-world applications.
Ivo Jimenez, Noah Watkins, Michael Sevilla, Jay F. Lofstead, Carlos Maltzahn
ICPE1
2017 Malacology: A Programmable Storage System
abstract
Storage systems need to support high-performance for special-purpose data processing applications that run on an evolving storage device technology landscape. This puts tremendous pressure on storage systems to support rapid change both in terms of their interfaces and their performance. But adapting storage systems can be difficult because unprincipled changes might jeopardize years of code-hardening and performance optimization efforts that were necessary for users to entrust their data to the storage system. We introduce the programmable storage approach, which exposes internal services and abstractions of the storage stack as building blocks for higher-level services. We also build a prototype to explore how existing abstractions of common storage system services can be leveraged to adapt to the needs of new data processing systems and the increasing variety of storage devices. We illustrate the advantages and challenges of this approach by composing existing internal abstractions into two new higher-level services: a file system metadata load balancer and a high-performance distributed shared-log. The evaluation demonstrates that our services inherit desirable qualities of the back-end storage system, including the ability to balance load, efficiently propagate service metadata, recover from failure, and navigate trade-offs between latency and throughput using leases.
Michael Sevilla, Noah Watkins, Ivo Jimenez, Peter Alvaro, Shel Finkelstein, Jeff LeFevre, Carlos Maltzahn
EuroSys3
2017 DeclStore: Layering Is for the Faint of Heart
Noah Watkins, Michael Sevilla, Ivo Jimenez, Kathryn Dahlgren, Peter Alvaro, Shel Finkelstein, Carlos Maltzahn
HotStorage3
2016 DAOS and friends: a proposal for an exascale storage system
abstract
The DOE Extreme-Scale Technology Acceleration Fast Forward Storage and IO Stack project is going to have significant impact on storage systems design within and beyond the HPC community. With phase two of the project starting, it is an excellent opportunity to explore the complete design and how it will address the needs of extreme scale platforms. This paper examines each layer of the proposed stack in some detail along with cross-cutting topics, such as transactions and metadata management. This paper not only provides a timely summary of important aspects of the design specifications but also captures the underlying reasoning that is not available elsewhere. We encourage the broader community to understand the design, intent, and future directions to foster discussion guiding phase two and the ultimate production storage stack based on this work. An initial performance evaluation of the early prototype implementation is also provided to validate the presented design.
Jay F. Lofstead, Ivo Jimenez, Carlos Maltzahn, Quincey Koziol, John Bent, Eric Barton
SC2
2015 The Role of Container Technology in Reproducible Computer Systems Research
abstract
Evaluating experimental results in the field of computer systems is a challenging task, mainly due to the many changes in software and hardware that computational environments go through. In this position paper, we analyze salient features of container technology that, if leveraged correctly, can help reduce the complexity of reproducing experiments in systems research. We present a use case in the area of distributed storage systems to illustrate the extensions that we envision, mainly in terms of container management infrastructure. We also discuss the benefits and limitations of using containers as a way of reproducing research in other areas of experimental systems research.
Ivo Jimenez, Carlos Maltzahn, Adam Moody, Kathryn Mohror, Jay F. Lofstead, Remzi H. Arpaci-Dusseau, Andrea C. Arpaci-Dusseau
IC2E1
2015 RITA: an index-tuning advisor for replicated databases
abstract
Given a replicated database, a divergent design tunes the indexes in each replica differently in order to specialize it for a specific subset of the workload. Empirical studies have shown that this specialization brings significant performance gains compared to the common practice of having the same indexes in all replicas. However, reaping the benefits of divergent designs requires the development of new tuning tools for database administrators, and the existing tools unfortunately suffer from severe shortcomings: they assume a fixed number of replicas and a known workload distribution, and ignore the possibility of replica failures and the subsequent effect on load imbalance.
Quoc Trung Tran, Ivo Jimenez, Neoklis Polyzotis, Anastasia Ailamaki
SSDBM2
2014 An innovative storage stack addressing extreme scale platforms and Big Data applications
abstract
Current production HPC IO stack design is unlikely to offer sufficient features and performance to adequately serve extreme scale science platform requirements as well as Big Data problems. A joint effort between the US Department of Energy's Office of Advanced Simulation and Computing and Advanced Scientific Computing Research commissioned a project to develop a design and prototype for an IO stack suitable for the extreme scale environment. It will be referred to as the Fast Forward Storage and IO (FFSIO) project. This is a joint effort led by Lawrence Livermore National Laboratory, with the DOE Data Management Nexus leads Rob Ross and Gary Grider as coordinators and contract lead Mark Gary.
Jay F. Lofstead, Ivo Jimenez, Carlos Maltzahn, Quincey Koziol, John Bent, Eric Barton
CLUSTER2
2012 Kaizen: a semi-automatic index advisor
abstract
Index tuning; i.e., selecting indexes that are appropriate for the workload to obtain good system performance, is a crucial task for database administrators. Administrators rely on automated index advisors for this task, but existing advisors work either offline, requiring a-priori knowledge of the workload, or online, taking the administrator out of the picture and assuming total control of the index tuning task. Semi-automatic index tuning is a new paradigm that achieves a middle ground: the advisor analyzes the workload online and provides recommendations tailored to the current workload, and the administrator is able to provide feedback to refine future recommendations. In this demonstration we present Kaizen, an index tuning tool that implements semi-automatic tuning.
Ivo Jimenez, Huascar Sanchez, Quoc Trung Tran, Neoklis Polyzotis
SIGMOD Conference1
2010 Data desensitization of customer data for use in optimizer performance experiments
abstract
Improving the performance and functionality of database system optimizers requires experimentation on real customer data. Often these data are of sensitive nature and the only way to keep them is by applying a non-reversible transformation to obfuscate them. However, in order that the database optimizer generates exactly the same query plans as for the sensitive data, the transformation has to preserve the order and some important properties of the data distribution. Unfortunately, existing data obfuscation techniques do not preserve all of these properties and therefore are not applicable in this context. In this paper we present a Desensitizer tool that we have developed for optimizer performance experiments of HP's Neoview high availability data warehousing product. The tool is based on novel numeric and string desensitization algorithms which are agnostic to the database system. We explain the core concepts behind the algorithms, how they preserve the required data properties and important implementation considerations that were made. We present the architecture of the Desensitizer tool and results of the extensive validation that we conducted.
Malú Castellanos, Bin Zhang 0004, Ivo Jimenez, Perla Ruiz, Miguel Durazo, Umeshwar Dayal, Lily Jow
ICDE3
2009 QuickStart: An Upfront Client-Based Design Advisor for Parallel Data Warehouses
abstract
QuickStart is a tool to automate the physical design of data warehouses for HP's Neoview system. It has been researched and prototyped at HP Labs with close interaction from Neoview design experts. It embodies heuristics and best practices of the experts to search for candidate physical features and uses cost calculations to recommend the features that result in good designs. It has some unique characteristics that differentiate it from other physical design advisors. In particular, it is the only advisor that is client-based, does not require a DBMS server installation and can work off a laptop by simply connecting to the customers' flat files. Another unique characteristic of QuickStart is that it provides the rationale for its recommendations.
Malú Castellanos, Ivo Jimenez, Neal Coddington, Hansjörg Zeller, Steven Euijong Whang, Umeshwar Dayal
ICDE2