VLDB 2026 Research / reviewers in the wild / expert
Pedro Holanda
dblp:119/6830 · also Pedro H. F. Holanda
· DBLP profile ↗
10ranked-venue papers
6as first author
2since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Query processing and optimization · 51% Indexing and storage engines · 49% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Indexing and storage engines
multidimensional indexing |
0.5 | 1 | 2021 | Multidimensional Adaptive & Progressive Indexes · ICDE 2021 |
Query processing and optimization
interactive query processing |
0.4 | 1 | 2019 | Progressive Indexes: Indexing for Interactive Data Analysis · Proc. VLDB Endow. 2019 |
Methods — techniques the papers use, named apart from their topics
pareto optimization · 0.5kd-tree · 0.5performance-driven indexing · 0.4budget tuning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Progressive Mergesort: Merging Batches of Appends into Progressive IndexesabstractInteractive exploratory data analysis consists of workloads that are composed of filter-aggregate queries with highly selective filters [1]. Hence, their performance is dependent on how much data they can skip during their scans, with indexes being the most efficient technique for aggressive data-skipping. Progressive Indexes are the state-of-the-art on automatic index creation for interactive exploratory data analysis. These indexes are partially constructed during query execution, eventually refining to a full index. However, progressive indexes have been designed for static databases, while in exploratory data analysis updates - usually batch-appends of newly acquired data - are frequent. In this paper, we propose Progressive Mergesort, a novel merging technique to make Progressive Indexes cope with updates. Progressive Mergesort differs from other merging techniques for partial indexes as it incorporates the index budget strategy design from Progressive Indexing. It follows the same three principles as Progressive Indexes: (1) fast query execution, (2) high robustness,(3) guaranteed convergence. Our experimental evaluation demonstrates that Progressive Mergesort is capable of achieving a 2x speedup when merging updates and up to 3 orders of magnitude lower variance than the state of the art. Pedro Holanda, Stefan Manegold |
EDBT | 1 |
| 2021 | Multidimensional Adaptive & Progressive IndexesabstractExploratory data analysis is the primary technique used by data scientists to extract knowledge from new data sets. This type of workload is composed of trial-and-error hypothesis-driven queries with a human in the loop. To keep up with the data scientist's productivity, the system must be capable of answering queries in interactive times. Given that these queries are highly selective multidimensional queries, multidimensional indexes are necessary to ensure low latency. However, creating the appropriate indexes is not a given due to the highly exploratory and interactive nature of such human-in-the-loop scenarios.In this paper, we identify four main objectives that are desirable for exploratory data analysis workloads: (1) low overhead over the initial queries, (2) low query variance (i.e., high robustness), (3) predictable index convergence, and (4) low total workload time. Given that not all of them can be achieved at the same time, we present three novel incremental multidimensional indexing techniques that represent three sample points on a Pareto front for this multi-objective optimization problem. (a) The Adaptive KD-Tree is designed to achieve the lowest total workload time at the expense of a higher indexing penalty for the initial queries, lack of robustness, and unpredictable convergence. (b) The Progressive KD-Tree has predictable convergence and a user-defined indexing cost for the initial queries. However, total workload time can be higher than with Adaptive KD-Trees, and per-query time still varies. (c) The Greedy Progressive KD-Tree aims at full robustness at the expense of only improving the per-query cost after full index convergence.Our extensive experimental evaluation using both synthetic and real-life data sets and workloads shows that (a) the Adaptive KD-Tree reduces total workload time by up to a factor 2 compared to the state-of-the-art, (b) the Progressive KD-Tree achieves predictable convergence with up to one order of magnitude lower initial query cost, and (c) the Greedy Progressive KDTree exhibits the lowest query variance up to three orders of magnitude lower than the state-of-the-art. Matheus Nerone, Pedro Holanda, Eduardo C. de Almeida, Stefan Manegold |
ICDE | 2 |
| 2019 | Relational Queries with a Tensor Processing UnitabstractTensor Processing Units are specialized hardware devices built to train and apply Machine Learning models at high speed through high-bandwidth memory and massive instruction parallelism. In this short paper, we investigate how relational operations can be translated to those devices. We present mapping of relational operators to TPU-supported TensorFlow operations and experimental results comparing with GPU and CPU implementations. Results show that while raw speeds are enticing, TPUs are unlikely to improve relational query processing for now due to a variety of issues. Pedro Holanda, Hannes Mühleisen |
DaMoN | 1 |
| 2019 | devUDF: Increasing UDF development efficiency through IDE Integration. It works like a PyCharm!abstractUser-defined functions (UDFs) facilitate the execution of analytics pipelines inside the database. They provide many advantages over traditional methods, such as close-to-data execution and automatic parallelization. However, the standard workflow for developing and debugging UDFs does not allow developers to use their regular toolchains and Integrated Development Environments (IDEs). As a result, writing functional UDFs is challenging. In this demo, we present the devUDF, a plugin to the PyCharm IDE that allows developers to develop and debug their MonetDB/Python UDFs directly from within the IDE. Mark Raasveldt, Pedro Holanda, Stefan Manegold |
EDBT | 2 |
| 2019 | Progressive Indexes: Indexing for Interactive Data AnalysisabstractInteractive exploration of large volumes of data is increasingly common, as data scientists attempt to extract interesting information from large opaque data sets. This scenario presents a difficult challenge for traditional database systems, as (1) nothing is known about the query workload in advance, (2) the query workload is constantly changing, and (3) the system must provide interactive responses to the issued queries. This environment is challenging for index creation, as traditional database indexes require upfront creation, hence a priori workload knowledge, to be efficient. In this paper, we introduce Progressive Indexing , a novel performance-driven indexing technique that focuses on automatic index creation while providing interactive response times to incoming queries. Its design allows queries to have a limited budget to spend on index creation. The indexing budget is automatically tuned to each query before query processing. This allows for systems to provide interactive answers to queries during index creation while being robust against various workload patterns and data distributions. Pedro Holanda, Stefan Manegold, Hannes Mühleisen, Mark Raasveldt |
Proc. VLDB Endow. | 1 |
| 2018 | Cracking KD-Tree: The First Multidimensional Adaptive Indexing (Position Paper)abstractWorkload-aware physical data access structures are crucial to achieve short response time with (exploratory) data analysis tasks as commonly required for Big Data and Data Science applications. Recently proposed techniques such as automatic index advisers (for a priori known static workloads) and query-driven adaptive incremental indexing (for a priori unknown dynamic workloads) form the state-of-the-art to build single-dimensional indexes for single-attribute query predicates. However, similar techniques for more demanding multi-attribute query predicates, which are vital for any data analysis task, have not been proposed, yet. In this paper, we present our on-going work on a new set of workload-adaptive indexing techniques that focus on creating multidimensional indexes. We present our proof-of-concept, the Cracking KD-Tree, an adaptive indexing approach that generates a KD-Tree based on multidimensional range query predicates. It works by incrementally creating partial multidimensional indexes as a by-product of query processing. The indexes are produced only on those parts of the data that are accessed, and their creation cost is effectively distributed across a stream of queries. Experimental results show that the Cracking KD-Tree is three times faster than creating a full KD-Tree, one order of magnitude faster than executing full scans and two orders of magnitude faster than using uni-dimensional full or adaptive indexes on multiple columns. Pedro Holanda, Matheus Nerone, Eduardo C. de Almeida, Stefan Manegold |
DATA | 1 |
| 2018 | Deep Integration of Machine Learning Into Column StoresabstractWe leverage vectorized User-Defined Functions (UDFs) to efficiently integrate unchanged machine learning pipelines into an analytical data management system.The entire pipelines including data, models, parameters and evaluation outcomes are stored and executed inside the database system.Experiments using our MonetDB/Python UDFs show greatly improved performance due to reduced data movement and parallel processing opportunities.In addition, this integration enables meta-analysis of models using relational queries. Mark Raasveldt, Pedro Holanda, Hannes Mühleisen, Stefan Manegold |
EDBT | 2 |
| 2017 | SPST-Index: A Self-Pruning Splay Tree Index for Caching Database Cracking
Pedro Holanda, Eduardo C. de Almeida |
EDBT | 1 |
| 2017 | Mixtape: Using Real-Time User Feedback to Navigate Large Media CollectionsabstractIn this work, we explore the increasing demand for novel user interfaces to navigate large media collections. We implement a geometric data structure to store and retrieve item-to-item similarity information and propose a novel navigation framework that uses vector operations and real-time user feedback to direct the outcome. The framework is scalable to large media collections and is suitable for computationally constrained devices. In particular, we implement this framework in the domain of music. To evaluate the effectiveness of the navigation process, we propose an automatic evaluation framework, based on synthetic user profiles, which allows us to quickly simulate and compare navigation paths using different algorithms and datasets. Moreover, we perform a real user study. To do that, we developed and launched Mixtape , a simple web application that allows users to create playlists by providing real-time feedback through liking and skipping patterns. Luciana Fujii Pontello, Pedro Holanda, Bruno Guilherme, João Paulo V. Cardoso, Olga Goussevskaia, Ana Paula Couto da Silva |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2015 | TV Goes Social: Characterizing User Interaction in an Online Social Network for TV Fans
Pedro Holanda, Bruno Guilherme, Ana Paula Couto da Silva, Olga Goussevskaia |
ICWE | 1 |