EDBT 2026 Demo / reviewers in the wild / expert
Miguel Branco
dblp:61/2926
· DBLP profile ↗
11ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7Systems, architecture and hardware · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Query processing and optimization · 35% Transaction processing and concurrency control · 32% Database system architecture and tuning · 19% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 55% Processor architecture and microarchitecture · 29% Parallel and multicore computing · 16% |
Topics — the 7 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization › query execution
raw data querying |
0.3 | 2 | 2014 | Adaptive Query Processing on RAW Data · Proc. VLDB Endow. 2014 NoDB in Action: Adaptive Query Processing on Raw Data · Proc. VLDB Endow. 2012 |
Transaction processing and concurrency control
OLTP |
0.3 | 2 | 2016 | Characterization of the Impact of Hardware Islands on OLTP · VLDB J. 2016 A data-oriented transaction execution engine and supporting tools · SIGMOD Conference 2011 |
Query processing and optimization
adaptive query processing |
0.1 | 1 | 2012 | NoDB in Action: Adaptive Query Processing on Raw Data · Proc. VLDB Endow. 2012 |
Memory systems
non-uniform memory access |
0.1 | 1 | 2012 | OLTP on Hardware Islands · Proc. VLDB Endow. 2012 |
Transaction processing and concurrency control
transaction execution |
0.1 | 1 | 2011 | A data-oriented transaction execution engine and supporting tools · SIGMOD Conference 2011 |
Processor architecture and microarchitecture
multicore design |
0.1 | 1 | 2016 | Characterization of the Impact of Hardware Islands on OLTP · VLDB J. 2016 |
Data integration and cleaning
heterogeneous data sources |
0.1 | 1 | 2014 | Adaptive Query Processing on RAW Data · Proc. VLDB Endow. 2014 |
Methods — techniques the papers use, named apart from their topics
performance analysis · 0.3just-in-time access paths · 0.2column shreds · 0.2incremental parsing · 0.1adaptive indexing · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Characterization of the Impact of Hardware Islands on OLTP
Danica Porobic, Ippokratis Pandis, Miguel Branco, Pinar Tözün, Anastasia Ailamaki |
VLDB J. | 3 |
| 2015 | Just-In-Time Data Virtualization: Lightweight Data Management with ViDa
Manos Karpathiotakis, Ioannis Alagiannis, Thomas Heinis, Miguel Branco, Anastasia Ailamaki |
CIDR | 4 |
| 2014 | Adaptive Query Processing on RAW DataabstractDatabase systems deliver impressive performance for large classes of workloads as the result of decades of research into optimizing database engines. High performance, however, is achieved at the cost of versatility. In particular, database systems only operate efficiently over loaded data, i.e., data converted from its original raw format into the system's internal data format. At the same time, data volume continues to increase exponentially and data varies increasingly, with an escalating number of new formats. The consequence is a growing impedance mismatch between the original structures holding the data in the raw files and the structures used by query engines for efficient processing. In an ideal scenario, the query engine would seamlessly adapt itself to the data and ensure efficient query processing regardless of the input data formats, optimizing itself to each instance of a file and of a query by leveraging information available at query time. Today's systems, however, force data to adapt to the query engine during data loading. This paper proposes adapting the query engine to the formats of raw data. It presents RAW, a prototype query engine which enables querying heterogeneous data sources transparently. RAW employs Just-In-Time access paths, which efficiently couple heterogeneous raw files to the query engine and reduce the overheads of traditional general-purpose scan operators. There are, however, inherent overheads with accessing raw data directly that cannot be eliminated, such as converting the raw values. Therefore, RAW also uses column shreds, ensuring that we pay these costs only for the subsets of raw data strictly needed by a query. We use RAW in a real-world scenario and achieve a two-order of magnitude speedup against the existing hand-written solution. Manos Karpathiotakis, Miguel Branco, Ioannis Alagiannis, Anastasia Ailamaki |
Proc. VLDB Endow. | 2 |
| 2012 | NoDB: efficient query execution on raw data filesabstractAs data collections become larger and larger, data loading evolves to a major bottleneck. Many applications already avoid using database systems, e.g., scientific data analysis and social networks, due to the complexity and the increased data-to-query time. For such applications data collections keep growing fast, even on a daily basis, and we are already in the era of data deluge where we have much more data than what we can move, store, let alone analyze. Ioannis Alagiannis, Renata Borovica, Miguel Branco, Stratos Idreos, Anastasia Ailamaki |
SIGMOD Conference | 3 |
| 2012 | NoDB in Action: Adaptive Query Processing on Raw DataabstractAs data collections become larger and larger, users are faced with increasing bottlenecks in their data analysis. More data means more time to prepare the data, to load the data into the database and to execute the desired queries. Many applications already avoid using traditional database systems, e.g., scientific data analysis and social networks, due to their complexity and the increaseddata-to-querytime, i.e. the time between getting the data and retrieving its first useful results. For many applications data collections keep growing fast, even on a daily basis, and thisdata delugewill only increase in the future, where it is expected to have much more data than what we can move or store, let alone analyze. In this demonstration, we will showcase a new philosophy for designing database systems called NoDB. NoDB aims at minimizing the data-to-query time, most prominently by removing the need to load data before launching queries. We will present our prototype implementation, PostgresRaw, built on top of PostgreSQL, which allows for efficient query execution over raw data files with zero initialization overhead. We will visually demonstrate how PostgresRaw incrementally and adaptively touches, parses, caches and indexes raw data files autonomously and exclusively as a side-effect of user queries. Ioannis Alagiannis, Renata Borovica, Miguel Branco, Stratos Idreos, Anastasia Ailamaki |
Proc. VLDB Endow. | 3 |
| 2012 | OLTP on Hardware IslandsabstractModern hardware is abundantly parallel and increasingly heterogeneous. The numerous processing cores have nonuniform access latencies to the main memory and to the processor caches, which causes variability in the communication costs. Unfortunately, database systems mostly assume that all processing cores are the same and that microarchitecture differences are not significant enough to appear in critical database execution paths. As we demonstrate in this paper, however, hardware heterogeneity does appear in the critical path and conventional database architectures achieve suboptimal and even worse, unpredictable performance. We perform a detailed performance analysis of OLTP deployments in servers with multiple cores per CPU ( multicore ) and multiple CPUs per server ( multisocket ). We compare different database deployment strategies where we vary the number and size of independent database instances running on a single server, from a single shared-everything instance to fine-grained shared-nothing configurations. We quantify the impact of non-uniform hardware on various deployments by (a) examining how efficiently each deployment uses the available hardware resources and (b) measuring the impact of distributed transactions and skewed requests on different workloads. Finally, we argue in favor of shared-nothing deployments that are topology- and workload-aware and take advantage of fast on-chip communication between islands of cores on the same socket. Danica Porobic, Ippokratis Pandis, Miguel Branco, Pinar Tözün, Anastasia Ailamaki |
Proc. VLDB Endow. | 3 |
| 2011 | A data-oriented transaction execution engine and supporting toolsabstractConventional OLTP systems assign each transaction to a worker thread and that thread accesses data, depending on what the transaction dictates. This thread-to-transaction work assignment policy leads to unpredictable accesses. The unpredictability forces each thread to enter a large number of critical sections for the completion of even the simplest of the transactions; leading to poor performance and scalability on modern manycore hardware. Ippokratis Pandis, Pinar Tözün, Miguel Branco, Dimitris Karampinas, Danica Porobic, Ryan Johnson 0001, Anastasia Ailamaki |
SIGMOD Conference | 3 |
| 2010 | Identification, Modelling and Prediction of Non-periodic Bursts in WorkloadsabstractNon-periodic bursts are prevalent in workloads of large scale applications. Existing workload models do not predict such non-periodic bursts very well because they mainly focus on repeatable base functions. We begin by showing the necessity to include bursts in workload models by investigating their detrimental effects in a petabyte-scale distributed data management system. This work then makes three contributions. First, we analyse the accuracy of five existing prediction models on workloads of data and computational grids, as well as derived synthetic workloads. Second, we introduce a novel averages-based model to predict bursts in arbitrary workloads. Third, we present a novel metric, mean absolute estimated distance, to assess the prediction accuracy of the model. Using our model and metric, we show that burst behaviour in workloads can be identified, quantified and predicted independently of the underlying base functions. Furthermore, our model and metric are applicable to arbitrary kinds of burst prediction for time series. Mario Lassnig, Thomas Fahringer, Vincent Garonne, Angelos Molfetas, Miguel Branco |
CCGRID | 5 |
| 2010 | Managing very large distributed data sets on a data gridabstractAbstract In this work we address the management of very large data sets, which need to be stored and processed across many computing sites. The motivation for our work is the ATLAS experiment for the Large Hadron Collider (LHC), where the authors have been involved in the development of the data management middleware. This middleware, called DQ2, has been used for the last several years by the ATLAS experiment for shipping petabytes of data to research centres and universities worldwide. We describe our experience in developing and deploying DQ2 on the Worldwide LHC computing Grid, a production Grid infrastructure formed of hundreds of computing sites. From this operational experience, we have identified an important degree of uncertainty that underlies the behaviour of large Grid infrastructures. This uncertainty is subjected to a detailed analysis, leading us to present novel modelling and simulation techniques for Data Grids. In addition, we discuss what we perceive as practical limits to the development of data distribution algorithms for Data Grids given the underlying infrastructure uncertainty, and propose future research directions. Copyright © 2009 John Wiley & Sons, Ltd. Miguel Branco, Ed Zaluska, David De Roure, Mario Lassnig, Vincent Garonne |
Concurr. Comput. Pract. Exp. | 1 |
| 2009 | Stream Monitoring in Large-Scale Distributed Concealed EnvironmentsabstractWe present a probabilistic tracing method that captures both user and system behaviour for large-scale distributed applications. Our method extends the notion of data stream monitoring to work within what we define as concealed environments. We detail the conceptual design and implementation of our method. Additionally, we evaluate the scalability of the tracing method in a real petabyte-scale distributed data management system. Finally, we demonstrate the usefulness of the collected trace data in three scenarios. First, we use collected trace data to examine the arrival of user events and find self-similar processes. Second, we examine the behaviour and performance of mass storage systems in a grid under concurrent requests. Third, we develop a model for prediction of user event arrivals based on historical data. Our results suggest that a probabilistic tracing method is scalable, straightforward to integrate with existing applications, and provides useful insight into the behaviour of very large-scale applications. Mario Lassnig, Thomas Fahringer, Vincent Garonne, Angelos Molfetas, Miguel Branco |
eScience | 5 |
| 2007 | The Requirements of Using Provenance in e-Science Experiments
Simon Miles, Paul Groth, Miguel Branco, Luc Moreau 0001 |
J. Grid Comput. | 3 |