EDBT 2026 Demo / reviewers in the wild / expert
Romulo Goncalves
dblp:32/935
· DBLP profile ↗
19ranked-venue papers
5as first author
0since 2021 · last 2020
0000-0003-2225-1428ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 13 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSystems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Query processing and optimization · 41% Indexing and storage engines · 19% Distributed and cloud data management · 18% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 50% Distributed systems · 50% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Indexing and storage engines
column store |
0.4 | 3 | 2015 | GIS Navigation Boosted by Column Stores · Proc. VLDB Endow. 2015 An architecture for recycling intermediates in a column-store · SIGMOD Conference 2009 Column-store support for RDF data management: not all swans are white · Proc. VLDB Endow. 2008 |
Query processing and optimization
join processing |
0.2 | 1 | 2016 | Building a Hybrid Warehouse: Efficient Joins between Data Stored in HDFS and Enterprise Warehouse · ACM Trans. Database Syst. 2016 |
Spatial and temporal data management
spatial query processing |
0.2 | 1 | 2015 | GIS Navigation Boosted by Column Stores · Proc. VLDB Endow. 2015 |
Query processing and optimization › result reuse
intermediate result reuse |
0.2 | 2 | 2010 | An architecture for recycling intermediates in a column-store · ACM Trans. Database Syst. 2010 An architecture for recycling intermediates in a column-store · SIGMOD Conference 2009 |
Distributed and cloud data management
distributed query processing |
0.1 | 1 | 2011 | The data cyclotron query processing scheme · ACM Trans. Database Syst. 2011 |
Query processing and optimization
query processing architecture |
0.1 | 1 | 2011 | The data cyclotron query processing scheme · ACM Trans. Database Syst. 2011 |
Distributed systems
distributed coordination |
0.1 | 1 | 2011 | The data cyclotron query processing scheme · ACM Trans. Database Syst. 2011 |
Cloud and datacenter computing
resource management |
0.1 | 1 | 2011 | The data cyclotron query processing scheme · ACM Trans. Database Syst. 2011 |
Graph data management
RDF data management |
0.1 | 1 | 2008 | Column-store support for RDF data management: not all swans are white · Proc. VLDB Endow. 2008 |
Database system architecture and tuning › database design › physical database design
vertical partitioning |
0.1 | 1 | 2008 | Column-store support for RDF data management: not all swans are white · Proc. VLDB Endow. 2008 |
Query processing and optimization
cost model |
0.1 | 1 | 2016 | Building a Hybrid Warehouse: Efficient Joins between Data Stored in HDFS and Enterprise Warehouse · ACM Trans. Database Syst. 2016 |
Query processing and optimization
query optimization |
0.1 | 1 | 2016 | Building a Hybrid Warehouse: Efficient Joins between Data Stored in HDFS and Enterprise Warehouse · ACM Trans. Database Syst. 2016 |
Spatial and temporal data management
spatial databases |
0.1 | 1 | 2015 | GIS Navigation Boosted by Column Stores · Proc. VLDB Endow. 2015 |
Information retrieval › evaluation
benchmark evaluation |
0.0 | 1 | 2008 | Column-store support for RDF data management: not all swans are white · Proc. VLDB Endow. 2008 |
Methods — techniques the papers use, named apart from their topics
zigzag join algorithm · 0.2simulation · 0.2remote DMA · 0.2bloom filter · 0.2imprints index · 0.2column storage · 0.2result caching · 0.1materialized views · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Exploring Spring Onset at Continental Scales: Mapping Phenoregions and Correlating Temperature and Satellite-Based PhenometricsabstractEach spring many plants put on new leaves and/or open their flowers creating a “green-wave” that can be tracked using phenological data. Various phenological datasets can be used to study spring onset at continental to global scales. Here we present a novel exploratory analysis where we link two multi-decadal and high-spatial resolution datasets: temperature-based phenological indices and land surface phenological metrics derived from satellite images. Our exploratory analysis, illustrated with data for the conterminous US, focuses on identifying regions with similar spring onset, and on mapping the coherence between these phenological products. Our results show that the spring onset patterns captured by the satellite are more complex than the ones identified using temperature-based phenological indices. They also highlight areas with stable and unstable spring onsets (i.e., areas that tend to remain or change of phenoregion from year to year). Finally, our results reveal that temperature-based indices are both positively and negatively correlated with the phenological information that can be derived from satellites. This opens the door to the definition of rules to integrate multi-source phenological data. To cope with the computational challenges of analyzing big geospatial rasters, we executed our analysis on a cloud platform running Apache Spark and various of its extensions (e.g., Geotrellis, SparkMLlib). This platform performed well and allowed the execution of user-tailored analyses. Hence, we believe that our computational platform paves the path towards the efficient analysis of global vegetation phenology at very high spatial resolution and, more generally, to the analysis of the ever-increasing collections of geospatial data about our planet. Raúl Zurita-Milla, Romulo Goncalves, Emma Izquierdo-Verdiguier, Frank O. Ostermann |
IEEE Trans. Big Data | 2 |
| 2018 | A Spark-Based Platform to Extract Phenological Information from Satellite ImagesabstractPhenology is the study of periodic plant and animal life cycle events and how these are influenced by seasonal and inter-annual variations in weather and climate, as well as in other environmental factors. Time series of remote sensing (RS) images can be used to characterize land surface phenology at continental to global scales. For this, the RS images are typically transformed into various vegetation indices (VI) such as the normalized difference vegetation index (NDVI) or the enhanced vegetation index (EVI). These indices can then be used to extract various phenological metrics. In our previous work we used cloud computing to generate temperature-based phenological indices [1], [2], and to relate one phenological metric, namely the Start-of-Season (SOS), with those indices [3], [4]. Here we present an extension of our work where we use a Spark-based platform to efficiently extract phenological metrics from time series of NDVI and EVI. This platform allows obtaining and analyzing high spatial resolution metrics (in this case 1km) from 10-day composites. The platform uses the same architecture as in [3], i.e., it is organized into three layers: a storage layer, a processing layer, and JupyterHub services for user-interaction. It is designed to store the data in well-known file formats like GeoTiffs and Hierarchical Data Format (HDF). For the data analysis the user expresses the operations in Jupyter notebooks as Python, R, or Scala code (Fig. 1). Hence, with a browser and remote connection, the user can express a research question and/or collect insights from large data sets. All computations are pushed down to the computational platform, and results fetched back for data visualization. To extract the phenological metrics, we rely on TimeSat [5]. TimeSat is a software package that can be used to fit a function (e.g. double logistic) to time series of VIs. After that, it uses various approaches to extract vegetation seasonality metrics such as SOS. The programs numerical and graphical routines are coded in Matlab and Fortran. These routines are highly vectorized and efficient for use with large data sets. However, distributed processing is required to determine SOS at continental scales. Through an efficient partition of the data, and Spark’s scheduling policies, these single-core routines are scheduled for parallel execution over multiple machines. The study evaluates which VIs and fitting functions are most suitable for certain vegetation types by comparing the SOS metrics to volunteered phenological observations curated by the USA national phenological network [6]. Our preliminary results show there can be up to 20-30 days differences in the SOS depending on the fitting function, the VI and the approach used to extract the SOS metric. In the South, SOS is around mid-February or March whereas in mountainous regions and the North, the SOS can be as late as June-July. We are to further evaluate how our results compare to the ground volunteered observations. This work is then a first stepping stone towards being able to systematically analyze and map the impact of climate change on the seasonality of plants. Our tests show that the platform is scalable and can be extended to work with even higher resolution VIs, such as those that can be derived from Sentinel-2 images (10 m resolution). Because of this, our work opens the door to studies at continental to global scales, and to the use of high and very high spatial resolution data. Viktor Bakayov, Romulo Goncalves, Raúl Zurita-Milla, Emma Izquierdo-Verdiguier |
eScience | 2 |
| 2016 | A columnar architecture for modern risk management systemsabstract3D digital city models form the basis for flow simulations (e.g. wind flow and water runoff), urban planning, under-and over-ground formation analysis, and they are very important for automated anomaly detection on man made structures. They consist of large collections of semantically rich objects which have many properties such as material and color. Such user's data structure perception is leading to complex storage schemas. The number of table relations to manage and the large data storage footprint drawbacks are then extended with the fact that not all the systems have a “real” 3D data type. In this work we would like to show our efforts to develop a new kind of Spatial Data Management System (SDBMS) where topological and geometric functionality for 3D raster manipulation will become part of the relational kernel and not an add-on. With it spatial analysis tailored to different use case scenarios is done on-demand and fast enough to support real-time interaction in modern risk management systems. Romulo Goncalves, Sisi Zlatanova, Kostis Kyzirakos, Pirouz Nourian, Foteini Alvanaki, Willem Robert van Hage |
eScience | 1 |
| 2016 | A spatial column-store to triangulate the Netherlands on the flyabstract3D digital city models, important for urban planning, are currently constructed from massive point clouds obtained through airborne LiDAR (Light Detection and Ranging). They are semantically enriched with information obtained from auxiliary GIS data like Cadastral data which contains information about the boundaries of properties, road networks, rivers, lakes etc. Romulo Goncalves, Tom van Tilburg, Kostis Kyzirakos, Foteini Alvanaki, Panagiotis Koutsourakis, Ben van Werkhoven, Willem Robert van Hage |
SIGSPATIAL/GIS | 1 |
| 2016 | Building a Hybrid Warehouse: Efficient Joins between Data Stored in HDFS and Enterprise WarehouseabstractThe Hadoop Distributed File System (HDFS) has become an important data repository in the enterprise as the center for all business analytics, from SQL queries and machine learning to reporting. At the same time, enterprise data warehouses (EDWs) continue to support critical business analytics. This has created the need for a new generation of a special federation between Hadoop-like big data platforms and EDWs, which we call the hybrid warehouse . There are many applications that require correlating data stored in HDFS with EDW data, such as the analysis that associates click logs stored in HDFS with the sales data stored in the database. All existing solutions reach out to HDFS and read the data into the EDW to perform the joins, assuming that the Hadoop side does not have efficient SQL support. In this article, we show that it is actually better to do most data processing on the HDFS side, provided that we can leverage a sophisticated execution engine for joins on the Hadoop side. We identify the best hybrid warehouse architecture by studying various algorithms to join database and HDFS tables. We utilize Bloom filters to minimize the data movement and exploit the massive parallelism in both systems to the fullest extent possible. We describe a new zigzag join algorithm and show that it is a robust join algorithm for hybrid warehouses that performs well in almost all cases. We further develop a sophisticated cost model for the various join algorithms and show that it can facilitate query optimization in the hybrid warehouse to correctly choose the right algorithm under different predicate and join selectivities. Yuanyuan Tian 0001, Fatma Özcan 0001, Romulo Goncalves, Hamid Pirahesh |
ACM Trans. Database Syst. | 4 |
| 2015 | A Round Table for Multi-disciplinary Research on Geospatial and Climate DataabstractEarth observation sciences produce large sets of data which are inherently rich in spatial and geo-spatial information. Together with live data collected from monitoring systems and large collections of semantically rich objects they provide new opportunities for advanced eScience research on climatology, urban planing and smart cities. Such combination of heterogeneous data sets forms a new source of knowledge. Efficient knowledge extraction from them is an eScience challenge. It requires efficient bulk data injection from both static and streaming data sources, dynamic adaptation of the physical and logical schema, efficient methods to correlate spatial and temporal data, and flexibility to (re-)formulate the research question at any time. In this work, we present a data management layer over a column-oriented relational data management system that provides efficient analysis of spatiotemporal data. It provides fast data ingestion through different data loaders, tabular and array based storage, and a dynamic step-wise exploration. Romulo Goncalves, Milena Ivanova, Foteini Alvanaki, Jason Maassen, Kostis Kyzirakos, Oscar Martinez-Rubi, Hannes Mühleisen |
e-Science | 1 |
| 2015 | Joins for Hybrid Warehouses: Exploiting Massive Parallelism in Hadoop and Enterprise Data WarehousesabstractHDFS has become an important data repository in the enterprise as the center for all business analytics, from SQL queries, machine learning to reporting. At the same time, enterprise data warehouses (EDWs) continue to support critical business analytics. This has created the need for a new generation of special federation between Hadoop-like big data platforms and EDWs, which we call the hybrid warehouse. There are many applications that require correlating data stored in HDFS with EDW data, such as the analysis that associates click logs stored in HDFS with the sales data stored in the database. All existing solutions reach out to HDFS and read the data into the EDW to perform the joins, assuming that the Hadoop side does not have the efficient SQL support. In this paper, we show that it is actually better to do most data processing on the HDFS side, provided that we can leverage a sophisticated execution engine for joins on the Hadoop side. We identify the best hybrid warehouse architecture by studying various algorithms to join database and HDFS tables. We utilize Bloom filters to minimize the data movement, and exploit the massive parallelism in both systems to the fullest extent possible. We describe a new zigzag join algorithm, and show that it is a robust join algorithm for hybrid warehouses which performs well in almost all cases. Yuanyuan Tian 0001, Fatma Özcan 0001, Romulo Goncalves, Hamid Pirahesh |
EDBT | 4 |
| 2015 | Massive point cloud data management: Design, implementation and execution of a point cloud benchmark
Peter van Oosterom, Oscar Martinez-Rubi, Milena Ivanova, Mike Hörhammer, Daniel Geringer, Siva Ravada, Theo Tijssen, Martin Kodde, Romulo Goncalves |
Comput. Graph. | 9 |
| 2015 | GIS Navigation Boosted by Column StoresabstractEarth observation sciences, astronomy, and seismology have large data sets which have inherently rich spatial and geospatial information. In combination with large collections of semantically rich objects which have a large number of thematic properties, they form a new source of knowledge for urban planning, smart cities and natural resource management. Modeling and storing these properties indicating the relationships between them is best handled in a relational database. Furthermore, the scalability requirements posed by the latest 26-attribute light detection and ranging (LIDAR) data sets are a challenge for file-based solutions. In this demo we show how to query a 640 billion point data set using a column store enriched with GIS functionality. Through a lightweight and cache conscious secondary index called Imprints, spatial queries performance on a flat table storage is comparable to traditional file-based solutions. All the results are visualised in real time using QGIS. Foteini Alvanaki, Romulo Goncalves, Milena Ivanova, Martin L. Kersten, Kostis Kyzirakos |
Proc. VLDB Endow. | 2 |
| 2013 | Peak performance: remote memory revisitedabstractMany database systems share a need for large amounts of fast storage. However, economies of scale limit the utility of extending a single machine with an arbitrary amount of memory. The recent broad availability of the zero-copy data transfer protocol RDMA over low-latency and high-throughput network connections such as InfiniBand prompts us to revisit the long-proposed usage of memory provided by remote machines. In this paper, we present a solution to make use of remote memory without manipulation of the operating system, and investigate the impact on database performance. Hannes Mühleisen, Romulo Goncalves, Martin L. Kersten |
DaMoN | 2 |
| 2011 | The data cyclotron query processing schemeabstractA grand challenge of distributed query processing is to devise a self-organizing architecture which exploits all hardware resources optimally to manage the database hot set, minimize query response time, and maximize throughput without single point global coordination. The Data Cyclotron architecture [Goncalves and Kersten 2010] addresses this challenge using turbulent data movement through a storage ring built from distributed main memory and capitalizing on the functionality offered by modern remote-DMA network facilities. Queries assigned to individual nodes interact with the storage ring by picking up data fragments, which are continuously flowing around, that is, the hot set. The storage ring is steered by the Level Of Interest ( LOI ) attached to each data fragment, which represents the cumulative query interest as it passes around the ring multiple times. A fragment with LOI below a given threshold, inversely proportional to the ring load, is pulled out to free up resources. This threshold is dynamically adjusted in a fully distributed manner based on ring characteristics and locally observed query behavior. It optimizes resource utilization by keeping the average data access latency low. The approach is illustrated using an extensive and validated simulation study. The results underpin the fragment hot set management robustness in turbulent workload scenarios. A fully functional prototype of the proposed architecture has been implemented using modest extensions to MonetDB and runs within a multirack cluster equipped with Infiniband. Extensive experimentation using both microbenchmarks and high-volume workloads based on TPC-H demonstrates its feasibility. The Data Cyclotron architecture and experiments open a new vista for modern distributed database architectures with a plethora of new research challenges. Romulo Goncalves, Martin L. Kersten |
ACM Trans. Database Syst. | 1 |
| 2010 | The Data Cyclotron query processing schemeabstractDistributed database systems exploit static workload characteristics to steer data fragmentation and data allocation schemes. However, the grand challenge of distributed query processing is to come up with a self-organizing architecture, which exploits all resources to manage the hot data set, minimize query response time, and maximize throughput without global co-ordination. Romulo Goncalves, Martin L. Kersten |
EDBT | 1 |
| 2010 | A Spinning Join That Does Not Get DizzyabstractAs network infrastructures with 10 Gb/s bandwidth and beyond have become pervasive and as cost advantages of large commodity-machine clusters continue to increase, research and industry strive to exploit the available processing performance for large-scale database processing tasks. In this work we look at the use of high-speed networks for distributed join processing. We propose Data Roundabout as alight weight transport layer that uses Remote Direct Memory Access (RDMA) to gain access to the throughput opportunities in modern networks. The essence of Data Roundabout is a ring shaped network in which each host stores one portion of a large database instance. We leverage the available bandwidth to (continuously) pump data through the high-speed network. Based on Data Roundabout, we demonstrate cyclo-join, which exploits the cycling flow of data to execute distributed joins. The study uses different join algorithms (hash join and sort-merge join) to expose the pitfalls and the advantages of each algorithm in the data cycling arena. The experiments show the potential of a large distributed main-memory cache glued together with RDMA into a novel distributed database architecture. Philip Werner Frey, Romulo Goncalves, Martin L. Kersten, Jens Teubner |
ICDCS | 2 |
| 2010 | An architecture for recycling intermediates in a column-storeabstractAutomatic recycling of intermediate results to improve both query response time and throughput is a grand challenge for state-of-the-art databases. Tuples are loaded and streamed through a tuple-at-a-time processing pipeline, avoiding materialization of intermediates as much as possible. This limits the opportunities for reuse of overlapping computations to DBA-defined materialized views and function/result cache tuning. In contrast, the operator-at-a-time execution paradigm produces fully materialized results in each step of the query plan. To avoid resource contention, these intermediates are evicted as soon as possible. In this article we study an architecture that harvests the byproducts of the operator-at-a-time paradigm in a column-store system using a lightweight mechanism, the recycler. The key challenge then becomes the selection of the policies to admit intermediates to the resource pool, to determine their retention period, and devise the eviction strategy when facing resource limitations. The proposed recycling architecture has been implemented in an open-source system. An experimental analysis against the TPC-H ad-hoc decision support benchmark and a complex, real-world application (SkyServer) demonstrates its effectiveness in terms of self-organizing behavior and its significant performance gains. The results indicate the potentials of recycling intermediates and charts a route for further development of database kernels. Milena Ivanova, Martin L. Kersten, Niels Nes, Romulo Goncalves |
ACM Trans. Database Syst. | 4 |
| 2009 | Spinning relations: high-speed networks for distributed join processingabstractBy leveraging modern networking hardware (RDMA-enabled network cards), we can shift priorities in distributed database processing significantly. Complex and sophisticated mechanisms to avoid network traffic can be replaced by a scheme that takes advantage of the bandwidth and low latency offered by such interconnects. Philip Werner Frey, Romulo Goncalves, Martin L. Kersten, Jens Teubner |
DaMoN | 2 |
| 2009 | Exploiting the power of relational databases for efficient stream processingabstractStream applications gained significant popularity over the last years that lead to the development of specialized stream engines. These systems are designed from scratch with a different philosophy than nowadays database engines in order to cope with the stream applications requirements. However, this means that they lack the power and sophisticated techniques of a full fledged database system that exploits techniques and algorithms accumulated over many years of database research. Erietta Liarou, Romulo Goncalves, Stratos Idreos |
EDBT | 2 |
| 2009 | An architecture for recycling intermediates in a column-storeabstractAutomatically recycling (intermediate) results is a grand challenge for state-of-the-art databases to improve both query response time and throughput. Tuples are loaded and streamed through a tuple-at-a-time processing pipeline avoiding materialization of intermediates as much as possible. This limits the opportunities for reuse of overlapping computations to DBA-defined materialized views and function/result cache tuning. Milena Ivanova, Martin L. Kersten, Niels Nes, Romulo Goncalves |
SIGMOD Conference | 4 |
| 2008 | Column-store support for RDF data management: not all swans are whiteabstractThis paper reports on the results of an independent evaluation of the techniques presented in the VLDB 2007 paper "Scalable Semantic Web Data Management Using Vertical Partitioning", authored by D. Abadi, A. Marcus, S. R. Madden, and K. Hollenbach [1]. We revisit the proposed benchmark and examine both the data and query space coverage. The benchmark is extended to cover a larger portion of the query space in a canonical way. Repeatability of the experiments is assessed using the code base obtained from the authors. Inspired by the proposed vertically-partitioned storage solution for RDF data and the performance figures using a column-store, we conduct a complementary analysis of state-of-the-art RDF storage solutions. To this end, we employ MonetDB/SQL, a fully-functional open source column-store, and a well-known -- for its performance -- commercial row-store DBMS. We implement two relational RDF storage solutions -- triple-store and vertically-partitioned -- in both systems. This allows us to expand the scope of [1] with the performance characterization along both dimensions -- triple-store vs. vertically-partitioned and row-store vs. column-store -- individually, before analyzing their combined effects. A detailed report of the experimental test-bed, as well as an in-depth analysis of the parameters involved, clarify the scope of the solution originally presented and position the results in a broader context by covering more systems. Lefteris Sidirourgos, Romulo Goncalves, Martin L. Kersten, Niels Nes, Stefan Manegold |
Proc. VLDB Endow. | 2 |
| 2007 | MonetDB/SQL Meets SkyServer: the Challenges of a Scientific DatabaseabstractThis paper presents our experiences in porting the Sloan Digital Sky Survey(SDSS)/ SkyServer to the state-of- the-art open source database system MonetDB/SQL. SDSS acts as a well-documented benchmark for scientific database management. We have achieved a fully functional prototype for the personal SkyServer, to be downloaded from our site. The lessons learned are 1) the column store approach of MonetDB demonstrates a great potential in the world of scientific databases. However, the application also challenged the functionality of our implementation and revealed that a fully operational SQL environment is needed, e.g. including persistent stored modules; 2) the initial performance is competitive to the reference platform, MS SQL Server 2005, and 3) the analysis of SDSS query traces hints at several techniques to boost performance by utilizing repetitive behavior and zoom-in/zoom-out access patterns, that are currently not captured by the system. Milena Ivanova, Niels Nes, Romulo Goncalves, Martin L. Kersten |
SSDBM | 3 |