Fabian Nagel

dblp:131/4779 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Query processing and optimization · 40% Data integration and cleaning · 20% Transaction processing and concurrency control · 14%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning › data warehouse
cloud data warehouse
0.612022
Amazon Redshift Re-invented · SIGMOD Conference 2022
Query processing and optimization › query compilation
code generation for query execution
0.212014
Code Generation for Efficient Query Processing in Managed Runtimes · Proc. VLDB Endow. 2014
Indexing and storage engines
index maintenance
0.212014
Reducing Database Locking Contention Through Multi-version Concurrency · Proc. VLDB Endow. 2014
Query processing and optimization › query execution
in-memory query processing
0.212014
Code Generation for Efficient Query Processing in Managed Runtimes · Proc. VLDB Endow. 2014
Data models and query languages › database programming language
language-integrated query
0.212014
Code Generation for Efficient Query Processing in Managed Runtimes · Proc. VLDB Endow. 2014
Transaction processing and concurrency control › concurrency control
multiversion concurrency control
0.212014
Reducing Database Locking Contention Through Multi-version Concurrency · Proc. VLDB Endow. 2014
Indexing and storage engines
multiversion index
0.212014
Reducing Database Locking Contention Through Multi-version Concurrency · Proc. VLDB Endow. 2014
Query processing and optimization
query compilation
0.212014
Code Generation for Efficient Query Processing in Managed Runtimes · Proc. VLDB Endow. 2014
Transaction processing and concurrency control › isolation levels
snapshot isolation
0.212014
Reducing Database Locking Contention Through Multi-version Concurrency · Proc. VLDB Endow. 2014
Query processing and optimization › result reuse
intermediate result reuse
0.212013
Recycling in pipelined query evaluation · ICDE 2013
Query processing and optimization
result reuse
0.212013
Recycling in pipelined query evaluation · ICDE 2013
Runtime systems and virtual machines
managed runtime
0.112014
Code Generation for Efficient Query Processing in Managed Runtimes · Proc. VLDB Endow. 2014
Query processing and optimization
query processing architecture
0.012013
Recycling in pipelined query evaluation · ICDE 2013

Methods — techniques the papers use, named apart from their topics

query compilation · 0.4language-integrated query · 0.4pessimistic concurrency control · 0.2optimistic concurrency control · 0.2kv-indirection structure · 0.2subsumption · 0.2proactive query rewriting · 0.2intermediate result caching · 0.2
YearPublicationVenuePosition
2022 Amazon Redshift Re-invented
abstract
In 2013, AmazonWeb Services revolutionized the data warehousing industry by launching Amazon Redshift, the first fully-managed, petabyte-scale, enterprise-grade cloud data warehouse. Amazon Redshift made it simple and cost-effective to efficiently analyze large volumes of data using existing business intelligence tools. This cloud service was a significant leap from the traditional on-premise data warehousing solutions, which were expensive, not elastic, and required significant expertise to tune and operate. Customers embraced Amazon Redshift and it became the fastest growing service in AWS. Today, tens of thousands of customers use Redshift in AWS's global infrastructure to process exabytes of data daily.
Nikos Armenatzoglou, Sanuj Basu, Naga Bhanoori, Mengchu Cai, Naresh Chainani, Kiran Chinta, Venkatraman Govindaraju, Todd J. Green, Monish Gupta, Sebastian Hillig, Eric Hotinger, Yan Leshinksy, Jintian Liang, Michael McCreedy, Fabian Nagel, Ippokratis Pandis, Panos Parchas, Rahul Pathak, Orestis Polychroniou, Foyzur Rahman, Gokul Soundararajan, Sriram Subramanian, Douglas B. Terry
SIGMOD Conference15
2017 Self-managed collections: Off-heap memory management for scalable query-dominated collections
abstract
Explosive growth in DRAM capacities and the emergence of language-integrated query enable a new class of managed applications that perform complex query processing on huge volumes of data stored as collections of objects in the memory space of the application. While more flexible in terms of schema design and application development, this approach typically experiences sub-par query execution performance when compared to specialized systems like DBMS. To address this issue, we propose self-managed collections, which utilize off-heap memory management and dynamic query compilation to improve the performance of querying managed data through language-integrated query. We evaluate self-managed collections using both microbenchmarks and enumeration-heavy queries from the TPC-H business intelligence benchmark. Our results show that self-managed collections outperform ordinary managed collections in both query processing and memory management by up to an order of magnitude and even outperform an optimized in memory columnar database system for the vast majority of queries.
Fabian Nagel, Gavin M. Bierman, Aleksandar Dragojevic, Stratis Viglas
EDBT1
2014 Code Generation for Efficient Query Processing in Managed Runtimes
abstract
In this paper we examine opportunities arising from the convergence of two trends in data management: in-memory database systems (imdbs), which have received renewed attention following the availability of affordable, very large main memory systems; and language-integrated query, which transparently integrates database queries with programming languages (thus addressing the famous 'impedance mismatch' problem). Language-integrated query not only gives application developers a more convenient way to query external data sources like imdbs, but also to use the same querying language to query an application's in-memory collections. The latter offers further transparency to developers as the query language and all data is represented in the data model of the host programming language. However, compared to imdbs, this additional freedom comes at a higher cost for query evaluation. Our vision is to improve in-memory query processing of application objects by introducing database technologies to managed runtimes. We focus on querying and we leverage query compilation to improve query processing on application objects. We explore different query compilation strategies and study how they improve the performance of query processing over application data. We take C# as the host programming language as it supports language-integrated query through the linq framework. Our techniques deliver significant performance improvements over the default linq implementation. Our work makes important first steps towards a future where data processing applications will commonly run on machines that can store their entire datasets in-memory, and will be written in a single programming language employing language-integrated query and imdb-inspired runtimes to provide transparent and highly efficient querying.
Fabian Nagel, Gavin M. Bierman, Stratis Viglas
Proc. VLDB Endow.1
2014 Reducing Database Locking Contention Through Multi-version Concurrency
abstract
In multi-version databases, updates and deletions of records by transactions require appending a new record to tables rather than performing in-place updates. This mechanism incurs non-negligible performance overhead in the presence of multiple indexes on a table, where changes need to be propagated to all indexes. Additionally, an uncommitted record update will block other active transactions from using the index to fetch the most recently committed values for the updated record. In general, in order to support snapshot isolation and/or multi-version concurrency, either each active transaction is forced to search a database temporary area (e.g., roll-back segments) to fetch old values of desired records, or each transaction is forced to scan the entire table to find the older versions of the record in a multi-version database (in the absence of specialized temporal indexes). In this work, we describe a novel kV-Indirection structure to enable efficient (parallelizable) optimistic and pessimistic multi-version concurrency control by utilizing the old versions of records (at most two versions of each record) to provide direct access to the recent changes of records without the need of temporal indexes. As a result, our technique results in higher degree of concurrency by reducing the clashes between readers and writers of data and avoiding extended lock delays. We have a working prototype of our concurrency model and kV-Indirection structure in a commercial database and conducted an extensive evaluation to demonstrate the benefits of our multi-version concurrency control, and we obtained orders of magnitude speed up over the single-version concurrency control.
Mohammad Sadoghi, Mustafa Canim, Bishwaranjan Bhattacharjee, Fabian Nagel, Kenneth A. Ross
Proc. VLDB Endow.4
2013 Recycling in pipelined query evaluation
abstract
Database systems typically execute queries in isolation. Sharing recurring intermediate and final results between successive query invocations is ignored or only exploited by caching final query results. The DBA is kept in the loop to make explicit sharing decisions by identifying and/or defining materialized views. Thus decisions are made only after a long time and sharing opportunities may be missed. Recycling intermediate results has been proposed as a method to make database query engines profit from opportunities to reuse fine-grained partial query results, that is fully autonomous and is able to continuously adapt to changes in the workload. The technique was recently revisited in the context of MonetDB, a system that by default materializes all intermediate results. Materializing intermediate results can consume significant system resources, therefore most other database systems avoid this where possible, following a pipelined query architecture instead. The novelty of this paper is to show how recycling can successfully be applied in pipelined query executors, by tracking the benefit of materializing possible intermediate results and then choosing the ones making best use of a limited intermediate result cache. We present ways to maximize the potential of recycling by leveraging subsumption and proactive query rewriting. We have implemented our approach in the Vectorwise database engine and have experimentally evaluated its potential using both synthetic and real-world datasets. Our results show that intermediate result recycling significantly improves performance.
Fabian Nagel, Peter Boncz, Stratis Viglas
ICDE1