Franz Färber

dblp:36/7116 · also Franz Faerber · DBLP profile ↗
← Back
28ranked-venue papers
2as first author
0since 2021 · last 2020
0009-0006-5771-3951ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 28 · 2 first-authorArtificial intelligence and machine learning · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
17 papers
Indexing and storage engines · 37% Database system architecture and tuning · 20% Transaction processing and concurrency control · 12%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 24 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Indexing and storage engines
hierarchical index
0.732017
Order Indexes: supporting highly dynamic hierarchical data in relational main-memory database systems · VLDB J. 2017
Indexing Highly Dynamic Hierarchical Data · Proc. VLDB Endow. 2015
DeltaNI: an efficient labeling scheme for versioned hierarchical data · SIGMOD Conference 2013
Database system architecture and tuning
main-memory database
0.432017
Order Indexes: supporting highly dynamic hierarchical data in relational main-memory database systems · VLDB J. 2017
DeltaNI: an efficient labeling scheme for versioned hierarchical data · SIGMOD Conference 2013
SAP HANA distributed in-memory database system: Transaction, session, and metadata management · ICDE 2013
Indexing and storage engines
temporal indexing
0.322013
Comprehensive and Interactive Temporal Query Processing with SAP HANA · Proc. VLDB Endow. 2013
Timeline index: a unified data structure for processing queries on temporal data in SAP HANA · SIGMOD Conference 2013
Spatial and temporal data management
temporal query processing
0.322013
Comprehensive and Interactive Temporal Query Processing with SAP HANA · Proc. VLDB Endow. 2013
Timeline index: a unified data structure for processing queries on temporal data in SAP HANA · SIGMOD Conference 2013
Indexing and storage engines › access methods
ordered index
0.312017
Order Indexes: supporting highly dynamic hierarchical data in relational main-memory database systems · VLDB J. 2017
Query processing and optimization
hierarchical query processing
0.322015
Supporting hierarchical data in SAP HANA · ICDE 2015
Indexing Highly Dynamic Hierarchical Data · Proc. VLDB Endow. 2015
Query processing and optimization
aggregation
0.212015
Cache-Efficient Aggregation: Hashing Is Sorting · SIGMOD Conference 2015
Indexing and storage engines › index maintenance
dynamic indexing
0.212015
Indexing Highly Dynamic Hierarchical Data · Proc. VLDB Endow. 2015
Data models and query languages › data modeling
hierarchical data model
0.212015
Supporting hierarchical data in SAP HANA · ICDE 2015
Indexing and storage engines
labeling scheme
0.212015
Indexing Highly Dynamic Hierarchical Data · Proc. VLDB Endow. 2015
Transaction processing and concurrency control
distributed transaction processing
0.212014
Distributed snapshot isolation: global transactions pay globally, local transactions pay locally · VLDB J. 2014
Transaction processing and concurrency control › isolation levels
snapshot isolation
0.212014
Distributed snapshot isolation: global transactions pay globally, local transactions pay locally · VLDB J. 2014
Distributed and cloud data management
distributed query processing
0.212013
SAP HANA distributed in-memory database system: Transaction, session, and metadata management · ICDE 2013
Indexing and storage engines
column store
0.112012
Efficient transaction processing in SAP HANA database: the end of a column store myth · SIGMOD Conference 2012
Database system architecture and tuning › database design
physical database design
0.112012
A Storage Advisor for Hybrid-Store Databases · Proc. VLDB Endow. 2012
Database system architecture and tuning › main-memory database
in-memory database engine
0.112011
Bridging Two Worlds with RICE Integrating R into the SAP In-Memory Computing Engine · Proc. VLDB Endow. 2011
Indexing and storage engines › column store
main-memory column store
0.122013
Comprehensive and Interactive Temporal Query Processing with SAP HANA · Proc. VLDB Endow. 2013
Timeline index: a unified data structure for processing queries on temporal data in SAP HANA · SIGMOD Conference 2013
Indexing and storage engines › data compression
dictionary compression
0.112009
Dictionary-based order-preserving string compression for main memory column stores · SIGMOD Conference 2009
Indexing and storage engines › data compression
order-preserving string compression
0.112009
Dictionary-based order-preserving string compression for main memory column stores · SIGMOD Conference 2009
Indexing and storage engines
external memory algorithms
0.112015
Cache-Efficient Aggregation: Hashing Is Sorting · SIGMOD Conference 2015
Distributed and cloud data management › large-scale data management
web-scale data management
0.112015
Towards a web-scale data management ecosystem demonstrated by SAP HANA · ICDE 2015
Distributed and cloud data management
data placement
0.012013
SAP HANA distributed in-memory database system: Transaction, session, and metadata management · ICDE 2013
Database system architecture and tuning
hybrid transactional and analytical processing
0.012012
Efficient transaction processing in SAP HANA database: the end of a column store myth · SIGMOD Conference 2012
Data integration and cleaning
data fusion
0.012006
Panel: One Platform for Mining Structured & Unstructured Data: Dream or Reality? · VLDB 2006

Methods — techniques the papers use, named apart from their topics

logical timestamps · 0.4asynchronous update propagation · 0.4sorting · 0.2order-index · 0.2indexing · 0.2hashing · 0.2external memory model · 0.2domain-specific engine integration · 0.2scale-out architecture · 0.2labeling scheme · 0.2
YearPublicationVenuePosition
2020 Data Management for Enterprise Applications
Franz Färber
CIDR1
2017 Order Indexes: supporting highly dynamic hierarchical data in relational main-memory database systems
Jan Finis, Robert Brunel, Alfons Kemper, Thomas Neumann 0001, Norman May, Franz Färber
VLDB J.6
2016 Footprint Reduction and Uniqueness Enforcement with Hash Indices in SAP HANA
Martin Faust, Martin Boissier 0001, Marvin Keller, David Schwalb, Holger Bischoff, Katrin Eisenreich, Franz Färber, Hasso Plattner
DEXA (2)7
2015 Supporting hierarchical data in SAP HANA
abstract
Managing hierarchies is an ever-recurring challenge for relational database systems. Through investigations of customer scenarios at SAP we found that today's RDBMSs still leave a lot to be desired in order to meet the requirements of typical applications. Our research puts a new twist on handling hierarchies in SQL-based systems. We present an approach for modeling hierarchical data natively, and we extend the SQL language with expressive constructs for creating, manipulating, and querying a hierarchy. The constructs can be evaluated efficiently by leveraging existing indexing and query processing techniques. We demonstrate the feasibility of our concepts with initial measurements on a HANA-based prototype.
Robert Brunel, Jan Finis, Gerald Franz, Norman May, Alfons Kemper, Thomas Neumann 0001, Franz Färber
ICDE7
2015 Towards a web-scale data management ecosystem demonstrated by SAP HANA
abstract
Over the years, data management has diversified and moved into multiple directions, mainly caused by a significant growth in the application space with different usage patterns, a massive change in the underlying hardware characteristics, and-last but not least-growing data volumes to be processed. A solution matching these constraints has to cope with a multidimensional problem space including techniques dealing with a large number of domain-specific data types, data and consistency models, deployment scenarios, and processing, storage, and communication infrastructures on a hardware level. Specialized database engines are available and are positioned in the market optimizing a particular dimension on the one hand while relaxing other aspects (e.g. web-scale deployment with relaxed consistency). Today it is common sense, that there is no single engine which can handle all the different dimensions equally well and therefore we have very good reasons to tackle this problem and optimize the dimensions with specialized approaches in a first step. However, we argue for a second step (reflecting in our opinion on the even harder problem) of a deep integration of individual engines into a single coherent and consistent data management ecosystem providing not only shared components but also a common understanding of the overall business semantics. More specifically, a data management ecosystem provides common “infrastructure” for software and data life cycle management, backup/recovery, replication and high availability, accounting and monitoring, and many other operational topics, where administrators and users expect a harmonized experience. More importantly from an application perspective however, customer experience teaches us to provide a consistent business view across all different components and the ability to seamlessly combine different capabilities. For example, within recent customer-based Internet of Things scenarios, a huge potential exists in combining graph-processing functionality with temporal and geospatial information and keywords extracted from high-throughput twitter streams. Using SAP HANA as the running example, we want to demonstrate what moving a set of individual engines and infra-structural components towards a holistic but also flexible data management ecosystem could look like. Although there are some solutions for some problems already visible on the horizon, we encourage the database research community in general to focus more on the Big Picture providing a holistic/integrated approach to efficiently deal with different types of data, with different access methods, and different consistency requirements-research in this field would push the envelope far beyond the traditional notion of data management.
Franz Färber, Jonathan Dees, Martin Weidner, Stefan Bäuerle, Wolfgang Lehner
ICDE1
2015 Cache-Efficient Aggregation: Hashing Is Sorting
abstract
For decades researchers have studied the duality of hashing and sorting for the implementation of the relational operators, especially for efficient aggregation. Depending on the underlying hardware and software architecture, the specifically implemented algorithms, and the data sets used in the experiments, different authors came to different conclusions about which is the better approach. In this paper we argue that in terms of cache efficiency, the two paradigms are actually the same. We support our claim by showing that the complexity of hashing is the same as the complexity of sorting in the external memory model. Furthermore we make the similarity of the two approaches obvious by designing an algorithmic framework that allows to switch seamlessly between hashing and sorting during execution. The fact that we mix hashing and sorting routines in the same algorithmic framework allows us to leverage the advantages of both approaches and makes their similarity obvious. On a more practical note, we also show how to achieve very low constant factors by tuning both the hashing and the sorting routines to modern hardware. Since we observe a complementary dependency of the constant factors of the two routines to the locality of the input, we exploit our framework to switch to the faster routine where appropriate. The result is a novel relational aggregation algorithm that is cache-efficient---independently and without prior knowledge of input skew and output cardinality---, highly parallelizable on modern multi-core systems, and operating at a speed close to the memory bandwidth, thus outperforming the state-of-the-art by up to 3.7x.
Ingo Müller 0002, Peter Sanders 0001, Arnaud Lacurie, Wolfgang Lehner, Franz Färber
SIGMOD Conference5
2015 Indexing Highly Dynamic Hierarchical Data
abstract
Maintaining and querying hierarchical data in a relational database system is an important task in many business applications. This task is especially challenging when considering dynamic use cases with a high rate of complex, possibly skewed structural updates. Labeling schemes are widely considered the indexing technique of choice for hierarchical data, and many different schemes have been proposed. However, they cannot handle dynamic use cases well due to various problems which we investigate in this paper. We therefore propose our dynamic Order Indexes , which offer competitive query performance, unprecedented update efficiency, and robustness for highly dynamic workloads.
Jan Finis, Robert Brunel, Alfons Kemper, Thomas Neumann 0001, Norman May, Franz Färber
Proc. VLDB Endow.6
2015 Towards Scalable Real-time Analytics: An Architecture for Scale-out of OLxP Workloads
abstract
We present an overview of our work on the SAP HANA Scale-out Extension, a novel distributed database architecture designed to support large scale analytics over real-time data. This platform permits high performance OLAP with massive scale-out capabilities, while concurrently allowing OLTP workloads. This dual capability enables analytics over real-time changing data and allows fine grained user-specified service level agreements (SLAs) on data freshness. We advocate the decoupling of core database components such as query processing, concurrency control, and persistence, a design choice made possible by advances in high-throughput low-latency networks and storage devices. We provide full ACID guarantees and build on a logical timestamp mechanism to provide MVCC-based snapshot isolation, while not requiring synchronous updates of replicas. Instead, we use asynchronous update propagation guaranteeing consistency with timestamp validation. We provide a view into the design and development of a large scale data management platform for real-time analytics, driven by the needs of modern enterprise customers.
Anil K. Goel, Jeffrey Pound, Nathan Auch, Peter Bumbulis, Scott MacLean, Franz Färber, Francis Gropengießer, Christian Mathis, Thomas Bodner 0001, Wolfgang Lehner
Proc. VLDB Endow.6
2014 Adaptive String Dictionary Compression in In-Memory Column-Store Database Systems
abstract
Domain encoding is a common technique to compress the columns of a column store and to accelerate many types of queries at the same time. It is based on the assumption that most columns contain a relatively small set of distinct values, in particular string columns. In this paper, we argue that domain encoding is not the end of the story. In real world systems, we observe that a substantial amount of the columns are of string types. Moreover, most of the memory space is consumed by only a small fraction of these columns. To address this issue, we make three main contributions: First we survey several approaches and variants for dictionary compression, i. e., data structures that store the dictionary of domain encoding in a compressed way. As expected, there is a trade-off between size of the data structure and its access performance. This observation can be used to compress rarely accessed data more than frequently ac-cessed data. Furthermore the question which approach has the best compression ratio for a certain column heavily depends on specific characteristics of its content. Consequently, as a second contribu-tion, we present non-trivial sampling schemes for all our dictionary formats, enabling us to estimate their size for a given column. This way it is possible to identify compression schemes specialized for the content of a specific column. Third, we draft how to fully automate the decision of the dic-tionary format. We sketch a compression manager that selects the most appropriate dictionary format based on column access and up-date patterns, characteristics of the underlying data, and costs for set-up and access of the different data structures. We evaluate an off-line prototype of a compression manager using a variation of the TPC-H benchmark [15]. The compression manager can con-figure the database system to be anywhere in a large range of the space / time trade-off with a fine granularity, providing significantly better trade-offs than any fixed dictionary format. 1.
Ingo Müller 0002, Cornelius Ratsch, Franz Färber
EDBT3
2014 Distributed snapshot isolation: global transactions pay globally, local transactions pay locally
Carsten Binnig, Stefan Hildenbrand, Franz Färber, Donald Kossmann, Juchang Lee, Norman May
VLDB J.3
2013 Multi-level Parallel Query Execution Framework for CPU and GPU
Hannes Rauhe, Jonathan Dees, Kai-Uwe Sattler, Franz Färber
ADBIS4
2013 RWS-Diff: flexible and efficient change detection in hierarchical data
abstract
The problem of generating a cost-minimal edit script between two trees has many important applications. However, finding such a cost-minimal script is computationally hard, thus the only methods that scale are approximate ones. Various approximate solutions have been proposed recently. However, most of them still show quadratic or worse runtime complexity in the tree size and thus do not scale well either. The only solutions with log-linear runtime complexity use simple matching algorithms that only find corresponding subtrees as long as these subtrees are equal. Consequently, such solutions are not robust at all, since small changes in the leaves which occur frequently can make all subtrees that contain the changed leaves unequal and thus prevent the matching of large portions of the trees. This problem could be avoided by searching for similar instead of equal subtrees but current similarity approaches are too costly and thus also show quadratic complexity. Hence, currently no robust log-linear method exists.
Jan Finis, Martin Raiber, Nikolaus Augsten, Robert Brunel, Alfons Kemper, Franz Färber
CIKM6
2013 SAP HANA distributed in-memory database system: Transaction, session, and metadata management
abstract
One of the core principles of the SAP HANA database system is the comprehensive support of distributed query facility. Supporting scale-out scenarios was one of the major design principles of the system from the very beginning. Within this paper, we first give an overview of the overall functionality with respect to data allocation, metadata caching and query routing. We then dive into some level of detail for specific topics and explain features and methods not common in traditional disk-based database systems. In summary, the paper provides a comprehensive overview of distributed query processing in SAP HANA database to achieve scalability to handle large databases and heterogeneous types of workloads.
Juchang Lee, Yongsik Kwon, Franz Färber, Michael Muehle, Chulwon Lee, Christian Bensberg, Joo-Yeon Lee, Arthur H. Lee, Wolfgang Lehner
ICDE3
2013 DeltaNI: an efficient labeling scheme for versioned hierarchical data
abstract
Main-memory database systems are emerging as the new backbone of business applications. Besides flat relational data representations also hierarchical ones are essential for these modern applications; therefore we devise a new indexing and versioning approach for hierarchies that is deeply integrated into the relational kernel.
Jan Finis, Robert Brunel, Alfons Kemper, Thomas Neumann 0001, Franz Färber, Norman May
SIGMOD Conference5
2013 Timeline index: a unified data structure for processing queries on temporal data in SAP HANA
abstract
Managing temporal data is becoming increasingly important for many applications. Several database systems already support the time dimension, but provide only few temporal operators, which also often exhibit poor performance characteristics. On the academic side, a large number of algorithms and data structures have been proposed, but they often address a subset of these temporal operators only. In this paper, we develop the Timeline Index as a novel, unified data structure that efficiently supports temporal operators such as temporal aggregation, time travel, and temporal joins. As the Timeline Index is independent of the physical order of the data, it provides flexibility in physical design; e.g., it supports any kind of compression scheme, which is crucial for main memory column stores. Our experiments show that the Timeline Index has predictable performance and beats state-of-the-art approaches significantly, sometimes by orders of magnitude.
Martin Kaufmann, Amin Amiri Manjili, Panagiotis Vagenas, Peter M. Fischer 0001, Donald Kossmann, Franz Färber, Norman May
SIGMOD Conference6
2013 Comprehensive and Interactive Temporal Query Processing with SAP HANA
abstract
In this demo, we present a prototype of a main memory database system which provides a wide range of temporal operators featuring predictable and interactive response times. Much of real-life data is temporal in nature, and there is an increasing application demand for temporal models and operations in databases. Nevertheless, SQL:2011 has only recently overcome a decade-long standstill on standardizing temporal features. As a result, few database systems provide any temporal support, and even those only have limited expressiveness and poor performance. Our prototype combines an in-memory column store and a novel, generic temporal index structure named Timeline Index. As we will show on a workload based on real customer use cases, it achieves predictable and interactive query performance for a wide range of temporal query types and data sizes.
Martin Kaufmann, Panagiotis Vagenas, Peter M. Fischer 0001, Donald Kossmann, Franz Färber
Proc. VLDB Endow.5
2013 SAP HANA: The Evolution from a Modern Main-Memory Data Platform to an Enterprise Application Platform
abstract
SAP HANA is a pioneering, and one of the best performing, data platform designed from the grounds up to heavily exploit modern hardware capabilities, including SIMD, and large memory and CPU footprints. As a comprehensive data management solution, SAP HANA supports the complete data life cycle encompassing modeling, provisioning, and consumption. This extended abstract outlines the vision and planned next step of the SAP HANA evolution growing from a core data platform into an innovative enterprise application platform as the foundation for current as well as novel business applications in both on-premise and on-demand scenarios. We argue that only a holistic system design rigorously applying co-design at different levels may yield a highly optimized and sustainable platform for modern enterprise applications.
Vishal Sikka, Franz Färber, Anil K. Goel, Wolfgang Lehner
Proc. VLDB Endow.2
2012 Data management with SAPs in-memory computing engine
abstract
We present some architectural and technological insights on SAP's HANA database and derive research challenges for future enterprise application development.
Joos-Hendrik Böse, Cafer Tosun, Christian Mathis, Franz Färber
EDBT4
2012 Efficient transaction processing in SAP HANA database: the end of a column store myth
abstract
The SAP HANA database is the core of SAP's new data management platform. The overall goal of the SAP HANA database is to provide a generic but powerful system for different query scenarios, both transactional and analytical, on the same data representation within a highly scalable execution environment. Within this paper, we highlight the main features that differentiate the SAP HANA database from classical relational database engines. Therefore, we outline the general architecture and design criteria of the SAP HANA in a first step. In a second step, we challenge the common belief that column store data structures are only superior in analytical workloads and not well suited for transactional workloads. We outline the concept of record life cycle management to use different storage formats for the different stages of a record. We not only discuss the general concept but also dive into some of the details of how to efficiently propagate records through their life cycle and moving database entries from write-optimized to read-optimized storage formats. In summary, the paper aims at illustrating how the SAP HANA database is able to efficiently work in analytical as well as transactional workload environments.
Vishal Sikka, Franz Färber, Wolfgang Lehner, Sang Kyun Cha, Thomas Peh, Christof Bornhövd
SIGMOD Conference2
2012 A Storage Advisor for Hybrid-Store Databases
abstract
With the SAP HANA database, SAP offers a high-performance in-memory hybrid-store database. Hybrid-store databases---that is, databases supporting row- and column-oriented data management---are getting more and more prominent. While the columnar management offers high-performance capabilities for analyzing large quantities of data, the row-oriented store can handle transactional point queries as well as inserts and updates more efficiently. To effectively take advantage of both stores at the same time the novel question whether to store the given data row- or column-oriented arises. We tackle this problem with a storage advisor tool that supports database administrators at this decision. Our proposed storage advisor recommends the optimal store based on data and query characteristics; its core is a cost model to estimate and compare query execution times for the different stores. Besides a per-table decision, our tool also considers to horizontally and vertically partition the data and manage the partitions on different stores. We evaluated the storage advisor for the use in the SAP HANA database; we show the recommendation quality as well as the benefit of having the data in the optimal store with respect to increased query performance.
Philipp Rösch, Lars Dannecker, Gregor Hackenbroich, Franz Färber
Proc. VLDB Endow.4
2011 Hybrid Data-Flow Graphs for Procedural Domain-Specific Query Languages
Bernhard Jäcksch, Franz Färber, Frank Rosenthal, Wolfgang Lehner
SSDBM2
2011 Bridging Two Worlds with RICE Integrating R into the SAP In-Memory Computing Engine
Philipp Große, Wolfgang Lehner, Thomas Weichert, Franz Färber, Wen-Syan Li
Proc. VLDB Endow.4
2010 Optimizing Write Performance for Read Optimized Databases
Jens Krüger 0003, Martin Grund, Christian Tinnefeld, Hasso Plattner, Alexander Zeier, Franz Färber
DASFAA (2)6
2010 Speeding Up Queries in Column Stores - A Case for Compression
Christian Lemke, Kai-Uwe Sattler, Franz Färber, Alexander Zeier
DaWak3
2010 A plan for OLAP
abstract
So far, data warehousing has often been discussed in the light of complex OLAP queries and as reporting facility for operative data. We argue that business planning as a means to generate plan data is an equally important cornerstone of a data warehouse system, and we propose it to be a first-class citizen within an OLAP engine. We introduce an abstract model describing relevant aspects of the planning process in general and the requirements it poses to a planning engine. Furthermore, we show that business planning lends itself well to parallelization and benefits from a column-store much like traditional OLAP does. We then develop a physical model specifically targeted at a highly parallel column-store, and with our implementation, we show nearly linear scaling behavior.
Bernhard Jäcksch, Wolfgang Lehner, Franz Färber
EDBT3
2010 Cherry picking in database languages
abstract
To avoid expensive round-trips between the application layer and the database layer it is crucial that data-intensive processing and calculations happen close to where the data resides -- ideally within the database engine. However, each application has its own domain and provides domain-specific languages (DSL) as a user interface to keep interactions confined within the well-known metaphors of the respective domain. Revealing the innards of the underlying data layer by forcing users to formulate problems in terms of a general database language is often not an option. To bridge that gap, we propose an approach to transform and directly compile a DSL into a general database execution plan using graph transformations. We identify the commonalities and mismatches between different models and show which parts can be cherry-picked for direct translation. Finally, we argue that graph transformations can be used in general to translate a DSL into an executable plan for a database.
Bernhard Jäcksch, Franz Färber, Wolfgang Lehner
IDEAS2
2009 Dictionary-based order-preserving string compression for main memory column stores
abstract
Column-oriented database systems [19, 23] perform better than traditional row-oriented database systems on analytical workloads such as those found in decision support and business intelligence applications. Moreover, recent work [1, 24] has shown that lightweight compression schemes significantly improve the query processing performance of these systems. One such a lightweight compression scheme is to use a dictionary in order to replace long (variable-length) values of a certain domain with shorter (fixedlength) integer codes. In order to further improve expensive query operations such as sorting and searching, column-stores often use order-preserving compression schemes.
Carsten Binnig, Stefan Hildenbrand, Franz Färber
SIGMOD Conference3
2006 Panel: One Platform for Mining Structured & Unstructured Data: Dream or Reality?
Dina Bitton, Franz Färber, Laura M. Haas, Jayavel Shanmugasundaram
VLDB2