Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hasso Plattner

dblp:58/3831 · DBLP profile ↗
← Back
42ranked-venue papers
3as first author
3since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 27 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6Software engineering, systems software and programming languages · 5 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
10 papers
Database system architecture and tuning · 48% Indexing and storage engines · 23% Data mining · 14%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Cloud and datacenter computing · 48% Performance modeling and evaluation · 34% Memory systems · 6%
Software engineering, system software, and programming languages
2 papers
Empirical software engineering · 72% Software maintenance and evolution · 23% Requirements engineering and software design · 4%

Topics — the 25 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › causal inference › causal modeling
causal discovery
0.512021
MPCSL - A Modular Pipeline for Causal Structure Learning · KDD 2021
Database system architecture and tuning
configuration tuning
0.512021
A Cockpit for the Development and Evaluation of Autonomous Database Systems · ICDE 2021
Database system architecture and tuning
self-managing database systems
0.512021
A Cockpit for the Development and Evaluation of Autonomous Database Systems · ICDE 2021
Indexing and storage engines › column store
main-memory column store
0.432014
The Impact of Columnar In-Memory Databases on Enterprise Systems · Proc. VLDB Endow. 2014
Fast Updates on Read-Optimized Databases Using Multi-Core CPUs · Proc. VLDB Endow. 2011
A common database approach for OLTP and OLAP using an in-memory column database · SIGMOD Conference 2009
Database system architecture and tuning
hybrid transactional and analytical processing
0.332014
The Impact of Columnar In-Memory Databases on Enterprise Systems · Proc. VLDB Endow. 2014
A common database approach for OLTP and OLAP using an in-memory column database · SIGMOD Conference 2009
Fast Updates on Read-Optimized Databases Using Multi-Core CPUs · Proc. VLDB Endow. 2011
Cloud and datacenter computing › multi-tenancy
tenant placement
0.322013
RTP: robust tenant placement for elastic in-memory database clusters · SIGMOD Conference 2013
Predicting in-memory database performance for automating cluster management tasks · ICDE 2011
Database system architecture and tuning
main-memory database
0.212016
Leveraging non-volatile memory for instant restarts of in-memory database systems · ICDE 2016
Empirical software engineering
mining software repositories
0.212016
Lightweight collection and storage of software repository data with DataRover · ASE 2016
Indexing and storage engines
column store
0.222009
SIMD-Scan: Ultra Fast in-Memory Table Scan using on-Chip Vector Processing Units · Proc. VLDB Endow. 2009
A common database approach for OLTP and OLAP using an in-memory column database · SIGMOD Conference 2009
Cloud and datacenter computing
cluster resource management and scheduling
0.212013
RTP: robust tenant placement for elastic in-memory database clusters · SIGMOD Conference 2013
Performance modeling and evaluation
benchmarking
0.112012
Interactive performance monitoring of a composite OLTP and OLAP workload · SIGMOD Conference 2012
Performance modeling and evaluation
performance monitoring
0.112012
Interactive performance monitoring of a composite OLTP and OLAP workload · SIGMOD Conference 2012
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.112011
Predicting in-memory database performance for automating cluster management tasks · ICDE 2011
Performance modeling and evaluation › performance prediction
response time estimation
0.112011
Predicting in-memory database performance for automating cluster management tasks · ICDE 2011
Indexing and storage engines › storage management
hybrid storage engine
0.112010
HYRISE - A Main Memory Hybrid Storage Engine · Proc. VLDB Endow. 2010
Indexing and storage engines
in-memory storage
0.112010
HYRISE - A Main Memory Hybrid Storage Engine · Proc. VLDB Endow. 2010
Query processing and optimization › query execution › scan processing
sequential scan
0.112009
SIMD-Scan: Ultra Fast in-Memory Table Scan using on-Chip Vector Processing Units · Proc. VLDB Endow. 2009
Software maintenance and evolution
software ecosystems
0.112016
Lightweight collection and storage of software repository data with DataRover · ASE 2016
Storage systems › storage reliability
durability
0.112016
Leveraging non-volatile memory for instant restarts of in-memory database systems · ICDE 2016
Memory systems
non-volatile memory
0.112016
Leveraging non-volatile memory for instant restarts of in-memory database systems · ICDE 2016
Query processing and optimization
OLAP
0.012012
Interactive performance monitoring of a composite OLTP and OLAP workload · SIGMOD Conference 2012
Transaction processing and concurrency control
OLTP
0.012012
Interactive performance monitoring of a composite OLTP and OLAP workload · SIGMOD Conference 2012
Processor architecture and microarchitecture
SIMD
0.012009
SIMD-Scan: Ultra Fast in-Memory Table Scan using on-Chip Vector Processing Units · Proc. VLDB Endow. 2009
Processor architecture and microarchitecture › vector processor
vector processing unit
0.012009
SIMD-Scan: Ultra Fast in-Memory Table Scan using on-Chip Vector Processing Units · Proc. VLDB Endow. 2009
Software maintenance and evolution › software reengineering › software modernization › software migration
legacy system migration
0.011996
A Standard Software Application Development: SAP R/3 (Abstract) · ICSE 1996

Methods — techniques the papers use, named apart from their topics

interactive assessment · 0.5causal structure learning · 0.5algorithm benchmarking · 0.5elastic scaling algorithms · 0.3mapping-based schema definition · 0.2data transformation and linkage · 0.2API querying · 0.2performance modeling · 0.1parallelization · 0.1architecture-aware optimization · 0.1cost modeling · 0.1cache miss modeling · 0.1vector processing · 0.1in-memory column storage · 0.1SQL · 0.1SIMD · 0.1software development methodology · 0.0
YearPublicationVenuePosition
2021 A Cockpit for the Development and Evaluation of Autonomous Database Systems
abstract
Databases are highly optimized complex systems with a multitude of configuration options. Especially in cloud scenarios with thousands of database deployments, determining optimized database configurations in an automated fashion is of increasing importance for database providers. At the same time, due to increased system complexity, it becomes more challenging to identify well-performing configurations. Therefore, research interest in autonomous or self-driving database systems has increased enormously in recent years. Such systems promise both performance improvements and cost reductions. In the literature, various fully or partially autonomous optimization mechanisms exist that optimize single aspects, e.g., index selection. However, database administrators and developers often distrust autonomous approaches, and there is a lack of practical experimentation opportunities that could create a better understanding. Moreover, the interplay of different autonomous mechanisms under complex workloads remains an open question. The presented cockpit enables an interactive assessment of the impact of autonomous components for database systems by comparing (autonomous) systems with different configurations side by side. Thereby, the cockpit enables users to build trust in autonomous solutions by experimenting with such technologies and observing their effects in practice.
Jan Kossmann, Martin Boissier 0001, Alexander Dubrawski, Fabian Heseding, Caterina Mandel, Udo Pigorsch, Max Schneider, Til Schniese, Mona Sobhani, Petr Tsayun, Katharina Wille, Michael Perscheid, Matthias Uflacker, Hasso Plattner
ICDE14
2021 MPCSL - A Modular Pipeline for Causal Structure Learning
abstract
The examination of causal structures is crucial for data scientists in a variety of machine learning application scenarios. In recent years, the corresponding interest in methods of causal structure learning has led to a wide spectrum of independent implementations, each having specific accuracy characteristics and introducing implementation-specific overhead in the runtime. Hence, considering a selection of algorithms or different implementations in different programming languages utilizing different hardware setups becomes a tedious manual task with high setup costs. Consequently, a tool that enables to plug in existing methods from different libraries into a single system to compare and evaluate the results is substantial support for data scientists in their research efforts.
Johannes Hügle, Christopher Hagedorn, Michael Perscheid, Hasso Plattner
KDD4
2021 ESPBench: The Enterprise Stream Processing Benchmark
abstract
Growing data volumes and velocities in fields such as Industry 4.0 or the Internet of Things have led to the increased popularity of data stream processing systems. Enterprises can leverage these developments by enriching their core business data and analyses with up-to-date streaming data. Comparing streaming architectures for these complex use cases is challenging, as existing benchmarks do not cover them. ESPBench is a new enterprise stream processing benchmark that fills this gap. We present its architecture, the benchmarking process, and the query workload. We employ ESPBench on three state-of-the-art stream processing systems, Apache Spark, Apache Flink, and Hazelcast Jet, using provided query implementations developed with Apache Beam. Our results highlight the need for the provided ESPBench toolkit that supports benchmark execution, as it enables query result validation and objective latency measures.
Günter Hesse, Christoph Matthies, Michael Perscheid, Matthias Uflacker, Hasso Plattner
ICPE5
2019 Hyrise Re-engineered: An Extensible Database System for Research in Relational In-Memory Data Management
Markus Dreseler, Jan Kossmann, Martin Boissier 0001, Stefan Halfpap, Matthias Uflacker, Hasso Plattner
EDBT6
2016 IMDBfs: Bridging the gap between in-memory database technology and file-based tools for life sciences
abstract
Many established processing and analysis tools in life sciences are still operating on files, e.g. alignment of genome data. Their integration into optimized workflows requires time-consuming transformation, import and export of data. In the given contribution, we introduce IMDBfs: a shared file system operating on top of an in-memory database system. Our IMDBfs provides transparent access to database resources via traditional file system operations. For the first time, it allows seamless integration of file-based life science tools into processes optimized for latest IMDB technology. As a result, existing tools neither need to be ported nor modified whilst transparent IMDB data access is provided.
Matthieu-P. Schapranow, Milena Kraus, Marius Danner, Hasso Plattner
BIBM4
2016 Hyrise-NV: Instant Recovery for In-Memory Databases Using Non-Volatile Memory
David Schwalb, Markus Dreseler, Anusha S., Martin Faust, Adolf Hohl, Tim Berning, Gaurav Makkar, Hasso Plattner, Parag Deshmukh
DASFAA (2)9
2016 Footprint Reduction and Uniqueness Enforcement with Hash Indices in SAP HANA
Martin Faust, Martin Boissier 0001, Marvin Keller, David Schwalb, Holger Bischoff, Katrin Eisenreich, Franz Färber, Hasso Plattner
DEXA (2)8
2016 Agile metrics for a university software engineering course
abstract
Teaching agile software development by pairing lectures with hands-on projects has become the norm. This approach poses the problem of grading and evaluating practical project work as well as process conformance during development. Yet, few best practices exist for measuring the success of students in implementing agile practices. Most university courses rely on observations during the course or final oral exams. In this paper, we propose a set of metrics which give insights into the adherence to agile practices in teams. The metrics identify instances in development data, e.g. commits or user stories, where agile processes were not followed. The identified violations can serve as starting points for further investigation and team discussions. With contextual knowledge of the violation, the executed process or the metric itself can be refined. The metrics reflect our experiences with running a software engineering course over the last five years. They measure aspects which students frequently have issues with and that diminish process adoption and student engagement. We present the proposed metrics, which were tested in the latest course installment, alongside tutoring, lectures, and oral exams.
Christoph Matthies, Thomas Kowark, Matthias Uflacker, Hasso Plattner
FIE4
2016 Leveraging non-volatile memory for instant restarts of in-memory database systems
abstract
Emerging non-volatile memory technologies (NVM) offer fast and byte-addressable access, allowing to rethink the durability mechanisms of in-memory databases. Hyrise-NV is a database storage engine that maintains table and index structures on NVM. Our architecture updates the database state and index structures transactionally consistent on NVM using multi-version data structures, allowing to instantly recover data-bases independent of their size. In this paper, we demonstrate the instant restart capabilities of Hyrise-NV, storing all data on non-volatile memory. Recovering a dataset of size 92.2 GB takes about 53 seconds using our log-based approach, whereas Hyrise-NV recovers in under one second.
David Schwalb, Martin Faust, Markus Dreseler, Pedro Flemming, Hasso Plattner
ICDE5
2016 Lightweight collection and storage of software repository data with DataRover
abstract
The ease of setting up collaboration infrastructures for software engineering projects creates a challenge for researchers that aim to analyze the resulting data. As teams can choose from various available software-as-a-service solutions and can configure them with a few clicks, researchers have to create and maintain multiple implementations for collecting and aggregating the collaboration data in order to perform their analyses across different setups. The DataRover system presented in this paper simplifies this task by only requiring custom source code for API authentication and querying. Data transformation and linkage is performed based on mappings, which users can define based on sample responses through a graphical front end. This allows storing the same input data in formats and databases most suitable for the intended analysis without requiring additional coding. Furthermore, API responses are continuously monitored to detect changes and allow users to update their mappings and data collectors accordingly. A screencast of the described use cases is available at https://youtu.be/mt4ztff4SfU
Thomas Kowark, Christoph Matthies, Matthias Uflacker, Hasso Plattner
ASE4
2015 The Medical Knowledge Cockpit: Real-time analysis of big medical data enabling precision medicine
abstract
Significant medical knowledge has been generated over past decades, but is stored in databases distributed all over the globe using individual data formats. Detailed diagnostic tests result in steadily growing patient data. Consequently, medical experts are facing challenges outside of their field of expertise, e.g. analyzing, interpreting and linking medical data. In this contribution, we share details about our Medical Knowledge Cockpit, an application designed in an interdisciplinary cooperation with medical experts to improve the identification of relevant knowledge. Built upon latest in-memory database technology, it offers medical experts a unique starting point to link and analyze medical data on their own in real time. Result sets providing links back to primary data sources are assembled using patient specifics.
Matthieu-P. Schapranow, Milena Kraus, Cindy Perscheid, Cornelius Bock, Franz Liedke, Hasso Plattner
BIBM6
2015 Interactive, Flexible, and Generic What-If Analyses Using In-Memory Column Stores
Stefan Halfpap, Lars Butzmann, Stephan von Schorlemer, Martin Faust, David Schwalb, Matthias Uflacker, Werner Sinzig, Hasso Plattner
DASFAA (2)8
2015 Using Object-Awareness to Optimize Join Processing in the SAP HANA Aggregate Cache
abstract
The introduction of columnar in-memory databases, along with hardware evolution, has made the execution of transactional and analytical workloads on a single system both feasible and viable. Yet, doing analytics directly on the transactional data introduces an increasing amount of resourceintensive aggregate queries which can slow down the overall system performance in a multi-user environment. To increase the scalability of a system in the presence of multiple such queries, we propose an aggregate cache in the general delta-main architecture that provides an ecient means to handle costly aggregate queries by applying incremental materialized view maintenance and query compensation techniques. Handling aggregate queries based on joins of multiple tables however is still a challenge as query compensation can be very expensive in the delta-main architecture of columnar in-memory databases. Our analysis of enterprise applications has revealed several data schema and workload patterns that can be leveraged for addressing performance of query processing using the aggregate cache. We contribute by presenting an approach to transport the application object semantics into the database system, becoming object-aware, and optimize the query processing using the aggregate cache by applying partition pruning and predicate pushdown in such general delta-main architecture. Our experimental validation using customer data and workloads confirms that this type of optimizations enables ecient usage of the aggregate cache for an even higher share of aggregate queries as one mean to scale the system.
Stephan von Schorlemer, Anisoara Nica, Lars Butzmann, Stefan Halfpap, Hasso Plattner
EDBT5
2014 Towards integrating the detection of genetic variants into an in-memory database
abstract
Next-generation sequencing enables whole genome sequencing within a few hours at a minimum of cost, entailing advanced medical applications such as personalized treatments. However, this recent technology imposes new challenges to alignment and variant calling as subsequent analysis steps. Compared to former sequencing, both must deal with an increasing amount of data to process at a significantly lower data quality - and are currently not capable of that. In this work, we focus on addressing these challenges for identifying Single Nucleotide Polymorphisms, i.e. SNP calling, in genome data as one subtask of variant calling. We propose the application of a column-store in-memory database for efficient data processing and apply the statistical model that is provided by the Genome Analysis Toolkit's UnifiedGenotyper. Comparisons with the UnifiedGenotyper show that our approach can exploit all computational resources available and accelerates SNP calling up to a factor of 22x.
Cindy Fähnrich, Matthieu-P. Schapranow, Hasso Plattner
IEEE BigData3
2014 Concurrent Execution of Mixed Enterprise Workloads on In-Memory Databases
Johannes Wust, Martin Grund, Kai Hoewelmeyer, David Schwalb, Hasso Plattner
DASFAA (1)5
2014 In-memory technology enables interactive drug response analysis
abstract
Latest medical diagnostics generate increasing amounts of big medical data. Specific software tools optimized for the use by healthcare experts and researchers as well as systematic processes for data processing and analysis in clinical and research environments are still missing. Our work focuses on the integration of high-throughput next-generation sequencing data and its systematic processing and its instantaneous analysis to use them in the course of precision medicine. We share our research results on designing a generic research process for drug response analysis including specific software tools built on top of our distributed in-memory computing platform for processing of big medical data. Furthermore, we present our technical foundations as well as process aspects of integrating and combining heterogeneous data sources, such as genome, patient, and experimental data.
Matthieu-P. Schapranow, Konrad Klinghammer, Cindy Fähnrich, Hasso Plattner
Healthcom4
2014 The Impact of Columnar In-Memory Databases on Enterprise Systems
abstract
Five years ago I proposed a common database approach for transaction processing and analytical systems using a columnar in-memory database, disputing the common belief that column stores are not suitable for transactional workloads. Today, the concept has been widely adopted in academia and industry and it is proven that it is feasible to run analytical queries on large data sets directly on a redundancy-free schema, eliminating the need to maintain pre-built aggregate tables during data entry transactions. The resulting reduction in transaction complexity leads to a dramatic simplification of data models and applications, redefining the way we build enterprise systems. First analyses of productive applications adopting this concept confirm that system architectures enabled by in-memory column stores are conceptually superior for business transaction processing compared to row-based approaches. Additionally, our analyses show a shift of enterprise workloads to even more read-oriented processing due to the elimination of updates of transaction-maintained aggregates.
Hasso Plattner
Proc. VLDB Endow.1
2013 Workload-aware aggregate maintenance in columnar in-memory databases
abstract
Database workloads generated by enterprise applications are comprised of short-running transactional as well as long-running analytical queries with resource-intensive aggregations. The expensive aggregate queries can be significantly accelerated by using materialized views. This speed-up, however, comes with the cost of materialized view maintenance which is necessary to guarantee consistency when the underlying data changes. While several view maintenance strategies are applicable in the context of an in-memory column store, their performance depends on various factors, most importantly the ratio between queries accessing the materialized view and queries altering the base data, called insert ratio. As a contribution in this paper, we propose algorithms that determine the best-performing view maintenance strategy based on the currently monitored factors. Using our novel materialized aggregate engine, we are able to switch between view maintenance strategies on demand. We have created cost models for the identified view maintenance strategies that determine at which insert ratio it is advisable to switch to another strategy. Our benchmarks in SanssouciDB reveal that for all identified workloads, switching between maintenance strategies is more beneficial than staying with a single strategy.
Stephan von Schorlemer, Lars Butzmann, Stefan Halfpap, Hasso Plattner
IEEE BigData4
2013 HIG - An in-memory database platform enabling real-time analyses of genome data
abstract
Costs and time required for sequencing of DNA and RNA declined through use of next generation sequencing technology, e.g. up to 30-times coverage reads are generated in less than two days. However, its interpretation and analysis is still a time-consuming process potentially taking weeks. In this work, we present a completely new architecture for processing and analyzing genome data. It builds on the in-memory database technology to eliminate time-consuming file-based data operations and to enable real-time data analysis. We found out that the use of in-memory technology as an integral component for genome data processing and its analysis significantly reduces time and costs to obtain relevant results, e.g. in the course of personalized medicine.
Matthieu-P. Schapranow, Hasso Plattner
IEEE BigData2
2013 Physical Column Organization in In-Memory Column Stores
David Schwalb, Martin Faust, Jens Krüger 0003, Hasso Plattner
DASFAA (2)4
2013 Elastic online analytical processing on RAMCloud
abstract
A shared-nothing architecture is state-of-the-art for deploying a distributed analytical in-memory database management system: it preserves the in-memory performance advantage by processing data locally on each node but is difficult to scale out. Modern switched fabric communication links such as InfiniBand narrow the performance gap between local and remote DRAM data access to a single order of magnitude. Based on these premises, we introduce a distributed in-memory database architecture that separates the query execution engine and data access: this enables a) the usage of a large-scale DRAM-based storage system such as Stanford's RAMCloud and b) the push-down of bandwidth-intensive database operators into the storage system. We address the resulting challenges such as finding the optimal operator execution strategy and partitioning scheme. We demonstrate that such an architecture delivers both: the elasticity of a shared-storage approach and the performance characteristics of operating on local DRAM.
Christian Tinnefeld, Donald Kossmann, Martin Grund, Joos-Hendrik Böse, Frank Renkes, Vishal Sikka, Hasso Plattner
EDBT7
2013 Efficient View Maintenance for Enterprise Applications in Columnar In-Memory Databases
abstract
Enterprise applications such as available-to-promise (ATP), financial accounting, and dunning typically employ a mixed database workload with short-running transactional as well as analytical queries with resource-intensive aggregations. The latter type of queries can be significantly accelerated by using materialized views with pre-calculated aggregates. However, this speed-up comes with the cost of view maintenance which is necessary to guarantee consistency when the underlying data changes. In this paper, we evaluate existing view maintenance strategies in the context of a columnar in-memory database that is designed for mixed workloads. We propose a novel view maintenance strategy that takes the main-delta architecture and resulting merge process of columnar storage into account. A further contribution is a cost model which determines the best maintenance strategy given a specific workload. Our experiments using an ATP application show that our novel strategy outperforms other strategies in mixed workloads with an insert-ratio of more than 40 percent.
Stephan von Schorlemer, Lars Butzmann, Kai Hoewelmeyer, Stefan Halfpap, Hasso Plattner
EDOC5
2013 RTP: robust tenant placement for elastic in-memory database clusters
abstract
In the cloud services industry, a key issue for cloud operators is to minimize operational costs. In this paper, we consider algorithms that elastically contract and expand a cluster of in-memory databases depending on tenants' behavior over time while maintaining response time guarantees.
Jan Schaffner, Tim Januschowski, Megan Kercher, Tim Kraska, Hasso Plattner, Michael J. Franklin, Dean Jacobs
SIGMOD Conference5
2012 Efficient logging for enterprise workloads on column-oriented in-memory databases
abstract
The introduction of a 64 bit address space in commodity operating systems and the constant drop in hardware prices made large capacities of main memory in the order of terabytes technically feasible and economically viable. Especially column-oriented in-memory databases are a promising platform to improve data management for enterprise applications. As in-memory databases hold the primary persistence in volatile memory, some form of recovery mechanism is required to prevent potential data loss in case of failures. Two desirable characteristics of any recovery mechanism are (1) that it has a minimal impact on the running system, and (2) that the system recovers quickly and without any data loss after a failure. This paper introduces an efficient logging mechanism for dictionary-compressed column structures that addresses these two characteristics by (1) reducing the overall log size by writing dictionary-compressed values and (2) allowing for parallel writing and reading of log files. We demonstrate the efficiency of our logging approach by comparing the resulting log-file size with traditional logical logging on a workload produced by a productive enterprise system.
Johannes Wust, Joos-Hendrik Böse, Frank Renkes, Sebastian Blessing, Jens Krüger 0003, Hasso Plattner
CIKM6
2012 An in-depth analysis of data aggregation cost factors in a columnar in-memory database
abstract
Precise prediction of query execution performance is the basis for various database optimization strategies. With columnar in-memory databases, cost modeling changes in two dimensions: First, models for disk-based databases are not well-suited as the new bottleneck is main memory access. Second, the possibility to execute mixed workloads creates new challenges. For transactional and analytical queries with aggregation operations, memory access patterns and thus execution times vary significantly. This paper discusses the influences of data characteristics on aggregation operations and elevates not considered factors by existing cost model approaches. Further, we present benchmarks implemented and executed on a columnar in-memory research database to underline our assumptions.
Stephan von Schorlemer, Hasso Plattner
DOLAP2
2012 Interactive performance monitoring of a composite OLTP and OLAP workload
abstract
Online transaction processing (OLTP) and online analytical processing (OLAP) are thought of as two separate domains, despite sharing the same business data to operate on. This is the result of performance impairments encountered in the past when running on the same system, the workloads becoming ever more sophisticated, leading to contradictory optimization in database design. Recent developments in hardware and database systems are bringing forth research prototypes supporting mixed OLTP and OLAP workloads, challenging this separation. At the same time new benchmarks are proposed to assess these mixed workload systems. In the demonstration, we show an interactive performance monitor and benchmark driver developed for the Composite Benchmark for Transaction Processing and Reporting. The performance monitor allows us to directly determine the impact of changing shares within the workload and to interactively assess behavioral characteristics of different database systems under changing mixed workload conditions.
Anja Bog, Kai Sachs, Hasso Plattner
SIGMOD Conference3
2012 Costs of authentic pharmaceuticals: research on qualitative and quantitative aspects of enabling anti-counterfeiting in RFID-aided supply chains
Matthieu-P. Schapranow, Jürgen Müller 0003, Alexander Zeier, Hasso Plattner
Pers. Ubiquitous Comput.4
2011 Security Extensions for Improving Data Security of Event Repositories in EPCglobal Networks
abstract
Location-based event data is captured in RFID-aided supply chains for tacking individual goods. They are stored in distributed event repositories by involved supply chain parties. Performing anti-counterfeiting checks involves exchange of event data without exposure of sensitive business secrets. Current EPC global standards leave the definition of security strategies open for concrete implementation. We consider data security as the major aspect that needs to be clarified before wide adaption of EPC global standards will be considered by industries. The given work contributes by defining security extensions for EPC global networks and sharing implementation details of our research prototype. We show that incorporating in-memory technology enables history-based access control while keeping response times fast.
Matthieu-P. Schapranow, Alexander Zeier, Hasso Plattner
EUC3
2011 Predicting in-memory database performance for automating cluster management tasks
abstract
In Software-as-a-Service, multiple tenants are typically consolidated into the same database instance to reduce costs. For analytics-as-a-service, in-memory column databases are especially suitable because they offer very short response times. This paper studies the automation of operational tasks in multi-tenant in-memory column database clusters. As a prerequisite, we develop a model for predicting whether the assignment of a particular tenant to a server in the cluster will lead to violations of response time goals. This model is then extended to capture drops in capacity incurred by migrating tenants between servers. We present an algorithm for moving tenants around the cluster to ensure that response time goals are met. In so doing, the number of servers in the cluster may be dynamically increased or decreased. The model is also extended to manage multiple copies of a tenant's data for scalability and availability. We validated the model with an implementation of a multi-tenant clustering framework for SAP's in-memory column database TREX.
Jan Schaffner, Benjamin Eckart, Dean Jacobs, Christian Schwarz 0001, Hasso Plattner, Alexander Zeier
ICDE5
2011 Two Algorithms for Locating Ancestors of a Large Set of Vertices in a Tree
Oleksandr Panchenko, Arian Treffer, Hasso Plattner, Alexander Zeier
ICSOFT (1)3
2011 Cache-conscious data placement in an in-memory key-value store
abstract
Key-value stores which keep the data entirely in main memory can serve applications whose performance criteria cannot be met by disk-based key-value stores. This paper evaluates the performance implications of cache-conscious data placement in an in-memory key-value store by examining how many values have to be stored consecutively in blocks in order to fully exploit memory locality during bandwidth-bound operations. We contribute by introducing a random block traversal main memory access pattern, by describing the corresponding memory access costs as well as by formally and experimentally deriving the correlation between block size and throughput. Our calculations and experiments vary the value and block sizes as well as their placement in the memory and derive their impact on cache-misses throughout the different memory hierarchies, the ability to prefetch data, and the number of needed CPU cycles to perform a certain set of data operations. The paper closes with the insight that a block-wise grouping of relatively few key-value pairs increases the throughput up to a factor six and with a discussion which implications a block-wise grouping of data has on the system design key-value store.
Christian Tinnefeld, Alexander Zeier, Hasso Plattner
IDEAS3
2011 Precise and Scalable Querying of Syntactical Source Code Patterns Using Sample Code Snippets and a Database
abstract
While analyzing a log file of a text-based source code search engine we discovered that developers search for fine-grained syntactical patterns in 36% of queries. Currently, to cope with queries of this kind developers need to use regular expressions, to add redundant terms to the query or to combine searching with other tools provided by the development environment. To improve the expressiveness of the queries, these can be formulated as tree patterns of abstract syntax trees. These search patterns can be expressed by using query languages, such as XPath. However, developers usually do not work with either XPath or with AST. To shield developers from the complexity of query formulation we propose using sample code snippets as queries. The novelty of our approach is the combination of a query language that is very close to the surface programming language and a special database technology to store a large amount of abstract syntax trees. The advantage of this approach over existing source code query languages and search engines is the performance of both query formulation and query execution. This paper describes the technical details of the method and illustrates the value of this approach with performance measures and an industrial controlled experiment. All developers were able to complete the tasks of the experiment faster and more accurately by using our tool (ACS) than by using a text-based search engine. The number of false positives in the result lists was significantly decreased.
Oleksandr Panchenko, Jan Karstens, Hasso Plattner, Alexander Zeier
ICPC3
2011 Fast Updates on Read-Optimized Databases Using Multi-Core CPUs
abstract
Read-optimized columnar databases use differential updates to handle writes by maintaining a separate write-optimized delta partition which is periodically merged with the read-optimized and compressed main partition. This merge process introduces significant overheads and unacceptable downtimes in update intensive systems, aspiring to combine transactional and analytical workloads into one system. In the first part of the paper, we report data analyses of 12 SAP Business Suite customer systems. In the second half, we present an optimized merge process reducing the merge overhead of current systems by a factor of 30. Our linear-time merge algorithm exploits the underlying high compute and bandwidth resources of modern multi-core CPUs with architecture-aware optimizations and efficient parallelization. This enables compressed in-memory column stores to handle the transactional update rate required by enterprise applications, while keeping properties of read-optimized databases for analytic-style queries.
Jens Krüger 0003, Changkyu Kim, Martin Grund, Nadathur Satish, David Schwalb, Jatin Chhugani, Hasso Plattner, Pradeep Dubey, Alexander Zeier
Proc. VLDB Endow.7
2010 Optimizing Write Performance for Read Optimized Databases
Jens Krüger 0003, Martin Grund, Christian Tinnefeld, Hasso Plattner, Alexander Zeier, Franz Färber
DASFAA (2)4
2010 Enterprise Application-Specific Data Management
abstract
Enterprise applications are presently built on a 20-year old data management infrastructure that was designed to meet a specific set of requirements for OLTP systems. In the meantime, enterprise applications have become more sophisticated, data set sizes have increased, requirements on the freshness of input data have been strengthened, and the time allotted for completing business processes has been reduced. To meet these challenges, enterprise applications have become increasingly complicated to make up for short-comings in the data management infrastructure. To address this issue we investigate recent trends in data management such as main memory databases, column stores and compression techniques with regards to the workload requirements and data characteristics derived from actual customer systems. We show that a main memory column store is better suited for to days enterprise systems, which we validate by using SAP's Net Weaver Business Warehouse Accelerator and a realistic set of data from an inventory management application.
Jens Krüger 0003, Martin Grund, Alexander Zeier, Hasso Plattner
EDOC4
2010 How to juggle columns: an entropy-based approach for table compression
abstract
Many relational databases exhibit complex dependencies between data attributes, caused either by the nature of the underlying data or by explicitly denormalized schemas. In data warehouse scenarios, calculated key figures may be materialized or hierarchy levels may be held within a single dimension table. Such column correlations and the resulting data redundancy may result in additional storage requirements. They may also result in bad query performance if inappropriate independence assumptions are made during query compilation. In this paper, we tackle the specific problem of detecting functional dependencies between columns to improve the compression rate for column-based database systems, which both reduces main memory consumption and improves query performance. Although a huge variety of algorithms have been proposed for detecting column dependencies in databases, we maintain that increased data volumes and recent developments in hardware architectures demand novel algorithms with much lower runtime overhead and smaller memory footprint. Our novel approach is based on entropy estimations and exploits a combination of sampling and multiple heuristics to render it applicable for a wide range of use cases. We demonstrate the quality of our approach by means of an implementation within the SAP NetWeaver Business Warehouse Accelerator. Our experiments indicate that our approach scales well with the number of columns and produces reliable dependence structure information. This both reduces memory consumption and improves performance for nontrivial queries.
Marcus Paradies, Christian Lemke, Hasso Plattner, Wolfgang Lehner, Kai-Uwe Sattler, Alexander Zeier, Jens Krüger 0003
IDEAS3
2010 Enabling real-time charging for smart grid infrastructures using in-memory databases
abstract
The emerging trend towards smart grids defines new requirements for designing enterprise applications for the energy market. Current solutions were built to process single billing runs as time-consuming batch jobs. Rather than processing some readings per year and household, a constant stream of meter readings has to be processed in context of a smart power grid. Additionally, consumers demand for convenient ways to monitor power consumption while getting real-time charging information on a daily basis. In this paper, we share our experiences of integrating meter readings in an industry-specific enterprise application for utilities to support real-time billing. Our prototype enables customers to track consumptions and to view billing details online. Besides, we discuss new business scenarios enabled by this real-timeliness.
Matthieu-P. Schapranow, Ralph Kühne, Alexander Zeier, Hasso Plattner
LCN4
2010 A Dynamic Mutual RFID Authentication Model Preventing Unauthorized Third Party Access
abstract
RFID implementations leverage competitive business advantages in processing, tracking and tracing of fast-moving goods. Most of them suffer from security threats and the resulting privacy risks as RFID technology was not designed for exchange of sensible data. Emerging global RFID-aided supply chains require open interfaces for data exchange of confidential business data between business partners. We present a mutual authentication model based on one-time passwords preventing tag access by unauthorized third parties. Compared to models using complex on-tag encryption methods our implementation focuses on reducing tag-manufacturing costs while increasing customers' acceptance for RFID technology.
Matthieu-P. Schapranow, Alexander Zeier, Hasso Plattner
NSS3
2010 HYRISE - A Main Memory Hybrid Storage Engine
abstract
In this paper, we describe a main memory hybrid database system called HYRISE, which automatically partitions tables into vertical partitions of varying widths depending on how the columns of the table are accessed. For columns accessed as a part of analytical queries (e.g., via sequential scans), narrow partitions perform better, because, when scanning a single column, cache locality is improved if the values of that column are stored contiguously. In contrast, for columns accessed as a part of OLTP-style queries, wider partitions perform better, because such transactions frequently insert, delete, update, or access many of the fields of a row, and co-locating those fields leads to better cache locality. Using a highly accurate model of cache misses, HYRISE is able to predict the performance of different partitionings, and to automatically select the best partitioning using an automated database design algorithm. We show that, on a realistic workload derived from customer applications, HYRISE can achieve a 20% to 400% performance improvement over pure all-column or all-row designs, and that it is both more scalable and produces better designs than previous vertical partitioning approaches for main memory systems.
Martin Grund, Jens Krüger 0003, Hasso Plattner, Alexander Zeier, Philippe Cudré-Mauroux, Samuel Madden 0001
Proc. VLDB Endow.3
2009 A common database approach for OLTP and OLAP using an in-memory column database
abstract
When SQL and the relational data model were introduced 25 years ago as a general data management concept, enterprise software migrated quickly to this new technology. It is fair to say that SQL and the various implementations of RDBMSs became the backbone of enterprise systems. In those days. we believed that business planning, transaction processing and analytics should reside in one single system. Despite the incredible improvements in computer hardware, high-speed networks, display devices and the associated software, speed and flexibility remained an issue.
Hasso Plattner
SIGMOD Conference1
2009 SIMD-Scan: Ultra Fast in-Memory Table Scan using on-Chip Vector Processing Units
abstract
The availability of huge system memory, even on standard servers, generated a lot of interest in main memory database engines. In data warehouse systems, highly compressed column-oriented data structures are quite prominent. In order to scale with the data volume and the system load, many of these systems are highly distributed with a shared-nothing approach. The fundamental principle of all systems is a full table scan over one or multiple compressed columns. Recent research proposed different techniques to speedup table scans like intelligent compression or using an additional hardware such as graphic cards or FPGAs. In this paper, we show that utilizing the embedded Vector Processing Units (VPUs) found in standard superscalar processors can speed up the performance of mainmemory full table scan by factors. This is achieved without changing the hardware architecture and thereby without additional power consumption. Moreover, as on-chip VPUs directly access the system's RAM, no additional costly copy operations are needed for using the new SIMD-scan approach in standard main memory database engines. Therefore, we propose this scan approach to be used as the standard scan operator for compressed column-oriented main memory storage. We then discuss how well our solution scales with the number of processor cores; consequently, to what degree it can be applied in multi-threaded environments. To verify the feasibility of our approach, we implemented the proposed techniques on a modern Intel multi-core processor using Intel® Streaming SIMD Extensions (Intel® SSE). In addition, we integrated the new SIMD-scan approach into SAP® Netweaver® Business Warehouse Accelerator. We conclude with describing the performance benefits of using our approach for processing and scanning compressed data using VPUs in column-oriented main memory database systems.
Thomas Willhalm, Nicolae Popovici, Yazan Boshmaf, Hasso Plattner, Alexander Zeier, Jan Schaffner
Proc. VLDB Endow.4
1996 A Standard Software Application Development: SAP R/3 (Abstract)
Hasso Plattner
ICSE1