Veit Köppen

dblp:56/4614 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
1since 2021 · last 2021
0000-0002-6068-3275ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Indexing and storage engines · 52% Query processing and optimization · 48%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Indexing and storage engines
in-memory index
0.722019
Efficient Evaluation of Multi-Column Selection Predicates in Main-Memory · IEEE Trans. Knowl. Data Eng. 2019
Accelerating Multi-Column Selection Predicates in Main-Memory - The Elf Approach · ICDE 2017
Query processing and optimization › shared computation
shared filter evaluation
0.722019
Efficient Evaluation of Multi-Column Selection Predicates in Main-Memory · IEEE Trans. Knowl. Data Eng. 2019
Accelerating Multi-Column Selection Predicates in Main-Memory - The Elf Approach · ICDE 2017
Indexing and storage engines › multidimensional indexing
high-dimensional indexing
0.212013
QuEval: Beyond high-dimensional indexing a la carte · Proc. VLDB Endow. 2013
Query processing and optimization › query execution › scan processing
scan acceleration
0.112017
Accelerating Multi-Column Selection Predicates in Main-Memory - The Elf Approach · ICDE 2017
Performance modeling and evaluation
benchmarking
0.012013
QuEval: Beyond high-dimensional indexing a la carte · Proc. VLDB Endow. 2013
Performance modeling and evaluation › benchmarking › database system benchmarking
index benchmarking
0.012013
QuEval: Beyond high-dimensional indexing a la carte · Proc. VLDB Endow. 2013

Methods — techniques the papers use, named apart from their topics

data compression · 0.7cache-sensitive layout · 0.4SIMD · 0.4empirical evaluation · 0.3cache-sensitive storage layout · 0.3
YearPublicationVenuePosition
2021 Towards multi-purpose main-memory storage structures: Exploiting sub-space distance equalities in totally ordered data sets for exact knn queries
abstract
Efficient knn computation for high-dimensional data is an important, yet challenging task. Today, most information systems use a column-store back-end for relational data. For such systems, multi-dimensional indexes accelerating selections are known. However, they cannot be used to accelerate knn queries. Consequently, one relies on sequential scans, specialized knn indexes, or trades result quality for speed. To avoid storing one specialized index per query type, we envision multipurpose indexes allowing to efficiently compute multiple query types. In this paper, we focus on additionally supporting knn queries as first step towards this goal. To this end, we study how to exploit total orders for accelerating knn queries based on the sub-space distance equalities observation. It means that non-equal points in the full space, which are projected to the same point in a sub space, have the same distance to every other point in this sub space. In case one can easily find these equalities and tune storage structures towards them, this offers two effects one can exploit to accelerate knn queries. The first effect allows pruning of point groups based on a cascade of lower bounds. The second allows to re-use previously computed sub-space distances between point groups. This results in a worst-case execution bound, which is independent of the distance function. We present knn algorithms exploiting both effects and show how to tune a storage structure already known to work well for multi-dimensional selections. Our investigations reveal that the effects are robust to increasing, e.g., the dimensionality, suggesting generally good knn performance. Comparing our knn algorithms to well-known competitors reveals large performance improvements up to one order of magnitude. Furthermore, the algorithms deliver at least comparable performance as the next fastest competitor suggesting that the algorithms are only marginally affected by the curse of dimensionality.
Martin Schäler, Christine Schäler, Veit Köppen, David Broneske, Gunter Saake
Inf. Syst.3
2020 Combining Two Worlds: MonetDB with Multi-Dimensional Index Structure Support to Efficiently Query Scientific Data
abstract
Reproducibility and generalizability are important criteria for today’s data management society. Hence, stand-alone solutions that work well in isolation, but cannot convince at system level lead to a frustrating user experience. As a consequence, in our demo, we take the step of accelerating queries on scientific data by integrating the multi-dimensional index structure Elf into the main-memory-optimized database management system MonetDB. The overall intention is to show that the stand-alone speed ups of using Elf can also be observed when integrated into a holistic system storing scientific data sets. In our prototypical implementation, we demonstrate the performance of an Elf-backed MonetDB on the standard OLAP-benchmark, TPC-H, and the genomic multi-dimensional range query benchmark from the scientific data community. Queries can be run live on both benchmarks by the audience, while they are able to create different indexes to accelerate selection performance.
Paul Blockhaus, David Broneske, Martin Schäler, Veit Köppen, Gunter Saake
SSDBM4
2019 Efficient Evaluation of Multi-Column Selection Predicates in Main-Memory
abstract
Efficient evaluation of selection predicates is a performance-critical task, for instance to reduce intermediate result sizes being the input for further operations. With analytical queries getting more and more complex, the number of evaluated selection predicates per query and table rises, too. This leads to numerous multi-column selection predicates. Recent approaches to increase the performance of main-memory databases for selection-predicate evaluation aim at optimally exploiting the speed of the CPU by using accelerated scans. However, scanning each column one by one leaves tuning opportunities open that arise if all predicates are considered together. To this end, we introduce Elf, an index structure that is able to exploit the relation between several selection predicates. Elf features cache sensitivity, an optimized storage layout, fixed search paths, and slight data compression. In a large-scale evaluation, we compare its query performance to state-of-the-art approaches and a sequential scan using SIMD capabilities. Our results indicate a clear superiority of our approach for queries returning less than 10 percent of all tuples - a selectivity almost one order of magnitude larger than observed for related indexing approaches. For TPC-H queries with multi-column selection predicates, we achieve a speedup between factor five and two orders of magnitude, mainly depending on the selectivity of the predicates. Further scaling experiments reveal that for large data sets, these speedup factors are expected to increase, due to more densely populated data spaces. Finally, our results indicate that using a delta-store like concept to support periodic insertions results in virtually no performance penalty for reasonable sizes of a write-optimized Elf as delta store.
David Broneske, Veit Köppen, Gunter Saake, Martin Schäler
IEEE Trans. Knowl. Data Eng.2
2017 Accelerating Multi-Column Selection Predicates in Main-Memory - The Elf Approach
abstract
Evaluating selection predicates is a data-intensive task that reduces intermediate results, which are the input for further operations. With analytical queries getting more and more complex, the number of evaluated selection predicates per query and table rises, too. This leads to numerous multicolumn selection predicates. Recent approaches to increase the performance of main-memory databases for selection-predicate evaluation aim at optimally exploiting the speed of the CPU by using accelerated scans. However, scanning each column one by one leaves tuning opportunities open that arise if all predicates are considered together. To this end, we introduce Elf, an index structure that is able to exploit the relation between several selection predicates. Elf features cache sensitivity, an optimized storage layout, fixed search paths, and slight data compression. In our evaluation, we compare its query performance to state-of the-art approaches and a sequential scan using SIMD capabilities. Our results indicate a clear superiority of our approach for multicolumn selection predicate queries with a low combined selectivity. For TPC-H queries with multi-column selection predicates, we achieve a speed-up between a factor of five and two orders of magnitude, mainly depending on the selectivity of the predicates.
David Broneske, Veit Köppen, Gunter Saake, Martin Schäler
ICDE2
2015 An analytical model for data persistence in Business Data Warehouses
abstract
Redundancy of data persistence in Data Warehouses is mostly justified with better performance when accessing data for analysis. However, there are other reasons to store data redundantly, which are often not recognized when designing data warehouses. Especially in Business Data Warehouses, data management via multiple persistence levels is necessary to condition the huge amount of data into an adequate format for its final usage. Redundant data allocates additional disk space and requires time-consuming processing and huge effort for complex maintenance. That means in reverse: avoiding data persistence leads to less effort. The question arises: What data for what purposes do really need to be stored? In this paper, we discuss decision support and evaluation approaches beyond cost-based comparisons. We use a compendium of purposes for data persistence. We define a model that includes objective indicators and subjective user preferences for decision making on data persistence in Business Data Warehouses. We develop an indicator system that enables the measurement of technical as well as business-related facts. With multi-criteria decision methodology, we present a framework to objectively compare different alternatives for data persistence. Finally, we apply our developed method to a real world Business Data Warehouse and show applicability and integration of our model in an existing system.
Veit Köppen, Thorsten Winsemann, Gunter Saake
RCIS1
2014 Toward variability management to tailor high dimensional index implementations
abstract
The increasing amount of complex data requires a solution to store and query these data efficiently. One possibility to speed-up various query types is the application of high dimensional index structures. In prior work, we introduced QuEval as platform to evaluate these indexes for user-defined use cases. Our design allows to easily extend QuEval with new index structure implementations. However, based on our experiences, we encountered severe challenges by tailoring index structure implementations to specific use cases. In particular, we face challenges to manage several similar implementation variants of the same index. In this paper, we consequently show benefits and drawbacks that emphasize the necessity to tailor index structure implementations with the help of a short evaluation study. Finally, we outline approaches for adequate variability management to address the aforementioned drawbacks.
Veit Köppen, Martin Schäler, Reimar Schröter
RCIS1
2014 Relational on demand data management for IT-services
abstract
Database systems are widely used in technical applications. However, it is difficult to decide which database management system fits best for a certain application. For many applications, different workload types often blend to mixed workloads that cause mixed requirements. The selection of an appropriate database management system is more critical for mixed workloads because classical domains with complementary requirements are combined, e.g., OLTP and OLAP. A definite decision for a database management system is not possible. Hybrid database system are developed to accept this challenge, i.e., these systems combine different storage approaches. However, a mutual optimization in hybrid systems is not available for mixed workloads. We develop a decision-support framework to provide application-performance estimation on a certain database management system on the one hand and to provide query optimization for hybrid database systems on the other hand. In this paper, we combine heuristics to a rule-based query optimization framework for hybrid relational database systems. That is, we aim on support for IT-services with volatile requirements. We evaluate the Aqua2framework on standard database benchmarks. and show an acceleration of query execution on hybrid database systems.
Andreas Lübcke, Martin Schäler, Veit Köppen, Gunter Saake
RCIS3
2013 QuEval: Beyond high-dimensional indexing a la carte
abstract
In the recent past, the amount of high-dimensional data, such as feature vectors extracted from multimedia data, increased dramatically. A large variety of indexes have been proposed to store and access such data efficiently. However, due to specific requirements of a certain use case, choosing an adequate index structure is a complex and time-consuming task. This may be due to engineering challenges or open research questions. To overcome this limitation, we present QuEval, an open-source framework that can be flexibly extended w.r.t. index structures, distance metrics, and data sets. QuEval provides a unified environment for a sound evaluation of different indexes, for instance, to support tuning of indexes. In an empirical evaluation, we show how to apply our framework, motivate benefits, and demonstrate analysis possibilities.
Martin Schäler, Alexander Grebhahn, Reimar Schröter, Sandro Schulze, Veit Köppen, Gunter Saake
Proc. VLDB Endow.5
2012 Persistence in Data Warehousing
abstract
Persistence of redundant data in Data Warehouses is often simply justified with an achievement of better performance when accessing data for analysis and reporting. However, there are other reasons to store data persistently, which are often not recognized when designing Data Warehouses. As processing and maintenance of data is complex and requires huge effort, less redundancy downsizes effort. Latest in-memory technologies enable good response times for data access. That arises the question, what data for what purposes really need to be stored persistently. We present a compendium of purposes for data persistence and use it as a basis for decision-making whether to store data or not.
Thorsten Winsemann, Veit Köppen
RCIS2
2011 Using background colors to support program comprehension in software product lines
abstract
Background: Software product line engineering provides an effective mechanism to implement variable software.However, the usage of preprocessors, which is typical in industry, is heavily criticized, because it often leads to obfuscated code.Using background colors to support comprehensibility has shown effective, however, scalability to large software product lines (SPLs) is questionable.Aim: Our goal is to implement and evaluate scalable usage of background colors for industrial-sized SPLs.Method: We designed and implemented scalable concepts in a tool called FeatureCommander.To evaluate its effectiveness, we conducted a controlled experiment with a large real-world SPL with over 160,000 lines of code and 340 features.We used a within-subjects design with treatments colors and no colors.We compared correctness and response time of tasks for both treatments.Results: For certain kinds of tasks, background colors improve program comprehension.Furthermore, subjects generally favor background colors.Conclusion: We show that background colors can improve program comprehension in large SPLs.Based on these encouraging results, we will continue our work improving program comprehension in large SPLs.Difficulty U value 20.5 24.5 17.5 18 10.
Janet Siegmund, Michael Schulze, Maria Papendieck, Christian Kästner, Raimund Dachselt, Veit Köppen, Mathias Frisch
EASE6
2011 A decision model to select the optimal storage architecture for relational databases
abstract
Requirements for database systems differ from small-scale database programs for embedded devices with minimal footprint to large-scale on-line analytical processing applications. For relational database management systems, two storage architectures have been introduced: a) row-oriented architecture and b) column-oriented architecture. In this paper, we present a query decomposition approach to evaluate database operations with respect to their performance according to the storage architecture. We map decomposed queries to workload patterns which contain aggregated database statistics. Further, we develop our complementary decision models which advise the selection of the optimal storage architecture for a given application domain. The first decision model improves the performance of running systems (on-line). The second and third decision model advise an efficient database design or decide which architecture is more suitable for a given application domain (off-line).
Andreas Lübcke, Veit Köppen, Gunter Saake
RCIS2
2010 Combining Schema and Level-Based Matching for Web Service Discovery
Alsayed Algergawy, Richi Nayak, Norbert Siegmund, Veit Köppen, Gunter Saake
ICWE4
2006 Edits - Data Cleansing at the Data Entry to assert semantic Consistency of metric Data
abstract
It is a matter of fact that the input of numeric data into databases needs careful screening to avoid semantic incoherency with respect to the knowledge at hand. In nearly all real applications such knowledge exists as models, i.e. as balance equations, behavioral equations or simply as definitions. The representation of those objects is possible by validation rules ("edits"), which are roughly speaking specially tailored tests. The methodology is presented, recent work in progress is shown, and a business application is presented.
Hans-Joachim Lenz, Veit Köppen, Roland M. Müller
SSDBM2