Stephan Baumann 0002

dblp:16/6503-2 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
3since 2021 · last 2024
0009-0006-2557-9321ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2024 So Far and yet so Near - Accelerating Distributed Joins with CXL
abstract
Distributed partitioned joins are one of the most expensive operators in distributed DBMSs where a major part of the execution is attributed to network transfer costs. Although high-speed network technologies, such as RDMA, can lower this cost, they still come with significantly higher latency than local DRAM access. The emerging CXL interconnect protocol promises to provide direct and cache-coherent access to remote memory while offering byte-addressable memory access without CPU intervention. For short-distance communication in distributed DBMSs, CXL represents an interesting alternative for low-latency requirements. In this work, we explore how CXL can be leveraged for engine-internal communication and data exchange. We discuss and apply communication strategies to distributed joins. We emulate various CXL characteristics based on optimistic and pessimistic assumptions on the real performance of upcoming CXL devices and evaluate their impact on the execution of distributed joins. Our results show that CXL has the potential to improve distributed join performance.
Alexander Baumstark, Marcus Paradies, Kai-Uwe Sattler, Steffen Kläbe, Stephan Baumann 0002
DaMoN5
2021 Updatable Materialization of Approximate Constraints
abstract
Modern big data applications integrate data from various sources. As a result, these datasets may not satisfy perfect constraints, leading to sparse schema information and non-optimal query performance. The existing approach of PatchIndexes enable the definition of approximate constraints and improve query performance by exploiting the materialized constraint information. As real world data warehouse workloads are often not limited to read-only queries, we enhance the PatchIndex structure towards an update-conscious design in this paper. Therefore, we present a sharded bitmap as the underlying data structure which offers efficient update operations, and describe approaches to maintain approximate constraints under updates, avoiding index recomputations and full table scans. In our evaluation, we prove that PatchIndexes provide more lightweight update support than traditional materialization approaches.
Steffen Kläbe, Kai-Uwe Sattler, Stephan Baumann 0002
ICDE3
2021 PatchIndex: exploiting approximate constraints in distributed databases
abstract
Abstract Cloud data warehouse systems lower the barrier to access data analytics. These applications often lack a database administrator and integrate data from various sources, potentially leading to data not satisfying strict constraints. Automatic schema optimization in self-managing databases is difficult in these environments without prior data cleaning steps. In this paper, we focus on constraint discovery as a subtask of schema optimization. Perfect constraints might not exist in these unclean datasets due to a small set of values violating the constraints. Therefore, we introduce the concept of a generic PatchIndex structure, which handles exceptions to given constraints and enables database systems to define these approximate constraints. We apply the concept to the environment of distributed databases, providing parallel index creation approaches and optimization techniques for parallel queries using PatchIndexes. Furthermore, we describe heuristics for automatic discovery of PatchIndex candidate columns and prove the performance benefit of using PatchIndexes in our evaluation.
Steffen Kläbe, Kai-Uwe Sattler, Stephan Baumann 0002
Distributed Parallel Databases3
2020 Elastic Scaling in VectorH
abstract
Cloud infrastructures allow to dynamically adapt to workload changes by provisioning additional resources or deprovisioning resources to reduce costs. This offers also opportunities for scalable distributed data management. However, elastic scaling in databases requires to migrate or even repartition data. In this work, we present an approach implemented in Actian’s MPP solution VectorH that speeds up the elastic resizing process by minimizing partition reassignments while still achieving load balancing. Moreover, we describe a buffer matching and prefilling technique to further increase performance after the resize step. The experimental evaluation shows that our solution significantly outperforms the non-elastic way of scaling using a system restart by a factor of 2 up to 4 and reduces downtimes during resizing to less than one minute.
Steffen Kläbe, Kai-Uwe Sattler, Stephan Baumann 0002, Michael Rink 0001
EDBT3
2016 Bitwise dimensional co-clustering for analytical workloads
Stephan Baumann 0002, Peter Boncz, Kai-Uwe Sattler
VLDB J.1
2012 Data3 - A Kinect Interface for OLAP Using Complex Event Processing
abstract
Motion sensing input devices like Microsoft's Kinect offer an alternative to traditional computer input devices like keyboards and mouses. Daily new applications using this interface appear. Most of them implement their own gesture detection. In our demonstration we show a new approach using the data stream engine Andu IN. The gesture detection is done based on Andu IN's complex event processing functionality. This way we build a system that allows to define new and complex gestures on the basis of a declarative programming interface. On this basis our demonstration data3provides a basic natural interaction OLAP interface for a sample star schema database using Microsoft's Kinect.
Steffen Hirte, Andreas Seifert, Stephan Baumann 0002, Daniel Klan, Kai-Uwe Sattler
ICDE3
2010 Flashing databases: expectations and limitations
abstract
Flash devices (solid state disks) promise a significant performance improvement for disk-based database processing. However, database storage structures and processing strategies originally designed for magnetic disks prevent the optimal utilization of SSDs. Based on previous work on bench-marking SSDs and a detailed discussion of I/O methods, in this paper, we analyze appropriate execution methods for database processing as well as important parameters and boundaries and present a tool which helps to derive these parameters.
Stephan Baumann 0002, Giel de Nijs, Kai-Uwe Sattler
DaMoN1