Thomas Legler

dblp:57/5597 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
2since 2021 · last 2022
0009-0006-8775-0446ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Query processing and optimization · 30% Database system architecture and tuning · 28% Indexing and storage engines · 15%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 41% Storage systems · 41% Performance modeling and evaluation · 12%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Database system architecture and tuning
main-memory database
0.732022
Cost Modelling for Optimal Data Placement in Heterogeneous Main Memory · Proc. VLDB Endow. 2022
Interleaving with Coroutines: A Practical Approach for Robust Index Joins · Proc. VLDB Endow. 2017
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Storage systems
data placement
0.612022
Cost Modelling for Optimal Data Placement in Heterogeneous Main Memory · Proc. VLDB Endow. 2022
Memory systems
hybrid memory
0.612022
Cost Modelling for Optimal Data Placement in Heterogeneous Main Memory · Proc. VLDB Endow. 2022
Indexing and storage engines
b+-tree
0.212014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Transaction processing and concurrency control › transactional memory
hardware transactional memory
0.212014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Indexing and storage engines
in-memory index
0.212014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Transaction processing and concurrency control
synchronization
0.212014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Performance modeling and evaluation
cost modeling
0.212022
Cost Modelling for Optimal Data Placement in Heterogeneous Main Memory · Proc. VLDB Endow. 2022
Programming languages and type systems › control operators
coroutines
0.112019
Interleaving with coroutines: a systematic and practical approach to hide memory latency in index joins · VLDB J. 2019
Data mining › pattern mining
frequent pattern mining
0.112009
Robust Distributed Top-N Frequent Pattern Mining Using the SAP BW Accelerator · Proc. VLDB Endow. 2009
Data mining › pattern mining › itemset mining › frequent itemset mining
top-k frequent items
0.112009
Robust Distributed Top-N Frequent Pattern Mining Using the SAP BW Accelerator · Proc. VLDB Endow. 2009
Distributed systems › distributed data processing
distributed data mining
0.112009
Robust Distributed Top-N Frequent Pattern Mining Using the SAP BW Accelerator · Proc. VLDB Endow. 2009
Data stream processing
continuous query processing
0.112005
Robust Real-time Query Processing with QStream · VLDB 2005
Query processing and optimization › online query processing
real-time query processing
0.112005
Robust Real-time Query Processing with QStream · VLDB 2005
Query processing and optimization › query optimization
robust query processing
0.112005
Robust Real-time Query Processing with QStream · VLDB 2005
Data mining › pattern mining
association rule mining
0.012009
Robust Distributed Top-N Frequent Pattern Mining Using the SAP BW Accelerator · Proc. VLDB Endow. 2009
Query processing and optimization
OLAP
0.012006
Data Mining with the SAP Netweaver BI Accelerator · VLDB 2006

Methods — techniques the papers use, named apart from their topics

workload characterization · 1.1pareto optimization · 1.1cost modeling · 1.1coroutines · 1.0instruction stream interleaving · 0.3minimum absolute-support threshold · 0.2heuristic search · 0.2hardware transactional memory · 0.2Intel TSX · 0.2partition · 0.1TPUT · 0.1PARTITION · 0.1
YearPublicationVenuePosition
2022 Cost Modelling for Optimal Data Placement in Heterogeneous Main Memory
abstract
The cost of DRAM contributes significantly to the operating costs of in-memory database management systems (IMDBMS). Persistent memory (PMEM) is an alternative type of byte-addressable memory that offers --- in addition to persistence --- higher capacities than DRAM at a lower price with the disadvantage of increased latencies and reduced bandwidth. This paper evaluates PMEM as a cheaper alternative to DRAM for storing table base data, which can make up a significant fraction of an IMDBMS' total memory footprint. Using a prototype implementation in the SAP HANA IMDBMS, we find that placing all table data in PMEM can reduce query performance in analytical benchmarks by more than a factor of two, while transactional workloads are less affected. To quantify the performance impact of placing individual data structures in PMEM, we propose a cost model based on a lightweight workload characterization. Using this model, we show how to place data pareto-optimally in the heterogeneous memory. Our evaluation demonstrates the accuracy of the model and shows that it is possible to place more than 75% of table data in PMEM while keeping performance within 10% of the DRAM baseline for two analytical benchmarks.
Robert Lasch, Thomas Legler, Norman May, Bernhard Scheirle, Kai-Uwe Sattler
Proc. VLDB Endow.2
2021 Workload-Driven Placement of Column-Store Data Structures on DRAM and NVM
abstract
Non-volatile memory (NVM) offers lower costs per capacity and higher total capacities than DRAM. However, NVM cannot simply be used as a drop-in replacement for DRAM in database management systems due to its different performance characteristics. We thus investigate the placement of column-store data structures in a hybrid hierarchy of DRAM and NVM, with the goal of placing as much data as possible in NVM without compromising performance. After analyzing how different memory access patterns affect query runtimes when columns are placed in NVM, we propose a heuristic that leverages lightweight access counters to suggest which structures should be placed in DRAM and which in NVM. Our evaluation using TPC-H shows that more than 80% of the data touched by queries can be placed in NVM with almost no slowdown, while naively placing all data in NVM would increase runtime by 53%.
Robert Lasch, Robert Schulze, Thomas Legler, Kai-Uwe Sattler
DaMoN3
2019 Bridging the Latency Gap between NVM and DRAM for Latency-bound Operations
abstract
Non-Volatile Memory (NVM) technologies exhibit 4X the read access latency of conventional DRAM. When the working set does not fit in the processor cache, this latency gap between DRAM and NVM leads to more than 2X runtime increase for queries dominated by latency-bound operations such as index joins and tuple reconstruction. We explain how to easily hide NVM latency by interleaving the execution of parallel work in index joins and tuple reconstruction using coroutines. Our evaluation shows that interleaving applied to the non-trivial implementations of these two operations in a production-grade codebase accelerates end-to-end query runtimes on both NVM and DRAM by up to 1.7X and 2.6X respectively, thereby reducing the performance difference between DRAM and NVM by more than 60%.
Georgios Psaropoulos, Ismail Oukid, Thomas Legler, Norman May, Anastasia Ailamaki
DaMoN3
2019 Interleaving with coroutines: a systematic and practical approach to hide memory latency in index joins
Georgios Psaropoulos, Thomas Legler, Norman May, Anastasia Ailamaki
VLDB J.2
2017 Interleaving with Coroutines: A Practical Approach for Robust Index Joins
abstract
Index join performance is determined by the efficiency of the lookup operation on the involved index. Although database indexes are highly optimized to leverage processor caches, main memory accesses inevitably increase lookup runtime when the index outsizes the last-level cache; hence, index join performance drops. Still, robust index join performance becomes possible with instruction stream interleaving : given a group of lookups, we can hide cache misses in one lookup with instructions from other lookups by switching among their respective instruction streams upon a cache miss. In this paper, we propose interleaving with coroutines for any type of index join. We showcase our proposal on SAP HANA by implementing binary search and CSB + -tree traversal for an instance of index join related to dictionary compression. Coroutine implementations not only perform similarly to prior interleaving techniques, but also resemble the original code closely, while supporting both interleaved and non-interleaved execution. Thus, we claim that coroutines make interleaving practical for use in real DBMS codebases.
Georgios Psaropoulos, Thomas Legler, Norman May, Anastasia Ailamaki
Proc. VLDB Endow.2
2014 Improving in-memory database index performance with Intel® Transactional Synchronization Extensions
abstract
The increasing number of cores every generation poses challenges for high-performance in-memory database systems. While these systems use sophisticated high-level algorithms to partition a query or run multiple queries in parallel, they also utilize low-level synchronization mechanisms to synchronize access to internal database data structures. Developers often spend significant development and verification effort to improve concurrency in the presence of such synchronization. The Intel®Transactional Synchronization Extensions (Intel®TSX) in the 4th Generation Core™ Processors enable hardware to dynamically determine whether threads actually need to synchronize even in the presence of conservatively used synchronization. This paper evaluates the effectiveness of such hardware support in a commercial database. We focus on two index implementations: a B+Tree Index and the Delta Storage Index used in the SAP HANA®database system. We demonstrate that such support can improve performance of database data structures such as index trees and presents a compelling opportunity for the development of simpler, scalable, and easy-to-verify algorithms.
Tomas Karnagel, Roman Dementiev, Ravi Rajwar, Konrad Lai, Thomas Legler, Benjamin Schlegel, Wolfgang Lehner
HPCA5
2009 Robust Distributed Top-N Frequent Pattern Mining Using the SAP BW Accelerator
abstract
Mining for association rules and frequent patterns is a central activity in data mining. However, most existing algorithms are only moderately suitable for real-world scenarios. Most strategies use parameters like minimum support, for which it can be very difficult to define a suitable value for unknown datasets. Since most untrained users are unable or unwilling to set such technical parameters, we address the problem of replacing the minimum-support parameter with top- n strategies. In our paper, we start by extending a top- n implementation of the ECLAT algorithm to improve its performance by using heuristic search strategy optimizations. Also, real-world datasets are often distributed and modern database architectures are switching from expensive SMPs to cheaper shared-nothing blade servers. Thus, most mining queries require distribution handling. Since partitioning can be forced by user-defined semantics, it is often forbidden to transform the data. Therefore, we developed an adaptive top- n frequent-pattern mining algorithm that simplifies the mining process on real distributions by relaxing some requirements on the results. We first combine the PARTITION and the TPUT algorithms to handle distributed top- n frequent-pattern mining. Then, we extend this new algorithm for distributions with real-world data characteristics. For frequent-pattern mining algorithms, equal distributions are important conditions, and tiny partitions can cause performance bottlenecks. Hence, we implemented an approach called MAST that defines a minimum absolute-support threshold. MAST prunes patterns with low chances of reaching the global top- n result set and high computing costs. In total, our approach simplifies the process of frequent-pattern mining for real customer scenarios and data sets. This may make frequent-pattern mining accessible for very new user groups. Finally, we present results of our algorithms when run on the SAP NetWeaver BW Acceleratorwith standard and real business datasets.
Thomas Legler, Wolfgang Lehner, Jan Schaffner, Jens Krüger 0003
Proc. VLDB Endow.1
2006 Data Mining with the SAP Netweaver BI Accelerator
Thomas Legler, Wolfgang Lehner, Andrew Ross
VLDB1
2005 Real-Time Scheduling for Data Stream Management Systems
abstract
Quality-aware management of data streams is gaining more and more importance with the amount of data produced by streams growing continuously. The resources required for data stream processing depend on different factors and are limited by the environment of the data stream management system (DSMS). Thus, with a potentially unbounded amount of stream data and limited processing resources, some of the data stream processing tasks (originating from different users) may not be satisfyingly answered, and therefore, users should be enabled to negotiate a certain quality for the execution of their stream processing tasks. After the negotiation process, it is the responsibility of the Data Stream Management System to meet the quality constraints by using adequate resource reservation and scheduling techniques. Within this paper, we consider different aspects of real-time scheduling for operations within a DSMS. We propose a scheduling concept which enables us to meet certain time-dependent quality of service requirements for user-given processing tasks. Furthermore, we describe the implementation of our scheduling concept within a real-time capable data stream management system, and we give experimental results on that.
Sven Schmidt, Thomas Legler, Daniel Schaller, Wolfgang Lehner
ECRTS2
2005 Robust Real-time Query Processing with QStream
Sven Schmidt, Thomas Legler, Sebastian Schär, Wolfgang Lehner
VLDB2