Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Tomas Karnagel

dblp:127/0442 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 5 first-authorSystems, architecture and hardware · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Query processing and optimization · 48% Indexing and storage engines · 24% Transaction processing and concurrency control · 24%
Artificial intelligence
1 paper
Efficient and distributed learning · 77% Learning theory · 23%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 69% GPUs and heterogeneous computing · 31%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › parallel query processing
heterogeneous hardware query processing
0.522017
Adaptive Work Placement for Query Processing on Heterogeneous Computing Resources · Proc. VLDB Endow. 2017
Demonstrating efficient query processing in heterogeneous environments · SIGMOD Conference 2014
Machine learning › Efficient and distributed learning
automated machine learning
0.412020
Oracle AutoML: A Fast and Predictive AutoML Pipeline · Proc. VLDB Endow. 2020
Indexing and storage engines
b+-tree
0.212014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Transaction processing and concurrency control › transactional memory
hardware transactional memory
0.212014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Indexing and storage engines
in-memory index
0.212014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Query processing and optimization
query execution
0.212014
Demonstrating efficient query processing in heterogeneous environments · SIGMOD Conference 2014
Transaction processing and concurrency control
synchronization
0.212014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Distributed systems › distributed scheduling
operator placement
0.212014
Demonstrating efficient query processing in heterogeneous environments · SIGMOD Conference 2014
Machine learning › Learning theory
model selection
0.112020
Oracle AutoML: A Fast and Predictive AutoML Pipeline · Proc. VLDB Endow. 2020
Query processing and optimization
cardinality estimation
0.112017
Adaptive Work Placement for Query Processing on Heterogeneous Computing Resources · Proc. VLDB Endow. 2017
GPUs and heterogeneous computing
heterogeneous resources
0.112017
Adaptive Work Placement for Query Processing on Heterogeneous Computing Resources · Proc. VLDB Endow. 2017
Database system architecture and tuning
main-memory database
0.112014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014

Methods — techniques the papers use, named apart from their topics

virtualization layer · 0.6adaptive placement · 0.6proxy model prediction · 0.4meta-learning · 0.4hardware transactional memory · 0.2Intel TSX · 0.2
YearPublicationVenuePosition
2020 Oracle AutoML: A Fast and Predictive AutoML Pipeline
abstract
Machine learning (ML) is at the forefront of the rising popularity of data-driven software applications. The resulting rapid proliferation of ML technology, explosive data growth, and shortage of data science expertise have caused the industry to face increasingly challenging demands to keep up with fast-paced develop-and-deploy model lifecycles. Recent academic and industrial research efforts have started to address this problem through automated machine learning (AutoML) pipelines and have focused on model performance as the first-order design objective. We present Oracle AutoML, a novel iteration-free AutoML pipeline designed to not only provide accurate models, but also in a shorter runtime. We are able to achieve these objectives by eliminating the need to continuously iterate over various pipeline configurations. In our feed-forward approach, each pipeline stage makes decisions based on metalearned proxy models that can predict candidate pipeline configuration performances before building the full final model. Our approach, which builds and tunes only the best candidate pipeline, achieves better scores at a fraction of the time compared to state-of-the-art open source AutoML tools, such as H2O and Auto-sklearn. This makes Oracle AutoML a prime candidate for addressing current industry challenges.
Anatoly Yakovlev, Hesam Fathi Moghadam, Ali Moharrer, Jingxiao Cai, Nikan Chavoshi, Venkatanathan Varadarajan, Sandeep R. Agrawal, Tomas Karnagel, Sam Idicula, Sanjay Jinturkar, Nipun Agarwal
Proc. VLDB Endow.8
2017 Big data causing big (TLB) problems: taming random memory accesses on the GPU
abstract
GPUs are increasingly adopted for large-scale database processing, where data accesses represent the major part of the computation. If the data accesses are irregular, like hash table accesses or random sampling, the GPU performance can suffer. Especially when scaling such accesses beyond 2GB of data, a performance decrease of an order of magnitude is encountered. This paper analyzes the source of the slowdown through extensive micro-benchmarking, attributing the root cause to the Translation Lookaside Buffer (TLB). Using the micro-benchmarks, the TLB hierarchy and structure are fully analyzed on two different GPU architectures, identifying never-before-published TLB sizes that can be used for efficient large-scale application tuning. Based on the gained knowledge, we propose a TLB-conscious approach to mitigate the slowdown for algorithms with irregular memory access. The proposed approach is applied to two fundamental database operations - random sampling and hash-based grouping - showing that the slowdown can be dramatically reduced, and resulting in a performance increase of up to 13×.
Tomas Karnagel, Tal Ben-Nun, Matthias Werner 0004, Dirk Habich, Wolfgang Lehner
DaMoN1
2017 Adaptive Work Placement for Query Processing on Heterogeneous Computing Resources
abstract
The hardware landscape is currently changing from homogeneous multi-core systems towards heterogeneous systems with many different computing units, each with their own characteristics. This trend is a great opportunity for data-base systems to increase the overall performance if the heterogeneous resources can be utilized efficiently. To achieve this, the main challenge is to place the right work on the right computing unit. Current approaches tackling this placement for query processing assume that data cardinalities of intermediate results can be correctly estimated. However, this assumption does not hold for complex queries. To overcome this problem, we propose an adaptive placement approach being independent of cardinality estimation of intermediate results. Our approach is incorporated in a novel adaptive placement sequence. Additionally, we implement our approach as an extensible virtualization layer, to demonstrate the broad applicability with multiple database systems. In our evaluation, we clearly show that our approach significantly improves OLAP query processing on heterogeneous hardware, while being adaptive enough to react to changing cardinalities of intermediate query results.
Tomas Karnagel, Dirk Habich, Wolfgang Lehner
Proc. VLDB Endow.1
2016 Limitations of Intra-operator Parallelism Using Heterogeneous Computing Resources
Tomas Karnagel, Dirk Habich, Wolfgang Lehner
ADBIS1
2016 HW/SW-database-codesign for compressed bitmap index processing
abstract
Compressed bitmap indices are heavily used in scientific and commercial database systems because they largely improve query performance for various workloads. Early research focused on finding tailor-made index compression schemes that are amenable for modern processors. Improving performance further typically comes at the expense of a lower compression rate, which is in many applications not acceptable because of memory limitations. Alternatively, tailor-made hardware allows to achieve a performance that can only hardly be reached with software running on general-purpose CPUs. In this paper, we will show how to create a custom instruction set framework for compressed bitmap processing that is generic enough to implement most of the major compressed bitmap indices. For evaluation, we implemented WAH, PLWAH, and COMPAX operations using our framework and compared the resulting implementation to multiple state-of-the-art processors. We show that the custom-made bitmap processor achieves speedups of up to one order of magnitude by also using two orders of magnitude less energy compared to a modern energy-efficient Intel processor. Finally, we discuss how to embed our processor with database-specific instruction sets into database system environments.
Sebastian Haas, Tomas Karnagel, Oliver Arnold, Erik Laux, Benjamin Schlegel, Gerhard P. Fettweis, Wolfgang Lehner
ASAP2
2014 Improving in-memory database index performance with Intel® Transactional Synchronization Extensions
abstract
The increasing number of cores every generation poses challenges for high-performance in-memory database systems. While these systems use sophisticated high-level algorithms to partition a query or run multiple queries in parallel, they also utilize low-level synchronization mechanisms to synchronize access to internal database data structures. Developers often spend significant development and verification effort to improve concurrency in the presence of such synchronization. The Intel®Transactional Synchronization Extensions (Intel®TSX) in the 4th Generation Core™ Processors enable hardware to dynamically determine whether threads actually need to synchronize even in the presence of conservatively used synchronization. This paper evaluates the effectiveness of such hardware support in a commercial database. We focus on two index implementations: a B+Tree Index and the Delta Storage Index used in the SAP HANA®database system. We demonstrate that such support can improve performance of database data structures such as index trees and presents a compelling opportunity for the development of simpler, scalable, and easy-to-verify algorithms.
Tomas Karnagel, Roman Dementiev, Ravi Rajwar, Konrad Lai, Thomas Legler, Benjamin Schlegel, Wolfgang Lehner
HPCA1
2014 Demonstrating efficient query processing in heterogeneous environments
abstract
The increasing heterogeneity in hardware systems gives developers many opportunities to add more functionality and computational power to the system. As a consequence, modern database systems will need to be able to adapt to a wide variety of heterogeneous architectures. While porting single operators to accelerator architectures is well-understood, a more generic approach is needed for the whole database system. In prior work, we presented a generic hardware-oblivious database system, where the operators can be executed on the main processor as well as on a large number of accelerator architectures. However, to achieve fully heterogeneous query processing, placement decisions are needed for the database operators. We enhance the presented system with heterogeneity-aware operator placement (HOP) to take a major step towards designing a database system that can efficiently exploit highly heterogeneous hardware environments. In this demonstration, we are focusing on the placement-integration aspect as well as presenting the resulting database system.
Tomas Karnagel, Matthias Hille, Mario Ludwig, Dirk Habich, Wolfgang Lehner, Max Heimel, Volker Markl
SIGMOD Conference1
2013 The HELLS-join: a heterogeneous stream join for extremely large windows
abstract
Upcoming processors are combining different computing units in a tightly-coupled approach using a unified shared memory hierarchy. This tightly-coupled combination leads to novel properties with regard to cooperation and interaction. This paper demonstrates the advantages of those processors for a stream-join operator as an important data-intensive example. In detail, we propose our HELLS-Join approach employing all heterogeneous devices by outsourcing parts of the algorithm on the appropriate device. Our HELLS-Join performs better than CPU stream joins, allowing wider time windows, higher stream frequencies, and more streams to be joined as before.
Tomas Karnagel, Dirk Habich, Benjamin Schlegel, Wolfgang Lehner
DaMoN1
2013 Scalable frequent itemset mining on many-core processors
abstract
Frequent-itemset mining is an essential part of the association rule mining process, which has many application areas. It is a computation and memory intensive task with many opportunities for optimization. Many efficient sequential and parallel algorithms were proposed in the recent years. Most of the parallel algorithms, however, cannot cope with the huge number of threads that are provided by large multiprocessor or many-core systems. In this paper, we provide a highly parallel version of the well-known Eclat algorithm. It runs on both, multiprocessor systems and many-core coprocessors, and scales well up to a very large number of threads---244 in our experiments. To evaluate mcEclat's performance, we conducted many experiments on realistic datasets. mcEclat achieves high speedups of up to 11.5x and 100x on a 12-core multiprocessor system and a 61-core Xeon Phi many-core coprocessor, respectively. Furthermore, mcEclat is competitive with highly optimized existing frequent-itemset mining implementations taken from the FIMI repository.
Benjamin Schlegel, Tomas Karnagel, Tim Kiefer, Wolfgang Lehner
DaMoN2