David Sauerwein

dblp:327/8089 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2023
0009-0006-9191-8395ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%
Databases, data mining, and information retrieval
1 paper
Distributed and cloud data management · 50% Indexing and storage engines · 50%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Indexing and storage engines
columnar storage
0.712023
BtrBlocks: Efficient Columnar Compression for Data Lakes · Proc. ACM Manag. Data 2023
Distributed and cloud data management
data lake
0.712023
BtrBlocks: Efficient Columnar Compression for Data Lakes · Proc. ACM Manag. Data 2023
Storage systems › data management › database storage
columnar storage
0.712023
BtrBlocks: Efficient Columnar Compression for Data Lakes · Proc. ACM Manag. Data 2023
Storage systems
data compression
0.712023
BtrBlocks: Efficient Columnar Compression for Data Lakes · Proc. ACM Manag. Data 2023

Methods — techniques the papers use, named apart from their topics

lightweight encoding schemes · 1.3
YearPublicationVenuePosition
2023 BtrBlocks: Efficient Columnar Compression for Data Lakes
abstract
Analytics is moving to the cloud and data is moving into data lakes. These reside on object storage services like S3 and enable seamless data sharing and system interoperability. To support this, many systems build on open storage formats like Apache Parquet. However, these formats are not optimized for remotely-accessed data lakes and today's high-throughput networks. Inefficient decompression makes scans CPU-bound and thus increases query time and cost. With this work we present BtrBlocks, an open columnar storage format designed for data lakes. BtrBlocks uses a set of lightweight encoding schemes, achieving fast and efficient decompression and high compression ratios.
Maximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor Leis
Proc. ACM Manag. Data2