Nico Schäfer

dblp:261/1970 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 5 first-author · 3 since 2021
YearPublicationVenuePosition
2023 To UDFs and Beyond: Demonstration of a Fully Decomposed Data Processor for General Data Wrangling Tasks
abstract
While existing data management solutions try to keep up with novel data formats and features, a myriad of valuable functionality is often only accessible via programming language libraries. Particularly for machine learning tasks, there is a wealth of pre-trained models and easy-to-use libraries that allow a wide audience to harness state-of-the-art machine learning. We propose the demonstration of a highly modularized data processor for semi-structured data that can be extended by means of plain Python scripts. Next to commonly supported user-defined functions, the deep decomposition allows augmenting the core engine with additional index structures, customized import and export routines, and custom aggregation functions. For several use cases, we detail how user-defined modules can be quickly realized and invite the audience to write and apply custom code, to tailor provided code snippets that we bring along to own preferences to solve data analytics tasks involving sentiment analysis of Twitter tweets.
Nico Schäfer, Damjan Gjurovski, Angjela Davitkova, Sebastian Michel 0001
Proc. VLDB Endow.1
2022 BETZE: Benchmarking Data Exploration Tools with (Almost) Zero Effort
abstract
In this paper, we propose BETZE, a benchmark generator to evaluate the performance of data exploration solutions for semi-structured data. It is tailored to the typical query capabilities of modern JSON document stores and can be extended to match more. At its core, the query generator mimics the behavior of a data scientist through a model similar to the random surfer idea known from PageRank. We propose preset parameters that pose different query loads to the system, intended to reflect novice, intermediate, and expert users interacting with the system. The proposed approach analyzes a given JSON dataset and generates queries into an intermediate representation that is then translated to system-specific query syntax. We have implemented support for MongoDB, PostgreSQL, jq, and our own JSON processor JODA, and describe how additional tools can be supported. To get started, we report on a first experimental study, showing the versatility of the benchmark generator, using the NoBench dataset, and real-world data obtained from Twitter and Reddit.
Nico Schäfer, Sebastian Michel 0001
ICDE1
2021 Utilizing Delta Trees for Efficient, Iterative Exploration and Transformation of Semi-Structured Contents
abstract
The keywords data exploration or data wrangling summarize various different query workload scenarios in which users aim to explore or tailor data to their needs. For semi-structured data, next to commonly used SQL-style select-from-where and aggregation queries, also the structure of the possibly-nested schema-free data can be altered, schema attributes renamed, and so on. This typically involves various rounds of refining or discarding queries-imposing that intermediate results as well as the original sources cannot be eliminated. In this work, we extend our prior work on JODA, a vertically scalable, versatile JSON data processor, to make use of so-called delta trees for the succinct representation of incrementally created query results.
Nico Schäfer, Sebastian Michel 0001
ICDE1
2020 Partially Materializable Delta Trees for Efficient Data Wrangling of Semi-Structured Contents
Nico Schäfer, Sebastian Michel 0001
EDBT1
2020 JODA: A Vertically Scalable, Lightweight JSON Processor for Big Data Transformations
abstract
We describe the demonstration of JODA (Json On Demand Analytics), an approach to handling large amounts of JSON documents in a vertically scalable manner. With JODA, the user can import, filter, transform, aggregate, group, and export documents with a simple PIG-style query language, offering fast execution speed. This is achieved by utilizing a multithreaded architecture over disjoint, read-only containers of data that are processed in parallel, similar to what RDDs are to Spark. Containers are augmented with auxiliary information like Bloom filters and adaptive indices and all containers are processed in parallel by individual threads. By avoiding locks, latches, and synchronization beyond simple thread pooling, we do not risk contention and therefore maximize resource utilization. The demonstration scenarios aim at engaging visitors with several data analytics tasks around large, real-world datasets that are to be solved with the help of JODA, and further gives insights on system internals and the installation/configuration process.
Nico Schäfer, Sebastian Michel 0001
ICDE1