Mark Raasveldt

dblp:182/7109 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
3since 2021 · last 2025
0000-0001-5005-6844ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 7 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Runtime-Extensible Parsers
Hannes Mühleisen, Mark Raasveldt
CIDR2
2024 How to Make your Duck Fly: Advanced Floating Point Compression to the Rescue
Panagiotis Liakos, Katia Papakonstantinopoulou, Thijs Bruineman, Mark Raasveldt, Yannis Kotidis
EDBT4
2022 DuckDB-Wasm: Fast Analytical Processing for the Web
abstract
We introduce DuckDB-Wasm, a WebAssembly version of the database system DuckDB, to provide fast analytical processing for the Web. DuckDB-Wasm evaluates SQL queries asynchronously in web workers, supports efficient user-defined functions written in JavaScript, and features a browser-agnostic filesystem that reads local and remote data in pages. DuckDB-Wasm outperforms previous data processing libraries for the Web in the TPC-H benchmark at multiple scale factors. We demonstrate the capabilities of an analytical database in the browser using an interactive SQL shell.
André Kohn 0001, Dominik Moritz, Mark Raasveldt, Hannes Mühleisen, Thomas Neumann 0001
Proc. VLDB Endow.3
2020 Data Management for Data Science - Towards Embedded Analytics
Mark Raasveldt, Hannes Mühleisen
CIDR1
2019 devUDF: Increasing UDF development efficiency through IDE Integration. It works like a PyCharm!
abstract
User-defined functions (UDFs) facilitate the execution of analytics pipelines inside the database. They provide many advantages over traditional methods, such as close-to-data execution and automatic parallelization. However, the standard workflow for developing and debugging UDFs does not allow developers to use their regular toolchains and Integrated Development Environments (IDEs). As a result, writing functional UDFs is challenging. In this demo, we present the devUDF, a plugin to the PyCharm IDE that allows developers to develop and debug their MonetDB/Python UDFs directly from within the IDE.
Mark Raasveldt, Pedro Holanda, Stefan Manegold
EDBT1
2019 DuckDB: an Embeddable Analytical Database
abstract
The immense popularity of SQLite shows that there is a need for unobtrusive in-process data management solutions. However, there is no such system yet geared towards analytical workloads. We demonstrate DuckDB, a novel data management system designed to execute analytical SQL queries while embedded in another process. In our demonstration, we pit DuckDB against other data management solutions to showcase its performance in the embedded analytics scenario. DuckDB is available as Open Source software under a permissive license.
Mark Raasveldt, Hannes Mühleisen
SIGMOD Conference1
2019 Progressive Indexes: Indexing for Interactive Data Analysis
abstract
Interactive exploration of large volumes of data is increasingly common, as data scientists attempt to extract interesting information from large opaque data sets. This scenario presents a difficult challenge for traditional database systems, as (1) nothing is known about the query workload in advance, (2) the query workload is constantly changing, and (3) the system must provide interactive responses to the issued queries. This environment is challenging for index creation, as traditional database indexes require upfront creation, hence a priori workload knowledge, to be efficient. In this paper, we introduce Progressive Indexing , a novel performance-driven indexing technique that focuses on automatic index creation while providing interactive response times to incoming queries. Its design allows queries to have a limited budget to spend on index creation. The indexing budget is automatically tuned to each query before query processing. This allows for systems to provide interactive answers to queries during index creation while being robust against various workload patterns and data distributions.
Pedro Holanda, Stefan Manegold, Hannes Mühleisen, Mark Raasveldt
Proc. VLDB Endow.4
2018 Deep Integration of Machine Learning Into Column Stores
abstract
We leverage vectorized User-Defined Functions (UDFs) to efficiently integrate unchanged machine learning pipelines into an analytical data management system.The entire pipelines including data, models, parameters and evaluation outcomes are stored and executed inside the database system.Experiments using our MonetDB/Python UDFs show greatly improved performance due to reduced data movement and parallel processing opportunities.In addition, this integration enables meta-analysis of models using relational queries.
Mark Raasveldt, Pedro Holanda, Hannes Mühleisen, Stefan Manegold
EDBT1
2018 MonetDBLite: An Embedded Analytical Database
abstract
No abstract available.
Mark Raasveldt
SIGMOD Conference1
2017 Don't Hold My Data Hostage - A Case For Client Protocol Redesign
abstract
Transferring a large amount of data from a database to a client program is a surprisingly expensive operation. The time this requires can easily dominate the query execution time for large result sets. This represents a significant hurdle for external data analysis, for example when using statistical software. In this paper, we explore and analyse the result set serialization design space. We present experimental results from a large chunk of the database market and show the inefficiencies of current approaches. We then propose a columnar serialization method that improves transmission performance by an order of magnitude.
Mark Raasveldt, Hannes Mühleisen
Proc. VLDB Endow.1
2016 Vectorized UDFs in Column-Stores
abstract
Data Scientists rely on vector-based scripting languages such as R, Python and MATLAB to perform ad-hoc data analysis on potentially large data sets. When facing large data sets, they are only efficient when data is processed using vectorized or bulk operations. At the same time, overwhelming volume and variety of data as well as parsing overhead suggests that the use of specialized analytical data management systems would be beneficial. Data might also already be stored in a database. Efficient execution of data analysis programs such as data mining directly inside a database greatly improves analysis efficiency.
Mark Raasveldt, Hannes Mühleisen
SSDBM1