Charlotte Felius

dblp:399/4329 · also Lotte Felius · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0000-4271-2055ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 G-ALP: Rethinking Light-weight Encodings for GPUs
abstract
This paper introduces G-ALP, a GPU-optimized version of ALP, which is a recent and state-of-the-art compression scheme for floating-point.This GPU-optimization is based on two core ideas.First, all parts of the decoding process must be fully data-parallelized.In this paper, we fully data-parallelize exception patching, which typically applies to only 1% of the data.While patching has negligible performance cost on CPUs, it can become the main bottleneck on GPUs if it is not data-parallel.Second, the decoding API must minimize its register footprint, a highly scarce resource on GPUs, and hence deliver just one value-at-a-time.Our unique aim is to integrate G-ALP decoding into GPU kernels that consume data from global memory, rather than let decompression be a separate kernel.We consider these two ideas general guidelines for future GPU-optimized lightweight encodings, and a significant evolution of our new FastLanes file format, making it GPU-friendly.We extensively test G-ALP in a series of microbenchmarks and evaluate its performance on an NVIDIA V100 GPU and an NVIDIA RTX4070 Super Ti GPU, demonstrating superior performance compared to NVIDIA nvCOMP and ndzip in both decoding and filtering queries.
Sven Hepkema, Azim Afroozeh, Charlotte Felius, Peter Boncz, Stefan Manegold
DaMoN3
2025 VCrypt: Leveraging Vectorized and Compressed Execution for Client-side Encryption
Charlotte Felius, Peter Boncz
EDBT1
2024 Accelerating GPU Data Processing using FastLanes Compression
abstract
We show that compression can be a win-win for GPU data processing: it not only allows to store more data in GPU global memory, but can also accelerate data processing. We show that the complete redesign of compressed columnar storage in FastLanes, with its fully data-parallel bit-packing and encodings, also benefits GPU hardware. We micro-benchmark the performance of FastLanes on two GPU architectures (Nvidia T4 and V100) and integrate FastLanes in the Crystal GPU query processing prototype. Our experiments show that FastLanes decompression significantly outperforms previous decompression methods in micro-benchmarks, and can make end-to-end SSB queries up to twice faster compared to uncompressed query processing - in contrast to previous work where GPU decompression caused execution to slow down. We further discovered that an access granularity of decoding vectors of 1024 values is too large for a single GPU warp due to register pressure. We mitigate this here using mini-vectors - a future work question is how to further reduce this granularity with minimal impact on efficiency.
Azim Afroozeh, Charlotte Felius, Peter Boncz
DaMoN2
2024 DuckDB-SGX2: The Good, The Bad and The Ugly within Confidential Analytical Query Processing
abstract
We provide an evaluation of an analytical workload in a confidential computing environment, combining DuckDB with two technologies: modular columnar encryption in Parquet files (data at rest) and the newest version of the Intel SGX Trusted Execution Environment (TEE), providing a hardware enclave where data in flight can be (more) securely decrypted and processed. One finding is that the "performance tax" for such confidential analytical processing is acceptable compared to not using these technologies. We eventually manage to run TPC-H SF30 with under 2x overhead compared to non-encrypted, non-enclave execution; we show that, specifically, columnar compression and encryption are a good combination. Our second finding consists of dos and don'ts to tune DuckDB to work effectively in this environment. There are various performance hazards: potentially 5x higher cache miss costs due to memory encryption inside the enclave, NUMA penalties, and highly elevated cost of swapping pages in and out of the enclave - which is also triggered indirectly by using a non-SGX-aware malloc library.
Ilaria Battiston, Charlotte Felius, Sam Ansmink, Laurens Kuiper, Peter Boncz
DaMoN2