Simon Andreas Frimann Lund

dblp:120/7437 · also Simon A. F. Lund · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0002-1942-3818ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Flexible I/O for Database Management Systems with xNVMe
Emil Houlborg, Andreas Nicolaj Tietgen, Simon Andreas Frimann Lund, Marcel Weisgut, Tilmann Rabl, Javier González 0006, Vivek Shah 0001, Pinar Tözün
CIDR3
2025 Path to GPU-Initiated I/O for Data-Intensive Systems
abstract
The process of training and serving deep learning (DL) models is computationally expensive, mandating the use of powerful and expensive accelerators such as GPUs and TPUs.Furthermore, the prevalence of GPUs in data centers today motivate developing database systems that can leverage the available GPU resources.Both the latency of DL tasks and database queries and high utilization of these accelerators depend on how efficiently we can move the data to the accelerators.Given today's dataset sizes, fitting everything in GPU or even CPU memory is not always feasible or can be expensive.The I/O path while fetching the data from disks, however, still dominantly relies on CPUs.In this work, we take a step toward understanding today's landscape for optimizing the I/O path for reading data to GPUs from disks, with a focus on SSDs.First, we review the prominent technologies that target GPU-centric storage accesses.Then, we dive deeper into BaM [38], as the state-of-the-art method for GPU-centric storage, and evaluate its performance in comparison to the state-of-theart CPU-centric storage interface SPDK.Our results demonstrate that while BaM is able to match the performance of SPDK without involving CPUs on the I/O path, this comes at the cost of a very high GPU use.Finally, we highlight future research directions to enable an I/O path that is both efficient and easy-to-adopt for data-intensive systems that use GPUs.
Karl B. Torp, Simon Andreas Frimann Lund, Pinar Tözün
DaMoN2
2024 I/O Passthru: Upstreaming a flexible and efficient I/O Path in Linux
Kanchan Joshi, Javier González 0006, Krishna Kanth Reddy, Arun George, Simon Andreas Frimann Lund, Jens Axboe
FAST7
2022 I/O interface independence with xNVMe
abstract
The tight coupling of data-intensive systems and I/O interface has been a problem for years. A database system, relying on an specific I/O backend for direct asynchronous I/Os such as libaio, inherits its limitations in terms of portability, expressiveness and performance. The emergence of high-performance NVMe Solid-State Drives (SSDs), enabling new command sets, compounds this problem. Indeed, efforts to streamline the I/O stack have led to the introduction of new, complex and idiosyncratic I/O interfaces such as SPDK, io_uring or asynchronous ioctls. What is the appropriate I/O interface for a given system? How can applications effectively leverage SSD and end-to-end I/O interface innovations? Is I/O interface lock-in a necessary evil for data-intensive systems and storage services? Our answer to the latter question is no. Our answer to the former questions is xNVMe, a cross-platform user-space library that provides I/O-interface independence to user-space software. In this paper, we present the xNVMe API, we detail its design and we show that xNVMe has a negligible cost atop the most efficient I/O interfaces on Linux, FreeBSD and Windows.
Simon Andreas Frimann Lund, Philippe Bonnet, Klaus B. A. Jensen, Javier González 0006
SYSTOR1
2016 Fusion of Parallel Array Operations
abstract
We address the problem of fusing array operations based on criteria such as shape compatibility, data reuse, and minimizing for data reuse, the fusion problem has been formulated as a static weighted graph partitioning problem (known as the Weighted Loop Fusion problem). We show that this scheme cannot accurately track data reuse between multiple independent loops, since it overestimates total data reuse of certain cases. Our formulation in terms of partitions allows use of realistic cost functions that can track resource usage accurately. We give correctness proofs, and prove that WSP can maximize data reuse in programs exactly, in contrast to prior work.
Mads Ruben Burgdorff Kristensen, Simon Andreas Frimann Lund, Troels Blum, James Avery
PACT2