VLDB 2026 Research / reviewers in the wild / expert
Florian Stock
dblp:92/1248
· DBLP profile ↗
14ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0001-9411-0267ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Update NDP: On Offloading Modifications to Smart Storage with Transactional Guarantees in Near-Data Processing DBMSabstractThe performance and scalability of modern data-intensive systems processing large datasets are limited by unnecessary data movement. Even though near-data processing (NDP) can provably reduce data transfers and increase performance, at present, NDP is utilized primarily in read-only settings. Near-data execution of data-intensive modification operations is currently infeasible due to the lack of transactional consistency and the absence of practicable low-latency synchronization mechanisms between the host database engine and the NDP-engine on smart storage. In this article, we introduce update NDP as an approach to offloading modifications to computational storage with transactional guarantees in an NDP database system called neoDBMS . To ensure consistency, we introduce a low-latency shared lock table between the host and computational storage, based on novel cache-coherent interconnects . We also introduce a novel locking protocol that seamlessly integrates the shared lock table within the lock manager of the host NDP-engine. To handle failure recovery, while preserving high and robust performance, we introduce novel extended locking and logging mechanisms that allow the host and computational storage to perform useful work during log-movement. Our evaluation indicates that in-storage modifications in neoDBMS in mixed workload settings are ≥ 6.52× faster than host-only executions and exhibit robust performance due to lower data movement and better resource utilization. Arthur Bernhardt, Sajjad Tamimi, Florian Stock, Andreas Koch 0001, Ilia Petrov 0001 |
ACM Trans. Database Syst. | 3 |
| 2025 | PUL: Pre-load in Software for Caches Wouldn't Always Play Along
Arthur Bernhardt, Sajjad Tamimi, Florian Stock, Andreas Koch 0001, Ilia Petrov 0001 |
ADBIS | 3 |
| 2025 | CINDA: Using Cache-Coherent Interconnects for Accelerating Databases by Enabling Near-Data Processing of Update TransactionsabstractNear-Data Processing (NDP) has been proven useful to accelerate Database Management Systems (DBMS) that handle infrequently accessed data stored in slow persistent storage. A key challenge for such an architecture is the synchronization of host-based and NDP operations, which require fine-grained interactions especially when the NDP device can also update (modify) the DBMS data autonomously.This paper introduces CINDA, the first full-stack computational storage capable of acceleratingbothread and update (write) database transactions using NDP. The proposed system relies on a hybrid host-device interface to enable the DBMS accessing persisted data, offloading computation to the storage device, and coordinating concurrent device-update operations with the host-update ones. A hybrid interface utilizes a cache-coherent inter-connect such as CCIX or CXL for low-latency synchronization using a shared-lock table, and PCIe DMA for high-throughput bulk I/O. We evaluated the effectiveness of the proposed approach in a CCIX-based system by realizing an FPGA-based NDP-capable computational storage device and customizing an NDP-capable DBMS based on PostgreSQL to support update NDP operations. Our full-stack evaluation using the YCSB benchmark demonstrates that CINDA can deliver ≈4.2× end-to-end speedup when executing long-running update transactions directly on the storage device, while the host DBMS performs frequent short updates. Sajjad Tamimi, Arthur Bernhardt, Florian Stock, Ilia Petrov 0001, Andreas Koch 0001 |
IEEE Trans. Computers | 3 |
| 2025 | DANSEN: Database Acceleration on Native Computational Storage by Exploiting NDPabstractThis article introduces DANSEN , the hardware accelerator component for neoDBMS, a full-stack computational storage system designed to manage on-device execution of database queries/transactions as a Near-Data Processing (NDP)-operation. The proposed system enables Database Management Systems (DBMS) to offload NDP-operations to the storage while maintaining control over data through a native storage interface . DANSEN provides an NDP-engine that enables DBMS to perform both low-level database tasks, such as performing database administration, as well as high-level tasks like executing SQL, on the smart storage device while observing the DBMS concurrency control. Furthermore, DANSEN enables the incorporation of custom accelerators as an NDP-operation (e.g., to perform hardware-accelerated ML inference directly on the stored data). We built the DANSEN storage prototype and interface on an UltraScale+HBM FPGA and fully integrated it with PostgreSQL 12. Experimental results demonstrate that the proposed NDP approach outperforms software-only PostgreSQL using a fast off-the-shelf NVMe drive and significantly improves the end-to-end execution time of an aggregation operation (similar to Q6 from CH-benCHmark, 150 million records) by ≈ 10.6×. The versatility of the proposed approach is also validated by integrating a compute-intensive data analytics application with multi-row results, outperforming PostgreSQL by ≈ 1.5×. Sajjad Tamimi, Arthur Bernhardt, Florian Stock, Ilia Petrov 0001, Andreas Koch 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2023 | Game Physics Engine Using Optimised Geometric Algebra RISC-V Vector Extensions Code Using Fourier Series Data
Ed Saribatir, Niko Zurstraßen, Dietmar Hildenbrand, Florian Stock, Atilio Morillo Piña, Frederic von Wegner, Zheng Yan 0001, Shiping Wen 0001, Matthew Arnold |
CGI (4) | 4 |
| 2022 | Cache-Coherent Shared Locking for Transactionally Consistent Updates in Near-Data Processing DBMS on Smart Storage
Arthur Bernhardt, Sajjad Tamimi, Florian Stock, Tobias Vinçon, Andreas Koch 0001, Ilia Petrov 0001 |
EDBT | 3 |
| 2022 | An Evaluation of Using CCIX for Cache-Coherent Host-FPGA InterfacingabstractFor a long time, most discrete accelerators have been attached to host systems using various generations of the PCI Express interface. However, with its lack of support for coherency between accelerator and host caches, fine-grained interactions require frequent cache-flushes, or even the use of inefficient uncached memory regions. The Cache Coherent Interconnect for Accelerators (CCIX) was the first multi-vendor standard for enabling cache-coherent host-accelerator attachments, and already is indicative of the capabilities of upcoming standards such as Compute Express Link (CXL). In our work, we compare-and-contrast the use of CCIX with PCIe when interfacing an ARM-based host with two generations of CCIX-enabled FPGAs. We provide both low-level throughput and latency measurements for accesses and address translation, as well as examine an application-level use-case of using CCIX for fine-grained synchronization in an FPGA-accelerated database system. We can show that especially smaller reads from the FPGA to the host can benefit from CCIX by having roughly 33% shorter latency than PCIe. Small writes to the host have a latency roughly 32% higher than PCIe, though, since they carry a higher coherency overhead. For the database use-case, the use of CCIX allowed to maintain a constant synchronization latency even with heavy host-FPGA parallelism. Sajjad Tamimi, Florian Stock, Andreas Koch 0001, Arthur Bernhardt, Ilia Petrov 0001 |
FCCM | 2 |
| 2022 | neoDBMS: In-situ Snapshots for Multi-Version DBMS on Native Computational StorageabstractMulti-versioning and MVCC are the foundations of many modern DBMSs. Under mixed workloads and large datasets, the creation of the transactional snapshot can become very expensive, as long-running analytical transactions may request old versions, residing on cold storage, for reasons of transactional consistency. Furthermore, analytical queries operate on cold data, stored on slow persistent storage. Due to the poor data locality, snapshot creation may cause massive data transfers and thus lower performance. Given the current trend towards computational storage and near-data processing, it has become viable to perform such operations in-storage to reduce data transfers and improve scalability. neoDBMS is a DBMS designed for near-data processing and computational storage. In this paper, we demonstrate how neoDBMS performs snapshot computation in-situ. We showcase different interactive scenarios, where neoDBMS outperforms PostgreSQL 12 by up to 5×. Arthur Bernhardt, Sajjad Tamimi, Tobias Vinçon, Christian Knödler, Florian Stock, Carsten Heinz, Andreas Koch 0001, Ilia Petrov 0001 |
ICDE | 5 |
| 2022 | Near-Data Processing in Database Systems on Native Computational Storage under HTAP WorkloadsabstractToday's Hybrid Transactional and Analytical Processing (HTAP) systems, tackle the ever-growing data in combination with a mixture of transactional and analytical workloads. While optimizing for aspects such as data freshness and performance isolation, they build on the traditional data-to-code principle and may trigger massive cold data transfers that impair the overall performance and scalability. Firstly, in this paper we show that Near-Data Processing (NDP) naturally fits in the HTAP design space. Secondly, we propose an NDP database architecture, allowing transactionally consistent in-situ executions of analytical operations in HTAP settings. We evaluate the proposed architecture in state-of-the-art key/value-stores and multi-versioned DBMS. In contrast to traditional setups, our approach yields robust, resource- and cost-efficient performance. Tobias Vinçon, Christian Knödler, Leonardo Solis-Vasquez, Arthur Bernhardt, Sajjad Tamimi, Lukas Weber, Florian Stock, Andreas Koch 0001, Ilia Petrov 0001 |
Proc. VLDB Endow. | 7 |
| 2020 | Using Parallel Programming Models for Automotive Workloads on Heterogeneous Systems - a Case StudyabstractDue to the ever-increasing computational demand of automotive applications, and in particular autonomous driving functionalities, the automotive industry and supply vendors are starting to adopt parallel and heterogeneous embedded platforms for their products.However, C and C++, the currently dominating programming languages in this industry, do not provide sufficient mechanisms to target such platforms. Established parallel programming models such as OpenMP and OpenCL on the other hand are tailored towards HPC systems.In this case study, we investigate the applicability of established parallel programming models to automotive workloads on heterogeneous platforms. We pursue a practical approach by re-enacting a typical development process for typical embedded platforms and representative benchmarks. Lukas Sommer, Florian Stock, Leonardo Solis-Vasquez, Andreas Koch 0001 |
PDP | 2 |
| 2009 | Optimizations and Performance of a Robotics Grasping Algorithm Described in Geometric Algebra
Florian Wörsdörfer, Florian Stock, Eduardo Bayro-Corrochano, Dietmar Hildenbrand |
CIARP | 2 |
| 2009 | Acceleration and Energy Efficiency of a Geometric Algebra Computation using Reconfigurable Computers and GPUsabstractGeometric algebra (GA) is a mathematical framework that allows the compact description of geometric relationships and algorithms in many fields of science and engineering. The execution of these algorithms, however, requires significant computational power that made the use of GA impractical for many real-world applications. We describe how a GA-based formulation of the inverse kinematics problem from computer animation and robotics can be accelerated using reconfigurable FPGA-based computing and using a graphics processing unit (GPU). The practical evaluation covers not only the sheer compute performance, but also the energy efficiency. Holger Lange, Florian Stock, Andreas Koch 0001, Dietmar Hildenbrand |
FCCM | 2 |
| 2008 | Memory access parallelisation in high-level language compilation for reconfigurable adaptive computersabstractControl-memory-data flow graphs (CMDFGs) are a unified intermediate representation for compiling high-level languages onto reconfigurable adaptive computing systems. We present both their initial construction as well as transformations for parallel memory accesses. The impact on a number of applications is examined, also considering the effect of caches on acceleration efficiency. Hagen Gädke-Lütjens, Florian Stock, Andreas Koch 0001 |
FPL | 2 |
| 2006 | Architecture Exploration and Tools for Pipelined Coarse-Grained Reconfigurable ArraysabstractThe paper presents a heavily parametrized tool suite that allows the modeling and exploration of heterogeneous, coarse-grained, heavily pipelined reconfigurable architectures. Our tools perform a simultaneous mapping and pipelining-aware placement, which is then followed by a congestion-avoiding router. Initial experiments show that this flow can succeed in implementing applications with smaller track count and reduced connectivity than existing commercial tools, suggesting changes to the original array architecture. The placer can reduce pipeline latency mismatches on converging paths, simplifying the problem for a pipelining-aware routing step Florian Stock, Andreas Koch 0001 |
FPL | 1 |