EDBT 2026 Demo / reviewers in the wild / expert
Oana Balmau
dblp:157/1134
· DBLP profile ↗
13ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-6822-8891ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MinatoLoader: Accelerating Machine Learning Training Through Efficient Data PreprocessingabstractMachine learning (ML) frameworks, such as PyTorch and TensorFlow, rely on data loaders to preprocess data before feeding it to accelerators. When preprocessing is inefficiently pipelined, GPUs can remain idle over long periods of time, leading to substantial training delays. For example, PyTorch's default data loaders can cause up to 76% GPU idleness. A key bottleneck is the variability in preprocessing time across samples within the same dataset. Existing data loaders are oblivious to this variability, training all samples uniformly. In this case, a single slow sample can stall the entire batch, causing head-of-line blocking. Rahma Nouaji, Stella Bitchebe, Ricardo Macedo, Oana Balmau |
EuroSys | 4 |
| 2026 | HarMoEny: Efficient Inference of MoE Models
Zachary Doucet, Rishi Sharma 0001, Martijn de Vos, Rafael Pires 0001, Anne-Marie Kermarrec, Oana Balmau |
IPDPS | 6 |
| 2025 | Keigo: Co-designing Log-Structured Merge Key-Value Stores with a Non-Volatile, Concurrency-aware Storage HierarchyabstractWe present Keigo, a concurrency- and workload-aware storage middleware that enhances the performance of log-structured merge key-value stores (LSM KVS) when they are deployed on a hierarchy of storage devices. The key observation behind Keigo is that there is no one-size-fits-all placement of data across the storage hierarchy that optimizes for all workloads. Hence, to leverage the benefits of combining different storage devices, Keigo places files across different devices based on their parallelism, I/O bandwidth, and capacity. We introduce three techniques - concurrency-aware data placement, persistent read-only caching , and context-based I/O differentiation. Keigo is portable across different LSMs, is adaptable to dynamic workloads, and does not require extensive profiling. Our system enables established production KVS such as RocksDB, LevelDB, and Speedb to benefit from heterogeneous storage setups. We evaluate Keigo using synthetic and realistic workloads, showing that it improves the throughput of production-grade LSMs up to 4X for write- and 18X for read-heavy workloads when compared to general-purpose storage systems and specialized LSM KVS. Rúben Adão, Zhongjie Wu, Changjun Zhou, Oana Balmau, João Paulo 0001, Ricardo Macedo |
Proc. VLDB Endow. | 4 |
| 2024 | Falcon: Live Reconfiguration for Stateful Stream Processing on the EdgeabstractStream processing is an attractive paradigm for deploying applications in geo-distributed edge-cloud environments. However, the reverse economics of scale in edge networks and the movement of data sources between edges require the ability to dynamically reconfigure the deployment of stateful applications to adapt to workload variations and user mobility. Unfortunately, existing stream processing engines either do not support the reconfiguration of stateful operators or are ill-suited to edge-cloud environments since they stop application processing during reconfiguration or require costly duplication of application state. We propose Falcon, a new stream processing engine. At its core lies a live key migration approach to allow reconfiguration to occur with minimal disruption to processing, even across distant datacenters. Falcon supports the reconfiguration of stateful operators including different windowing approaches and source mobility across different edge regions. It scales gracefully with network latency, the number of datacenters, and the size and number of keys. Our evaluation in geo-distributed edge-cloud deployments shows that Falcon reduces the length of processing interruptions and their impact on latency by 2 to 4 orders of magnitude compared to the existing state-of-the-art frameworks such as Apache Flink, Trisk, and Meces. Pritish Mishra, Nelson Bore, Brian Ramprasad, Myles Thiessen, Moshe Gabel, Alexandre da Silva Veith, Oana Balmau, Eyal de Lara |
SEC | 7 |
| 2024 | vPIM: Processing-in-Memory VirtualizationabstractData movement is the leading cause of performance degradation and energy consumption in modern data centers. Processing inmemory (PIM) is an architecture that addresses data movement by bringing computation inside the memory chips. This paper is the first to study the virtualization of PIM devices by designing and implementing vPIM, an open-source UPMEM-based virtualization system for the cloud. Our vPIM design considers four requirements: Compatibility such that no hardware and no hypervisor changes are needed; Multiplexing and isolation for a higher utilization ratio; Utilizability and transparency such that applications written for PIM can be efficiently run out-of-the-box, leading to rapid adoption; Minimalization of virtualization performance overhead. Dufy Teguia, Stella Bitchebe, Oana Balmau, Alain Tchana |
Middleware | 4 |
| 2022 | Shepherd: Seamless Stream Processing on the EdgeabstractNext generation applications such as augmented/vir-tual reality, autonomous driving, and Industry 4.0, have tight latency constraints and produce large amounts of data. To address the real-time nature and high bandwidth usage of new applications, edge computing provides an extension to the cloud infrastructure through a hierarchy of datacenters located between the edge devices and the cloud. Outside of the cloud and closer to the edge, the network becomes more dynamic requiring stream processing frameworks to adapt more frequently. Cloud based frameworks adapt very slowly because they employ a stop-the-world approach and it can take several minutes to reconfigure jobs resulting in downtime. In this paper, we propose Shepherd, a new stream processing framework for edge computing. Shepherd minimizes downtime during application reconfiguration, with almost no impact on data processing latency. Our experiments show that, compared to Apache Storm, Shepherd reduces application downtime from several minutes to a few tens of milliseconds. Brian Ramprasad, Pritish Mishra, Myles Thiessen, Alexandre da Silva Veith, Moshe Gabel, Oana Balmau, Abelard Chow, Eyal de Lara |
SEC | 7 |
| 2020 | Kvell+: Snapshot Isolation without Snapshots
Baptiste Lepers, Oana Balmau, Willy Zwaenepoel |
OSDI | 2 |
| 2019 | KVell: the design and implementation of a fast persistent key-value storeabstractModern block-addressable NVMe SSDs provide much higher bandwidth and similar performance for random and sequential access. Persistent key-value stores (KVs) designed for earlier storage devices, using either Log-Structured Merge (LSM) or B trees, do not take full advantage of these new devices. Logic to avoid random accesses, expensive operations for keeping data sorted on disk, and synchronization bottlenecks make these KVs CPU-bound on NVMe SSDs. Baptiste Lepers, Oana Balmau, Willy Zwaenepoel |
SOSP | 2 |
| 2019 | SILK: Preventing Latency Spikes in Log-Structured Merge Key-Value Stores
Oana Balmau, Florin Dinu, Willy Zwaenepoel, Ravishankar Chandhiramoorthi, Diego Didona |
USENIX ATC | 1 |
| 2018 | SILK+ Preventing Latency Spikes in Log-Structured Merge Key-Value Stores Running Heterogeneous WorkloadsabstractLog-Structured Merge Key-Value stores (LSM KVs) are designed to offer good write performance, by capturing client writes in memory, and only later flushing them to storage. Writes are later compacted into a tree-like data structure on disk to improve read performance and to reduce storage space use. It has been widely documented that compactions severely hamper throughput. Various optimizations have successfully dealt with this problem. These techniques include, among others, rate-limiting flushes and compactions, selecting among compactions for maximum effect, and limiting compactions to the highest level by so-called fragmented LSMs. In this article, we focus on latencies rather than throughput. We first document the fact that LSM KVs exhibit high tail latencies. The techniques that have been proposed for optimizing throughput do not address this issue, and, in fact, in some cases, exacerbate it. The root cause of these high tail latencies is interference between client writes, flushes, and compactions. Another major cause for tail latency is the heterogeneous nature of the workloads in terms of operation mix and item sizes whereby a few more computationally heavy requests slow down the vast majority of smaller requests. We introduce the notion of an Input/Output (I/O) bandwidth scheduler for an LSM-based KV store to reduce tail latency caused by interference of flushing and compactions and by workload heterogeneity. We explore three techniques as part of this I/O scheduler: (1) opportunistically allocating more bandwidth to internal operations during periods of low load, (2) prioritizing flushes and compactions at the lower levels of the tree, and (3) separating client requests by size and by data access path. SILK+ is a new open-source LSM KV that incorporates this notion of an I/O scheduler. Oana Balmau, Florin Dinu, Willy Zwaenepoel, Ravishankar Chandhiramoorthi, Diego Didona |
ACM Trans. Comput. Syst. | 1 |
| 2017 | FloDB: Unlocking Memory in Persistent Key-Value StoresabstractLog-structured merge (LSM) data stores enable to store and process large volumes of data while maintaining good performance. They mitigate the I/O bottleneck by absorbing updates in a memory layer and transferring them to the disk layer in sequential batches. Yet, the LSM architecture fundamentally requires elements to be in sorted order. As the amount of data in memory grows, maintaining this sorted order becomes increasingly costly. Contrary to intuition, existing LSM systems could actually lose throughput with larger memory components. Oana Balmau, Rachid Guerraoui, Vasileios Trigonakis, Igor Zablotchi |
EuroSys | 1 |
| 2017 | TRIAD: Creating Synergies Between Memory, Disk and Log in Log Structured Key-Value Stores
Oana Balmau, Diego Didona, Rachid Guerraoui, Willy Zwaenepoel, Huapeng Yuan, Aashray Arora, Pavan Konka |
USENIX ATC | 1 |
| 2016 | Fast and Robust Memory Reclamation for Concurrent Data StructuresabstractIn concurrent systems without automatic garbage collection, it is challenging to determine when it is safe to reclaim memory, especially for lock-free data structures. Existing concurrent memory reclamation schemes are either fast but do not tolerate process delays, robust to delays but with high overhead, or both robust and fast but narrowly applicable. This paper proposes QSense, a novel concurrent memory reclamation technique. QSense is a hybrid technique with a fast path and a fallback path. In the common case (without process delays), a high-performing memory reclamation scheme is used (fast path). If process delays block memory reclamation through the fast path, a robust fallback path is used to guarantee progress. The fallback path uses hazard pointers, but avoids their notorious need for frequent and expensive memory fences. Oana Balmau, Rachid Guerraoui, Maurice Herlihy, Igor Zablotchi |
SPAA | 1 |