EDBT 2026 Demo / reviewers in the wild / expert
Aleix Roca
dblp:286/1963 · also Aleix Roca Nonell
· DBLP profile ↗
4ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0002-6715-3605ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Thread Scheduling under Oversubscription: A User-Space Framework for Coordinating Multi-runtime and Multi-process WorkloadsabstractThe convergence of high-performance computing (HPC) and artificial intelligence (AI) is driving the emergence of increasingly complex parallel applications and workloads. These workloads often combine multiple parallel runtimes within the same application or across co-located jobs, creating scheduling demands that place significant stress on traditional OS schedulers. When oversubscribed (there are more ready threads than cores), OS schedulers rely on periodic preemptions to multiplex cores, often introducing interference that may degrade performance. In this paper, we present: (1) The User-space Scheduling Framework (USF), a novel seamless process scheduling framework completely implemented in user-space. USF enables users to implement their own process scheduling algorithms without requiring special permissions. We evaluate USF with its default cooperative policy, (2) SCHED_COOP, designed to reduce interference by switching threads only upon blocking. This approach mitigates well-known issues such as Lock-Holder Preemption (LHP), Lock-Waiter Preemption (LWP), and scalability collapse. We implement USF and SCHED_COOP by extending the GNU C library with the nOS-V runtime, enabling seamless coordination across multiple runtimes (e.g., OpenMP) without requiring invasive application changes. Evaluations show gains up to 2.4x in oversubscribed multi-process scenarios, including nested BLAS workloads, multi-process PyTorch inference with LLaMA-3, and Molecular Dynamics (MD) simulations. Aleix Roca, Vicenç Beltran 0001 |
PPoPP | 1 |
| 2025 | Enhancing OmpSs-2 Suspendable Tasks by Combining Operating System and User-Level Threads with C++ CoroutinesabstractThis paper explores three methods for implementing suspendable tasks within task-based programming models: OS threads (pthreads), User-Level Threads (ULTs), and C++ coroutines. We enhance the OmpSs-2 programming model, originally supporting suspendable tasks via pthreads, to also accommodate ULTs and C++ coroutines. This unified approach facilitates a comprehensive comparative analysis using various benchmarks that includes recursive fork-join and data-flow parallelization strategies. Additionally, we contrast these suspension methods with the Cilk and OpenMP task-based programming models, which, despite their efficiency, lack support for suspendable tasks. Key contributions of this study include the novel integration of C++20 coroutines into the OmpSs-2 programming model, which can be combined with pthreads or ULTs. Furthermore, we introduce a new Linux kernel syscall that accelerates pthread context switches by an order of magnitude, thus narrowing the performance gap between ULTs and pthreads in context-switch times from two to one order of magnitude. C++ Coroutines are the ideal solution for scenarios where many tasks are simultaneously in a suspended state, or the frequency of task suspension and resumption is high because they have the smallest memory footprint and minor contextswitch overhead. However, they are limited to C++ programs, do not support TLS, and only allow task suspension at top-level functions. Still, pthreads and ULTs can bring remarkable benefits where C++ Coroutines cannot. We conclude that combining the strengths of coroutines with pthreads or ULTs brings productivity and performance benefits for programming models. Arnau Cinca, Aleix Roca, Kevin Sala, Raúl Peñacoba Veigas, David Álvarez 0006, Vicenç Beltran 0001 |
IPDPS | 2 |
| 2021 | Advanced synchronization techniques for task-based runtime systemsabstractTask-based programming models like OmpSs-2 and OpenMP provide a flexible data-flow execution model to exploit dynamic, irregular and nested parallelism. Providing an efficient implementation that scales well with small granularity tasks remains a challenge, and bottlenecks can manifest in several runtime components. In this paper, we analyze the limiting factors in the scalability of a task-based runtime system and propose individual solutions for each of the challenges, including a wait-free dependency system and a novel scalable scheduler design based on delegation. We evaluate how the optimizations impact the overall performance of the runtime, both individually and in combination. We also compare the resulting runtime against state of the art OpenMP implementations, showing equivalent or better performance, especially for fine-grained tasks. David Álvarez 0006, Kevin Sala, Marcos Maronas, Aleix Roca, Vicenç Beltran 0001 |
PPoPP | 4 |
| 2019 | A Linux Kernel Scheduler Extension for Multi-core SystemsabstractThe Linux kernel is mostly designed for multi-programed environments, but high-performance applications have other requirements. Such applications are run standalone, and usually rely on runtime systems to distribute the application's workload on worker threads, one per core. However, due to current OSes limitations, it is not feasible to track whether workers are actually running or blocked due to, for instance, a requested resource. For I/O intensive applications, this leads to a significant performance degradation given that the core of a blocked thread becomes idle until it is able to run again. In this paper, we present the proof-of-concept of a Linux kernel extension denoted User-Monitored Threads (UMT) which tackles this problem. Our extension allows a user-space process to be notified of when the selected threads become blocked or unblocked, making it possible for a runtime to schedule additional work on the idle core. We implemented the extension on the Linux Kernel 5.1 and adapted the Nanos6 runtime of the OmpSs-2 programming model to take advantage of it. The whole prototype was tested on two applications which, on the tested hardware and the appropriate conditions, reported speedups of almost 2x. Aleix Roca, Samuel Rodríguez, Albert Segura, Kevin Marquet, Vicenç Beltran 0001 |
HiPC | 1 |