Kevin Sala

dblp:218/4661 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0001-8233-1185ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Enhancing OmpSs-2 Suspendable Tasks by Combining Operating System and User-Level Threads with C++ Coroutines
abstract
This paper explores three methods for implementing suspendable tasks within task-based programming models: OS threads (pthreads), User-Level Threads (ULTs), and C++ coroutines. We enhance the OmpSs-2 programming model, originally supporting suspendable tasks via pthreads, to also accommodate ULTs and C++ coroutines. This unified approach facilitates a comprehensive comparative analysis using various benchmarks that includes recursive fork-join and data-flow parallelization strategies. Additionally, we contrast these suspension methods with the Cilk and OpenMP task-based programming models, which, despite their efficiency, lack support for suspendable tasks. Key contributions of this study include the novel integration of C++20 coroutines into the OmpSs-2 programming model, which can be combined with pthreads or ULTs. Furthermore, we introduce a new Linux kernel syscall that accelerates pthread context switches by an order of magnitude, thus narrowing the performance gap between ULTs and pthreads in context-switch times from two to one order of magnitude. C++ Coroutines are the ideal solution for scenarios where many tasks are simultaneously in a suspended state, or the frequency of task suspension and resumption is high because they have the smallest memory footprint and minor contextswitch overhead. However, they are limited to C++ programs, do not support TLS, and only allow task suspension at top-level functions. Still, pthreads and ULTs can bring remarkable benefits where C++ Coroutines cannot. We conclude that combining the strengths of coroutines with pthreads or ULTs brings productivity and performance benefits for programming models.
Arnau Cinca, Aleix Roca, Kevin Sala, Raúl Peñacoba Veigas, David Álvarez 0006, Vicenç Beltran 0001
IPDPS3
2024 nOS-V: Co-Executing HPC Applications Using System-Wide Task Scheduling
abstract
Future Exascale systems will feature massive parallelism, many-core processors and heterogeneous architectures. In this scenario, it is increasingly difficult for HPC applications to fully and efficiently utilize the resources in system nodes. Moreover, the increased parallelism exacerbates the effects of existing inefficiencies in current applications. Research has shown that co-scheduling applications to share system nodes instead of executing each application exclusively can increase resource utilization and efficiency. Nevertheless, the current oversubscription and co-location techniques to share nodes have several drawbacks which limit their applicability and make them very application-dependent.This paper presents co-execution through system-wide scheduling. Co-execution is a novel fine-grained technique to execute multiple HPC applications simultaneously on the same node, outperforming current state-of-the-art approaches. We implement this technique in nOS-V, a lightweight tasking library that supports co-execution through system-wide task scheduling. Moreover, nOS-V can be easily integrated with existing programming models, requiring no changes to user applications. We showcase how co-execution with nOS-V significantly reduces schedule makespan for several applications on different scenarios, outperforming prior node-sharing techniques.
David Álvarez 0006, Kevin Sala, Vicenç Beltran 0001
IPDPS2
2021 Combining One-Sided Communications with Task-Based Programming Models
abstract
Hybrid programming combining task-based and message-passing models is an increasingly popular technique to exploit multi-core clusters. The Task-Aware MPI (TAMPI) library integrates both models enabling the safe overlap of computation and communication tasks using two-sided MPI communications. Two-sided primitives combine data transfers with implicit synchronizations, but one-sided models usually offer more efficient data transfers decoupling synchronizations. MPI offers four distinct one-sided synchronization modes, while GASPI is a PGAS API providing one-sided operations with remote notifications for fine inter-process synchronizations.In this paper, we study the challenges of integrating MPI and GASPI one-sided operations with the OpenMP and OmpSs-2 tasking models. We propose and implement several extensions to the GASPI and OmpSs-2 programming models, which are leveraged by a new library called Task-Aware GASPI (TAGASPI). The TAGASPI library allows the efficient and safe use of one-sided operations with remote notifications inside tasks. Both TAGASPI and TAMPI transparently manage communications issued by tasks and allow these to overlap with computation tasks naturally, following a data-flow model. These libraries are complementary and can be mixed in the same application.Our experience porting several mini-apps to this hybrid model shows that TAGASPI helps leverage one-sided communications with similar complexity to pure and hybrid two-sided MPI approaches. We show that our hybrid one-sided approach outperforms the pure MPI strategies, but it also surpasses the TAMPI’s performance when stressing communication phases, e.g., increasing the communication parallelism and reducing the communication tasks’ sizes.
Kevin Sala, Sandra Macià, Vicenç Beltran 0001
CLUSTER1
2021 Advanced synchronization techniques for task-based runtime systems
abstract
Task-based programming models like OmpSs-2 and OpenMP provide a flexible data-flow execution model to exploit dynamic, irregular and nested parallelism. Providing an efficient implementation that scales well with small granularity tasks remains a challenge, and bottlenecks can manifest in several runtime components. In this paper, we analyze the limiting factors in the scalability of a task-based runtime system and propose individual solutions for each of the challenges, including a wait-free dependency system and a novel scalable scheduler design based on delegation. We evaluate how the optimizations impact the overall performance of the runtime, both individually and in combination. We also compare the resulting runtime against state of the art OpenMP implementations, showing equivalent or better performance, especially for fine-grained tasks.
David Álvarez 0006, Kevin Sala, Marcos Maronas, Aleix Roca, Vicenç Beltran 0001
PPoPP2
2020 Towards Data-Flow Parallelization for Adaptive Mesh Refinement Applications
abstract
Adaptive Mesh Refinement (AMR) is a prevalent method used by distributed-memory simulation applications to adapt the accuracy of their solutions depending on the turbulent conditions in each of their domain regions. These applications are usually dynamic since their domain areas are refined or coarsened in various refinement stages during their execution. Thus, they periodically redistribute their workloads among processes to avoid load imbalance. Although the defacto standard for scientific computing in distributed environments is MPI, in recent years, pure MPI applications are being ported to hybrid ones, attempting to cope with modern multi-core systems. Recently, the Task-Aware MPI library was proposed to efficiently integrate MPI communications and tasking models, providing also the transparent management of communications issued by tasks. In this paper, we demonstrate the benefits of porting AMR applications to data-flow programming models leveraging that novel hybrid approach. We exploit most of the application parallelism by taskifying all stages, allowing their natural overlap. We employ these techniques on the miniAMR proxy application, which mimics the refinement, load balancing, communication, and computation patterns of general AMR applications. We evaluate how this approach reduces the time in its computation and communication phases while achieving better programmability than other conventional hybrid techniques.
Kevin Sala, Alejandro Rico, Vicenç Beltran 0001
CLUSTER1
2019 Worksharing Tasks: An Efficient Way to Exploit Irregular and Fine-Grained Loop Parallelism
abstract
Shared memory programming models usually provide worksharing and task constructs. The former relies on the efficient fork-join execution model to exploit structured parallelism; while the latter relies on fine-grained synchronization among tasks and a flexible data-flow execution model to exploit dynamic, irregular, and nested parallelism. On applications that show both structured and unstructured parallelism, both worksharing and task constructs can be combined. However, it is difficult to mix both execution models without penalizing the data-flow execution model. Hence, on many applications structured parallelism is also exploited using tasks to leverage the full benefits of a pure data-flow execution model. However, task creation and management might introduce a non-negligible overhead that prevents the efficient exploitation of fine-grained structured parallelism, especially on many-core processors. In this work, we propose worksharing tasks. These are tasks that internally leverage worksharing techniques to exploit fine-grained structured loop-based parallelism. The evaluation shows promising results on several benchmarks and platforms.
Marcos Maronas, Kevin Sala, Sergi Mateo, Eduard Ayguadé, Vicenç Beltran 0001
HiPC2
2019 Integrating blocking and non-blocking MPI primitives with task-based programming models
Kevin Sala, Xavier Teruel, Josep M. Pérez, Antonio J. Peña, Vicenç Beltran 0001, Jesús Labarta
Parallel Comput.1
2018 Improving the Interoperability between MPI and Task-Based Programming Models
abstract
In this paper we propose an API to pause and resume task execution depending on external events. We leverage this generic API to improve the interoperability between MPI synchronous communication primitives and tasks. When an MPI operation blocks, the task running is paused so that the runtime system can schedule a new task on the core that became idle. Once the MPI operation is completed, the paused task is put again on the runtime system's ready queue. We expose our proposal through a new MPI threading level which we implement through two approaches.
Kevin Sala, Jorge Bellón, Pau Farré, Xavier Teruel, Josep M. Pérez, Antonio J. Peña, Daniel J. Holmes, Vicenç Beltran 0001, Jesús Labarta
EuroMPI1
2018 On the adequacy of lightweight thread approaches for high-level parallel programming models
Adrián Castelló 0001, Rafael Mayo 0002, Kevin Sala, Vicenç Beltran 0001, Pavan Balaji, Antonio J. Peña
Future Gener. Comput. Syst.3