Juliette Fournis d'Albiat

dblp:417/0136 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0005-6575-6989ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 90% High-performance computing · 10%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel programming models
task-based programming
0.912025
Leveraging iterative applications to improve the scalability of task-based programming models on distributed systems · ACM Trans. Archit. Code Optim. 2025
Parallel and multicore computing › task scheduling
task graph scheduling
0.912025
Leveraging iterative applications to improve the scalability of task-based programming models on distributed systems · ACM Trans. Archit. Code Optim. 2025
High-performance computing
distributed memory systems
0.312025
Leveraging iterative applications to improve the scalability of task-based programming models on distributed systems · ACM Trans. Archit. Code Optim. 2025
Parallel and multicore computing › parallel programming models › message passing
MPI communication
0.312025
Leveraging iterative applications to improve the scalability of task-based programming models on distributed systems · ACM Trans. Archit. Code Optim. 2025
Parallel and multicore computing
parallel programming runtimes
0.312025
Leveraging iterative applications to improve the scalability of task-based programming models on distributed systems · ACM Trans. Archit. Code Optim. 2025

Methods — techniques the papers use, named apart from their topics

taskiter directive · 0.9cyclic task graph partitioning · 0.9
YearPublicationVenuePosition
2025 Leveraging iterative applications to improve the scalability of task-based programming models on distributed systems
abstract
Distributed tasking models such as OmpSs-2@Cluster, StarPU-MPI, and PaRSEC express HPC applications as task graphs with explicit dependencies. The single task graph unifies the representation of parallelism across CPU cores, accelerators, and distributed-memory nodes, offering higher programmer productivity compared to traditional MPI + X. Most task-based models construct the task graph sequentially, which provides a clear and familiar programming model, simplifying code development, maintenance, and porting. However, this design introduces a bottleneck in task creation and dependency management, limiting performance and scalability. As a result, unless the tasks are very coarse-grained, current distributed sequential tasking models cannot match the performance of MPI + X. Many scientific applications, however, are iterative in nature, constructing the same directed acyclic task graph at each timestep. We exploit this structure to eliminate the sequential bottleneck and control message overhead in a sequentially-constructed distributed tasking model, while preserving its simplicity and productivity. Our approach builts on the recently proposed taskiter directive for OpenMP and OmpSs-2, allowing a single iteration to be expressed as a cyclic graph. The runtime partitions the cyclic graph across nodes, precomputes the MPI transfers, and then executes the loop body at low overhead. By integrating the MPI communications directly into the application’s task graph, our approach naturally overlaps computation and communication, in some cases exposing dramatically more parallelism than fork–join MPI + OpenMP. We define the programming model and describe the full runtime implementation, and integrate our proposal into OmpSs-2@Cluster. We evaluate it using five benchmarks on up to 128 nodes of the MareNostrum 5 supercomputer. For applications with fork–join parallelism, our approach has performance similar to fork–join MPI + OpenMP, making it a viable productive alternative, unlike the existing OmpSs-2@Cluster model, which is up to 7.7 times slower than MPI + OpenMP. For a 2D Gauss–Seidel stencil computation, our approach enables 3D wavefront computation, giving performance up to 22 times faster than fork–join MPI + OpenMP and on-a-par with state-of-the-art TAMPI + OmpSs-2. All software, comprising the compiler, runtime, and benchmarks, is released open source. 1
Omar Shaaban Ibrahim ali, Juliette Fournis d'Albiat, Isabel Piedrahita, Vicenç Beltran 0001, Xavier Martorell, Paul M. Carpenter, Eduard Ayguadé, Jesús Labarta
ACM Trans. Archit. Code Optim.2