VLDB 2026 Research / reviewers in the wild / expert
Sara Royuela
dblp:55/11418
· DBLP profile ↗
18ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-7644-0868ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GuardianOMP: A Framework for Highly Productive Fault Tolerance Via OpenMP Task-Level ReplicationabstractAdvanced critical real-time embedded systems (CRTES) like autonomous cars impose dependability and highperformance requirements. Parallel architectures are increasingly adopted to cope with the performance needs of these systems. In this context, the OpenMP parallel programming model is becoming popular for effectively exploiting complex hardware due to its productivity and time-predictability, but it lacks the mechanisms to provide fail-operational behaviour. This paper presents GuardianOMP, a complete open-source framework based on extensions to the OpenMP language, compiler and runtime tools built on top of LLVM framework to improve the fault tolerance of critical systems while minimizing overheads. GuardianOMP offers software-based task-level replication and user-directed fault detection to mitigate the impact of transient faults. The proposed framework provides flexibility in defining parameters, such as the number of replicas, the replication type, and the voting method, incurring minimal efforts to provide programmability and adaptability. Moreover, it exploits the available parallel resources of shared memory systems to minimize the impact on the system's overall performance. This work evaluates GuardianOMP in terms of accuracy, performance, and programmability by employing BOTS (Barcelona OpenMP Task Suite) benchmarks, showing the elimination of false positives, a slight reduction in overhead and an enhancement in the programmability and flexibility compared to the state-of-the-art. The framework is also evaluated in a real-world railway obstacle detection application, reinforcing its effectiveness. Furthermore, this work contributes with a new open-source fault injector tool that simulates software transient faults in processes memory and hardware registers and enables the evaluation of this work and its further use and extension. Adrian Munera, Eduardo Quiñones, Sara Royuela |
IPDPS | 3 |
| 2025 | Profiler-Guided Execution of Recurrent OpenMP Task Graphs on Heterogeneous ClustersabstractDistributed task-based execution models are well-suited for parallelizing irregular applications across clusters. OpenMP Cluster (OMPC) extends the traditional OpenMP tasking model to support distributed memory systems, leveraging a HEFT-based scheduler to improve resource utilization. However, the efficiency of such a scheduler depends heavily on accurate estimates of task execution and communication costs – information that is often difficult to obtain reliably and efficiently. To address this limitation, we propose a novel scheduling framework that combines the recent taskgraph directive introduced in OpenMP 6.0 with partial online profiling of iterative applications. Our approach performs quasi-static scheduling by recording task graphs at runtime and selectively profiling representative iterations to estimate performance. This information is interpolated and fed back into the scheduler to enhance decision-making. We demonstrate that our framework can improve scheduling quality with minimal overhead, making it suitable for long-running or repetitive workloads commonly found in High-Performance Computing (HPC) applications. We achieve up to 20% speedup for the total application and 4× speedup for scheduling. Rémy Neveu, Rodrigo Ceccato, Adrian Munera, Sara Royuela, José Monsalve Diaz, Hervé Yviquel |
SBAC-PAD | 4 |
| 2024 | Enhancing Heterogeneous Computing Through OpenMP and GPU GraphabstractModern computing platforms are increasingly heterogeneous, most of them include accelerators such as GPU. OpenMP as the de-facto standard to parallelize CPU applications, incorporates target construct allowing users to offload work onto such accelerators, to leverage the computing capability provided by these accelerators while maintaining a simple, directive based programming style. However, due to architectural differences between CPUs and GPUs, OpenMP is not always able to exploit the latest computing optimizations available on GPUs. Chenle Yu, Sara Royuela, Eduardo Quiñones |
ICPP | 2 |
| 2024 | Evaluation of Heuristic Task-to-Thread Mapping Using Static and Dynamic Approaches
Tiago Carvalho 0001, Luís Miguel Pinho, Sara Royuela |
JSSPP | 4 |
| 2024 | Fine-grained adaptive parallelism for automotive systems through AMALTHEA and OpenMPabstractThe software development complexity of automotive systems has significantly increased during the last decade due to the latest Advanced Driving Assistance System (ADAS) functionalities. To effectively address this complexity, domain specific modeling languages (DSMLs) like AUTOSAR or an open-source system performance model for AUTOSAR-aligned systems, APP4MC, have become a common trend in the automotive industry. DSMLs allow for easily capturing the functional and non-functional requirements of the system without needing to master low level details of the programming model or the processor architecture. Unfortunately, current DSMLs do not support the parallel programming models, like OpenMP and CUDA, that are used to exploit parallel heterogeneous architectures featuring acceleration devices such as GPUs and FPGAs required. These architectures are however essential to cope with the performance needs of ADAS. This exposes a gap between the DSMLs used by automotive designers to enhance software productivity and leverage verification and validation processes, and the parallel processor architectures used in this domain. This paper presents a complete framework to safely exploit the inherent parallelism exposed by the AMALTHEA system description, supported in APP4MC, by: (1) automatically transforming the high-level design into the OpenMP parallel programming model targeting both host and accelerator parallelism, and (2) using compiler analysis techniques to prove the correctness of the model transformed to OpenMP code. The paper contributes also with (3) an analysis of the parallel execution model allowed by the AMALTHEA DSML and that of OpenMP, and (4) a performance plus productivity evaluation of the proposed framework on real automotive systems executed on an embedded GPU-based processor architecture. Adrian Munera, Sara Royuela, Michael Pressler, Harald Mackamul, Dirk Ziegenbein, Eduardo Quiñones |
J. Syst. Archit. | 2 |
| 2024 | Time-predictable task-to-thread mapping in multi-core processorsabstractThe performance of time-predictable systems can be improved in multi-core processors using parallel programming models (e.g., OpenMP). However, schedulability analysis of parallel applications is a big challenge due to their sophisticated structure. The common drawbacks of current task-to-thread mapping approaches in OpenMP are that they (i) utilize a global queue in the mapping process, which may increase contention, (ii) do not apply heuristic techniques, which may reduce the predictability and performance of the system, and (iii) use basic analytical techniques, which may cause notable pessimism in the temporal conditions. Accordingly, this paper proposes a task-to-thread mapping method in multi-core processors based on the OpenMP framework. The mapping process is carried out through two phases: allocation and dispatching. Each thread has an allocation queue in order to minimize contention, and the allocation and dispatching processes are performed using several heuristic algorithms to enhance predictability. In the allocation phase, each task-part from the OpenMP DAG is allocated to one of the allocation queues, which includes both sibling and child task-parts. A suitable thread (i.e., allocation queue) is selected using one of the suggested heuristic allocation algorithms. In the dispatching phase, when a thread is idle, a task-part is selected from its allocation queue using one of the suggested heuristic dispatching algorithms and then dispatched to and executed by the thread. The performance of the proposed method is evaluated under different conditions (e.g., varying the number of tasks and the number of threads) in terms of application response time and overhead of the mapping process. The simulation results show that the proposed method surpasses the other methods, especially in the scenario that includes overhead of the mapping. In addition, a prototype implementation of the main heuristics is evaluated using two kernels from real-world applications, showing that the methods work better than LLVM's default scheduler in most of the configurations. Sara Royuela, Luís Miguel Pinho, Tiago Carvalho 0001, Eduardo Quiñones |
J. Syst. Archit. | 2 |
| 2023 | Framework for the Analysis and Configuration of Real-Time OpenMP ApplicationsabstractHigh-performance cyber-physical applications impose several requirements with respect to performance, functional correctness and non-functional aspects. Nowadays, the design of these systems usually follows a model-driven approach, where models generate executable applications, usually with an automated approach. As these applications might execute in different parallel environments, their behavior becomes very hard to predict, and making the verification of non-functional requirements complicated. In this regard, it is crucial to analyse and understand the impact that the mapping and scheduling of computation have on the real-time response of the applications. In fact, different strategies in these steps of the parallel orchestration may produce significantly different interference, leading to different timing behaviour.Tuning the application parameters and the system configuration proves to be one of the most fitting solutions. The design space can however be very cumbersome for a developer to test manually all combinations of application and system configurations. This paper presents a methodology and a toolset to profile, analyse, and configure the timing behaviour of high-performance cyber-physical applications and the target platforms. The methodology leverages on the possibility of generating a task dependency graph representing the parallel computation to evaluate, through measurements, different mapping configurations and select the one that minimizes response time. Tiago Carvalho 0001, Luís Miguel Pinho, Sara Royuela, Adrian Munera, Eduardo Quiñones |
INDIN | 4 |
| 2023 | Taskgraph: A Low Contention OpenMP Tasking FrameworkabstractOpenMP is the de-facto standard for shared memory systems in High-Performance Computing (HPC). It includes a tasking model that offers a high-level of abstraction to effectively exploit structured (loop-based) and highly dynamic unstructured (task-based) parallelism in an easy and flexible way. Unfortunately, the run-time overheads introduced to manage tasks are (very) high in most common OpenMP frameworks (e.g., GCC, LLVM), which defeats the potential benefits of the tasking model, and makes it suitable for coarse-grained tasks only. This paper presentstaskgraph, a framework that uses a task dependency graph (TDG) to represent a region of code implemented with OpenMP tasks in order to reduce the run-time overheads associated with the management of tasks, i.e., contention and parallel orchestration, including task creation and synchronization. The TDG avoids the overheads related to the resolution of task dependencies and greatly reduces those deriving from accesses to shared resources. Moreover, the taskgraph framework introduces in OpenMP therecord-and-replayexecution model that accelerates the taskgraph region from its second execution. Overall, the multiple optimizations presented in this paper allow exploiting fine-grained OpenMP tasks to cope with the trend in current applications pointing to leverage massive on-node parallelism, fine-grained and dynamic scheduling paradigms. The framework is implemented on LLVM 15.0. Results show that the taskgraph implementation outperforms the vanilla OpenMP system in terms of performance and scalability, for all structured and unstructured parallelism, and considering coarse and fine grained tasks. Furthermore, the proposed framework makes the tasking model a competitive alternative to the OpenMP thread model in most cases. Chenle Yu, Sara Royuela, Eduardo Quiñones |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Heuristic-based Task-to-Thread Mapping in Multi-Core ProcessorsabstractOpenMP can be used in real-time applications to enhance system performance. However, predictability of OpenMP applications is still a challenge. This paper investigates heuristics for the mapping of OpenMP task graphs in underlying threads, for the development of time-predictable OpenMP programs. These approaches are based on a global scheduling queue, as well as per-thread allocation queues. The proposed method is divided into scheduling and allocation phases. In the former phase, OpenMP task-parts are discovered from OpenMP graph and placed in the scheduling queue. Afterwards, an appropriate allocation queue is selected for each task-part using four heuristic algorithms. In the latter phase, the best task-part is selected from the allocation queue to be allocated to and executed by an idle thread. Preliminary simulation results show that the new method overcomes BFS and WFS in terms of scheduling time and idle time. Sara Royuela, Luís Miguel Pinho, Tiago Carvalho 0001, Eduardo Quiñones |
ETFA | 2 |
| 2020 | Towards a Qualifiable OpenMP Framework for Embedded SystemsabstractOpenMP is a very convenient programming model for critical real-time parallel applications due to its powerful tasking model and its proven time predictability. However, current implementations are not suitable for critical environments based on the intensive use of dynamically allocated memory needed to efficiently manage the parallel execution. This jeopardizes the qualification processes needed to ensure that the integrated software stack is compliant with system requirements. This paper proposes a novel OpenMP framework that statically allocates the data structures needed to efficiently manage the parallel execution of OpenMP tasks. Our framework is composed of a compiler that captures the environment of the OpenMP tasks instantiated along the parallel execution and bounds the exposed parallelism, and a runtime implementing a lazy task creation policy that significantly reduces the runtime memory requirements, whilst exploiting parallelism efficiently. The evaluation shows that our tool achieves the same performance as current OpenMP implementations, while bounds and drastically reduces the dynamic memory requirements at run-time. Adrian Munera, Sara Royuela, Eduardo Quiñones |
DATE | 2 |
| 2020 | A Toolchain to Verify the Parallelization of OmpSs-2 Applications
Simone Economo, Sara Royuela, Eduard Ayguadé, Vicenç Beltran 0001 |
Euro-Par | 2 |
| 2020 | Experiences on the characterization of parallel applications in embedded systems with Extrae/ParaverabstractCutting-edge functionalities in embedded systems require the use of parallel architectures to meet their performance requirements. This imposes the introduction of a new layer in the software stacks of embedded systems: the parallel programming model. Unfortunately, the tools used to analyze embedded systems fall short to characterize the performance of parallel applications at a parallel programming model level, and correlate this with information about non-functional requirements such as real-time, energy, memory usage, etc. HPC tools, like Extrae, are designed with that level of abstraction in mind, but their main focus is on performance evaluation. Overall, providing insightful information about the performance of parallel embedded applications at the parallel programming model level, and relate it to the non-functional requirements, is of paramount importance to fully exploit the performance capabilities of parallel embedded architectures. Adrian Munera, Sara Royuela, Germán Llort, Estanislao Mercadal, Franck Wartel, Eduardo Quiñones |
ICPP | 2 |
| 2020 | The AMPERE Project: : A Model-driven development framework for highly Parallel and EneRgy-Efficient computation supporting multi-criteria optimizationabstractThe high-performance requirements needed to implement the most advanced functionalities of current and future Cyber-Physical Systems (CPSs) are challenging the development processes of CPSs. On one side, CPSs rely on model-driven engineering (MDE) to satisfy the non-functional constraints and to ensure a smooth and safe integration of new features. On the other side, the use of complex parallel and heterogeneous embedded processor architectures becomes mandatory to cope with the performance requirements. In this regard, parallel programming models, such as OpenMP or CUDA, are a fundamental brick to fully exploit the performance capabilities of these architectures. However, parallel programming models are not compatible with current MDE approaches, creating a gap between the MDE used to develop CPSs and the parallel programming models supported by novel and future embedded platforms.The AMPERE project will bridge this gap by implementing a novel software architecture for the development of advanced CPSs. To do so, the proposed software architecture will be capable of capturing the definition of the components and communications described in the MDE framework, together with the non-functional properties, and transform it into key parallel constructs present in current parallel models, which may require extensions. These features will allow for making an efficient use of underlying parallel and heterogeneous architectures, while ensuring compliance with non-functional requirements, including those on real-time performance of the system. Eduardo Quiñones, Sara Royuela, Claudio Scordino, Paolo Gai, Luís Miguel Pinho, Luís Nogueira, Jan Rollo, Tommaso Cucinotta, Alessandro Biondi 0001, Arne Hamann 0001, Dirk Ziegenbein, Hadi Saoud, Romain Soulat, Björn Forsberg, Luca Benini, Gianluca Mandò, Luigi Rucher |
ISORC | 2 |
| 2020 | OpenMP to CUDA graphs: a compiler-based transformation to enhance the programmability of NVIDIA devicesabstractHeterogeneous computing is increasingly being used in a diversity of computing systems, ranging from HPC to the real-time embedded domain, to cope with the performance requirements. Due to the variety of accelerators, e.g., FPGAs, GPUs, the use of high-level parallel programming models is desirable to exploit the performance capabilities of them, while maintaining an adequate productivity level. In that regard, OpenMP is a well-known high-level programming model that incorporates powerful task and accelerator models capable of efficiently exploiting structured and unstructured parallelism in heterogeneous computing. This paper presents a novel compiler transformation technique that automatically transforms OpenMP code into CUDA graphs, combining the benefits of programmability of a high-level programming model such as OpenMP, with the performance benefits of a low-level programming model such as CUDA. Evaluations have been performed on two NVIDIA GPUs from the HPC and embedded domains, i.e., the V100 and the Jetson AGX respectively. Chenle Yu, Sara Royuela, Eduardo Quiñones |
SCOPES | 2 |
| 2020 | Enabling Ada and OpenMP runtimes interoperability through template-based executionabstractThe growing trend to support parallel computation to enable the performance gains of the recent hardware architectures is increasingly present in more conservative domains, such as safety-critical systems. Applications such as autonomous driving require levels of performance only achievable by fully leveraging the potential parallelism in these architectures. To address this requirement, the Ada language, designed for safety and robustness, is considering to support parallel features in the next revision of the standard (Ada 202X). Recent works have motivated the use of OpenMP, a de facto standard in high-performance computing, to enable parallelism in Ada, showing the compatibility of the two models, and proposing static analysis to enhance reliability. This paper summarizes these previous efforts towards the integration of OpenMP into Ada to exploit its benefits in terms of portability, programmability and performance, while providing the safety benefits of Ada in terms of correctness. The paper extends those works proposing and evaluating an application transformation that enables the OpenMP and the Ada runtimes to operate (under certain restrictions) as they were integrated. The objective is to allow Ada programmers to (naturally) experiment and evaluate the benefits of parallelizing concurrent Ada tasks with OpenMP while ensuring the compliance with both specifications. Sara Royuela, Luís Miguel Pinho, Eduardo Quiñones |
J. Syst. Archit. | 1 |
| 2018 | Converging safety and high-performance domains: Integrating OpenMP into AdaabstractThe use of parallel heterogeneous embedded architectures is needed to implement the level of performance required in advanced safety-critical systems. Hence, there is a demand for using high level parallel programming models capable of efficiently exploiting the performance opportunities. In this paper, we evaluate the incorporation of OpenMP, a parallel programming model used in HPC, into Ada, a language spread in safety-critical domains. We demonstrate that the execution model of OpenMP is compatible with the recently proposed Ada tasklet model, meant to exploit fine-grain structured parallelism. Moreover, we show the compatibility of the OpenMP and tasklet models, enabling the use of OpenMP directives in Ada to further exploit unstructured parallelism and heterogeneous computation. Finally, we state the safety properties of OpenMP and analyze the interoperability between the OpenMP and Ada runtimes. Overall, we conclude that OpenMP can be effectively incorporated into Ada without jeopardizing its safety properties. Sara Royuela, Luís Miguel Pinho, Eduardo Quiñones |
DATE | 1 |
| 2016 | A lightweight OpenMP4 run-time for embedded systemsabstractOpenMP is increasingly being adopted by current many-core embedded processors to exploit their parallel computation capabilities. Unfortunately, current run-time implementations of the latest specification (v4.0) are not suitable for processors relying on small and fast on-chip memories, due to its memory consumption. This paper proposes an OpenMP4 run-time that reduces the memory consumption while providing the same performance. Our run-time relies on a new compiler pass capable to generate the task dependency graph of OpenMP programs, which is then efficiently stored in memory. Roberto Vargas, Sara Royuela, Maria A. Serrano, Xavier Martorell, Eduardo Quiñones |
ASP-DAC | 2 |
| 2015 | Optimizing Overlapped Memory Accesses in User-directed VectorizationabstractCurrent processors incorporate wide and powerful vector units whose optimal exploitation is crucial to reach peak performance. However, present autovectorizing compilers fall short of that goal. Exploiting some vector instructions requires aggressive approaches that are not affordable in production compilers. Thus, advanced programmers pursuing the best performance from their applications are compelled to manually vectorize them using low-level SIMD intrinsics. Diego Caballero, Sara Royuela, Roger Ferrer, Alejandro Duran, Xavier Martorell |
ICS | 2 |