EDBT 2026 Demo / reviewers in the wild / expert
Adrian Munera
dblp:268/1941
· DBLP profile ↗
6ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0001-9031-3606ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GuardianOMP: A Framework for Highly Productive Fault Tolerance Via OpenMP Task-Level ReplicationabstractAdvanced critical real-time embedded systems (CRTES) like autonomous cars impose dependability and highperformance requirements. Parallel architectures are increasingly adopted to cope with the performance needs of these systems. In this context, the OpenMP parallel programming model is becoming popular for effectively exploiting complex hardware due to its productivity and time-predictability, but it lacks the mechanisms to provide fail-operational behaviour. This paper presents GuardianOMP, a complete open-source framework based on extensions to the OpenMP language, compiler and runtime tools built on top of LLVM framework to improve the fault tolerance of critical systems while minimizing overheads. GuardianOMP offers software-based task-level replication and user-directed fault detection to mitigate the impact of transient faults. The proposed framework provides flexibility in defining parameters, such as the number of replicas, the replication type, and the voting method, incurring minimal efforts to provide programmability and adaptability. Moreover, it exploits the available parallel resources of shared memory systems to minimize the impact on the system's overall performance. This work evaluates GuardianOMP in terms of accuracy, performance, and programmability by employing BOTS (Barcelona OpenMP Task Suite) benchmarks, showing the elimination of false positives, a slight reduction in overhead and an enhancement in the programmability and flexibility compared to the state-of-the-art. The framework is also evaluated in a real-world railway obstacle detection application, reinforcing its effectiveness. Furthermore, this work contributes with a new open-source fault injector tool that simulates software transient faults in processes memory and hardware registers and enables the evaluation of this work and its further use and extension. Adrian Munera, Eduardo Quiñones, Sara Royuela |
IPDPS | 1 |
| 2025 | Profiler-Guided Execution of Recurrent OpenMP Task Graphs on Heterogeneous ClustersabstractDistributed task-based execution models are well-suited for parallelizing irregular applications across clusters. OpenMP Cluster (OMPC) extends the traditional OpenMP tasking model to support distributed memory systems, leveraging a HEFT-based scheduler to improve resource utilization. However, the efficiency of such a scheduler depends heavily on accurate estimates of task execution and communication costs – information that is often difficult to obtain reliably and efficiently. To address this limitation, we propose a novel scheduling framework that combines the recent taskgraph directive introduced in OpenMP 6.0 with partial online profiling of iterative applications. Our approach performs quasi-static scheduling by recording task graphs at runtime and selectively profiling representative iterations to estimate performance. This information is interpolated and fed back into the scheduler to enhance decision-making. We demonstrate that our framework can improve scheduling quality with minimal overhead, making it suitable for long-running or repetitive workloads commonly found in High-Performance Computing (HPC) applications. We achieve up to 20% speedup for the total application and 4× speedup for scheduling. Rémy Neveu, Rodrigo Ceccato, Adrian Munera, Sara Royuela, José Monsalve Diaz, Hervé Yviquel |
SBAC-PAD | 3 |
| 2024 | Fine-grained adaptive parallelism for automotive systems through AMALTHEA and OpenMPabstractThe software development complexity of automotive systems has significantly increased during the last decade due to the latest Advanced Driving Assistance System (ADAS) functionalities. To effectively address this complexity, domain specific modeling languages (DSMLs) like AUTOSAR or an open-source system performance model for AUTOSAR-aligned systems, APP4MC, have become a common trend in the automotive industry. DSMLs allow for easily capturing the functional and non-functional requirements of the system without needing to master low level details of the programming model or the processor architecture. Unfortunately, current DSMLs do not support the parallel programming models, like OpenMP and CUDA, that are used to exploit parallel heterogeneous architectures featuring acceleration devices such as GPUs and FPGAs required. These architectures are however essential to cope with the performance needs of ADAS. This exposes a gap between the DSMLs used by automotive designers to enhance software productivity and leverage verification and validation processes, and the parallel processor architectures used in this domain. This paper presents a complete framework to safely exploit the inherent parallelism exposed by the AMALTHEA system description, supported in APP4MC, by: (1) automatically transforming the high-level design into the OpenMP parallel programming model targeting both host and accelerator parallelism, and (2) using compiler analysis techniques to prove the correctness of the model transformed to OpenMP code. The paper contributes also with (3) an analysis of the parallel execution model allowed by the AMALTHEA DSML and that of OpenMP, and (4) a performance plus productivity evaluation of the proposed framework on real automotive systems executed on an embedded GPU-based processor architecture. Adrian Munera, Sara Royuela, Michael Pressler, Harald Mackamul, Dirk Ziegenbein, Eduardo Quiñones |
J. Syst. Archit. | 1 |
| 2023 | Framework for the Analysis and Configuration of Real-Time OpenMP ApplicationsabstractHigh-performance cyber-physical applications impose several requirements with respect to performance, functional correctness and non-functional aspects. Nowadays, the design of these systems usually follows a model-driven approach, where models generate executable applications, usually with an automated approach. As these applications might execute in different parallel environments, their behavior becomes very hard to predict, and making the verification of non-functional requirements complicated. In this regard, it is crucial to analyse and understand the impact that the mapping and scheduling of computation have on the real-time response of the applications. In fact, different strategies in these steps of the parallel orchestration may produce significantly different interference, leading to different timing behaviour.Tuning the application parameters and the system configuration proves to be one of the most fitting solutions. The design space can however be very cumbersome for a developer to test manually all combinations of application and system configurations. This paper presents a methodology and a toolset to profile, analyse, and configure the timing behaviour of high-performance cyber-physical applications and the target platforms. The methodology leverages on the possibility of generating a task dependency graph representing the parallel computation to evaluate, through measurements, different mapping configurations and select the one that minimizes response time. Tiago Carvalho 0001, Luís Miguel Pinho, Sara Royuela, Adrian Munera, Eduardo Quiñones |
INDIN | 5 |
| 2020 | Towards a Qualifiable OpenMP Framework for Embedded SystemsabstractOpenMP is a very convenient programming model for critical real-time parallel applications due to its powerful tasking model and its proven time predictability. However, current implementations are not suitable for critical environments based on the intensive use of dynamically allocated memory needed to efficiently manage the parallel execution. This jeopardizes the qualification processes needed to ensure that the integrated software stack is compliant with system requirements. This paper proposes a novel OpenMP framework that statically allocates the data structures needed to efficiently manage the parallel execution of OpenMP tasks. Our framework is composed of a compiler that captures the environment of the OpenMP tasks instantiated along the parallel execution and bounds the exposed parallelism, and a runtime implementing a lazy task creation policy that significantly reduces the runtime memory requirements, whilst exploiting parallelism efficiently. The evaluation shows that our tool achieves the same performance as current OpenMP implementations, while bounds and drastically reduces the dynamic memory requirements at run-time. Adrian Munera, Sara Royuela, Eduardo Quiñones |
DATE | 1 |
| 2020 | Experiences on the characterization of parallel applications in embedded systems with Extrae/ParaverabstractCutting-edge functionalities in embedded systems require the use of parallel architectures to meet their performance requirements. This imposes the introduction of a new layer in the software stacks of embedded systems: the parallel programming model. Unfortunately, the tools used to analyze embedded systems fall short to characterize the performance of parallel applications at a parallel programming model level, and correlate this with information about non-functional requirements such as real-time, energy, memory usage, etc. HPC tools, like Extrae, are designed with that level of abstraction in mind, but their main focus is on performance evaluation. Overall, providing insightful information about the performance of parallel embedded applications at the parallel programming model level, and relate it to the non-functional requirements, is of paramount importance to fully exploit the performance capabilities of parallel embedded architectures. Adrian Munera, Sara Royuela, Germán Llort, Estanislao Mercadal, Franck Wartel, Eduardo Quiñones |
ICPP | 1 |