EDBT 2026 Demo / reviewers in the wild / expert
Eduardo César
dblp:61/1918 · also Eduardo César Galobardes
· DBLP profile ↗
30ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-9729-8557ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 4 first-author · 4 since 2021Computer networks · 1Security and privacy · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automatic tuning based on hardware performance counters and machine learningabstractThis paper presents a Machine Learning (ML) methodology for automatically tuning parallel applications in heterogeneous High Performance Computing (HPC) environments using Hardware Performance Counters (HwPCs). The methodology addresses three critical challenges: counter quantity versus accessibility tradeoff, data interpretation complexity, and dynamic optimization needs. The introduced ensemble-based methodology automatically identifies minimal yet informative HwPC sets for code region identification and tuning parameter optimization. Experimental validation demonstrates high accuracy in predicting optimal thread allocation ( > 0.90 K-fold accuracy) and thread affinity ( > 0.95 accuracy) while requiring only 4–6 HwPCs. Compared to search-based methods like OpenTuner, the methodology achieves competitive performance with dramatically reduced optimization time. The architecture-agnostic design enables consistent performance across CPU and GPU platforms. These results establish a foundation for efficient, portable, automatic, and scalable tuning of parallel applications. Suren Harutyunyan Gevorgyan, Eduardo César, Anna Sikora, Jiri Filipovic, Jordi Alcaraz |
Future Gener. Comput. Syst. | 2 |
| 2026 | Towards analysis and refinement of auto-tuning spacesabstractSource code-level auto-tuning enables applications to adapt their implementation to maintain peak performance under varying execution environments (i. e.hardware, input, or application settings). However, the performance of the auto-tuned code is inherently tied to the design of the tuning space (the space of possible changes to the code). An ideal tuning space must include configurations diverse enough to ensure high performance across all targeted environments while simultaneously eliminating redundant or inefficient regions that slow the tuning space search process. Traditional research has focused primarily on identifying optimization opportunities in the code and on efficient tuning space search. However, there is no rigorous methodology or tool supporting analysis and refinement of the tuning spaces, allowing for the addition of configurations that perform well in an unseen environment or the removal of configurations that perform poorly in any realistic environment. In this short communication, we argue that hardware performance counters should be used to analyze tuning spaces, and that such an analysis would allow programmers to refine the tuning spaces by adding configurations that unlock additional performance in unseen environments and removing those unlikely to produce efficient code in any realistic environment. While our primary goal is to introduce this research question and foster discussion, we also present a preliminary methodology for tuning-space analysis. We validate our approach through a case study using a GPU implementation of an N-body simulation. Our results demonstrate that the proposed analysis can detect the weaknesses of a tuning space: based on its outcomes, we refined the tuning space, improving the average configuration performance 3 . 3 × , and the best-performing configuration by 2 − 18 % . Jiri Filipovic, Suren Harutyunyan Gevorgyan, Eduardo César, Anna Sikora |
Parallel Comput. | 3 |
| 2025 | Hierarchical Dynamic Multilevel Graph Partitioning for Load Balancing in Distributed Agent-Based SimulationsabstractApplications handling massive graphs deployed within distributed High-Performance Computing (HPC) systems require careful allocation of vertices across processing elements (PEs) to maximize the utilization of the available resources. This allocation should minimize the number of edges cut between PEs while ensuring a balanced workload. Different graph partitioning tools are available to address this issue, providing both static and dynamic methods for efficiently distributing the application graph across the system.Among existing partitioners, multilevel graph partitioning (MGP) approaches produce high-quality partitions with even workload distribution across PEs while minimizing inter-process communication. However, state-of-the-art MGP frameworks such as Zoltan and ParMETIS often struggle with real-world networks. Although ParHIP provides better support for these cases, it lacks dynamic workload balancing, making it unsuitable for dynamic graphs.This paper presents a novel methodology for a distributed, hierarchical, and dynamic multilevel graph partitioning (HDMGP) framework designed for load-balancing large-scale simulations handling real-world graphs. The HDMGP framework is validated through a proof of concept implementation using ParHIP as a baseline. Initial tests demonstrate that the tool maintains repartition quality comparable to the baseline MGP while achieving a repartitioning time up to 8,8 times faster than recomputing the entire graph partition. Cristina Quesada Peralta, Eduardo César, Andreu Moreno Vendrell, Anna Sikora |
SBAC-PAD | 2 |
| 2024 | Efficient Code Region Characterization Through Automatic Performance Counters Reduction Using Machine Learning TechniquesabstractAbstract Leveraging hardware performance counters provides valuable insights into system resource utilization, aiding performance analysis and tuning for parallel applications. The available counters vary with architecture and are collected at execution time. Their abundance and the limited number of registers for measurement make gathering laborious and costly. Efficient characterization of parallel regions necessitates a dimension reduction strategy. While recent efforts have focused on manually reducing the number of counters for specific architectures, this paper introduces a novel approach: an automatic dimension reduction technique for efficiently characterizing parallel code regions across diverse architectures. The methodology is based on Machine Learning ensembles because of their precision and ability at capturing different relationships between the input features and the target variables. Evaluation results show that ensembles can successfully reduce the number of hardware performance counters that characterize a code region. We validate our approach on CPUs using a comprehensive dataset of OpenMP regions, showing that any region can be accurately characterized by 8 relevant hardware performance counters. In addition, we also apply the proposed methodology on GPUs using a reduced set of kernels, demonstrating its effectiveness across various hardware configurations and workloads. Suren Harutyunyan Gevorgyan, Eduardo César, Anna Sikora, Jiri Filipovic, Akash Dutta, Ali Jannesari, Jordi Alcaraz |
Euro-Par (1) | 2 |
| 2023 | Performance Optimization using Multimodal Modeling and Heterogeneous GNNabstractGrowing heterogeneity and configurability in HPC architectures has made auto-tuning applications and runtime parameters on these systems very complex. Users are presented with a multitude of options to configure parameters. In addition to application specific solutions, a common approach is to use general purpose search strategies, which often might not identify the best configurations or their time to convergence is a significant barrier. There is, thus, a need for a general purpose and efficient tuning approach that can be easily scaled and adapted to various tuning tasks. We propose a technique for tuning parallel code regions that is general enough to be adapted to multiple tasks. In this paper, we analyze IR-based programming models to make task-specific performance optimizations. To this end, we propose the Multimodal Graph Neural Network and Autoencoder (MGA) tuner, a multimodal deep learning based approach that adapts Heterogeneous Graph Neural Networks and Denoising Autoencoders for modeling IR-based code representations that serve as separate modalities. This approach is used as part of our pipeline to model a syntax, semantics, and structure-aware IR-based code representation for tuning parallel code regions/kernels. We extensively experiment on OpenMP and OpenCL code regions/kernels obtained from PolyBench, Rodinia, STREAM, DataRaceBench, AMD SDK, NPB, NVIDIA SDK, Parboil, SHOC, LULESH, XSBench, RSBench, miniFE, miniAMR, and Quicksilver benchmarks and applications. We apply our multimodal learning techniques to the tasks of (i) optimizing the number of threads, scheduling policy and chunk size in OpenMP loops and, (ii) identifying the best device for heterogeneous device mapping of OpenCL kernels. Our experiments show that this multimodal learning based approach outperforms the state-of-the-art in almost all experiments. Akash Dutta, Jordi Alcaraz, Ali TehraniJamsaz, Eduardo César, Anna Sikora, Ali Jannesari |
HPDC | 4 |
| 2021 | Building representative and balanced datasets of OpenMP parallel regionsabstractIncorporating machine learning into automatic performance analysis and tuning tools is a promising path to tackle the increasing heterogeneity of current HPC applications. However, this introduces the need for generating balanced and representative datasets of parallel applications' executions. This work proposes a methodology for building datasets of OpenMP parallel code regions patterns. It allows for determining whether a given code region covers a unique part of the pattern input space not covered by the patterns already included in the dataset. The proposed methodology uses hardware performance counters to represent the execution of the region, which is referred to as the region signature for a given number of cores. Then, a complete representation of the region is built by joining the signatures for every different thread configuration in the system. Next, correlation analysis is performed between this representation and the representation of all the patterns already in the training set. Finally, if this correlation is below a given threshold, the region is considered to cover a unique part of the pattern input space and is subsequently added to the dataset. For validating this methodology, an example dataset, obtained from well known benchmarks, has been used to train a carefully designed neural network model to demonstrate that it is able to classify different patterns of OpenMP parallel regions. Jordi Alcaraz, Steven Sleder, Ali TehraniJamsaz, Anna Sikora, Ali Jannesari, Joan Sorribes, Eduardo César |
PDP | 7 |
| 2019 | Hardware Counters' Space Reduction for Code Region Characterization
Jordi Alcaraz, Anna Sikora, Eduardo César |
Euro-Par | 3 |
| 2019 | Designing a benchmark for the performance evaluation of agent-based simulation applications on HPCabstractAgent-based modeling and simulation (ABMS) is a class of computational models for simulating the actions and interactions of autonomous agents with the goal of assessing their effects on a system as a whole. Several frameworks for generating parallel ABMS applications have been developed taking advantage of their common characteristics, but there is a lack of a general benchmark for comparing the performance of the generated applications. We propose and design a benchmark that takes into consideration the most common characteristics of this type of applications and includes parameters for influencing their relevant performance aspects. We provide an initial implementation of the benchmark for FLAME, FLAME GPU, Repast HPC and EcoLab, some of the most popular parallel ABMS platforms, and use it for comparing the applications generated by these platforms. The obtained results are mostly in agreement with previous studies, but the designed and implemented specification has allowed for testing a wider set of aspects, such as the number of interacting agents, the amount of interchanged data or the evolution of the workload and obtaining more reliable results. Andreu Moreno, Juan J. Rodríguez, Daniel Beltrán, Anna Sikora, Josep Jorba 0001, Eduardo César |
J. Supercomput. | 6 |
| 2018 | Evaluating a formal methodology for dynamic tuning of large-scale parallel applicationsabstractSummary Large‐scale parallel applications performance is usually far from the expected. Dynamic tuning is a powerful technique that helps to improve the performance of parallel applications. To bring this technique to large‐scale computers, this work presents a model that enables decentralized dynamic tuning of large‐scale parallel applications. In this model, applications are decomposed into disjoint subsets of tasks that can be tuned individually but also abstracted to obtain a global view of the parallel application. The proposed model has been designed as a hierarchical tuning network of distributed analysis modules and implemented in the form of ELASTIC, an environment for large‐scale dynamic tuning. Using ELASTIC an experimental evaluation has been conducted over a synthetic large‐scale parallel application and a real agent‐based parallel application. The results show that the proposed model, embodied in ELASTIC, is able to scale to meet the demands of dynamic tuning over thousands of processes, while effectively improving the performance of large‐scale applications. Andrea Martínez, Anna Sikora, Eduardo César, Joan Sorribes |
Concurr. Comput. Pract. Exp. | 3 |
| 2018 | A multi-aspect online tuning framework for HPC applications
Michael Gerndt, Siegfried Benkner, Eduardo César, Carmen B. Navarrete, Enes Bajrovic, Jirí Dokulil, Carla Guillén, Robert Mijakovic, Anna Sikora |
Softw. Qual. J. | 3 |
| 2017 | Colorectal tumour simulation using agent based modelling and high performance computingabstract450,000 European citizens are diagnosed every year with colorectal cancer (CRC) and more than 230,000 succumb to the disease annually. For this reason, significant resources are dedicated to the identification of more effective therapies for this disease. However, classical assessment techniques for these treatments are slow and costly. Consequently, systems biology researchers at the Royal College of Surgeons in Ireland (RCSI) are developing computational agent-based models simulating tumour growth and treatment responses with the objective of speeding up the therapeutic development process while, at the same time, producing a tool for adapting treatments to patient-specific characteristics. However, the model complexity and the high number of agents to be simulated require a thorough optimisation of the process in order to execute realistic simulations of tumour growth on currently available platforms. We propose to apply the most advanced HPC techniques to achieve the efficient and realistic simulation of a virtual tissue model that mimics tumour growth or regression in space and time. These techniques combine extensions of the previously developed agent-based simulation software platform (FLAME) with autotuning capabilities and optimisation strategies for the current tumour model. Development of such a platform could advance the development of novel therapeutic approaches for the treatment of CRC which can also be applied other solid tumours. Guiyeom Kang, Claudio Márquez, Ana Barat, Annette T. Byrne, Jochen H. M. Prehn, Joan Sorribes, Eduardo César |
Future Gener. Comput. Syst. | 7 |
| 2017 | Introducing computational thinking, parallel programming and performance engineering in interdisciplinary studies
Eduardo César, Ana Cortés, Antonio Espinosa 0001, Tomàs Margalef, Juan C. Moure, Anna Sikora, Remo Suppi |
J. Parallel Distributed Comput. | 1 |
| 2017 | HeDPM: load balancing of linear pipeline applications on heterogeneous systemsabstractThis work presents a new algorithm, called Heterogeneous Dynamic Pipeline Mapping, that allows for dynamically improving the performance of pipeline applications running on heterogeneous systems. It is aimed at balancing the application load by determining the best replication (of slow stages) and gathering (of fast stages) combination taking into account processors computation and communication capacities. In addition, the algorithm has been designed with the requirement of keeping complexity low to allow its usage in a dynamic tuning tool. For this reason, it uses an analytical performance model of pipeline applications that addresses hardware heterogeneity and which depends on parameters that can be known in advance or measured at run-time. A wide experimentation is presented, including the comparison with the optimal brute force algorithm, a general comparison with the Binary Search Closest algorithm, and an application example with the Ferret pipeline included in the PARSEC benchmark suite. Results, matching those of the best existing algorithms, show significant performance improvements with lower complexity ( $$O(N^3$$ ), where N is the number of pipeline stages). Andreu Moreno, Anna Sikora, Eduardo César, Joan Sorribes, Tomàs Margalef |
J. Supercomput. | 3 |
| 2014 | Dynamic tuning of the workload partition factor and the resource utilization in data-intensive applications
Claudia Rosas, Anna Sikora, Josep Jorba 0001, Andreu Moreno, Antonio Espinosa 0001, Eduardo César |
Future Gener. Comput. Syst. | 6 |
| 2013 | Increasing Automated Vulnerability Assessment Accuracy on Cloud and Grid Middleware
Jairo Serrano, Eduardo César, Elisa Heymann, Barton P. Miller |
ISPEC | 2 |
| 2013 | Methodology for MPI applications autotuningabstractThis paper proposes a methodology designed to tackle the most common problems of MPI parallel programs. By developing a methodology that applies simple steps in a systematic way, we expect to obtain the basis for a successful autotuning approach of MPI applications based on measurements taken from their own execution. As part of the Au-toTune project, our work is ultimately aimed at extending Periscope to apply automatic tuning to parallel applications and thus provide a straightforward way of tuning MPI parallel codes. Experimental tests demonstrate that this methodology could lead to significant performance improvements. Antonio Pimenta, Eduardo César, Anna Sikora |
EuroMPI | 2 |
| 2012 | Hierarchical MATE's approach for dynamic performance tuning of large-scale parallel applicationsabstractCurrently, performance analysis support tools are required to exploit the potential performance of large-scale computers. However, in this context, scalability becomes a major problem for this kind of tools. Nowadays, there are automatic performance analysis tools, such as Scalasca [1] or Periscope [2], capable of scaling and looking for performance problems of parallel applications. Nevertheless, if the behaviour of a parallel application varies during the execution according to the data evolution, then dynamic analysis and tuning of the application during its execution, such as that performed by MATE [3] tool, is necessary. Andrea Martínez, Anna Sikora, Eduardo César, Joan Sorribes |
IPCCC | 3 |
| 2012 | Load balancing in homogeneous pipeline based applications
Andreu Moreno, Eduardo César, Andreu Guevara, Joan Sorribes, Tomàs Margalef |
Parallel Comput. | 2 |
| 2011 | Workload Balancing Methodology for Data-Intensive Applications with Divisible LoadabstractData-intensive applications are those that explore, query, analyze, and, in general, process very large data sets. Generally in High Performance Computing (HPC), the main performance problem associated to these applications is the load unbalance or inefficient resources utilization. This paper proposes a methodology for improving performance of data-intensive applications based on performing multiple data partitions prior to the execution, and ordering the data chunks according to their processing times during the application execution. As a first step, we consider that a single execution includes multiple related explorations on the same data set. Consequently, we propose to monitor the processing of each exploration and use the data gathered to dynamically tune the performance of the application. The tuning parameters included in the methodology are the partition factor of the data set, the distribution of these data chunks, and the number of processing nodes to be used by the application. The methodology has been initially tested using the well-known bioinformatics tool BLAST, obtaining encouraging results (up to a 40% of improvement). Claudia Rosas, Anna Sikora, Josep Jorba 0001, Eduardo César |
SBAC-PAD | 4 |
| 2010 | A Performance Tuning Strategy for Complex Parallel ApplicationabstractDefining performance models associated with the application structure has been proven a useful strategy for implementing dynamic tuning tools. However, for extending this strategy to more complex applications (those composed by different structures) it must integrate a policy for the distribution of the resources among the different application components. Consequently, we propose to take advantage of the knowledge of these models and combine them with a resource management policy for obtaining a global model. In this sense, this work constitutes the ongoing effort in the development of performance models for dynamic tuning. Jose Alexander Guevara, Eduardo César, Joan Sorribes, Andreu Moreno, Tomàs Margalef, Emilio Luque |
PDP | 2 |
| 2009 | Task distribution using factoring load balancing in Master-Worker applications
Andreu Moreno, Eduardo César, Joan Sorribes, Tomàs Margalef, Emilio Luque |
Inf. Process. Lett. | 2 |
| 2008 | Dynamic Pipeline Mapping (DPM)
Andreu Moreno, Eduardo César, Andreu Guevara, Joan Sorribes, Tomàs Margalef, Emilio Luque |
Euro-Par | 2 |
| 2008 | Performance Model for Parallel Mathematical Libraries Based on Historical Knowledgebase
Ihab Salawdeh, Eduardo César, Anna Sikora, Tomàs Margalef, Emilio Luque |
Euro-Par | 2 |
| 2006 | Modeling Master/Worker applications for automatic performance tuning
Eduardo César, Andreu Moreno, Joan Sorribes, Emilio Luque |
Parallel Comput. | 1 |
| 2005 | Modeling Pipeline Applications in POETRIES
Eduardo César, Joan Sorribes, Emilio Luque |
Euro-Par | 1 |
| 2005 | Automatic Tuning of Master/Worker Applications
Anna Sikora, Eduardo César, Paola Caymes-Scutari, Tomàs Margalef, Joan Sorribes, Emilio Luque |
Euro-Par | 2 |
| 2004 | Modeling Master-Worker Applications in POETRIESabstractParallel/distributed application development is a very difficult task for non-expert programmers, and therefore support tools are needed for all phases of this kind of application development cycle. This means that developing applications using predefined programming structures (frameworks) should be easier than doing it from scratch. We propose to take advantage of the knowledge about the structure of the application in order to develop a dynamic and automatic tuning tool. In this sense, we have designed POETRIES, which is a dynamic performance tuning tool based on the idea that a performance model could be associated to the high-level structure of the application. This way, the tool could efficiently make better tuning decisions. Specifically, we focus this work on the definition of the performance model associated to applications developed with the master-worker framework. Eduardo César, José G. Mesa, Joan Sorribes, Emilio Luque |
HIPS | 1 |
| 2003 | POETRIES: Performance Oriented Environment for Transparent Resource-Management, Implementing End-User Parallel/Distributed Applications
Eduardo César, José G. Mesa, Joan Sorribes, Emilio Luque |
Euro-Par | 1 |
| 2001 | Dynamic Performance Tuning Environment
Anna Sikora, Eduardo César, Tomàs Margalef, Joan Sorribes, Emilio Luque |
Euro-Par | 2 |
| 1996 | Parallel systems development in education: a guided methodabstractArticle Free Access Share on Parallel systems development in education: a guided method Authors: E. Luque Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, Spain Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, SpainView Profile , J. Sorribes Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, Spain Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, SpainView Profile , R. Suppi Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, Spain Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, SpainView Profile , E. Cesar Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, Spain Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, SpainView Profile , J. L. Falguera Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, Spain Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, SpainView Profile , M. Serrano Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, Spain Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, SpainView Profile Authors Info & Claims ITiCSE '96: Proceedings of the 1st conference on Integrating technology into computer science educationJune 1996 Pages 156–158https://doi.org/10.1145/237466.237629Online:01 January 1996Publication History 2citation169DownloadsMetricsTotal Citations2Total Downloads169Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Emilio Luque, Joan Sorribes, Remo Suppi, Eduardo César, J. Falguera, Massimo Serranó |
ITiCSE | 4 |