Anna Sikora

dblp:87/6629 · also Anna Morajko · DBLP profile ↗
← Back
32ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0003-0090-4109ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 6 first-author · 5 since 2021Software engineering, systems software and programming languages · 2Computer networks · 1
YearPublicationVenuePosition
2026 Automatic tuning based on hardware performance counters and machine learning
abstract
This paper presents a Machine Learning (ML) methodology for automatically tuning parallel applications in heterogeneous High Performance Computing (HPC) environments using Hardware Performance Counters (HwPCs). The methodology addresses three critical challenges: counter quantity versus accessibility tradeoff, data interpretation complexity, and dynamic optimization needs. The introduced ensemble-based methodology automatically identifies minimal yet informative HwPC sets for code region identification and tuning parameter optimization. Experimental validation demonstrates high accuracy in predicting optimal thread allocation ( > 0.90 K-fold accuracy) and thread affinity ( > 0.95 accuracy) while requiring only 4–6 HwPCs. Compared to search-based methods like OpenTuner, the methodology achieves competitive performance with dramatically reduced optimization time. The architecture-agnostic design enables consistent performance across CPU and GPU platforms. These results establish a foundation for efficient, portable, automatic, and scalable tuning of parallel applications.
Suren Harutyunyan Gevorgyan, Eduardo César, Anna Sikora, Jiri Filipovic, Jordi Alcaraz
Future Gener. Comput. Syst.3
2026 Towards analysis and refinement of auto-tuning spaces
abstract
Source code-level auto-tuning enables applications to adapt their implementation to maintain peak performance under varying execution environments (i. e.hardware, input, or application settings). However, the performance of the auto-tuned code is inherently tied to the design of the tuning space (the space of possible changes to the code). An ideal tuning space must include configurations diverse enough to ensure high performance across all targeted environments while simultaneously eliminating redundant or inefficient regions that slow the tuning space search process. Traditional research has focused primarily on identifying optimization opportunities in the code and on efficient tuning space search. However, there is no rigorous methodology or tool supporting analysis and refinement of the tuning spaces, allowing for the addition of configurations that perform well in an unseen environment or the removal of configurations that perform poorly in any realistic environment. In this short communication, we argue that hardware performance counters should be used to analyze tuning spaces, and that such an analysis would allow programmers to refine the tuning spaces by adding configurations that unlock additional performance in unseen environments and removing those unlikely to produce efficient code in any realistic environment. While our primary goal is to introduce this research question and foster discussion, we also present a preliminary methodology for tuning-space analysis. We validate our approach through a case study using a GPU implementation of an N-body simulation. Our results demonstrate that the proposed analysis can detect the weaknesses of a tuning space: based on its outcomes, we refined the tuning space, improving the average configuration performance 3 . 3 × , and the best-performing configuration by 2 − 18 % .
Jiri Filipovic, Suren Harutyunyan Gevorgyan, Eduardo César, Anna Sikora
Parallel Comput.4
2025 Hierarchical Dynamic Multilevel Graph Partitioning for Load Balancing in Distributed Agent-Based Simulations
abstract
Applications handling massive graphs deployed within distributed High-Performance Computing (HPC) systems require careful allocation of vertices across processing elements (PEs) to maximize the utilization of the available resources. This allocation should minimize the number of edges cut between PEs while ensuring a balanced workload. Different graph partitioning tools are available to address this issue, providing both static and dynamic methods for efficiently distributing the application graph across the system.Among existing partitioners, multilevel graph partitioning (MGP) approaches produce high-quality partitions with even workload distribution across PEs while minimizing inter-process communication. However, state-of-the-art MGP frameworks such as Zoltan and ParMETIS often struggle with real-world networks. Although ParHIP provides better support for these cases, it lacks dynamic workload balancing, making it unsuitable for dynamic graphs.This paper presents a novel methodology for a distributed, hierarchical, and dynamic multilevel graph partitioning (HDMGP) framework designed for load-balancing large-scale simulations handling real-world graphs. The HDMGP framework is validated through a proof of concept implementation using ParHIP as a baseline. Initial tests demonstrate that the tool maintains repartition quality comparable to the baseline MGP while achieving a repartitioning time up to 8,8 times faster than recomputing the entire graph partition.
Cristina Quesada Peralta, Eduardo César, Andreu Moreno Vendrell, Anna Sikora
SBAC-PAD4
2025 Content delivery network solutions for the CMS experiment: The evolution towards HL-LHC
abstract
The Large Hadron Collider at CERN in Geneva is poised for a transformative upgrade, preparing to enhance both its accelerator and particle detectors. This strategic initiative is driven by the tenfold increase in proton-proton collisions anticipated for the forthcoming high-luminosity phase scheduled to start by 2029. The vital role played by the underlying computational infrastructure, the World-Wide LHC Computing Grid, in processing the data generated during these collisions underlines the need for its expansion and adaptation to meet the demands of the new accelerator phase. The provision of these computational resources by the worldwide community remains essential, all within a constant budgetary framework. While technological advancements offer some relief for the expected increase, numerous research and development projects are underway. Their aim is to bring future resources to manageable levels and provide cost-effective solutions to effectively handle the expanding volume of generated data. In the quest for optimized data access and resource utilization, the LHC community is actively investigating Content Delivery Network (CDN) techniques. These techniques serve as a mechanism for the cost-effective deployment of lightweight storage systems that support both, traditional and opportunistic compute resources. Furthermore, they aim to enhance the performance of executing tasks by facilitating the efficient reading of input data via caching content near the end user. A comprehensive study is presented to assess the benefits of implementing data cache solutions for the Compact Muon Solenoid (CMS) experiment. This in-depth examination serves as a use-case study specifically conducted for the Spanish compute facilities, playing a crucial role in supporting CMS activities. Data access patterns and popularity studies suggest that user analysis tasks benefit the most from CDN techniques. Consequently, a data cache has been introduced in the region to acquire a deeper understanding of these effects. In this paper, the details of the implementation of a data cache system in the PIC Tier-1 compute facility are presented. It includes insights into the developed monitoring tools and discusses the positive impact on CPU usage for analysis tasks executed in the region. The study is augmented by simulations of data caches, with the objective of discerning the most optimal requirements in both size and network connectivity for a data cache serving the Spanish region. Additionally, the study delves into the cost benefits associated with deploying such a solution in a production environment. Furthermore, it investigates the potential impact of incorporating this solution into other regions of the CMS computing infrastructure. • XRootD's XCache-based CDNs strategically reduces latency, improving CMS job performance. • Deployed in production, it shows 75% reduction of latency and 30% overall performance improvement. • The effectiveness is demonstrated through data usage analyses and benchmarking of CMS Analysis jobs logs and simulations. • XRootD's XCache provides a cost-effective solution, reducing WAN data transfer and task execution time. • It minimizes WAN use by caching remote data, economically managing data retrieval during task execution.
Carlos Perez Dengra, Josep Flix, Anna Sikora
J. Parallel Distributed Comput.3
2024 Efficient Code Region Characterization Through Automatic Performance Counters Reduction Using Machine Learning Techniques
abstract
Abstract Leveraging hardware performance counters provides valuable insights into system resource utilization, aiding performance analysis and tuning for parallel applications. The available counters vary with architecture and are collected at execution time. Their abundance and the limited number of registers for measurement make gathering laborious and costly. Efficient characterization of parallel regions necessitates a dimension reduction strategy. While recent efforts have focused on manually reducing the number of counters for specific architectures, this paper introduces a novel approach: an automatic dimension reduction technique for efficiently characterizing parallel code regions across diverse architectures. The methodology is based on Machine Learning ensembles because of their precision and ability at capturing different relationships between the input features and the target variables. Evaluation results show that ensembles can successfully reduce the number of hardware performance counters that characterize a code region. We validate our approach on CPUs using a comprehensive dataset of OpenMP regions, showing that any region can be accurately characterized by 8 relevant hardware performance counters. In addition, we also apply the proposed methodology on GPUs using a reduced set of kernels, demonstrating its effectiveness across various hardware configurations and workloads.
Suren Harutyunyan Gevorgyan, Eduardo César, Anna Sikora, Jiri Filipovic, Akash Dutta, Ali Jannesari, Jordi Alcaraz
Euro-Par (1)3
2023 Performance Optimization using Multimodal Modeling and Heterogeneous GNN
abstract
Growing heterogeneity and configurability in HPC architectures has made auto-tuning applications and runtime parameters on these systems very complex. Users are presented with a multitude of options to configure parameters. In addition to application specific solutions, a common approach is to use general purpose search strategies, which often might not identify the best configurations or their time to convergence is a significant barrier. There is, thus, a need for a general purpose and efficient tuning approach that can be easily scaled and adapted to various tuning tasks. We propose a technique for tuning parallel code regions that is general enough to be adapted to multiple tasks. In this paper, we analyze IR-based programming models to make task-specific performance optimizations. To this end, we propose the Multimodal Graph Neural Network and Autoencoder (MGA) tuner, a multimodal deep learning based approach that adapts Heterogeneous Graph Neural Networks and Denoising Autoencoders for modeling IR-based code representations that serve as separate modalities. This approach is used as part of our pipeline to model a syntax, semantics, and structure-aware IR-based code representation for tuning parallel code regions/kernels. We extensively experiment on OpenMP and OpenCL code regions/kernels obtained from PolyBench, Rodinia, STREAM, DataRaceBench, AMD SDK, NPB, NVIDIA SDK, Parboil, SHOC, LULESH, XSBench, RSBench, miniFE, miniAMR, and Quicksilver benchmarks and applications. We apply our multimodal learning techniques to the tasks of (i) optimizing the number of threads, scheduling policy and chunk size in OpenMP loops and, (ii) identifying the best device for heterogeneous device mapping of OpenCL kernels. Our experiments show that this multimodal learning based approach outperforms the state-of-the-art in almost all experiments.
Akash Dutta, Jordi Alcaraz, Ali TehraniJamsaz, Eduardo César, Anna Sikora, Ali Jannesari
HPDC5
2021 Building representative and balanced datasets of OpenMP parallel regions
abstract
Incorporating machine learning into automatic performance analysis and tuning tools is a promising path to tackle the increasing heterogeneity of current HPC applications. However, this introduces the need for generating balanced and representative datasets of parallel applications' executions. This work proposes a methodology for building datasets of OpenMP parallel code regions patterns. It allows for determining whether a given code region covers a unique part of the pattern input space not covered by the patterns already included in the dataset. The proposed methodology uses hardware performance counters to represent the execution of the region, which is referred to as the region signature for a given number of cores. Then, a complete representation of the region is built by joining the signatures for every different thread configuration in the system. Next, correlation analysis is performed between this representation and the representation of all the patterns already in the training set. Finally, if this correlation is below a given threshold, the region is considered to cover a unique part of the pattern input space and is subsequently added to the dataset. For validating this methodology, an example dataset, obtained from well known benchmarks, has been used to train a carefully designed neural network model to demonstrate that it is able to classify different patterns of OpenMP parallel regions.
Jordi Alcaraz, Steven Sleder, Ali TehraniJamsaz, Anna Sikora, Ali Jannesari, Joan Sorribes, Eduardo César
PDP4
2019 Hardware Counters' Space Reduction for Code Region Characterization
Jordi Alcaraz, Anna Sikora, Eduardo César
Euro-Par2
2019 Designing a benchmark for the performance evaluation of agent-based simulation applications on HPC
abstract
Agent-based modeling and simulation (ABMS) is a class of computational models for simulating the actions and interactions of autonomous agents with the goal of assessing their effects on a system as a whole. Several frameworks for generating parallel ABMS applications have been developed taking advantage of their common characteristics, but there is a lack of a general benchmark for comparing the performance of the generated applications. We propose and design a benchmark that takes into consideration the most common characteristics of this type of applications and includes parameters for influencing their relevant performance aspects. We provide an initial implementation of the benchmark for FLAME, FLAME GPU, Repast HPC and EcoLab, some of the most popular parallel ABMS platforms, and use it for comparing the applications generated by these platforms. The obtained results are mostly in agreement with previous studies, but the designed and implemented specification has allowed for testing a wider set of aspects, such as the number of interacting agents, the amount of interchanged data or the evolution of the workload and obtaining more reliable results.
Andreu Moreno, Juan J. Rodríguez, Daniel Beltrán, Anna Sikora, Josep Jorba 0001, Eduardo César
J. Supercomput.4
2018 Evaluating a formal methodology for dynamic tuning of large-scale parallel applications
abstract
Summary Large‐scale parallel applications performance is usually far from the expected. Dynamic tuning is a powerful technique that helps to improve the performance of parallel applications. To bring this technique to large‐scale computers, this work presents a model that enables decentralized dynamic tuning of large‐scale parallel applications. In this model, applications are decomposed into disjoint subsets of tasks that can be tuned individually but also abstracted to obtain a global view of the parallel application. The proposed model has been designed as a hierarchical tuning network of distributed analysis modules and implemented in the form of ELASTIC, an environment for large‐scale dynamic tuning. Using ELASTIC an experimental evaluation has been conducted over a synthetic large‐scale parallel application and a real agent‐based parallel application. The results show that the proposed model, embodied in ELASTIC, is able to scale to meet the demands of dynamic tuning over thousands of processes, while effectively improving the performance of large‐scale applications.
Andrea Martínez, Anna Sikora, Eduardo César, Joan Sorribes
Concurr. Comput. Pract. Exp.2
2018 A multi-aspect online tuning framework for HPC applications
Michael Gerndt, Siegfried Benkner, Eduardo César, Carmen B. Navarrete, Enes Bajrovic, Jirí Dokulil, Carla Guillén, Robert Mijakovic, Anna Sikora
Softw. Qual. J.9
2017 Introducing computational thinking, parallel programming and performance engineering in interdisciplinary studies
Eduardo César, Ana Cortés, Antonio Espinosa 0001, Tomàs Margalef, Juan C. Moure, Anna Sikora, Remo Suppi
J. Parallel Distributed Comput.6
2017 HeDPM: load balancing of linear pipeline applications on heterogeneous systems
abstract
This work presents a new algorithm, called Heterogeneous Dynamic Pipeline Mapping, that allows for dynamically improving the performance of pipeline applications running on heterogeneous systems. It is aimed at balancing the application load by determining the best replication (of slow stages) and gathering (of fast stages) combination taking into account processors computation and communication capacities. In addition, the algorithm has been designed with the requirement of keeping complexity low to allow its usage in a dynamic tuning tool. For this reason, it uses an analytical performance model of pipeline applications that addresses hardware heterogeneity and which depends on parameters that can be known in advance or measured at run-time. A wide experimentation is presented, including the comparison with the optimal brute force algorithm, a general comparison with the Binary Search Closest algorithm, and an application example with the Ferret pipeline included in the PARSEC benchmark suite. Results, matching those of the best existing algorithms, show significant performance improvements with lower complexity ( $$O(N^3$$ ), where N is the number of pipeline stages).
Andreu Moreno, Anna Sikora, Eduardo César, Joan Sorribes, Tomàs Margalef
J. Supercomput.2
2015 Online root-cause performance analysis of parallel applications
Anna Sikora, Tomàs Margalef, Josep Jorba 0001
Parallel Comput.1
2014 Dynamic tuning of the workload partition factor and the resource utilization in data-intensive applications
Claudia Rosas, Anna Sikora, Josep Jorba 0001, Andreu Moreno, Antonio Espinosa 0001, Eduardo César
Future Gener. Comput. Syst.2
2014 GMATE: Dynamic Tuning of Parallel Applications in Grid Environment
Genaro Costa, Anna Sikora, Josep Jorba 0001, Tomàs Margalef
J. Grid Comput.2
2014 Enhancing multi-model forest fire spread prediction by exploiting multi-core parallelism
Carlos Brun, Tomàs Margalef, Ana Cortés, Anna Sikora
J. Supercomput.4
2013 Methodology for MPI applications autotuning
abstract
This paper proposes a methodology designed to tackle the most common problems of MPI parallel programs. By developing a methodology that applies simple steps in a systematic way, we expect to obtain the basis for a successful autotuning approach of MPI applications based on measurements taken from their own execution. As part of the Au-toTune project, our work is ultimately aimed at extending Periscope to apply automatic tuning to parallel applications and thus provide a straightforward way of tuning MPI parallel codes. Experimental tests demonstrate that this methodology could lead to significant performance improvements.
Antonio Pimenta, Eduardo César, Anna Sikora
EuroMPI3
2012 Hierarchical MATE's approach for dynamic performance tuning of large-scale parallel applications
abstract
Currently, performance analysis support tools are required to exploit the potential performance of large-scale computers. However, in this context, scalability becomes a major problem for this kind of tools. Nowadays, there are automatic performance analysis tools, such as Scalasca [1] or Periscope [2], capable of scaling and looking for performance problems of parallel applications. Nevertheless, if the behaviour of a parallel application varies during the execution according to the data evolution, then dynamic analysis and tuning of the application during its execution, such as that performed by MATE [3] tool, is necessary.
Andrea Martínez, Anna Sikora, Eduardo César, Joan Sorribes
IPCCC2
2012 A methodology for transparent knowledge specification in a dynamic tuning environment
abstract
SUMMARY The increasing use of parallel/distributed applications demands a continuous support to take significant advantages from parallel power. This includes the evolution of performance analysis and tuning tools which automatically allows for obtaining a better behavior of the applications. Different approaches and tools have been proposed and they are continuously evolving to cover the requirements and expectations of users. One such tool is MATE (Monitoring Analysis and Tuning Environment), which provides automatic and dynamic tuning for parallel/distributed applications. The knowledge used by MATE to analyze and take decisions is based on performance models which include a set of performance parameters and a set of mathematical expressions modeling the solution of the performance problem. These elements are used by the tuning environment to conduct the monitoring and analysis steps, respectively. The tuning phase depends on the results of the performance analysis. This paper presents a methodology to specify performance models. Each performance model specification can be automatically and transparently translated into a piece of software code encapsulating the knowledge to be straightforwardly included in MATE. Applying this methodology, the user does not have to be involved in the implementation details of MATE, which makes the usage of the tool more transparent. Copyright © 2011 John Wiley & Sons, Ltd.
Paola Caymes-Scutari, Anna Sikora, Tomàs Margalef, Emilio Luque
Softw. Pract. Exp.2
2011 Workload Balancing Methodology for Data-Intensive Applications with Divisible Load
abstract
Data-intensive applications are those that explore, query, analyze, and, in general, process very large data sets. Generally in High Performance Computing (HPC), the main performance problem associated to these applications is the load unbalance or inefficient resources utilization. This paper proposes a methodology for improving performance of data-intensive applications based on performing multiple data partitions prior to the execution, and ordering the data chunks according to their processing times during the application execution. As a first step, we consider that a single execution includes multiple related explorations on the same data set. Consequently, we propose to monitor the processing of each exploration and use the data gathered to dynamically tune the performance of the application. The tuning parameters included in the methodology are the partition factor of the data set, the distribution of these data chunks, and the number of processing nodes to be used by the application. The methodology has been initially tested using the well-known bioinformatics tool BLAST, obtaining encouraging results (up to a 40% of improvement).
Claudia Rosas, Anna Sikora, Josep Jorba 0001, Eduardo César
SBAC-PAD2
2010 Scalable dynamic Monitoring, Analysis and Tuning Environment for parallel applications
Paola Caymes-Scutari, Anna Sikora, Tomàs Margalef, Emilio Luque
J. Parallel Distributed Comput.2
2009 Introduction
Felix Wolf 0001, Andy D. Pimentel, Luiz De Rose, Soonhoi Ha, Thilo Kielmann, Anna Sikora
Euro-Par6
2008 Performance models for dynamic tuning of parallel applications on Computational Grids
abstract
Performance is a main issue in parallel application development. Dynamic tuning is a technique that acts over application parameters to raise execution performance indexes. To perform that, it is necessary to collect measurements, analyze application behavior using a performance model and carry out tuning actions. Computational Grids present proclivity for dynamic changes on their features during application execution. Thus, dynamic tuning tools are indispensable to reach the expected performance indexes on those environments. A particular problem which provokes performance bottlenecks is the load distribution in master/worker applications. This paper addresses the performance modeling of such applications on Computational Grids for the perspective of dynamic tuning. It is inferred that grain size and number of workers are critical parameters to reduce execution time while raising the efficiency of resources usage. A heuristic to dynamically tune granularity and number of workers is proposed. The experimental simulated results of a matrix multiplication application in a heterogeneous Grid environment are shown.
Genaro Costa, Josep Jorba 0001, Anna Sikora, Tomàs Margalef, Emilio Luque
CLUSTER3
2008 On-Line Performance Modeling for MPI Applications
Oleg Morajko, Anna Sikora, Tomàs Margalef, Emilio Luque
Euro-Par2
2008 Performance Model for Parallel Mathematical Libraries Based on Historical Knowledgebase
Ihab Salawdeh, Eduardo César, Anna Sikora, Tomàs Margalef, Emilio Luque
Euro-Par3
2007 Automatic Generation of Dynamic Tuning Techniques
Paola Caymes-Scutari, Anna Sikora, Tomàs Margalef, Emilio Luque
Euro-Par2
2007 MATE: Monitoring, Analysis and Tuning Environment for parallel/distributed applications
abstract
Abstract The main goal of parallel/distributed applications is to solve the considered problem as fast as possible using the available resources. In this context, the application performance becomes a crucial issue. Developers of these applications must optimize them if they are to fulfill the promise of high‐performance computation. To improve performance, developers search for bottlenecks by analyzing application behavior, try to identify performance problems, determine their causes and overcome them by changing the source code of the application. Current approaches require developers to do these tasks manually and imply a high degree of expertise. Therefore, another approach is needed to help developers during the optimization process. This paper presents the dynamic tuning approach that addresses these issues. In this approach, many tasks are automated and the user intervention and required experience may be significantly reduced. An application is monitored, its performance bottlenecks are detected and it is modified automatically during execution, without recompiling or re‐running it. The introduced modifications adapt the application behavior to changing conditions. We present an environment called MATE (Monitoring, Analysis and Tuning Environment) that has been developed to provide dynamic tuning of parallel/distributed applications. We also show practical experiments conducted with MATE to prove its effectiveness and profitability. Copyright © 2006 John Wiley & Sons, Ltd.
Anna Sikora, Paola Caymes-Scutari, Tomàs Margalef, Emilio Luque
Concurr. Comput. Pract. Exp.1
2007 Design and implementation of a dynamic tuning environment
Anna Sikora, Tomàs Margalef, Emilio Luque
J. Parallel Distributed Comput.1
2005 Automatic Tuning of Master/Worker Applications
Anna Sikora, Eduardo César, Paola Caymes-Scutari, Tomàs Margalef, Joan Sorribes, Emilio Luque
Euro-Par1
2004 MATE: Dynamic Performance Tuning Environment
Anna Sikora, Oleg Morajko, Tomàs Margalef, Emilio Luque
Euro-Par1
2001 Dynamic Performance Tuning Environment
Anna Sikora, Eduardo César, Tomàs Margalef, Joan Sorribes, Emilio Luque
Euro-Par1