EDBT 2026 Demo / reviewers in the wild / expert
Hiago Rocha
dblp:274/6042 · also Hiago Mayk G. de A. Rocha
· DBLP profile ↗
10ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-0827-0131ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Assuming the best: Towards a reliable protocol for resource usage prediction for high-performance computing based on machine learningabstractIn High-Performance Computing (HPC) systems, multiple processes simultaneously consume resources such as CPU time, memory, and electrical power, among others. Accurately predicting the resource consumption of a process based on its execution parameters enables more efficient resource allocation, ultimately improving the overall performance of the HPC system. While many studies have explored this topic, fewer explicitly examine the underlying assumptions of their approaches. This work contributes to filling that gap by proposing, experimenting with, and discussing a protocol to approach this problem, covering from the collection of processes footprint data to the experimental evaluation of Machine Learning models based on such data. The reported results of the assessment of this protocol in a case study of the RAxML bioinformatics application on a real supercomputer highlight not only its effectiveness ( R 2 values greater than 0.9 were achieved in most tests) but also the reasonableness of the assumptions considered. Alexandre H. L. Porto, Micaella Coelho, Hiago Rocha, Carla Osthoff, Kary A. C. S. Ocaña, Douglas de O. Cardoso |
Future Gener. Comput. Syst. | 3 |
| 2025 | Integration framework for online thread throttling with thread and page mapping on NUMA systems
Janaina Schwarzrock, Hiago Rocha, Arthur Francisco Lorenzon, Samuel Xavier de Souza, Antonio Carlos Schneider Beck |
J. Parallel Distributed Comput. | 2 |
| 2024 | Allok: a machine learning approach for efficient graph execution on CPU-GPU clusters
Marcelo K. Moori, Hiago Rocha, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
J. Supercomput. | 2 |
| 2023 | Automatic CPU-GPU Allocation for Graph ExecutionabstractAlthough advances in modern GPUs have accelerated the execution of heavy data processing applications, speeding up graph processing on these systems is not a trivial task: graph applications are characterized by their high volume of irregular memory access that varies with the graph structure so that they do not reach their peak performance when executing on GPUs in many times. In these cases, the CPU execution is more suitable. Given that graph structures can be identified through high-level metrics (e.g., diameter and average clustering coefficient), they may assist the designer in deciding where to execute a given input graph (GPU or CPU). Based on that, in this work, we propose GraCo: a graph processing framework to help the decision-making on where to process a batch of graph applications. Whenever a new batch is submitted to the target HPC system, GraCo decides the best machine to execute each application based only on the available high-level features, precluding any additional applications' execution. Our experimental results comparing GraCo with three other strategies executed on an HPC system comprised of 4 CPUs and 3 GPUs showed that GraCo outperforms the other strategies by at least 34.94×, 13.59×, and 492.31× in total execution time, energy, and energy-delay product. Marcelo K. Moori, Hiago Rocha, Matheus A. Silva, Janaina Schwarzrock, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
PDP | 2 |
| 2023 | Improving the efficiency of graph algorithm executions on high-performance computingabstractSummary The growing need for extracting information from large graphs has been pushing the development of parallel graph algorithms. However, the highly irregular structure of the real‐world graphs limits the performance and energy improvements of graph applications. In this paper, we show that, in most cases, using all the available cores of the multiprocessor is not the best option in terms of the aforementioned non‐functional requirements. Based on that, we proposeGraphKat, a framework that enables the simultaneous processing of several algorithms/graphs instead of executing them serially (i.e., one after another), increasing efficiency in terms of performance and energy.GraphKatworks in two steps: (i) it characterizes the graph applications with a specific number of threads based on their efficiency levels; and (ii) it defines the execution order of all graph applications in the target system. Experimental results on three multicore processors (Intel and AMD) show thatGraphKatimproves the overall system's efficiency related to performance (up to ) and energy‐saving (up to 245.21), and reduces the graph applications' execution time (up to ) and energy consumption (up to 6.64) compared to the default execution of parallel applications on HPC systems. Marcelo K. Moori, Hiago Rocha, Janaina Schwarzrock, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
Concurr. Comput. Pract. Exp. | 2 |
| 2023 | Smart resource allocation of concurrent execution of parallel applicationsabstractAbstract Thread‐level parallelism (TLP) has been widely exploited to optimize computational resource usage in high‐performance systems. However, as many applications do not scale as the number of threads increase, resources will be wasted when the application executes with the maximum possible number of threads (i.e., the default execution) rather than fewer threads (thread throttling) that may use the resources more efficiently. Hence, instead of executing only one application with as many threads as possible, one can run more applications simultaneously by applying thread throttling to each one. The primary outcome of this strategy is a significant reduction in the total execution time and energy consumption when the system needs to execute a list of applications. Given that, we propose a smart resource allocation (SRA) for concurrent parallel application execution. It automatically finds the ideal degree of TLP for each application and guides the simultaneous parallel applications execution. When running 25 well‐known benchmarks on three multicore systems and comparing SRA to state‐of‐the‐art strategies (e.g., Batch, Equal policy, and Scalability), SRA improves the EDP by 87.4% over the Batch strategy; 75.5% over the Equal policy; and 38.8% over the scalability strategy. Vinicius S. da Silva, Angelo Gaspar Diniz Nogueira, Everton Camargo de Lima, Hiago Rocha, Matheus S. Serpa, Marcelo Caggiani Luizelli, Fábio D. Rossi, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | Using machine learning to optimize graph execution on NUMA machinesabstractThis paper proposes PredG, a Machine Learning framework to enhance the graph processing performance by finding the ideal thread and data mapping on NUMA systems. PredG is agnostic to the input graph: it uses the available graphs' features to train an ANN to perform predictions as new graphs arrive - without any application execution after being trained. When evaluating PredG over representative graphs and algorithms on three NUMA systems, its solutions are up to 41% faster than the Linux OS Default and the Best Static - on average 2% far from the Oracle -, and it presents lower energy consumption. Hiago Rocha, Janaina Schwarzrock, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
DAC | 1 |
| 2021 | Boosting Graph Analytics by Tuning Threads and Data Affinity on NUMA SystemsabstractThe execution of large real-world graphs, such as web searches and social networks, has been boosting by modern HPC systems. However, their irregular communication patterns and poor data locality impose many challenges, mainly when executed on NUMA systems. As we show in this paper, there is no one-fits-all configuration for threads/data mapping, and the best combination will vary according to the NUMA system, graph algorithm, and input graph at hand. Based on that, we propose Graphith: a framework that automatically enhances graph processing performance by adapting its execution considering the variables mentioned above. Graphith also goes one step further and improves the existing policies: it uses a Genetic Algorithm to fine-tune the thread-to-core allocation combined with data mapping policies. With that, Graphith improves in 21%, on average, the default execution, and is, on average, 7% better than the best possible combination of standard policies. Hiago Rocha, Janaina Schwarzrock, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
PDP | 1 |
| 2020 | A Machine Learning Approach for Reliability-Aware Application Mapping for Heterogeneous MulticoresabstractWe propose a transparent and runtime methodology to increase the system's Mean Workload to Failure (MWTF) in heterogeneous multicore processors. For that, we leverage an Artificial Neural Network that makes online predictions of the core's Architectural Vulnerability Factor (AVF), which allows for reliability-aware application-to-core mappings. We experiment with different configurations of RISC-V cores and compare the MWTF of prediction-based mappings against the optimal oracle, showing that our proposed model provides MWTF as close as 5.6% to the oracle. We also compare homogeneous and heterogeneous multicores, showing that heterogeneity provides room for increasing the MWTF in up to 19.4%. Rafael Billig Tonetto, Hiago Rocha, Gabriel L. Nazar, Antonio Carlos Schneider Beck |
DAC | 2 |
| 2020 | A Reliability-Oriented Machine Learning Strategy for Heterogeneous Multicore Application MappingabstractWe propose a methodology to transparently estimate near-optimal application mappings aiming at increasing the Mean Workload to Failure (MWTF) in heterogeneous multicore processors. For that, we leverage an Artificial Neural Network (ANN) capable of estimating the vulnerability factor of RISC-V cores at runtime, which allows for efficient and dynamic application-to-core mappings targeting better MWTF and MWTF/energy tradeoffs. Results show that our ANN-based mapping yields very close-to-optimal solutions, with a difference in MWTF of only 3% when compared to the optimal mapping. When compared to a homogeneous architecture composed of only big cores, heterogeneous architectures may provide improvement in MWTF of up to 20.5% while impacting 12.2% on performance. Rafael Billig Tonetto, Hiago Rocha, Bruno Zatt, Antonio Carlos Schneider Beck, Gabriel L. Nazar |
ISCAS | 2 |