Murilo Boratto

dblp:02/5996 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-3908-4451ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards a hierarchical approach for autotuning task-based libraries
abstract
Abstract This work proposes a hierarchical approach to reduce the training time of task-based routines by reusing previously obtained autotuning information. This approach has been integrated into a working prototype of Chameleon, a dense linear algebra software whose tile-based routines are executed on the available computational resources by means of a runtime system. The results show that this approach provides a high degree of scalability to the entire self-optimization process, achieving a reduction in training time of up to 80% and an appropriate selection of values for the adjustable parameters.
Jesús Cámara, Javier Cuenca 0001, Murilo Boratto
J. Supercomput.3
2025 DeSAO: A new approach for De Novo Drug using Simulated Annealing Optimization
Rosalvo F. O. Neto, Murilo Boratto, Edilson Beserra de Alencar Filho
Expert Syst. Appl.2
2025 An autotuning approach to select the inter-GPU communication library on heterogeneous systems
abstract
Abstract In this work, an automatic optimisation approach for parallel routines on multi-GPU systems is presented. Several inter-GPU communication libraries (such as CUDA-Aware MPI or NCCL) are used with a set of routines to perform the numerical operations among the GPUs located on the compute nodes. The main objective is the selection of the most appropriate communication library, the number of GPUs to be used and the workload to be distributed among them in order to reduce the cost of data movements, which represent a large percentage of the total execution time. To this end, a hierarchical modelling of the execution time of each routine to be optimised is proposed, combining experimental and theoretical approaches. The results show that near-optimal decisions are taken in all the scenarios analysed.
Jesús Cámara, Javier Cuenca 0001, Victor Galindo, Arturo Vicente, Murilo Boratto
J. Supercomput.5
2023 GPU performance analysis for viscoacoustic wave equations using fast stencil computation from the symbolic specification
Lauê Jesus, Peterson Nogueira, João Speglich, Murilo Boratto
J. Supercomput.4
2022 Parallel signal detection for generalized spatial modulation MIMO systems
abstract
Abstract Generalized Spatial Modulation is a recently developed technique that is designed to enhance the efficiency of transmissions in MIMO Systems. However, the procedure for correctly retrieving the sent signal at the receiving end is quite demanding. Specifically, the computation of the maximum likelihood solution is computationally very expensive. In this paper, we propose a parallel method for the computation of the maximum likelihood solution using the parallel computing library OpenMP. The proposed parallel algorithm computes the maximum likelihood solution faster than the sequential version, and substantially reduces the worst-case computing times.
Víctor M. García 0001, M. Ángeles Simarro, Francisco-Jose Martínez-Zaldívar, Murilo Boratto, Pedro Alonso 0002, Alberto González 0001
J. Supercomput.4
2019 Exploring hybrid parallel systems for probabilistic record linkage
Murilo Boratto, Pedro Alonso 0002, Clícia Pinto, Pedro Melo, Marcos E. Barreto, Spiros C. Denaxas
J. Supercomput.1
2017 Accelerating Docking Simulation Using Multicore and GPU Systems
Everton Mendonça, Marcos E. Barreto, Vinícius Guimarães, Nelci Santos, Samuel Pita, Murilo Boratto
ICCSA (1)6
2017 Automatic tuning to performance modelling of matrix polynomials on multicore and multi-GPU systems
Murilo Boratto, Pedro Alonso 0002, Domingo Giménez, Alexey L. Lastovetsky
J. Supercomput.1
2016 Auto-tuning TRSM with an asynchronous task assignment model on multicore, multi-GPU and coprocessor systems
abstract
The increasing need for computing power today justifies the continuous search for techniques that decrease the time to answer usual computational problems. To take advantage of new hybrid parallel architectures composed by multithreading and multiprocessor hardware, our current efforts involve the design and validation of highly parallel algorithms that efficiently explore the characteristics of such architectures. In this paper, we propose an automatic tuning methodology to easily exploit multicore, multi-GPU and coprocessor systems. We present an optimization of an algorithm for solving triangular systems (TRSM), based on block decomposition and asynchronous task assignment, and discuss some results.
Clícia Pinto, Marcos E. Barreto, Murilo Boratto
AICCSA3
2014 Automatic routine tuning to represent landform attributes on multicore and multi-GPU systems
Murilo Boratto, Pedro Alonso 0002, Domingo Giménez, Marcos E. Barreto
J. Supercomput.1
2014 Solving time-invariant differential matrix Riccati equations using GPGPU computing
Jesús Peinado-Pinilla, Pedro Alonso 0002, Jacinto Javier Ibáñez, Vicente Hernández, Murilo Boratto
J. Supercomput.5
2012 Parallel Algorithm for Landform Attributes Representation on Multicore and Multi-GPU Systems
Murilo Boratto, Pedro Alonso 0002, Carla Ramiro, Marcos E. Barreto, Leandro dos Santos Coelho
ICCSA (1)1
2008 A Threaded Divide and Conquer Symmetric Tridiagonal Eigensolver on Multicore Systems
abstract
The increasing power of computation of modern processors rely on the increasing number of cores per chip. The challenge of software developers is to keep this power with the legacy code. Although commercial and non commercial libraries are improving their codes step by step, there exits probably insurmountable scalability issues for standard programming models due to the fact that using locks to implement synchronisation is inherently a bottleneck. We propose an implementation of the divide and conquer algorithm to compute the eigenpairs of symmetric tridiagonal matrices on multicore systems. We take advantage of the natural parallelism of the method by using pthreads. We avoided as much as possible the negative impact of synchronisation in the performance by overlapping operations of different classes. Furthermore, the unevenly workload distribution of the computational cost of the elemental tasks yields in a speedup even larger than expected.
Antonio M. Vidal, Murilo Boratto, Pedro Alonso 0002
ISPDC2