Ricardo Nobre

dblp:91/11134 · DBLP profile ↗
← Back
20ranked-venue papers
11as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 7 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorArtificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 Thievory: Graph Processing with Multi-GPU Memory Stealing
abstract
Graph workloads have gained widespread attention across diverse applications. These workloads often display large memory footprints while offering significant parallelism opportunities in graph traversals. To exploit this parallelism, the state-of-the-art approaches focus on single-GPU environments and/or out-of-memory processing. However, these approaches are individually insufficient to address the challenges of modern GPU-accelerated graph processing. First, the ever-increasing graph size can span well beyond the available GPU memory. Second, modern computing platforms are de facto multi-GPU, thus bringing a new set of challenges to be taken into account for taming the highly irregular nature of modern large graphs. Novel approaches must be devised to achieve efficient multi-device graph processing by exploiting the fast cross-device interconnects and seamless peer-to-peer communication. This work proposes Thievory – the first out-of-memory graph processing framework that explores novel execution patterns in multi-GPU environments. In particular, we explore the use of GPUs as data buffers, effectively using spare neighbor GPUs’ memory to boost out-of-memory graph traversals, thus increasing resource efficiency in these multi-GPU environments. Moreover, we also analyse a rich set of data transfer optimizations and caching mechanisms. Thievory shows up to 10.5× speedups over existing state-of-the-art approaches when evaluated with a representative set of graphs.
João Brotas, Ricardo Nobre, Aleksandar Ilic
ICPP2
2025 EPIClear: Exploiting Domain-Specific Features for Epistasis Detection Acceleration on Tensor Cores
abstract
High-order epistasis detection is challenging, making it important to efficiently leverage today's supercomputers.The fastest approaches are those relying on binary precision tensorized operations on modern GPUs.This paper presents a novel approach that significantly surpasses the state-of-theart in high-order epistasis detection by leveraging previously unexplored domain-specific features on the genotype distribution patterns in the dataset.It accelerates time-to-solution with a computational step that reduces the volume of data that needs to be processed to count genotypes.The proposed approach achieves 4× higher performance on a A100 GPU than the previously fastest approach when processing balanced genotype distributions.Evaluation on datasets with unbalanced genotype distributions, which is something that is bound to happen in real datasets, results in significantly higher performance.The proposed accelerating scheme exhibits high scalability.Epistasis detection searches on the MeluXina supercomputer with 32 A100 GPUs resulted in a speedup of up to 30× in comparison to a single GPU, and in achieving a performance scaled to sample size of up to 13 Peta SNP combinations per second for the genotype distribution most unfavorable to the proposed accelerating scheme.
Ricardo Nobre, Miguel Graça, Leonel Sousa, Aleksandar Ilic
ICS1
2024 IPU-EpiDet: Identifying Gene Interactions on Massively Parallel Graph-Based AI Accelerators
abstract
Epistasis detection is a bioinformatics application that searches for associations between sets of single nucleotide polymorphisms (SNPs) and a given trait of a population. Epistasis detection is a computationally complex problem, especially when tackling interaction orders above two. This paper presents a pioneering approach for performing third-order epistasis searches that has been devised around the bulk-synchronous parallel (BSP) model of execution, which is used in processor designs that target artificial intelligence (AI) workloads, such as Graphcore’s Intelligence Processing Unit (IPU). We propose a parallelization approach and a set of optimizations to efficiently exploit the computation and communication resources of these novel AI processors to perform precise bioinformatics searches. The proposed approach achieves a performance of 3.9 Tera SNP combinations evaluated per second, scaled to sample size, on an IPU-M2000 AI accelerator, while close to a linear speedup (up to 15.01×) is achieved with 16 IPU-M2000 accelerators.
Ricardo Nobre, Aleksandar Ilic, Sergio Santander-Jiménez, Leonel Sousa
IPDPS1
2022 Tensor-Accelerated Fourth-Order Epistasis Detection on GPUs
abstract
The improved accessibility of gene sequencing technologies has led to creation of huge datasets, i.e. patient records related to certain human diseases (phenotypes). Hence, deriving fast and accurate algorithms for efficiently processing these datasets is a paramount concern to enable some key healthcare scenarios, such as personalizing treatments, explaining the occurrence of and/or susceptibility to complex conditions and reducing the spread of infectious diseases. This is especially true for high-order epistasis detection, one of the most computationally challenging problems in bioinformatics, where associations between a given phenotype and single nucleotide polymorphisms (SNPs) of a population can often only be uncovered through evaluation of a large number of SNP combinations. To tackle this challenge, we propose a novel fourth-order epistasis detection algorithm that leverages tensor processing capabilities of two distinct accelerator architectures by efficiently mapping core computations related to processing quads of SNPs to binary tensor-accelerated matrix operations. Experimental results show that the proposed approach delivers very high performance even in single-GPU environments, e.g., 27.8 and 90.9 tera quads of SNPs per second, scaled to the sample size, were processed on Titan RTX (Turing) and A100 (Ampere) PCIe GPUs, respectively. Being the first approach that exploits tensor cores for accelerating searches with interaction order above three, the proposed method achieved a performance of up to 835.4 tera quads of SNPs per second on the 8-GPU HGX A100 server, which represents performance two or more orders of magnitude higher than that of related art.
Ricardo Nobre, Aleksandar Ilic, Sergio Santander-Jiménez, Leonel Sousa
ICPP1
2021 Fourth-Order Exhaustive Epistasis Detection for the xPU Era
abstract
The investigation of highly-efficient parallel algorithms targeting modern heterogeneous systems can provide bioinformaticians new mechanisms to find relations between genetics, phenotype and environment. This paper proposes an approach for fourth-order exhaustive, i.e. as precise as possible, epistasis detection targeting modern heterogeneous systems. Being implemented in Data Parallel C++ / SYCL, the proposed approach relies on technologies and tools built around open standards, making it able to target different types of architectures and devices. As a means to show interoperability with hardware from different sources, we have included performance results obtained from execution on different systems. Scaled to the number of samples, the proposed approach achieved a performance per GPU stream core of up to 236, 472 or 487 mega quads of SNPs processed per second, on GPUs with the Gen9.5, Gen12 and Turing architectures. This metric is reported for different implementation variants combining different considered optimizations. The proposed approach is able to target a wide range of CPU and GPU devices, enabling more users to access high-throughput epistasis detection software.
Ricardo Nobre, Aleksandar Ilic, Sergio Santander-Jiménez, Leonel Sousa
ICPP1
2021 Retargeting Tensor Accelerators for Epistasis Detection
abstract
The substitution of nucleotides at specific positions in the genome of a population, known as single-nucleotide polymorphisms (SNPs), has been correlated with a number of important diseases. Complex conditions such as Alzheimer's disease or Crohn's disease are significantly linked to genetics when the impact of multiple SNPs is considered. SNPs often interact in an epistatic manner, where the joint effect of multiple SNPs may not be simply mapped to a linear additive combination of individual effects. Genome-wide association studies considering epistasis are computationally challenging, especially when performing triplet searches is required. Some contemporary computer architectures support fused XOR and population count as the highest throughput operations as part of tensor operations. This article presents a new approach for efficiently repurposing this capability to accelerate 2-way (pairs) and 3-way (triplets) epistasis detection searches. Experimental evaluation targeting the Turing GPU architecture resulted in previously unattainable levels of performance, with the proposal being able to evaluate up to 108.1 and 54.5 tera unique sets of SNPs per second, scaled to the sample size, in 2-way and 3-way searches, respectively.
Ricardo Nobre, Aleksandar Ilic, Sergio Santander-Jiménez, Leonel Sousa
IEEE Trans. Parallel Distributed Syst.1
2020 Exploring the Binary Precision Capabilities of Tensor Cores for Epistasis Detection
abstract
Genome-wide association studies are performed to correlate a number of diseases and other physical or even psychological conditions (phenotype) with substitutions of nucleotides at specific positions in the human genome, mainly single-nucleotide polymorphisms (SNPs). Some conditions, possibly because of the complexity of the mechanisms that give rise to them, have been identified to be more statistically correlated with genotype when multiple SNPs are jointly taken into account. However, the discovery of new associations between genotype and phenotype is exponentially slowed down by the increase of computational power required when epistasis, i.e., interactions between SNPs, is considered. This paper proposes a novel graphics processing unit (GPU)-based approach for epistasis detection that combines the use of modern tensor cores with native support for processing binarized inputs with algorithmic and target-focused optimizations. Using only a single mid-range Turing-based GPU, the proposed approach is able to evaluate 64.8×1012and 25.4×1012sets of SNPs per second, normalized to the number of patients, when considering 2-way and 3-way epistasis detection, respectively. This proposal is able to surpass the state-of-the-art approach by 6× and 8.2× in terms of the number of pairs and triplets of SNP allelic patient data evaluated per unit of time per GPU.
Ricardo Nobre, Aleksandar Ilic, Sergio Santander-Jiménez, Leonel Sousa
IPDPS1
2020 Accelerating 3-Way Epistasis Detection with CPU+GPU Processing
Ricardo Nobre, Sergio Santander-Jiménez, Leonel Sousa, Aleksandar Ilic
JSSPP1
2018 Autotuning and adaptivity in energy efficient HPC systems: the ANTAREX toolbox
abstract
Designing and optimizing applications for energy-efficient High Performance Computing systems up to the Exascale era is an extremely challenging problem. This paper presents the toolbox developed in the ANTAREX European project for autotuning and adaptivity in energy efficient HPC systems. In particular, the modules of the ANTAREX toolbox are described as well as some preliminary results of the application to two target use cases. 1
Cristina Silvano, Gianluca Palermo, Giovanni Agosta, Amir H. Ashouri, Davide Gadioli, Stefano Cherubin, Emanuele Vitali, Luca Benini, Andrea Bartolini, Daniele Cesarini, João M. P. Cardoso, João Bispo, Pedro Pinto 0002, Ricardo Nobre, Erven Rohou, Loïc Besnard, Imane Lasri, Nico Sanna, Carlo Cavazzoni, Radim Cmar, Jan Martinovic, Katerina Slaninová, Martin Golasowski, Andrea Beccari, Candida Manelfi
CF14
2018 SOCRATES - A seamless online compiler and system runtime autotuning framework for energy-aware applications
abstract
Configuring program parallelism and selecting optimal compiler options according to the underlying platform architecture is a difficult task. Tipically, this task is either assigned to the programmer or done by a standard one-fits-all policy generated by the compiler or runtime system. A runtime selection of the best configuration requires the insertion of a lot of glue code for profiling and runtime selection. This represents a programming wall for application developers. This paper presents a structured approach, called SOCRATES, based on an aspect-oriented language (LARA) and a runtime autotuner (mARGOt) to mitigate this problem. LARA has been used to hide the glue code insertion, thus separating the pure functional application description from extra-functional requirements. mARGOT has been used for the automatic selection of the best configuration according to the runtime evolution of the application.1
Davide Gadioli, Ricardo Nobre, Pedro Pinto 0002, Emanuele Vitali, Amir H. Ashouri, Gianluca Palermo, João M. P. Cardoso, Cristina Silvano
DATE2
2016 A graph-based iterative compiler pass selection and phase ordering approach
abstract
Nowadays compilers include tens or hundreds of optimization passes, which makes it difficult to find sequences of optimizations that achieve compiled code more optimized than the one obtained using typical compiler options such as -O2 and -O3. The problem involves both the selection of the compiler passes to use and their ordering in the compilation pipeline. The improvement achieved by the use of custom phase orders for each function can be significant, and thus important to satisfy strict requirements such as the ones present in high-performance embedded computing systems. In this paper we present a new and fast iterative approach to the phase selection and ordering challenges resulting in compiled code with higher performance than the one achieved with the standard optimization levels of the LLVM compiler. The obtained performance improvements are comparable with the ones achieved by other iterative approaches while requiring considerably less time and resources. Our approach is based on sampling over a graph representing transitions between compiler passes. We performed a number of experiments targeting the LEON3 microarchitecture using the Clang/LLVM 3.7 compiler, considering 140 LLVM passes and a set of 42 representative signal and image processing C functions. An exhaustive cross-validation shows our new exploration method is able to achieve a geometric mean performance speedup of 1.28x over the best individually selected -OX flag when considering 100,000 iterations; versus geometric mean speedups from 1.16x to 1.25x obtained with state-of-the-art iterative methods not using the graph. From the set of exploration methods tested, our new method is the only one consistently finding compiler sequences that result in performance improvements when considering 100 or less exploration iterations. Specifically, it achieved geometric mean speedups of 1.08x and 1.16x for 10 and 100 iterations, respectively.
Ricardo Nobre, Luiz G. A. Martins, João M. P. Cardoso
LCTES1
2016 SAR Image Segmentation With Rényi's Entropy
abstract
Synthetic aperture radar (SAR) image segmentation is an important task in image processing. However, classic segmentation techniques are inadequate due to the presence of speckle noise. In this paper, we present a methodology for SAR image segmentation that uses the matrix of Rényi's entropy. This matrix arises from SAR data that follows the GA0model, and here, it is an input to segmentation methods. For performance evaluation of the proposed methodology, we employ the error of segmentation, the cross-region fitting index, the Dice measure, as well as the rates of false positives and negatives. Tests have been performed on synthetic and real SAR images and the matrix of entropy has improved the results, regardless of the increase of the number of looks. The Otsu's method of global thresholding produced good segmentation results when applied to this matrix.
Ricardo Nobre, Francisco Alixandre Ávila Rodrigues, Régis C. P. Marques, Juvêncio S. Nobre, Jeová Farias Sales Rocha Neto, Fátima N. S. de Medeiros
IEEE Signal Process. Lett.1
2016 Clustering-Based Selection for the Exploration of Compiler Optimization Sequences
abstract
A large number of compiler optimizations are nowadays available to users. These optimizations interact with each other and with the input code in several and complex ways. The sequence of application of optimization passes can have a significant impact on the performance achieved. The effect of the optimizations is both platform and application dependent. The exhaustive exploration of all viable sequences of compiler optimizations for a given code fragment is not feasible. As this exploration is a complex and time-consuming task, several researchers have focused on Design Space Exploration (DSE) strategies both to select optimization sequences to improve the performance of each function of the application and to reduce the exploration time. In this article, we present a DSE scheme based on a clustering approach for grouping functions with similarities and exploration of a reduced search space resulting from the combination of optimizations previously suggested for the functions in each group. The identification of similarities between functions uses a data mining method that is applied to a symbolic code representation. The data mining process combines three algorithms to generate clusters: the Normalized Compression Distance, the Neighbor Joining, and a new ambiguity-based clustering algorithm. Our experiments for evaluating the effectiveness of the proposed approach address the exploration of optimization sequences in the context of the ReflectC compiler, considering 49 compilation passes while targeting a Xilinx MicroBlaze processor, and aiming at performance improvements for 51 functions and four applications. Experimental results reveal that the use of our clustering-based DSE approach achieves a significant reduction in the total exploration time of the search space (20× over a Genetic Algorithm approach) at the same time that considerable performance speedups (41% over the baseline) were obtained using the optimized codes. Additional experiments were performed considering the LLVM compiler, considering 124 compilation passes, and targeting a LEON3 processor. The results show that our approach achieved geometric mean speedups of 1.49 × , 1.32 × , and 1.24 × for the best 10, 20, and 30 functions, respectively, and a global improvement of 7% over the performance obtained when compiling with -O2.
Luiz G. A. Martins, Ricardo Nobre, João M. P. Cardoso, Alexandre C. B. Delbem, Eduardo Marques
ACM Trans. Archit. Code Optim.2
2015 Use of Previously Acquired Positioning of Optimizations for Phase Ordering Exploration
abstract
This paper presents a new approach to efficiently search for suitable compiler pass sequences, a challenge known as phase ordering. Our approach relies on information about the relative positions of compiler passes in compiler pass sequences previously generated for a set of functions when compiling for a specific processor. We enhanced two iterative compiler pass exploration schemes, one relying on simple sequential compiler pass insertion and other implementing an auto-tuned simulated annealing process, with a data structure that holds information about the relative positions of compiler sequences; in order to reduce the set of compiler passes considered for insertion in a given position of a given candidate compiler pass sequence to include only the passes that have a higher probability of performing well on that relative position in the compiler sequence, speeding up the exploration time as a result. We tested our approach with two different compilers and two different targets; the ReflectC and the LLVM compilers, targeting a MicroBlaze processor and a LEON3 processor, respectively. The experimental results show that we can considerably reduce the number of algorithm iterations by a factor of up to more than an order of magnitude when targeting the MicroBlaze or the LEON3, while finding compiler sequences that result in binaries that when executed on the target processor/simulator are able to outperform (i.e. use less CPU cycles) all the standard optimization levels (i.e., we compare against the most performing optimization level flag on each kernel, e.g. -O1, -O2 or -O3 in the case of LLVM) by a geometric mean performance improvement of 1.23x and 1.20x when targeting the MicroBlaze processor, and 1.94x and 2.65x when targetting the LEON3 processor; for each of the two exploration algorithms and two kernel sets considered.
Ricardo Nobre, Luiz G. A. Martins, João M. P. Cardoso
SCOPES1
2014 A clustering-based approach for exploring sequences of compiler optimizations
abstract
In this paper we present a clustering-based selection approach for reducing the number of compilation passes used in search space during the exploration of optimizations aiming at increasing the performance of a given function and/or code fragment. The basic idea is to identify similarities among functions and to use the passes previously explored each time a new function is being compiled. This subset of compiler optimizations is then used by a Design Space Exploration (DSE) process. The identification of similarities is obtained by a data mining method which is applied to a symbolic code representation that translates the main structures of the source code to a sequence of symbols based on transformation rules. Experiments were performed for evaluating the effectiveness of the proposed approach. The selection of compiler optimization sequences considering a set of 49 compilation passes and targeting a Xilinx MicroBlaze processor was performed aiming at latency improvements for 41 functions from Texas Instruments benchmarks. The results reveal that the passes selection based on our clustering method achieves a significant gain on execution time over the full search space still achieving important performance speedups.
Luiz G. A. Martins, Ricardo Nobre, Alexandre C. B. Delbem, Eduardo Marques, João M. P. Cardoso
IEEE Congress on Evolutionary Computation2
2014 Exploration of compiler optimization sequences using clustering-based selection
abstract
Due to the large number of optimizations provided in modern compilers and to compiler optimization specific opportunities, a Design Space Exploration (DSE) is necessary to search for the best sequence of compiler optimizations for a given code fragment (e.g., function). As this exploration is a complex and time consuming task, in this paper we present DSE strategies to select optimization sequences to both improve the performance of each function and reduce the exploration time. The DSE is based on a clustering approach which groups functions with similarities and then explore the reduced search space provided by the optimizations previously suggested for the functions in each group. The identification of similarities between functions uses a data mining method which is applied to a symbolic code representation of the source code. The DSE process uses the reduced set identified by clustering in two ways: as the design space or as the initial configuration. In both ways, the adoption of a pre-selection based on clustering allows the use of simple and fast DSE algorithms. Our experiments for evaluating the effectiveness of the proposed approach address the exploration of compiler optimization sequences considering 49 compilation passes and targeting a Xilinx MicroBlaze processor, and were performed aiming performance improvements for 41 functions. Experimental results reveal that the use of our new clustering-based DSE approach achieved a significant reduction on the total exploration time of the search space (18x over a Genetic Algorithm approach for DSE) at the same time that important performance speedups (43% over the baseline) were obtained by the optimized codes.
Luiz G. A. Martins, Ricardo Nobre, Alexandre C. B. Delbem, Eduardo Marques, João M. P. Cardoso
LCTES2
2013 Identifying sequences of optimizations for HW/SW compilation
abstract
This PhD work focuses on techniques to devise sequences of compiler optimizations able to increase the performance of a given function/program when mapped to software or to hardware. The techniques are being evaluated using an integrated hardware/software compiler controlled by strategies expressed in LARA, a domain-specific language specially developed to guide/control code instrumentation, code transformations and compiler optimizations. We started by evaluating a Design Space Exploration (DSE) scheme based on a Simulated Annealing approach expressed in LARA. This scheme is able to explore compilation design space points generated by distinct compilation sequences. In our first studies we apply compilation strategies to two hot-spot functions from two industrial applications provided in the context of the REFLECT project, Grid Iterate, from a 3-D Path Planning application, and Filter Subband, part of an MPEG audio encoder. A first evaluation of the DSE scheme resulted in speed-ups of 1.75 and 1.23 for Filter Subband, and 2.24 and 1.13 for Grid Iterate, when targeting a MicroBlaze processor and reconfigurable hardware, respectively.
Ricardo Nobre
FPL1
2013 The MATISSE MATLAB compiler
abstract
This paper describes MATISSE, a MATLAB to C compiler targeting embedded systems that is based on Strategic and Aspect-Oriented Programming concepts. MATISSE takes as input: (1) MATLAB code and (2) LARA aspects related to types and shapes, code insertion/removal, and specialization based directives defining default variable values. In this paper we also illustrate the use of MATISSE in leveraging data types and shapes to generate customized C code suitable for high-level hardware synthesis tools. The preliminary experimental results presented here reveal the described approach to yield performance results for the resulting hardware and software references implementations that are comparable in terms of performance with hand-crafted solutions but derived automatically at a fraction of the cost.
João Bispo, Pedro Pinto 0002, Ricardo Nobre, Tiago Carvalho 0001, João M. P. Cardoso, Pedro C. Diniz
INDIN3
2013 Enriching MATLAB with aspect-oriented features for developing embedded systems
João M. P. Cardoso, João M. Fernandes 0001, Miguel P. Monteiro 0001, Tiago Carvalho 0001, Ricardo Nobre
J. Syst. Archit.5
2012 Specifying Compiler Strategies for FPGA-based Systems
abstract
The development of applications for high-performance Field Programmable Gate Array (FPGA) based embedded systems is a long and error-prone process. Typically, developers need to be deeply involved in all the stages of the translation and optimization of an application described in a high-level programming language to a lower-level design description to ensure the solution meets the required functionality and performance. This paper describes the use of a novel aspect-oriented hardware/software design approach for FPGA-based embedded platforms. The design-flow uses LARA, a domain-specific aspect-oriented programming language designed to capture high-level specifications of compilation and mapping strategies, including sequences of data/computation transformations and optimizations. With LARA, developers are able to guide a design-flow to partition and map an application between hardware and software components. We illustrate the use of LARA on two complex real-life applications using high-level compilation and synthesis strategies for achieving complete hardware/software implementations with speedups of 2.5× and 6.8× over software-only implementations. By allowing developers to maintain a single application source code, this approach promotes developer productivity as well as code and performance portability.
João M. P. Cardoso, José C. Alves, Ricardo Nobre, Pedro C. Diniz, José Gabriel F. Coutinho, Wayne Luk
FCCM4