Sandro Rigo

dblp:79/2208 · DBLP profile ↗
← Back
24ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-9539-6874ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Multi -FPGA streaming using OpenMP
abstract
The growth in the demand for high-performance and power-efficient applications has led to an increasing interest in FPGA-based acceleration. FPGAs have been applied to a wide range of applications. Still, programming them can be a complex task, requiring extensive knowledge of tools and libraries, especially in multi-FPGA architectures. As such, the desire for tools and frameworks to ease the burden and abstract the knowledge of using FPGAs has increased. OpenMP, an already dominant parallel programming model in HPC, has been shown to be a successful approach to program multi-FPGA architecture. This work is based on the OMPC-F framework, which leverages the capability of OpenMP to offload computation to FPGAs. Although OMPC-F abstracts FPGA handling and task distribution from the final user, it does not support streaming computation based on multi-FPGA architectures. Streaming is widely used in FPGA designs to create a pipeline of computation between FPGA kernels. This work uses FPGA kernel binary information to synthesize streams as OpenMP buffers, while adapting the OpenMP dependency system accordingly. The proposal was evaluated in an AMD/Xilinx multi-FPGA system and shows speedups of the order of 6.84x in 8 FPGAs, scaling well with the addition of more kernels and FPGAs to the architecture. Moreover, compared to the regular approach to developing FPGA applications based on MPI+XRT communication, the proposed approach reduces the programming effort by 59% according to various code analysis metrics, resulting in a small average overhead of 4.3% when compared to the MPI+XRT programming model.
Pedro Henrique Di Francia Rosso, Rémy Neveu, Nusrat Jahan Lisa, Lucas B. da Silva, Hervé Yviquel, Sandro Rigo, Vanderlei Bonato, Guido Araujo
J. Parallel Distributed Comput.6
2025 A Distributed and Storage-Aware Approach to Large-Scale Cholesky Factorization
abstract
Cholesky factorization is a core operation in scientific computing, yet its scalability is often constrained by memory limitations when processing extremely large dense matrices. This work introduces an out-of-core Cholesky factorization algorithm for symmetric positive-definite matrices that integrates GPU acceleration, block-wise lossless compression, and parallel I/O to overcome these limitations. The approach leverages the OMPC runtime for asynchronous task scheduling and employs HDF5 to store the matrix on disk, taking advantage of Lustre’s parallel I/O capabilities in distributed environments. Tiles are decompressed just-in-time on the GPU, significantly reducing host memory usage, storage footprint, and end-to-end data movement overhead—from disk through the CPU to the GPU—without compromising numerical accuracy. Experimental results show that the proposed method scales across 8 GPU nodes, successfully factorizing matrices up to 3M × 3M. In comparison, SLATE could only handle sizes up to 700K × 700K, with the proposed algorithm achieving up to 41% higher throughput. These results demonstrate the algorithm’s scalability and competitiveness beyond memory-constrained in-core solutions, offering a practical path for enabling extreme-scale scientific applications.
Carla Cusihuallpa, Rodrigo Ceccato, Sandro Rigo, Guido Araujo, Hervé Yviquel
SBAC-PAD3
2024 Combining Compression and Prefetching to Improve Checkpointing for Inverse Seismic Problems in GPUs
Thiago Maltempi, Sandro Rigo, Márcio Machado Pereira, Hervé Yviquel, Jessé Costa, Guido Araujo
Euro-Par (3)2
2024 DeepWave: A Software Stack for Parallelizing Deep Learning Models Used in Geophysics
abstract
This paper introduces DeepWave, a novel software stack, and methodology designed to integrate generative artificial intelligence into traditional seismic surveying techniques, significantly enhancing the computational efficiency of geophysical exploration. By utilizing advanced machine learning frameworks such as JAX, FLAX, and ALPA, DeepWave employs a parallelization strategy for image-to-image translation networks, optimizing the seismic data interpretation process. DeepWave reduces the computational demands of intensive geophysical algorithms, such as Full-waveform Inversion (FWI), while maintaining the accuracy required for detailed subsurface analysis. This method enables faster and more efficient processing of large seismic datasets, providing deeper insights into the Earth’s subsurface structures with reduced computational resources. The results demonstrate a substantial improvement in processing speed and resource management, establishing a new geophysical research and exploration standard.
Allan Pinto, Gustavo Leite, Márcio Machado Pereira, Hervé Yviquel, Sandro Rigo, Guido Araujo
SBAC-PAD5
2024 Ion-molecule collision cross-section calculations using trajectory parallelization in distributed systems
Samuel Cajahuaringa, Leandro Zanotto, Sandro Rigo, Hervé Yviquel, Munir S. Skaf, Guido Araujo
J. Parallel Distributed Comput.3
2022 Ion-Molecule Collision Cross-Section Simulation using Linked-cell and Trajectory Parallelization
abstract
Ion Mobility coupled to Mass Spectrometry (IMMS) has become a highly valued tool for structural characterization of biological samples. In IM-MS the protein under investigation is ionized and accelerated by an electric field into a drift tube where it collides against a buffer gas. The separation of the gas-phase ions is then measured through the differences in their rotationally averaged Collision Cross-Section (CCS) values. The utility of the measured CCS for structural characterization critically depends on the validation against its theoretical calculation, which relies on intensive molecular mechanics simulation. Increasing the performance of CCS simulation is thus a relevant computational-chemistry research problem. This work shows that the combination of a linked-cell based algorithm with parallelization techniques can considerably increase the performance of CCS simulation. Experimental results reveal speedups from$\sim \mathbf{10}\times$up to$\sim \mathbf{400}\times$and parallelization efficiency greater than 0.98 when compared to High Performance Collision Cross Section (HPCCS), an optimized solution for CCS simulation. This reduces the CCS computation time from hours to minutes for a large range of proteins, making the proposed method the most performant approach to this problem nowadays, to the best of our knowledge.
Samuel Cajahuaringa, Leandro Zanotto, Daniel L. Z. Caetano, Sandro Rigo, Hervé Yviquel, Munir S. Skaf, Guido Araujo
SBAC-PAD4
2021 Employing Simulation to Facilitate the Design of Dynamic Binary Translators
abstract
Dynamic Binary Translation (DBT) is a sophisticated technique that allows the implementation of highperformance ISA emulators. In this technique, the guest code is compiled dynamically at runtime. Consequently, achieving good performance depends on several design decisions, including the shape of the regions of code being translated. Researchers and engineers explore these decisions to bring the best performance possible. However, a real DBT engine is a very sophisticated piece of software, and modifying one is a challenging and demanding task. Hence, we propose using simulation to evaluate the impact of design decisions on dynamic binary translators and present RAIn, an open-source DBT simulator that facilitates the test of DBT's design decisions, such as Region Formation Techniques (RFTs). RAIn outputs several statistics that support the analysis of how design decisions may affect the behavior and the performance of a real DBT. We validated RAIn running a set of experiments with six well-known RFTs (NET, MRET2, LEI, NETPlus, NET-R, and NETPlus-e-r) and showed that it could reproduce well-known results from the literature without the effort of implementing them on a real and thus complex dynamic binary translator engine.
Vanderson Martins do Rosário, Raphael Zinsly, Sandro Rigo, Edson Borin
SBAC-PAD3
2020 Simulating Smart Campus Applications in Edge and Fog Computing
abstract
Due to the rapid increase of IoT applications and their use in many different areas, large amounts of data have been generated to be processed and stored. In this scenario, some applications are sensitive to high latency and response times. In order to fulfil these requirements, Edge and Fog Computing appear with the objective of bringing processing and storage devices closer to applications and management mechanisms. In this context, due to limitations related to high cost, scalability and planning, several mechanisms and algorithms need to be simulated before being implemented in the real world. This paper presents a comparison between two simulation tools and their main characteristics (EdgeCloudSim and iFogSim) using a smart campus scenario deployed at the University of Campinas, where the sensors collect data from water meters and smart energy marker watches, in addition to smart public transportation and battery disposal bins. Our evaluation shows that the information processing in edge and fog can efficiently serve the applications, however, each simulation tool has its specificities, and should be used according to the researcher's objectives and needs.
Denis Contini, Lucas Fernando Souza de Castro, Edmundo Roberto Mauro Madeira, Sandro Rigo, Luiz Fernando Bittencourt
SMARTCOMP4
2019 Approximation with Error Bounds in Spark
abstract
Many decision-making queries are based on aggregating massive amounts of data, where sampling is an important approximation technique for reducing execution times. It is important to estimate error bounds when sampling to help users balance between precision and performance. However, error bound estimation is challenging because data processing pipelines often transform the input dataset in complex ways before computing the final aggregated values. In this paper, we introduce a sampling framework to support approximate computing with estimated error bounds in Spark. Our framework allows sampling to be performed at multiple arbitrary points within a sequence of transformations preceding an aggregation operation. The framework constructs a data provenance tree to maintain information about how transformations are clustering output data items to be aggregated. It then uses the tree and multi-stage sampling theories to compute the approximate aggregate values and corresponding error bounds. When information about output keys are available early, the framework can also use adaptive stratified reservoir sampling to avoid (or reduce) key losses in the final output and to achieve more consistent error bounds across popular and rare keys. Finally, the framework includes an algorithm to dynamically choose sampling rates to meet user-specified constraints on the CDF of error bounds in the outputs. We have implemented a prototype of our framework called ApproxSpark and used it to implement five approximate applications from different domains. Evaluation results show that ApproxSpark can (a) significantly reduce execution time if users can tolerate small amounts of uncertainties and, in many cases, loss of rare keys, and (b) automatically find sampling rates to meet user-specified constraints on error bounds. We also explore and discuss extensively tradeoffs between sampling rates, execution time, accuracy and key loss.
Guangyan Hu, Sandro Rigo, Desheng Zhang 0002, Thu D. Nguyen
MASCOTS2
2018 Uncertainty Propagation in Data Processing Systems
abstract
We are seeing an explosion of uncertain data---i.e., data that is more properly represented by probability distributions or estimated values with error bounds rather than exact values---from sensors in IoT, sampling-based approximate computations and machine learning algorithms. In many cases, performing computations on uncertain data as if it were exact leads to incorrect results. Unfortunately, developing applications for processing uncertain data is a major challenge from both the mathematical and performance perspectives. This paper proposes and evaluates an approach for tackling this challenge in DAG-based data processing systems. We present a framework for uncertainty propagation (UP) that allows developers to modify precise implementations of DAG nodes to process uncertain inputs with modest effort. We implement this framework in a system called UP-MapReduce, and use it to modify ten applications, including AI/ML, image processing and trend analysis applications to process uncertain data. Our evaluation shows that UP-MapReduce propagates uncertainties with high accuracy and, in many cases, low performance overheads. For example, a social network trend analysis application that combines data sampling with UP can reduce execution time by 2.3x when the user can tolerate a maximum relative error of 5% in the final answer. These results demonstrate that our UP framework presents a compelling approach for handling uncertain data in DAG processing.
Ioannis Manousakis, Íñigo Goiri, Ricardo Bianchini, Sandro Rigo, Thu D. Nguyen
SoCC4
2018 Exploring Power Budget Scheduling Opportunities and Tradeoffs for AMR-Based Applications
abstract
Computational demand has brought major changes to Advanced Cyber-Infrastructure (ACI) architectures. It is now possible to run scientific simulations faster and obtain more accurate results. However, power and energy have become critical concerns. Also, the current roadmap toward the new generation of ACI includes power budget as one of the main constraints. Current research efforts have studied power and performance tradeoffs and how to balance these (e.g., using Dynamic Voltage and Frequency Scaling (DVFS) and power capping for meeting power constraints, which can impact performance). However, applications may not tolerate degradation in performance, and other tradeoffs need to be explored to meet power budgets (e.g., involving the application in making energy-performance-quality tradeoff decisions). This paper proposes using the properties of AMR-based algorithms (e.g., dynamically adjusting the resolution of a simulation in combination with power capping techniques) to schedule or re-distribute the power budget. It specifically explores the opportunities to realize such an approach using checkpointing as a proof-of-concept use case and provides a characterization of a representative set of applications that use Adaptive Mesh Refinement (AMR) methods, including a Low-Mach-Number Combustion (LMC) application. It also explores the potential of utilizing power capping to understand power-quality tradeoffs via simulation.
Yubo Qin, Ivan Rodero, Pradeep Subedi, Manish Parashar, Sandro Rigo
SBAC-PAD5
2014 Leveraging Optimization Methods for Dynamically Assisted Control-Flow Integrity Mechanisms
abstract
Dynamic Binary Modification (DBM) tools are useful for cross-platform execution of binaries and are powerful run time environments that allow execution optimizations, instrumentation and profiling. These tools have also been used as enablers for control-flow integrity verification, a process that consists in the observation and analysis of a program's execution path focusing on the detection of anomalies, such as those arising from flow corruption based software attacks. Even though this class of tools helps us in identifying a myriad of attacks, it is typically expensive at run time and introduce significant overhead to the program execution. Considering their inherent high cost, further expanding the capabilities of such tools for detection of program flow anomalies can slow down the analysis to the point that it is unfeasible to run it in real world workflows. In this paper we present a mechanism for including program flow verification in DBMs that uses asynchronous analysis and applies different parallel-programming techniques that leverage current multi-core systems to control the overhead of our analysis. Our mechanism was tested against synthetic program flow corruption use cases and correctly detected all detours. With our new optimizations, we show that our system achieves an slowdown of only 1.46x, while a naively implemented verification system face 4.22x of overhead.
Lucas Teixeira, Edson Borin, Sandro Rigo
SBAC-PAD4
2014 Adaptive global power optimization for Web servers
Leonardo Piga, Reinaldo A. Bergamaschi, Maurício Breternitz, Sandro Rigo
J. Supercomput.4
2013 Assessing computer performance with stocs
abstract
Several aspects of a computer system cause performance measurements to include random errors. Moreover, these systems are typically composed of a non-trivial combination of individual components that may cause one system to perform better or worse than another depending on the workload. Hence, properly measuring and comparing computer systems performance are non-trivial tasks.
Leonardo Piga, Gabriel F. T. Gomes, Rafael Auler, Bruno Rosa 0001, Sandro Rigo, Edson Borin
ICPE5
2010 STM versus lock-based systems: an energy consumption perspective
abstract
The shift towards multicore processors and the well-known drawbacks imposed by lock-based synchronization have forced researchers to devise new alternatives for building concurrent software, of which transactional memory is a promising one. This work presents a comprehensive study on the energy consumption of a state-of-the-art STM (Software Transactional Memory) implementation using STAMP, a representative set of transactional workloads, comparing it to its lock-based counterpart. Our results show that STM can be up to 22x (~3x on average) more energy-inefficient when compared to locks. This work is a novel step towards a better understanding of the energy behavior of STM systems.
Felipe Klein, Alexandro Baldassin, Paulo Centoducatte, Sandro Rigo, Rodolfo Azevedo
ISLPED5
2009 A novel verification technique to uncover out-of-order DUV behaviors
abstract
Post-partitioning verification has to deal with abstract data, implementation artifacts, and the order of events may not be preserved in the DUV due to the concurrency treatment in the golden model. Existing techniques are limited either by the use of greedy heuristics (jeopardizing verification guarantees) or by black-box approaches (impairing observability). This work proposes a novel white-box technique that overcomes those limitations by casting the problem as an extended bipartite graph matching. By relying on proven properties, solid verification guarantees are provided. Experimental validation was performed upon platforms built around contemporary real-life applications.
Gabriel Marcilio, Luiz Cláudio Villar dos Santos, Bruno de Carvalho Albertini, Sandro Rigo
DAC4
2008 An open-source binary utility generator
abstract
Electronic system level (ESL) modeling allows early hardware-dependent software (HDS) development. Due to broad CPU diversity and shrinking time-to-market, HDS development can neither rely on hand-retargeting binary tools, nor can it rely on pre-existent tools within standard packages. As a consequence, binary utilities which can be easily adapted to new CPU targets are of increasing interest. We present in this article a framework for automatic generation of binary utilities. It relies on two innovative ideas: platform-aware modeling and more inclusive relocation handling. Generated assemblers, linkers, disassemblers and debuggers were validated for MIPS, SPARC, PowerPC, i8051 and PIC16F84. An open-source prototype generator is available for download.
Alexandro Baldassin, Paulo Centoducatte, Sandro Rigo, Daniel C. Casarotto, Luiz Cláudio Villar dos Santos, Max R. de O. Schultz, Olinto J. V. Furtado
ACM Trans. Design Autom. Electr. Syst.3
2005 Extending the ArchC Language for Automatic Generation of Assemblers
abstract
In this paper, we extend the ArchC language with new constructs to describe the assembly language syntax and operand encoding of an instruction set architecture. Based on the extended language we have created a tool which can automatically generate assemblers. Our tool uses the GNU Binutils framework in order to produce the assembler, generating the architecture dependent files necessary to retarget the GNU assembler and the Binutils libraries. We have generated assemblers for the MIPS-I and SPARC-V8 architectures based on ArchC models using our tool. The assemblers generated for both architectures were compared with the default gas assemblers for a set of files taken from the MiBench benchmark, and the ELF object files generated by each pair of assemblers were equivalent in both cases.
Alexandro Baldassin, Paulo Centoducatte, Sandro Rigo
SBAC-PAD3
2004 Modeling and Simulating Memory Hierarchies in a Platform-Based Design Methodology
abstract
This paper presents an environment based on SystemC for architecture specification of programmable systems. Making use of the new architecture description language ArchC, able to capture the processor description as well as the memory subsystem configuration, this environment offers support for system-level specification, intended for platform-based design. As a case study, it is presented the memory architecture exploration for a simple image processing application, yet a more robust environment evaluation is performed through the execution of some real-world benchmarks.
Pablo Viana, Edna Barros, Sandro Rigo, Rodolfo Azevedo, Guido Araujo
DATE3
2004 Optimizations for Compiled Simulation Using Instruction Type Information
abstract
The design of new architectures can be simplified with the use of retargetable instruction set simulation tools, which can validate the design decisions in the design exploration cycle with high flexibility and reduced cost. The growing system complexity makes the traditional approach inefficient for today's architectures. Compiled simulation techniques make use of a priori knowledge to accelerate the simulation, with the highest efficiency achieved by employing static scheduling techniques. This paper presents our approach to the static scheduling compiled simulation technique that is 90% faster than the best published performance results. It also introduces two novel optimization techniques based on instruction type information that further increase the simulation speed by more than 100%. The so-called fast static compiled simulation (FSCS) technique applicability will be demonstrated by the use of the SPARC and MIPS architectures.
Marcus Bartholomeu, Rodolfo Azevedo, Sandro Rigo, Guido Araujo
SBAC-PAD3
2004 ArchC: A SystemC-Based Architecture Description Language
abstract
This paper presents an architecture description language (ADL) called ArchC, which is an open-source SystemC-based language that is specialized for processor architecture description. Its main goal is to provide enough information, at the right level of abstraction, in order to allow users to explore and verify new architectures, by automatically generating software tools like simulators and co-verification interfaces. ArchC's key features are a storage-based co-verification mechanism that automatically checks the consistency of a refined ArchC model against a reference (functional) description, memory hierarchy modeling capability, the possibility of integration with other SystemC IPs and the automatic generation of high-level SystemC simulators. We have used ArchC to synthesize both functional and cycle-based simulators for the MIPS, Intel 8051 and SPARC V8 processors, as well as functional models of modern architectures like TMS320C62x, XScale and PowerPC.
Sandro Rigo, Guido Araujo, Marcus Bartholomeu, Rodolfo Azevedo
SBAC-PAD1
2003 Exploring Memory Hierarchy with ArchC
abstract
We present the cache configuration exploration of a programmable system, in order to find the best matching between the architecture and a given application. Here, programmable systems composed by processor and memories may be rapidly simulated making use of ArchC, an architecture description language (ADL) based on SystemC. Initially designed to model processor architectures, ArchC was extended to support a more detailed description of the memory subsystem, allowing the design space exploration of the whole programmable system. As an example, it is shown an image processing application, running on a SPARC-V8 processor-based architecture, which had its memory organization adjusted to minimize cache misses.
Pablo Viana, Edna Barros, Sandro Rigo, Rodolfo Azevedo, Guido Araujo
SBAC-PAD3
2001 Optimal Live Range Merge for Address Register Allocation in Embedded Programs
Guilherme Ottoni, Sandro Rigo, Guido Araujo, Subramanian Rajagopalan, Sharad Malik
CC2
2001 A retargetable VLIW compiler framework for DSPs withinstruction-level parallelism
abstract
A standard design methodology for embedded processors today is the system-on-a-chip design with potentially multiple heterogeneous processing elements on a chip, such as a very long instruction word (VLIW) processor, digital signal processor (DSP), and field-programmable gate array. To be able to program these devices, we need compilers that are capable of generating efficient code for the different types of processing elements with efficiency measured in terms of power, area, and execution time. In addition, the compilers should also be highly retargetable to enable the system designer to quickly evaluate different cores for the application on hand and reduce the time to market. In this paper, we show that we can extend a conventional VLIW compilation environment to develop highly retargetable optimizing compilers for DSPs with irregular architectures. We have used the second generation Fujitsu Hiperion fixed-point DSP as our primary example to evaluate the compiler framework. We demonstrate through experimental results that execution time for the assembly code generated using our framework is roughly two times better than that of the code generated by a widely used commercially available DSP compiler. Even without incorporating DSP-specific optimizations in our extended VLIW framework, we demonstrate that the compiled code has a better performance than the code generated by a commercial DSP-specific compiler in all our examples.
Subramanian Rajagopalan, Sreeranga P. Rajan, Sharad Malik, Sandro Rigo, Guido Araujo, Koichiro Takayama
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4