EDBT 2026 Demo / reviewers in the wild / expert
Márcio Machado Pereira
dblp:144/4917
· DBLP profile ↗
19ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Using Task Graph Caching to Accelerate TVM Code GenerationabstractDeep Learning (DL) models are at the core of a growing number of applications, making fast, low-latency execution across diverse device architectures both a critical requirement and a challenge. DL compilers, such as TVM, address this challenge by automatically translating high-level models into optimized low-level code that effectively exploits device architectures. However, search-space-based algorithms face difficulties in exploring the vast optimization sequence space, often resulting in lengthy compilation times that significantly impact the design cycle. This article introduces the Task Graph Caching (TGC) algorithm, 1 which aims to reduce the high compilation time while preserving the quality of the code generated by TVM. In particular, TGC enhances TVM auto-tuning by exploiting the fact that similar DL subgraphs appear both within and across models, thus enabling optimization sequences discovered in past compilations to guide future executions. To achieve this, TGC introduces a cache structure that stores high-performance optimization sequences found in previous TVM executions. This information is then used to seed the population of TVM evolutionary search, avoiding redundant exploration of the optimization space and accelerating convergence. Experimental results on twelve DL models show that TGC can significantly speed up the search for efficient optimization sequences, reducing auto-tuning time by up to 2.89× for Ansor and 3.13× for MetaSchedule on CPU. Moreover, TGC reduces auto-tuning time by up to 3.25× for MetaSchedule on GPU. Furthermore, on average, TGC can maintain the inference time achieved by the default TVM, making it a promising solution for accelerating the compilation of DL models. Thais Aparecida Silva Camacho, Lucas Fernando Alvarenga e Silva, Márcio Machado Pereira, Guido Araujo |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | Scalable OpenMP Remote Offloading via Asynchronous MPI and Coroutine-Driven Communication
Jhonatan Cléto, Guilherme Valarini, Márcio Machado Pereira, Guido Araujo, Hervé Yviquel |
Euro-Par (3) | 3 |
| 2024 | Combining Compression and Prefetching to Improve Checkpointing for Inverse Seismic Problems in GPUs
Thiago Maltempi, Sandro Rigo, Márcio Machado Pereira, Hervé Yviquel, Jessé Costa, Guido Araujo |
Euro-Par (3) | 3 |
| 2024 | DeepWave: A Software Stack for Parallelizing Deep Learning Models Used in GeophysicsabstractThis paper introduces DeepWave, a novel software stack, and methodology designed to integrate generative artificial intelligence into traditional seismic surveying techniques, significantly enhancing the computational efficiency of geophysical exploration. By utilizing advanced machine learning frameworks such as JAX, FLAX, and ALPA, DeepWave employs a parallelization strategy for image-to-image translation networks, optimizing the seismic data interpretation process. DeepWave reduces the computational demands of intensive geophysical algorithms, such as Full-waveform Inversion (FWI), while maintaining the accuracy required for detailed subsurface analysis. This method enables faster and more efficient processing of large seismic datasets, providing deeper insights into the Earth’s subsurface structures with reduced computational resources. The results demonstrate a substantial improvement in processing speed and resource management, establishing a new geophysical research and exploration standard. Allan Pinto, Gustavo Leite, Márcio Machado Pereira, Hervé Yviquel, Sandro Rigo, Guido Araujo |
SBAC-PAD | 3 |
| 2023 | Tensor slicing and optimization for multicore NPUs
Rafael C. F. Sousa, Márcio Machado Pereira, Yongin Kwon, Namsoon Jung, Michael Frank 0008, Guido Araujo |
J. Parallel Distributed Comput. | 2 |
| 2023 | Advancing Direct Convolution Using Convolution Slicing Optimization and ISA ExtensionsabstractConvolution is one of the most computationally intensive operations that must be performed for machine learning model inference. A traditional approach to computing convolutions is known as the Im2Col + BLAS method. This article proposes SConv: a direct-convolution algorithm based on an MLIR/LLVM code-generation toolchain that can be integrated into machine-learning compilers. This algorithm introduces: (a) Convolution Slicing Analysis (CSA)—a convolution-specific 3D cache-blocking analysis pass that focuses on tile reuse over the cache hierarchy; (b) Convolution Slicing Optimization—a code-generation pass that uses CSA to generate a tiled direct-convolution macro-kernel; and (c) Vector-based Packing—an architecture-specific optimized input-tensor packing solution based on vector-register shift instructions for convolutions with unitary stride. Experiments conducted on 393 convolutions from full ONNX-MLIR machine learning models indicate that the elimination of the Im2Col transformation and the use of fast packing routines result in a total packing time reduction, on full model inference, of 2.3×–4.0× on Intel x86 and 3.3×–5.9× on IBM POWER10. The speed-up over an Im2Col + BLAS method based on current BLAS implementations for end-to-end machine-learning model inference is in the range of 11%–27% for Intel x86 and 11%–34% for IBM POWER10 architectures. The total convolution speedup for model inference is 13%–28% on Intel x86 and 23%–39% on IBM POWER10. SConv also outperforms BLAS GEMM, when computing pointwise convolutions in more than 82% of the 219 tested instances. Victor Ferrari, Rafael C. F. Sousa, Márcio Machado Pereira, João P. L. de Carvalho, José Nelson Amaral, José E. Moreira, Guido Araujo |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | Improving Convolution via Cache Hierarchy Tiling and Reduced PackingabstractConvolution is one of the most computationally intensive machine learning model operations, usually solved by the known Im2Col + BLAS method. This work proposes a novel convolution-algorithm to improve upon Im2Col + BLAS by introducing (a) CSA: a convolution specific 3D cache-blocking analysis that focuses on tile reuse over the cache hierarchy, (b) CSO: a macro-kernel that follows CSA to compute the convolution by tiling it, (c) a specialized microkernel that seeks to achieve peak hardware performance, and (d) packing routines for the input tensor and filters to bridge the gap between tiling and micro-kernel. Our approach speeds up end-to-end machine learning model inference by up to 26% and 21% for x86 and POWER10 architectures, respectively. Victor Ferrari, Rafael C. F. Sousa, Márcio Machado Pereira, João P. L. de Carvalho, José Nelson Amaral, Guido Araujo |
PACT | 3 |
| 2021 | Enabling OpenMP Task Parallelism on Multi-FPGAsabstractFPGA-based accelerators have received increasing attention recently. Nevertheless, the amount of resources available on even the most powerful FPGA is still not enough to speed up very large workloads. To achieve that, FPGAs need to be interconnected in a Multi-FPGA architecture. However, programming such architecture is a challenging endeavor. This paper extends the OpenMP task-based computation offloading model to enable several FPGAs to work as a single Multi-FPGA architecture. Experimental results, for a set of OpenMP stencil applications running on a Multi-FPGA platform, have shown close to linear speedups as the number of FPGAs and IP-cores per FPGA increases. Ramon Nepomuceno, Renan Sterle, Guilherme Valarini, Márcio Machado Pereira, Hervé Yviquel, Guido Araujo |
FCCM | 4 |
| 2020 | OmpTracing: Easy Profiling of OpenMP ProgramsabstractOne of the greatest challenges of modern computing is the development of software for parallel execution. To address such challenge, programmers use profiling tools to record relevant operations, like the communications that the different parts of an application carried out during its execution. Profilers can be used to analyze the execution of the application as they enable the programmer to check its performance hot spots and sources of overhead. This paper introduces the OmpTracing library, a lightweight tool that eases the task of profiling OpenMP based applications without the need to inject expensive profiling code into the program. OmpTracing leverages on OMPT, an application programming interface that provides an introspection mechanism of the OpenMP runtime, and that enables the programmer to capture execution details of the parallelized application while generating notifications about significant program events. Vitoria Pinho, Hervé Yviquel, Márcio Machado Pereira, Guido Araujo |
SBAC-PAD | 3 |
| 2019 | Data-flow analysis and optimization for data coherence in heterogeneous architectures
Rafael C. F. Sousa, Márcio Machado Pereira, Fernando Magno Quintão Pereira, Guido Araujo |
J. Parallel Distributed Comput. | 2 |
| 2018 | Automatic Offloading of Cluster AcceleratorsabstractThe sheer amount of computing resources required to run modern cloud workloads has put a lot of pressure on the design of power efficient cluster nodes. To address this problem, Intel (HARP) and Microsoft (Catapult) have proposed CPU-FPGA integrated architectures that can deliver efficient power-performance executions. Unfortunately, the integration of FPGA acceleration modules to software is a challenging endeavor that does not have a seamless programming model. This paper proposes HardCloud (www.hardcloud.org), an extension of the OpenMP 4.X standard that eases the task of offloading FPGA modules to cluster accelerators. Ciro Ceissler, Ramon Nepomuceno, Márcio Machado Pereira, Guido Araujo |
FCCM | 3 |
| 2018 | DOACROSS Parallelization Based on Component Annotation and Loop-Carried ProbabilityabstractAlthough modern compilers implement many loop parallelization techniques, their application is typically restricted to loops that have no loop-carried dependences (DOALL) or that contain well-known structured dependence patterns (e.g. reduction). These restrictions preclude the parallelization of many computational intensive DOACROSS loops. In such loops, either the compiler finds at least one loop-carried dependence or it cannot prove, at compile-time, that the loop is free of such dependences, even though they might never show-up at runtime. In any case, most compilers end-up not parallelizing DOACROSS loops. This paper brings three contributions to address this problem. First, it integrates three algorithms (TLS, DOAX, and BDX) into a simple openMP clause that enables the programmer to select the best algorithm for a given loop. Second, it proposes an annotation approach to separate the sequential components of a loop, thus exposing other components to parallelization. Finally, it shows that loop-carried probability is an effective metric to decide when to use TLS or other non-speculative techniques (e.g. DOAX or BDX) to parallelize DOACROSS loops. Experimental results reveal that, for certain loops, slow-downs can be transformed in 2× speed-ups by quickly selecting the appropriate algorithm. Luis Mattos, Divino Cesar S. Lucas, Juan Salamanca 0001, João P. L. de Carvalho, Márcio Machado Pereira, Guido Araujo |
SBAC-PAD | 5 |
| 2017 | Data Coherence Analysis and Optimization for Heterogeneous ComputingabstractAlthough heterogeneous computing has enabled impressive program speed-ups, knowledge about the architecture of the target device is still critical to reap full hardware benefits. Programming such architectures is complex and is usually done by means of specialized languages (e.g. CUDA, OpenCL). The cost of moving and keeping host/device data coherent may easily eliminate any performance gains achieved by acceleration. Although this problem has been extensively studied for multicore architectures and was recently tackled in discrete GPUs through CUDA8, no generic solution exists for integrated CPU/GPUs architectures like those found in mobile devices (e.g. ARM Mali). This paper proposes Data Coherence Analysis (DCA), a set of two data-flow analyses that determine how variables are used by host/device at each program point. It also introduces Data Coherence Optimization (DCO), a code optimization technique that uses DCA information to: (a) allocate OpenCL shared buffers between host and devices; and (b) insert appropriate OpenCL function calls into program points so as to minimize the number of data coherence operations. DCO was implemented in AClang LLVM (www.aclang.org) a compiler capable of translating OpenMP 4.X annotated loops to OpenCL kernels, thus hiding the complexity of directly programming in OpenCL. Experimental results using DCA and DCO in AClang to compile programs from the Parboil, Polybench and Rodinia benchmarks reveal performance speed-ups of up to 5.25x on an Exynos 8890 Octacore CPU with ARM Mali-T880 MP12 GPU and up to 2.03x on a 2.4 GHz dual-core Intel Core i5 processor equipped with an Intel Iris GPU unit. Rafael C. F. Sousa, Márcio Machado Pereira, Fernando Magno Quintão Pereira, Guido Araujo |
SBAC-PAD | 2 |
| 2017 | DawnCC: Automatic Annotation for Data Parallelism and OffloadingabstractDirective-based programming models, such as OpenACC and OpenMP, allow developers to convert a sequential program into a parallel one with minimum human intervention. However, inserting pragmas into production code is a difficult and error-prone task, often requiring familiarity with the target program. This difficulty restricts the ability of developers to annotate code that they have not written themselves. This article provides a suite of compiler-related methods to mitigate this problem. Such techniques rely on symbolic range analysis, a well-known static technique, to achieve two purposes: populate source code with data transfer primitives and to disambiguate pointers that could hinder automatic parallelization due to aliasing. We have materialized our ideas into a tool, DawnCC, which can be used stand-alone or through an online interface. To demonstrate its effectiveness, we show how DawnCC can annotate the programs available in PolyBench without any intervention from users. Such annotations lead to speedups of over 100× in an Nvidia architecture and over 50× in an ARM architecture. Gleison Souza Diniz Mendonca, Breno Campos Ferreira Guimarães, Péricles Rafael Oliveira Alves, Márcio Machado Pereira, Guido Araujo, Fernando Magno Quintão Pereira |
ACM Trans. Archit. Code Optim. | 4 |
| 2016 | Automatic Insertion of Copy Annotation in Data-Parallel ProgramsabstractDirective-based programming models, such as OpenACC and OpenMP arise today as promising techniques to support the development of parallel applications. These systems allow developers to convert a sequential program into a parallel one with minimum human intervention. However, inserting pragmas into production code is a difficult and error-prone task, often requiring familiarity with the target program. This difficulty restricts the ability of developers to annotate code that they have not written themselves. This paper provides one fundamental component in the solution of this problem. We introduce a static program analysis that infers the bounds of memory regions referenced in source code. Such bounds allow us to automatically insert data-transfer primitives, which are needed when the parallelized code is meant to be executed in an accelerator device, such as a GPU. To validate our ideas, we have applied them onto Polybench, using two different architectures: Nvidia and Qualcomm-based. We have successfully analyzed 98% of all the memory accesses in Polybench. This result has enabled us to insert automatic annotations into those benchmarks leading to speedups of over 100x. Gleison Souza Diniz Mendonca, Breno Campos Ferreira Guimarães, Péricles Rafael Oliveira Alves, Fernando Magno Quintão Pereira, Márcio Machado Pereira, Guido Araujo |
SBAC-PAD | 5 |
| 2016 | Study of hardware transactional memory characteristics and serialization policies on Haswell
Márcio Machado Pereira, Matthew Gaudet, José Nelson Amaral, Guido Araujo |
Parallel Comput. | 1 |
| 2014 | Measuring Effective Work to Reward Success in Dynamic Transaction SchedulingabstractOne of the greatest challenges of modern computing is the development of software optimized for parallel execution in multi-core processors. Transactional Memory (TM) is a new trend in concurrency control that has emerged to address these challenges. TM promises the performance of finer grain locks combined with lower programming complexity. However, transactional memories are speculative and rely on contention managers to resolve conflicts between transactions. This paper explores a complementary approach to boost the performance of TM through the use of schedulers. A TM scheduler is a software component that decides when a particular transaction should be executed. TM scheduling mechanisms are typically restricted to either serialization or yielding. Moreover, their effectiveness is very sensitive to the accuracy of the metric used to predict transaction behavior, particularly in high-contention scenarios. This paper proposes a new Dynamic Transaction Scheduler (DTS) to select a transaction to execute next, based on a new policy that rewards success and uses an improved metric that measures the amount of effective work performed by a transaction. An experimental evaluation indicates that scheduling transactions based on DTS can provide good average-case performance. Márcio Machado Pereira, José Nelson Amaral, Guido Araujo |
ICPP | 1 |
| 2014 | Multi-dimensional Evaluation of Haswell's Transactional Memory PerformanceabstractThis paper presents an extensive performance study of the implementation of Hardware Transactional Memory (HTM) in the Haswell generation of Intel x86 core processors. This study evaluates the strengths and weaknesses of this new architecture exploring several dimensions in the space of Transactional Memory (TM) application characteristics using the Eigenbench [1] and the CLOMP-TM [2] benchmarks. This detailed performance study provides insights on the constraints imposed by the Intel's Transaction Synchronization Extension (Intel's TSX) and introduces a simple, but efficient policy for guaranteeing forward progress on top of the besteffort Intel's HTM and also was critical to achieving performance. The evaluation also shows that there are a number of potential improvements for designers of TM applications and software systems that use Intel's TM and provides recommendations to extract maximum benefit from the current TM support available in Haswell. Márcio Machado Pereira, Matthew Gaudet, José Nelson Amaral, Guido Araujo |
SBAC-PAD | 1 |
| 2013 | Transaction scheduling using conflict avoidance and Contention IntensityabstractIn the last few years, Transactional Memories (TMs) have been shown to be a parallel programming model that can effectively combine performance improvement with ease of programming. Moreover, the recent introduction of TM-based ISA extensions, by major microprocessor manufacturers, also seems to endorse TM as a programming model for today's parallel applications. One of the central issues in designing Software TM (STM) systems is to identify mechanisms/heuristics that can minimize contention arising from conflicting transactions. Although a number of mechanisms have been proposed to tackle contention, such techniques have a limited scope, as conflict is avoided by either interrupting or serializing transaction execution, thus considerably impacting performance. To deal with this limitation, we have proposed a new effective transaction scheduler, along with a conflict-avoidance heuristic, that implements a fully cooperative scheduler that switches a conflicting transaction by another with a lower conflicting probability. This paper extends such framework and introduces a new heuristic, built from the combination of our previous conflict avoidance technique with the Contention Intensity heuristic proposed by Yoo and Lee. Experimental results, obtained using the STMBench7 and STAMP benchmarks atop tinySTM, show that the proposed heuristic produces significant speedups when compared to other four solutions. Márcio Machado Pereira, Alexandro Baldassin, Guido Araujo, Luiz Eduardo Buzato |
HiPC | 1 |