VLDB 2026 Research / reviewers in the wild / expert
Rodrigo Caetano Rocha
dblp:156/4621 · also Rodrigo C. O. Rocha, Rodrigo Caetano de Oliveira Rocha
· DBLP profile ↗
18ranked-venue papers
10as first author
8since 2021 · last 2026
0000-0002-6317-3908ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 13 · 8 first-author · 7 since 2021Systems, architecture and hardware · 11 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Arancini: A Hybrid Binary Translator for Weak Memory Model ArchitecturesabstractBinary translation is a powerful approach to support cross-architecture emulation of unmodified binaries in increasingly heterogeneous computing environments. However, binary translation systems face correctness issues, due to the strong-on-weak memory model mismatch (e.g., from x86-64 to Arm/RISC-V) for concurrent programs. Besides, the current landscape of binary translation systems is fundamentally limited in terms of completeness for static systems and performance for dynamic ones. Sebastian Reimers, Dennis Sprokholt, Martin Fink 0004, Theofilos Augoustis, Simon Kammermeier, Rodrigo Caetano Rocha, Tom Spink, Redha Gouicem, Soham Chakraborty 0001, Pramod Bhatotia |
ASPLOS (2) | 6 |
| 2024 | Toast: A Heterogeneous Memory Management SystemabstractModern applications employ several heterogeneous memory types for improved performance, security, and reliability. To manage them, programmers must currently digress from the traditional load/store interface and rely on various custom libraries specific to each memory type, thus introducing programmability, performance, portability, and protection challenges. Maurice Bailleu, Dimitrios Stavrakakis, Rodrigo Caetano Rocha, Soham Chakraborty 0001, Deepak Garg 0001, Pramod Bhatotia |
PACT | 3 |
| 2023 | Risotto: A Dynamic Binary Translator for Weak Memory Model ArchitecturesabstractDynamic Binary Translation (DBT) is a powerful approach to support cross-architecture emulation of unmodified binaries. However, DBT systems face correctness and performance challenges, when emulating concurrent binaries from strong to weak memory consistency architectures. As a matter of fact, we report several translation errors in QEMU, when emulating x86 binaries on Arm hosts. Redha Gouicem, Dennis Sprokholt, Jasper Ruehl, Rodrigo Caetano Rocha, Tom Spink, Soham Chakraborty 0001, Pramod Bhatotia |
ASPLOS (1) | 4 |
| 2023 | HyBF: A Hybrid Branch Fusion Strategy for Code Size ReductionabstractBinary code size is a first-class design consideration in many computing domains and a critical factor in many more, but compiler optimizations targeting code size are few and often limited in functionality. When size reduction opportunities are left unexploited, it results in higher downstream costs such as memory, storage, bandwidth, or programmer time. Rodrigo Caetano Rocha, Charitha Saumya, Kirshanthan Sundararajah, Pavlos Petoumenos, Milind Kulkarni 0001, Michael F. P. O'Boyle |
CC | 1 |
| 2022 | Loop Rolling for Code Size ReductionabstractCode size is critical for resource-constrained devices, where memory and storage are limited. Compilers, therefore, should offer optimizations aimed at code reduction. One such optimization is loop rerolling, which transforms a partially unrolled loop into a fully rolled one. However, existing techniques are limited and rarely applicable to real-world programs. They are incapable of handling partial rerolling or straight-line code. In this paper, we propose RoLAG, a novel code-size optimization that creates loops out of straight-line code. It identifies isomorphic code by aligning SSA graphs in a bottom-up fashion. The aligned code is later rolled into a loop. In addition, we propose several optimizations that increase the amount of aligned code by identifying specific patterns of code. Finally, an analysis is used to estimate the profitability of the rolled loop before deciding which version should be kept in the code.Our evaluation of RoLAG on full programs from MiBench and SPEC 2017 show absolute reductions of up to 88 KB while LLVM’s technique is hardly distinguishable from the baseline with no rerolling. Finally, our results show that RoLAG is highly applicable to real-world code extracted from popular GitHub repositories. RoLAG is triggered several orders of magnitude more often than LLVM’s rerolling, resulting in meaningful reductions on real-world functions. Rodrigo Caetano Rocha, Pavlos Petoumenos, Björn Franke, Pramod Bhatotia, Michael F. P. O'Boyle |
CGO | 1 |
| 2022 | F3M: Fast Focused Function MergingabstractFrom IoT devices to datacenters, code size is important, motivating ongoing research in binary reduction. A key technique is the merging of similar functions to reduce code redundancy. Success, however, depends on accurately identifying functions that can be profitably merged. Attempting to merge all function pairs is prohibitively expensive. Current approaches, therefore, employ summaries to estimate similarity. However these summaries often give little information about how well two programs will merge. To make things worse, they rely on exhaustive search across all summaries; impractical for realworld programs. In this work, we propose a new technique for matching similar functions. We use a hash-based approach that better captures code similarity and, at the same time, significantly reduces the search space by focusing on the most promising candidates. Experimental results show that our similarity metric has a better correlation with merging profitability. This improves the average code size reduction by 6 percentage points, while it reduces the overhead of function merging by 1.8x on average and by as much as 597x for large applications. Faster merging and reduced code size to compile at later stages mean that our approach introduces little to no compile time overhead, while in many cases it makes compilation faster by up to 30%. Sean Stirling, Rodrigo Caetano Rocha, Kim M. Hazelwood, Hugh Leather, Michael F. P. O'Boyle, Pavlos Petoumenos |
CGO | 2 |
| 2022 | Lasagne: a static binary translator for weak memory model architecturesabstractThe emergence of new architectures create a recurring challenge to ensure that existing programs still work on them. Manually porting legacy code is often impractical. Static binary translation (SBT) is a process where a program’s binary is automatically translated from one architecture to another, while preserving their original semantics. However, these SBT tools have limited support to various advanced architectural features. Importantly, they are currently unable to translate concurrent binaries. The main challenge arises from the mismatches of the memory consistency model specified by the different architectures, especially when porting existing binaries to a weak memory model architecture. Rodrigo Caetano Rocha, Dennis Sprokholt, Martin Fink 0004, Redha Gouicem, Tom Spink, Soham Chakraborty 0001, Pramod Bhatotia |
PLDI | 1 |
| 2021 | HyFM: function merging for freeabstractFunction merging is an important optimization for reducing code size. It merges multiple functions into a single one, eliminating duplicate code among them. The existing state-of-the-art relies on a well-known sequence alignment algorithm to identify duplicate code across whole functions. However, this algorithm is quadratic in time and space on the number of instructions. This leads to very high time overheads and prohibitive levels of memory usage even for medium-sized benchmarks. For larger programs, it becomes impractical. Rodrigo Caetano Rocha, Pavlos Petoumenos, Zheng Wang 0001, Murray Cole, Kim M. Hazelwood, Hugh Leather |
LCTES | 1 |
| 2020 | Vectorization-aware loop unrolling with seed forwardingabstractLoop unrolling is a widely adopted loop transformation, commonly used for enabling subsequent optimizations. Straight-line-code vectorization (SLP) is an optimization that benefits from unrolling. SLP converts isomorphic instruction sequences into vector code. Since unrolling generates repeatead isomorphic instruction sequences, it enables SLP to vectorize more code. However, most production compilers apply these optimizations independently and uncoordinated. Unrolling is commonly tuned to avoid code bloat, not maximizing the potential for vectorization, leading to missed vectorization opportunities. Rodrigo Caetano Rocha, Vasileios Porpodas, Pavlos Petoumenos, Fabrício Góes, Zheng Wang 0001, Murray Cole, Hugh Leather |
CC | 1 |
| 2020 | Effective function merging in the SSA formabstractFunction merging is an important optimization for reducing code size. This technique eliminates redundant code across functions by merging them into a single function. While initially limited to identical or trivially similar functions, the most recent approach can identify all merging opportunities in arbitrary pairs of functions. However, this approach has a serious limitation which prevents it from reaching its full potential. Because it cannot handle phi-nodes, the state-of-the-art applies register demotion to eliminate them before applying its core algorithm. While a superficially minor workaround, this has a three-fold negative effect: by artificially lengthening the instruction sequences to be aligned, it hinders the identification of mergeable instruction; it prevents a vast number of functions from being profitably merged; it increases compilation overheads, both in terms of compile-time and memory usage. Rodrigo Caetano Rocha, Pavlos Petoumenos, Zheng Wang 0001, Murray Cole, Hugh Leather |
PLDI | 1 |
| 2019 | Super-Node SLP: Optimized Vectorization for Code Sequences Containing Operators and Their Inverse ElementsabstractSLP Auto-vectorization converts straight-line code into vector code. It scans input code for groups of instructions that can be combined into vectors and replaces them with their corresponding vector instructions. This work introduces Super-Node SLP (SN-SLP), a new SLP-style algorithm, optimized for expressions that include a commutative operator (such as addition) and its corresponding inverse element (subtraction). SN-SLP uses the algebraic properties of commutative operators and their inverse elements to enable additional transformations that extend auto-vectorization to cases difficult for state-of-the-art auto-vectorizing compilers. We implemented SN-SLP in LLVM. Our evaluation on a real system demonstrates considerable performance improvements of benchmark code with no significant change in compilation time. Vasileios Porpodas, Rodrigo Caetano Rocha, Evgueni Brevnov, Fabrício Góes, Timothy G. Mattson |
CGO | 2 |
| 2019 | Function Merging by Sequence AlignmentabstractResource-constrained devices for embedded systems are becoming increasingly important. In such systems, memory is highly restrictive, making code size in most cases even more important than performance. Compared to more traditional platforms, memory is a larger part of the cost and code occupies much of it. Despite that, compilers make little effort to reduce code size. One key technique attempts to merge the bodies of similar functions. However, production compilers only apply this optimization to identical functions, while research compilers improve on that by merging the few functions with identical control-flow graphs and signatures. Overall, existing solutions are insufficient and we end up having to either increase cost by adding more memory or remove functionality from programs. We introduce a novel technique that can merge arbitrary functions through sequence alignment, a bioinformatics algorithm for identifying regions of similarity between sequences. We combine this technique with an intelligent exploration mechanism to direct the search towards the most promising function pairs. Our approach is more than 2.4x better than the state-of-the-art, reducing code size by up to 25%, with an overall average of 6%, while introducing an average compilation-time overhead of only 15%. When aided by profiling information, this optimization can be deployed without any significant impact on the performance of the generated code. Rodrigo Caetano Rocha, Pavlos Petoumenos, Zheng Wang 0001, Murray Cole, Hugh Leather |
CGO | 1 |
| 2019 | Automatic parallelization of recursive functions with rewriting rules
Rodrigo Caetano Rocha, Fabrício Góes, Fernando Magno Quintão Pereira |
Sci. Comput. Program. | 1 |
| 2018 | VW-SLP: auto-vectorization with adaptive vector widthabstractAuto-vectorization techniques allow the compiler to automatically generate SIMD vector code out of scalar code. SLP is a commonly-used algorithm for converting straight-line code into vector code, which complements the loop-based traditional vectorizers. It works by scanning the input code looking for groups of instructions that can be combined into vectors and replacing them with the corresponding vector instructions. The state-of-the-art SLP algorithm works by attempting to vectorize blocks of code with a fixed vector width and falling back to smaller widths for the whole block upon failure. Vasileios Porpodas, Rodrigo Caetano Rocha, Fabrício Góes |
PACT | 2 |
| 2018 | Look-ahead SLP: auto-vectorization in the presence of commutative operationsabstractAuto-vectorizing compilers automatically generate vector (SIMD) instructions out of scalar code. The state-of-the-art algorithm for straight-line code vectorization is Superword-Level Parallelism (SLP). In this work we identify a major limitation at the core of the SLP algorithm, in the performance-critical step of collecting the vectorization candidate instructions that form the SLP-graph data structure. SLP lacks global knowledge when building its vectorization graph, which negatively affects its local decisions when it encounters commutative instructions. We propose LSLP, an improved algorithm that can plug-in to existing SLP implementations, and can effectively vectorize code with arbitrarily long chains of commutative operations. LSLP relies on short-depth look-ahead for better-informed local decisions. Our evaluation on a real machine shows that LSLP can significantly improve the performance of real-world code with little compilation-time overhead. Vasileios Porpodas, Rodrigo Caetano Rocha, Fabrício Góes |
CGO | 2 |
| 2017 | TOAST: Automatic tiling for iterative stencil computations on GPUsabstractSummary The stencil pattern is important in many scientific and engineering domains, spurring great interest from researchers and industry. In recent years, various optimizations have been proposed for parallel stencil applications running on graphics processing units (GPUs). In particular, tiling is a technique that can significantly enhance application performance by improving data locality and by reducing the volume of communication between host memory and GPU. In addition, tiling enables stencil applications to process inputs that are larger than the physical GPU memory. However, implementing tiling efficiently is complex, time‐consuming, and error‐prone. In this paper, we propose transparently optimized automatic stencil tiling (TOAST), an automatic tiling mechanism for iterative stencil computations running on GPUs; TOAST has 3 main benefits: (1) It incorporates an optimization model that seeks to maximize data reuse within tiles while respecting the amount of dynamically available GPU memory; (2) it offers a virtualized GPU memory for stencil computations, allowing for large input data; and (3) it performs optimal tiling transparently to the developer of the parallel stencil application. The current implementation of TOAST augments the PSkel framework with an internal solver based on genetic algorithms. Our experimental results show that TOAST improves the performance of iterative stencil applications by up to 13 × compared with their multithreaded (central processing unit–based) optimized versions and up to 48 × compared with a naive tiling approach on GPU. The TOAST mechanism is able to automatically achieve a low percentual overhead of data management compared with actual stencil computation. Rodrigo Caetano Rocha, Alyson D. Pereira, Fabrício Góes |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Regent-Dependent Creativity: A Domain Independent Metric for the Assessment of Creative Artifacts
Celso França, Fabrício Góes, Alvaro Amorim, Rodrigo Caetano Rocha, Alysson Ribeiro Da Silva |
ICCC | 4 |
| 2016 | Watershed-ng: an extensible distributed stream processing frameworkabstractSummary Most high‐performance data processing (a.k.a. big data) systems allow users to express their computation using abstractions (like MapReduce), which simplify the extraction of parallelism from applications. Most frameworks, however, do not allow users to specify how communication must take place: That element is deeply embedded into the run‐time system abstractions, making changes hard to implement. In this work, we describe Wathershed‐ng, our re‐engineering of the Watershed system, a framework based on the filter–stream paradigm and originally focused on continuous stream processing. Like other big‐data environments, Watershed provided object‐oriented abstractions to express computation (filters), but the implementation of streams was a run‐time system element. By isolating stream functionality into appropriate classes, combination of communication patterns and reuse of common message handling functions (like compression and blocking) become possible. The new architecture even allows the design of new communication patterns, for example, allowing users to choose MPI, TCP, or shared memory implementations of communication channels as their problem demands. Applications designed for the new interface showed reductions in code size on the order of 50%and above in some cases. The performance results also showed significant improvements, because some implementation bottlenecks were removed in the re‐engineering process. Copyright © 2016 John Wiley & Sons, Ltd. Rodrigo Caetano Rocha, Bruno Hott, Vinícius Vitor dos Santos Dias, Renato Ferreira 0001, Wagner Meira Jr., Dorgival O. Guedes |
Concurr. Comput. Pract. Exp. | 1 |