EDBT 2026 Demo / reviewers in the wild / expert
Vanderson Martins do Rosário
dblp:177/3157 · also Vanderson Martins Rosario, Vanderson Martins do Rosario
· DBLP profile ↗
10ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0002-8737-0252ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fusion of Operators of Computational Graphs via Greedy Clustering: The XNNC ExperienceabstractTensor compilers like XLA, TVM, and TensorRT operate on computational graphs, where vertices represent operations and edges represent data flow between these operations. Operator fusion is an optimization that merges operators to improve their efficiency. This paper presents the operator fusion algorithm recently deployed in the Xtensa Neural Network Compiler (XNNC) - Cadence Tensilica's tensor compiler. The algorithm clusters nodes within the computational graph and iteratively grows these clusters until reaching a fixed point. A priority queue, sorted by the estimated profitability of merging cluster candidates, guides this iterative process. It balances precision and practicality, producing models 39% faster than XNNC's previous fusion approach, which was based on a depth-first traversal of the computational graph. Moreover, unlike recently proposed exhaustive or evolutionary search methods, this algorithm terminates quickly while often yielding equally efficient models. Michael Canesche, Vanderson Martins do Rosário, Edson Borin, Fernando Magno Quintão Pereira |
CC | 2 |
| 2024 | The Droplet Search Algorithm for Kernel SchedulingabstractKernel scheduling is the problem of finding the most efficient implementation for a computational kernel. Identifying this implementation involves experimenting with the parameters of compiler optimizations, such as the size of tiling windows and unrolling factors. This article shows that it is possible to organize these parameters as points in a coordinate space. The function that maps these points to the running time of kernels, in general, will not determine a convex surface. However, this article provides empirical evidence that the origin of this surface (an unoptimized kernel) and its global optimum (the fastest kernel) reside on a convex region. We call this hypothesis the “droplet expectation.” Consequently, a search method based on the Coordinate Descent algorithm tends to find the optimal kernel configuration quickly if the hypothesis holds. This approach—called Droplet Search—has been available in Apache TVM since April of 2023. Experimental results with six large deep learning models on various computing devices (ARM, Intel, AMD, and NVIDIA) indicate that Droplet Search is not only as effective as other AutoTVM search techniques but also 2 to 10 times faster. Moreover, models generated by Droplet Search are competitive with those produced by TVM’s AutoScheduler (Ansor), despite the latter using 4 to 5 times more code transformations than AutoTVM. Michael Canesche, Vanderson Martins do Rosário, Edson Borin, Fernando Magno Quintão Pereira |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | PB3Opt: Profile-based biased Bayesian optimization to select computing clusters on the cloudabstractSummary Given the wide variety of cloud computing resources for creating high‐performance computer clusters and their complex performance relationship with applications, finding the optimal, or near‐optimal, cluster is a complex problem. As a result, several approaches have been proposed to find the optimal, or near‐optimal, cluster for a given high‐performance computing workload, while reducing the search cost. Among the approaches found in the literature, Bayesian optimization is one of the most known and applied. However, it is still possible to increase its performance by integrating it with historical data related to workload behavior. In this context, we suggest the approach, which introduces a bias in the Bayesian optimization expected improvement acquisition function. The new acquisition function uses the ranking of computer clusters of previously explored workloads that have the same behavior as the workload being optimized. Our experimental results show that classifies the behavior of workloads in groups so that the average‐ranking has 88.7% similarity with the ranking of the workload. With this, finds, for almost 95% of workloads, a solution that is less than or equal to 1.2 worse than the optimal computer cluster. In addition, the works well when combined with the paramount iterations technique and is capable of reducing the search cost significantly. Thais Aparecida Silva Camacho, Vanderson Martins do Rosário, Otávio O. Napoli, Edson Borin |
Concurr. Comput. Pract. Exp. | 2 |
| 2023 | Fast selection of compiler optimizations using performance prediction with graph neural networksabstractAbstract Tuning application performance on modern computing infrastructures involves choices in a vast design space as modern computing architectures can have several complex structures impacting performance. Moreover, different applications use these structures in different ways, leading to a challenging performance function. Consequently, it is hard for compilers or experts to find optimal compilation parameters for an application that maximizes such performance function. One approach to tackle this problem is to evaluate many possible optimization plans and select the best among them. However, executing an application to measure its performance for every plan can be very expensive. To tackle this problem, previous work has investigated the use of Machine Learning techniques to predict the performance of the applications without executing them quickly. In this work, we evaluate the use of graph neural networks (GNN) to make fast predictions without executing the application to guide the selection of good optimization sequences. We propose a GNN architecture to make such predictions. We train and test it using 30 thousand different compilation plans applied to 300 different applications, using ARM64 and LLVM IR code representations as input. Our results indicate that the control and data flow graph can then learn features from the control and data flow graph to outperform nongraph‐aware Machine Learning models. Our GNN architecture achieved 91% accuracy in our dataset compared to 79% when using a nongraph‐aware architecture–taking only 16ms to predict a given input. If the application been optimized took an average of 10 s to execute, and we evaluated 1000 optimization sequences, it would take almost 9 h to assess all pairs, but only 16 s with our GNN . Vanderson Martins do Rosário, Anderson Faustino da Silva, André Felipe Zanella, Otávio O. Napoli, Edson Borin |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Employing Simulation to Facilitate the Design of Dynamic Binary TranslatorsabstractDynamic Binary Translation (DBT) is a sophisticated technique that allows the implementation of highperformance ISA emulators. In this technique, the guest code is compiled dynamically at runtime. Consequently, achieving good performance depends on several design decisions, including the shape of the regions of code being translated. Researchers and engineers explore these decisions to bring the best performance possible. However, a real DBT engine is a very sophisticated piece of software, and modifying one is a challenging and demanding task. Hence, we propose using simulation to evaluate the impact of design decisions on dynamic binary translators and present RAIn, an open-source DBT simulator that facilitates the test of DBT's design decisions, such as Region Formation Techniques (RFTs). RAIn outputs several statistics that support the analysis of how design decisions may affect the behavior and the performance of a real DBT. We validated RAIn running a set of experiments with six well-known RFTs (NET, MRET2, LEI, NETPlus, NET-R, and NETPlus-e-r) and showed that it could reproduce well-known results from the literature without the effort of implementing them on a real and thus complex dynamic binary translator engine. Vanderson Martins do Rosário, Raphael Zinsly, Sandro Rigo, Edson Borin |
SBAC-PAD | 1 |
| 2021 | Smart selection of optimizations in dynamic compilersabstractSummary Dynamic compilers perform compilation and generation of target code during runtime, implying that the compilation time is added into the program runtime. Thus, to build a high‐performing dynamic compilation system, it is crucial to be able to generate high‐quality code and, at the same time, have a small compilation cost. In this article, we present an approach that uses machine learning to select sequences of optimization for dynamic compilation that considers both code quality and compilation overhead. Our approach starts by training a model, offline, with a knowledge bank of those sequences with low overhead and high‐quality code generation capability using a genetic heuristic. Then, this bank is used to guide the smart selection of optimizations sequences for the compilation of code fragments during the emulation of an application. We evaluate the proposed strategy in two LLVM‐based dynamic binary translators, namely OI‐DBT and HQEMU, and show that these two translators can achieve average speedups of 1.26x and 1.15x in MiBench and Spec Cpu benchmarks, respectively. Vanderson Martins do Rosário, Anderson Faustino da Silva, Thais Aparecida Silva Camacho, Otávio O. Napoli, Maurício Breternitz, Edson Borin |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Efficiency and scalability of multi-lane capsule networks (MLCN)
Vanderson Martins do Rosário, Maurício Breternitz, Edson Borin |
J. Parallel Distributed Comput. | 1 |
| 2019 | Efficiency and Scalability of Multi-lane Capsule Networks (MLCN)
Vanderson Martins do Rosário, Maurício Breternitz, Edson Borin |
SBAC-PAD | 1 |
| 2019 | The Multi-Lane Capsule NetworkabstractWe introduce multi-lane capsule networks (MLCN), which are a separable and resource efficient organization of capsule networks (CapsNet) that allows parallel processing while achieving high accuracy at reduced cost. A MLCN is composed of a number of (distinct) parallel lanes, each contributing to a dimension of the result, trained using the routing-by-agreement organization of CapsNet. Our results indicate similar accuracy with a much-reduced cost in number of parameters for the Fashion-MNIST and Cifar10 datasets. They also indicate that the MLCN outperforms the original CapsNet when using a proposed novel configuration for the lanes. MLCN also has faster training and inference times, being more than two-fold faster than the original CapsNet in a same accelerator. Vanderson Martins do Rosário, Edson Borin, Maurício Breternitz |
IEEE Signal Process. Lett. | 1 |
| 2017 | Beyond the Fog: Bringing Cross-Platform Code Execution to Constrained IoT DevicesabstractConsidering the prediction that there will be over 50 billion devices connected to the Internet of Things (IoT) in the near future, the demand for efficient ways to process data streams generated by sensors grows ever larger, highlighting the necessity to re-evaluate current approaches, such as sending all data to the cloud for processing and analysis. In this paper, we explore one of the methods for improving this scenario: bringing the computation closer to data sources. By executing the code on the IoT devices themselves instead of on the network edge or the cloud, solutions can better meet the latency requirements of several applications, avoid problems with slow and intermittent network connections, prevent network congestion, and potentially save energy by reducing communication. To this end, we propose the LMC framework and compare it with Edgent, an open-source project that is under development by the Apache Incubator. By using a DragonBoard 410c to execute a simple filter, an outlier detector, and a program that calculates the FFT, we obtained results that indicate that LMC outperforms Edgent when dynamic translation is disabled for both of them and is more suitable for lightweight quick queries otherwise. More importantly, the LMC also enables us to perform cross-platform code execution on small, cheap devices that do not have enough resources to run Edgent, like the NodeMCU 1.0. Flávia Pisani, Jeferson Rech Brunetta, Vanderson Martins do Rosário, Edson Borin |
SBAC-PAD | 3 |