Edson Borin

dblp:55/608 · DBLP profile ↗
← Back
54ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0003-1783-4231ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 27 · 4 first-author · 9 since 2021Software engineering, systems software and programming languages · 10 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 On the Components That Enable Robust Generalization in HAR Models
abstract
Generalizing human activity recognition (HAR) models across datasets remains challenging due to variations in sensors, environments, and user behavior.Domain Generalization (DG) methods attempt to address these shifts through objective-level modifications, architecture-level augmentations, and model pretraining strategies, but prior HAR studies often evaluate these components in isolation using suboptimal baselines.We systematically assess the contribution of each DG component across multiple HAR architectures, from CNNs to Transformers, using the DAGHAR benchmark.Our results show that pretraining the model with a self-supervised learning technique provides the most substantial and consistent gains in cross-dataset generalization, while architecture-level augmentations offer complementary improvements, and objective-level methods alone yield limited benefits across architectures.This suggests that DG studies should treat model pretraining as a standard baseline rather than an optional enhancement.* This project was supported by the
Otávio O. Napoli, Edson Borin
ESANN2
2026 Applying ViT Masked Autoencoders to Seismic Data for Feature Extraction and Few-Shot Learning
abstract
We apply the self-supervised learning technique of vision transformers masked autoencoder (ViT MAE) models with the goal of to producing a feature extractor vision transformer (ViT) backbone for neural networks that receive seismic data as input. We then evaluate the quality of these backbones by coupling them to a simple linear prediction head and fine-tuning these models in a seismic semantic segmentation task. We compare domain-specific ViT MAE against cross-domain pretrained and randomly initialized ViTs, and show that it yields superior performance in low-data regimes. Furthermore, we also demonstrate that pretraining loss correlates with downstream performance, supporting its use as a proxy for feature quality.
Fernando G. Marques, Carlos A. Astudillo, Alan Souza, Daniel Miranda, Edson Borin
IEEE Geosci. Remote. Sens. Lett.5
2025 Homomorphic WiSARDs: Efficient Weightless Neural Network Training over Encrypted Data
Leonardo Neumann, Antonio Guimarães, Diego F. Aranha, Edson Borin
ACNS (3)4
2025 Fusion of Operators of Computational Graphs via Greedy Clustering: The XNNC Experience
abstract
Tensor compilers like XLA, TVM, and TensorRT operate on computational graphs, where vertices represent operations and edges represent data flow between these operations. Operator fusion is an optimization that merges operators to improve their efficiency. This paper presents the operator fusion algorithm recently deployed in the Xtensa Neural Network Compiler (XNNC) - Cadence Tensilica's tensor compiler. The algorithm clusters nodes within the computational graph and iteratively grows these clusters until reaching a fixed point. A priority queue, sorted by the estimated profitability of merging cluster candidates, guides this iterative process. It balances precision and practicality, producing models 39% faster than XNNC's previous fusion approach, which was based on a depth-first traversal of the computational graph. Moreover, unlike recently proposed exhaustive or evolutionary search methods, this algorithm terminates quickly while often yielding equally efficient models.
Michael Canesche, Vanderson Martins do Rosário, Edson Borin, Fernando Magno Quintão Pereira
CC3
2025 On Domain Generalization for Human Activity Recognition with Mix-Based Methods
abstract
Domain generalization (DG) is a challenging problem that involves adapting a model trained on source domains to an unseen target domain.In human activity recognition (HAR), domain shifts often arise from differences in sensor placement, device specifications, or environmental factors, making generalization difficult.In this work, we investigate the effectiveness of mix-based methods like MixStyle and Exact Feature Distribution Mixing (EFDM) when integrated into state-of-the-art models like ResNet and TS2Vec for DG in HAR tasks, leveraging the DAGHAR benchmark.Our results demonstrate that MixStyle significantly outperforms both EFDM and Empirical Risk Minimization approaches, highlighting its effectiveness in addressing domain shifts.* This project was
Otávio O. Napoli, Edson Borin
ESANN2
2025 SPINN: a Tool for Distributed Patch Inference on Massive Data Samples
abstract
Patched inference is a widely used technique in machine learning (ML) that enables fixed-shape models to process arbitrarily large or variably sized inputs by dividing them into smaller, compatible patches. This approach is particularly useful in domains such as seismic processing, medical imaging, and electron microscopy, where data samples often exceed the memory capacity of individual computing nodes. While patched inference is effective for leveraging pre-trained models and operating on resource-constrained hardware, there remains a lack of tools supporting its efficient, distributed execution at scale.To address this gap, we introduce SPINN (Scalable Parallel INference Network), a Python library designed to streamline and accelerate patched inference on high-performance computing (HPC) systems. SPINN supports data partitioning, patch-wise processing using user-defined ML models, and result aggregation, all while leveraging distributed computing frameworks such as Dask and Ray.We validate SPINN on two seismic interpretation tasks, fault detection and facies segmentation, using both public and large-scale private data (up to 272 GB). Experiments demonstrate that SPINN enables smoother prediction outputs via overlapping patches and achieves superlinear scalability with Dask in HPC environments, significantly outperforming conventional solutions such as the NVIDIA Triton Inference Server in large-scale scenarios. SPINN thus emerges as a robust and scalable solution for applying deep learning inference to massive data samples in memory-constrained or compute-intensive settings.
João Seródio, Júlio César Faracco, Fernando Gubitoso, Otávio O. Napoli, Alan Souza, Daniel Miranda, Carlos A. Astudillo, Edson Borin
SBAC-PAD8
2024 A lightweight performance proxy for deep-learning model training on Amazon SageMaker
abstract
Summary Cloud computing has become popular for training deep‐learning (DL) models, avoiding the costs of acquiring and maintaining on‐premise systems. SageMaker is a cloud service that automates the execution of DL workloads. Its features include automatic hyperparameter optimization and use of spot instances. Nonetheless, it does not assist in selecting the right instance type for a workload. In public clouds, rent price depends on the configuration of the chosen instance type. Advanced and faster instances are typically more expensive, but not always the best choice. To select the optimal instance type, users must compare the workload's relative performance (and hence cost) on several candidates. Building on the execution profiles of multiple DL applications, we model the performance and cost of training DL applications on SageMaker and propose a lightweight technique to estimate these at low temporal and monetary cost. This method is a performance proxy that can be used to replace more expensive performance measurement procedures. So, it could speed up any technique that relies on such measurements. We show how it can help cloud customers seeking suitable instance types to train DL models, and that it can accurately predict the performance of different instance types when training these models on SageMaker.
Rafael Keller Tesser, Alvaro Marques, Edson Borin
Concurr. Comput. Pract. Exp.3
2024 Memory-efficient DRASiW Models
Otávio O. Napoli, Ana de Almeida 0002, Edson Borin, Maurício Breternitz
Neurocomputing3
2024 BrkgaCuda 2.0: a framework for fast biased random-key genetic algorithms on GPUs
Bruno A. Oliveira, Eduardo C. Xavier, Edson Borin
Soft Comput.3
2024 The Droplet Search Algorithm for Kernel Scheduling
abstract
Kernel scheduling is the problem of finding the most efficient implementation for a computational kernel. Identifying this implementation involves experimenting with the parameters of compiler optimizations, such as the size of tiling windows and unrolling factors. This article shows that it is possible to organize these parameters as points in a coordinate space. The function that maps these points to the running time of kernels, in general, will not determine a convex surface. However, this article provides empirical evidence that the origin of this surface (an unoptimized kernel) and its global optimum (the fastest kernel) reside on a convex region. We call this hypothesis the “droplet expectation.” Consequently, a search method based on the Coordinate Descent algorithm tends to find the optimal kernel configuration quickly if the hypothesis holds. This approach—called Droplet Search—has been available in Apache TVM since April of 2023. Experimental results with six large deep learning models on various computing devices (ARM, Intel, AMD, and NVIDIA) indicate that Droplet Search is not only as effective as other AutoTVM search techniques but also 2 to 10 times faster. Moreover, models generated by Droplet Search are competitive with those produced by TVM’s AutoScheduler (Ansor), despite the latter using 4 to 5 times more code transformations than AutoTVM.
Michael Canesche, Vanderson Martins do Rosário, Edson Borin, Fernando Magno Quintão Pereira
ACM Trans. Archit. Code Optim.3
2024 Rank-based Hashing for Effective and Efficient Nearest Neighbor Search for Image Retrieval
abstract
The large and growing amount of digital data creates a pressing need for approaches capable of indexing and retrieving multimedia content. A traditional and fundamental challenge consists of effectively and efficiently performing nearest-neighbor searches. After decades of research, several different methods are available, including trees, hashing, and graph-based approaches. Most of the current methods exploit learning to hash approaches based on deep learning. In spite of effective results and compact codes obtained, such methods often require a significant amount of labeled data for training. Unsupervised approaches also rely on expensive training procedures usually based on a huge amount of data. In this work, we propose an unsupervised data-independent approach for nearest neighbor searches, which can be used with different features, including deep features trained by transfer learning. The method uses a rank-based formulation and exploits a hashing approach for efficient ranked list computation at query time. A comprehensive experimental evaluation was conducted on seven public datasets, considering deep features based on CNNs and Transformers. Both effectiveness and efficiency aspects were evaluated. The proposed approach achieves remarkable results in comparison to traditional and state-of-the-art methods. Hence, it is an attractive and innovative solution, especially when costly training procedures need to be avoided.
Vinicius Atsushi Sato Kawai, Lucas Pascotti Valem, Alexandro Baldassin, Edson Borin, Daniel C. G. Pedronette, Longin Jan Latecki
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Efficient Knowledge Aggregation Methods for Weightless Neural Networks
abstract
Weightless Neural Networks (WNN) are good candidates for Federated Learning scenarios due to their robustness and computational lightness.In this work, we show that it is possible to aggregate the knowledge of multiple WNNs using more compact data structures, such as Bloom Filters, to reduce the amount of data transferred between devices.Finally, we explore variations of Bloom Filters and found that a particular data-structure, the Count-Min Sketch (CMS), is a good candidate for aggregation.Costing at most 3% of accuracy, CMS can be up to 3x smaller when compared to previous approaches, specially for large datasets.
Otávio O. Napoli, Ana de Almeida 0002, José Miguel Salles Dias, Luís Brás Rosário, Edson Borin, Maurício Breternitz
ESANN5
2023 A Self-Distributing System Framework for the Computing Continuum
abstract
Applications such as autonomous vehicles, virtual reality, augmented reality, and heavy machine learning-based applications are becoming popular and demanding more flexible deployment environments. The computing continuum, a hierarchical hybrid infrastructure comprehending user devices (smartphones, sensors, laptops, etc.), edge data centers, and cloud platforms, offers a wide range of deployment possibilities with a full range of varying computing resources. To take full advantage of such infrastructure, application development is faced with many challenges, the most important being the implementation of a transparent and generalized mechanism for code offloading and mobility throughout the continuum. To tackle such issues, this paper presents the Self-Distributing Systems (SDS) framework, a self-distribution framework that supports generalized code-offloading capabilities at the application level with a machine learning agent for deciding where to place components and a component-based model to enable seamless distribution of an application's components at runtime. We describe the framework, show its applicability in different application scenarios, and report our preliminary results. We conclude the paper with a list of challenges and invite the systems community to join the effort to further investigate them.
Roberto Rodrigues Filho, Renato S. Dias, João Seródio, Barry Porter, Fábio M. Costa, Edson Borin, Luiz Fernando Bittencourt
ICCCN6
2023 PB3Opt: Profile-based biased Bayesian optimization to select computing clusters on the cloud
abstract
Summary Given the wide variety of cloud computing resources for creating high‐performance computer clusters and their complex performance relationship with applications, finding the optimal, or near‐optimal, cluster is a complex problem. As a result, several approaches have been proposed to find the optimal, or near‐optimal, cluster for a given high‐performance computing workload, while reducing the search cost. Among the approaches found in the literature, Bayesian optimization is one of the most known and applied. However, it is still possible to increase its performance by integrating it with historical data related to workload behavior. In this context, we suggest the approach, which introduces a bias in the Bayesian optimization expected improvement acquisition function. The new acquisition function uses the ranking of computer clusters of previously explored workloads that have the same behavior as the workload being optimized. Our experimental results show that classifies the behavior of workloads in groups so that the average‐ranking has 88.7% similarity with the ranking of the workload. With this, finds, for almost 95% of workloads, a solution that is less than or equal to 1.2 worse than the optimal computer cluster. In addition, the works well when combined with the paramount iterations technique and is capable of reducing the search cost significantly.
Thais Aparecida Silva Camacho, Vanderson Martins do Rosário, Otávio O. Napoli, Edson Borin
Concurr. Comput. Pract. Exp.4
2023 Fast selection of compiler optimizations using performance prediction with graph neural networks
abstract
Abstract Tuning application performance on modern computing infrastructures involves choices in a vast design space as modern computing architectures can have several complex structures impacting performance. Moreover, different applications use these structures in different ways, leading to a challenging performance function. Consequently, it is hard for compilers or experts to find optimal compilation parameters for an application that maximizes such performance function. One approach to tackle this problem is to evaluate many possible optimization plans and select the best among them. However, executing an application to measure its performance for every plan can be very expensive. To tackle this problem, previous work has investigated the use of Machine Learning techniques to predict the performance of the applications without executing them quickly. In this work, we evaluate the use of graph neural networks (GNN) to make fast predictions without executing the application to guide the selection of good optimization sequences. We propose a GNN architecture to make such predictions. We train and test it using 30 thousand different compilation plans applied to 300 different applications, using ARM64 and LLVM IR code representations as input. Our results indicate that the control and data flow graph can then learn features from the control and data flow graph to outperform nongraph‐aware Machine Learning models. Our GNN architecture achieved 91% accuracy in our dataset compared to 79% when using a nongraph‐aware architecture–taking only 16ms to predict a given input. If the application been optimized took an average of 10 s to execute, and we evaluated 1000 optimization sequences, it would take almost 9 h to assess all pairs, but only 16 s with our GNN .
Vanderson Martins do Rosário, Anderson Faustino da Silva, André Felipe Zanella, Otávio O. Napoli, Edson Borin
Concurr. Comput. Pract. Exp.5
2023 Containers in HPC: a survey
Rafael Keller Tesser, Edson Borin
J. Supercomput.2
2022 An evaluation of fast segmented sorting implementations on GPUs
Rafael F. Schmid, Flávia Pisani, Edson Cáceres, Edson Borin
Parallel Comput.4
2021 Employing Simulation to Facilitate the Design of Dynamic Binary Translators
abstract
Dynamic Binary Translation (DBT) is a sophisticated technique that allows the implementation of highperformance ISA emulators. In this technique, the guest code is compiled dynamically at runtime. Consequently, achieving good performance depends on several design decisions, including the shape of the regions of code being translated. Researchers and engineers explore these decisions to bring the best performance possible. However, a real DBT engine is a very sophisticated piece of software, and modifying one is a challenging and demanding task. Hence, we propose using simulation to evaluate the impact of design decisions on dynamic binary translators and present RAIn, an open-source DBT simulator that facilitates the test of DBT's design decisions, such as Region Formation Techniques (RFTs). RAIn outputs several statistics that support the analysis of how design decisions may affect the behavior and the performance of a real DBT. We validated RAIn running a set of experiments with six well-known RFTs (NET, MRET2, LEI, NETPlus, NET-R, and NETPlus-e-r) and showed that it could reproduce well-known results from the literature without the effort of implementing them on a real and thus complex dynamic binary translator engine.
Vanderson Martins do Rosário, Raphael Zinsly, Sandro Rigo, Edson Borin
SBAC-PAD4
2021 High-performance IO for seismic processing on the cloud
abstract
Summary Most of the applications in the seismology field rely on the processing of up to hundreds of terabytes of data and their performance is strongly affected by IO operations. In this article, we analyze the main file structures currently used to store seismic data and propose a new intermediate data structure to improve IO performance while still complying with established standards. We show that, throughout a common workflow in seismic data analysis, our IO performance gain greatly surpasses the overhead of translating data to the intermediate structure. This approach enables a speedup of up to 208 times in reading time when using classical standards (e.g., SEG‐Y) and our intermediate structure is up to 1.8 times more efficient than modern formats (e.g., ASDF). Considering cache‐friendly applications, our speedups over the direct use of SEG‐Y reach 8000 times. We also performed a cost analysis on the AWS cloud showing that, in our approach, HDDs can be 1.25 times more cost‐effective than SSDs.
Antonio Guimarães, Luis Lacalle, Charles Boulhosa Rodamilans, Edson Borin
Concurr. Comput. Pract. Exp.4
2021 Smart selection of optimizations in dynamic compilers
abstract
Summary Dynamic compilers perform compilation and generation of target code during runtime, implying that the compilation time is added into the program runtime. Thus, to build a high‐performing dynamic compilation system, it is crucial to be able to generate high‐quality code and, at the same time, have a small compilation cost. In this article, we present an approach that uses machine learning to select sequences of optimization for dynamic compilation that considers both code quality and compilation overhead. Our approach starts by training a model, offline, with a knowledge bank of those sequences with low overhead and high‐quality code generation capability using a genetic heuristic. Then, this bank is used to guide the smart selection of optimizations sequences for the compilation of code fragments during the emulation of an application. We evaluate the proposed strategy in two LLVM‐based dynamic binary translators, namely OI‐DBT and HQEMU, and show that these two translators can achieve average speedups of 1.26x and 1.15x in MiBench and Spec Cpu benchmarks, respectively.
Vanderson Martins do Rosário, Anderson Faustino da Silva, Thais Aparecida Silva Camacho, Otávio O. Napoli, Maurício Breternitz, Edson Borin
Concurr. Comput. Pract. Exp.6
2021 Efficiency and scalability of multi-lane capsule networks (MLCN)
Vanderson Martins do Rosário, Maurício Breternitz, Edson Borin
J. Parallel Distributed Comput.3
2020 A unified model for accelerating unsupervised iterative re-ranking algorithms
abstract
Summary Despite the continuous advances in image retrieval technologies, performing effective and efficient content‐based searches remains a challenging task. Unsupervised iterative re‐ranking algorithms have emerged as a promising solution and have been widely used to improve the effectiveness of multimedia retrieval systems. Although substantially more efficient than related approaches based on diffusion processes, these re‐ranking algorithms can still be computationally costly, demanding the specification and implementation of efficient big multimedia analysis approaches. Such demand associated with the significant potential for parallelization and highly effective results achieved by recently proposed re‐ranking algorithms creates the need for exploiting efficiency vs effectiveness trade‐offs. In this article, we introduce a class of unsupervised iterative re‐ranking algorithms and present a model that can be used to guide their implementation and optimization for parallel architectures. We also analyze the impact of the parallelization on the performance of four algorithms that belong to the proposed class: Contextual Spaces, RL‐Sim, Contextual Re‐ranking, and Cartesian Product of Ranking References. The experiments show speedups that reach up to 6.0×, 16.1×, 3.3×, and 7.1× for each algorithm, respectively. These results demonstrate that the proposed parallel programming model can be successfully applied to various algorithms and used to improve the performance of multimedia retrieval systems.
Flávia Pisani, Lucas Pascotti Valem, Daniel C. G. Pedronette, Ricardo da Silva Torres, Edson Borin, Maurício Breternitz
Concurr. Comput. Pract. Exp.5
2019 Improving Virtual Machine Consolidation for Heterogeneous Cloud Computing Datacenters
abstract
Datacenters in cloud computing systems may consist of thousands of computing nodes and communication devices, requiring virtual machine placement to optimize resources utilization. Heterogeneous environments of current datacenters have not been properly evaluated yet when the number of processing nodes and network devices is considered. This paper presents an approach to place virtual machines to minimize both the number of employed computing nodes and communication costs in heterogeneous environments. We evaluate our results against other known algorithms to show the efficiency of the proposed approach, which outperforms a state-of-the-art algorithm with a result of up to 13% better for network improvement while achieves the same number of computing nodes.
João Antonio Magri Rodrigues, Fabíola Martins Campos de Oliveira, Renata Spolon Lobato, Roberta Spolon Ulson, Aleardo Manacero, Edson Borin
SBAC-PAD6
2019 Efficiency and Scalability of Multi-lane Capsule Networks (MLCN)
Vanderson Martins do Rosário, Maurício Breternitz, Edson Borin
SBAC-PAD3
2019 Optimized implementation of QC-MDPC code-based cryptography
abstract
Summary This paper presents a new enhanced version of the QcBits key encapsulation mechanism, which is a constant‐time implementation of the Niederreiter cryptosystem using QC‐MDPC codes. In this version, we updated the implementation parameters to meet the 128‐bit quantum security level, replaced some of the core algorithms to avoid using slower instructions, vectorized the entire code using the AVX‐512 instruction set extension, and applied several other techniques to achieve a competitive performance level. Our implementation takes 928, 259, and 5008 thousand Skylake cycles to perform batch key generation (cost per key), encryption, and uniform decryption, respectively. Comparing with the current state‐of‐the‐art implementation for QC‐MDPC codes, BIKE, our code is 1.9 times faster when decrypting messages.
Antonio Guimarães, Diego F. Aranha, Edson Borin
Concurr. Comput. Pract. Exp.3
2019 The Multi-Lane Capsule Network
abstract
We introduce multi-lane capsule networks (MLCN), which are a separable and resource efficient organization of capsule networks (CapsNet) that allows parallel processing while achieving high accuracy at reduced cost. A MLCN is composed of a number of (distinct) parallel lanes, each contributing to a dimension of the result, trained using the routing-by-agreement organization of CapsNet. Our results indicate similar accuracy with a much-reduced cost in number of parameters for the Fashion-MNIST and Cifar10 datasets. They also indicate that the MLCN outperforms the original CapsNet when using a proposed novel configuration for the lanes. MLCN also has faster training and inference times, being more than two-fold faster than the original CapsNet in a same accelerator.
Vanderson Martins do Rosário, Edson Borin, Maurício Breternitz
IEEE Signal Process. Lett.2
2018 Evaluating the Performance and Cost of Accelerating Seismic Processing with CUDA, OpenCL, OpenACC, and OpenMP
abstract
The Common Midpoint and Common Reflection Surface methods are computationally demanding seismic processing techniques for improving signal-to-noise ratios. In this paper, we discuss the performance results and the cost-benefit of accelerating these two procedures using CUDA, OpenCL, OpenACC, and OpenMP on CPUs and GPUs. We obtained results on server-class CPUs and state-of-the-art GPUs that show that, while OpenCL and CUDA present the best performance results on GPUs, OpenACC can also be an interesting choice due to its programmability. Among the tested accelerators, GPUs with the Pascal microarchitecture showed the best results for the tested seismic processing methods in terms of raw performance, energy efficiency, and performance per price.
Tiago Lobato Gimenes, Flávia Pisani, Edson Borin
IPDPS3
2018 The Alberta Workloads for the SPEC CPU 2017 Benchmark Suite
abstract
A proper evaluation of techniques that require multiple training and evaluation executions of a benchmark, such as Feedback-Directed Optimization (FDO), requires multiple workloads that can be used to characterize variations on the behaviour of a program based on the workload. This paper aims to improve the performance evaluation of computer systems - including compilers, computer architecture simulation, and operating-system prototypes - that rely on the industrystandard SPEC CPU benchmark suite. A main concern with the use of this suite in research is that it is distributed with a very small number of workloads. This paper describes the process to create additional workloads for this suite and offers useful insights in many of its benchmarks. The set of additional workloads created, named the Alberta Workloads for the SPEC CPU 2017 Benchmark Suite1 is made freely available with the goal of providing additional data points for the exploration of learning in computing systems. These workloads should also contribute to ameliorate the hidden learning problem where a researcher sets parameters to a system during development based on a set of benchmarks and then evaluates the system using the very same set of benchmarks with the very same workloads.
José Nelson Amaral, Edson Borin, Dylan R. Ashley, Caian Benedicto, Elliot Colp, Joao Henrique Stange Hoffmam, Marcus Karpoff, Erick Ochoa 0001, Morgan Redshaw, Raphael Ernani Rodrigues
ISPASS2
2018 Partitioning Convolutional Neural Networks for Inference on Constrained Internet-of-Things Devices
abstract
With the prospects of a world in which the IoT will be pervasive in a near future, the great amount of data produced by its devices will have to be processed and interpreted in an efficient and intelligent way. One approach to do that is the use of fog computing, in which the network infrastructure and the devices themselves can process data. Deep learning techniques have been successfully applied to the interpretation of the kind of data generated by the IoT, however, even the inference execution of convolutional neural networks may be computationally costly when resource-limited devices are considered. In order to enable the execution of neural network models on resource-constrained IoT systems, the code may be partitioned and distributed among multiple devices. Different partitioning approaches are possible, nonetheless, some of them increase the amount of communication that needs to be performed between the IoT devices. In this work, we propose KLP, a Kernighan-and-Lin-based partitioning algorithm that partitions neural network models for efficient distributed execution on multiple IoT devices. Our results show that KLP is capable of producing partitions that require up to 4.5 times less communication than partitioning approaches used by TensorFlow and other frameworks.
Fabíola Martins Campos de Oliveira, Edson Borin
SBAC-PAD2
2018 Special issue on Computer Architecture and High Performance Computing
Lúcia M. A. Drummond, Edson Borin
J. Parallel Distributed Comput.2
2017 The Case for Flexible ISAs: Unleashing Hardware and Software
abstract
For a long time the Instruction Set Architecture (ISA) has been the firm contract between software and hardware. This firm contract plays an important role by decoupling the development of software from hardware micro-architectural features, enabling both to evolve independently. Nonetheless, it also condemns the ISA to become larger, more cluttered and inefficient as new instructions are incorporated over the years and deprecated instructions are left untouched to keep legacy compatibility. In this work we propose OpenISA, a flexible ISA that enables both the software and the hardware to evolve independently and discuss how OpenISA 1.0 was designed to enable efficient OpenISA software emulation on alien ISAs, which is key to free the user from hardware lock-ins. Our results show that software compiled to OpenISA can be latter emulated on x86 and ARM processors with very little overhead achieving near native performance, under 10% for the majority of programs.
Rafael Auler, Edson Borin
SBAC-PAD2
2017 Beyond the Fog: Bringing Cross-Platform Code Execution to Constrained IoT Devices
abstract
Considering the prediction that there will be over 50 billion devices connected to the Internet of Things (IoT) in the near future, the demand for efficient ways to process data streams generated by sensors grows ever larger, highlighting the necessity to re-evaluate current approaches, such as sending all data to the cloud for processing and analysis. In this paper, we explore one of the methods for improving this scenario: bringing the computation closer to data sources. By executing the code on the IoT devices themselves instead of on the network edge or the cloud, solutions can better meet the latency requirements of several applications, avoid problems with slow and intermittent network connections, prevent network congestion, and potentially save energy by reducing communication. To this end, we propose the LMC framework and compare it with Edgent, an open-source project that is under development by the Apache Incubator. By using a DragonBoard 410c to execute a simple filter, an outlier detector, and a program that calculates the FFT, we obtained results that indicate that LMC outperforms Edgent when dynamic translation is disabled for both of them and is more suitable for lightweight quick queries otherwise. More importantly, the LMC also enables us to perform cross-platform code execution on small, cheap devices that do not have enough resources to run Edgent, like the NodeMCU 1.0.
Flávia Pisani, Jeferson Rech Brunetta, Vanderson Martins do Rosário, Edson Borin
SBAC-PAD4
2017 Handling IoT platform heterogeneity with COISA, a compact OpenISA virtual platform
abstract
Summary In face of the high number of different hardware platforms we need to program with Internet‐of‐Things (IoT), virtual machines (VMs) pose as a promising technology to allow a program once, deploy everywhere strategy. Unfortunately, existing VMs are either too heavy or use a stripped‐down version to work on resource‐constrained IoT devices. We present COISA, a compact virtual platform that relies on OpenISA, an instruction set architecture (ISA) that strives for easy emulation, to allow a single program to be deployed on many platforms, including tiny microcontrollers. By exploring the benefits of using a concrete ISA as our VM language, our experimental results indicate that COISA is easily portable and is capable of running unmodified guest applications in highly heterogeneous host platforms, including one with only 2 kB of RAM. For time‐critical IoT applications on constrained platforms where extracting performance is of paramount importance, we propose the use of cloud‐assisted translations, which employ static binary translation to deliver a binary fully converted to the native ISA used in the IoT device. Copyright © 2016 John Wiley & Sons, Ltd.
Rafael Auler, Carlos Eduardo Millani, Alexandre Brisighello, Alisson Linhares, Edson Borin
Concurr. Comput. Pract. Exp.5
2017 Contextual Spaces Re-Ranking: accelerating the Re-sort Ranked Lists step on heterogeneous systems
abstract
Summary Re‐ranking algorithms have been proposed to improve the effectiveness of content‐based image retrieval systems by exploiting contextual information encoded in distance measures and ranked lists. In this paper, we show how we improved the efficiency of one of these algorithms, called Contextual Spaces Re‐Ranking (CSRR). One of our approaches consists in parallelizing the algorithm with OpenCL to use the central and graphics processing units of an accelerated processing unit. The other is to modify the algorithm to a version that, when compared with the original CSRR, not only reduces the total running time of our implementations by a median of 1.6 × but also increases the accuracy score in most of our test cases. Combining both parallelization and algorithm modification results in a median speedup of 5.4 × from the original serial CSRR to the parallelized modified version. Different implementations for CSRR's Re‐sort Ranked Lists step were explored as well, providing insights into graphics processing unit sorting, the performance impact of image descriptors, and the trade‐offs between effectiveness and efficiency. Copyright © 2016 John Wiley & Sons, Ltd.
Flávia Pisani, Daniel C. G. Pedronette, Ricardo da Silva Torres, Edson Borin
Concurr. Comput. Pract. Exp.4
2015 SHRINK: reducing the ISA complexity via instruction recycling
abstract
Microprocessor manufacturers typically keep old instruction sets in modern processors to ensure backward compatibility with legacy software. The introduction of newer extensions to the ISA increases the design complexity of microprocessor front-ends, exacerbates the consumption of precious on-chip resources (e.g., silicon area and energy), and demands more efforts for hardware verification and debugging. We analyzed several x86 applications and operating systems deployed between 1995 and 2012 and observed that many instructions stop being used over time, and more than 500 instructions were never used in these applications. We also investigate the impact of including these unused instructions in the design of the x86 decoders and propose SHRINK, a mechanism to remove old instructions without breaking backward compatibility with legacy code. SHRINK allows us to remove 40% of the instructions from the x86 ISA and improve the critical path, area, and power consumption of the instruction decoder, respectively, by 23%, 48%, and 49%, on average.
Bruno Cardoso Lopes, Rafael Auler, Edson Borin, Rodolfo Azevedo
ISCA4
2015 Effective, Efficient, and Scalable Unsupervised Distance Learning in Image Retrieval Tasks
abstract
Various unsupervised learning methods have been proposed with significant improvements in the effectiveness of image search systems. However, despite the relevant effectiveness gains, these approaches commonly require high computation efforts, not addressing properly efficiency and scalability requirements. In this paper, we present a novel unsupervised learning approach for improving the effectiveness of image retrieval tasks. The proposed method is also scalable and efficient as it exploits parallel and heterogeneous computing on CPU and GPU devices. Extensive experiments were conducted considering five different public image collections and several descriptors. This rigorous experimental protocol evaluates the effectiveness, efficiency, and scalability of the proposed approach, and compares it with previous methods. Experimental results demonstrate that high effectiveness gains (up to +29%) can be obtained requiring small run times.
Lucas Pascotti Valem, Daniel C. G. Pedronette, Ricardo da Silva Torres, Edson Borin, Jurandy Almeida
ICMR4
2015 Performance implications of dynamic memory allocators on transactional memory systems
abstract
Although dynamic memory management accounts for a significant part of the execution time on many modern software systems, its impact on the performance of transactional memory systems has been mostly overlooked. In order to shed some light into this subject, this paper conducts a thorough investigation of the interplay between memory allocators and software transactional memory (STM) systems. We show that allocators can interfere with the way memory addresses are mapped to versioned locks on state-of-the-art software transactional memory implementations. Moreover, we observed that key aspects of allocators such as false sharing avoidance, scalability, and locality have a drastic impact on the final performance. For instance, we have detected performance differences of up to 171% in the STAMP applications when using distinct allocators. Moreover, we show that optimizations at the STM-level (such as caching transactional objects) are not effective when a modern allocator is already in use. All in all, our study highlights the importance of reporting the allocator utilized in the performance evaluation of transactional memory systems.
Alexandro Baldassin, Edson Borin, Guido Araujo
PPoPP2
2014 Compiler support for selective page migration in NUMA architectures
abstract
Current high-performance multicore processors provide users with a non-uniform memory access model (NUMA). These systems perform better when threads access data on memory banks next to the core where they run. However, ensuring data locality is difficult. In this paper, we propose compiler analyses and code generation methods to support a lightweight runtime system that dynamically migrates memory pages to improve data locality. Our technique combines static and dynamic analyses and is capable of identifying the most promising pages to migrate. Statically, we infer the size of arrays, plus the amount of reuse of each memory access instruction in a program. These estimates rely on a simple, yet accurate, trip count predictor of our own design. This knowledge lets us build templates of dynamic checks, to be filled with values known only at runtime. These checks determine when it is profitable to migrate data closer to the processors where this data is used. Our static analyses are quadratic on the number of variables in a program, and the dynamic checks are O(1) in practice. Our technique does not require any form of user intervention, neither the support of a third-party middleware, nor modifications in the operating system's kernel. We have applied our technique on several parallel algorithms, which are completely oblivious to the asymmetric memory topology, and have observed speedups of up to 4x, compared to static heuristics. We compare our approach against Minas, a middleware that supports NUMA-aware data allocation, and show that we can outperform it by up to 50% in some cases.
Guilherme Piccoli, Henrique Nazaré, Raphael Ernani Rodrigues, Christiane Pousa, Edson Borin, Fernando Magno Quintão Pereira
PACT5
2014 Addressing JavaScript JIT Engines Performance Quirks: A Crowdsourced Adaptive Compiler
Rafael Auler, Edson Borin, Jonathan de Halleux, Michal Moskal, Nikolai Tillmann
CC2
2014 Leveraging Optimization Methods for Dynamically Assisted Control-Flow Integrity Mechanisms
abstract
Dynamic Binary Modification (DBM) tools are useful for cross-platform execution of binaries and are powerful run time environments that allow execution optimizations, instrumentation and profiling. These tools have also been used as enablers for control-flow integrity verification, a process that consists in the observation and analysis of a program's execution path focusing on the detection of anomalies, such as those arising from flow corruption based software attacks. Even though this class of tools helps us in identifying a myriad of attacks, it is typically expensive at run time and introduce significant overhead to the program execution. Considering their inherent high cost, further expanding the capabilities of such tools for detection of program flow anomalies can slow down the analysis to the point that it is unfeasible to run it in real world workflows. In this paper we present a mechanism for including program flow verification in DBMs that uses asynchronous analysis and applies different parallel-programming techniques that leverage current multi-core systems to control the overhead of our analysis. Our mechanism was tested against synthetic program flow corruption use cases and correctly detected all detours. With our new optimizations, we show that our system achieves an slowdown of only 1.46x, while a naively implemented verification system face 4.22x of overhead.
Lucas Teixeira, Edson Borin, Sandro Rigo
SBAC-PAD3
2013 Image Re-ranking Acceleration on GPUs
abstract
Huge image collections are becoming available lately. In this scenario, the use of Content-Based Image Retrieval (CBIR) systems has emerged as a promising approach to support image searches. The objective of CBIR systems is to retrieve the most similar images in a collection, given a query image, by taking into account image visual properties such as texture, color, and shape. In these systems, the effectiveness of the retrieval process depends heavily on the accuracy of ranking approaches. Recently, re-ranking approaches have been proposed to improve the effectiveness of CBIR systems by taking into account the relationships among images. The re-ranking approaches consider the relationships among all images in a given dataset. These approaches typically demands a huge amount of computational power, which hampers its use in practical situations. On the other hand, these methods can be massively parallelized. In this paper, we propose to speedup the computation of the RL-Sim algorithm, a recently proposed image re-ranking approach, by using the computational power of Graphics Processing Units (GPU). GPUs are emerging as relatively inexpensive parallel processors that are becoming available on a wide range of computer systems. We address the image re-ranking performance challenges by proposing a parallel solution designed to fit the computational model of GPUs. We conducted an experimental evaluation considering different implementations and devices. Experimental results demonstrate that significant performance gains can be obtained. Our approach achieves speedups of 7x from serial implementation considering the overall algorithm and up to 36x on its core steps.
Daniel C. G. Pedronette, Ricardo da Silva Torres, Edson Borin, Maurício Breternitz
SBAC-PAD3
2013 Assessing computer performance with stocs
abstract
Several aspects of a computer system cause performance measurements to include random errors. Moreover, these systems are typically composed of a non-trivial combination of individual components that may cause one system to perform better or worse than another depending on the workload. Hence, properly measuring and comparing computer systems performance are non-trivial tasks.
Leonardo Piga, Gabriel F. T. Gomes, Rafael Auler, Bruno Rosa 0001, Sandro Rigo, Edson Borin
ICPE6
2013 An automatic energy consumption characterization of processors using ArchC
Marcelo Guedes, Rafael Auler, Liana Dessandre Duenha, Edson Borin, Rodolfo Azevedo
J. Syst. Archit.4
2012 Efficient Image Re-Ranking Computation on GPUs
abstract
The huge growth of image collections and multimedia resources available is remarkable. One of the most common approaches to support image searches relies on the use of Content-Based Image Retrieval (CBIR) systems. CBIR systems aim at retrieving the most similar images in a collection, given a query image. Since the effectiveness of those systems is very dependent on the accuracy of ranking approaches, re-ranking algorithms have been proposed to exploit contextual information and improve the effectiveness of CBIR systems. Image re-ranking algorithms typically consider the relationship among every image in a given dataset when computing the new ranking. This approach demands a huge amount of computational power, which may render it prohibitive on very large data sets. In order to mitigate this problem, we propose using the computational power of Graphics Processing Units (GPU) to speedup the computation of image re-ranking algorithms. GPUs are fast emerging and relatively inexpensive parallel processors that are becoming available on a wide range of computer systems. In this paper, we propose a parallel implementation of an image re-ranking algorithm designed to fit the computational model of GPUs. Experimental results demonstrate that relevant performance gains can be obtained by our approach.
Daniel C. G. Pedronette, Ricardo da Silva Torres, Edson Borin, Maurício Breternitz
ISPA3
2012 An ArchC approach for automatic energy consumption characterization of processors
abstract
This paper presents acSynth, an ArchC framework for energy characterization and simulation. Based on Tiwari's Method, a subject processor is characterized in an affordable time and the information is fed into acSynth to bring architecture level power analysis. The framework provides power reports and energy profiling. The experimental results show the characterization flow for the Plasma processor, a MIPS-I HDL description. The acSynth can provide power analysis at 35.1 million instructions per second in simulation with small accuracy diversion and without loss of generality. The system allows the execution of large tests in minutes, which would otherwise take years in a standard HDL methodology.
Marcelo Guedes, Rafael Auler, Edson Borin, Rodolfo Azevedo
RSP3
2012 ACCGen: An Automatic ArchC Compiler Generator
abstract
The current level of circuit integration led to complex designs encompassing full systems on a single chip, known as System-on-a-Chip (SoC). In order to predict the best design options and reduce the design costs, designers are required to perform a large design space exploration on early stages of the design. To speed up this process, Electronic Design Automation (EDA) tools are employed to model and experiment with the system. ArchC is an "Architecture Description Language" (ADL) and a set of tools that can be leveraged to automatically build SoC simulators based on high-level system models, enabling easy and fast design space exploration in early stages of the design. Currently, ArchC is capable of automatically generating hardware simulators, assemblers, and linkers for a given architecture model. In this work, we present ACCGen, an automatic Compiler Generator for ArchC, the missing link on the automatic generation of compiler tool chains for ArchC. Our experimental results show that compilers generated by ACCGen are correct for Mibench applications. They compare, as well, the generated code quality with LLVM and gcc, two well-known open-source compilers. We also show that ACCGen is fast and has little impact on the design space exploration turnaround time, allowing the designer to, using an easy and fully automated workflow, completely assess the outcome of architectural changes in less than 2 minutes.
Rafael Auler, Paulo Centoducatte, Edson Borin
SBAC-PAD3
2011 LAR-CC: Large atomic regions with conditional commits
abstract
HW/SW Co-designed systems rely on dynamic binary translation and optimizations for efficient execution of binary code. Due to memory ordering properties and other architectural constraints, most binary optimizations are applied to regions of code that are atomically executed. To ensure that the underlying hardware has enough speculative resources to execute the whole atomic region, these systems typically form short atomic regions, with only 20 to 30 instructions. However, the shorter is the atomic region the smaller is the scope for optimizations. We present LAR-CC, a novel technique that enables HW/SW co-designed systems to optimize large atomic regions and dynamically fit them into the available speculative hardware resources by means of conditional commits. The LAR-CC technique consists of two major components: 1) conditional branch instructions to conditionally skip commit operations; 2) code transformations that replace commit operations by conditional commits and enable optimizations to be applied on the large atomic regions. Our experiments show that LAR-CC can effectively achieve dynamic atomic region sizes larger than 1000 instructions, providing sufficiently large scope to apply many advanced optimizations on HW/SW co-designed systems.
Edson Borin, Youfeng Wu, Maurício Breternitz, Cheng Wang 0013
CGO1
2011 A HW/SW co-designed heterogeneous multi-core virtual machine for energy-efficient general purpose computing
abstract
It is increasingly challenging to improve single thread performance because power/energy consumption becomes a major barrier to achieve significantly higher performance for general purpose cores. General purpose processors are designed to perform well in a wide variety of market segments, at the cost of having significantly lower performance-per-watt than special purpose processors targeting limited applications or market segments. In this paper, we propose a HW/SW co-designed heterogeneous multi-core virtual machine, called TwinPeaks, which integrates a set of less general but power efficient cores and uses dynamic binary optimization to schedule code regions to run on the most efficient cores. Our experiment and analysis indicate that TwinPeaks with a wide in-order core and a narrow out-of-order core may achieve 108% performance at ~71% energy of a big 4-wide out-of-order core.
Youfeng Wu, Shiliang Hu, Edson Borin, Cheng Wang 0013
CGO3
2011 Structure-Constrained Microcode Compression
abstract
Microcode enables programmability of (micro) architectural structures to enhance functionality and to apply patches to an existing design. As more features get added to a CPU core, the area and power costs associated with microcode increase. One solution to address the microcode size issue is to store the microcode in a compressed form and decompress it during execution. Furthermore, the reuse of a single hardware building block layout to implement different dictionaries in the two-level microcode compression reduces the cost and the design time of the decompression engine. However, the reuse of the hardware building block imposes structural constraints to the compression algorithm, and existing algorithms may yield poor compression. In this paper, we develop the SC2 algorithm that considers the structural constraint in its objective function and reduces the area expansion when reusing hardware building blocks to implement different dictionaries. Our experimental results show that the SC2 algorithm is able to produce similar sized dictionaries and achieves the similar compression ratio to the non-constrained algorithm.
Edson Borin, Guido Araujo, Maurício Breternitz, Youfeng Wu
SBAC-PAD1
2010 TAO: two-level atomicity for dynamic binary optimizations
abstract
Dynamic binary translation is a key component of Hardware/Software (HW/SW) co-design, which is an enabling technology for processor microarchitecture innovation. There are two well-known dynamic binary optimization techniques based on atomic execution support. Frame-based optimizations leverage processor pipeline support to enable atomic execution of hot traces. Region level optimizations employ transactional-memory-like atomicity support to aggressively optimize large regions of code. In this paper we propose a two-level atomic optimization scheme which not only overcomes the limitations of the two approaches, but also boosts the benefits of the two approaches effectively. Our experiment shows that the combined approach can achieve a total of 21.5% performance improvement over an aggressive out-of-order baseline machine and improve the performance over the frame-based approach by an additional 5.3%.
Edson Borin, Youfeng Wu, Cheng Wang 0013, Wei Liu 0014, Maurício Breternitz, Shiliang Hu, Esfir Natanzon, Shai Rotem, Roni Rosner
CGO1
2009 Dynamic parallelization of single-threaded binary programs using speculative slicing
abstract
The performance of single-threaded programs and legacy binary code is of critical importance in many everyday applications. However, neither can hardware multi-core processors directly speed up single-threaded programs, nor can software automatic parallelizing compilers effectively parallelize legacy binary code and irregular applications. In this paper, we propose a framework and a set of algorithms to dynamically parallelize single-threaded binary programs. Our parallelization is based on program slicing and explores both instruction-level parallelism (ILP) and thread-level parallelism (TLP). To significantly reduce the critical path of the parallel slices, our slicing algorithms exploit speculation to cut rare dependences, and use well-designed program transformations to expose parallelism. Furthermore, because we transparently parallelize binary code at runtime, we perform slicing only on program hot regions. Our experiments demonstrate that the proposed speculative slicing approach extracts more parallelism than any known slicing based parallelization schemes. For the SPEC2000 benchmarks, we can achieve 3x parallelism with infinite number of threads, and 1.8x parallelism with 4 threads.
Cheng Wang 0013, Youfeng Wu, Edson Borin, Shiliang Hu, Wei Liu 0014, Dave Sager, Tin-Fook Ngai, Jesse Fang
ICS3
2006 Software-Based Transparent and Comprehensive Control-Flow Error Detection
abstract
Shrinking microprocessor feature size and growing transistor density may increase the soft-error rates to unacceptable levels in the near future. While reliable systems typically employ hardware techniques to address soft-errors, software-based techniques can provide a less expensive and more flexible alternative. This paper presents a control-flow error classification and proposes two new software-based comprehensive control-flow error detection techniques. The new techniques are better than the previous ones in the sense that they detect errors in all the branch-error categories. We implemented the techniques in our dynamic binary translator so that the techniques can be applied to existing x86 binaries transparently. We compared our new techniques with the previous ones and we show that our methods cover more errors while has similar performance overhead.
Edson Borin, Cheng Wang 0013, Youfeng Wu, Guido Araujo
CGO1
2006 Clustering-Based Microcode Compression
abstract
Microcode enables programmability of (micro) architectural structures to enhance functionality and to apply patches to an existing design. As more features get added to a CPU core, the area and power costs associated with microcode increase. A recent Intel internal design targeted at low power and small footprint has estimated the costs of the microcode ROM to approach 20% of the total die area (and associated power consumption). Therefore, it is desirable to apply compression techniques to microcode. Microcode poses unique challenges for compression due to the long instruction format, the hand-coded nature of the programs and the stringent performance requirements that require fast decompression. This paper describes techniques for microcode compression that achieve .significant area and power savings, while presenting a streamlined architecture that enables high throughput within the constraints of a high performance CPU. The paper presents results for microcode compression on several commercial CPU designs which demonstrates compression ratios ranging from 50% to 62%.
Edson Borin, Maurício Breternitz, Youfeng Wu, Guido Araujo
ICCD1
2005 Efficient datapath merging for partially reconfigurable architectures
abstract
Reconfigurable systems have been shown to achieve significant performance speedup through architectures that map the most time-consuming application kernel modules or inner loops to a reconfigurable datapath. As each portion of the application starts to execute, the system partially reconfigures the datapath so as to perform the corresponding computation. The reconfigurable datapath should have as few and simple hardware blocks and interconnections as possible, in order to reduce its cost, area, and reconfiguration overhead. To achieve that, hardware blocks and interconnections should be reused as much as possible across the application. We represent each piece of the application as a data-flow graph (DFG). The DFG merging process identifies similarities among the DFGs, and produces a single datapath that can be dynamically reconfigured and has a minimum area cost, when considering both hardware blocks and interconnections. In this paper we present a novel technique for the DFG merge problem, and we evaluate it using programs from the MediaBench benchmark. Our algorithm execution time approaches the fastest previous solution to this problem and produces datapaths with an average area reduction of 20%. When compared to the best known area solution, our approach produces datapaths with area costs equivalent to (and in many cases better than) it, while achieving impressive speedups.
Nahri Moreano, Edson Borin, Cid C. de Souza, Guido Araujo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2