VLDB 2026 Research / reviewers in the wild / expert
Maurício Breternitz
dblp:b/MauricioBreternitz · also Maurício Breternitz Jr.
· DBLP profile ↗
41ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0003-1752-6255ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 8 since 2021Software engineering, systems software and programming languages · 4Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Weightless Neural Networks on Flexible Substrates: A Novel Approach to Wearable Machine LearningabstractIn this article, we present a novel approach that seamlessly integrates machine learning (ML) algorithms into wearable technology through the use of weightless neural networks (WNNs) and flexible integrated circuits (FlexICs). Our methodology employs combinational intelligent networks (COIN) for edge inference on resource-constrained devices, highlighting the advantages of WNNs in terms of power efficiency and minimal hardware requirements. We propose an automated design flow for implementing COIN as FlexICs aimed at developing scalable, cost-effective, and environmentally sustainable wearable monitoring solutions. As a proof-of-concept demonstrator, an arrhythmia detection FlexIC was fabricated using COIN to meet the stringent requirements of medium-complexity wearable applications, offering a promising path toward personalized and accessible healthcare solutions. Igor D. S. Miranda, Velu Pillai, Tejas Musale, Mugdha P. Jadhao, Paulo C. R. Souza Neto, Zachary Susskind, Alan T. L. Bacellar, Mael Lhostis, Priscila M. V. Lima, Diego Leonel Cadette Dutra, Eugene John, Maurício Breternitz, Felipe M. G. França, Emre Ozer 0001, Lizy Kurian John |
IEEE Trans. Very Large Scale Integr. Syst. | 12 |
| 2024 | Differentiable Weightless Neural NetworksabstractWe introduce the Differentiable Weightless Neural Network (DWN), a model based on interconnected lookup tables. Training of DWNs is enabled by a novel Extended Finite Difference technique for approximate differentiation of binary values. We propose Learnable Mapping, Learnable Reduction, and Spectral Regularization to further improve the accuracy and efficiency of these models. We evaluate DWNs in three edge computing contexts: (1) an FPGA-based hardware accelerator, where they demonstrate superior latency, throughput, energy efficiency, and model area compared to state-of-the-art solutions, (2) a low-power microcontroller, where they achieve preferable accuracy to XGBoost while subject to stringent memory constraints, and (3) ultra-low-cost chips, where they consistently outperform small models in both accuracy and projected hardware area. DWNs also compare favorably against leading approaches for tabular datasets, with higher average rank. Overall, our work positions DWNs as a pioneering solution for edge-compatible high-throughput neural networks. Alan T. L. Bacellar, Zachary Susskind, Maurício Breternitz, Eugene John, Lizy Kurian John, Priscila M. V. Lima, Felipe M. G. França |
ICML | 3 |
| 2024 | Soon Filter: Advancing Tiny Neural Architectures for High Throughput Edge InferenceabstractAs Deep Neural Networks become more complex and computationally demanding, efficient models for inference at the edge, particularly multiplication-free ones, have gained significant attention. The Ultra Low-Energy Edge Neural Network (ULEEN) is a notable architecture optimized for high throughput edge designs. ULEEN uniquely employs Bloom Filters with binary values to compute neuron activation, boasting better efficiency metrics than Binary Neural Networks (BNNs). This work uncovers a gradient back-propagation bottleneck within ULEEN’s Bloom filters and introduces a simplified version of it as a solution: the "Soon Filter". Both theoretically and empirically, we demonstrate that our approach improves gradient back-propagation efficiency. Tests on MLPerf Tiny, MNIST and various UCI datasets reveal that our method surpasses ULEEN, BNN, and DeepShift. Notably, with MLPerf KWS (Key Word Spotting) dataset, we achieve 69.6% accuracy with only 101KiB, while ULEEN, BNN and DeepShift achieve only 67.4%, 55.9%, and 24.9% respectively. Remarkably, we also achieve 67.7% accuracy with only 50KiB, resulting in a 2x model size reduction compared to ULEEN while maintaining similar accuracy (+0.3%). This results underscores the promising potential of our solution for efficient inference at the edge in applications that rely on high throughput architectures. Alan T. L. Bacellar, Zachary Susskind, Maurício Breternitz, Lizy Kurian John, Felipe M. G. França, Priscila M. V. Lima |
IJCNN | 3 |
| 2024 | Memory-efficient DRASiW Models
Otávio O. Napoli, Ana de Almeida 0002, Edson Borin, Maurício Breternitz |
Neurocomputing | 4 |
| 2023 | COIN: Combinational Intelligent NetworksabstractWe introduce Combinational Intelligent Networks (COIN), a machine learning technique that targets edge inference using low-resourced FPGAs or ASICs. COIN is an improvement on LogicWiSARD, a recent weightless neural network that achieves low power, small area, and high throughput. We convert the LogicWiSARD model into a binary neural network, train it using backpropagation, and then convert it to a COIN model. As a result, COIN can achieve higher accuracy than LogicWiSARD or it can require significantly fewer hardware resources when comparing models with similar accuracies. In comparison to a BNN implementation, FINN, small and large COIN models are more energy efficient demonstrating up to 11.5x higher inferences/Joule at similar accuracy. Our tool executes the complete flow, from training to RTL. and is publicly available. Igor D. S. Miranda, Aman Arora 0001, Zachary Susskind, Josias S. A. Souza, Mugdha P. Jadhao, Luis A. Q. Villon, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Maurício Breternitz, Lizy Kurian John |
ASAP | 10 |
| 2023 | Efficient Knowledge Aggregation Methods for Weightless Neural NetworksabstractWeightless Neural Networks (WNN) are good candidates for Federated Learning scenarios due to their robustness and computational lightness.In this work, we show that it is possible to aggregate the knowledge of multiple WNNs using more compact data structures, such as Bloom Filters, to reduce the amount of data transferred between devices.Finally, we explore variations of Bloom Filters and found that a particular data-structure, the Count-Min Sketch (CMS), is a good candidate for aggregation.Costing at most 3% of accuracy, CMS can be up to 3x smaller when compared to previous approaches, specially for large datasets. Otávio O. Napoli, Ana de Almeida 0002, José Miguel Salles Dias, Luís Brás Rosário, Edson Borin, Maurício Breternitz |
ESANN | 6 |
| 2023 | An FPGA-Based Weightless Neural Network for Edge Network Intrusion DetectionabstractAlgorithms for mobile networking are increasingly being moved from centralized servers towards the edge in order to decrease latency and improve the user experience. While much of this work is traditionally done using ASICs, 6G emphasizes the adaptability of algorithms for specific user scenarios, which motivates broader adoption of FPGAs. In this paper, we propose the FPGA-based Weightless Intrusion Warden (FWIW), a novel solution for detecting anomalous network traffic on edge devices. While prior work in this domain is based on conventional deep neural networks (DNNs), FWIW incorporates a weightless neural network (WNN), a table lookup-based model which learns sophisticated nonlinear behaviors. This allows FWIW to achieve accuracy far superior to prior FPGA-based work at a very small fraction of the model footprint, enabling deployment on small, low-cost devices. FWIW achieves a prediction accuracy of 98.5% on the UNSW-NB15 dataset with a total model parameter size of just 192 bytes, reducing error by 7.9x and model size by 262x vs. LogicNets, the best prior edge-optimized implementation. Implemented on a Xilinx Virtex UltraScale+ FPGA, FWIW demonstrates a 59x reduction in LUT usage with a 1.6x increase in throughput. The accuracy of FWIW comes within 0.6% of the best-reported result in literature (Edge-Detect), a model several orders of magnitude larger. Our results make it clear that WNNs are worth exploring in the emerging domain of edge networking, and suggest that FPGAs are capable of providing the extreme throughput needed. Zachary Susskind, Aman Arora 0001, Alan T. L. Bacellar, Diego Leonel Cadette Dutra, Igor D. S. Miranda, Maurício Breternitz, Priscila M. V. Lima, Felipe M. G. França, Lizy Kurian John |
FPGA | 6 |
| 2023 | A conditional branch predictor based on weightless neural networks
Luis A. Q. Villon, Zachary Susskind, Alan T. L. Bacellar, Igor D. S. Miranda, Leandro Santiago de Araújo, Priscila M. V. Lima, Maurício Breternitz, Lizy Kurian John, Felipe M. G. França, Diego Leonel Cadette Dutra |
Neurocomputing | 7 |
| 2023 | ULEEN: A Novel Architecture for Ultra-low-energy Edge Neural Networksabstract‘‘Extreme edge” 1 devices, such as smart sensors, are a uniquely challenging environment for the deployment of machine learning. The tiny energy budgets of these devices lie beyond what is feasible for conventional deep neural networks, particularly in high-throughput scenarios, requiring us to rethink how we approach edge inference. In this work, we propose ULEEN, a model and FPGA-based accelerator architecture based on weightless neural networks (WNNs). WNNs eliminate energy-intensive arithmetic operations, instead using table lookups to perform computation, which makes them theoretically well-suited for edge inference. However, WNNs have historically suffered from poor accuracy and excessive memory usage. ULEEN incorporates algorithmic improvements and a novel training strategy inspired by binary neural networks (BNNs) to make significant strides in addressing these issues. We compare ULEEN against BNNs in software and hardware using the four MLPerf Tiny datasets and MNIST. Our FPGA implementations of ULEEN accomplish classification at 4.0–14.3 million inferences per second, improving area-normalized throughput by an average of 3.6× and steady-state energy efficiency by an average of 7.1× compared to the FPGA-based Xilinx FINN BNN inference platform. While ULEEN is not a universally applicable machine learning model, we demonstrate that it can be an excellent choice for certain applications in energy- and latency-critical edge environments. Zachary Susskind, Aman Arora 0001, Igor D. S. Miranda, Alan T. L. Bacellar, Luis A. Q. Villon, Rafael Fontella Katopodis, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Maurício Breternitz, Lizy Kurian John |
ACM Trans. Archit. Code Optim. | 11 |
| 2022 | Weightless Neural Networks for Efficient Edge InferenceabstractWeightless neural networks (WNNs) are a class of machine learning model which use table lookups to perform inference, rather than the multiply-accumulate operations typical of deep neural networks (DNNs). Individual weightless neurons are capable of learning non-linear functions of their inputs, a theoretical advantage over the linear neurons in DNNs, yet state-of-the-art WNN architectures still lag behind DNNs in accuracy on common classification tasks. Additionally, many existing WNN architectures suffer from high memory requirements, hindering implementation. In this paper, we propose a novel WNN architecture, BTHOWeN, with key algorithmic and architectural improvements over prior work, namely counting Bloom filters, hardware-friendly hashing, and Gaussian-based nonlinear thermometer encodings. These enhancements improve model accuracy while reducing size and energy per inference. BTHOWeN targets the large and growing edge computing sector by providing superior latency and energy efficiency to both prior WNNs and comparable quantized DNNs. Compared to state-of-the-art WNNs across nine classification datasets, BTHOWeN on average reduces error by more than 40% and model size by more than 50%. We demonstrate the viability of a hardware implementation of BTHOWeN by presenting an FPGA-based inference accelerator, and compare its latency and resource usage against similarly accurate quantized DNN inference accelerators, including multi-layer perceptron (MLP) and convolutional models. The proposed BTHOWeN models consume almost 80% less energy than the MLP models, with nearly 85% reduction in latency. In our quest for efficient ML on the edge, WNNs are clearly deserving of additional attention. Zachary Susskind, Aman Arora 0001, Igor D. S. Miranda, Luis A. Q. Villon, Rafael Fontella Katopodis, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Maurício Breternitz, Lizy Kurian John |
PACT | 10 |
| 2022 | LogicWiSARD: Memoryless Synthesis of Weightless Neural NetworksabstractWeightless neural networks (WNNs) are an alternative pattern recognition technique where RAM nodes function as neurons. As both training and inference require mostly table lookups, few additions, and no multiplications, WNNs are suitable for high-performance and low-power embedded applications. This work introduces a novel approach to implement WiSARD, the leading WNN state-of-the-art architecture, completely eliminating memories and arithmetic circuits and utilizing only logic functions. The approach creates compressed minimized implementations by converting trained WNN nodes from lookup tables to logic functions. The proposed LogicWiSARD is implemented in FPGA and ASIC technologies to illustrate its suitability for edge inference. Experimental results show more than 80% reduction in energy consumption when the proposed LogicWiSARD model is compared with a multilayer perceptron network (MLP) of equivalent accuracy. Compared to previous work on FPGA implementations for WNNs, convolutional neural networks, and binary neural networks, the energy savings of LogicWiSARD range between 32.2% and 99.6%. Igor D. S. Miranda, Aman Arora 0001, Zachary Susskind, Luis A. Q. Villon, Rafael Fontella Katopodis, Diego Leonel Cadette Dutra, Leandro Santiago de Araújo, Priscila M. V. Lima, Felipe M. G. França, Lizy Kurian John, Maurício Breternitz |
ASAP | 11 |
| 2022 | Distributive Thermometer: A New Unary Encoding for Weightless Neural NetworksabstractThe binary encoding of real valued inputs is a crucial part of Weightless Neural Networks.The Linear Thermometer and its variations are the most prominent methods to determine binary encoding for input data but, as they make assumptions about the input distribution, the resulting encoding is sub-optimal and possibly wasteful when the assumption is incorrect.We propose a new thermometer approach that doesn't require such assumptions.Our results show that it achieves similar or better accuracy when compared to a thermometer that correctly assumes the distribution, and accuracy gains up to 26.3% when other thermometer representations assume an unsound distribution. Alan T. L. Bacellar, Zachary Susskind, Luis A. Q. Villon, Igor D. S. Miranda, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Maurício Breternitz, Lizy Kurian John, Priscila M. V. Lima, Felipe M. G. França |
ESANN | 7 |
| 2022 | Pruning Weightless Neural NetworksabstractWeightless neural networks (WNNs) are a type of machine learning model which perform prediction using lookup tables (LUTs) instead of arithmetic operations.Recent advancements in WNNs have reduced model sizes and improved accuracies, reducing the gap in accuracy with deep neural networks (DNNs).Modern DNNs leverage "pruning" techniques to reduce model size, but this has not previously been explored for WNNs.We propose a WNN pruning strategy based on identifying and culling the LUTs which contribute least to overall model accuracy.We demonstrate an average 40% reduction in model size with at most 1% reduction in accuracy. Zachary Susskind, Alan T. L. Bacellar, Aman Arora 0001, Luis A. Q. Villon, Renan Mendanha, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Igor D. S. Miranda, Maurício Breternitz, Lizy Kurian John |
ESANN | 11 |
| 2022 | A WiSARD-based conditional branch predictorabstractConditional branch prediction is a technique used to speculatively execute instructions before knowing the direction of conditional branch statements. Perceptron-based predictors have been extensively studied, however, they need large input sizes for the data to be linearly separable. To learn nonlinear functions from the inputs, we propose a conditional branch predictor based on the WiSARD model and compare it with two state-of-the-art predictors, the TAGE-SC-L and the Multiperspective Perceptron. We show that the WiSARD-based predictor with a smaller input size outperforms the perceptron-based predictor by about 0.09% and achieves similar accuracy to that of TAGE-SC-L. Luis A. Q. Villon, Zachary Susskind, Alan T. L. Bacellar, Igor D. S. Miranda, Leandro Santiago de Araújo, Priscila M. V. Lima, Maurício Breternitz, Lizy Kurian John, Felipe M. G. França, Diego Leonel Cadette Dutra |
ESANN | 7 |
| 2021 | Smart selection of optimizations in dynamic compilersabstractSummary Dynamic compilers perform compilation and generation of target code during runtime, implying that the compilation time is added into the program runtime. Thus, to build a high‐performing dynamic compilation system, it is crucial to be able to generate high‐quality code and, at the same time, have a small compilation cost. In this article, we present an approach that uses machine learning to select sequences of optimization for dynamic compilation that considers both code quality and compilation overhead. Our approach starts by training a model, offline, with a knowledge bank of those sequences with low overhead and high‐quality code generation capability using a genetic heuristic. Then, this bank is used to guide the smart selection of optimizations sequences for the compilation of code fragments during the emulation of an application. We evaluate the proposed strategy in two LLVM‐based dynamic binary translators, namely OI‐DBT and HQEMU, and show that these two translators can achieve average speedups of 1.26x and 1.15x in MiBench and Spec Cpu benchmarks, respectively. Vanderson Martins do Rosário, Anderson Faustino da Silva, Thais Aparecida Silva Camacho, Otávio O. Napoli, Maurício Breternitz, Edson Borin |
Concurr. Comput. Pract. Exp. | 5 |
| 2021 | Efficiency and scalability of multi-lane capsule networks (MLCN)
Vanderson Martins do Rosário, Maurício Breternitz, Edson Borin |
J. Parallel Distributed Comput. | 2 |
| 2020 | A unified model for accelerating unsupervised iterative re-ranking algorithmsabstractSummary Despite the continuous advances in image retrieval technologies, performing effective and efficient content‐based searches remains a challenging task. Unsupervised iterative re‐ranking algorithms have emerged as a promising solution and have been widely used to improve the effectiveness of multimedia retrieval systems. Although substantially more efficient than related approaches based on diffusion processes, these re‐ranking algorithms can still be computationally costly, demanding the specification and implementation of efficient big multimedia analysis approaches. Such demand associated with the significant potential for parallelization and highly effective results achieved by recently proposed re‐ranking algorithms creates the need for exploiting efficiency vs effectiveness trade‐offs. In this article, we introduce a class of unsupervised iterative re‐ranking algorithms and present a model that can be used to guide their implementation and optimization for parallel architectures. We also analyze the impact of the parallelization on the performance of four algorithms that belong to the proposed class: Contextual Spaces, RL‐Sim, Contextual Re‐ranking, and Cartesian Product of Ranking References. The experiments show speedups that reach up to 6.0×, 16.1×, 3.3×, and 7.1× for each algorithm, respectively. These results demonstrate that the proposed parallel programming model can be successfully applied to various algorithms and used to improve the performance of multimedia retrieval systems. Flávia Pisani, Lucas Pascotti Valem, Daniel C. G. Pedronette, Ricardo da Silva Torres, Edson Borin, Maurício Breternitz |
Concurr. Comput. Pract. Exp. | 6 |
| 2020 | Weightless Neural Networks as Memory Segmented Bloom Filters
Leandro Santiago de Araújo, Letícia Dias Verona, Fábio Medeiros Rangel, Fabrício Firmino de Faria, Daniel Sadoc Menasché, Wouter Caarls, Maurício Breternitz, Sandip Kundu, Priscila M. V. Lima, Felipe M. G. França |
Neurocomputing | 7 |
| 2019 | Memory Efficient Weightless Neural Network using Bloom Filter
Leandro Santiago de Araújo, Letícia Dias Verona, Fábio Medeiros Rangel, Fabrício Firmino de Faria, Daniel Sadoc Menasché, Wouter Caarls, Maurício Breternitz, Sandip Kundu, Priscila M. V. Lima, Felipe M. G. França |
ESANN | 7 |
| 2019 | Efficiency and Scalability of Multi-lane Capsule Networks (MLCN)
Vanderson Martins do Rosário, Maurício Breternitz, Edson Borin |
SBAC-PAD | 2 |
| 2019 | The Multi-Lane Capsule NetworkabstractWe introduce multi-lane capsule networks (MLCN), which are a separable and resource efficient organization of capsule networks (CapsNet) that allows parallel processing while achieving high accuracy at reduced cost. A MLCN is composed of a number of (distinct) parallel lanes, each contributing to a dimension of the result, trained using the routing-by-agreement organization of CapsNet. Our results indicate similar accuracy with a much-reduced cost in number of parameters for the Fashion-MNIST and Cifar10 datasets. They also indicate that the MLCN outperforms the original CapsNet when using a proposed novel configuration for the lanes. MLCN also has faster training and inference times, being more than two-fold faster than the original CapsNet in a same accelerator. Vanderson Martins do Rosário, Edson Borin, Maurício Breternitz |
IEEE Signal Process. Lett. | 3 |
| 2018 | ComP-net: command processor networking for efficient intra-kernel communications on GPUsabstractCurrent state-of-the-art in GPU networking advocates a host-centric model that reduces performance and increases code complexity. Recently, researchers have explored several techniques for networking within a GPU kernel itself. These approaches, however, suffer from high latency, waste energy on the host, and are not scalable with larger/more GPUs on a node. In this work, we introduce Command Processor Networking (ComP-Net), which leverages the availability of scalar cores integrated on the GPU itself to provide high-performance intra-kernel networking. ComP-Net enables efficient synchronization between the Command Processors and Compute Units on the GPU through a line locking scheme implemented in the GPU's shared last-level cache. We illustrate that ComP-Net can improve application performance by up to 20% and provide up to 50% reduction in energy consumption vs. competing networking techniques across a Jacobi stencil, allreduce collective, and machine learning applications. Michael LeBeane, Khaled Hamidouche, Brad Benton, Maurício Breternitz, Steven K. Reinhardt, Lizy Kurian John |
PACT | 4 |
| 2017 | GPU triggered networking for intra-kernel communicationsabstractGPUs are widespread across clusters of compute nodes due to their attractive performance for data parallel codes. However, communicating between GPUs across the cluster is cumbersome when compared to CPU networking implementations. A number of recent works have enabled GPUs to more naturally access the network, but suffer from performance problems, require hidden CPU helper threads, or restrict communications to kernel boundaries. Michael LeBeane, Khaled Hamidouche, Brad Benton, Maurício Breternitz, Steven K. Reinhardt, Lizy Kurian John |
SC | 4 |
| 2016 | Extended task queuing: active messages for heterogeneous systemsabstractAccelerators have emerged as an important component of modern cloud, datacenter, and HPC computing environments. However, launching tasks on remote accelerators across a network remains unwieldy, forcing programmers to send data in large chunks to amortize the transfer and launch overhead. By combining advances in intra-node accelerator unification with one-sided Remote Direct Memory Access (RDMA) communication primitives, it is possible to efficiently implement lightweight tasking across distributed-memory systems. This paper introduces Extended Task Queuing (XTQ), an RDMA-based active messaging mechanism for accelerators in distributed-memory systems. XTQ's direct NIC-to-accelerator communication decreases inter-node GPU task launch latency by 10-15% for small-to-medium sized messages and ameliorates CPU message servicing overheads. These benefits are shown in the context of MPI accumulate, reduce, and allreduce operations with up to 64 nodes. Finally, we illustrate how XTQ can improve the performance of popular deep learning workloads implemented in the Computational Network Toolkit (CNTK). Michael LeBeane, Brandon Potter, Abhisek Pan, Alexandru Dutu, Vinay Agarwala, Wonchan Lee, Deepak Majeti, Bibek Ghimire, Eric Van Tassell, Samuel Wasmundt, Brad Benton, Maurício Breternitz, Michael L. Chu, Mithuna Thottethodi, Lizy Kurian John, Steven K. Reinhardt |
SC | 12 |
| 2016 | HadoopCL2: Motivating the Design of a Distributed, Heterogeneous Programming System With Machine-Learning ApplicationsabstractMachine learning (ML) algorithms have garnered increased interest as they demonstrate improved ability to extract meaningful trends from large, diverse, and noisy data sets. While research is advancing the state-of-the-art in ML algorithms, it is difficult to drastically improve the real-world performance of these algorithms. Porting new and existing algorithms from single-node systems to multi-node clusters, or from architecturally homogeneous systems to heterogeneous systems, is a promising optimization technique. However, performing optimized ports is challenging for domain experts who may lack experience in distributed and heterogeneous software development. This work explores how challenges in ML application development on heterogeneous, distributed systems shaped the development of the HadoopCL2 (HCL2) programming system. ML applications guide this work because they exhibit features that make application development difficult: large & diverse datasets, complex algorithms, and the need for domain-specific knowledge. The goal of this work is a general, MapReduce programming system that outperforms existing programming systems. This work evaluates the performance and portability of HCL2 against five ML applications from the Mahout ML framework on two hardware platforms. HCL2 demonstrates speedups of greater than 20x relative to Mahout for three computationally heavy algorithms and maintains minor performance improvements for two I/O bound algorithms. Max Grossman, Maurício Breternitz, Vivek Sarkar |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | Adaptive global power optimization for Web servers
Leonardo Piga, Reinaldo A. Bergamaschi, Maurício Breternitz, Sandro Rigo |
J. Supercomput. | 3 |
| 2013 | Image Re-ranking Acceleration on GPUsabstractHuge image collections are becoming available lately. In this scenario, the use of Content-Based Image Retrieval (CBIR) systems has emerged as a promising approach to support image searches. The objective of CBIR systems is to retrieve the most similar images in a collection, given a query image, by taking into account image visual properties such as texture, color, and shape. In these systems, the effectiveness of the retrieval process depends heavily on the accuracy of ranking approaches. Recently, re-ranking approaches have been proposed to improve the effectiveness of CBIR systems by taking into account the relationships among images. The re-ranking approaches consider the relationships among all images in a given dataset. These approaches typically demands a huge amount of computational power, which hampers its use in practical situations. On the other hand, these methods can be massively parallelized. In this paper, we propose to speedup the computation of the RL-Sim algorithm, a recently proposed image re-ranking approach, by using the computational power of Graphics Processing Units (GPU). GPUs are emerging as relatively inexpensive parallel processors that are becoming available on a wide range of computer systems. We address the image re-ranking performance challenges by proposing a parallel solution designed to fit the computational model of GPUs. We conducted an experimental evaluation considering different implementations and devices. Experimental results demonstrate that significant performance gains can be obtained. Our approach achieves speedups of 7x from serial implementation considering the overall algorithm and up to 36x on its core steps. Daniel C. G. Pedronette, Ricardo da Silva Torres, Edson Borin, Maurício Breternitz |
SBAC-PAD | 4 |
| 2012 | Efficient Image Re-Ranking Computation on GPUsabstractThe huge growth of image collections and multimedia resources available is remarkable. One of the most common approaches to support image searches relies on the use of Content-Based Image Retrieval (CBIR) systems. CBIR systems aim at retrieving the most similar images in a collection, given a query image. Since the effectiveness of those systems is very dependent on the accuracy of ranking approaches, re-ranking algorithms have been proposed to exploit contextual information and improve the effectiveness of CBIR systems. Image re-ranking algorithms typically consider the relationship among every image in a given dataset when computing the new ranking. This approach demands a huge amount of computational power, which may render it prohibitive on very large data sets. In order to mitigate this problem, we propose using the computational power of Graphics Processing Units (GPU) to speedup the computation of image re-ranking algorithms. GPUs are fast emerging and relatively inexpensive parallel processors that are becoming available on a wide range of computer systems. In this paper, we propose a parallel implementation of an image re-ranking algorithm designed to fit the computational model of GPUs. Experimental results demonstrate that relevant performance gains can be obtained by our approach. Daniel C. G. Pedronette, Ricardo da Silva Torres, Edson Borin, Maurício Breternitz |
ISPA | 4 |
| 2012 | Cloud Workload Analysis with SWATabstractThis note describes the Synthetic Workload Application Toolkit (SWAT) and presents the results from a set of experiments on some key cloud workloads. SWAT is a software platform that automates the creation, deployment, provisioning, execution, and (most importantly) data gathering of synthetic compute workloads on clusters of arbitrary size. SWAT collects and aggregates data from application execution logs, operating system call interfaces, and micro architecture-specific program counters. The data collected by SWAT are used to characterize the effects of network traffic, file I/O, and computation on program performance. The output is analyzed to provide insight into the design and deployment of cloud workloads and systems. Each workload is characterized according to its scalability with the number of server nodes and Hadoop server jobs, sensitivity to network characteristics (bandwidth, latency, statistics on packet size), and computation vs. I/O intensity as these values adjusted via workload-specific parameters. (In the future, we will use SWAT's benchmark synthesizer capability.) We also characterize micro-architectural characteristics that give insight on the micro architecture of processors better suited for this class of workloads. We contrast our results with prior work on Cloud Suite [5], validating some conclusions and providing further insight into others. This illustrates SWAT's data collection capabilities and usefulness to obtain insight on cloud applications and systems. Maurício Breternitz, Keith Lowery, Anton Charnoff, Patryk Kaminski, Leonardo Piga |
SBAC-PAD | 1 |
| 2011 | LAR-CC: Large atomic regions with conditional commitsabstractHW/SW Co-designed systems rely on dynamic binary translation and optimizations for efficient execution of binary code. Due to memory ordering properties and other architectural constraints, most binary optimizations are applied to regions of code that are atomically executed. To ensure that the underlying hardware has enough speculative resources to execute the whole atomic region, these systems typically form short atomic regions, with only 20 to 30 instructions. However, the shorter is the atomic region the smaller is the scope for optimizations. We present LAR-CC, a novel technique that enables HW/SW co-designed systems to optimize large atomic regions and dynamically fit them into the available speculative hardware resources by means of conditional commits. The LAR-CC technique consists of two major components: 1) conditional branch instructions to conditionally skip commit operations; 2) code transformations that replace commit operations by conditional commits and enable optimizations to be applied on the large atomic regions. Our experiments show that LAR-CC can effectively achieve dynamic atomic region sizes larger than 1000 instructions, providing sufficiently large scope to apply many advanced optimizations on HW/SW co-designed systems. Edson Borin, Youfeng Wu, Maurício Breternitz, Cheng Wang 0013 |
CGO | 3 |
| 2011 | Structure-Constrained Microcode CompressionabstractMicrocode enables programmability of (micro) architectural structures to enhance functionality and to apply patches to an existing design. As more features get added to a CPU core, the area and power costs associated with microcode increase. One solution to address the microcode size issue is to store the microcode in a compressed form and decompress it during execution. Furthermore, the reuse of a single hardware building block layout to implement different dictionaries in the two-level microcode compression reduces the cost and the design time of the decompression engine. However, the reuse of the hardware building block imposes structural constraints to the compression algorithm, and existing algorithms may yield poor compression. In this paper, we develop the SC2 algorithm that considers the structural constraint in its objective function and reduces the area expansion when reusing hardware building blocks to implement different dictionaries. Our experimental results show that the SC2 algorithm is able to produce similar sized dictionaries and achieves the similar compression ratio to the non-constrained algorithm. Edson Borin, Guido Araujo, Maurício Breternitz, Youfeng Wu |
SBAC-PAD | 3 |
| 2010 | TAO: two-level atomicity for dynamic binary optimizationsabstractDynamic binary translation is a key component of Hardware/Software (HW/SW) co-design, which is an enabling technology for processor microarchitecture innovation. There are two well-known dynamic binary optimization techniques based on atomic execution support. Frame-based optimizations leverage processor pipeline support to enable atomic execution of hot traces. Region level optimizations employ transactional-memory-like atomicity support to aggressively optimize large regions of code. In this paper we propose a two-level atomic optimization scheme which not only overcomes the limitations of the two approaches, but also boosts the benefits of the two approaches effectively. Our experiment shows that the combined approach can achieve a total of 21.5% performance improvement over an aggressive out-of-order baseline machine and improve the performance over the frame-based approach by an additional 5.3%. Edson Borin, Youfeng Wu, Cheng Wang 0013, Wei Liu 0014, Maurício Breternitz, Shiliang Hu, Esfir Natanzon, Shai Rotem, Roni Rosner |
CGO | 5 |
| 2008 | A Segmented Bloom Filter Algorithm for Efficient PredictorsabstractBloom Filters are a technique to reduce the effects of conflicts/interference in hash table-like structures. Conventional hash tables store information in a single location which is susceptible to destructive interference through hash conflicts. A Bloom Filter uses multiple hash functions to store information in several locations, and recombines the information through some voting mechanism. Many microarchitectural predictors use simple single-index hash tables to make binary 0/1 predictions, and Bloom Filters help improve predictor accuracy. However, implementing a true Bloom Filter requires k hash functions, which in turn implies a k-ported hash table, or k sequential accesses. Unfortunately,the area of a hardware table increases quadratically with the port count, increasing costs of area, latency and power consumption. We propose a simple but elegant modification to the Bloom Filter algorithm that uses banking combined with special hash functions that guarantee all hash indexes fall into non-conflicting banks. We evaluate several applications of our Banked Bloom Filter (BBF) prediction in processors: BBF branch prediction, BBF load hit/miss prediction, and BBF last-tag prediction. We show that BBF predictors can provide accurate predictions with substantially less cost than previous techniques. Maurício Breternitz, Gabriel H. Loh, Bryan Black, Jeff Rupley, Peter G. Sassone, Wesley Attrot, Youfeng Wu |
SBAC-PAD | 1 |
| 2007 | Impacts of Multiprocessor Configurations on Workloads in BioinformaticsabstractBioinformatics is among the most active research areas in computer science. In this study, we investigate a suite of workloads in bioinformatics on two multiprocessor systems with different configurations, and examine the effects of the configurations on the performance of the workloads. Our result indicates that the configurations of the multiprocessor systems have significant impact on the performance and scalability of the workloads. For example, a number of workloads have significantly higher scalability on one of the systems, but poorer absolute performance than on the other system. However, traditional scalability failed to capture the impacts of the system configurations on the workloads. We present insights on what kinds of workloads will run faster on which systems and propose new metrics to capture the impacts of multiple processor configurations on the workloads. These findings not only provide an easy way to compare results running on different systems, but also enable re-configuration of the underlying systems to run specific workloads efficiently. We also show how processor mapping and loop spreading may help map the workoads to the underlining multiprocessor configuration and achieve consistent scalability for these workloads. Youfeng Wu, Maurício Breternitz, Victor Ying |
SBAC-PAD | 2 |
| 2006 | Clustering-Based Microcode CompressionabstractMicrocode enables programmability of (micro) architectural structures to enhance functionality and to apply patches to an existing design. As more features get added to a CPU core, the area and power costs associated with microcode increase. A recent Intel internal design targeted at low power and small footprint has estimated the costs of the microcode ROM to approach 20% of the total die area (and associated power consumption). Therefore, it is desirable to apply compression techniques to microcode. Microcode poses unique challenges for compression due to the long instruction format, the hand-coded nature of the programs and the stringent performance requirements that require fast decompression. This paper describes techniques for microcode compression that achieve .significant area and power savings, while presenting a streamlined architecture that enables high throughput within the constraints of a high performance CPU. The paper presents results for microcode compression on several commercial CPU designs which demonstrates compression ratios ranging from 50% to 62%. Edson Borin, Maurício Breternitz, Youfeng Wu, Guido Araujo |
ICCD | 2 |
| 2004 | The Accuracy of Initial Prediction in Two-Phase Dynamic Binary TranslatorsabstractDynamic binary translators use a two-phase approach to identify and optimize frequently executed code dynamically. In the first step (profiling phase), blocks of code are interpreted or quickly translated to collect execution frequency information for the blocks. In the second phase (optimization phase), frequently executed blocks are grouped into regions and advanced optimizations are applied on them. This approach implicitly assumes that the initial profile of each block is representative of the block throughout its lifetime. We investigate the ability of the initial profile to predict the average program behavior. We compare the predicted behavior of varying lengths of the initial execution with the average program behavior for the whole program execution, and use the prediction from the training input as the reference. Our result indicates that, for the SPEC2000 benchmarks, even very short initial profiles have comparable prediction accuracy to the traditional profile-guided optimizations using the training input, although the initial profile is inadequate for predicting loop trip count information for some integer programs and several benchmarks can benefit from phase-awareness during dynamic binary translation. Youfeng Wu, Maurício Breternitz, Justin Quek, Orna Etzion, Jesse Fang |
CGO | 2 |
| 1997 | Enhanced Compression Techniques to Simplify Programm Decompression and ExecutionabstractCompressing instruction sequences can reduce the cost of embedded systems by reducing program ROM-size requirements. Compression also facilitates the use of RISC core architectures, like the PowerPC/sup TM/ architecture, in embedded systems. Compression techniques are presented which enable decompression and execution of compressed code to occur without the need of a lookaside table (LAT) or cache lookaside buffer (CLB). These techniques successfully merge code modification and compression into a single software preprocessing step. Decompression and execution of compressed code are made very simple. An application of these techniques to about 120000 instructions of PowerPC firmware code is described. Maurício Breternitz |
ICCD | 1 |
| 1996 | Design Tradeoffs and Experience with Motorola PowerPC? Migration ToolabstractThe Motorola PowerPC migration tools enable the conversion of assembly programs from other architectures to PowerPC. This paper describes the design approach and experience with the tool to translate x86 assembly programs to PowerPC. The key problems of handling 16-bit code, the effects of masking 16-bit operations into 32-bit registers and optimization of condition flags are discussed. The efficiency of translation and the effects of architectural constraints on design tradeoffs are analyzed. Maurício Breternitz, A. Manikonda, M. Ommerman, W. Su, A. Thornto |
ICCD | 1 |
| 1991 | Implementation Optimization Techniques for Architecture Synthesis of Application-Specific ProcessorsabstractArticle Free Access Share on Implementation optimization techniques for architecture synthesis of application-specific processors Authors: Mauricio Breternitz Advanced Workstations Division, IBM-Austin, TX Advanced Workstations Division, IBM-Austin, TXView Profile , John Paul Shen ECE Department, Carnegie-Mellon University ECE Department, Carnegie-Mellon UniversityView Profile Authors Info & Claims MICRO 24: Proceedings of the 24th annual international symposium on MicroarchitectureSeptember 1991 Pages 114–123https://doi.org/10.1145/123465.123488Published:01 September 1991Publication History 0citation252DownloadsMetricsTotal Citations0Total Downloads252Last 12 Months11Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Maurício Breternitz, John Paul Shen |
MICRO | 1 |
| 1990 | Architecture Synthesis of High-Performance Application-Specific ProcessorsabstractThe key principles of the Application-Specific Processor Design (ASPD) methodology include: a semi-custom compilation-driven design/implementation approach, the exploitation of fine-grained parallelism for high performance, and the adaptation of datapath topology to the data transfers required by the application. The powerful microcode compilation techniques of Percolation Scheduling and Pipeline Scheduling extract and enhance the parallelism in the application object code to generate an optimized specification of the target processor. Implementation optimization is performed to allocate functional units and register files. Graph-coloring algorithms minimize the amount of hardware needed to exploit available parallelism. Data memory employs an organization with multiple banks. Compilation techniques are used to allocate data over the memory banks to enhance parallel access. Maurício Breternitz, John Paul Shen |
DAC | 1 |
| 1988 | The White Dwarf: A High-Performance Application-Specific ProcessorabstractThe design and implementation of a high-performance special-purpose processor, called the White Dwarf, or accelerating finite-element analysis algorithms is presented. The White Dwarf CPU contains two Am2935 32-bit floating-point processors and one Am29332 32-bit arithmetic logic unit (ALU), and uses a wide-instruction-word architecture in which the application algorithm is directly implemented in microcode. The entire system is VME-bus compatible and interfaces with a Sun 3/160 host. The system's potential peak performance is 20 MFLOPS (million floating-point operations per second) a sustained computation rate in excess of 15 MFLOPS is expected. A potential speedup of between one and two orders of magnitude is possible. With a fully populated memory subsystem, the White Dwarf can accommodate finite-element problems involving up to half a million nodes. The system is designed using an approach called application-specific processor design (ASPD). A retargetable compiler has been developed which is capable of generating highly parallel and efficient code for the White Dwarf and other processors with similar architecture. System debug/integration is in progress; a highly useful system is expected.> Andrew Wolfe, Maurício Breternitz, Chriss Stephens, A. L. Ting, David Blair Kirk, Ronald P. Bianchini Jr., John Paul Shen |
ISCA | 2 |