Maurício Breternitz

dblp:b/MauricioBreternitz · also Maurício Breternitz Jr. · DBLP profile ↗
← Back
41ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0003-1752-6255ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 8 since 2021Software engineering, systems software and programming languages · 4Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Weightless Neural Networks on Flexible Substrates: A Novel Approach to Wearable Machine Learning
abstract
In this article, we present a novel approach that seamlessly integrates machine learning (ML) algorithms into wearable technology through the use of weightless neural networks (WNNs) and flexible integrated circuits (FlexICs). Our methodology employs combinational intelligent networks (COIN) for edge inference on resource-constrained devices, highlighting the advantages of WNNs in terms of power efficiency and minimal hardware requirements. We propose an automated design flow for implementing COIN as FlexICs aimed at developing scalable, cost-effective, and environmentally sustainable wearable monitoring solutions. As a proof-of-concept demonstrator, an arrhythmia detection FlexIC was fabricated using COIN to meet the stringent requirements of medium-complexity wearable applications, offering a promising path toward personalized and accessible healthcare solutions.
Igor D. S. Miranda, Velu Pillai, Tejas Musale, Mugdha P. Jadhao, Paulo C. R. Souza Neto, Zachary Susskind, Alan T. L. Bacellar, Mael Lhostis, Priscila M. V. Lima, Diego Leonel Cadette Dutra, Eugene John, Maurício Breternitz, Felipe M. G. França, Emre Ozer 0001, Lizy Kurian John
IEEE Trans. Very Large Scale Integr. Syst.12
2024 Differentiable Weightless Neural Networks
abstract
We introduce the Differentiable Weightless Neural Network (DWN), a model based on interconnected lookup tables. Training of DWNs is enabled by a novel Extended Finite Difference technique for approximate differentiation of binary values. We propose Learnable Mapping, Learnable Reduction, and Spectral Regularization to further improve the accuracy and efficiency of these models. We evaluate DWNs in three edge computing contexts: (1) an FPGA-based hardware accelerator, where they demonstrate superior latency, throughput, energy efficiency, and model area compared to state-of-the-art solutions, (2) a low-power microcontroller, where they achieve preferable accuracy to XGBoost while subject to stringent memory constraints, and (3) ultra-low-cost chips, where they consistently outperform small models in both accuracy and projected hardware area. DWNs also compare favorably against leading approaches for tabular datasets, with higher average rank. Overall, our work positions DWNs as a pioneering solution for edge-compatible high-throughput neural networks.
Alan T. L. Bacellar, Zachary Susskind, Maurício Breternitz, Eugene John, Lizy Kurian John, Priscila M. V. Lima, Felipe M. G. França
ICML3
2024 Soon Filter: Advancing Tiny Neural Architectures for High Throughput Edge Inference
abstract
As Deep Neural Networks become more complex and computationally demanding, efficient models for inference at the edge, particularly multiplication-free ones, have gained significant attention. The Ultra Low-Energy Edge Neural Network (ULEEN) is a notable architecture optimized for high throughput edge designs. ULEEN uniquely employs Bloom Filters with binary values to compute neuron activation, boasting better efficiency metrics than Binary Neural Networks (BNNs). This work uncovers a gradient back-propagation bottleneck within ULEEN’s Bloom filters and introduces a simplified version of it as a solution: the "Soon Filter". Both theoretically and empirically, we demonstrate that our approach improves gradient back-propagation efficiency. Tests on MLPerf Tiny, MNIST and various UCI datasets reveal that our method surpasses ULEEN, BNN, and DeepShift. Notably, with MLPerf KWS (Key Word Spotting) dataset, we achieve 69.6% accuracy with only 101KiB, while ULEEN, BNN and DeepShift achieve only 67.4%, 55.9%, and 24.9% respectively. Remarkably, we also achieve 67.7% accuracy with only 50KiB, resulting in a 2x model size reduction compared to ULEEN while maintaining similar accuracy (+0.3%). This results underscores the promising potential of our solution for efficient inference at the edge in applications that rely on high throughput architectures.
Alan T. L. Bacellar, Zachary Susskind, Maurício Breternitz, Lizy Kurian John, Felipe M. G. França, Priscila M. V. Lima
IJCNN3
2024 Memory-efficient DRASiW Models
Otávio O. Napoli, Ana de Almeida 0002, Edson Borin, Maurício Breternitz
Neurocomputing4
2023 COIN: Combinational Intelligent Networks
abstract
We introduce Combinational Intelligent Networks (COIN), a machine learning technique that targets edge inference using low-resourced FPGAs or ASICs. COIN is an improvement on LogicWiSARD, a recent weightless neural network that achieves low power, small area, and high throughput. We convert the LogicWiSARD model into a binary neural network, train it using backpropagation, and then convert it to a COIN model. As a result, COIN can achieve higher accuracy than LogicWiSARD or it can require significantly fewer hardware resources when comparing models with similar accuracies. In comparison to a BNN implementation, FINN, small and large COIN models are more energy efficient demonstrating up to 11.5x higher inferences/Joule at similar accuracy. Our tool executes the complete flow, from training to RTL. and is publicly available.
Igor D. S. Miranda, Aman Arora 0001, Zachary Susskind, Josias S. A. Souza, Mugdha P. Jadhao, Luis A. Q. Villon, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Maurício Breternitz, Lizy Kurian John
ASAP10
2023 Efficient Knowledge Aggregation Methods for Weightless Neural Networks
abstract
Weightless Neural Networks (WNN) are good candidates for Federated Learning scenarios due to their robustness and computational lightness.In this work, we show that it is possible to aggregate the knowledge of multiple WNNs using more compact data structures, such as Bloom Filters, to reduce the amount of data transferred between devices.Finally, we explore variations of Bloom Filters and found that a particular data-structure, the Count-Min Sketch (CMS), is a good candidate for aggregation.Costing at most 3% of accuracy, CMS can be up to 3x smaller when compared to previous approaches, specially for large datasets.
Otávio O. Napoli, Ana de Almeida 0002, José Miguel Salles Dias, Luís Brás Rosário, Edson Borin, Maurício Breternitz
ESANN6
2023 An FPGA-Based Weightless Neural Network for Edge Network Intrusion Detection
abstract
Algorithms for mobile networking are increasingly being moved from centralized servers towards the edge in order to decrease latency and improve the user experience. While much of this work is traditionally done using ASICs, 6G emphasizes the adaptability of algorithms for specific user scenarios, which motivates broader adoption of FPGAs. In this paper, we propose the FPGA-based Weightless Intrusion Warden (FWIW), a novel solution for detecting anomalous network traffic on edge devices. While prior work in this domain is based on conventional deep neural networks (DNNs), FWIW incorporates a weightless neural network (WNN), a table lookup-based model which learns sophisticated nonlinear behaviors. This allows FWIW to achieve accuracy far superior to prior FPGA-based work at a very small fraction of the model footprint, enabling deployment on small, low-cost devices. FWIW achieves a prediction accuracy of 98.5% on the UNSW-NB15 dataset with a total model parameter size of just 192 bytes, reducing error by 7.9x and model size by 262x vs. LogicNets, the best prior edge-optimized implementation. Implemented on a Xilinx Virtex UltraScale+ FPGA, FWIW demonstrates a 59x reduction in LUT usage with a 1.6x increase in throughput. The accuracy of FWIW comes within 0.6% of the best-reported result in literature (Edge-Detect), a model several orders of magnitude larger. Our results make it clear that WNNs are worth exploring in the emerging domain of edge networking, and suggest that FPGAs are capable of providing the extreme throughput needed.
Zachary Susskind, Aman Arora 0001, Alan T. L. Bacellar, Diego Leonel Cadette Dutra, Igor D. S. Miranda, Maurício Breternitz, Priscila M. V. Lima, Felipe M. G. França, Lizy Kurian John
FPGA6
2023 A conditional branch predictor based on weightless neural networks
Luis A. Q. Villon, Zachary Susskind, Alan T. L. Bacellar, Igor D. S. Miranda, Leandro Santiago de Araújo, Priscila M. V. Lima, Maurício Breternitz, Lizy Kurian John, Felipe M. G. França, Diego Leonel Cadette Dutra
Neurocomputing7
2023 ULEEN: A Novel Architecture for Ultra-low-energy Edge Neural Networks
abstract
‘‘Extreme edge” 1 devices, such as smart sensors, are a uniquely challenging environment for the deployment of machine learning. The tiny energy budgets of these devices lie beyond what is feasible for conventional deep neural networks, particularly in high-throughput scenarios, requiring us to rethink how we approach edge inference. In this work, we propose ULEEN, a model and FPGA-based accelerator architecture based on weightless neural networks (WNNs). WNNs eliminate energy-intensive arithmetic operations, instead using table lookups to perform computation, which makes them theoretically well-suited for edge inference. However, WNNs have historically suffered from poor accuracy and excessive memory usage. ULEEN incorporates algorithmic improvements and a novel training strategy inspired by binary neural networks (BNNs) to make significant strides in addressing these issues. We compare ULEEN against BNNs in software and hardware using the four MLPerf Tiny datasets and MNIST. Our FPGA implementations of ULEEN accomplish classification at 4.0–14.3 million inferences per second, improving area-normalized throughput by an average of 3.6× and steady-state energy efficiency by an average of 7.1× compared to the FPGA-based Xilinx FINN BNN inference platform. While ULEEN is not a universally applicable machine learning model, we demonstrate that it can be an excellent choice for certain applications in energy- and latency-critical edge environments.
Zachary Susskind, Aman Arora 0001, Igor D. S. Miranda, Alan T. L. Bacellar, Luis A. Q. Villon, Rafael Fontella Katopodis, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Maurício Breternitz, Lizy Kurian John
ACM Trans. Archit. Code Optim.11
2022 Weightless Neural Networks for Efficient Edge Inference
abstract
Weightless neural networks (WNNs) are a class of machine learning model which use table lookups to perform inference, rather than the multiply-accumulate operations typical of deep neural networks (DNNs). Individual weightless neurons are capable of learning non-linear functions of their inputs, a theoretical advantage over the linear neurons in DNNs, yet state-of-the-art WNN architectures still lag behind DNNs in accuracy on common classification tasks. Additionally, many existing WNN architectures suffer from high memory requirements, hindering implementation. In this paper, we propose a novel WNN architecture, BTHOWeN, with key algorithmic and architectural improvements over prior work, namely counting Bloom filters, hardware-friendly hashing, and Gaussian-based nonlinear thermometer encodings. These enhancements improve model accuracy while reducing size and energy per inference. BTHOWeN targets the large and growing edge computing sector by providing superior latency and energy efficiency to both prior WNNs and comparable quantized DNNs. Compared to state-of-the-art WNNs across nine classification datasets, BTHOWeN on average reduces error by more than 40% and model size by more than 50%. We demonstrate the viability of a hardware implementation of BTHOWeN by presenting an FPGA-based inference accelerator, and compare its latency and resource usage against similarly accurate quantized DNN inference accelerators, including multi-layer perceptron (MLP) and convolutional models. The proposed BTHOWeN models consume almost 80% less energy than the MLP models, with nearly 85% reduction in latency. In our quest for efficient ML on the edge, WNNs are clearly deserving of additional attention.
Zachary Susskind, Aman Arora 0001, Igor D. S. Miranda, Luis A. Q. Villon, Rafael Fontella Katopodis, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Maurício Breternitz, Lizy Kurian John
PACT10
2022 LogicWiSARD: Memoryless Synthesis of Weightless Neural Networks
abstract
Weightless neural networks (WNNs) are an alternative pattern recognition technique where RAM nodes function as neurons. As both training and inference require mostly table lookups, few additions, and no multiplications, WNNs are suitable for high-performance and low-power embedded applications. This work introduces a novel approach to implement WiSARD, the leading WNN state-of-the-art architecture, completely eliminating memories and arithmetic circuits and utilizing only logic functions. The approach creates compressed minimized implementations by converting trained WNN nodes from lookup tables to logic functions. The proposed LogicWiSARD is implemented in FPGA and ASIC technologies to illustrate its suitability for edge inference. Experimental results show more than 80% reduction in energy consumption when the proposed LogicWiSARD model is compared with a multilayer perceptron network (MLP) of equivalent accuracy. Compared to previous work on FPGA implementations for WNNs, convolutional neural networks, and binary neural networks, the energy savings of LogicWiSARD range between 32.2% and 99.6%.
Igor D. S. Miranda, Aman Arora 0001, Zachary Susskind, Luis A. Q. Villon, Rafael Fontella Katopodis, Diego Leonel Cadette Dutra, Leandro Santiago de Araújo, Priscila M. V. Lima, Felipe M. G. França, Lizy Kurian John, Maurício Breternitz
ASAP11
2022 Distributive Thermometer: A New Unary Encoding for Weightless Neural Networks
abstract
The binary encoding of real valued inputs is a crucial part of Weightless Neural Networks.The Linear Thermometer and its variations are the most prominent methods to determine binary encoding for input data but, as they make assumptions about the input distribution, the resulting encoding is sub-optimal and possibly wasteful when the assumption is incorrect.We propose a new thermometer approach that doesn't require such assumptions.Our results show that it achieves similar or better accuracy when compared to a thermometer that correctly assumes the distribution, and accuracy gains up to 26.3% when other thermometer representations assume an unsound distribution.
Alan T. L. Bacellar, Zachary Susskind, Luis A. Q. Villon, Igor D. S. Miranda, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Maurício Breternitz, Lizy Kurian John, Priscila M. V. Lima, Felipe M. G. França
ESANN7
2022 Pruning Weightless Neural Networks
abstract
Weightless neural networks (WNNs) are a type of machine learning model which perform prediction using lookup tables (LUTs) instead of arithmetic operations.Recent advancements in WNNs have reduced model sizes and improved accuracies, reducing the gap in accuracy with deep neural networks (DNNs).Modern DNNs leverage "pruning" techniques to reduce model size, but this has not previously been explored for WNNs.We propose a WNN pruning strategy based on identifying and culling the LUTs which contribute least to overall model accuracy.We demonstrate an average 40% reduction in model size with at most 1% reduction in accuracy.
Zachary Susskind, Alan T. L. Bacellar, Aman Arora 0001, Luis A. Q. Villon, Renan Mendanha, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Igor D. S. Miranda, Maurício Breternitz, Lizy Kurian John
ESANN11
2022 A WiSARD-based conditional branch predictor
abstract
Conditional branch prediction is a technique used to speculatively execute instructions before knowing the direction of conditional branch statements. Perceptron-based predictors have been extensively studied, however, they need large input sizes for the data to be linearly separable. To learn nonlinear functions from the inputs, we propose a conditional branch predictor based on the WiSARD model and compare it with two state-of-the-art predictors, the TAGE-SC-L and the Multiperspective Perceptron. We show that the WiSARD-based predictor with a smaller input size outperforms the perceptron-based predictor by about 0.09% and achieves similar accuracy to that of TAGE-SC-L.
Luis A. Q. Villon, Zachary Susskind, Alan T. L. Bacellar, Igor D. S. Miranda, Leandro Santiago de Araújo, Priscila M. V. Lima, Maurício Breternitz, Lizy Kurian John, Felipe M. G. França, Diego Leonel Cadette Dutra
ESANN7
2021 Smart selection of optimizations in dynamic compilers
abstract
Summary Dynamic compilers perform compilation and generation of target code during runtime, implying that the compilation time is added into the program runtime. Thus, to build a high‐performing dynamic compilation system, it is crucial to be able to generate high‐quality code and, at the same time, have a small compilation cost. In this article, we present an approach that uses machine learning to select sequences of optimization for dynamic compilation that considers both code quality and compilation overhead. Our approach starts by training a model, offline, with a knowledge bank of those sequences with low overhead and high‐quality code generation capability using a genetic heuristic. Then, this bank is used to guide the smart selection of optimizations sequences for the compilation of code fragments during the emulation of an application. We evaluate the proposed strategy in two LLVM‐based dynamic binary translators, namely OI‐DBT and HQEMU, and show that these two translators can achieve average speedups of 1.26x and 1.15x in MiBench and Spec Cpu benchmarks, respectively.
Vanderson Martins do Rosário, Anderson Faustino da Silva, Thais Aparecida Silva Camacho, Otávio O. Napoli, Maurício Breternitz, Edson Borin
Concurr. Comput. Pract. Exp.5
2021 Efficiency and scalability of multi-lane capsule networks (MLCN)
Vanderson Martins do Rosário, Maurício Breternitz, Edson Borin
J. Parallel Distributed Comput.2
2020 A unified model for accelerating unsupervised iterative re-ranking algorithms
abstract
Summary Despite the continuous advances in image retrieval technologies, performing effective and efficient content‐based searches remains a challenging task. Unsupervised iterative re‐ranking algorithms have emerged as a promising solution and have been widely used to improve the effectiveness of multimedia retrieval systems. Although substantially more efficient than related approaches based on diffusion processes, these re‐ranking algorithms can still be computationally costly, demanding the specification and implementation of efficient big multimedia analysis approaches. Such demand associated with the significant potential for parallelization and highly effective results achieved by recently proposed re‐ranking algorithms creates the need for exploiting efficiency vs effectiveness trade‐offs. In this article, we introduce a class of unsupervised iterative re‐ranking algorithms and present a model that can be used to guide their implementation and optimization for parallel architectures. We also analyze the impact of the parallelization on the performance of four algorithms that belong to the proposed class: Contextual Spaces, RL‐Sim, Contextual Re‐ranking, and Cartesian Product of Ranking References. The experiments show speedups that reach up to 6.0×, 16.1×, 3.3×, and 7.1× for each algorithm, respectively. These results demonstrate that the proposed parallel programming model can be successfully applied to various algorithms and used to improve the performance of multimedia retrieval systems.
Flávia Pisani, Lucas Pascotti Valem, Daniel C. G. Pedronette, Ricardo da Silva Torres, Edson Borin, Maurício Breternitz
Concurr. Comput. Pract. Exp.6
2020 Weightless Neural Networks as Memory Segmented Bloom Filters
Leandro Santiago de Araújo, Letícia Dias Verona, Fábio Medeiros Rangel, Fabrício Firmino de Faria, Daniel Sadoc Menasché, Wouter Caarls, Maurício Breternitz, Sandip Kundu, Priscila M. V. Lima, Felipe M. G. França
Neurocomputing7
2019 Memory Efficient Weightless Neural Network using Bloom Filter
Leandro Santiago de Araújo, Letícia Dias Verona, Fábio Medeiros Rangel, Fabrício Firmino de Faria, Daniel Sadoc Menasché, Wouter Caarls, Maurício Breternitz, Sandip Kundu, Priscila M. V. Lima, Felipe M. G. França
ESANN7
2019 Efficiency and Scalability of Multi-lane Capsule Networks (MLCN)
Vanderson Martins do Rosário, Maurício Breternitz, Edson Borin
SBAC-PAD2
2019 The Multi-Lane Capsule Network
abstract
We introduce multi-lane capsule networks (MLCN), which are a separable and resource efficient organization of capsule networks (CapsNet) that allows parallel processing while achieving high accuracy at reduced cost. A MLCN is composed of a number of (distinct) parallel lanes, each contributing to a dimension of the result, trained using the routing-by-agreement organization of CapsNet. Our results indicate similar accuracy with a much-reduced cost in number of parameters for the Fashion-MNIST and Cifar10 datasets. They also indicate that the MLCN outperforms the original CapsNet when using a proposed novel configuration for the lanes. MLCN also has faster training and inference times, being more than two-fold faster than the original CapsNet in a same accelerator.
Vanderson Martins do Rosário, Edson Borin, Maurício Breternitz
IEEE Signal Process. Lett.3
2018 ComP-net: command processor networking for efficient intra-kernel communications on GPUs
abstract
Current state-of-the-art in GPU networking advocates a host-centric model that reduces performance and increases code complexity. Recently, researchers have explored several techniques for networking within a GPU kernel itself. These approaches, however, suffer from high latency, waste energy on the host, and are not scalable with larger/more GPUs on a node. In this work, we introduce Command Processor Networking (ComP-Net), which leverages the availability of scalar cores integrated on the GPU itself to provide high-performance intra-kernel networking. ComP-Net enables efficient synchronization between the Command Processors and Compute Units on the GPU through a line locking scheme implemented in the GPU's shared last-level cache. We illustrate that ComP-Net can improve application performance by up to 20% and provide up to 50% reduction in energy consumption vs. competing networking techniques across a Jacobi stencil, allreduce collective, and machine learning applications.
Michael LeBeane, Khaled Hamidouche, Brad Benton, Maurício Breternitz, Steven K. Reinhardt, Lizy Kurian John
PACT4
2017 GPU triggered networking for intra-kernel communications
abstract
GPUs are widespread across clusters of compute nodes due to their attractive performance for data parallel codes. However, communicating between GPUs across the cluster is cumbersome when compared to CPU networking implementations. A number of recent works have enabled GPUs to more naturally access the network, but suffer from performance problems, require hidden CPU helper threads, or restrict communications to kernel boundaries.
Michael LeBeane, Khaled Hamidouche, Brad Benton, Maurício Breternitz, Steven K. Reinhardt, Lizy Kurian John
SC4
2016 Extended task queuing: active messages for heterogeneous systems
abstract
Accelerators have emerged as an important component of modern cloud, datacenter, and HPC computing environments. However, launching tasks on remote accelerators across a network remains unwieldy, forcing programmers to send data in large chunks to amortize the transfer and launch overhead. By combining advances in intra-node accelerator unification with one-sided Remote Direct Memory Access (RDMA) communication primitives, it is possible to efficiently implement lightweight tasking across distributed-memory systems. This paper introduces Extended Task Queuing (XTQ), an RDMA-based active messaging mechanism for accelerators in distributed-memory systems. XTQ's direct NIC-to-accelerator communication decreases inter-node GPU task launch latency by 10-15% for small-to-medium sized messages and ameliorates CPU message servicing overheads. These benefits are shown in the context of MPI accumulate, reduce, and allreduce operations with up to 64 nodes. Finally, we illustrate how XTQ can improve the performance of popular deep learning workloads implemented in the Computational Network Toolkit (CNTK).
Michael LeBeane, Brandon Potter, Abhisek Pan, Alexandru Dutu, Vinay Agarwala, Wonchan Lee, Deepak Majeti, Bibek Ghimire, Eric Van Tassell, Samuel Wasmundt, Brad Benton, Maurício Breternitz, Michael L. Chu, Mithuna Thottethodi, Lizy Kurian John, Steven K. Reinhardt
SC12
2016 HadoopCL2: Motivating the Design of a Distributed, Heterogeneous Programming System With Machine-Learning Applications
abstract
Machine learning (ML) algorithms have garnered increased interest as they demonstrate improved ability to extract meaningful trends from large, diverse, and noisy data sets. While research is advancing the state-of-the-art in ML algorithms, it is difficult to drastically improve the real-world performance of these algorithms. Porting new and existing algorithms from single-node systems to multi-node clusters, or from architecturally homogeneous systems to heterogeneous systems, is a promising optimization technique. However, performing optimized ports is challenging for domain experts who may lack experience in distributed and heterogeneous software development. This work explores how challenges in ML application development on heterogeneous, distributed systems shaped the development of the HadoopCL2 (HCL2) programming system. ML applications guide this work because they exhibit features that make application development difficult: large & diverse datasets, complex algorithms, and the need for domain-specific knowledge. The goal of this work is a general, MapReduce programming system that outperforms existing programming systems. This work evaluates the performance and portability of HCL2 against five ML applications from the Mahout ML framework on two hardware platforms. HCL2 demonstrates speedups of greater than 20x relative to Mahout for three computationally heavy algorithms and maintains minor performance improvements for two I/O bound algorithms.
Max Grossman, Maurício Breternitz, Vivek Sarkar
IEEE Trans. Parallel Distributed Syst.2
2014 Adaptive global power optimization for Web servers
Leonardo Piga, Reinaldo A. Bergamaschi, Maurício Breternitz, Sandro Rigo
J. Supercomput.3
2013 Image Re-ranking Acceleration on GPUs
abstract
Huge image collections are becoming available lately. In this scenario, the use of Content-Based Image Retrieval (CBIR) systems has emerged as a promising approach to support image searches. The objective of CBIR systems is to retrieve the most similar images in a collection, given a query image, by taking into account image visual properties such as texture, color, and shape. In these systems, the effectiveness of the retrieval process depends heavily on the accuracy of ranking approaches. Recently, re-ranking approaches have been proposed to improve the effectiveness of CBIR systems by taking into account the relationships among images. The re-ranking approaches consider the relationships among all images in a given dataset. These approaches typically demands a huge amount of computational power, which hampers its use in practical situations. On the other hand, these methods can be massively parallelized. In this paper, we propose to speedup the computation of the RL-Sim algorithm, a recently proposed image re-ranking approach, by using the computational power of Graphics Processing Units (GPU). GPUs are emerging as relatively inexpensive parallel processors that are becoming available on a wide range of computer systems. We address the image re-ranking performance challenges by proposing a parallel solution designed to fit the computational model of GPUs. We conducted an experimental evaluation considering different implementations and devices. Experimental results demonstrate that significant performance gains can be obtained. Our approach achieves speedups of 7x from serial implementation considering the overall algorithm and up to 36x on its core steps.
Daniel C. G. Pedronette, Ricardo da Silva Torres, Edson Borin, Maurício Breternitz
SBAC-PAD4
2012 Efficient Image Re-Ranking Computation on GPUs
abstract
The huge growth of image collections and multimedia resources available is remarkable. One of the most common approaches to support image searches relies on the use of Content-Based Image Retrieval (CBIR) systems. CBIR systems aim at retrieving the most similar images in a collection, given a query image. Since the effectiveness of those systems is very dependent on the accuracy of ranking approaches, re-ranking algorithms have been proposed to exploit contextual information and improve the effectiveness of CBIR systems. Image re-ranking algorithms typically consider the relationship among every image in a given dataset when computing the new ranking. This approach demands a huge amount of computational power, which may render it prohibitive on very large data sets. In order to mitigate this problem, we propose using the computational power of Graphics Processing Units (GPU) to speedup the computation of image re-ranking algorithms. GPUs are fast emerging and relatively inexpensive parallel processors that are becoming available on a wide range of computer systems. In this paper, we propose a parallel implementation of an image re-ranking algorithm designed to fit the computational model of GPUs. Experimental results demonstrate that relevant performance gains can be obtained by our approach.
Daniel C. G. Pedronette, Ricardo da Silva Torres, Edson Borin, Maurício Breternitz
ISPA4
2012 Cloud Workload Analysis with SWAT
abstract
This note describes the Synthetic Workload Application Toolkit (SWAT) and presents the results from a set of experiments on some key cloud workloads. SWAT is a software platform that automates the creation, deployment, provisioning, execution, and (most importantly) data gathering of synthetic compute workloads on clusters of arbitrary size. SWAT collects and aggregates data from application execution logs, operating system call interfaces, and micro architecture-specific program counters. The data collected by SWAT are used to characterize the effects of network traffic, file I/O, and computation on program performance. The output is analyzed to provide insight into the design and deployment of cloud workloads and systems. Each workload is characterized according to its scalability with the number of server nodes and Hadoop server jobs, sensitivity to network characteristics (bandwidth, latency, statistics on packet size), and computation vs. I/O intensity as these values adjusted via workload-specific parameters. (In the future, we will use SWAT's benchmark synthesizer capability.) We also characterize micro-architectural characteristics that give insight on the micro architecture of processors better suited for this class of workloads. We contrast our results with prior work on Cloud Suite [5], validating some conclusions and providing further insight into others. This illustrates SWAT's data collection capabilities and usefulness to obtain insight on cloud applications and systems.
Maurício Breternitz, Keith Lowery, Anton Charnoff, Patryk Kaminski, Leonardo Piga
SBAC-PAD1
2011 LAR-CC: Large atomic regions with conditional commits
abstract
HW/SW Co-designed systems rely on dynamic binary translation and optimizations for efficient execution of binary code. Due to memory ordering properties and other architectural constraints, most binary optimizations are applied to regions of code that are atomically executed. To ensure that the underlying hardware has enough speculative resources to execute the whole atomic region, these systems typically form short atomic regions, with only 20 to 30 instructions. However, the shorter is the atomic region the smaller is the scope for optimizations. We present LAR-CC, a novel technique that enables HW/SW co-designed systems to optimize large atomic regions and dynamically fit them into the available speculative hardware resources by means of conditional commits. The LAR-CC technique consists of two major components: 1) conditional branch instructions to conditionally skip commit operations; 2) code transformations that replace commit operations by conditional commits and enable optimizations to be applied on the large atomic regions. Our experiments show that LAR-CC can effectively achieve dynamic atomic region sizes larger than 1000 instructions, providing sufficiently large scope to apply many advanced optimizations on HW/SW co-designed systems.
Edson Borin, Youfeng Wu, Maurício Breternitz, Cheng Wang 0013
CGO3
2011 Structure-Constrained Microcode Compression
abstract
Microcode enables programmability of (micro) architectural structures to enhance functionality and to apply patches to an existing design. As more features get added to a CPU core, the area and power costs associated with microcode increase. One solution to address the microcode size issue is to store the microcode in a compressed form and decompress it during execution. Furthermore, the reuse of a single hardware building block layout to implement different dictionaries in the two-level microcode compression reduces the cost and the design time of the decompression engine. However, the reuse of the hardware building block imposes structural constraints to the compression algorithm, and existing algorithms may yield poor compression. In this paper, we develop the SC2 algorithm that considers the structural constraint in its objective function and reduces the area expansion when reusing hardware building blocks to implement different dictionaries. Our experimental results show that the SC2 algorithm is able to produce similar sized dictionaries and achieves the similar compression ratio to the non-constrained algorithm.
Edson Borin, Guido Araujo, Maurício Breternitz, Youfeng Wu
SBAC-PAD3
2010 TAO: two-level atomicity for dynamic binary optimizations
abstract
Dynamic binary translation is a key component of Hardware/Software (HW/SW) co-design, which is an enabling technology for processor microarchitecture innovation. There are two well-known dynamic binary optimization techniques based on atomic execution support. Frame-based optimizations leverage processor pipeline support to enable atomic execution of hot traces. Region level optimizations employ transactional-memory-like atomicity support to aggressively optimize large regions of code. In this paper we propose a two-level atomic optimization scheme which not only overcomes the limitations of the two approaches, but also boosts the benefits of the two approaches effectively. Our experiment shows that the combined approach can achieve a total of 21.5% performance improvement over an aggressive out-of-order baseline machine and improve the performance over the frame-based approach by an additional 5.3%.
Edson Borin, Youfeng Wu, Cheng Wang 0013, Wei Liu 0014, Maurício Breternitz, Shiliang Hu, Esfir Natanzon, Shai Rotem, Roni Rosner
CGO5
2008 A Segmented Bloom Filter Algorithm for Efficient Predictors
abstract
Bloom Filters are a technique to reduce the effects of conflicts/interference in hash table-like structures. Conventional hash tables store information in a single location which is susceptible to destructive interference through hash conflicts. A Bloom Filter uses multiple hash functions to store information in several locations, and recombines the information through some voting mechanism. Many microarchitectural predictors use simple single-index hash tables to make binary 0/1 predictions, and Bloom Filters help improve predictor accuracy. However, implementing a true Bloom Filter requires k hash functions, which in turn implies a k-ported hash table, or k sequential accesses. Unfortunately,the area of a hardware table increases quadratically with the port count, increasing costs of area, latency and power consumption. We propose a simple but elegant modification to the Bloom Filter algorithm that uses banking combined with special hash functions that guarantee all hash indexes fall into non-conflicting banks. We evaluate several applications of our Banked Bloom Filter (BBF) prediction in processors: BBF branch prediction, BBF load hit/miss prediction, and BBF last-tag prediction. We show that BBF predictors can provide accurate predictions with substantially less cost than previous techniques.
Maurício Breternitz, Gabriel H. Loh, Bryan Black, Jeff Rupley, Peter G. Sassone, Wesley Attrot, Youfeng Wu
SBAC-PAD1
2007 Impacts of Multiprocessor Configurations on Workloads in Bioinformatics
abstract
Bioinformatics is among the most active research areas in computer science. In this study, we investigate a suite of workloads in bioinformatics on two multiprocessor systems with different configurations, and examine the effects of the configurations on the performance of the workloads. Our result indicates that the configurations of the multiprocessor systems have significant impact on the performance and scalability of the workloads. For example, a number of workloads have significantly higher scalability on one of the systems, but poorer absolute performance than on the other system. However, traditional scalability failed to capture the impacts of the system configurations on the workloads. We present insights on what kinds of workloads will run faster on which systems and propose new metrics to capture the impacts of multiple processor configurations on the workloads. These findings not only provide an easy way to compare results running on different systems, but also enable re-configuration of the underlying systems to run specific workloads efficiently. We also show how processor mapping and loop spreading may help map the workoads to the underlining multiprocessor configuration and achieve consistent scalability for these workloads.
Youfeng Wu, Maurício Breternitz, Victor Ying
SBAC-PAD2
2006 Clustering-Based Microcode Compression
abstract
Microcode enables programmability of (micro) architectural structures to enhance functionality and to apply patches to an existing design. As more features get added to a CPU core, the area and power costs associated with microcode increase. A recent Intel internal design targeted at low power and small footprint has estimated the costs of the microcode ROM to approach 20% of the total die area (and associated power consumption). Therefore, it is desirable to apply compression techniques to microcode. Microcode poses unique challenges for compression due to the long instruction format, the hand-coded nature of the programs and the stringent performance requirements that require fast decompression. This paper describes techniques for microcode compression that achieve .significant area and power savings, while presenting a streamlined architecture that enables high throughput within the constraints of a high performance CPU. The paper presents results for microcode compression on several commercial CPU designs which demonstrates compression ratios ranging from 50% to 62%.
Edson Borin, Maurício Breternitz, Youfeng Wu, Guido Araujo
ICCD2
2004 The Accuracy of Initial Prediction in Two-Phase Dynamic Binary Translators
abstract
Dynamic binary translators use a two-phase approach to identify and optimize frequently executed code dynamically. In the first step (profiling phase), blocks of code are interpreted or quickly translated to collect execution frequency information for the blocks. In the second phase (optimization phase), frequently executed blocks are grouped into regions and advanced optimizations are applied on them. This approach implicitly assumes that the initial profile of each block is representative of the block throughout its lifetime. We investigate the ability of the initial profile to predict the average program behavior. We compare the predicted behavior of varying lengths of the initial execution with the average program behavior for the whole program execution, and use the prediction from the training input as the reference. Our result indicates that, for the SPEC2000 benchmarks, even very short initial profiles have comparable prediction accuracy to the traditional profile-guided optimizations using the training input, although the initial profile is inadequate for predicting loop trip count information for some integer programs and several benchmarks can benefit from phase-awareness during dynamic binary translation.
Youfeng Wu, Maurício Breternitz, Justin Quek, Orna Etzion, Jesse Fang
CGO2
1997 Enhanced Compression Techniques to Simplify Programm Decompression and Execution
abstract
Compressing instruction sequences can reduce the cost of embedded systems by reducing program ROM-size requirements. Compression also facilitates the use of RISC core architectures, like the PowerPC/sup TM/ architecture, in embedded systems. Compression techniques are presented which enable decompression and execution of compressed code to occur without the need of a lookaside table (LAT) or cache lookaside buffer (CLB). These techniques successfully merge code modification and compression into a single software preprocessing step. Decompression and execution of compressed code are made very simple. An application of these techniques to about 120000 instructions of PowerPC firmware code is described.
Maurício Breternitz
ICCD1
1996 Design Tradeoffs and Experience with Motorola PowerPC? Migration Tool
abstract
The Motorola PowerPC migration tools enable the conversion of assembly programs from other architectures to PowerPC. This paper describes the design approach and experience with the tool to translate x86 assembly programs to PowerPC. The key problems of handling 16-bit code, the effects of masking 16-bit operations into 32-bit registers and optimization of condition flags are discussed. The efficiency of translation and the effects of architectural constraints on design tradeoffs are analyzed.
Maurício Breternitz, A. Manikonda, M. Ommerman, W. Su, A. Thornto
ICCD1
1991 Implementation Optimization Techniques for Architecture Synthesis of Application-Specific Processors
abstract
Article Free Access Share on Implementation optimization techniques for architecture synthesis of application-specific processors Authors: Mauricio Breternitz Advanced Workstations Division, IBM-Austin, TX Advanced Workstations Division, IBM-Austin, TXView Profile , John Paul Shen ECE Department, Carnegie-Mellon University ECE Department, Carnegie-Mellon UniversityView Profile Authors Info & Claims MICRO 24: Proceedings of the 24th annual international symposium on MicroarchitectureSeptember 1991 Pages 114–123https://doi.org/10.1145/123465.123488Published:01 September 1991Publication History 0citation252DownloadsMetricsTotal Citations0Total Downloads252Last 12 Months11Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Maurício Breternitz, John Paul Shen
MICRO1
1990 Architecture Synthesis of High-Performance Application-Specific Processors
abstract
The key principles of the Application-Specific Processor Design (ASPD) methodology include: a semi-custom compilation-driven design/implementation approach, the exploitation of fine-grained parallelism for high performance, and the adaptation of datapath topology to the data transfers required by the application. The powerful microcode compilation techniques of Percolation Scheduling and Pipeline Scheduling extract and enhance the parallelism in the application object code to generate an optimized specification of the target processor. Implementation optimization is performed to allocate functional units and register files. Graph-coloring algorithms minimize the amount of hardware needed to exploit available parallelism. Data memory employs an organization with multiple banks. Compilation techniques are used to allocate data over the memory banks to enhance parallel access.
Maurício Breternitz, John Paul Shen
DAC1
1988 The White Dwarf: A High-Performance Application-Specific Processor
abstract
The design and implementation of a high-performance special-purpose processor, called the White Dwarf, or accelerating finite-element analysis algorithms is presented. The White Dwarf CPU contains two Am2935 32-bit floating-point processors and one Am29332 32-bit arithmetic logic unit (ALU), and uses a wide-instruction-word architecture in which the application algorithm is directly implemented in microcode. The entire system is VME-bus compatible and interfaces with a Sun 3/160 host. The system's potential peak performance is 20 MFLOPS (million floating-point operations per second) a sustained computation rate in excess of 15 MFLOPS is expected. A potential speedup of between one and two orders of magnitude is possible. With a fully populated memory subsystem, the White Dwarf can accommodate finite-element problems involving up to half a million nodes. The system is designed using an approach called application-specific processor design (ASPD). A retargetable compiler has been developed which is capable of generating highly parallel and efficient code for the White Dwarf and other processors with similar architecture. System debug/integration is in progress; a highly useful system is expected.>
Andrew Wolfe, Maurício Breternitz, Chriss Stephens, A. L. Ting, David Blair Kirk, Ronald P. Bianchini Jr., John Paul Shen
ISCA2