Leandro Santiago de Araújo

dblp:230/3564 · also Leandro Santiago 0001 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-3631-8761ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 A Ship Detection Technique Using Weightless Neural Networks
abstract
Maritime vessel detection plays a critical role in navigation safety, surveillance, and environmental monitoring. While deep learning-based models, such as YOLOv8, offer high detection accuracy, they have the need for greater computational resources, making them less suited for embedded and real-time systems. This paper presents an efficient and lightweight vessel detection system by utilizing weightless neural networks for optimal object recognition and computational efficiency. Our approach was evaluated against YOLOv8 in terms of precision, recall, execution time, and energy consumption, using a publicly available ship dataset. The results demonstrate that our method achieves higher precision, comparable F1-score, and significantly lower computational overhead, reducing both inference time and power consumption. While YOLOv8 outperforms our approach in object localization via bounding box estimation, our model accurately localizes vessel positions with small error. Additionally, our framework reduces preprocessing and training time by up to$\mathbf{1 2}$times, making it highly effective for Edge AI and real-time maritime surveillance. By significantly lowering energy demands and computational latency, this research provides a scalable and sustainable alternative for vessel detection in autonomous maritime systems, smart surveillance networks, and embedded vision applications. Future work will explore hybrid models integrating weightless neural networks with deep learning to further enhance detection accuracy while maintaining energy efficiency.
Adriano G. Pereira, Claudio M. de Farias, Leandro Santiago de Araújo
FUSION3
2023 A conditional branch predictor based on weightless neural networks
Luis A. Q. Villon, Zachary Susskind, Alan T. L. Bacellar, Igor D. S. Miranda, Leandro Santiago de Araújo, Priscila M. V. Lima, Maurício Breternitz, Lizy Kurian John, Felipe M. G. França, Diego Leonel Cadette Dutra
Neurocomputing5
2023 ULEEN: A Novel Architecture for Ultra-low-energy Edge Neural Networks
abstract
‘‘Extreme edge” 1 devices, such as smart sensors, are a uniquely challenging environment for the deployment of machine learning. The tiny energy budgets of these devices lie beyond what is feasible for conventional deep neural networks, particularly in high-throughput scenarios, requiring us to rethink how we approach edge inference. In this work, we propose ULEEN, a model and FPGA-based accelerator architecture based on weightless neural networks (WNNs). WNNs eliminate energy-intensive arithmetic operations, instead using table lookups to perform computation, which makes them theoretically well-suited for edge inference. However, WNNs have historically suffered from poor accuracy and excessive memory usage. ULEEN incorporates algorithmic improvements and a novel training strategy inspired by binary neural networks (BNNs) to make significant strides in addressing these issues. We compare ULEEN against BNNs in software and hardware using the four MLPerf Tiny datasets and MNIST. Our FPGA implementations of ULEEN accomplish classification at 4.0–14.3 million inferences per second, improving area-normalized throughput by an average of 3.6× and steady-state energy efficiency by an average of 7.1× compared to the FPGA-based Xilinx FINN BNN inference platform. While ULEEN is not a universally applicable machine learning model, we demonstrate that it can be an excellent choice for certain applications in energy- and latency-critical edge environments.
Zachary Susskind, Aman Arora 0001, Igor D. S. Miranda, Alan T. L. Bacellar, Luis A. Q. Villon, Rafael Fontella Katopodis, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Maurício Breternitz, Lizy Kurian John
ACM Trans. Archit. Code Optim.7
2022 Weightless Neural Networks for Efficient Edge Inference
abstract
Weightless neural networks (WNNs) are a class of machine learning model which use table lookups to perform inference, rather than the multiply-accumulate operations typical of deep neural networks (DNNs). Individual weightless neurons are capable of learning non-linear functions of their inputs, a theoretical advantage over the linear neurons in DNNs, yet state-of-the-art WNN architectures still lag behind DNNs in accuracy on common classification tasks. Additionally, many existing WNN architectures suffer from high memory requirements, hindering implementation. In this paper, we propose a novel WNN architecture, BTHOWeN, with key algorithmic and architectural improvements over prior work, namely counting Bloom filters, hardware-friendly hashing, and Gaussian-based nonlinear thermometer encodings. These enhancements improve model accuracy while reducing size and energy per inference. BTHOWeN targets the large and growing edge computing sector by providing superior latency and energy efficiency to both prior WNNs and comparable quantized DNNs. Compared to state-of-the-art WNNs across nine classification datasets, BTHOWeN on average reduces error by more than 40% and model size by more than 50%. We demonstrate the viability of a hardware implementation of BTHOWeN by presenting an FPGA-based inference accelerator, and compare its latency and resource usage against similarly accurate quantized DNN inference accelerators, including multi-layer perceptron (MLP) and convolutional models. The proposed BTHOWeN models consume almost 80% less energy than the MLP models, with nearly 85% reduction in latency. In our quest for efficient ML on the edge, WNNs are clearly deserving of additional attention.
Zachary Susskind, Aman Arora 0001, Igor D. S. Miranda, Luis A. Q. Villon, Rafael Fontella Katopodis, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Maurício Breternitz, Lizy Kurian John
PACT6
2022 LogicWiSARD: Memoryless Synthesis of Weightless Neural Networks
abstract
Weightless neural networks (WNNs) are an alternative pattern recognition technique where RAM nodes function as neurons. As both training and inference require mostly table lookups, few additions, and no multiplications, WNNs are suitable for high-performance and low-power embedded applications. This work introduces a novel approach to implement WiSARD, the leading WNN state-of-the-art architecture, completely eliminating memories and arithmetic circuits and utilizing only logic functions. The approach creates compressed minimized implementations by converting trained WNN nodes from lookup tables to logic functions. The proposed LogicWiSARD is implemented in FPGA and ASIC technologies to illustrate its suitability for edge inference. Experimental results show more than 80% reduction in energy consumption when the proposed LogicWiSARD model is compared with a multilayer perceptron network (MLP) of equivalent accuracy. Compared to previous work on FPGA implementations for WNNs, convolutional neural networks, and binary neural networks, the energy savings of LogicWiSARD range between 32.2% and 99.6%.
Igor D. S. Miranda, Aman Arora 0001, Zachary Susskind, Luis A. Q. Villon, Rafael Fontella Katopodis, Diego Leonel Cadette Dutra, Leandro Santiago de Araújo, Priscila M. V. Lima, Felipe M. G. França, Lizy Kurian John, Maurício Breternitz
ASAP7
2022 Distributive Thermometer: A New Unary Encoding for Weightless Neural Networks
abstract
The binary encoding of real valued inputs is a crucial part of Weightless Neural Networks.The Linear Thermometer and its variations are the most prominent methods to determine binary encoding for input data but, as they make assumptions about the input distribution, the resulting encoding is sub-optimal and possibly wasteful when the assumption is incorrect.We propose a new thermometer approach that doesn't require such assumptions.Our results show that it achieves similar or better accuracy when compared to a thermometer that correctly assumes the distribution, and accuracy gains up to 26.3% when other thermometer representations assume an unsound distribution.
Alan T. L. Bacellar, Zachary Susskind, Luis A. Q. Villon, Igor D. S. Miranda, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Maurício Breternitz, Lizy Kurian John, Priscila M. V. Lima, Felipe M. G. França
ESANN5
2022 Pruning Weightless Neural Networks
abstract
Weightless neural networks (WNNs) are a type of machine learning model which perform prediction using lookup tables (LUTs) instead of arithmetic operations.Recent advancements in WNNs have reduced model sizes and improved accuracies, reducing the gap in accuracy with deep neural networks (DNNs).Modern DNNs leverage "pruning" techniques to reduce model size, but this has not previously been explored for WNNs.We propose a WNN pruning strategy based on identifying and culling the LUTs which contribute least to overall model accuracy.We demonstrate an average 40% reduction in model size with at most 1% reduction in accuracy.
Zachary Susskind, Alan T. L. Bacellar, Aman Arora 0001, Luis A. Q. Villon, Renan Mendanha, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Igor D. S. Miranda, Maurício Breternitz, Lizy Kurian John
ESANN6
2022 A WiSARD-based conditional branch predictor
abstract
Conditional branch prediction is a technique used to speculatively execute instructions before knowing the direction of conditional branch statements. Perceptron-based predictors have been extensively studied, however, they need large input sizes for the data to be linearly separable. To learn nonlinear functions from the inputs, we propose a conditional branch predictor based on the WiSARD model and compare it with two state-of-the-art predictors, the TAGE-SC-L and the Multiperspective Perceptron. We show that the WiSARD-based predictor with a smaller input size outperforms the perceptron-based predictor by about 0.09% and achieves similar accuracy to that of TAGE-SC-L.
Luis A. Q. Villon, Zachary Susskind, Alan T. L. Bacellar, Igor D. S. Miranda, Leandro Santiago de Araújo, Priscila M. V. Lima, Maurício Breternitz, Lizy Kurian John, Felipe M. G. França, Diego Leonel Cadette Dutra
ESANN5
2021 Gamma - General Abstract Model for Multiset mAnipulation and dynamic dataflow model: An equivalence study
abstract
Abstract With the increase of the search for computational models where the expression of parallelism occurs naturally, some paradigms arise as options for the current generation of computers. In this context, dynamic dataflow and Gamma—GeneralAbstractModel forMultiset mAnipulation—emerge as interesting computational model choices. In dynamic dataflow model, operations are performed as soon as their associated operands are available, without rely on a Program Counter to dictate the execution order of instructions. The Gamma paradigm is based on a parallel multiset rewriting scheme. It provides a nondeterministic execution model inspired by anabstract chemical machinemetaphor, where operations are formulated as reactions that occur freely among matching elements belonging to the multiset. In this work, equivalence relations between the dynamic dataflow and Gamma paradigms are exposed and explored, while methods to convert from dataflow to Gamma paradigm and vice versa are provided. It is shown that vertices and edges of a dynamic dataflow graph can correspond, respectively, to reactions and multiset elements in the Gamma paradigm. This work provides the scientific community with the possibility of taking profit of both parallel programming models, contributing with a versatility component to researchers and developers.
Rui Rodrigues de Mello Junior, Leandro Santiago de Araújo, Tiago A. O. Alves, Leandro A. J. Marzulo, Gabriel Antoine Louis Paillard, Felipe M. G. França
Concurr. Comput. Pract. Exp.2
2020 Building a portable deeply-nested implicit information flow tracking
abstract
Dynamic Information Flow Tracking has been successfully used to prevent a wide range of attacks and detect illegal access to sensitive information. Most proposed solutions only track the explicit information flow where the taint is propagated through data dependencies. However, recent evasion attacks exploit implicit flows, that use control flow in the application, to manipulate the data thus making the malicious activity undetectable. We propose NIFT - a nested implicit flow tracking mechanism that extends explicit propagation to instructions affected by a control dependency. Our technique generates taint instructions at compile time which are executed by specialized hardware to propagate taint implicitly even in cases of deeply-nested branches. In addition, we propose a restricted taint propagation for data executed in conditional branches that affects only immediate instructions instead of all instructions inside the branch scope. Our technique efficiently locates implicit flows and resolves them with negligible performance overhead. Moreover, it mitigates the over-tainting problem.
Leandro Santiago de Araújo, Leandro A. J. Marzulo, Tiago A. O. Alves, Felipe M. G. França, Israel Koren, Sandip Kundu
CF1
2020 Fluid computing: interest-based communication in dataflow/multiset rewriting computing
abstract
Among the existing computational parallel models, Gamma and Dynamic Dataflow are equivalent models where parallel programs can be developed in a natural way. However, the implementation of Gamma paradigm poses several communication and architectural challenges. On the other hand, interest-based protocols have emerged as possible solutions of communication routing enabling an efficient communication in IoT environments. An interesting property of Gamma is the possibility of locality exploration in a non-complex way that sounds quite suitable for IoT applications. This work proposes a novel execution model for Gamma programs integrating the Radnet --- an interest-based protocol --- as communication protocol. Since the interest processing using Gamma paradigm is not a trivial task, a dynamic dataflow graph is used to express the interest as edges between vertices. We explore the equivalence results between Gamma and dataflow and, also, provide experiments showing the potential of Gamma when used to implement approximate computing techniques.
Rui Rodrigues de Mello Junior, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Gabriel Antoine Louis Paillard, Claudio Luis de Amorim, Felipe M. G. França
EATIS2
2020 Fast Deep Neural Networks Convergence using a Weightless Neural Model
Alan T. L. Bacellar, Brunno F. Goldstein, Victor da Cruz Ferreira, Leandro Santiago de Araújo, Priscila M. V. Lima, Felipe M. G. França
ESANN4
2020 Weightless Neural Networks as Memory Segmented Bloom Filters
Leandro Santiago de Araújo, Letícia Dias Verona, Fábio Medeiros Rangel, Fabrício Firmino de Faria, Daniel Sadoc Menasché, Wouter Caarls, Maurício Breternitz, Sandip Kundu, Priscila M. V. Lima, Felipe M. G. França
Neurocomputing1
2020 Cost-effective, Energy-efficient, and Scalable Storage Computing for Large-scale AI Applications
abstract
The growing volume of data produced continuously in the Cloud and at the Edge poses significant challenges for large-scale AI applications to extract and learn useful information from the data in a timely and efficient way. The goal of this article is to explore the use of computational storage to address such challenges by distributed near-data processing. We describe Newport, a high-performance and energy-efficient computational storage developed for realizing the full potential of in-storage processing. To the best of our knowledge, Newport is the first commodity SSD that can be configured to run a server-like operating system, greatly minimizing the effort for creating and maintaining applications running inside the storage. We analyze the benefits of using Newport by running complex AI applications such as image similarity search and object tracking on a large visual dataset. The results demonstrate that data-intensive AI workloads can be efficiently parallelized and offloaded, even to a small set of Newport drives with significant performance gains and energy savings. In addition, we introduce a comprehensive taxonomy of existing computational storage solutions together with a realistic cost analysis for high-volume production, giving a good big picture of the economic feasibility of the computational storage technology.
Jaeyoung Do, Victor da Cruz Ferreira, Hossein Bobarshad, Mahdi Torabzadehkashi, Siavash Rezaei, Ali Heydarigorji, Diego Fonseca Pereira de Souza, Brunno F. Goldstein, Leandro Santiago de Araújo, Min Soo Kim 0009, Priscila M. V. Lima, Felipe M. G. França, Vladimir Castro Alves
ACM Trans. Storage9
2019 Efficient Testing of Physically Unclonable Functions for Uniqueness
abstract
Physically unclonable functions (PUFs) have emerged as lightweight hardware security primitives for implementing secure authentication. Strong PUFs rely on random manufacturing process variation to create unique Boolean mappings from input (challenge) to output (output). For secure authentication, challenge to response mappings are required to be unique for each device. However, uniqueness is not guaranteed by design or manufacturing. Testing for uniqueness and weeding out non-unique parts are the only way to ensure uniqueness of devices. Uniqueness testing can be expensive in time as the challenge-responses of the N th device under-test, must be proven to be different from previously tested N - 1 devices, or the device must be discarded. To reduce the time complexity of uniqueness testing, Multi-Index hashing (MIH) was proposed for online testing in high volume manufacturing. Database search using MIH was shown to be fast, but it suffers from high memory cost. In this paper, we address the memory problem of MIH based uniqueness testing by proposing alternative MIH strategies. Our results indicate that the proposed search strategies can significantly reduce the memory cost without sacrificing performance, requiring ≈ 3.35× less memory with just a 17% performance overhead when testing the uniqueness of 1 million PUFs.
Leandro Santiago de Araújo, Vinay C. Patil, Leandro A. J. Marzulo, Felipe M. G. França, Sandip Kundu
ATS1
2019 Memory Efficient Weightless Neural Network using Bloom Filter
Leandro Santiago de Araújo, Letícia Dias Verona, Fábio Medeiros Rangel, Fabrício Firmino de Faria, Daniel Sadoc Menasché, Wouter Caarls, Maurício Breternitz, Sandip Kundu, Priscila M. V. Lima, Felipe M. G. França
ESANN1
2019 A Feasible FPGA Weightless Neural Accelerator
abstract
AI applications have recently driven the computer architecture industry towards novel and more efficient dedicated hardware accelerators and tools. Weightless Neural Networks (WNNs) is a class of Artificial Neural Networks (ANNs) often applied to pattern recognition problems. It uses a set of Random Access Memories (RAMs) as the main mechanism for training and classifying information regarding a given input pattern. Due to its memory-based architecture, it can be easily mapped onto hardware and also greatly accelerated by means of a dedicated Register Transfer-Level (RTL) architecture designed to enable multiple memory accesses in parallel. On the other hand, a straightforward WNN hardware implementation requires too much memory resources for both ASIC and FPGA variants. This work aims at designing and evaluating a Weightless Neural accelerator designed in High-Level Synthesis (HLS). Our WNN accelerator implements Hash Tables, instead of regular RAMs, to substantially reduce its memory requirements, so that it can be implemented in a fairly small-sized Xilinx FPGA. Performance, circuit-area and power consumption results show that our accelerator can efficiently learn and classify the MNIST dataset in about 8 times faster than the system's embedded ARM processor.
Victor da Cruz Ferreira, Alexandre Solon Nery, Leandro A. J. Marzulo, Leandro Santiago de Araújo, Diego Fonseca Pereira de Souza, Brunno F. Goldstein, Felipe M. G. França, Vladimir Castro Alves
ISCAS4