Lucas B. da Silva

dblp:214/1143 · also Lucas Bragança, Lucas Bragança da Silva · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-9626-0274ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Multi -FPGA streaming using OpenMP
abstract
The growth in the demand for high-performance and power-efficient applications has led to an increasing interest in FPGA-based acceleration. FPGAs have been applied to a wide range of applications. Still, programming them can be a complex task, requiring extensive knowledge of tools and libraries, especially in multi-FPGA architectures. As such, the desire for tools and frameworks to ease the burden and abstract the knowledge of using FPGAs has increased. OpenMP, an already dominant parallel programming model in HPC, has been shown to be a successful approach to program multi-FPGA architecture. This work is based on the OMPC-F framework, which leverages the capability of OpenMP to offload computation to FPGAs. Although OMPC-F abstracts FPGA handling and task distribution from the final user, it does not support streaming computation based on multi-FPGA architectures. Streaming is widely used in FPGA designs to create a pipeline of computation between FPGA kernels. This work uses FPGA kernel binary information to synthesize streams as OpenMP buffers, while adapting the OpenMP dependency system accordingly. The proposal was evaluated in an AMD/Xilinx multi-FPGA system and shows speedups of the order of 6.84x in 8 FPGAs, scaling well with the addition of more kernels and FPGAs to the architecture. Moreover, compared to the regular approach to developing FPGA applications based on MPI+XRT communication, the proposed approach reduces the programming effort by 59% according to various code analysis metrics, resulting in a small average overhead of 4.3% when compared to the MPI+XRT programming model.
Pedro Henrique Di Francia Rosso, Rémy Neveu, Nusrat Jahan Lisa, Lucas B. da Silva, Hervé Yviquel, Sandro Rigo, Vanderlei Bonato, Guido Araujo
J. Parallel Distributed Comput.4
2025 Reconfigurable Domain-Specific Architectures Based on Coarse-Grained Operators in High-Performance FPGAs
abstract
ABSTRACT This work presents the development of reconfigurable processing units capable of encapsulating various operations to design new domain‐specific reconfigurable accelerators. These processing units are known as coarse‐grained operators because they can execute multiple operations. We validated the new operators within the HPCGRA framework, a design environment for creating custom coarse‐grained reconfigurable arrays that run as a virtual layer on commercial field‐programmable gate arrays. The environment is parameterized and employs an intermediate portable format that abstracts away low‐level specific bitstream details, enabling hardware‐agnostic reconfiguration. In this paper, we present three case studies. The first is a domain‐specific accelerator for the K‐means algorithm, which can be entirely reconfigured in less than 2.34 ms to explore various clustering values and attributes, achieving up to 159 Gop/s performance for considering 8 features. The reconfigurability of our accelerator has no impact on performance compared to static versions implemented directly in HLS and RTL. The second case study extends K‐means with a Gini calculation operator for dimensionality reduction, quantization, and quality classification, achieving up to 278 Gops/s. The third case study presents a systolic matrix multiplier, demonstrating the versatility of the environment for designing different architectural domains.
Lucas B. da Silva, César Grandis, Jeronimo Penha, José A. M. Nacif, Ricardo S. Ferreira 0001
Concurr. Comput. Pract. Exp.1
2023 Fast flow cloud: A stream dataflow framework for cloud FPGA accelerator overlays at runtime
abstract
Abstract Cloud FPGAs provide new energy‐efficient opportunities to design dataflow accelerators. Nevertheless, FPGAs still have challenges to overcome for widespread usages, such as programmability, compilation time (minutes to hours), and hardware knowledge, mainly because it is highly challenging for beginners to learn and use FPGAs. The READY tool recently provides compilation time reduction to the range of microseconds using a CGRA overlay and a friendly, high‐level C++ interface for the Intel/Altera HARPv2 FPGA cloud platform. However, the HARPv2 is not available in any commercial cloud platform. This work extends READY by creating the fast flow cloud framework (FFC). First, FFC offers a simple browser‐based graphical interface for less experienced FPGA users. Second, we improve the CGRA overlay portability to include Xilinx FPGAs and a transparent design flow to deploy in the widespread commercial Amazon AWS F1 cloud. Third, we improve the CGRA reconfiguration engine. Also, we compare the overlay performance of HARPv2 and AWS F1 to an eight‐thread XEON processor. Finally, the framework is open‐source for collaborative development and has clearly defined application programming interfaces for future extensions.
Lucas B. da Silva, Michael Canesche, Jeronimo Costa Penha, Josué Campos, José A. M. Nacif, Ricardo S. Ferreira 0001
Concurr. Comput. Pract. Exp.1
2023 Gene regulatory accelerators on cloud FPGA
abstract
Summary Gene regulatory networks (GRN) are dynamic models in time and space. These models are used to predict diseases and in drugs research. GRN models are discrete, and Boolean graphs can efficiently represent them. However, GRN algorithms explore a large solution space with high computational complexity. This work proposes efficient FPGA‐based accelerators to implement two GRN algorithms: attractor computation and Derrida plot. Nevertheless, FPGA accelerator design and deployment are still a challenge. This work presents an accelerator design framework for AWS Amazon FPGA cloud. The framework simplifies the software (SW) and hardware (HW) generation for GRN accelerators. The user provides a high‐level model for the Boolean GRN, and our tool automatically creates AWS‐ready‐to‐deploy software and hardware components. For the attractor and the Derrida plot computation, the proposed FPGA accelerators are on average and faster than a V100 GPU.
Jeronimo Costa Penha, Lucas B. da Silva, Michael Canesche, Dener V. Ribeiro, José A. M. Nacif, Ricardo S. Ferreira 0001
Concurr. Comput. Pract. Exp.2
2021 Google Colab CAD4U: Hands-On Cloud Laboratories for Digital Design
abstract
Google Colab is a cloud Jupyter notebook widespread used to teach machine learning by writing text explanations and Python codes through the browser. This work introduces new Colab extensions to teach logic circuit design, Verilog language, processor, and GPU architectures. Colab allows us to share reproducible experiments on the Web. The students become motivated to do laboratory assignments without download/configure software packages and dependencies on their computers. Furthermore, almost all universities had to shut down due to the COVID-19 pandemic, forcing us to adapt to virtual learning scenarios. Colab provides portability and accessibility since it can even run on smartphones. The lab assignments include intermediate guided exercises, text explanations, figures, online quizzes, problem sets, and basic hands-on tasks. We develop a simple setup for Icarus Verilog, PyEDA, CUDA, Valgrind, and Gem5 frameworks. This work presents Verilog teaching and computer architecture simulation insights by using Valgrind and Gem5, and GPU computer architecture profiling at the thread and instruction assembly level.
Michael Canesche, Lucas B. da Silva, Omar P. Vilela Neto, José A. M. Nacif, Ricardo S. Ferreira 0001
ISCAS2
2021 RESHAPE: A Run-Time Dataflow Hardware-Based Mapping for CGRA Overlays
abstract
Coarse-grained reconfigurable architectures (CGRA) are a power-efficient approach for hardware accelerators. However, there are few EDA tools for CGRA. We develop hardware-based placement and routing (P&R) for fully-pipelined CGRA mapped as an FPGA overlay. The key idea is to use the available FPGA resources to replicate several mapping units, thus exploring parallel execution, area/execution time trade-offs, and achieving near-optimal mapping solutions. Furthermore, our P&R provides portability and an incremental run-time approach. In comparison to VPR and CGRA-ME tools and a time-multiplexer approach, our spatial mapping reduces the P&R execution time, and it improves the performance up to hundreds of Gops/s by using fully-pipelined architectures.
Maria D. Vieira, Michael Canesche, Lucas B. da Silva, Josué Campos, Mateus Silva, Ricardo S. Ferreira 0001, José A. M. Nacif
ISCAS3
2019 ADD: Accelerator Design and Deploy - A tool for FPGA high-performance dataflow computing
abstract
Summary Dataflow‐based FPGA accelerators have become a promising alternative to deliver energy‐efficient high‐performance computing. However, FPGA programming is still a challenge. This paper presents Accelerator Design and Deploy (ADD), a high‐level framework to specify, to simulate, and to implement dataflow accelerators for streaming applications. The framework includes an open dataflow operator library, and templates are provided to easily design new operators. The framework also provides a high‐level and an accurate simulation at circuit level with short execution times. Moreover, ADD provides software and hardware APIs to simplify the integration process, extending the benefits of portability from low‐cost FPGA boards to high performance datacenter FPGA platforms. Our framework supports coupling with high‐level programming languages, and it has been validated on two FPGA platforms: the Intel high‐performance CPU‐FPGA heterogeneous computing platform and an educational FPGA kit. We show that our simple approach presents competitive performance, both in time and energy, when compared to multi‐core and GPU accelerators.
Jeronimo Costa Penha, Lucas B. da Silva, Jansen Silva, Kristtopher Coelho, Hector P. Baranda, José A. M. Nacif, Ricardo S. Ferreira 0001
Concurr. Comput. Pract. Exp.2
2019 READY: A Fine-Grained Multithreading Overlay Framework for Modern CPU-FPGA Dataflow Applications
abstract
In this work, we propose a framework called REconfigurable Accelerator DeploY (READY), the first framework to support polynomial runtime mapping of dataflow applications in high-performance CPU-FPGA platforms. READY introduces an efficient mapping with fine-grained multithreading onto an overlay architecture that hides the latency of a global interconnection network. In addition to our overlay architecture, we show how this system helps solve some of the challenges for FPGA cloud computing adoption in high-performance computing. The framework encapsulates dataflow descriptions by using a target independent, high-level API, and a dataflow model that allows for explicit spatial and temporal parallelism. READY directly maps the dataflow kernels onto the accelerator. Our tool is flexible and extensible and provides the infrastructure to explore different accelerator designs. We validate READY on the Intel Harp platform, and our experimental results show an average 2x execution runtime improvement when compared to an 8-thread multi-core processor.
Lucas B. da Silva, Ricardo S. Ferreira 0001, Michael Canesche, Marcelo de Matos Menezes, Maria D. Vieira, Jeronimo Costa Penha, Peter Jamieson, José A. M. Nacif
ACM Trans. Embed. Comput. Syst.1
2018 From Java to FPGA: An Experience with the Intel HARP System
abstract
Recent years have seen a surge in the popularity of Field-Programmable Gate Arrays (FPGAs). Programmers can use them to develop high-performance systems that are not only efficient in time, but also in energy. Yet, programming FPGAs remains a difficult task. Even though there exist today OpenCL interfaces to synthesize such hardware, higher-level programming languages, such as Java, C# or Python remain distant from them. In this paper, we describe a compiler, and its supporting runtime environment, that reduces this distance, translating functional code written in Java to the Intel HARP platform. Thus, we bring two contributions. First, the insight that a functional-style library is a good starting point to bridge the gap between high-level programming idioms and FPGAs. Second, the implementation of this system itself, including the compiler, its intermediate representation, and all the runtime support necessary to shield developers from the task of transferring data back and forth between the host CPU and the accelerator. To demonstrate the effectiveness of our system, we have used it to implement different benchmarks, used in image processing and data-mining. For large inputs, we can observe consistent 20x speedups over the Java Virtual Machine across all our benchmarks. Depending on the target function that we compile, this speedup can achieve 280x.
Pedro Caldeira, Jeronimo Costa Penha, Lucas B. da Silva, Ricardo S. Ferreira 0001, José A. M. Nacif, Renato Ferreira 0001, Fernando Magno Quintão Pereira
SBAC-PAD3