EDBT 2026 Demo / reviewers in the wild / expert
Ricardo S. Ferreira 0001
dblp:12/10030 · also Ricardo Ferreira 0001, Ricardo Santos Ferreira 0001, Ricardo dos Santos Ferreira 0001
· DBLP profile ↗
32ranked-venue papers
9as first author
15since 2021 · last 2025
0000-0003-1802-7829ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 8 first-author · 13 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SmartMap: Architecture-Agnostic CGRA Mapping Using Graph Traversal and Reinforcement LearningabstractCoarse-Grained Reconfigurable Architectures (CGRAs) have been the subject of extensive research due to their balance between performance, energy efficiency, and flexibility. CGRAs must be capable of executing a dataflow graph (DFG), which depends on a compiler producing quality valid mappings with feasible running time performance and portable mapping DFGs on different CGRA architectures. Machine learning-based compilers have shown promising results by presenting high quality and performance but offer limited portability. Moreover, some approaches do not explore efficient placement methods or do not demonstrate whether scaling to more challenging, less connected architectures. This paper presents SmartMap, an architecture-agnostic framework that uses an actor-critic reinforcement learning method applied to a Monte-Carlo Tree Search (MCTS) to learn how to map a DFG onto a CGRA. This framework offers full portability using a state-action representation layer in the policy network instead of a probability distribution over actions. SmartMap uses a graph traversal placement method to provide scalability and improve the efficiency of MCTS by enabling more efficient exploration during the search. Our results show that SmartMap has 2.81x more mapping capacity, a 16.82x speed-up in compilation time, and consumes fewer resources compared to the state-of-the-art. Fabio Ramos 0001, Pedro E. F. Realino, Wagner A. Junior, Alex Borges Vieira, Ricardo S. Ferreira 0001, José A. M. Nacif |
DATE | 5 |
| 2025 | Unblocking Placement and Routing in Rearrangeable Multi-Stage NetworksabstractABSTRACT High‐performance computing demands efficient and scalable interconnections. Although crossbar networks offer high parallel bandwidth, their costs are prohibitively expensive. Multi‐stage networks provide scalability with cost, yet they may block certain routing patterns. Rearrangeable multistage networks (RMNs) have emerged as a cost‐effective solution, enabling internal connection rearrangements to make all paths accessible without network blocking. However, discovering optimal rearrangement strategies remains a challenge for unblocking large‐scale (256 connections) reconfigurable networks. We advance the state of the art by effectively managing workloads with over 50% without blocking. When nearly all connections are required, we introduce routing strategies to rearrange existing connections. In the presence of multicast, where the configuration space exceeds , we propose novel strategies employing simulated annealing for placement and Monte Carlo tree search for routing to prioritize multicast connections, which simultaneously maximize the number of connections by minimizing conflicts and reducing the number of extra stages. To the best of our knowledge, we first demonstrate that the Benes network is not rearrangeable under multicast conditions. We propose exploring the rearrangeability of shuffle exchanges with additional network stages. Caio Von Rondow Morais, Jeronimo Penha, José A. M. Nacif, Ricardo S. Ferreira 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2025 | Reconfigurable Domain-Specific Architectures Based on Coarse-Grained Operators in High-Performance FPGAsabstractABSTRACT This work presents the development of reconfigurable processing units capable of encapsulating various operations to design new domain‐specific reconfigurable accelerators. These processing units are known as coarse‐grained operators because they can execute multiple operations. We validated the new operators within the HPCGRA framework, a design environment for creating custom coarse‐grained reconfigurable arrays that run as a virtual layer on commercial field‐programmable gate arrays. The environment is parameterized and employs an intermediate portable format that abstracts away low‐level specific bitstream details, enabling hardware‐agnostic reconfiguration. In this paper, we present three case studies. The first is a domain‐specific accelerator for the K‐means algorithm, which can be entirely reconfigured in less than 2.34 ms to explore various clustering values and attributes, achieving up to 159 Gop/s performance for considering 8 features. The reconfigurability of our accelerator has no impact on performance compared to static versions implemented directly in HLS and RTL. The second case study extends K‐means with a Gini calculation operator for dimensionality reduction, quantization, and quality classification, achieving up to 278 Gops/s. The third case study presents a systolic matrix multiplier, demonstrating the versatility of the environment for designing different architectural domains. Lucas B. da Silva, César Grandis, Jeronimo Penha, José A. M. Nacif, Ricardo S. Ferreira 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | Fast flow cloud: A stream dataflow framework for cloud FPGA accelerator overlays at runtimeabstractAbstract Cloud FPGAs provide new energy‐efficient opportunities to design dataflow accelerators. Nevertheless, FPGAs still have challenges to overcome for widespread usages, such as programmability, compilation time (minutes to hours), and hardware knowledge, mainly because it is highly challenging for beginners to learn and use FPGAs. The READY tool recently provides compilation time reduction to the range of microseconds using a CGRA overlay and a friendly, high‐level C++ interface for the Intel/Altera HARPv2 FPGA cloud platform. However, the HARPv2 is not available in any commercial cloud platform. This work extends READY by creating the fast flow cloud framework (FFC). First, FFC offers a simple browser‐based graphical interface for less experienced FPGA users. Second, we improve the CGRA overlay portability to include Xilinx FPGAs and a transparent design flow to deploy in the widespread commercial Amazon AWS F1 cloud. Third, we improve the CGRA reconfiguration engine. Also, we compare the overlay performance of HARPv2 and AWS F1 to an eight‐thread XEON processor. Finally, the framework is open‐source for collaborative development and has clearly defined application programming interfaces for future extensions. Lucas B. da Silva, Michael Canesche, Jeronimo Costa Penha, Josué Campos, José A. M. Nacif, Ricardo S. Ferreira 0001 |
Concurr. Comput. Pract. Exp. | 6 |
| 2023 | Heterogeneous reconfigurable architectures for machine learning dataflowsabstractAbstract This work explores the placement and routing of machine learning applications' dataflow graphs on different heterogeneous coarse‐grained reconfigurable architectures (CGRA). We analyze three different types of processing element (PE) heterogeneity, the first concerning the interconnection pattern, the second being on the kind of operations a single PE can execute, and the last concerning the PE buffer resources. This analysis aim to propose a fair reduction to the overall cost in comparison to the homogeneous CGRA architecture. We compare our results with the homogeneous case and one of the state‐of‐the‐art tools for placement and routing (P&R). Our algorithm executed, on average, 52% faster than VPR 8.1 (Versatile Place and Route), which is an open‐source academic tool designed for the FPGA placement and routing phases, reaching better mapping in 66% of cases and achieving the same results in 26% of cases. Furthermore, a heterogeneous architecture reduces the cost without losing performance in 76% of the cases considering multiplier heterogeneity. We propose a novel heterogeneous buffer architecture that minimizes the buffer resources by 56.3% for K‐means dataflow patterns. We also show that a heterogeneous border chess architecture outperforms a homogeneous one. In addition, our mapping reaches optimal instances of single tree dataflows compared to classical Lee/Choi and H‐trees. Westerley Carvalho, Michael Canesche, Lucas Reis, José A. M. Nacif, Ricardo S. Ferreira 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | High-performance graphics processing unit-based strategy for tuning a unmanned aerial vehicle controller subject to time-delay constraintsabstractAbstract Recently, high‐performance computing strategies have been implemented to improve performance analysis and reduce the development time of new solutions in robotic applications, such as path planning, machine learning, and vision, which require massive matrix computations. In this sense, this work aims to study the aerial robots' behavior during their mission execution. Due to the large search space in the set of parameter combinations and the high computational cost required to perform such an analysis after sequentially executing thousands of simulations, this work proposes an open‐source graphics processing unit (GPU)‐based implementation to simulate the robot behavior. A GPU‐accelerated flight route analysis for multi‐unmanned aerial vehicle (UAV) systems is proposed for the tuning control problem in the parameters‘ space considering the problem of delay in sending information to a ground control station. Considering our implementation, the experimental results show a speedup up to 325, 629, and 5959 in comparison to the parallel version with 16 threads, C coder converter, and native Matlab code, respectively. The implementation is available in the Colab Google platform and it can easily be expanded for analyses involving larger amounts of different parameters, robot models, strategies, and controllers. Leonardo Fagundes-Junior, Michael Canesche, Ricardo S. Ferreira 0001, Alexandre Santos Brandão |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Gene regulatory accelerators on cloud FPGAabstractSummary Gene regulatory networks (GRN) are dynamic models in time and space. These models are used to predict diseases and in drugs research. GRN models are discrete, and Boolean graphs can efficiently represent them. However, GRN algorithms explore a large solution space with high computational complexity. This work proposes efficient FPGA‐based accelerators to implement two GRN algorithms: attractor computation and Derrida plot. Nevertheless, FPGA accelerator design and deployment are still a challenge. This work presents an accelerator design framework for AWS Amazon FPGA cloud. The framework simplifies the software (SW) and hardware (HW) generation for GRN accelerators. The user provides a high‐level model for the Boolean GRN, and our tool automatically creates AWS‐ready‐to‐deploy software and hardware components. For the attractor and the Derrida plot computation, the proposed FPGA accelerators are on average and faster than a V100 GPU. Jeronimo Costa Penha, Lucas B. da Silva, Michael Canesche, Dener V. Ribeiro, José A. M. Nacif, Ricardo S. Ferreira 0001 |
Concurr. Comput. Pract. Exp. | 6 |
| 2022 | A polynomial time exact solution to the bit-aware register binding problemabstractFinding the minimum register bank is an optimization problem related to the synthesis of hardware. Given a program, the problem asks for the minimum number of registers plus their minimum size, in bits, that suffices to compile said program. This problem is NP-complete; hence, usually solved via heuristics. In this paper, we show that this problem has an optimal solution in polynomial time, as long as swaps can be inserted in the program to move variables across registers. This observation sets a lower bound to heuristics that minimize the size of register banks. We have compared the optimal algorithm with two classic heuristics. Our approach uses, on average, 6 to 10% less bits than that previous work. Michael Canesche, Ricardo S. Ferreira 0001, José A. M. Nacif, Fernando Magno Quintão Pereira |
CC | 2 |
| 2021 | Personalizing Online Computer Engineering Resources and Labs for Digital, Embedded, and Computer System CoursesabstractThe immediate and new challenges of the current Covid-19 pandemic have made it hard for all of us; in this work in progress as an innovative practice, we look to both leverage the challenges and share our work so that others might see some benefit to these times and improve their courses. In particular, our focus is on creating automated tools to physically create exams and sample code for Digital Systems and Computer Architecture courses. Additionally, we focus on shifting, traditional in-person labs to online, personalized formats for Digital Systems and Embedded Systems so that both educator and learner can still provide/experience virtual computer engineering education. We focus on three courses (Digital System Design, Computer Architecture/Organization, and Embedded System Design) as they are fundamentally driven by the implementation and execution of “algorithms”. From this starting point, we have created tools to generate sample code and exams, and have found means to virtualize labs and hands-on activities. In particular, we have created Python tools that allow educators to personalize code and problems, create these codes/problems (as text files or incorporated in word documents), and email these documents to students. This provides the means to create problems and code examples that are different from their peers and can be assessed on a per individual basis to alleviate some of the challenges with live and proctored exams. Additionally, we have found tools and methods for students to virtually perform the hands-on portion of these three subjects without the need for traditional lab equipment. This requires students to spend less than 100 USD worth of equipment and software. Our goal is to share these resources and our methodologies to help in this time of crisis. Additionally, these tools and methods have forced us to innovate our teaching, and we will, likely, use these tools and methods in the future. We share these tools in hope that the computer engineering education community will join this process to help us all improve our student's education. Peter Jamieson, Ricardo S. Ferreira 0001, José A. M. Nacif |
FIE | 2 |
| 2021 | Google Colab CAD4U: Hands-On Cloud Laboratories for Digital DesignabstractGoogle Colab is a cloud Jupyter notebook widespread used to teach machine learning by writing text explanations and Python codes through the browser. This work introduces new Colab extensions to teach logic circuit design, Verilog language, processor, and GPU architectures. Colab allows us to share reproducible experiments on the Web. The students become motivated to do laboratory assignments without download/configure software packages and dependencies on their computers. Furthermore, almost all universities had to shut down due to the COVID-19 pandemic, forcing us to adapt to virtual learning scenarios. Colab provides portability and accessibility since it can even run on smartphones. The lab assignments include intermediate guided exercises, text explanations, figures, online quizzes, problem sets, and basic hands-on tasks. We develop a simple setup for Icarus Verilog, PyEDA, CUDA, Valgrind, and Gem5 frameworks. This work presents Verilog teaching and computer architecture simulation insights by using Valgrind and Gem5, and GPU computer architecture profiling at the thread and instruction assembly level. Michael Canesche, Lucas B. da Silva, Omar P. Vilela Neto, José A. M. Nacif, Ricardo S. Ferreira 0001 |
ISCAS | 5 |
| 2021 | NMLib: A Nanomagnetic Logic Standard Cell LibraryabstractThe Nanomagnetic Logic (NML) is a promising new technology that can build low-power devices at room temperature. Furthermore, this technology allows mixing logic and memory on the same device. The creation of Electronic Design Automation (EDA) tools and flows is an essential step towards developing NML for integrated designs. There is plenty of room for creating new EDA methodologies for this kind of emerging nanotechnologies since the scarce number of works in this field. Standard cells is an important step in this context since they strongly relate to the routing and placement algorithms. This work presents NMLib, an NML cell library developed for the NMLSim 2.0 simulator. In contrast to CMOS, the NML features require logic cells and interconnection cells, since all circuit is developed using the same building block. Moreover, we present a full-adder, and a ripple carry adder circuit designs using NMLib to demonstrate NMLib feasibility. Laysson Oliveira Luz, José A. M. Nacif, Ricardo S. Ferreira 0001, Omar P. Vilela Neto |
ISCAS | 3 |
| 2021 | Is It Time to Include High-Level Synthesis Design in Digital System Education for Undergraduate Computer Engineers?abstractWe ask the question, “should High-level Synthesis (HLS) design be part of an undergraduate computer engineering education while learning digital system design?”. Current trends in industry include an increasing demand for engineers who can build FPGA systems. FPGAs, just like other chips, continue to improve in terms of complexity, speed, available resources, and new features. The design complexity for using an FPGA continues to grow and Hardware Description Languages (HDLs), though a step up in design efficiency compared to schematic design, is a low-level approach akin to assembly language for programmers, and HDLs limits the productivity of an engineer. HLS tools, such as Legup, Intel HLS, and Xilinx's Vivado attempt to provide designers with a higher-level design abstraction providing a means to describe their computation in high-level languages - such as C. As these tools become more mainstream in industry, when should education follow? In this work, we explore how HLS tools might be used by an undergraduate by looking at exemplar designs, a simple RISC-V processor and a basic C loop, and implementing the design in both HDL and HLS. We then analyze the FPGA cost of each implementation. Next, we provide a philosophical discussion based on this experience on what the pros and cons of moving students to HLS design abstraction level are. Isaac Nelson, Ricardo S. Ferreira 0001, José A. M. Nacif, Peter Jamieson |
ISCAS | 2 |
| 2021 | RESHAPE: A Run-Time Dataflow Hardware-Based Mapping for CGRA OverlaysabstractCoarse-grained reconfigurable architectures (CGRA) are a power-efficient approach for hardware accelerators. However, there are few EDA tools for CGRA. We develop hardware-based placement and routing (P&R) for fully-pipelined CGRA mapped as an FPGA overlay. The key idea is to use the available FPGA resources to replicate several mapping units, thus exploring parallel execution, area/execution time trade-offs, and achieving near-optimal mapping solutions. Furthermore, our P&R provides portability and an incremental run-time approach. In comparison to VPR and CGRA-ME tools and a time-multiplexer approach, our spatial mapping reduces the P&R execution time, and it improves the performance up to hundreds of Gops/s by using fully-pipelined architectures. Maria D. Vieira, Michael Canesche, Lucas B. da Silva, Josué Campos, Mateus Silva, Ricardo S. Ferreira 0001, José A. M. Nacif |
ISCAS | 6 |
| 2021 | TRAVERSAL: A Fast and Adaptive Graph-Based Placement and Routing for CGRAsabstractCoarse grain reconfigurable architectures (CGRAs) are an emerging hybrid computational architecture that has the parallel customization benefits of low-level logic devices, such as FPGAs and ASICs, while the relative coarseness of these architectures makes CGRAs easier to design for, which is more similar to the traditional processor. In the process of mapping designs to CGRAs, flexible, fast, and adaptive placement and routing (P&R) is fundamental in order to implement efficient run-time reconfigurable frameworks. It is well-known that P&R is an NP-complete problem, and thus, solutions rely on heuristics to achieve quality results with acceptable execution times. CGRA P&R has different constraints compared to traditional VLSI P&R, e.g., path latency balancing and modulo scheduling of loops. In this work, we propose a graph-based P&R approach that uses graph traversals to map designs to CGRAs. Additionally, we parallelize our approach with a graph-based greedy heuristic that executes on a GPU. We compare our proposed P&R approach with the CGRA-ME framework, which implements simulated annealing and integer linear programming placement algorithms. Our results show that this new approach can generate optimal mappings and improve the execution run-time up to several orders of magnitude. Furthermore, considering spatial mapping at the millisecond scale, our GPU approach is one order of magnitude faster compared to the state-of-the-art tool VPR. Michael Canesche, Marcelo de Matos Menezes, Westerley Carvalho, Frank Sill, Peter Jamieson, José A. M. Nacif, Ricardo S. Ferreira 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2021 | You Only Traverse Twice: A YOTT Placement, Routing, and Timing Approach for CGRAsabstractCoarse-grained reconfigurable architecture (CGRA) mapping involves three main steps: placement, routing, and timing. The mapping is an NP-complete problem, and a common strategy is to decouple this process into its independent steps. This work focuses on the placement step, and its aim is to propose a technique that is both reasonably fast and leads to high-performance solutions. Furthermore, a near-optimal placement simplifies the following routing and timing steps. Exact solutions cannot find placements in a reasonable execution time as input designs increase in size. Heuristic solutions include meta-heuristics, such as Simulated Annealing (SA) and fast and straightforward greedy heuristics based on graph traversal. However, as these approaches are probabilistic and have a large design space, it is not easy to provide both run-time efficiency and good solution quality. We propose a graph traversal heuristic that provides the best of both: high-quality placements similar to SA and the execution time of graph traversal approaches. Our placement introduces novel ideas based on “you only traverse twice” (YOTT) approach that performs a two-step graph traversal. The first traversal generates annotated data to guide the second step, which greedily performs the placement, node per node, aided by the annotated data and target architecture constraints. We introduce three new concepts to implement this technique: I/O and reconvergence annotation, degree matching, and look-ahead placement. Our analysis of this approach explores the placement execution time/quality trade-offs. We point out insights on how to analyze graph properties during dataflow mapping. Our results show that YOTT is 60.6 , 9.7 , and 2.3 faster than a high-quality SA, bounding box SA VPR, and multi-single traversal placements, respectively. Furthermore, YOTT reduces the average wire length and the maximal FIFO size (additional timing requirement on CGRAs) to avoid delay mismatches in fully pipelined architectures. Michael Canesche, Westerley Carvalho, Lucas Reis, Matheus Aguilar de Oliveira, Salles V. G. Magalhães, Peter Jamieson, José A. M. Nacif, Ricardo S. Ferreira 0001 |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2020 | Mind the Gap: Bridging Verilog and Computer ArchitectureabstractWe present an approach to teach RISC processor design for an undergraduate computer architecture course specifically aimed to reduce the gap between a high-level datapath block diagram and a complete Verilog code specification. We propose a graphical approach designed to develop an understanding of the MIPS processor organization at the Verilog structural level by using an online browser-based simulator, from a single cycle design to pipeline design. The students are lead through a series of examples, step by step, and they can actively be involved in the processor design process. We believe that the best choice should not introduce excessive complexity that becomes a barrier in describing the interconnection of a high-level diagram and the Verilog implementation code. Fernando Passe, Michael Canesche, Omar P. Vilela Neto, José A. M. Nacif, Ricardo S. Ferreira 0001 |
ISCAS | 5 |
| 2019 | ADD: Accelerator Design and Deploy - A tool for FPGA high-performance dataflow computingabstractSummary Dataflow‐based FPGA accelerators have become a promising alternative to deliver energy‐efficient high‐performance computing. However, FPGA programming is still a challenge. This paper presents Accelerator Design and Deploy (ADD), a high‐level framework to specify, to simulate, and to implement dataflow accelerators for streaming applications. The framework includes an open dataflow operator library, and templates are provided to easily design new operators. The framework also provides a high‐level and an accurate simulation at circuit level with short execution times. Moreover, ADD provides software and hardware APIs to simplify the integration process, extending the benefits of portability from low‐cost FPGA boards to high performance datacenter FPGA platforms. Our framework supports coupling with high‐level programming languages, and it has been validated on two FPGA platforms: the Intel high‐performance CPU‐FPGA heterogeneous computing platform and an educational FPGA kit. We show that our simple approach presents competitive performance, both in time and energy, when compared to multi‐core and GPU accelerators. Jeronimo Costa Penha, Lucas B. da Silva, Jansen Silva, Kristtopher Coelho, Hector P. Baranda, José A. M. Nacif, Ricardo S. Ferreira 0001 |
Concurr. Comput. Pract. Exp. | 7 |
| 2019 | READY: A Fine-Grained Multithreading Overlay Framework for Modern CPU-FPGA Dataflow ApplicationsabstractIn this work, we propose a framework called REconfigurable Accelerator DeploY (READY), the first framework to support polynomial runtime mapping of dataflow applications in high-performance CPU-FPGA platforms. READY introduces an efficient mapping with fine-grained multithreading onto an overlay architecture that hides the latency of a global interconnection network. In addition to our overlay architecture, we show how this system helps solve some of the challenges for FPGA cloud computing adoption in high-performance computing. The framework encapsulates dataflow descriptions by using a target independent, high-level API, and a dataflow model that allows for explicit spatial and temporal parallelism. READY directly maps the dataflow kernels onto the accelerator. Our tool is flexible and extensible and provides the infrastructure to explore different accelerator designs. We validate READY on the Intel Harp platform, and our experimental results show an average 2x execution runtime improvement when compared to an 8-thread multi-core processor. Lucas B. da Silva, Ricardo S. Ferreira 0001, Michael Canesche, Marcelo de Matos Menezes, Maria D. Vieira, Jeronimo Costa Penha, Peter Jamieson, José A. M. Nacif |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Exploration of the Synchronization Constraint in Quantum-dot Cellular AutomataabstractQuantum-dot Cellular Automata (QCA) is a field-coupled nanotechnology which might enable design with high performance and extraordinary low energy dissipation. Infor-mation processing and flow in QCA is controlled by external clocks, which requires a proper synchronization already during circuit design phase. In this paper, we discuss the fundamental differences between local and global synchronicity in QCA circuits. Further, we show that it is possible to relax the global synchronicity constraint and discuss the consequent impact on the design performance. Simulation results indicate that the design size can be reduced by about 70% while the throughput performance declines by similar values. Frank Sill, Pedro Arthur Silva, Geraldo Fontes, José A. M. Nacif, Ricardo S. Ferreira 0001, Omar P. Vilela Neto, Jeferson F. Chaves, Rolf Drechsler |
DSD | 5 |
| 2018 | Fast analysis of upstream features on spatial networks (GIS cup)abstractWe present a fast linear time algorithm that uses a block-cut tree for identifying upstream features from a set of starting points in a network. Our implementation has been parallelized and it can process a dataset with 32 million features in less than 8 seconds on a 8-core workstation. This problem is the 2018 ACM SIGSPATIAL CUP challenge and presents several applications mainly on the field of utility networks. Our code is freely available for nonprofit research and education at https://github.com/sallesviana/FastUpstream Salles V. G. Magalhães, W. Randolph Franklin, Ricardo S. Ferreira 0001 |
SIGSPATIAL/GIS | 3 |
| 2018 | Placement and Routing by Overlapping and Merging QCA GatesabstractThe QCA gate-level design introduces new challenges to the traditional mapping, placement, and routing flow. First, the wires consume more than 90% of total area. In addition, even a combinational circuit requires a clock scheme, and all internal paths should be balanced. This work proposes a novel approach to automatically map a gate-level circuit onto a QCA layout by using a merge overlapping approach, and a universal clock scheme to provide scalability. First, the circuit is decomposed into a set of overlapping subgraph partitions. This decomposition is guided by reconvergent paths. For each subgraph, more than one customized QCA layout cell is generated on-the-fly. During the last step, a merge overlapping algorithm rebuilds the entire circuit to produce the final layout. All inter and intra-partition wires should be balanced according to the adopted clock scheme. Our approach reduces the total area in more than 50% in comparison to a previous approach based on standard-cell QCA libraries. Finally, all layouts were validated on QCA Designer simulator to verify the design rules. Geraldo Fontes, Pedro Arthur Silva, José A. M. Nacif, Omar P. Vilela Neto, Ricardo S. Ferreira 0001 |
ISCAS | 5 |
| 2018 | A Novel Five-input Multiple-function QCA Threshold GateabstractQCA (Quantum-dot Cellular Automata) is a promising new technology with low power consumption and high speed that allows the design of nanoscale integrated circuits. The 3-input/1-output majority gate is the basic building block in QCA circuits. This work presents a new design of a multi-output, 5-input majority gate. Our proposed gate is quite useful because its outputs can present different configurations with several logical functions at once, enabling the design of smaller circuits. Also, the proposed gate is a feasible solution for the recently proposed USE (Universal, Scalar and Efficient) clocking scheme. In order to demonstrate the flexibility and area efficiency of our 5-input majority gates we implement two designs: a full adder and a RAM cell block. These designs have been implemented using a free and a regular (USE) clock schemes. Our results show area reductions up to 50% compared to state-of-the-art designs. Pedro Arthur Silva, Juliana Rezende S. B. Alves, Ricardo S. Ferreira 0001, Omar P. Vilela Neto, José A. M. Nacif |
ISCAS | 3 |
| 2018 | From Java to FPGA: An Experience with the Intel HARP SystemabstractRecent years have seen a surge in the popularity of Field-Programmable Gate Arrays (FPGAs). Programmers can use them to develop high-performance systems that are not only efficient in time, but also in energy. Yet, programming FPGAs remains a difficult task. Even though there exist today OpenCL interfaces to synthesize such hardware, higher-level programming languages, such as Java, C# or Python remain distant from them. In this paper, we describe a compiler, and its supporting runtime environment, that reduces this distance, translating functional code written in Java to the Intel HARP platform. Thus, we bring two contributions. First, the insight that a functional-style library is a good starting point to bridge the gap between high-level programming idioms and FPGAs. Second, the implementation of this system itself, including the compiler, its intermediate representation, and all the runtime support necessary to shield developers from the task of transferring data back and forth between the host CPU and the accelerator. To demonstrate the effectiveness of our system, we have used it to implement different benchmarks, used in image processing and data-mining. For large inputs, we can observe consistent 20x speedups over the Java Virtual Machine across all our benchmarks. Depending on the target function that we compile, this speedup can achieve 280x. Pedro Caldeira, Jeronimo Costa Penha, Lucas B. da Silva, Ricardo S. Ferreira 0001, José A. M. Nacif, Renato Ferreira 0001, Fernando Magno Quintão Pereira |
SBAC-PAD | 4 |
| 2015 | Be a simulator developer and go beyond in computing engineeringabstractThis work presents a methodology based on simulator developers to be applied in several courses as well as research activities. This proposal is also based on “The Cathedral and the Bazaar” lessons. First, good programmers know what to write, however great ones know what to rewrite. Students could learn by developing new components to be added to the current simulator. Second, any tool should be useful in the expected way, but a truly great tool lends itself to uses you never expected. The proposed simulator has been developed to teach digital circuits. However, it can be used in many other computing engineering fields. Third point is motivation, where the lessons are driven by the following principle: To solve an interesting problem, start by finding a problem that is interesting to you. Nowadays, we are surrounded by Internet of things, where the simulator engine could be connected to learn and to motivate the students. Image processing could be taught by connecting your web-cam and/or your mobile device camera to your simulator engine. Moreover, mobile phones and/or Arduino devices have a set of sensors that could be captured inside of the virtual simulator lab. Ricardo S. Ferreira 0001, José A. M. Nacif, Salles V. G. Magalhães, Thales T. de Almeida, Racyus D. G. Pacífico |
FIE | 1 |
| 2015 | A Runtime FPGA Placement and Routing Using Low-Complexity Graph TraversalabstractDynamic Partial Reconfiguration (DPaR) enables efficient allocation of logic resources by adding new functionalities or by sharing and/or multiplexing resources over time. Placement and routing (P&R) is one of the most time-consuming steps in the DPaR flow. P&R are two independent NP-complete problems, and, even for medium size circuits, traditional P&R algorithms are not capable of placing and routing hardware modules at runtime. We propose a novel runtime P&R algorithm for Field-Programmable Gate Array (FPGA)-based designs. Our algorithm models the FPGA as an implicit graph with a direct correspondence to the target FPGA. The P&R is performed as a graph mapping problem by exploring the node locality during a depth-first traversal. We perform the P&R using a greedy heuristic that executes in polynomial time. Unlike state-of-the-art algorithms, our approach does not try similar solutions, thus allowing the P&R to execute in milliseconds. Our algorithm is also suitable for P&R in fragmented regions. We generate results for a manufacturer-independent virtual FPGA. Compared with the most popular P&R tool running the same benchmark suite, our algorithm is up to three orders of magnitude faster. Ricardo S. Ferreira 0001, Luciana Rocha, José A. M. Nacif, Stephan Wong, Luigi Carro |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2013 | A run-time graph-based Polynomial Placement and routing algorithm for virtual FPGASabstractDynamic partial reconfiguration enables efficient use of hardware resources by multiplexing system functionality in time. However, many challenges arise from partial reconfiguration implementation. The placement and routing (P&R) of the hardware modules is a computationally intensive task, and the state-of-art algorithms are not suitable to place and route modules at run-time. This paper makes several contributions: (1) Single Placement at run-time: we propose a novel P&R algorithm based on greedy heuristic where a single placement is performed at run-time in few milliseconds. (2) Implicit Graph Model: the FPGA is modelled as an implicit graph with a direct correspondence to the physical FPGA, and the P&R is performed as a graph mapping problem by exploring the node locality during the depth-first traversal. (3) Polynomial Placement: we show that even a single placement can be routed without critical path degradation. (4) Fragmented Regions: the graph approach is flexible, and it allows efficient placement even onto fragmented FPGA areas. Compared with the most popular P&R tool running the same benchmark suite our algorithm is on average 864x faster. Moreover, the bitstream for partial reconfiguration is also reduced by a factor of 4. Ricardo S. Ferreira 0001, Luciana Rocha, José A. M. Nacif, Stephan Wong, Luigi Carro |
FPL | 1 |
| 2011 | An FPGA-based heterogeneous coarse-grained dynamically reconfigurable architectureabstractCoarse-grained reconfigurable architecture has emerged as a promising model for embedded systems as a solution to reduce the complexity of FPGA synthesis and mapping steps, consequently reducing reconfiguration time. Despite these advantages, CGRA usage has been limited due to the lack of commercial CGRA circuits. This work proposes a virtual and dynamic CGRA implemented on top of an FPGA. This approach allows the usage of commercial-off-the-shelf FPGA devices combined with the advantages of CGRAs. The proposed architecture consists of a set of heterogeneous functional units (FU) and a global interconnection network. The global network allows any FU to be used at each cycle, which reduces significantly the placement complexity. In addition, we introduce a polynomial mapping algorithm which includes scheduling, placement and routing steps (SPR). Moreover, the proposed approach performs a very fast placement and routing in comparison to similar CGRA approaches. The three SPR steps are computed in few milliseconds. The feasibility of this approach is demonstrated for a suite of digital signal processing benchmarks. Ricardo S. Ferreira 0001, Julio C. Goldner Vendramini, Lucas Mucida, Monica Magalhães Pereira, Luigi Carro |
CASES | 1 |
| 2011 | Fast placement and routing by extending coarse-grained reconfigurable arrays with Omega Networks
Ricardo S. Ferreira 0001, João M. P. Cardoso, Alex Damiany, Julio C. Goldner Vendramini, Tiago Teixeira |
J. Syst. Archit. | 1 |
| 2010 | FPGA-accelerated Attractor Computation of Scale Free Gene Regulatory NetworksabstractDue to the large amount of experimental data provided by DNA microarrays, different kinds of computational methods have been proposed to study gene interactions. These methods are computationally very intensive. More specifically, this work focuses on an efficient approach to compute the attractor cycles in gene regulatory networks, which are modeled by Boolean graphs. This work proposes to explore the inherent parallelism of this approach by using a FPGA based implementation. In addition, we also propose a runtime framework to dynamically insert/delete edges at low cost by using a multistage interconnection network (MIN). We also show that MIN could be efficiently mapped on FPGAs by taking advantages of embedded memory modules. The MIN implementation is close to theoretical complexity O(n lg n). Moreover, we propose a heterogeneous node set, where few nodes are strongly connected and most nodes are poorly connected, which allow us to handle efficiently scale free topologies. Furthermore, our approach is synthesized once onto a FPGA and several simulations are performed, without the need of re-synthesis. Experimental results have shown acceleration gains up to three orders of magnitude compared to sequential approaches. Ricardo S. Ferreira 0001, Julio C. Goldner Vendramini |
FPL | 1 |
| 2009 | A low cost and adaptable routing network for reconfigurable systemsabstractNowadays, scalability, parallelism and fault-tolerance are key features to take advantage of last silicon technology advances, and that is why reconfigurable architectures are in the spotlight. However, one of the major problems in designing reconfigurable and parallel processing elements concerns the design of a cost-effective interconnection network. This way, considering that Multistage Interconnection Network (MIN) has been successfully used in several computer system levels and applications in the past, in this work we propose the use of a MIN, at the word level, on a coarse-grained reconfigurable architecture. More precisely, this work presents a novel parallel self-placement and routing mechanism for MIN on the circuit-switching mode. We take into account one-to-one as well as multicast (one-to-many) permutations. Our approach is scalable and it is targeted to be used in run-time environments where dynamic routing among functional units is required. In addition, our algorithm is embedded in the switch structure, and it is independent of the interstage interconnection pattern. Our approach can handle blocking and non-blocking networks, symmetrical or asymmetrical topologies. As case study, we use the proposed technique in a dynamic reconfigurable system, showing a major area reduction of 30% without performance overhead. Ricardo S. Ferreira 0001, Marcone Laure, Antonio Carlos Schneider Beck, Thiago Lo, Mateus B. Rutzig, Luigi Carro |
IPDPS | 1 |
| 2008 | Reducing interconnection cost in coarse-grained dynamic computing through multistage networkabstractCoarse-grained reconfigurable architectures appear as a scalable solution to embedded system design, with a reduced reconfiguration time, memory footprint, as well as placement and routing complexity. To ensure high performance, data must be efficiently delivered to the reconfigurable matrix. For that, several architectures propose the use of fully interconnected local networks, as crossbar or large multiplexers. However, these interconnections are very area consuming. Therefore, in order to reduce the interconnection complexity without losing performance, this work proposes to use Multistage Interconnection Networks. As a case study, we have implemented the proposed approach in a tightly coupled reconfigurable array, which works together with a MIPS processor. Simulation results over the Mibench Benchmark set show savings of up to 26% of the total area, with a decrease of only 1% on the average performance. Ricardo S. Ferreira 0001, Marcone Laure, Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
FPL | 1 |
| 2004 | An Environment for Exploring Data-Driven Architectures
Ricardo S. Ferreira 0001, João M. P. Cardoso, Horácio C. Neto |
FPL | 1 |