EDBT 2026 Demo / reviewers in the wild / expert
Westerley Carvalho
dblp:292/1821
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2023
0000-0002-4030-4098ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Electronic design automation · 61% Reconfigurable computing and FPGAs · 30% GPUs and heterogeneous computing · 9% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture |
0.5 | 1 | 2021 | TRAVERSAL: A Fast and Adaptive Graph-Based Placement and Routing for CGRAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021 |
Electronic design automation
physical design |
0.5 | 1 | 2021 | TRAVERSAL: A Fast and Adaptive Graph-Based Placement and Routing for CGRAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021 |
Electronic design automation › physical design
placement and routing |
0.5 | 1 | 2021 | TRAVERSAL: A Fast and Adaptive Graph-Based Placement and Routing for CGRAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2021 | TRAVERSAL: A Fast and Adaptive Graph-Based Placement and Routing for CGRAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021 |
Methods — techniques the papers use, named apart from their topics
simulated annealing · 0.5integer linear programming · 0.5greedy heuristic · 0.5graph traversal · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Heterogeneous reconfigurable architectures for machine learning dataflowsabstractAbstract This work explores the placement and routing of machine learning applications' dataflow graphs on different heterogeneous coarse‐grained reconfigurable architectures (CGRA). We analyze three different types of processing element (PE) heterogeneity, the first concerning the interconnection pattern, the second being on the kind of operations a single PE can execute, and the last concerning the PE buffer resources. This analysis aim to propose a fair reduction to the overall cost in comparison to the homogeneous CGRA architecture. We compare our results with the homogeneous case and one of the state‐of‐the‐art tools for placement and routing (P&R). Our algorithm executed, on average, 52% faster than VPR 8.1 (Versatile Place and Route), which is an open‐source academic tool designed for the FPGA placement and routing phases, reaching better mapping in 66% of cases and achieving the same results in 26% of cases. Furthermore, a heterogeneous architecture reduces the cost without losing performance in 76% of the cases considering multiplier heterogeneity. We propose a novel heterogeneous buffer architecture that minimizes the buffer resources by 56.3% for K‐means dataflow patterns. We also show that a heterogeneous border chess architecture outperforms a homogeneous one. In addition, our mapping reaches optimal instances of single tree dataflows compared to classical Lee/Choi and H‐trees. Westerley Carvalho, Michael Canesche, Lucas Reis, José A. M. Nacif, Ricardo S. Ferreira 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | TRAVERSAL: A Fast and Adaptive Graph-Based Placement and Routing for CGRAsabstractCoarse grain reconfigurable architectures (CGRAs) are an emerging hybrid computational architecture that has the parallel customization benefits of low-level logic devices, such as FPGAs and ASICs, while the relative coarseness of these architectures makes CGRAs easier to design for, which is more similar to the traditional processor. In the process of mapping designs to CGRAs, flexible, fast, and adaptive placement and routing (P&R) is fundamental in order to implement efficient run-time reconfigurable frameworks. It is well-known that P&R is an NP-complete problem, and thus, solutions rely on heuristics to achieve quality results with acceptable execution times. CGRA P&R has different constraints compared to traditional VLSI P&R, e.g., path latency balancing and modulo scheduling of loops. In this work, we propose a graph-based P&R approach that uses graph traversals to map designs to CGRAs. Additionally, we parallelize our approach with a graph-based greedy heuristic that executes on a GPU. We compare our proposed P&R approach with the CGRA-ME framework, which implements simulated annealing and integer linear programming placement algorithms. Our results show that this new approach can generate optimal mappings and improve the execution run-time up to several orders of magnitude. Furthermore, considering spatial mapping at the millisecond scale, our GPU approach is one order of magnitude faster compared to the state-of-the-art tool VPR. Michael Canesche, Marcelo de Matos Menezes, Westerley Carvalho, Frank Sill, Peter Jamieson, José A. M. Nacif, Ricardo S. Ferreira 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | You Only Traverse Twice: A YOTT Placement, Routing, and Timing Approach for CGRAsabstractCoarse-grained reconfigurable architecture (CGRA) mapping involves three main steps: placement, routing, and timing. The mapping is an NP-complete problem, and a common strategy is to decouple this process into its independent steps. This work focuses on the placement step, and its aim is to propose a technique that is both reasonably fast and leads to high-performance solutions. Furthermore, a near-optimal placement simplifies the following routing and timing steps. Exact solutions cannot find placements in a reasonable execution time as input designs increase in size. Heuristic solutions include meta-heuristics, such as Simulated Annealing (SA) and fast and straightforward greedy heuristics based on graph traversal. However, as these approaches are probabilistic and have a large design space, it is not easy to provide both run-time efficiency and good solution quality. We propose a graph traversal heuristic that provides the best of both: high-quality placements similar to SA and the execution time of graph traversal approaches. Our placement introduces novel ideas based on “you only traverse twice” (YOTT) approach that performs a two-step graph traversal. The first traversal generates annotated data to guide the second step, which greedily performs the placement, node per node, aided by the annotated data and target architecture constraints. We introduce three new concepts to implement this technique: I/O and reconvergence annotation, degree matching, and look-ahead placement. Our analysis of this approach explores the placement execution time/quality trade-offs. We point out insights on how to analyze graph properties during dataflow mapping. Our results show that YOTT is 60.6 , 9.7 , and 2.3 faster than a high-quality SA, bounding box SA VPR, and multi-single traversal placements, respectively. Furthermore, YOTT reduces the average wire length and the maximal FIFO size (additional timing requirement on CGRAs) to avoid delay mismatches in fully pipelined architectures. Michael Canesche, Westerley Carvalho, Lucas Reis, Matheus Aguilar de Oliveira, Salles V. G. Magalhães, Peter Jamieson, José A. M. Nacif, Ricardo S. Ferreira 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |