EDBT 2026 Demo / reviewers in the wild / expert
Vanderlei Bonato
dblp:57/3827
· DBLP profile ↗
15ranked-venue papers
5as first author
2since 2021 · last 2026
0000-0002-1743-8004ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Electronic design automation · 77% Reconfigurable computing and FPGAs · 16% Performance modeling and evaluation · 6% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
high-level synthesis |
0.9 | 2 | 2021 | Fast Resource and Timing Aware Design Optimisation for High-Level Synthesis · IEEE Trans. Computers 2021 Scaling Up Modulo Scheduling for High-Level Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Electronic design automation
design space exploration |
0.5 | 1 | 2021 | Fast Resource and Timing Aware Design Optimisation for High-Level Synthesis · IEEE Trans. Computers 2021 |
Electronic design automation › high-level synthesis › pipeline synthesis
loop pipelining |
0.4 | 1 | 2019 | Scaling Up Modulo Scheduling for High-Level Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Reconfigurable computing and FPGAs
modulo scheduling |
0.4 | 1 | 2019 | Scaling Up Modulo Scheduling for High-Level Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Electronic design automation
hardware verification and test |
0.1 | 1 | 2021 | Fast Resource and Timing Aware Design Optimisation for High-Level Synthesis · IEEE Trans. Computers 2021 |
Performance modeling and evaluation
performance prediction |
0.1 | 1 | 2021 | Fast Resource and Timing Aware Design Optimisation for High-Level Synthesis · IEEE Trans. Computers 2021 |
Methods — techniques the papers use, named apart from their topics
profiling-based estimation · 0.5lina estimator · 0.5modulo scheduling · 0.4integer linear programming · 0.4gaussian filter · 0.0edge detection · 0.0absolute difference mask · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi -FPGA streaming using OpenMPabstractThe growth in the demand for high-performance and power-efficient applications has led to an increasing interest in FPGA-based acceleration. FPGAs have been applied to a wide range of applications. Still, programming them can be a complex task, requiring extensive knowledge of tools and libraries, especially in multi-FPGA architectures. As such, the desire for tools and frameworks to ease the burden and abstract the knowledge of using FPGAs has increased. OpenMP, an already dominant parallel programming model in HPC, has been shown to be a successful approach to program multi-FPGA architecture. This work is based on the OMPC-F framework, which leverages the capability of OpenMP to offload computation to FPGAs. Although OMPC-F abstracts FPGA handling and task distribution from the final user, it does not support streaming computation based on multi-FPGA architectures. Streaming is widely used in FPGA designs to create a pipeline of computation between FPGA kernels. This work uses FPGA kernel binary information to synthesize streams as OpenMP buffers, while adapting the OpenMP dependency system accordingly. The proposal was evaluated in an AMD/Xilinx multi-FPGA system and shows speedups of the order of 6.84x in 8 FPGAs, scaling well with the addition of more kernels and FPGAs to the architecture. Moreover, compared to the regular approach to developing FPGA applications based on MPI+XRT communication, the proposed approach reduces the programming effort by 59% according to various code analysis metrics, resulting in a small average overhead of 4.3% when compared to the MPI+XRT programming model. Pedro Henrique Di Francia Rosso, Rémy Neveu, Nusrat Jahan Lisa, Lucas B. da Silva, Hervé Yviquel, Sandro Rigo, Vanderlei Bonato, Guido Araujo |
J. Parallel Distributed Comput. | 7 |
| 2021 | Fast Resource and Timing Aware Design Optimisation for High-Level SynthesisabstractField-Programmable Gate Arrays (FPGA) are often present in energy-efficient systems, although its non-trivial development flow is an obstacle for massive adoption. High-Level Synthesis (HLS) approaches attempt to mitigate the gap by targetting FPGAs from software languages, however manual tuning is still essential to meet performance demands. We present a high-level design space exploration framework with timing and resource awareness that uses an estimator named Lina to evaluate each design point. Lina is a profiling-based approach that avoids the costly static analyses performed by HLS compilers, allowing a significantly faster exploration of optimisations. Estimations are improved by supporting a continuous range of operating frequencies and by considering resource usage for both floating-point and integer datapaths. For a given set of C kernels, the estimated solutions are among the best 1% for execution time and resource footprint. The exploration of each kernel using Lina was performed on average two orders of magnitude faster than using early HLS compiler reports, and four orders of magnitude faster than fully compiling each design point. By considering the design spaces traversed, our solutions reached 70% of the maximum speed-up achievable. This represents an average speed-up of 14-16× compared to the baseline designs with no optimisations enabled. André Bannwart Perina, Arthur Silitonga, Jürgen Becker 0001, Vanderlei Bonato |
IEEE Trans. Computers | 4 |
| 2019 | Scaling Up Modulo Scheduling for High-Level SynthesisabstractHigh-Level Synthesis tools have been increasingly used within the hardware design community to bridge the gap between productivity and the need to design large and complex systems. When targeting heterogeneous systems, where the CPU and the FPGA fabric are both available to perform computations, a design space exploration is usually carried out for deciding which parts of the initial code should be mapped to the FPGA fabric such as the overall system’s performance is enhanced by accelerating its computation via dedicated processors. As the targeted systems become more complex and larger, leading to a large design space exploration, the fast estimative of the possible acceleration that can be obtained by mapping certain functionality into the FPGA fabric is of paramount importance. Loop pipelining, which is responsible for the majority of HLS compilation time, is a key optimization towards achieving high-performance acceleration kernels. A new modulo scheduling algorithm is proposed, which reformulates the classical modulo scheduling problem and leads to a reduced number of integer linear problems solved, resulting in large computational savings. Moreover, the proposed approach has a controlled trade-off between solution quality and computation time. Results show the scalability is improved efficiently from quadratic, for the state-of-the-art method, to linear, for the proposed approach, while the optimized loop suffers a 1% (geomean) increment in the total number of cycles. Leandro de Souza Rosa, Christos-Savvas Bouganis, Vanderlei Bonato |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Mapping Estimator for OpenCL Heterogeneous AcceleratorsabstractTo increase computing performance while keeping energy consumption to an acceptable budget, heterogeneous systems are currently investigated. By using dedicated compute units as accelerators to speedup specific parts of an application, hardware resources are better utilised resulting in a more energy efficient computing system. However, the task of performing such application mapping to accelerators is still a challenge, requiring knowledge beyond software domain in order to understand which part of the code fits better to the capability of the hardware available. Currently, there are tools supporting unified frontends and languages to simplify the programming of such heterogeneous systems, however there is still a high dependency of the user to manually perform the final mapping process. This work exposes a machine learning framework used to automatically infer the most suitable accelerator (between FPGA and GPU) for a given code by statically estimating energy efficiency. This framework can be used to assist the developer in deciding the best mapping for its application with an average hit-rate of 85 percent. André Bannwart Perina, Vanderlei Bonato |
FPT | 2 |
| 2018 | Scaling Up Loop Pipelining for High-Level Synthesis: A Non-iterative ApproachabstractHigh-level synthesis is a powerful tool for increasing productivity in digital hardware design. However, as digital systems become larger and more complex, designers have to consider an increased number of optimizations and directives offered by high-level synthesis tools to control the hardware generation process, resulting in a large design space to be explored. One of the most impactful optimizations is loop pipelining due to its large improvement in the hardware throughput. Nevertheless, the modulo scheduling algorithms that are used for loop pipelining are computationally expensive, and their application to the whole design space can make its exploration inviable, leading to sub-optimum solutions. Current state-of-the-art tools for modulo scheduling follow an iterative approach, which solves O(n2) optimization problems, where n is the loop code size. To address this problem, this work proposes a novel data-flow-based approach that solves exactly 2 optimization problems, independently of the loop code size. Results show orders-of-magnitude savings in the computation time, leading to significant design space exploration time savings when compared with the state-of-the-art. As such, the proposed method produces hardware designs of higher performance than the ones produced by the current state of the art for large and complex loops, maintaining a similar resource utilization. Leandro de Souza Rosa, Vanderlei Bonato, Christos-Savvas Bouganis |
FPT | 2 |
| 2017 | Exploiting Kant and Kimura's Matrix Inversion Algorithm on FPGAabstractMatrix inversion for real-time applications can be a challenge for the designers since its computational complexity is typically cubic. Parallelism has been widely exploited to reduce such complexity, however most traditional methods do not scale well with the matrix size leading to communication bottlenecks. In this paper we exploit a decentralised parallel hardware architecture based on a strongly non-singular matrix inversion algorithm proposed by Kant and Kimura in 1978, which is a parallel-orientated method with communication mode independent of the matrix size, mitigating the problem of matrix scalability. The hardware architecture is implemented in two different approaches using fixed-point arithmetic: dedicated and shared. In the first approach a matrix can be inverted in linear time while the latter, for the best case, has a square complexity. Experimental results are demonstrated using a Stratix V GX FPGA. For instance, in dedicated approach an 8x8 matrix is inverted in 1.27us, while in shared approach a 64x64 matrix is inverted in 153.40us using 64 pipelined processing elements. André Bannwart Perina, Paulo Matias, Eduardo Marques, Vanderlei Bonato, João Miguel Gago Pontes de Brito Lima |
DSD | 4 |
| 2015 | Parameterizable Ethernet Network-on-Chip Architecture on FPGAabstractWith the number of cores increase in systems-on-chip (SoC), bus-based approach began facing challenges to support internal communication. An alternative that has been explored is the network-on-chip (NoC), an approach that proposes to use common network knowledge on SoC projects internal communication. The standards non-adoption in the NoC components development however has delayed its wide diffusion. This paper focuses on providing a complete NoC architecture, configurable and customizable following the Ethernet standard. The three NoC basic modules, Network Adapter (NA), Link and Switch, are implemented. The results were obtained using a Stratix IV FPGA. The evaluation metrics used for NoC validation are silicon area and latency. The experiment using two NAs, two cores and one Switch needed 7310 FPGA ALUTs which corresponds to 4% of their logical resources. The Ethernet frame (64 Bytes) transmission spent 422 clock cycles on FPGA. Helio Fernandes da Cunha Junior, Bruno de Abreu Silva, Vanderlei Bonato |
DSD | 3 |
| 2012 | Power/performance optimization in FPGA-based asymmetric multi-core systemsabstractIn embedded systems, energy efficiency is the new fundamental performance limiter. Considering that, many techniques were applied at different development levels, such as co-design, compilers, schedulers, run-time management, and applications. The fusion of techniques from different levels has also been exploited to increase the optimization opportunities. In this paper, we present a work in progress tool to exploit power/performance optimization techniques in FPGA-based asymmetric multi-core systems. The tool performs optimizations in two phases: compilation and execution. In compilation, are generated from a compiler the hardware and software configurations of a multi-core architecture based on LEON3 processor. During execution, information about application properties, available hardware resources, and system behavior is used to do thread scheduling, clock gating, and dynamic frequency scaling (DFS). Bruno de Abreu Silva, Vanderlei Bonato |
FPL | 2 |
| 2012 | A tool to support Bluespec SystemVerilog coding based on UML diagramsabstractThe use of high level languages to support the development of embedded systems is a current trend. Such approach tends to reduce the development time and cost. The process of translating a high level representation to the final hardware and software architecture is desirable to be automatic encompassing as much as possible the requirements specified in high level model. This work proposes a new tool to support Bluespec SystemVerilog code generation based on models represented via Activity and State UML diagrams. The tool accepts as input the XMI format and the generation process is based on templates where the target language is represented. Sergio H. M. Durand, Vanderlei Bonato |
IECON | 2 |
| 2012 | Designing FPGA-based embedded systems with MARTE: A PIM to PSM converterabstractWith the significant increase in complexity of system-on-chip arises the necessity of embedded systems designers to work with more flexible and detailed modeling rules. The MARTE (Modeling and Analysis of Real Time and Embedded Systems) modeling language developed by OMG (Object Management Group) to design embedded systems, supporting real-time constraints, allows specification and integration of models designed on System (Hw/Sw), Hardware(HW) and Software(Sw) levels. This language enables the MDA methodology (Model-Driven Architecture) to be employed for the development of embedded systems in a standard way independently of implementation technology. In this context, this paper presents a converter, called I2S, in which platform-independent models (PIM) are automatically converted to specific models (PSM) of Altera [1] and Xilinx [2] platforms. It also describes how the models are specified and interpreted in MARTE and how the PIM to PSM transformations occur. The PSM models generated by the tool are synthesizable, which allows their application to real-world problems. Roberto de Medeiros, Marcilyanne Moreira Gois, Vanderlei Bonato |
IECON | 3 |
| 2008 | A Parallel Hardware Architecture for Scale and Rotation Invariant Feature DetectionabstractThis paper proposes a parallel hardware architecture for image feature detection based on the scale invariant feature transform algorithm and applied to the simultaneous localization and mapping problem. The work also proposes specific hardware optimizations considered fundamental to embed such a robotic control system on-a-chip. The proposed architecture is completely stand-alone; it reads the input data directly from a CMOS image sensor and provides the results via a field-programmable gate array coupled to an embedded processor. The results may either be used directly in an on-chip application or accessed through an Ethernet connection. The system is able to detect features up to 30 frames per second (320times240 pixels) and has accuracy similar to a PC-based implementation. The achieved system performance is at least one order of magnitude better than a PC-based solution, a result achieved by investigating the impact of several hardware-orientated optimizations on performance, area and accuracy. Vanderlei Bonato, Eduardo Marques, George A. Constantinides |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | A floating-point Extended Kalman Filter implementation for autonomous mobile robotsabstractLocalization and Mapping are two of the most important capabilities for autonomous mobile robots and have been receiving considerable attention from the scientific computing community over the last 10 years. One of the most efficient methods to address these problems is based on the use of the Extended Kalman Filter (EKF). The EKF simultaneously estimates a model of the environment (map) and the position of the robot based on odometric and exteroceptive sensor information. As this algorithm demands a considerable amount of computation, it is usually executed on high end PCs coupled to the robot. In this work we present an FPGA-based architecture for the EKF algorithm that is capable of processing two-dimensional maps containing up to 1.8k features at real time (14Hz) and is two orders of magnitude more power efficient than a general purpose processor. Vanderlei Bonato, Eduardo Marques, George A. Constantinides |
FPL | 1 |
| 2004 | A Real Time Gesture Recognition System for Mobile Robots
Vanderlei Bonato, Adriano K. Sanches, Marcio Merino Fernandes, João M. P. Cardoso, Eduardo do Valle Simões, Eduardo Marques |
ICINCO (2) | 1 |
| 2003 | Design of a fingerprint system using a hardware/software environmentabstractProcessing system of fingerprint are CPU time intensive, being normally implemented in software. This paper present a new algorithm for fingerprint features localization, that can be easily implemented in hardware (system-on-a-chip, FPGA). This algorithm is composed by 3 stages, first stage read a fingerprint image (255x255pixels, ash tones) and apply a Gaussian Filter, after this, apply a absolute difference mask (ADM) for detector the edges in the image filtered and the last stage look for fingerprint features into the image. The information showed by 3th stage are the coordinate X and Y for each feature detected, asked minutiae. For localization the minutiae, the system pursue the edge detected by ADM, this edge represent ridge edge, and analyzing the information from each pixel pursued is possible to locate the minutiae. The average time for localization all minutiae into the fingerprint image, implemented in hardware (FLEX10KE Family, Altera), was 306 milliseconds. Beyond hardware implementation be fast, is possible create embedded systems. Vanderlei Bonato, Rolf Fredi Molz, João Carlos Furtado, Marcos Flôres Ferrão, Fernando Gehm Moraes |
FPGA | 1 |
| 2003 | Propose of a Hardware Implementation for Fingerprint Systems
Vanderlei Bonato, Rolf Fredi Molz, João Carlos Furtado, Marcos Flôres Ferrão, Fernando Gehm Moraes |
FPL | 1 |