Andrea Guerrieri

dblp:193/4432 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
16since 2021 · last 2025
0000-0002-2104-7452ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 5 first-author · 16 since 2021
YearPublicationVenuePosition
2025 DRSA: Accelerating Macro Placement on Commercial FPGAs
abstract
FPGAs are highly versatile devices, but their backend compilation process is significantly time-consuming. DynaRapid has demonstrated the potential to drastically reduce compilation times to mere seconds by leveraging macro-component-based design hierarchies [1]. However, DynaRapid faces challenges in macro-component placement, often resulting in frequency degradation. In this work, we propose DRSA, a fast placer based on simulated annealing targeting DynaRapid's macros, capable of overcoming the frequency degradation of the previous greedy placement strategy [1].
Menzo Bouaissi, Paolo Ienne, Lana Josipovic, Andrea Guerrieri
FCCM4
2025 Compile in Seconds and Run on an FPGA with DynaRapid
abstract
FPGAs are highly versatile devices, but their backend compilation process is significantly time-consuming. While software code compilation may take just a few seconds, FPGA compilation times can often span from several minutes to hours due to the complexity of the underlying toolchain and the ever-growing device capacity. DynaRapid1is a very fast compilation methodology that generates, in a matter of seconds, placed-and-routed kernel designs for AMD FPGAs [1] The DynaRapid workflow is illustrated in Figure 1.
Andrea Guerrieri, Isaac John Wetenkamp, Chris Lavin, Eddie Hung, Lana Josipovic, Paolo Ienne
FPL1
2024 Survival of the Fastest: Enabling More Out-of-Order Execution in Dataflow Circuits
abstract
Dynamically scheduled HLS, through dataflow circuit generation, has proven successful at exploiting operation-level parallelism in several important situations where statically scheduled HLS fails. Yet, although existing dataflow circuits support out-of-order execution of different operations, they strictly confine successive instances of the same operation to execute sequentially in program order, which drastically affects the circuit's performance in the presence of a long-latency operation. This is in stark contrast with the reordering freedom customary in superscalar processors that naturally exploit qualitatively more parallelism in a broad class of applications. The goal of this work is to produce dataflow circuits that have reordering capabilities closer to those of out-of-order superscalar processors. This can bring dramatic improvements in some practically important cases, including when outer iterations in nested loops are independent and the inner loop execution has an unavoidable large initiation interval. In various cases, our technique increases throughput by a factor dependent on the initiation interval of the kernel, at a comparatively modest area cost.
Ayatallah Elakhras, Andrea Guerrieri, Lana Josipovic, Paolo Ienne
FPGA2
2024 DynaRapid: From C to FPGA in a Few Seconds
abstract
Advancements in design automation technologies, such as high-level synthesis (HLS), have raised the input abstraction level and made the design entry process for FPGAs more friendly to software programmers. In contrast, the backend compilation process for implementing designs on FPGAs is considerably more lengthy compared to software compilation. While software code compilation may take just a few seconds, FPGA compilation times can often span from several minutes to hours due to the complexity of the underlying toolchain and the ever-growing device capacity. In this paper, we present DynaRapid, a very fast compilation tool that generates in a matter of seconds fully-legal placed-and-routed designs for commercial FPGAs. We leverage the inherently modular nature of dataflow circuits created by the HLS tool Dynamatic and combine it with the implementation manipulation capabilities provided by RapidWright. Our approach accelerates the C-to-FPGA implementation process by up to 33× with only 20% of degradation in operating frequency compared to a conventional commercial off-the-shelf implementation flow.
Andrea Guerrieri, Srijeet Guha, Lana Josipovic, Paolo Ienne
FPGA1
2024 DynaRapid: Fast-Tracking from C to Routed Circuits
abstract
Advancements in design automation technologies, such as high-level synthesis (HLS), have raised the input abstraction level and made the design entry process for FPGAs more friendly to software programmers. In contrast, the backend compilation process for implementing designs on FPGAs is considerably more lengthy compared to software compilation: while software code compilation may take just a few seconds, FPGA compilation times can often span from several minutes to hours due to the complexity of the underlying toolchain and ever-growing device capacities. In this paper, we present DynaRapid, a fast compilation tool that generates—in a matter of seconds—fully legal placed-and-routed designs for commercial FPGAs. Elastic circuits created by the HLS tool Dynamatic are made exclusively of a limited number of reusable components; we exploit this fact to create a library of placed and routed building blocks, and then stitch together instances of them as needed through RapidWright. Our approach accelerates the C-to-FPGA implementation process by a geomean $20 \times$ with only 10% of degradation in operating frequency compared to a conventional commercial off-the-shelf implementation flow.
Andrea Guerrieri, Srijeet Guha, Chris Lavin, Eddie Hung, Lana Josipovic, Paolo Ienne
FPL1
2024 Fast Switching Activity Estimation for HLS-Produced Dataflow Circuits
abstract
High-level synthesis (HLS) tools generate hardware designs from high-level software languages while sidestepping intricate low-level hardware details. However, HLS tools struggle with precise dynamic power estimation and optimization: the high abstraction level they operate on typically contains no or limited information on low-level circuit details that power consumption depends on. Dataflow circuits have recently been explored in the HLS context; apart from their ability to achieve performance that is superior to standard HLS-generated circuits, their well-defined structure and computational model offer entirely new opportunities for reasoning about power at the HLS level. This paper exploits this insight to present an accurate switching activity estimator for HLS-produced dataflow circuits. Our estimator combines the knowledge about the dataflow circuit structure with software profiling and detailed glitching analysis to estimate the circuit’s switching activity with an average error rate of 1.8% and average speedup of $17.8 \times$ compared with a cycle-accurate simulator. Our technology-agnostic solution makes a critical advancement in HLS power estimation and sets the stage for integrating power optimization within the HLS process.
Maksymilian Graczyk, Andrea Guerrieri, Lana Josipovic
FPL3
2023 An Iterative Method for Mapping-Aware Frequency Regulation in Dataflow Circuits
abstract
Dataflow circuits promise to overcome the scheduling limitations of standard HLS solutions. However, their performance suffers due to timing overheads caused by their handshake communication protocol. Current pipelining solutions fail to account for logic optimizations that occur during FPGA synthesis, thus producing over-conservative results. In this work, we develop an FPGA mapping-aware timing regulation technique for dataflow circuits; it relies on FPGA synthesis information to identify the circuit’s critical path and optimize it through register placement. Our dataflow circuits Pareto-dominate state-of-the-art solutions, with up to 29% and 21% execution time and area reduction, respectively.
Carmine Rizzi, Andrea Guerrieri, Lana Josipovic
DAC2
2023 Straight to the Queue: Fast Load-Store Queue Allocation in Dataflow Circuits
abstract
Dynamically scheduled high-level synthesis can exploit high levels of parallelism in poorly-predictable control-dominated applications. Yet, dataflow circuits are often generated by literal conversion of basic blocks into circuits interconnected in such a way as to mimic the program's sequential execution. Although correct and quite effective in many cases, this adherence to control flow still significantly limits exploitable parallelism. Recent research introduced techniques to deliver data tokens directly from producers to consumers and achieved tangible benefits both in circuit complexity and execution time. Unfortunately, while this successfully addressed ordinary data dependencies, the problem of potential dependencies through memory remains open: When no technique can statically disambiguate accesses, circuits must be built with load-store queues (LSQs) which, to reorder accesses safely, need memory accesses to be allocated in the queues in program order. Such in-order allocation still demands control circuitry emulating sequential execution, with its negative impact on parallelization. In this paper, we transform potential memory dependencies into virtual data dependencies and use the new direct token delivery strategy to allocate accesses sequentially into the LSQ. In other words, we exploit more parallelism by constructing control circuitry to emulate exclusively those parts of the control flow strictly necessary for in-order allocation. Our results show that we can achieve up to a 74% reduction in execution time compared to prior work, in some cases, at no area cost.
Ayatallah Elakhras, Riya Sawhney, Andrea Guerrieri, Lana Josipovic, Paolo Ienne
FPGA3
2023 Resource Sharing in Dataflow Circuits
abstract
To achieve resource-efficient hardware designs, high-level synthesis (HLS) tools share (i.e., time-multiplex) functional units among operations of the same type. This optimization is typically performed in conjunction with operation scheduling to ensure the best possible unit usage at each point in time. Dataflow circuits have emerged as an alternative HLS approach to efficiently handle irregular and control-dominated code. However, these circuits do not have a predetermined schedule—in its absence, it is challenging to determine which operations can share a functional unit without a performance penalty. More critically, although sharing seems to imply only some trivial circuitry, time-multiplexing units in dataflow circuits may cause deadlock by blocking certain data transfers and preventing operations from executing. In this paper, we present a technique to automatically identify performance-acceptable resource sharing opportunities in dataflow circuits. More importantly, we describe a sharing mechanism which achieves functionally correct and deadlock-free dataflow designs. On a set of benchmarks obtained from C code, we show that our approach effectively implements resource sharing. It results in significant area savings at a minor performance penalty compared to dataflow circuits which do not support this feature (i.e., it achieves a 64%, 2%, and 18% average reduction in DSPs, LUTs, and FFs, respectively, with an average increase in total execution time of only 2%) and matches the sharing capabilities of a state-of-the-art HLS tool.
Lana Josipovic, Axel Marmet, Andrea Guerrieri, Paolo Ienne
ACM Trans. Reconfigurable Technol. Syst.3
2022 Optimizing Lattice-based Post-Quantum Cryptography Codes for High-Level Synthesis
abstract
High-level synthesis is a mature Electronics Design Automation (EDA) technology for building hardware design in a short time. It produces automatically HDL code for FPGAs out of C/C++, bridging the gap from algorithm to hardware. Nevertheless, sometimes the QoR (Quality of Results) can be sub-optimal due to the difficulties of HLS in handling general-purpose software code. In this paper, we explore the current difficulties of HLS while synthesizing Lattice-based Post-Quantum Cryptog-raphy (PQC) algorithms. We propose code-level optimizations to overcome the limitations of high-level synthesis increasing the QoR of generated hardware. We analyzed and improved the results for the algorithms competing in the 3rd round of the NIST standardization process. We show how, starting from the original reference code submitted for the competition, original performance and resource utilization can be improved, in some cases with a speedup factor up to$200\times$or an area reduction of 80%.
Andrea Guerrieri, Gabriel Da Silva Marques, Francesco Regazzoni 0001, Andres Upegui
DSD1
2022 Resource Sharing in Dataflow Circuits
abstract
To achieve resource-efficient hardware designs, HLS tools share (i.e., time-multiplex) functional units among operations of the same type. This optimization is typically performed together with operation scheduling to ensure the best possible unit usage at each point in time. Dataflow circuits have emerged as an alternative HLS approach to efficiently handle irregular and control-dominated code. Yet, these circuits do not have a predetermined schedule—in its absence, it is challenging to determine which operations can share a functional unit without a performance penalty. Furthermore, although sharing seems to imply only some trivial circuitry, time-multiplexing units in dataflow circuits may cause deadlock by blocking certain data transfers and preventing operations from executing. In this paper, we present a technique to automatically identify performance-acceptable resource sharing opportunities in dataflow circuits and we describe a sharing mechanism that achieves deadlock-free dataflow designs. On benchmarks obtained from C code, we show that our approach effectively implements resource sharing: it results in significant area savings (i.e., a DSP reduction of up to 81%) compared to dataflow circuits which do not support this feature and matches the sharing capabilities of a state-of-the-art HLS tool.
Lana Josipovic, Axel Marmet, Andrea Guerrieri, Paolo Ienne
FCCM3
2022 Unleashing Parallelism in Elastic Circuits with Faster Token Delivery
abstract
High-level synthesis (HLS) is the process of automatically generating circuits out of high-level language descriptions. Previous research has shown that dynamically scheduled HLS through elastic circuit generation is successful at exploiting parallelism in some important use-cases. Nevertheless, the literal conversion of a standard compiler's control-data flow graph into elastic circuits often produces circuits with notable resource demands and inferior performance. In this work, we present a methodology for generating more area- and timing-efficient elastic circuits. We show that our strategy results in significant area and timing improvements compared to previous circuit generation strategies.
Ayatallah Elakhras, Andrea Guerrieri, Lana Josipovic, Paolo Ienne
FPL2
2022 A Comprehensive Timing Model for Accurate Frequency Tuning in Dataflow Circuits
abstract
The ability of dataflow circuits to implement dynamic scheduling promises to overcome the conservatism of static scheduling techniques that high-level synthesis tools typically rely on. Yet, the same distributed control mechanism that allows dataflow circuits to achieve high-throughput pipelines when static scheduling cannot also causes long critical paths and frequency degradation. This effect reduces the overall performance benefits of dataflow circuits and makes them an undesirable solution in broad classes of applications. In this work, we provide an in-depth study of the timing of dataflow circuits. We develop a mathematical model that accurately captures combinational delays among different dataflow constructs and appropriately places buffers to control the critical path. On a set of benchmarks obtained from C code, we show that the circuits optimized by our technique accurately meet the clock period target and result in a critical path reduction of up to 38% compared to prior solutions.
Carmine Rizzi, Andrea Guerrieri, Paolo Ienne, Lana Josipovic
FPL2
2022 From C/C++ Code to High-Performance Dataflow Circuits
abstract
High-level synthesis (HLS) tools typically generate statically scheduled datapaths. Static scheduling implies that the resulting circuits have a hard time exploiting parallelism in code with potential memory dependences, with control dependences, or where performance is limited by long latency control decisions. In this work, we describe an HLS approach which generates dynamically scheduled, dataflow circuits out of imperative code. We detail a complete set of rules to transform a standard compiler intermediate representation into a high-performance dataflow circuit that is able to dynamically resolve memory dependences and adapt its behavior on the fly to particular control flow decisions and operation latencies. Compared to a traditional HLS tool, the result is a different tradeoff between performance and circuit complexity: statically scheduled circuits display the best performance per cost in regular applications, but general-purpose, irregular, and control-dominated computing tasks require the runtime flexibility of dynamic scheduling. Therefore, enabling dynamic behavior in HLS is key to dealing with the increasing computational demands of new contexts and broader application domains.
Lana Josipovic, Andrea Guerrieri, Paolo Ienne
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Buffer Placement and Sizing for High-Performance Dataflow Circuits
abstract
Commercial high-level synthesis tools typically produce statically scheduled circuits. Yet, effective C-to-circuit conversion of arbitrary software applications calls for dataflow circuits, as they can handle efficiently variable latencies (e.g., caches), unpredictable memory dependencies, and irregular control flow. Dataflow circuits exhibit an unconventional property: registers (usually referred to as “buffers”) can be placed anywhere in the circuit without changing its semantics, in strong contrast to what happens in traditional datapaths. Yet, although functionally irrelevant, this placement has a significant impact on the circuit’s timing and throughput. In this work, we show how to strategically place buffers into a dataflow circuit to optimize its performance. Our approach extracts a set of choice-free critical loops from arbitrary dataflow circuits and relies on the theory of marked graphs to optimize the buffer placement and sizing. Our performance optimization model supports important high-level synthesis features such as pipelined computational units, units with variable latency and throughput, and if-conversion. We demonstrate the performance benefits of our approach on a set of dataflow circuits obtained from imperative code.
Lana Josipovic, Shabnam Sheikhha, Andrea Guerrieri, Paolo Ienne, Jordi Cortadella
ACM Trans. Reconfigurable Technol. Syst.3
2021 Resource Sharing in Dataflow Circuits
abstract
To achieve resource-efficient hardware designs, high-level synthesis tools share functional units among operations of the same type. This optimization is typically performed in conjunction with operation scheduling to ensure the best possible unit usage at each point in time. Dataflow circuits have emerged as an alternative HLS approach to efficiently handle irregular and control-dominated code. However, these circuits do not have a predetermined schedule; in its absence, it is challenging to determine which operations can share a functional unit without a performance penalty. Additionally, although sharing seems to imply only trivial circuitry, sharing units in dataflow circuits may cause deadlock by blocking certain data transfers and preventing operations from executing. We developed a complete methodology to implement resource sharing in dataflow designs. Our approach automatically identifies performance-acceptable resource sharing opportunities based on average unit utilization with data tokens. Our sharing mechanism achieves functionally correct and deadlock-free circuits by regulating the multiplexing of tokens at the inputs of the shared unit. On a set of benchmarks obtained out of C code, we showed that our approach effectively implements resource sharing and results in significant area savings compared to dataflow circuits which do not support this feature. Our sharing mechanism is key to achieve different area-performance tradeoffs in dataflow designs and to make them competitive in terms of computational resources with circuits achieved using standard HLS techniques.
Lana Josipovic, Axel Marmet, Andrea Guerrieri, Paolo Ienne
FPGA3
2020 Invited Tutorial: Dynamatic: From C/C++ to Dynamically Scheduled Circuits
abstract
High-level synthesis tools, both commercial and academic, typically rely on static scheduling to produce high-throughput pipelines. However, in applications with unpredictable memory accesses or irregular control flow, these tools need to make pessimistic scheduling assumptions. In contrast, dataflow circuits implement dynamically scheduled circuits, in which components communicate locally using a handshake mechanism and exchange data as soon as all conditions for a transaction are satisfied. Due to their ability to adapt the schedule at runtime, dataflow circuits are suitable for handling irregular and control-dominated code. This paper describes Dynamatic, an open-source HLS framework which generates synchronous dataflow circuits out of C/C++ code. The purpose of this paper is to give an introductory overview of Dynamatic and demonstrate some of its use cases, in order to enable others to use the tool and participate in its development.
Lana Josipovic, Andrea Guerrieri, Paolo Ienne
FPGA2
2020 Buffer Placement and Sizing for High-Performance Dataflow Circuits
abstract
Commercial high-level synthesis tools typically produce statically scheduled circuits. Yet, effective C-to-circuit conversion of arbitrary software applications calls for dataflow circuits, as they can handle efficiently variable latencies (e.g., caches) and unpredictable memory dependencies. Dataflow circuits exhibit an unconventional property: registers (usually referred to as "buffers") can be placed anywhere in the circuit without changing its semantics, in strong contrast to what happens in traditional datapaths. Yet, although functionally irrelevant, this placement has a significant impact on the circuit's timing and throughput. In this work, we show how to strategically place buffers into a dataflow circuit to optimize its performance. Our approach extracts a set of choice-free critical loops from arbitrary dataflow circuits and relies on the theory of marked graphs to optimize the buffer placement and sizing. We demonstrate the performance benefits of our approach on a set of dataflow circuits obtained from imperative code.
Lana Josipovic, Shabnam Sheikhha, Andrea Guerrieri, Paolo Ienne, Jordi Cortadella
FPGA3
2020 CloudMoles: Surveillance of Power-Wasting Activities by Infiltrating Undercover Sensors
abstract
Recently, FPGA-accelerated cloud has emerged as a new computing environment. The inclusion of FPGAs in the cloud has created new security risks, some of which are due to circuits exercising excessive switching activity. These power-wasting tenants can cause timing faults in the collocated circuits or a denial-of-service attack by resetting the host FPGA. In this work, we present the idea of populating the FPGA with voltage sensors based on ring oscillators, to continuously monitor the core voltage fluctuations across the entire FPGA. To implement the sensors, we do not lock any FPGA resources; instead, we infiltrate the sensors undercover, by taking advantage of the logic and the routing resources unused by the tenants. Additionally, we infiltrate the sensors into the FPGA circuits after their implementation, but before their deployment on the cloud; the tenants are thus neither aware nor affected by our voltage monitoring system. Finally, we devise a novel metric that takes the sensor measurements to quantify the power wasting activity in the FPGA clock regions where the sensors are infiltrated. We use VTR benchmarks and a Xilinx Virtex-7 FPGA to test the feasibility of our approach. Experimental results demonstrate that, using the undercover voltage sensors and our novel metric, one can accurately locate the source of the malicious power-wasting activity.
Seyedeh Sharareh Mirzargar, Andrea Guerrieri, Mirjana Stojilovic
FPGA2
2019 Speculative Dataflow Circuits
abstract
With FPGAs facing broader application domains, the conversion of imperative languages into dataflow circuits has been recently revamped as a way to overcome the conservatism of statically scheduled high-level synthesis. Apart from the ability to extract parallelism in irregular and control-dominated applications, dynamic scheduling opens a door to speculative execution, one of the most powerful ideas in computer architecture. Speculation allows executing certain operations before it is known whether they are correct or required: it can significantly increase fine-grain parallelism in loops where the condition takes many cycles to compute; it can also increase the performance of circuits limited by potential dependencies by assuming independence early on and by reverting to the correct execution if the prediction was wrong. In this work, we detail our methodology to enable tentative and reversible execution in dynamically scheduled dataflow circuits. We create a generic framework for handling speculation in dataflow circuits and show that our approach can achieve significant performance improvements over traditional circuit generation techniques.
Lana Josipovic, Andrea Guerrieri, Paolo Ienne
FPGA2
2018 LEOSoC: An Open-Source Cross-Platform Embedded Linux Library for Managing Hardware Accelerators in Heterogeneous System-on-Chips(Abstract Only)
abstract
Modern heterogeneous SoCs (System-on-Chip) contain a set of Hard IPs (HIPs) surrounded by an FPGA fabric for hosting custom Hardware Accelerators (HAs). However, efficiently managing such HAs in an embedded Linux environment involves creating and building custom device drivers specific to the target platform, which negatively impacts development cost, portability and time-to-market. To address this issue, we present LEOSoC, an open-source cross-platform embedded Linux library. LEOSoC reduces the development effort required to interface HAs with applications and makes SoCs easy to use for an embedded software developer who is familiar with the semantics of standard POSIX threads. Using LEOSoC does not require any specific version of the Linux kernel, nor to rebuild a custom driver for each new kernel release. LEOSoC consists of a base hardware system and a software layer. Both hardware and software are portable across SoC from various vendors and the library recognizes and auto-adapts to the target SoC platform on which it is running. Furthermore, LEOSoC allows the application to partially or completely change the structure of the HAs at runtime without rebooting the system by leveraging the underlying platforms? support for dynamic full/partial FPGA reconfigurability. The system has been tested on multiple COTS (Commercial Off The Shelf) boards from different vendors, each one running different versions of Linux and, therefore, proving the real portability and usability of LEOSoC in a specific industrial design.
Andrea Guerrieri, Sahand Kashani-Akhavan, Mikhail Asiatici, Pasquale Lombardi, Bilel Belhadj, Paolo Ienne
FPGA1