Axel Marmet

dblp:286/1941 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2023
0009-0008-7588-2761ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2023 Resource Sharing in Dataflow Circuits
abstract
To achieve resource-efficient hardware designs, high-level synthesis (HLS) tools share (i.e., time-multiplex) functional units among operations of the same type. This optimization is typically performed in conjunction with operation scheduling to ensure the best possible unit usage at each point in time. Dataflow circuits have emerged as an alternative HLS approach to efficiently handle irregular and control-dominated code. However, these circuits do not have a predetermined schedule—in its absence, it is challenging to determine which operations can share a functional unit without a performance penalty. More critically, although sharing seems to imply only some trivial circuitry, time-multiplexing units in dataflow circuits may cause deadlock by blocking certain data transfers and preventing operations from executing. In this paper, we present a technique to automatically identify performance-acceptable resource sharing opportunities in dataflow circuits. More importantly, we describe a sharing mechanism which achieves functionally correct and deadlock-free dataflow designs. On a set of benchmarks obtained from C code, we show that our approach effectively implements resource sharing. It results in significant area savings at a minor performance penalty compared to dataflow circuits which do not support this feature (i.e., it achieves a 64%, 2%, and 18% average reduction in DSPs, LUTs, and FFs, respectively, with an average increase in total execution time of only 2%) and matches the sharing capabilities of a state-of-the-art HLS tool.
Lana Josipovic, Axel Marmet, Andrea Guerrieri, Paolo Ienne
ACM Trans. Reconfigurable Technol. Syst.2
2022 Resource Sharing in Dataflow Circuits
abstract
To achieve resource-efficient hardware designs, HLS tools share (i.e., time-multiplex) functional units among operations of the same type. This optimization is typically performed together with operation scheduling to ensure the best possible unit usage at each point in time. Dataflow circuits have emerged as an alternative HLS approach to efficiently handle irregular and control-dominated code. Yet, these circuits do not have a predetermined schedule—in its absence, it is challenging to determine which operations can share a functional unit without a performance penalty. Furthermore, although sharing seems to imply only some trivial circuitry, time-multiplexing units in dataflow circuits may cause deadlock by blocking certain data transfers and preventing operations from executing. In this paper, we present a technique to automatically identify performance-acceptable resource sharing opportunities in dataflow circuits and we describe a sharing mechanism that achieves deadlock-free dataflow designs. On benchmarks obtained from C code, we show that our approach effectively implements resource sharing: it results in significant area savings (i.e., a DSP reduction of up to 81%) compared to dataflow circuits which do not support this feature and matches the sharing capabilities of a state-of-the-art HLS tool.
Lana Josipovic, Axel Marmet, Andrea Guerrieri, Paolo Ienne
FCCM2
2021 Resource Sharing in Dataflow Circuits
abstract
To achieve resource-efficient hardware designs, high-level synthesis tools share functional units among operations of the same type. This optimization is typically performed in conjunction with operation scheduling to ensure the best possible unit usage at each point in time. Dataflow circuits have emerged as an alternative HLS approach to efficiently handle irregular and control-dominated code. However, these circuits do not have a predetermined schedule; in its absence, it is challenging to determine which operations can share a functional unit without a performance penalty. Additionally, although sharing seems to imply only trivial circuitry, sharing units in dataflow circuits may cause deadlock by blocking certain data transfers and preventing operations from executing. We developed a complete methodology to implement resource sharing in dataflow designs. Our approach automatically identifies performance-acceptable resource sharing opportunities based on average unit utilization with data tokens. Our sharing mechanism achieves functionally correct and deadlock-free circuits by regulating the multiplexing of tokens at the inputs of the shared unit. On a set of benchmarks obtained out of C code, we showed that our approach effectively implements resource sharing and results in significant area savings compared to dataflow circuits which do not support this feature. Our sharing mechanism is key to achieve different area-performance tradeoffs in dataflow designs and to make them competitive in terms of computational resources with circuits achieved using standard HLS techniques.
Lana Josipovic, Axel Marmet, Andrea Guerrieri, Paolo Ienne
FPGA2