VLDB 2026 Research / reviewers in the wild / expert
Lana Josipovic
dblp:200/2746
· DBLP profile ↗
41ranked-venue papers
13as first author
33since 2021 · last 2026
0000-0001-6659-8533ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 13 first-author · 33 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graphiti: Formally Verified Out-of-Order Execution in Dataflow CircuitsabstractHigh-level synthesis (HLS) tools automatically synthesise hardware from imperative programs and have seen a significant rise in adoption in both industry and academia. To deliver high-quality hardware designs for increasingly general purpose programs, HLS compilers have to become more aggressive. For the most irregular programs, HLS tools generating dataflow circuits show promising performance by adapting and specializing key ideas from processor architectures, like out-of-order execution and speculation. However, the complexity of these transformations makes them difficult to reason about, increasing the risk of subtle bugs and potentially delaying their adoption in a conservative industry where bugs can be extremely costly. This paper introduces Graphiti, a framework embedded in the Lean 4 proof assistant designed to formally reason about and manipulate dataflow circuits at the core of these HLS tools. We develop a metatheory of graph refinement that allows us to verify a general-purpose dataflow circuit rewriting algorithm. Using this framework, we formally verify a loop rewrite that introduces out-of-order execution into a dataflow circuit. Our evaluation shows that the resulting verified optimization pipeline achieves a 2.1× speedup over the in-order HLS flow and a 5.8× speedup over a verified HLS tool generating a static state machine. We also show that it achieves the same performance compared to an existing unverified approach which introduces out-of-order execution. Graphiti is a step toward a fully-verified HLS flow targeting dataflow circuits. In the interim, it can serve as an extensible, verified, optimizing engine that can be integrated into existing dataflow HLS compilers. Yann Herklotz, Ayatallah Elakhras, Martina Camaioni, Paolo Ienne, Lana Josipovic, Thomas Bourgeat |
ASPLOS (2) | 5 |
| 2026 | Bridging the gap between software and hardwareabstractCustom hardware accelerators, such as FPGAs and ASICs, are a promising solution to deal with our increasing computational demands, as they offer high parallelism and energy efficiency. However, a major barrier to their success and adoption is the difficulty of hardware design—a task available exclusively to a limited number of hardware experts. In this talk, I will discuss the challenges and limitations of current hardware design approaches. I will outline high-level and logic synthesis techniques that overcome these limitations, enabling non-expert designers to achieve performant and energy-efficient circuits. Finally, I will share my vision for future advancements in hardware design and its accessibility to users from various application domains. Lana Josipovic |
CF | 1 |
| 2026 | EagerlyElastic: Correct-by-Construction Eager Execution in Dynamically-Scheduled HLSabstractCircuits generated through dynamically-scheduled high-level synthesis (also known as elastic circuits) are highly performant, as circuit execution can adjust at runtime to variable control flow or memory dependencies. While there are many promising opportunities for further optimization, the resulting circuits can be difficult to analyze—the challenge of maintaining circuit correctness forms a major obstacle to the optimization process. Seeking to increase performance, we introduce for the first time eager execution—speculatively executing both sides of a conditional branch in parallel, with results committed based on the branch condition—to elastic circuits. We build on recent work to present eager execution as a set of formally verified circuit rewrite patterns, localized transformations which we prove preserve correctness. This allows us to pursue high performance, while simultaneously guaranteeing that global circuit correctness is maintained at every stage. Our eager execution strongly increases performance using only slight alterations to the control path of the circuit: unlike the prohibitive resource cost of processor eager execution, elastic-circuit eager execution can be effectively enabled without additional computational units. On a set of selected kernels, we achieve an average speed-up of 3.2x, at an average increase of 1.2x in LUTs and 1.3x in FFs, compared to standard elastic circuits. Excitingly, elastic eager execution opens up a new design space of high-performance circuits, generated through correct-by-construction approaches, which we have only just begun to explore. EagerlyElastic is open-sourced under 10.5281/zenodo.17703725. Shun Katsumi, Emmet Murphy, Lana Josipovic |
FPGA | 3 |
| 2026 | HACE: HLS-Tool-Agnostic CDFG Extraction from RTL DesignsabstractHigh-Level Synthesis (HLS) compilers translate programs written in high-level languages (e.g., C/C++) into hardware by scheduling control-data flow graph (CDFG) operations. While the CDFG is central to scheduling, it is often inaccessible---either hidden by closed-source tools or impossible to export from open-source tools. Having access to such a high-level representation is highly beneficial: it provides a consistent basis for comparing scheduling strategies across different HLS tools. To address this gap, we introduce HACE, an HLS tool-agnostic framework that reconstructs the CDFG from HLS-generated circuits. HACE's extracted CDFGs can serve as a common ground for fair evaluation of different schedulers. We demonstrate HACE on RTL designs produced by commercial and open-source HLS compilers, and show how it facilitates cross-HLS-tool comparisons. HACE is open-sourced at https://github.com/ETHZ-DYNAMO/hace, offering a practical foundation for research and experimentation with diverse HLS flows. Carmine Rizzi, Sebastian Pfeiler, Lana Josipovic |
FPGA | 3 |
| 2025 | ElasticMiter: Formally Verified Dataflow Circuit RewritesabstractDataflow circuits have been studied for decades as a way to implement both asynchronous and synchronous designs, and, more recently, have attracted attention as the target of high-level synthesis (HLS) compilers. Yet, little is known about mechanisms to systematically transform and optimize the datapaths of the obtained circuits into functionally equivalent but simpler ones. The main challenge is that of equivalence verification: The latency-insensitive nature of dataflow circuits is incompatible with the standard notion of sequential equivalence, which prevents the direct usage of standard sequential equivalence verification strategies and hinders the development of formally verified dataflow circuit transformations in HLS. In this paper, we devise a generic framework for verifying the equivalence of latency-insensitive circuits. To showcase the practical usefulness of our verification framework, we develop a graph rewriting system that systematically transforms dataflow circuits into simpler ones. We employ our framework to verify our graph rewriting patterns and thus prove that the obtained circuits are equivalent to the original ones. Our work is the first to formally verify dataflow circuit transformations and is a foundation for building formally verified dataflow HLS compilers. Ayatallah Elakhras, Martin Erhart, Paolo Ienne, Lana Josipovic |
ASPLOS (2) | 5 |
| 2025 | CRUSH: A Credit-Based Approach for Functional Unit Sharing in Dynamically Scheduled HLSabstractDynamically scheduled high-level synthesis (HLS) automatically translates software code (e.g., C/C++) to dataflow circuits-networks of compute units that communicate via handshake signals. These signals schedule the circuit during runtime, allowing them to handle irregular control flow or unpredictable memory accesses efficiently, thus giving them performance merit over statically scheduled circuits produced by standard HLS tools. Lana Josipovic |
ASPLOS (1) | 2 |
| 2025 | SimGen: Simulation Pattern Generation for Efficient Equivalence CheckingabstractCombinational equivalence checking for hardware design tends to be slow due to the number and complexity of in-termediate node equivalences considered by the SAT solver. This is because the solver often spends extensive time disproving nodes that appear equivalent under random simulation. We propose SimGen, an open-source and expressive simulation pattern generator inspired by Automatic Test Pattern Generation (ATPG); it exploits the circuit's structure to disprove the equivalence of circuit nodes and avoid excessive SAT calls. We demonstrate the effectiveness of SimGen's simulation patterns over those generated by state-of-the-art random and guided simulation. Carmine Rizzi, Sarah Brunner, Alan Mishchenko, Lana Josipovic |
DATE | 4 |
| 2025 | DRSA: Accelerating Macro Placement on Commercial FPGAsabstractFPGAs are highly versatile devices, but their backend compilation process is significantly time-consuming. DynaRapid has demonstrated the potential to drastically reduce compilation times to mere seconds by leveraging macro-component-based design hierarchies [1]. However, DynaRapid faces challenges in macro-component placement, often resulting in frequency degradation. In this work, we propose DRSA, a fast placer based on simulated annealing targeting DynaRapid's macros, capable of overcoming the frequency degradation of the previous greedy placement strategy [1]. Menzo Bouaissi, Paolo Ienne, Lana Josipovic, Andrea Guerrieri |
FCCM | 3 |
| 2025 | Compile in Seconds and Run on an FPGA with DynaRapidabstractFPGAs are highly versatile devices, but their backend compilation process is significantly time-consuming. While software code compilation may take just a few seconds, FPGA compilation times can often span from several minutes to hours due to the complexity of the underlying toolchain and the ever-growing device capacity. DynaRapid1is a very fast compilation methodology that generates, in a matter of seconds, placed-and-routed kernel designs for AMD FPGAs [1] The DynaRapid workflow is illustrated in Figure 1. Andrea Guerrieri, Isaac John Wetenkamp, Chris Lavin, Eddie Hung, Lana Josipovic, Paolo Ienne |
FPL | 5 |
| 2025 | Promise: Property Mining for Sequential SynthesisabstractModularity—composing a large system using individually designed units—is an essential practice in hardware design. Yet, modularity might compromise quality: when individually designed units are put together, some of their states may become unreachable and, consequently, the logic that implements them is redundant. Sequential synthesis aims to remove redundant circuit logic by leveraging state unreachability. It critically depends on invariants—relations between signals and registers that hold in all reachable states—to prove the validity of redundancies. Yet, existing invariant generation techniques are mostly problem-specific (for a particular circuit or a property) or reliant on localized reasoning. We propose Promise, a fast circuit redundancy removal strategy. Promise exploits the rich information from simulation traces and uses efficient polynomial-time algorithms to infer global circuit invariants, optimizing the circuit and aiding other sequential synthesis procedures. Experiments show that Promise effectively optimizes circuits produced by high-level synthesis tools. Promise is open-sourced and available at github.com/ETHZ-DYNAMO/promise. Jordi Cortadella, Lana Josipovic |
ICCAD | 3 |
| 2025 | OptiPIM: Optimizing Processing-in-Memory Acceleration Using Integer Linear ProgrammingabstractProcessing-in-memory (PIM) accelerators provide superior performance and energy efficiency to conventional architectures by minimizing off-chip data movement and exploiting extensive internal memory bandwidth for computation.However, efficient PIM acceleration requires careful software-hardware mapping that transforms application algorithms into PIM operations and data layout.Unfortunately, existing PIM accelerators adopt manually tuned heuristics or exhaustive search to determine the mappings on PIM accelerators, leading to under-optimized performance and/or long optimization time.In this work, we propose OptiPIM, a novel optimization framework based on Integer Linear Programming (ILP) to efficiently generate the optimal mapping for data-intensive applications on PIM accelerators.The proposed framework adopts a PIM-friendly mapping representation with accurate cost modeling and a concise description of the entire design space, allowing us to formulate an efficient and effective ILP problem and optimize the mapping on PIM architectures.We implement OptiPIM in the opensource MLIR framework, enabling OptiPIM to generate optimized mappings for PyTorch workloads on PIM accelerators.We evaluate widely used machine learning workloads on two state-of-the-art PIM accelerators.Our experiments show that OptiPIM can generate optimal mappings within 4 minutes.Mappings generated by OptiPIM are at least 1.9× faster than those generated by heuristics. Minxuan Zhou, Yue Pan 0002, Chien-Yi Yang, Lana Josipovic, Tajana Rosing |
ISCA | 5 |
| 2024 | Survival of the Fastest: Enabling More Out-of-Order Execution in Dataflow CircuitsabstractDynamically scheduled HLS, through dataflow circuit generation, has proven successful at exploiting operation-level parallelism in several important situations where statically scheduled HLS fails. Yet, although existing dataflow circuits support out-of-order execution of different operations, they strictly confine successive instances of the same operation to execute sequentially in program order, which drastically affects the circuit's performance in the presence of a long-latency operation. This is in stark contrast with the reordering freedom customary in superscalar processors that naturally exploit qualitatively more parallelism in a broad class of applications. The goal of this work is to produce dataflow circuits that have reordering capabilities closer to those of out-of-order superscalar processors. This can bring dramatic improvements in some practically important cases, including when outer iterations in nested loops are independent and the inner loop execution has an unavoidable large initiation interval. In various cases, our technique increases throughput by a factor dependent on the initiation interval of the kernel, at a comparatively modest area cost. Ayatallah Elakhras, Andrea Guerrieri, Lana Josipovic, Paolo Ienne |
FPGA | 3 |
| 2024 | DynaRapid: From C to FPGA in a Few SecondsabstractAdvancements in design automation technologies, such as high-level synthesis (HLS), have raised the input abstraction level and made the design entry process for FPGAs more friendly to software programmers. In contrast, the backend compilation process for implementing designs on FPGAs is considerably more lengthy compared to software compilation. While software code compilation may take just a few seconds, FPGA compilation times can often span from several minutes to hours due to the complexity of the underlying toolchain and the ever-growing device capacity. In this paper, we present DynaRapid, a very fast compilation tool that generates in a matter of seconds fully-legal placed-and-routed designs for commercial FPGAs. We leverage the inherently modular nature of dataflow circuits created by the HLS tool Dynamatic and combine it with the implementation manipulation capabilities provided by RapidWright. Our approach accelerates the C-to-FPGA implementation process by up to 33× with only 20% of degradation in operating frequency compared to a conventional commercial off-the-shelf implementation flow. Andrea Guerrieri, Srijeet Guha, Lana Josipovic, Paolo Ienne |
FPGA | 3 |
| 2024 | Suppressing Spurious Dynamism of Dataflow Circuits via Latency and Occupancy BalancingabstractDataflow circuits produced via high-level synthesis (HLS) adapt their schedule at runtime to unpredictable data and control outcomes, thus promising superior performance to standard HLS solutions. However, their distributed handshake mechanism is extremely resource-expensive-there is a clear benefit in simplifying or removing it whenever it is unneeded for correctness and performance. Yet, even in such situations, transient and spurious stalls and irregular data exchanges prevent the systematic removal of handshake logic, thus resulting in an unnecessary resource overhead. In this work, we present a scalable strategy based on linear programming (LP) that eliminates unnecessary and spurious stalls via latency and occupancy balancing; the data exchange periodicity and predictability in the resulting circuits uncover new handshake logic removal opportunities and enable the formation of simple local controllers to replace it. We show that, in cases where dynamism is unneeded, our circuits qualitatively match those produced by standard HLS tools. Otherwise, our strategy allows us to systematically trade off area and performance to exploit various degrees of dynamism depending on the optimization objective. Lana Josipovic |
FPGA | 2 |
| 2024 | DynaRapid: Fast-Tracking from C to Routed CircuitsabstractAdvancements in design automation technologies, such as high-level synthesis (HLS), have raised the input abstraction level and made the design entry process for FPGAs more friendly to software programmers. In contrast, the backend compilation process for implementing designs on FPGAs is considerably more lengthy compared to software compilation: while software code compilation may take just a few seconds, FPGA compilation times can often span from several minutes to hours due to the complexity of the underlying toolchain and ever-growing device capacities. In this paper, we present DynaRapid, a fast compilation tool that generates—in a matter of seconds—fully legal placed-and-routed designs for commercial FPGAs. Elastic circuits created by the HLS tool Dynamatic are made exclusively of a limited number of reusable components; we exploit this fact to create a library of placed and routed building blocks, and then stitch together instances of them as needed through RapidWright. Our approach accelerates the C-to-FPGA implementation process by a geomean $20 \times$ with only 10% of degradation in operating frequency compared to a conventional commercial off-the-shelf implementation flow. Andrea Guerrieri, Srijeet Guha, Chris Lavin, Eddie Hung, Lana Josipovic, Paolo Ienne |
FPL | 5 |
| 2024 | Fast Switching Activity Estimation for HLS-Produced Dataflow CircuitsabstractHigh-level synthesis (HLS) tools generate hardware designs from high-level software languages while sidestepping intricate low-level hardware details. However, HLS tools struggle with precise dynamic power estimation and optimization: the high abstraction level they operate on typically contains no or limited information on low-level circuit details that power consumption depends on. Dataflow circuits have recently been explored in the HLS context; apart from their ability to achieve performance that is superior to standard HLS-generated circuits, their well-defined structure and computational model offer entirely new opportunities for reasoning about power at the HLS level. This paper exploits this insight to present an accurate switching activity estimator for HLS-produced dataflow circuits. Our estimator combines the knowledge about the dataflow circuit structure with software profiling and detailed glitching analysis to estimate the circuit’s switching activity with an average error rate of 1.8% and average speedup of $17.8 \times$ compared with a cycle-accurate simulator. Our technology-agnostic solution makes a critical advancement in HLS power estimation and sets the stage for integrating power optimization within the HLS process. Maksymilian Graczyk, Andrea Guerrieri, Lana Josipovic |
FPL | 4 |
| 2024 | Balor: HLS Source Code Evaluator Based on Custom Graphs and Hierarchical GNNsabstractWhile High-Level Synthesis (HLS) enables circuit generation directly from languages like C/C++ and OpenCL, optimal implementations require additional design specification through compiler directives. Automated optimization of these directives requires evaluation of candidate designs, but is bottlenecked by the high computational cost. Graph Neural Networks (GNNs) have recently emerged as the state-of-the-art for estimating the required Quality of Results (QoR) directly from high-level source code, but they come with associated challenges: the difficulty of information propagation, the complexity of software-oriented graph representations, and high levels of extraneous computation increase both error and computational cost. We present Balor, an HLS source code evaluator, which consists of a graph compiler and a GNN-based QoR estimator. The modular graph compiler is tailor-made for HLS QoR estimation, producing smaller graphs that contain only HLS-specific information. Balor analytically propagates important information across the graph, allowing us to pivot to small, local GNNs, while outperforming more expensive networks. Additionally, we make use of the natural hierarchy of an HLS kernel, and cluster the graph by basic block, to further propagate and process information---without the need for additional datasets. Combined, our contributions simultaneously reduce estimation error by 41%, and computational cost by 82%. Balor is open-sourced at github.com/emmet-murphy/balor Emmet Murphy, Lana Josipovic |
ICCAD | 2 |
| 2023 | An Iterative Method for Mapping-Aware Frequency Regulation in Dataflow CircuitsabstractDataflow circuits promise to overcome the scheduling limitations of standard HLS solutions. However, their performance suffers due to timing overheads caused by their handshake communication protocol. Current pipelining solutions fail to account for logic optimizations that occur during FPGA synthesis, thus producing over-conservative results. In this work, we develop an FPGA mapping-aware timing regulation technique for dataflow circuits; it relies on FPGA synthesis information to identify the circuit’s critical path and optimize it through register placement. Our dataflow circuits Pareto-dominate state-of-the-art solutions, with up to 29% and 21% execution time and area reduction, respectively. Carmine Rizzi, Andrea Guerrieri, Lana Josipovic |
DAC | 3 |
| 2023 | Straight to the Queue: Fast Load-Store Queue Allocation in Dataflow CircuitsabstractDynamically scheduled high-level synthesis can exploit high levels of parallelism in poorly-predictable control-dominated applications. Yet, dataflow circuits are often generated by literal conversion of basic blocks into circuits interconnected in such a way as to mimic the program's sequential execution. Although correct and quite effective in many cases, this adherence to control flow still significantly limits exploitable parallelism. Recent research introduced techniques to deliver data tokens directly from producers to consumers and achieved tangible benefits both in circuit complexity and execution time. Unfortunately, while this successfully addressed ordinary data dependencies, the problem of potential dependencies through memory remains open: When no technique can statically disambiguate accesses, circuits must be built with load-store queues (LSQs) which, to reorder accesses safely, need memory accesses to be allocated in the queues in program order. Such in-order allocation still demands control circuitry emulating sequential execution, with its negative impact on parallelization. In this paper, we transform potential memory dependencies into virtual data dependencies and use the new direct token delivery strategy to allocate accesses sequentially into the LSQ. In other words, we exploit more parallelism by constructing control circuitry to emulate exclusively those parts of the control flow strictly necessary for in-order allocation. Our results show that we can achieve up to a 74% reduction in execution time compared to prior work, in some cases, at no area cost. Ayatallah Elakhras, Riya Sawhney, Andrea Guerrieri, Lana Josipovic, Paolo Ienne |
FPGA | 4 |
| 2023 | Eliminating Excessive Dynamism of Dataflow Circuits Using Model CheckingabstractRecent HLS efforts explore the generation of dynamically scheduled, dataflow circuits from high-level code; their ability to adapt the schedule at runtime to particular data and control outcomes promises superior performance to standard, statically scheduled HLS solutions. However, dataflow circuits are notoriously resource-expensive: their distributed handshake mechanism brings performance benefits in some cases, but causes an unneeded resource overhead when general dynamism is not required. In this work, we present a verification framework based on model checking to systematically reduce the hardware complexity of dataflow circuits. We devise a series of formal proofs that identify the absence of particular behavioral scenarios and use this information to replace the generic dataflow logic with simpler and cheaper control structures. On a set of benchmarks obtained from high-level code, we demonstrate that our technique significantly reduces the resource requirements of dataflow circuits (i.e., it results in LUT and FF reductions of up to 51% and 53%, respectively), while still reaping all performance benefits of dynamic scheduling. Emmet Murphy, Jordi Cortadella, Lana Josipovic |
FPGA | 4 |
| 2023 | MapBuf: Simultaneous Technology Mapping and Buffer Insertion for HLS Performance OptimizationabstractBuffer placement (i.e., pipelining) for frequency regulation is a fundamental step of high-level synthesis (HLS). Typical HLS approaches place buffers before technology mapping; as the circuit implementation details are unknown, the HLS tool must resort to precharacterized and conservative delay estimates when deciding on the buffer placement. An alternative is to place buffers after technology mapping when the circuit details are known. However, the buffers themselves may invalidate prior mapping assumptions and irreversibly impact the ultimate circuit frequency. In this work, we propose a methodology that simultaneously tackles technology mapping and buffer insertion in HLS-produced dataflow circuits. The source code of our approach is open-source and integrated into a complete HLS framework; it achieves a 13.32% and 11.14% average improvement in execution time and area compared to state-of-the-art approaches that handle buffering and technology mapping separately. Carmine Rizzi, Lana Josipovic |
ICCAD | 3 |
| 2023 | Automatic Inductive Invariant Generation for Scalable Dataflow Circuit VerificationabstractFormal verification via BDD-based reachability analysis has been shown to improve the quality of dataflow circuits produced via high-level synthesis (HLS): it can restrict the generality of the dataflow handshake logic only to provably required constructs and significantly improve their resource requirements. Unfortunately, BDD- based strategies are unscalable for larger circuits. A promising alternative is k-induction, which offers scalability in the presence of suitable inductive invariants. Yet, appropriate invariants are not straightforward to determine: they must provide exclusively relevant information that constrains the induction to a small number of steps without complexing the system under verification. In this paper, we propose a fully automated framework that systematically generates suitable inductive invariants for scalable dataflow circuit verification. Our framework systematically exploits a variety of HLS insights to convey relevant invariant information to the verifier and applies to any dataflow circuit generated from C code. On a set of representative benchmarks, we show that our method significantly outperforms prior BDD-based approaches (i.e., it takes minutes to prove properties that a BDD-based checker cannot prove in days) with only a minor reduction in verification capabilities. We also demonstrate the effectiveness of our approach over prior induction-based techniques by proving up to 4 x more properties in the same runtime. Lana Josipovic |
ICCAD | 2 |
| 2023 | Parallelising Control Flow in Dynamic-scheduling High-level SynthesisabstractRecently, there is a trend to use high-level synthesis (HLS) tools to generate dynamically scheduled hardware. The generated hardware is made up of components connected using handshake signals. These handshake signals schedule the components at runtime when inputs become available. Such approaches promise superior performance on “irregular” source programs, such as those whose control flow depends on input data. This is at the cost of additional area. Current dynamic scheduling techniques are well able to exploit parallelism among instructions within each basic block (BB) of the source program, but parallelism between BBs is under-explored, due to the complexity in runtime control flows and memory dependencies. Existing tools allow some of the operations of different BBs to overlap, but to simplify the analysis required at compile time they require the BBs to start in strict program order, thus limiting the achievable parallelism and overall performance. We formulate a general dependency model suitable for comparing the ability of different dynamic scheduling approaches to extract maximal parallelism at runtime. Using this model, we explore a variety of mechanisms for runtime scheduling, incorporating and generalising existing approaches. In particular, we precisely identify the restrictions in existing scheduling implementation and define possible optimisation solutions. We identify two particularly promising examples where the compile-time overhead is small and the area overhead is minimal and yet we are able to significantly speed up execution time: (1) parallelising consecutive independent loops; and (2) parallelising independent inner-loop instances in a nested loop as individual threads. Using benchmark sets from related works, we compare our proposed toolflow against a state-of-the-art dynamic-scheduling HLS tool called Dynamatic. Our results show that, on average, our toolflow yields a 4× speedup from (1) and a 2.9× speedup from (2), with a negligible area overhead. This increases to a 14.3× average speedup when combining (1) and (2). Jianyi Cheng, Lana Josipovic, John Wickerson, George A. Constantinides |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2023 | Resource Sharing in Dataflow CircuitsabstractTo achieve resource-efficient hardware designs, high-level synthesis (HLS) tools share (i.e., time-multiplex) functional units among operations of the same type. This optimization is typically performed in conjunction with operation scheduling to ensure the best possible unit usage at each point in time. Dataflow circuits have emerged as an alternative HLS approach to efficiently handle irregular and control-dominated code. However, these circuits do not have a predetermined schedule—in its absence, it is challenging to determine which operations can share a functional unit without a performance penalty. More critically, although sharing seems to imply only some trivial circuitry, time-multiplexing units in dataflow circuits may cause deadlock by blocking certain data transfers and preventing operations from executing. In this paper, we present a technique to automatically identify performance-acceptable resource sharing opportunities in dataflow circuits. More importantly, we describe a sharing mechanism which achieves functionally correct and deadlock-free dataflow designs. On a set of benchmarks obtained from C code, we show that our approach effectively implements resource sharing. It results in significant area savings at a minor performance penalty compared to dataflow circuits which do not support this feature (i.e., it achieves a 64%, 2%, and 18% average reduction in DSPs, LUTs, and FFs, respectively, with an average increase in total execution time of only 2%) and matches the sharing capabilities of a state-of-the-art HLS tool. Lana Josipovic, Axel Marmet, Andrea Guerrieri, Paolo Ienne |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2022 | Resource Sharing in Dataflow CircuitsabstractTo achieve resource-efficient hardware designs, HLS tools share (i.e., time-multiplex) functional units among operations of the same type. This optimization is typically performed together with operation scheduling to ensure the best possible unit usage at each point in time. Dataflow circuits have emerged as an alternative HLS approach to efficiently handle irregular and control-dominated code. Yet, these circuits do not have a predetermined schedule—in its absence, it is challenging to determine which operations can share a functional unit without a performance penalty. Furthermore, although sharing seems to imply only some trivial circuitry, time-multiplexing units in dataflow circuits may cause deadlock by blocking certain data transfers and preventing operations from executing. In this paper, we present a technique to automatically identify performance-acceptable resource sharing opportunities in dataflow circuits and we describe a sharing mechanism that achieves deadlock-free dataflow designs. On benchmarks obtained from C code, we show that our approach effectively implements resource sharing: it results in significant area savings (i.e., a DSP reduction of up to 81%) compared to dataflow circuits which do not support this feature and matches the sharing capabilities of a state-of-the-art HLS tool. Lana Josipovic, Axel Marmet, Andrea Guerrieri, Paolo Ienne |
FCCM | 1 |
| 2022 | Dynamic Inter-Block Scheduling for HLSabstractA recent theme in HLS research is the production of dynamically scheduled circuits, which are made up of components that use handshaking to schedule themselves at run time, as opposed to following a schedule determined statically at compile time. Dynamically scheduled circuits promise superior performance on ‘irregular’ source programs, such as those whose control flow depends on input data, at the cost of additional area. Current dynamic scheduling techniques are well able to exploit parallelism among instructions within each basic block (BB) of the source program, but parallelism between BBs is underexplored. Although current tools allow the operations of different BBs to overlap, they require the BBs to start in strict program order, thus limiting the achievable parallelism and overall performance. We seek to lift this restriction. Doing so involves developing a toolflow that tackles the following challenges: (1) finding consecutive subgraphs in the control-flow graph and using static analysis to identify those subgraphs that can be safely parallelised, and (2) adapting the circuit so that those subgraphs are executed in parallel while ensuring deterministic circuit behaviour and correct usage of memory interfaces. Using two benchmark sets from related works, we compare our proposed toolflow against a state-of-the-art dynamically scheduled HLS tool called Dynamatic. Our results show that after standard loop unrolling is applied, our toolflow yields a 4 x average speedup, with a negligible area overhead. This increases to a 7.3 x average speedup when our toolflow is further combined with C-slow pipelining. Jianyi Cheng, Lana Josipovic, George A. Constantinides, John Wickerson |
FPL | 2 |
| 2022 | Unleashing Parallelism in Elastic Circuits with Faster Token DeliveryabstractHigh-level synthesis (HLS) is the process of automatically generating circuits out of high-level language descriptions. Previous research has shown that dynamically scheduled HLS through elastic circuit generation is successful at exploiting parallelism in some important use-cases. Nevertheless, the literal conversion of a standard compiler's control-data flow graph into elastic circuits often produces circuits with notable resource demands and inferior performance. In this work, we present a methodology for generating more area- and timing-efficient elastic circuits. We show that our strategy results in significant area and timing improvements compared to previous circuit generation strategies. Ayatallah Elakhras, Andrea Guerrieri, Lana Josipovic, Paolo Ienne |
FPL | 3 |
| 2022 | A Comprehensive Timing Model for Accurate Frequency Tuning in Dataflow CircuitsabstractThe ability of dataflow circuits to implement dynamic scheduling promises to overcome the conservatism of static scheduling techniques that high-level synthesis tools typically rely on. Yet, the same distributed control mechanism that allows dataflow circuits to achieve high-throughput pipelines when static scheduling cannot also causes long critical paths and frequency degradation. This effect reduces the overall performance benefits of dataflow circuits and makes them an undesirable solution in broad classes of applications. In this work, we provide an in-depth study of the timing of dataflow circuits. We develop a mathematical model that accurately captures combinational delays among different dataflow constructs and appropriately places buffers to control the critical path. On a set of benchmarks obtained from C code, we show that the circuits optimized by our technique accurately meet the clock period target and result in a critical path reduction of up to 38% compared to prior solutions. Carmine Rizzi, Andrea Guerrieri, Paolo Ienne, Lana Josipovic |
FPL | 4 |
| 2022 | Load-Store Queue Sizing for Efficient Dataflow CircuitsabstractDataflow circuits implement dynamic scheduling and have recently been explored as an alternative to standard, statically scheduled high-level synthesis (HLS) solutions. In contrast to static HLS, dataflow circuits resolve memory dependencies during runtime by employing load-store queues (LSQs) at the memory interface. However, LSQs are extremely resource-expensive to implement in a spatial system and may cause notable frequency degradation. Therefore, there is a clear need to minimize their size and complexity, while still allowing the circuit to achieve a high computational rate. So far, designers resorted to manually tuning the LSQ depth (i.e., number of queue entries) to trade off area and performance; yet, this approach is evidently time-consuming and unfeasible for complex designs. In this work, we develop a strategy to automatically determine the most affordable LSQ depths in dataflow circuits while maintaining the best possible circuit throughput. We demonstrate our technique on benchmarks obtained from C code with different memory access patterns and show that it can effectively produce the desired Pareto-optimal design points. Carmine Rizzi, Lana Josipovic |
FPT | 3 |
| 2022 | DASS: Combining Dynamic & Static Scheduling in High-Level SynthesisabstractA central task in high-level synthesis isscheduling: the allocation of operations to clock cycles. The classic approach to scheduling isstatic, in which each operation is mapped to a clock cycle at compile-time, but recent years have seen the emergence ofdynamicscheduling, in which an operation’s clock cycle is only determined at runtime. Both approaches have their merits: static scheduling (SS) can lead to simpler circuitry and more resource sharing, while dynamic scheduling (DS) can lead to faster hardware when the computation has nontrivial control flow. In this work, we seek a scheduling approach that combines the best of both worlds. Our idea is to identify the parts of the input program, where DS does not bring any performance advantage and to use SS on those parts. These statically scheduled parts are then treated as black boxes when creating a dataflow circuit for the remainder of the program, which can benefit from the flexibility of DS. An empirical evaluation on a range of applications suggests that by using this approach, we can obtain 74% of the area savings that would be made by switching from DS to SS, and 135% of the performance benefits that would be made by switching from SS to DS. Jianyi Cheng, Lana Josipovic, George A. Constantinides, Paolo Ienne, John Wickerson |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | From C/C++ Code to High-Performance Dataflow CircuitsabstractHigh-level synthesis (HLS) tools typically generate statically scheduled datapaths. Static scheduling implies that the resulting circuits have a hard time exploiting parallelism in code with potential memory dependences, with control dependences, or where performance is limited by long latency control decisions. In this work, we describe an HLS approach which generates dynamically scheduled, dataflow circuits out of imperative code. We detail a complete set of rules to transform a standard compiler intermediate representation into a high-performance dataflow circuit that is able to dynamically resolve memory dependences and adapt its behavior on the fly to particular control flow decisions and operation latencies. Compared to a traditional HLS tool, the result is a different tradeoff between performance and circuit complexity: statically scheduled circuits display the best performance per cost in regular applications, but general-purpose, irregular, and control-dominated computing tasks require the runtime flexibility of dynamic scheduling. Therefore, enabling dynamic behavior in HLS is key to dealing with the increasing computational demands of new contexts and broader application domains. Lana Josipovic, Andrea Guerrieri, Paolo Ienne |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Buffer Placement and Sizing for High-Performance Dataflow CircuitsabstractCommercial high-level synthesis tools typically produce statically scheduled circuits. Yet, effective C-to-circuit conversion of arbitrary software applications calls for dataflow circuits, as they can handle efficiently variable latencies (e.g., caches), unpredictable memory dependencies, and irregular control flow. Dataflow circuits exhibit an unconventional property: registers (usually referred to as “buffers”) can be placed anywhere in the circuit without changing its semantics, in strong contrast to what happens in traditional datapaths. Yet, although functionally irrelevant, this placement has a significant impact on the circuit’s timing and throughput. In this work, we show how to strategically place buffers into a dataflow circuit to optimize its performance. Our approach extracts a set of choice-free critical loops from arbitrary dataflow circuits and relies on the theory of marked graphs to optimize the buffer placement and sizing. Our performance optimization model supports important high-level synthesis features such as pipelined computational units, units with variable latency and throughput, and if-conversion. We demonstrate the performance benefits of our approach on a set of dataflow circuits obtained from imperative code. Lana Josipovic, Shabnam Sheikhha, Andrea Guerrieri, Paolo Ienne, Jordi Cortadella |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2021 | Resource Sharing in Dataflow CircuitsabstractTo achieve resource-efficient hardware designs, high-level synthesis tools share functional units among operations of the same type. This optimization is typically performed in conjunction with operation scheduling to ensure the best possible unit usage at each point in time. Dataflow circuits have emerged as an alternative HLS approach to efficiently handle irregular and control-dominated code. However, these circuits do not have a predetermined schedule; in its absence, it is challenging to determine which operations can share a functional unit without a performance penalty. Additionally, although sharing seems to imply only trivial circuitry, sharing units in dataflow circuits may cause deadlock by blocking certain data transfers and preventing operations from executing. We developed a complete methodology to implement resource sharing in dataflow designs. Our approach automatically identifies performance-acceptable resource sharing opportunities based on average unit utilization with data tokens. Our sharing mechanism achieves functionally correct and deadlock-free circuits by regulating the multiplexing of tokens at the inputs of the shared unit. On a set of benchmarks obtained out of C code, we showed that our approach effectively implements resource sharing and results in significant area savings compared to dataflow circuits which do not support this feature. Our sharing mechanism is key to achieve different area-performance tradeoffs in dataflow designs and to make them competitive in terms of computational resources with circuits achieved using standard HLS techniques. Lana Josipovic, Axel Marmet, Andrea Guerrieri, Paolo Ienne |
FPGA | 1 |
| 2020 | Combining Dynamic & Static Scheduling in High-level SynthesisabstractA central task in high-level synthesis is scheduling: the allocation of operations to clock cycles. The classic approach to scheduling is static, in which each operation is mapped to a clock cycle at compile-time, but recent years have seen the emergence of dynamic scheduling, in which an operation's clock cycle is only determined at run-time. Both approaches have their merits: static scheduling can lead to simpler circuitry and more resource sharing, while dynamic scheduling can lead to faster hardware when the computation has non-trivial control flow. Jianyi Cheng, Lana Josipovic, George A. Constantinides, Paolo Ienne, John Wickerson |
FPGA | 2 |
| 2020 | Invited Tutorial: Dynamatic: From C/C++ to Dynamically Scheduled CircuitsabstractHigh-level synthesis tools, both commercial and academic, typically rely on static scheduling to produce high-throughput pipelines. However, in applications with unpredictable memory accesses or irregular control flow, these tools need to make pessimistic scheduling assumptions. In contrast, dataflow circuits implement dynamically scheduled circuits, in which components communicate locally using a handshake mechanism and exchange data as soon as all conditions for a transaction are satisfied. Due to their ability to adapt the schedule at runtime, dataflow circuits are suitable for handling irregular and control-dominated code. This paper describes Dynamatic, an open-source HLS framework which generates synchronous dataflow circuits out of C/C++ code. The purpose of this paper is to give an introductory overview of Dynamatic and demonstrate some of its use cases, in order to enable others to use the tool and participate in its development. Lana Josipovic, Andrea Guerrieri, Paolo Ienne |
FPGA | 1 |
| 2020 | Buffer Placement and Sizing for High-Performance Dataflow CircuitsabstractCommercial high-level synthesis tools typically produce statically scheduled circuits. Yet, effective C-to-circuit conversion of arbitrary software applications calls for dataflow circuits, as they can handle efficiently variable latencies (e.g., caches) and unpredictable memory dependencies. Dataflow circuits exhibit an unconventional property: registers (usually referred to as "buffers") can be placed anywhere in the circuit without changing its semantics, in strong contrast to what happens in traditional datapaths. Yet, although functionally irrelevant, this placement has a significant impact on the circuit's timing and throughput. In this work, we show how to strategically place buffers into a dataflow circuit to optimize its performance. Our approach extracts a set of choice-free critical loops from arbitrary dataflow circuits and relies on the theory of marked graphs to optimize the buffer placement and sizing. We demonstrate the performance benefits of our approach on a set of dataflow circuits obtained from imperative code. Lana Josipovic, Shabnam Sheikhha, Andrea Guerrieri, Paolo Ienne, Jordi Cortadella |
FPGA | 1 |
| 2019 | Speculative Dataflow CircuitsabstractWith FPGAs facing broader application domains, the conversion of imperative languages into dataflow circuits has been recently revamped as a way to overcome the conservatism of statically scheduled high-level synthesis. Apart from the ability to extract parallelism in irregular and control-dominated applications, dynamic scheduling opens a door to speculative execution, one of the most powerful ideas in computer architecture. Speculation allows executing certain operations before it is known whether they are correct or required: it can significantly increase fine-grain parallelism in loops where the condition takes many cycles to compute; it can also increase the performance of circuits limited by potential dependencies by assuming independence early on and by reverting to the correct execution if the prediction was wrong. In this work, we detail our methodology to enable tentative and reversible execution in dynamically scheduled dataflow circuits. We create a generic framework for handling speculation in dataflow circuits and show that our approach can achieve significant performance improvements over traditional circuit generation techniques. Lana Josipovic, Andrea Guerrieri, Paolo Ienne |
FPGA | 1 |
| 2018 | Dynamically Scheduled High-level SynthesisabstractHigh-level synthesis (HLS) tools almost universally generate statically scheduled datapaths. Static scheduling implies that circuits out of HLS tools have a hard time exploiting parallelism in code with potential memory dependencies, with control-dependent dependencies in inner loops, or where performance is limited by long latency control decisions. The situation is essentially the same as in computer architecture between Very-Long Instruction Word (VLIW) processors and dynamically scheduled superscalar processors; the former display the best performance per cost in highly regular embedded applications, but general purpose, irregular, and control-dominated computing tasks require the runtime flexibility of dynamic scheduling. In this work, we show that high-level synthesis of dynamically scheduled circuits is perfectly feasible by describing the implementation of a prototype synthesizer which generates a particular form of latency-insensitive synchronous circuits. Compared to a commercial HLS tool, the result is a different trade-off between performance and circuit complexity, much as superscalar processors represent a different trade-off compared to VLIW processors: in demanding applications, the performance is very significantly improved at an affordable cost. We here demonstrate only the first steps towards more performant high-level synthesis tools adapted to emerging FPGA applications and the demands of computing in broader application domains. Lana Josipovic, Radhika Ghosal, Paolo Ienne |
FPGA | 1 |
| 2017 | An Out-of-Order Load-Store Queue for Spatial ComputingabstractThe efficiency of spatial computing depends on the ability to achieve maximal parallelism. This needs memory interfaces that can correctly handle memory accesses arriving in arbitrary order while still respecting data dependencies and ensuring appropriate ordering for semantic correctness. However, a typical memory interface for out-of-order processors (i.e., a load-store queue) cannot immediately fulfill these requirements: a different allocation policy is needed to achieve out-of-order execution in a spatial system. We show a practical way to organize the allocation for an out-of-order load-store queue for spatial computing by dynamically allocating groups of memory accesses, where the access order within the group is statically predetermined (for instance by a high-level synthesis tool). Lana Josipovic, Philip Brisk, Paolo Ienne |
FCCM | 1 |
| 2017 | An Out-of-Order Load-Store Queue for Spatial ComputingabstractThe efficiency of spatial computing depends on the ability to achieve maximal parallelism. This necessitates memory interfaces that can correctly handle memory accesses that arrive in arbitrary order while still respecting data dependencies and ensuring appropriate ordering for semantic correctness. However, a typical memory interface for out-of-order processors (i.e., a load-store queue) cannot immediately meet these requirements: a different allocation policy is needed to achieve out-of-order execution in spatial systems that naturally omit the notion of sequential program order, a fundamental piece of information for correct execution. We show a novel and practical way to organize the allocation for an out-of-order load-store queue for spatial computing. The main idea is to dynamically allocate groups of memory accesses (depending on the dynamic behavior of the application), where the access order within the group is statically predetermined (for instance by a high-level synthesis tool). We detail the construction of our load-store queue and demonstrate on a few practical cases its advantages over standard accelerator-memory interfaces. Lana Josipovic, Philip Brisk, Paolo Ienne |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2016 | Enriching C-based High-Level Synthesis with parallel pattern templatesabstractDespite the popularity of C-based High-Level Synthesis (HLS) tools, their generic input programming languages make it challenging for the designer to find the expression that will result in adequate hardware quality and performance. Moreover, the syntactic variance of the input description often causes the inability of the HLS tool to fully identify and benefit from the properties of the computations. In this work, we propose extending standard C-based HLS tools with the concept of computational patterns. In particular, we present a template-based hardware generation strategy which enables complete exploitation of the pattern properties to produce high-quality hardware modules. The parametric templates allow us to automatically scale the implementation to the resource and data-bandwidth constraints of the target device, independent from the analysis abilities of the HLS tool. To demonstrate the benefits of our approach, we generated hardware implementations for six applications which we composed using a set of computational patterns (i.e. map, zip and reduce), achieving 1.3× to 2.8× speed-up over a state-of-the-art commercial HLS tool. Lana Josipovic, Nithin George, Paolo Ienne |
FPT | 1 |