Robert Szafarczyk

dblp:340/4356 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
5since 2021 · last 2025
0009-0007-8883-1747ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Compiler Support for Speculation in Decoupled Access/Execute Architectures
abstract
Irregular codes are bottlenecked by memory and communication latency. Decoupled access/execute (DAE) is a common technique to tackle this problem. It relies on the compiler to separate memory address generation from the rest of the program, however, such a separation is not always possible due to control and data dependencies between the access and execute slices, resulting in a loss of decoupling. In this paper, we present compiler support for speculation in DAE architectures that preserves decoupling in the face of control dependencies. We speculate memory requests in the access slice and poison mis-speculations in the execute slice without the need for replays or synchronization. Our transformation works on arbitrary, reducible control flow and is proven to preserve sequential consistency. We show that our approach applies to a wide range of architectural work on CPU/GPU prefetchers, CGRAs, and accelerators, enabling DAE on a wider range of codes than before.
Robert Szafarczyk, Syed Waqar Nabi, Wim Vanderbauwhede
CC1
2025 Dynamic Loop Fusion in High-Level Synthesis
abstract
Dynamic High-Level Synthesis (HLS) uses additional hardware to perform memory disambiguation at runtime, increasing loop throughput in irregular codes compared to static HLS. However, most irregular codes consist of multiple sibling loops, which currently have to be executed sequentially by all HLS tools. Static HLS performs loop fusion only on regular codes, while dynamic HLS relies on loops with dependencies to run to completion before the next loop starts.
Robert Szafarczyk, Syed Waqar Nabi, Wim Vanderbauwhede
FPGA1
2023 Dynamically Scheduled Memory Operations in Static High-Level Synthesis
abstract
Dynamically scheduled high-level synthesis (HLS) achieves higher throughput on codes with unpredictable memory accesses compared to static HLS. However, dynamic scheduling results in circuits that use more resources and have a slower critical path, even if only a small part of the circuit exhibits dynamic behavior. In this extended abstract, we propose to introduce dynamically scheduled memory operations into static HLS. Our goal is to reach the same throughput as dynamic HLS on codes with irregular memory accesses while achieving comparable resource usage and critical paths as static HLS.
Robert Szafarczyk, Syed Waqar Nabi, Wim Vanderbauwhede
FCCM1
2023 Compiler Discovered Dynamic Scheduling of Irregular Code in High-Level Synthesis
abstract
Dynamically scheduled high-level synthesis (HLS) achieves higher throughput than static HLS for codes with unpredictable memory accesses and control flow. However, excessive dataflow scheduling results in circuits that use more resources and have a slower critical path, even when only a part of the circuit exhibits dynamic behavior. Recent work has shown that marking parts of a dataflow circuit for static scheduling can save resources and improve performance (hybrid scheduling), but the dynamic part of the circuit still bottlenecks the critical path. We propose instead to selectively introduce dynamic scheduling into static HLS. This paper presents an algorithm for identifying code regions amenable to dynamic scheduling and shows a methodology for introducing dynamically scheduled basic blocks, loops, and memory operations into static HLS. Our algorithm is informed by modulo-scheduling and can be integrated into any modulo-scheduled HLS tool. On a set of ten benchmarks, we show that our approach achieves on average an up to 3.7× and 3× speedup against dynamic and hybrid scheduling, respectively, with an area overhead of 1.3× and frequency degradation of 0.74× when compared to static HLS.
Robert Szafarczyk, Syed Waqar Nabi, Wim Vanderbauwhede
FPL1
2022 Reducing FPGA Memory Footprint of Stencil Codes through Automatic Extraction of Memory Patterns
abstract
FPGAs are attractive for scientific high-performance computing due to their potential for high performance-per-Watt. Stencil codes in scientific applications are difficult to optimize on FPGAs, because of redundant, non-contiguous memory accesses to relatively low bandwidth DRAM. In this paper, we present an algorithm to aggressively reduce on-chip block RAM (BRAM) and off-chip DRAM utilisation of stencil codes running on FPGAs. The algorithm extracts memory accesses from computational pipelines and removes all redundant intermediate arrays, including those used for stencil buffering, by trading DRAM accesses for computation. The algorithm is based on rewrite-rules on a strict functional representation derived from Fortran code and generates provably correct, optimized code. Typical FPGA implementations store the stencil window in on-chip shift registers implemented in BRAMs; we use only DRAM and optimize the memory accesses instead. Our approach dramatically reduces BRAM usage so that the domain size is only limited by available DRAM. We report a drop of 78% and 18% in BRAM usage in 3-D and 2-D stencil codes compared to a manual implementation using shift registers while staying competitive in performance or even improving performance-per-Watt.
Robert Szafarczyk, Syed Waqar Nabi, Wim Vanderbauwhede
FPL1