EDBT 2026 Demo / reviewers in the wild / expert
Gabriel Rodriguez-Canal
dblp:306/6636
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
0009-0005-0511-3922ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lifting to Tensors when Compiling Scientific Computing Workloads for AI Engines
Nick Brown 0002, Gabriel Rodriguez-Canal |
CCGrid | 2 |
| 2026 | Towards Compiler-Driven Dynamic Partial Reconfiguration with MLIRabstractHigh-Level Synthesis (HLS) has democratised Field-Programmable Gate Array (FPGA) programming, yet Dynamic Partial Reconfiguration (DPR)—which enables runtime logic swapping for adaptive or oversized workloads—remains manual and expert-only. HiPR [1] adds limited compiler support but restricts modules to one-to-one region mappings without runtime management. MLIR-DPR introduces: (i) a dpr dialect in the Multi-Level Intermediate Representation (MLIR) infrastructure [2] for identifying mutually exclusive regions; (ii) automated interface synthesis, floorplanning, and multi-threaded scheduler generation; and (iii) demonstrated Software-Defined Radio (SDR), Design-Space Exploration (DSE), and virtual-area applications. Gabriel Rodriguez-Canal, Nick Brown 0002, Maurice Jamieson, Nicolas Bohm Agostini, Ankur Limaye, Vito Giovanni Castellana, Joseph B. Manzano, Antonino Tumeo |
FCCM | 1 |
| 2026 | Towards Scheduling of Pipelined Dataflow Graphs in MLIRabstractWe present an MLIR flow that partitions neural networks and schedules them as software-driven macro-dataflow pipelines for low-latency streaming on CPU–FPGA SoCs. A new dataflow dialect and token-based scheduler pipeline even cyclic graphs with external memory, overcoming HLS limits. On an AlphaData ADM-PA101 (Versal VM1802) we demonstrate low-latency streaming; to our knowledge this is the first HLS flow to pipeline cyclic NN graphs. Gabriel Rodriguez-Canal, Nicolas Bohm Agostini, Ankur Limaye, Vito Giovanni Castellana, Joseph B. Manzano, Antonino Tumeo, Maurice Jamieson, Nick Brown 0002 |
FPGA | 1 |
| 2025 | Seamless Acceleration of Fortran Intrinsics via AMD AI Engines
Nick Brown 0002, Gabriel Rodriguez-Canal |
FPGA | 2 |
| 2024 | A shared compilation stack for distributed-memory parallelism in stencil DSLsabstractDomain Specific Languages (DSLs) increase programmer productivity and provide high performance. Their targeted abstractions allow scientists to express problems at a high level, providing rich details that optimizing compilers can exploit to target current- and next-generation supercomputers. The convenience and performance of DSLs come with significant development and maintenance costs. The siloed design of DSL compilers and the resulting inability to benefit from shared infrastructure cause uncertainties around longevity and the adoption of DSLs at scale. By tailoring the broadly-adopted MLIR compiler framework to HPC, we bring the same synergies that the machine learning community already exploits across their DSLs (e.g. Tensorflow, PyTorch) to the finite-difference stencil HPC community. We introduce new HPC-specific abstractions for message passing targeting distributed stencil computations. We demonstrate the sharing of common components across three distinct HPC stencil-DSL compilers: Devito, PSyclone, and the Open Earth Compiler, showing that our framework generates high-performance executables based upon a shared compiler ecosystem. George Bisbas, Anton Lydike, Emilien Bauer, Nick Brown 0002, Mathieu Fehr, Lawrence Mitchell, Gabriel Rodriguez-Canal, Maurice Jamieson, Paul H. J. Kelly, Michel Steuwer, Tobias Grosser |
ASPLOS (3) | 7 |
| 2023 | Fortran High-Level Synthesis: Reducing the Barriers to Accelerating HPC Codes on FPGAsabstractIn recent years the use of FPGAs to accelerate scientific applications has grown, with numerous applications demonstrating the benefit of FPGAs for high performance workloads. However, whilst High Level Synthesis (HLS) has significantly lowered the barrier to entry in programming FPGAs by enabling programmers to use C++, a major challenge is that most often these codes are not originally written in C++. Instead, Fortran is the lingua franca of scientific computing and-so it requires a complex and time consuming initial step to convert into C++ even before considering the FPGA. In this paper we describe work enabling Fortran for AMD Xilinx FPGAs by connecting the LLVM Flang front end to AMD Xilinx's LLVM back end. This enables programmers to use Fortran as a first-class language for programming FPGAs, and as we demonstrate enjoy all the tuning and optimisation opportunities that HLS C++ provides. Furthermore, we demonstrate that certain language features of Fortran make it especially beneficial for programming FPGAs compared to C++. The result of this work is a lowering of the barrier to entry in using FPGAs for scientific computing, enabling programmers to leverage their existing codebase and language of choice on the FPGA directly. Gabriel Rodriguez-Canal, Nick Brown 0002, Timothy Dykes, Jessica R. Jones, Utz-Uwe Haus |
FPL | 1 |
| 2023 | Task-based preemptive scheduling on FPGAs leveraging partial reconfigurationabstractSummary Field‐programmable gate arrays (FPGAs) are an attractive type of accelerator for all‐purpose high performance computing computing systems due to the possibility of deploying tailored hardware on demand. However, the common tools for programming and operating FPGAs are still complex to use, especially in scenarios where diverse types of tasks should be dynamically executed. In this work, we present a programming abstraction with a simple interface that internally leverages high‐level synthesis, dynamic partial reconfiguration and synchronization mechanisms to use an FPGA as a multi‐tasking server with preemptive scheduling and priority queues. This leads to an improved use of the FPGA resources, allowing the execution of several different kernels concurrently and deploying the most urgent ones as fast as possible. The results of our experimental study show that our approach incurs only a 10 5% overhead in the worst case when using two reconfigurable regions, whilst providing a significant performance improvement of at least 24 21% over the traditional full reconfiguration approach. Gabriel Rodriguez-Canal, Nick Brown 0002, Yuri Torres, Arturo González-Escribano |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Efficient heterogeneous programming with FPGAs using the Controller model
Gabriel Rodriguez-Canal, Yuri Torres, Francisco J. Andujar, Arturo González-Escribano |
J. Supercomput. | 1 |