Anna Drewes

dblp:248/8199 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2023
0000-0002-8322-6747ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2023 A Flexible and Scalable Reconfigurable FPGA Overlay Architecture for Data-Flow Processing
abstract
We present a flexible and scalable FPGA overlay architecture for data-flow applications. The overlay consists of a 2D grid of tiles consisting of compute units that can be exchanged at runtime. The overlay can deal with unbalanced data-flow graphs and is based on AXI-Stream to facilitate extensibility using IP and High-Level Synthesis. To support compute units that also require random access to the system memory, AXI4-ports can be enabled per tile at design-time. We implement a prototype of the overlay architecture tailored to the application domain of analytical query processing on a Xilinx Alveo U280 board. In this system, an overlay grid of$11\times 4$compute unit tiles occupies one-third of the available resources. with a design I/O throughput of$11\times 3.75\ \text{GB}/\mathrm{s}$. The overlay and HLS-based SIMD compute units provide full-throughput data processing, but are limited by the memory subsystem implemented with vendor IPs.
Anna Drewes, Vitalii Burtsev, Bala Gurumurthy, Martin Wilhelm, David Broneske, Gunter Saake, Thilo Pionteck
FCCM1
2023 A comprehensive modeling approach for the task mapping problem in heterogeneous systems with dataflow processing units
abstract
Summary We introduce a new model for the task mapping problem to aid in the systematic design of algorithms for heterogeneous systems including, but not limited to, CPUs, GPUs, and FPGAs. A special focus is set on the communication between the devices, its influence on parallel execution, as well as on device‐specific differences regarding parallelizability and streamability. We give a comprehensive description on how a given task mapping can be abstractly evaluated including mappings to dataflow‐based hardware accelerators. We show how this model can be utilized in different system design phases and present two novel mixed‐integer linear programs to demonstrate the usage of the model, showing significant improvements compared to pure CPU mapping for randomly generated task graphs. To the best of our knowledge, we present the first ILP for task mapping that considers pipelining effects when streaming tasks on an FPGA.
Martin Wilhelm, Hanna Geppert, Anna Drewes, Thilo Pionteck
Concurr. Comput. Pract. Exp.3
2022 EmuNoC: Hybrid Emulation for Fast and Flexible Network-on-Chip Prototyping on FPGAs
abstract
Networks-on-Chips (NoCs) recently became widely used, from multi-core CPUs to edge-AI accelerators. Emulation on FPGAs promises to accelerate their RTL modeling compared to slow simulations. However, realistic test stimuli are challenging to generate in hardware for diverse applications. In other words, both a fast and flexible design framework is required. The most promising solution is hybrid emulation, in which parts of the design are simulated in software, and the other parts are emulated in hardware. This paper proposes a novel hybrid emulation framework called EmuNoC. We introduce a clock-synchronization method and software-only packet generation that improves the emulation speed by 36.3 × to 79.3 × over state-of-the-art frameworks while retaining the flexibility of a pure-software interface for stimuli simulation. We also increased the area efficiency to model up to an NoC with 169 routers on a single FPGA, while previous frameworks only achieved 64 routers.
Yee Yang Tan, Felix Staudigl, Lukas Jünger 0001, Anna Drewes, Rainer Leupers, Jan Moritz Joseph
FPL4