EDBT 2026 Demo / reviewers in the wild / expert
Dirk Koch
dblp:60/5993
· DBLP profile ↗
80ranked-venue papers
22as first author
20since 2021 · last 2025
0000-0002-2568-4432ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 78 · 22 first-author · 19 since 2021Software engineering, systems software and programming languages · 2Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Design Space Exploration of Fast RISC-V Processors for Scalable Kilo-Core FPGA SystemsabstractThis paper presents microarchitectural and physical design trade-offs as well as primitive mapping techniques for designing high-throughput area-efficient RISC-V soft-cores. These techniques facilitate the implementation of tiny, highly efficient RISC-V cores that can be scaled to thousands of instances on large FPGAs, enabling massively parallel overlays by pushing above$\mathbf{1. 5}$MIPS/LUT in a single-core setup and over$\mathbf{0. 8 1}$MIPS/LUT in a kilo-core configuration delivering$\mathbf{5 1 2, 0 0 0}$MIPS peak throughput. The proposed fully open-source architecture is analyzed in detail, with a focus on optimizing trade-offs between core pipeline depth, multithreading, clock speed, and configurable parameters, such as the use of Digital Signal Processing (DSP) blocks in the Arithmetic Logic Unit (ALU). The interplay between these parameters is carefully examined to demonstrate their impact on system scalability and throughput. A case study is presented to illustrate the implementation of a kilo-core system, showing how thousands of RISC-V cores can be deployed on a large FPGA platform while maintaining per-core efficiency. This work paves the way for scalable many-core designs on FPGAs (and overlays in general), as they can achieve both a compact area footprint, low power dissipation, and high operating clock speed. Riadh Ben Abdelhamid, Vladislav Válek, Kevin Klein 0005, Dirk Koch |
FPL | 4 |
| 2025 | Virtualization and Dynamic Reconfiguration of Custom Instruction Accelerators (CIA) in RISC-V Embedded SystemsabstractCustom hardware instructions are frequently used to accelerate software. However, the number of active custom instructions - and speedup opportunities - are limited in resourceconstrained systems, such as those using small embedded FPGAs (eFPGAs). Hence, it is desirable to be able to dynamically reconfigure this resource to allow an application to change the accelerator logic. While an existing method called CX Table is described the Draft CX Specification to manage accelerators, it is not suitable for lightweight embedded systems because (1) virtual memory is required to isolate processes, and (2) an implied table lookup is required when switching custom instruction sets. This paper presents CIA Direct, a lightweight alternative to CX Table. It moves the table into a privileged control register to avoid virtual memory and external memory accesses. It provides inter-process virtualization of accelerator state, but intra-process virtualization has performance limitations. This paper also describes CxBex, the first ASIC implementation of CIA Direct. It contains a hard RISC-V processor with an eFPGA to support multiple custom instruction accelerators under dynamic reconfiguration. Unique to CxBex, context switching keeps multiple independent state context intact in eFPGA block during reconfiguration. Intended for lightweight embedded systems, this work demonstrates that the CIA Direct approach is simpler, reduces CPU complexity, eliminates dependence on external memory, and manages dynamic reconfiguration using exceptions. Bea Healy, Brandon Freiberger, Jonas Kuenstler, King Lok Chung, Emil Cozac, Meinhard Kissich, Gennadiy Knis, Ron Sass, Dirk Koch, Jan Gray, Guy Lemieux |
FPL | 9 |
| 2025 | FPGAs with FABulous - Framework and ChipsabstractFABulous is an easy usable, yet powerful and complete eFPGA (embedded FPGA) Framework covering all aspects of an eFPGA ecosystem. FABulous eFPGAs had shown good area density in both standard-cell and custom cell flows and the framework allows various customizations, including the integration of custom primitives, I/O cells, or complex blocks like CPUs cores or ADCs. So far, more than 10 chips containing FABulous eFPGAs had been designed in technology nodes ranging from 28-180 nm by seven different universities. This demo shows how embedded FPGAs (eFPGAs) can be 1) specified, 2) verified and tested, 3) integrated into an ASIC, and 4) programmed in Verilog or VHDL using the all open FABulous eFPGA framework. Moreover, the demo will showcase two boards featuring FABulous FPGAs, including the first FABulous open-everything FPGA. Dirk Koch, Myrtle Shah, King Lok Chung, Jonas Kuenstler, Marcel Jung, Jakob Ternes, Asma Mohsin, Gennadiy Knis |
FPL | 1 |
| 2025 | PhySwin: An Efficient and Physically-Informed Foundation Model for Multispectral Earth ObservationabstractRecent progress on Remote Sensing Foundation Models (RSFMs) aims toward universal representations for Earth observation imagery. However, current efforts often scale up in size significantly without addressing efficiency constraints critical for real-world applications (e.g., onboard processing, rapid disaster response) or treat multispectral (MS) data as generic imagery, overlooking valuable physical priors. We introduce PhySwin, a foundation model for MS data that integrates physical priors with computational efficiency. PhySwin combines three innovations: (i) physics-informed pretraining objectives leveraging radiometric constraints to enhance feature learning; (ii) an efficient MixMAE formulation tailored to SwinV2 for low-FLOP, scalable pretraining; and (iii) token-efficient spectral embedding to retain spectral detail without increasing token counts. Pretrained on over 1M Sentinel-2 tiles, PhySwin achieves SOTA results (+1.32\% mIoU segmentation, +0.80\% F1 change detection) while reducing inference latency by up to 14.4$\times$ and computational complexity by up to 43.6$\times$ compared to ViT-based RSFMs. Chong Tang 0006, Joseph Powell, Dirk Koch, Robert Mullins 0001, Alex S. Weddell, Jagmohan Chauhan |
NeurIPS | 3 |
| 2024 | SPARKLE: A 1,024-Core/16,384-Thread Single FPGA Many-Core RISC-V Barrel Processor OverlayabstractSPARKLE (Scalable Parallel Architecture for RISC-V Kernel-Level Execution) is a RISC-V many-core architecture for scalable parallel software processing, currently running at 400MHz on a Xilinx VU9P FPGA, including a PCIe backplane, with 1,024 cores and 16,384 hardware threads. SPARKLE is highly flexible, allowing various configurations. Its core, BRISKI (Barrel RISC- V for Kilo-core Implementations), operates at 650+ MHz in standalone mode and is one of the fastest barrel processor softcores, supporting full user-mode RV32I and additional instructions such as LRISC atomic instructions and CSRRS for hart ID retrieval. Riadh Ben Abdelhamid, Vladislav Válek, Dirk Koch |
ASAP | 3 |
| 2024 | Ph.D. Project: ManAge: A Tool for Timing Characterization of FPGAsabstractRecent advancements in transistor scaling have intensified reliability concerns in semiconductor devices. Addressing these issues requires a thorough understanding of the electrical properties of resources within the chip. This paper presents ManAge, a tool tailored for FPGA characterization with sub-picosecond timing resolution. ManAge offers an automated workflow spanning from generating test schemes to FPGA testing and data processing. Bardia Babaei, Dirk Koch |
FCCM | 2 |
| 2023 | Efficient Resource Scheduling for Runtime Reconfigurable Systems on FPGAsabstractFPGAs have high computation capabilities that can overload the memory subsystem and harm system performance. This paper proposes runtime resource management techniques that consider system limitations, such as the available compute resources and the memory bandwidth. By orchestrating the acceleration tasks in a heterogeneous multitenant FPGA operating system, we enhance the system's performance. The proposed approaches isolate the parallel running acceleration services in time and memory space domains based on their memory access behavior to reduce DRAM organization and access pattern effects. The proposed schedulers are evaluated and implemented on a Zynq UltraScale+ FPGA board running both acceleration and software tasks. The evaluation shows (on average) improvements in memory read/write throughput (10%), task execution time (30%), and makespan time (29%) over an existing greedy FPGA scheduler. Shaden M. Alismail, Dirk Koch |
FPL | 2 |
| 2023 | FABulous Demo: Open Source FPGA on Sky130abstractOpen source hardware has gained significant momentum in recent years, in particular with projects including RISC-V and open PDKs. SoCs often want to include some amount of custom functionality, for increased flexibility and future-proofing, and FABulous provides a highly customisable, open FPGA fabric generator. This demo shows all the fabric functionality working on our first silicon taped out on a fully open PDK, Skywater 130nm with shuttle runs sponsored by Google. Myrtle Shah, Jakob Ternes, Dirk Koch |
FPL | 3 |
| 2022 | How to Shrink My FPGAs - Optimizing Tile Interfaces and the Configuration Logic in FABulous FPGA FabricsabstractCommercial FPGAs from major vendors are extensively optimized, and fabrics use many hand-crafted custom cells, including switch matrix multiplexers and configuration memory cells. The physical design optimizations commonly improve area, latency (=speed), and power consumption together. This paper is dedicated to improving the physical implementation of FPGA tiles and the configuration storage in SRAM FPGAs. This paper proposes to remap configuration bits and interface wires to implement tightly packed tiles. Using the FABulous FPGA framework, we show that our optimizations are virtually for free but can save over 20% in area and improve latency at the same time. We will evaluate our approach in different scenarios by changing the available metal layers or the requested channel capacity. Our optimizations consider all tiles and we propose a flow that resolves dependencies between the CLBs and other tiles. Moreover, we will show that frame-based reconfiguration is, in almost all cases, better than shift register configuration. King Lok Chung, Nguyen Dao, Jing Yu 0014, Dirk Koch |
FPGA | 4 |
| 2022 | Tunable Fine-grained Clock Phase-shifting for FPGAsabstractHigh-resolution phase shifters have important practical applications in PET scanners, time-to-digital converters, and characterizing of the FPGA resources. This paper presents a fine-grained clock phase-shifting technique based on the FPGAs' clock managers' dynamic phase shifting capability that is commonly available on all recent FPGAs. Our method allows adjusting the phase shift resolution in the sub-picosecond range independent of the operating frequency. Experiments carried out on a Xilinx UltraScale+ FPGA show that phase-shifting resolution can be adjusted down to 88 f s in these devices. To verify the performance of this method, we have deployed it in a delay characterization circuit to measure the FPGA's resources delays. The experiments show that we can measure path delays below 1 ns which is impossible in conventional frequency sweep-based methods and we reach a much finer time resolution. Bardia Babaei, Dirk Koch |
FPL | 2 |
| 2022 | Precise Characterizing of FPGAs in Production SystemsabstractThe deployment of FPGAs in cloud data centers has entailed new security concerns. Although several defensive mechanisms are employed to detect and prevent malicious designs, a health monitoring tool can warn cloud service providers about the failure of implemented defensive fences. This PhD project aims to monitor the health status of internal FPGA resources by performing a precise timing characterization. Bardia Babaei, Dirk Koch |
FPL | 2 |
| 2022 | FPL Demo: Runtime Stream Processing with Resource-Elastic Pipelines on FPGAsabstractFPGAs are efficient at dataflow applications, as demonstrated in various application domains, including machine learning, communication, and image processing. In this demo, we accelerate database management operations transparently to the user by stitching together partially reconfigurable stream processing modules that implement database operators. Our runtime system orchestrates this, which builds custom pipelines according to runtime conditions. This demo will showcase an acceleration of SQL queries using our dynamic stream processing system running on a ZCU102 FPGA board. Kaspar Matas, Kristiyan Manev, Joseph Powell, Dirk Koch |
FPL | 4 |
| 2022 | FPL Demo: FPGA Bitstream Virus ScanningabstractThe expansion of the FPGA into complex market sectors imposes new demands on the security model of the devices. This demonstration shows off a series of tools developed to decode and scan the contents of a bitstream for malicious designs. Joseph Powell, Kaspar Matas, Kristiyan Manev, Dirk Koch |
FPL | 4 |
| 2022 | byteman: A Bitstream Manipulation FrameworkabstractFrom better resource pooling for FPGA cloud providers to building dynamic execution pipelines at runtime, the capabilities of partial reconfiguration (PR) are waiting to be fully explored. However, the community still fails to materialize PR at scale, and FPGAs are only used as updatable ASICs, hence, omitting the opportunities offered by dynamically reconfiguring FPGAs at runtime. This work proposes a resourceful FPGA bitstream manipulation framework. The proposed tool provides means for parsing, modification, and generation of bitstream files, and it has been open-sourced and demonstrated in a working system. As a distinguished feature, it supports multidie FPGAs (among the 106 Xilinx 7 Series, UltraScale, and UltraScale+ devices), and enables datacenter FPGAs to be used for relocatable PR. Using the versatile tool's built-in (dis)assembler allows for manual bitstream manipulations. Bundled with an efficient bitstream manipulation core, the efficacy is demonstrated by two case studies where we observe 58 - 377x higher bitstream merging throughput than a current state-of-art tool. Kristiyan Manev, Joseph Powell, Kaspar Matas, Dirk Koch |
FPT | 4 |
| 2022 | Automated Generation and Orchestration of Stream Processing Pipelines on FPGAsabstractFPGAs have demonstrated substantial performance and energy efficiency advantages for workloads that fit a stream processing model with direct module-to-module communication. However, when the dataflow processing system is required to adapt to runtime conditions, current static acceleration solutions are limited. To better use FPGAs in dynamic scenarios, this paper proposes using partial reconfiguration to stitch together different physically implemented operator modules on-the-fly. Rather than using designated module slots, our system places all modules and routing wires into a shared region with more placement options to minimize fragmentation. Furthermore, we use a module library that provides different resource and performance trade-offs for faster execution while considering the configuration cost. Our system finds the optimal set of modules while scheduling multiple acceleration requests and managing all constraints transparently to the end-user. We demonstrate that the middleware is fast enough to compose accelerator pipelines at runtime with end-to- end execution times equal to hand-crafted static systems when processing small datasets. For large datasets, we found up to 7.2 x faster execution over static systems when using our runtime methods. We exemplified our approach for database acceleration, where the whole dynamic FPGA acceleration is inferred by directly executing SQL queries. Kaspar Matas, Kristiyan Manev, Joseph Powell, Dirk Koch |
FPT | 4 |
| 2022 | The Future of FPGA Acceleration in Datacenters and the CloudabstractIn this article, we survey existing academic and commercial efforts to provide Field-Programmable Gate Array (FPGA) acceleration in datacenters and the cloud. The goal is a critical review of existing systems and a discussion of their evolution from single workstations with PCI-attached FPGAs in the early days of reconfigurable computing to the integration of FPGA farms in large-scale computing infrastructures. From the lessons learned, we discuss the future of FPGAs in datacenters and the cloud and assess the challenges likely to be encountered along the way. The article explores current architectures and discusses scalability and abstractions supported by operating systems, middleware, and virtualization. Hardware and software security becomes critical when infrastructure is shared among tenants with disparate backgrounds. We review the vulnerabilities of current systems and possible attack scenarios and discuss mitigation strategies, some of which impact FPGA architecture and technology. The viability of these architectures for popular applications is reviewed, with a particular focus on deep learning and scientific computing. This work draws from workshop discussions, panel sessions including the participation of experts in the reconfigurable computing field, and private discussions among these experts. These interactions have harmonized the terminology, taxonomy, and the important topics covered in this manuscript. Christophe Bobda, Joel Mandebi, Paul Chow, Mohammad Ewais, Naif Tarafdar, Juan Camilo Vega, Kenneth Eguro, Dirk Koch, Suranga Handagala, Miriam Leeser, Martin C. Herbordt, Hafsah Shahzad, H. Peter Hofstee, Burkhard Ringlein, Jakub Szefer, Ahmed Sanaullah, Russell Tessier |
ACM Trans. Reconfigurable Technol. Syst. | 8 |
| 2021 | Trusted Configuration in Cloud FPGAsabstractIn this paper we tackle the open paradoxical challenge of FPGA-accelerated cloud computing: On one hand, clients aim to secure their Intellectual Property (IP) by encrypting their configuration bitstreams prior to uploading them to the cloud. On the other hand, cloud service providers disallow the use of encrypted bitstreams to mitigate rogue configurations from damaging or disabling the FPGA. Instead, cloud providers require a verifiable check on the hardware design that is intended to run on a cloud FPGA at the netlist-level before generating the bitstream and loading it onto the FPGA, therefore, contradicting the IP protection requirement of clients. Currently, there exist no practical solution that can adequately address this challenge.We present the first practical solution that, under reasonable trust assumptions, satisfies the IP protection requirement of the client and provides a bitstream sanity check to the cloud provider. Our proof-of-concept implementation uses existing tools and commodity hardware. It is based on a trusted FPGA shell that utilizes less than 1% of the FPGA resources on a Xilinx VCU118 evaluation board, and an Intel SGX machine running the design checks on the client bitstream. Shaza Zeitouni, Jo Vliegen, Tommaso Frassetto, Dirk Koch, Ahmad-Reza Sadeghi, Nele Mentens |
FCCM | 4 |
| 2021 | FABulous: An Embedded FPGA FrameworkabstractAt the end of CMOS-scaling, the role of architecture design is increasingly gaining importance. Supporting this trend, customizable embedded FPGAs are an ingredient in ASIC architectures to provide the advantages of reconfigurable hardware exactly where and how it is most beneficial. To enable this, we are introducing the FABulous embedded open-source FPGA framework. FABulous is designed to fulfill the objectives of ease of use, maximum portability to different process nodes, good control for customization, and delivering good area, power, and performance characteristics of the generated FPGA fabrics. The framework provides templates for logic, arithmetic, memory, and I/O blocks that can be easily stitched together, whilst enabling users to add their own fully customized blocks and primitives. The FABulous ecosystem generates the embedded FPGA fabric for chip fabrication, integrates Yosys, ABC, VPR and nextpnr as FPGA CAD tools, deals with the bitstream generation and after fabrication tests. Additionally, we provide an emulation path for system development. FABulous was demonstrated for an ASIC integrating a RISC-V core with an embedded FPGA fabric for custom instruction set extensions using a TSMC 180nm process and an open-source 45nm process node. Dirk Koch, Nguyen Dao, Bea Healy, Jing Yu 0014, Andrew Attwood |
FPGA | 1 |
| 2021 | The FABulous Open eFPGA Ecosystem in Action - From Specifications to Chips to Running BitsteamsabstractThis demonstration shows the steps a designer has to take to specify and implement a chip with an embedded FPGA (eFPGA) using the FABulous open-source toolchain. We also show how the architecture graph is automatically generated for the open-source FPGA CAD tools (Yosys, ABC, nextpnr) to compile Verilog all the way to a bitstream. Ultimately, we demonstrate such bitstreams running on our FlexBex chip, which integrates an Ibex RISC-V core from lowRISC together with a FABulous eFPGA. The system supports multiple partially reconfigurable regions for hosting reconfigurable instruction set extensions, and the fabric provides logic, DSP, and memory slices. Jing Yu 0014, Andrew Attwood, Nguyen Dao, Dirk Koch |
FPL | 4 |
| 2021 | Memristor-Based Pass Gate for FPGA Programmable Routing SwitchabstractThe advancement of memristor technologies has recently attracted huge interest in exploiting their superior potential properties for enabling hybrid memristor-CMOS systems. This work presents a memristor-based Pass Gate - mPG, as a primitive cell as well as its deployment for implementing programmable routing switches targeting FPGAs. The mPG, consists of a transistor and output buffer(s) and can be used to discriminate resistance states for both binary and multistate memristor technologies without additional circuitry. The proposed routing structure eliminates leakage current and avoids degrading memristor's characteristics due to voltage drops, which are essential factors for building reliable large-scale digital systems like FPGAs. Simulation results of mPG-based switches show that the gate can be deployed for a wide range of memristor resistance with a switching delay in the subnanosecond range. Physical implementations of the proposed primitive demonstrate savings of about 50% of the design in comparison to standard CMOS designs. Nguyen Cong Dao, Dirk Koch |
ISCAS | 2 |
| 2020 | Message from the Conference Chairs - ASAP 2020abstractPresents the introductory welcome message from the conference proceedings. May include the conference officers' congratulations to all involved with the conference event and publication of the proceedings record. Dirk Koch, Frank Hannig, Javier Navaridas |
ASAP | 1 |
| 2020 | Power-hammering through Glitch Amplification - Attacks and MitigationabstractRecent work on FPGA hardware security showed a substantial potential risk through power-hammering, which uses high switching activity in order to create excessive dynamic power loads. Virtually all present power-hammering attack scenarios are based on some kind of ring oscillators for which mitigation strategies exist. In this paper, we use a different strategy to create excessive dynamic power consumption: glitch amplification. By carefully designing XOR trees, fast switching wires can be implemented that, while driving high fan-out nets, can draw enough power to crash an FPGA. In addition to the attack (which is crashing an Ultra96 board), we will present a scanner for detecting malicious glitch amplifying FPGA designs. Kaspar Matas, Tuan Minh La, Khoa Dang Pham, Dirk Koch |
FCCM | 4 |
| 2020 | Invited Tutorial: FPGA Hardware Security for Datacenters and BeyondabstractSince FPGAs are now available in datacenters to accelerate applications, providing FPGA hardware security is a high priority. FPGA security is becoming more serious with the transition to FPGA-as-a-Service where users can upload their own bitstreams. Full control over FPGA hardware through the bitstream enables attacks to weaken an FPGA-based system. These include physically damaging the FPGA equipment and leaking of sensitive information such as the secret keys of crypto algorithms. While there is no known attacks in the commercial settings so far, it is not so much a question of if but more of when? The tutorial will show concrete attacks applicable on datacenter FPGAs. The goal of this tutorial is to prepare the FPGA community to impending security issues in order to pave way for a proactive security. First, we will give a tour through the FPGA hardware security jungle surveying practical attacks and potential threats. We will reinforce this with live demos of denial of service attacks. Less than 10% of the logic resources on an FPGA can draw enough dynamic power to crash a datacenter FPGA card. In the second part of the tutorial, we will show different mitigations that are either vendor supported or proposed by the academic community. In summary, the tutorial will communicate that while FPGA hardware security is complicated to bring about, there are acceptable solutions for known FPGA security problems. Kaspar Matas, Tuan La, Nikola Grunchevski, Khoa Dang Pham, Dirk Koch |
FPGA | 5 |
| 2020 | Securing FPGA Accelerators at the Electrical Level for Multi-tenant PlatformsabstractAs FPGAs are now offered on the cloud, this exposes many potential security issues. This PhD project investigates current security issues and challenges when deploying FPGAs in the cloud as well as using FPGAs in a multi-tenancy scenario. By addressing practical threats, and most importantly, proposing feasible countermeasures, this paper shows preliminary results on protecting FPGAs for multi-tenant scenarios. Tuan La, Kaspar Matas, Khoa Dang Pham, Dirk Koch |
FPL | 4 |
| 2020 | Demo: A Closer Look at Malicious BitstreamsabstractAs FPGAs are now offered on the cloud and widely used in critical applications (infrastructure, military, medical), this exposes many potential security issues in which attackers can deploy attacks remotely. This demo will take a closer look at malicious bitstreams and demo our FPGA bitstream virus scanner FPGADefender that can scan for signatures relating to malicious circuits and many forms of bitstream manipulations. Additionally, we show a deny-of-service power-hammering attack, which can serve as a template for hardware Trojans. Tuan La, Kaspar Matas, Joseph Powell, Khoa Dang Pham, Dirk Koch |
FPL | 5 |
| 2020 | Resource Elastic Database AccelerationabstractDatabase sizes are growing faster than the processing power in the post-Moore era due to the advent of big data applications, which make hardware acceleration mandatory. We propose dynamic stream processing using partial reconfiguration to provide a high-performance solution to runtime-known problems. This work researches the benefits of applying resource elastic techniques when building the execution pipeline. Kristiyan Manev, Dirk Koch |
FPL | 2 |
| 2020 | Transparent Integration of a Dynamic FPGA Database Acceleration SystemabstractA substantial amount of recent work has been conducted on accelerating different database operators or database management systems (DBMS) as a whole, both in proprietary and open-source services. The missing piece of work in this field is to transparently accelerate a widely-used open-source database on an FPGA without substantial changes to the core database code. This PhD project aims to use a partially reconfigurable stream processing architecture to compose different database operators into a pipeline orchestrated automatically by the database query optimizer and query planner. This way, the database users do not have to change any of the existing database management system code while allowing a runtime system to optimize queries with information known only during runtime. Kaspar Matas, Dirk Koch |
FPL | 2 |
| 2020 | A Self-Compilation Flow Demo on FOS - The FPGA Operating SystemabstractWith the introduction of Zynq UltraScale+ MPSoCs equipped with powerful 64-bit ARM CPUs along with a 16nm UltraScale+ FPGA fabric in the same die, we can now build full hardware-software programmable systems and configure the FPGA with accelerators, as needed. Traditionally, accelerators had been developed offline using powerful servers running heavy lifting CAD toolchains. In this demo, we show a self-compilation system supporting a user-friendly Jupyter Notebook GUI and multi-tenancy use of the FPGA for educational purposes. From a user perspective, this system compiles accelerators at run-time directly on the ARM CPU without any involvement of the vendor tools. The final bitstreams then execute on the FPGA fabric using PR. Khoa Dang Pham, Anuj Vaishnav, Joseph Powell, Dirk Koch |
FPL | 4 |
| 2020 | FPGADefender: Malicious Self-oscillator Scanning for Xilinx UltraScale + FPGAsabstractSharing configuration bitstreams rather than netlists is a very desirable feature to protect IP or to share IP without longer CAD tool processing times. Furthermore, an increasing number of systems could hugely benefit from serving multiple users on the same FPGA, for example, for resource pooling in cloud infrastructures. This article researches the threat that a malicious application can impose on an FPGA-based system in a multi-tenancy scenario from a hardware security point of view. In particular, this article evaluates the risk systematically for FPGA power-hammering through short-circuits and self-oscillating circuits, which potentially may cause harm to a system. This risk includes implementing, tuning, and evaluating all FPGA self-oscillators known from the literature but also developing a large number of new power-hammering designs that have not been considered before. Our experiments demonstrate that malicious circuits can be tuned to the point that just 3% of the logic available on an Ultra96 FPGA board can draw the power budget of the entire FPGA board. This fact suggests a waste power potential for datacenter FPGAs in the range of kilowatts. In addition to carefully analyzing FPGA hardware security threats, we present the FPGA virus scanner FPGAD efender , which can detect (possibly) any self-oscillating FPGA circuit, as well as detecting short-circuits, high fanout nets, and a tapping onto signals outside the scope of a module for protecting data center FPGAs, such as Xilinx UltraScale+ devices at the bitstream level. Tuan Minh La, Kaspar Matas, Nikola Grunchevski, Khoa Dang Pham, Dirk Koch |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2020 | FOS: A Modular FPGA Operating System for Dynamic WorkloadsabstractWith FPGAs now being deployed in the cloud and at the edge, there is a need for scalable design methods that can incorporate the heterogeneity present in the hardware and software components of FPGA systems. Moreover, these FPGA systems need to be maintainable and adaptable to changing workloads while improving accessibility for the application developers. However, current FPGA systems fail to achieve modularity and support for multi-tenancy due to dependencies between system components and the lack of standardised abstraction layers. To solve this, we introduce a modular FPGA operating system – FOS, which adopts a modular FPGA development flow to allow each system component to be changed and be agnostic to the heterogeneity of EDA tool versions, hardware and software layers. Further, to dynamically maximise the utilisation transparently from the users, FOS employs resource-elastic scheduling to arbitrate the FPGA resources in both time and spatial domain for any type of accelerators. Our evaluation on different FPGA boards shows that FOS can provide performance improvements in both single-tenant and multi-tenant environments while substantially reducing the development time and, at the same time, improving flexibility. Anuj Vaishnav, Khoa Dang Pham, Joseph Powell, Dirk Koch |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2019 | End-to-end Dynamic Stream Processing on Maxeler HLS PlatformsabstractHigh Level Synthesis (HLS) tools are enabling non hardware experts to implement applications and algorithms on FPGAs. However, the majority of stream processing application that are currently developed through HLS are implemented statically and they are not exploiting the benefits enabled by partial reconfiguration. In this paper, we propose a generic approach for implementing and using partial reconfiguration through an HLS design flow for Maxeler platforms. Our flow extracts HLS generated HDL code from the Maxeler compilation process in order to implement a static FPGA infrastructure as well as run-time reconfigurable stream processing modules. As a distinct feature, our infrastructure can accommodate multiple partial modules in a pipeline daisy-chained manner, which aligns directly to Maxeler's dataflow programming paradigm. This allows the decomposition of complicated problems into basic building blocks that can be easily stitched together. The basic building blocks are entirely developed using Maxeler's MaxJ language. The benefits of the proposed flow are demonstrated by a case study of a dynamically reconfigurable video processing pipeline delivering 6.4GB/s throughput. Charalampos Kritikakis, Dirk Koch |
ASAP | 2 |
| 2019 | EFCAD - An Embedded FPGA CAD Tool Flow for Enabling On-chip Self-CompilationabstractThis paper combines a chain of academic tools to form an FPGA compilation flow for building partially reconfigurable modules on lightweight embedded platforms. Our flow - EFCAD - supports the entire stack from RTL (Verilog) to (partial) bitstream, and we demonstrate early results from the onchip ARM processor of, and targeting, the latest 16nm generation of a Zynq UltraScale+ MPSoC device. With this, we complement Xilinx's PYNQ initiative to not only facilitate System-on-Chip research and education entirely within an embedded system, but also to allow building new and specialising existing customcomputing accelerators without needing access to a workstation. Khoa Dang Pham, Malte Vesper, Dirk Koch, Eddie Hung |
FCCM | 3 |
| 2019 | The FOS (FPGA Operating System) DemoabstractWith the introduction of Zynq FPGAs that provide an ARM SoC with an attached FPGA fabric, it is possible to build complex software-centric systems that are software and hardware programmable. To harness the full potential of this approach, we developed FOS an FPGA Operating System which is built on open-source FPGA community and Xilinx vendor components. A distinct feature shown in this demo is a heterogeneous resource elastic scheduler that can dynamically and automatically adjust the allocation of tasks to hardware and software resources with respect to the present load scenario. We will also show the FOS ecosystem that allows easily implementing relocatable partially reconfigurable modules directly from RTL or HLS. Anuj Vaishnav, Khoa Dang Pham, Kristiyan Manev, Dirk Koch |
FPL | 4 |
| 2018 | A Soft Dual-Processor System with a Partially Run-Time Reconfigurable Shared 128-Bit SIMD EngineabstractIn this work, we present a soft dual-processor system that, as a distinctive feature, seamlessly integrates a partially run-time reconfigurable 128-bit SIMD engine. Importantly, the SIMD engine is tightly coupled to both scalar CPUs and it is shared amongst them with the purpose of drastically improving overall area utilization. We show that the proposed SIMD engine increases performance-per-area and that it can be used to substantially accelerate time consuming kernels for a set of media applications. Jose Raul Garcia Ordaz, Dirk Koch |
ASAP | 2 |
| 2018 | A Survey on FPGA VirtualizationabstractFPGA accelerators are being applied in various types of systems ranging from embedded systems to cloud computing for their high performance and energy efficiency. Given the scale of deployment, there is a need for efficient application development, resource management, and scalable systems, which make FPGA virtualization extremely important. Consequently, FPGA virtualization methods and hardware infrastructures have frequently been proposed in both academia and industry for addressing multi-tenancy execution, multi-FPGA acceleration, flexibility, resource management and security. In this survey, we identify and classify the various techniques and approaches into three main categories: 1)Resource level, 2)Node level, and 3)Multi-node level. In addition, we identify current trends and developments and highlight important future directions for FPGA virtualization which require further work. Anuj Vaishnav, Khoa Dang Pham, Dirk Koch |
FPL | 3 |
| 2018 | Resource Elastic Virtualization for FPGAs Using OpenCLabstractFPGAs are rising in popularity for acceleration in all kinds of systems. However, even in cloud environments, FPGA devices are typically still used exclusively by one application only. To overcome this, and as an approach to manage FPGA resources with OS functionality, this paper introduces the concept of resource elastic virtualization which allows shrinking and growing of accelerators in the spatial domain with the help of partial reconfiguration. With this, we can serve multiple applications simultaneously on the same FPGA and optimize the resource utilization and consequently the overall system performance. We demonstrate how an implementation of resource elasticity can be realized for OpenCL accelerators along with how it can achieve 2.3x better FPGA utilization and 49% better performance on average while simultaneously lowering waiting time for tasks. Anuj Vaishnav, Khoa Dang Pham, Dirk Koch, Jim D. Garside |
FPL | 3 |
| 2018 | Large Utility Sorting on FPGAsabstractThis paper presents a merge sorter able of merging thousands of streams in a single run where the logic cost scales logarithmic with the number of streams merged. Moreover, we apply several performance tuning techniques, including speculative execution, deep pipelining and optimized communication schemes between processing elements. An end-to-end case study utilizing a Xilinx VC709 board merges 2048 sequences of the Graysort benchmark between two DRAMs at 9.5GB/s or 1024 sequences at 10.3GB/s effective throughput. Kristiyan Manev, Dirk Koch |
FPT | 2 |
| 2018 | Live Migration for OpenCL FPGA AcceleratorsabstractFPGAs are currently being deployed at a large scale across data-centres for various applications because of their performance and power benefits. In particular, cloud service operators are now offering FPGAs as a Service. However, to completely integrate FPGAs in a data-centre environment like standard software systems, support for fault tolerance and task migration is essential. In this paper, we propose a live migration technique for FPGA accelerators to provide support for fault tolerance, system maintenance, and resource management. Our technique allows migration of OpenCL accelerators not only within a single FPGA but also across FPGAs with zero downtime. It achieves this by overlapping the computation with datamovements transparently from the user for OpenCL kernels. Moreover, distributed check-pointing mechanisms can be employed to recover from unknown faults with minimal loss of completed work. Altogether it enables system updates such as changing the static FPGA configuration or upgrading the OS without an interruption of service. Anuj Vaishnav, Khoa Dang Pham, Dirk Koch |
FPT | 3 |
| 2017 | BITMAN: A tool and API for FPGA bitstream manipulationsabstractTo fully support the partial reconfiguration capabilities of FPGAs, this paper introduces the tool and API BitMan for generating and manipulating configuration bitstreams. Bit-Man supports recent Xilinx FPGAs that can be used by the ISE and Vivado tool suites of the FPGA vendor Xilinx, including latest Virtex-6, 7 Series, UltraScale and UltraScale− series FPGAs. The functionality includes high-level commands such as cutting out regions of a bitstream and placing or relocating modules on an FPGA as well as low-level commands for modifying primitives and for routing clock networks or rerouting signal connections at run-time. All this is possible without the vendor CAD tools for allowing BitMan to be used even with embedded CPUs. The paper describes the capabilities, API and performance evaluation of BitMan. Khoa Dang Pham, Edson Lemos Horta, Dirk Koch |
DATE | 3 |
| 2017 | Asynchronous interface FIFO design on FPGA for high-throughput NRZ synchronisationabstractNetworks-on-chip (NoCs) have become a new chip design paradigm as the size of transistors continues to shrink. Globally-asynchronous locally-synchronous (GALS) on-chip networks are proposed for solving issues such as large clock tree distribution and signal delay variations. More interestingly, for the GALS networks using m-of-n delay-insensitive interconnect, the asynchronous interconnect not only can be used for on-chip interconnection, but also provides a simple, direct and power-saving solution for off-chip interconnection. This paper presents an asynchronous interface FIFO design to improve throughput over asynchronous inter-chip links using 2-of-7 Non-Return-to-Zero (NRZ) encoding in an existing many-core system. The proposed design is suitable for implementation on commodity FPGAs without using the limited global clock buffer resources, but involves using the FPGA to implement asynchronous circuits. The interface FIFO is constructed from the transition detectors themselves rather than by employing a separate buffer in the more conventional fashion. The proposed solution has been demonstrated in an existing system and is suitable for adaptation to other asynchronous m-of-n NRZ coding protocols for high-throughput communication. Gengting Liu, Jim D. Garside, Steve Furber, Luis A. Plana, Dirk Koch |
FPL | 5 |
| 2017 | Making a case for an ARM Cortex-A9 CPU interlay replacing the NEON SIMD unitabstractAs an alternative of adding more and more instructions to CPU cores in order to address a wide range of applications, this paper examines to use a mixed grained CPU interlay fabric to provide reconfigurable instruction set extensions. In detail, we are examining to replace the hardened NEON SIMD unit of an ARM Cortex-A9 with an identical sized FPGA fabric. We show that by applying a set of optimizations, we are able to emulate original applications using NEON instructions at the same hardware cost and at very little performance drop by an interlay. Moreover we are demonstrating examples where special custom instructions running on a CPU-Interlay-hybrid are substantially outperforming the original hardened CPU-NEONsystem, hence making a strong case to embed reconfigurability as a beneficial feature in future processors. Jose Raul Garcia Ordaz, Dirk Koch |
FPL | 2 |
| 2017 | A security library for FPGA interlaysabstractMany CPU design houses have added dedicated support for cryptography in recent processor generations, including Intel, IBM, and ARM. While adding accelerators and/or dedicated instructions boosts performance on cryptography, we are investigating a different approach that is not adding extra silicon area: We study to replace the hardened NEON SIMD unit of an ARM Cortex-A9 with an identical sized FPGA fabric, called an interlay. This will be used for implementing cryptographic instructions in soft-logic. We show that this approach can outperform the hardened NEON by up to 7.7× on AES and provide functionality that is not available in the hardened ARM. Anuj Vaishnav, Jose Raul Garcia Ordaz, Dirk Koch |
FPL | 3 |
| 2016 | soft-NEON: A study on replacing the NEON engine of an ARM SoC with a reconfigurable fabricabstractPower is a limiting factor in the design of embedded processors. For this reason adding more instruction extensions is not a scalable option. To overcome this issue, we study the effects of replacing the NEON unit of an ARM SoC with an FPGA-like reconfigurable fabric. We measure the gap between the conventional hard-NEON and a soft-NEON implementation. We found that the soft-NEON has an overhead of 25.17× and 6.23× for area and latency, respectively. This overhead is reduced by exploiting the reconfigurability of the fabric by incorporating FPGA-specific optimization techniques. Moreover, we show that instead of implementing the pre-defined NEON instruction set, custom instructions can be loaded to the reconfigurable fabric by using a HLS compilation flow. With this approach performance gains of over 2.8× have been obtained for some kernels. Jose Raul Garcia Ordaz, Dirk Koch |
ASAP | 2 |
| 2016 | ECOSCALE: Reconfigurable computing and runtime system for future exascale systems
Iakovos Mavroidis, Ioannis Papaefstathiou, Luciano Lavagno, Dimitrios S. Nikolopoulos, Dirk Koch, John Goodacre, Ioannis Sourdis, Vassilis Papaefstathiou, Marcello Coppola, Manuel Palomino |
DATE | 5 |
| 2016 | Parallel Hardware Merge SorterabstractSorting has tremendous usage in the applications that handle massive amount of data. Existing techniques accelerate sorting using multiprocessors or GPGPUs where a data set is partitioned into disjunctive subsets to allow multiple sorting threads working in parallel. Hardware sorters implemented in FPGAs have the potential of providing high-speed and low-energy solutions but the partition algorithms used in software systems are so data dependent that they cannot be easily adopted. The speed of most current sequential sorters still hangs around 1 number/cycle. Recently a new hardware merge sorter broke this speed limit by merging a large number of sorted sequences at a speed proportional to the number of sequences. This paper significantly improves its area and speed scalability by allowing stalls and variable sorting rate. A 32-port parallel merge-tree that merges 32 sequences is implemented in a Virtex-7 FPGA. It merges sequences at an average rate of 31.05 number/cycle and reduces the total sorting time by 160 times compared with traditional sequential sorters. Wei Song 0002, Dirk Koch, Mikel Luján, Jim D. Garside |
FCCM | 2 |
| 2016 | JetStream: An open-source high-performance PCI Express 3 streaming library for FPGA-to-Host and FPGA-to-FPGA communicationabstractMany FPGA-based accelerators are constrained by the available resources and multi-FPGA solutions can be necessary for building more capable systems. Available PCIe solutions provide only FPGA-to-Host communication. In this paper we present JetStream, an open-source1modular PCIe 3 library, supporting not only fast FPGA-to-Host communication, but also allowing direct FPGA-to-FPGA communication which fully bypasses the memory subsystem. The direct mode saves memory bandwidth for multicast modes and permits to connect multiple FPGAs in various software defined topologies. We show the benefits of JetStream with a large FIR filter spanning four FPGA boards, achieving throughputs of up to 7.09 GB/s per link. Utilizing direct FPGA-to-FPGA transfers reduces the required memory bandwidth by up to 75%. Malte Vesper, Dirk Koch, Kizheppatt Vipin, Suhaib A. Fahmy |
FPL | 2 |
| 2016 | A partial reconfiguration controller for Altera Stratix V FPGAsabstractWith the introduction of the Stratix V family, the FPGA vendor Altera is now fully supporting partial reconfiguration in all their recent FPGA devices. A distinct feature in the Altera architecture is that reconfigurable regions can be arbitrarily defined which is possible by writing a configuration mask prior to writing the actual configuration data to the FPGA fabric. In this paper, we will present details and the flow for implementing partial reconfiguration using Altera FPGAs, as well as a study on configuration bitstream sizes and configuration speeds for various resource and bounding-box aspect ratio variants. The results are used to build a partial reconfiguration controller that is featuring a lightweight but effective bitstream decompression module for greatly improving configuration speed on a DE5-net board. Zhenzhong Xiao, Dirk Koch, Mikel Luján |
FPL | 2 |
| 2015 | Rapid Overlay Builder for Xilinx FPGAsabstractOverlays are emerging as useful design patterns for solving reconfigurable computing problems. Overlays consist of compiler-like tools and an architecture written in RTL, making it easier for users to quickly compile high-level languages into FPGAs. Despite a high degree of regularity and repetition present in most overlays, it takes a long time for FPGA tools to generate the configuration bit stream. This paper proposes a methodology called Rapid Overlay Builder, or ROB, that combines module relocation, module variants and an efficient form of "router less" module stitching that we call zipping. Our case study demonstrates up to 22 times speedup in compile-time over a regular Xilinx ISE compilation, while achieving higher clock speeds. By applying ROB, we anticipate that overlays can be implemented more quickly and with more consistent clock rates. Michael Xi Yue, Dirk Koch, Guy Lemieux |
FCCM | 2 |
| 2015 | Placing partially reconfigurable stream processing applications on FPGAsabstractFinding placement locations for modules on an FPGA in a limited amount of time is a crucial task that determines the efficiency of a dynamic partially reconfigurable system. In this work, we will define a placement method based on transforming the inherent two dimensional (2D) structure of the FPGA into a one dimensional string and employing string matching. Moreover, our model is suited to compute a module placement over multiple chained reconfigurable regions. Our algorithm is based on a hybrid approach consisting of an offline precompute phase at design-time which in turn is used to speed-up module placement at run-time. Nicolae Bogdan Grigore, Dirk Koch |
FPL | 2 |
| 2014 | Portable module relocation and bitstream compression for Xilinx FPGAsabstractThis paper presents a novel methodology for generating and compressing configuration bitstreams for modules that can be executed at different positions of an FPGA. The presented methodology for bitstream generation and compression does not need deep knowledge of the bitstream format and it is independent of the target (Xilinx) FPGA family. The approach consists of a design phase where partial bitstreams are decomposed into sequences of module dependent and module independent pieces of configuration data. At run-time, this data can then be recomposed for the individual placement positions by a special DMA configuration controller as one atomic operation without any further software interaction. Our experiments demonstrate that module relocation and fast partial reconfiguration can be implemented at low logic cost. Christian Beckhoff, Dirk Koch, Jim Tørresen |
FPL | 2 |
| 2014 | Hierarchical reconfiguration of FPGAsabstractPartial reconfiguration allows some applications to substantially save FPGA area by time sharing resources among multiple modules. In this paper, we push this approach further by introducing hierarchical reconfiguration where reconfigurable modules can have reconfigurable submodules. This is useful for complex systems where many modules have common parts or where modules can share components. For such systems, we show that the number of bitstreams and the bitstream storage requirements can be scaled down from a multiplicative to an additive behavior with respect to the number of modules and submodules. A case study consisting of different reconfigurable softcore CPUs and hierarchically reconfigurable custom instruction set extensions demonstrates a 18.7× lower bitstream storage requirement and up to 10× faster reconfiguration speed when using hierarchical reconfiguration instead of using conventional single-level module-based reconfiguration. Dirk Koch, Christian Beckhoff |
FPL | 1 |
| 2014 | Design Tools for Implementing Self-Aware and Fault-Tolerant Systems on FPGAsabstractTo fully exploit the capabilities of runtime reconfigurable FPGAs in self-aware systems, design tools are required that exceed the capabilities of present vendor design tools. Such tools must allow the implementation of scalable reconfigurable systems with various different partial modules that might be loaded to different positions of the device at runtime. This comprises several complex tasks, including floorplanning, communication architecture synthesis, physical constraints generation, physical implementation, and timing verification all the way down to the final bitstream generation. In this article, we present how our GoAhead framework helps in implementing self-aware systems on FPGAs with a minimum of user interaction. Christian Beckhoff, Dirk Koch, Jim Tørresen |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2013 | Building partial systems with GoAheadabstractGOAHEAD is a tool for easily building complex run-time reconfigurable systems. The tool provides sophisticated features like module relocation, hierarchical reconfiguration, or reusing modules among different systems. This demonstration shows 1) how reconfigurable systems can be built using GOAHEAD with only a few mouse clicks. In addition, 2) we will show how a partial module can be compiled all the way to the final bitstream running on an Atlys Spartan-6 FPGA board in one single batch job. The demonstrated system can simultaneously host up to 75 video overlay modules or 10 partially reconfigurable MIPS CPU systems. In the latter case, the CPUs feature reconfigurable custom instruction set extensions, hence demonstrating a hierarchically reconfigurable multi-core system. Christian Beckhoff, Alexander Wold, Anders Fritzell, Dirk Koch, Jim Tørresen |
FPL | 4 |
| 2013 | An efficient FPGA overlay for portable custom instruction set extensionsabstractCustom instruction set extensions can substantially boost performance of reconfigurable softcore CPUs. While this approach is commonly tailored to one specific FPGA system, we are presenting a fine-grained FPGA-like overlay architecture which can be implemented in the user logic of various FPGA families from different vendors. This allows the execution of a portable application consisting of a program binary and an overlay configuration in a completely heterogeneous environment. Furthermore, we are presenting different optimizations for dramatically reducing the implementation cost of the proposed overlay architecture. In particular, this includes the mapping of the overlay interconnection network directly into the switch fabric of the hosting FPGA. Our case study demonstrates an overhead reduction of an order of magnitude as compared to related approaches. Dirk Koch, Christian Beckhoff, Guy Lemieux |
FPL | 1 |
| 2013 | Remote FPGA design through eDiViDe - European Digital Virtual Design LababstractThe design and development of digital electronic systems is mainly performed by use of a hardware description language. To prepare students in electrical engineering for a career in hardware design many universities provide courses on VHDL. The traditional approach in teaching VHDL is mainly by means of textbook examples and simulation provided by software applications. These exercises are perceived as monotonous by the students and do not or only very slightly correspond with actual real-life applications based on FPGAs. Moreover, most real-life applications are too expensive to be equipped in student laboratories. To bridge the gap between a simulation-only environment and affordable real-life applications students should be provided access to remote real-life setups with a 24/7 availability and preferably shared between multiple institutes. The eDiViDe platform (European Digital Virtual Design Lab, http://www.edivide.eu), see Fig. 1, provides students with this unlimited and exciting access to FPGA based setups. Instead of theory-only courses and a quick basic lab, they can work their way through digital design courses testing their skills on real-life setups to trigger their interest. The platform hosts multiple FPGA setups at different European institutes. These setups are accessible through a web-based interface with video feedback. VHDL development is performed offline, given an entity and specific setup information. All further steps of the FPGA toolchain are performed on the platform. A reservation system takes care of the FPGA programming and student interaction with the setups. Similar initiatives provide stable solutions with educational support [1,2,3]. The eDiViDe platform differentiates with a distributed platform across several institutes and with the support for advanced setups. It is the result of a joint effort and easily expandable with additional setups at any location. At this moment following setups are available: greenhouse, stepper motor control, sea noise emulator, state machine workshop, Geffe generator, pong / game of life, traffic light control, MIPS CPU. This set will be extended with more advanced setups that include e.g. a partial reconfiguration workshop for audio/video filters, a side-channel analysis setup and a mars rover playfield. Besides promoting digital design education, the eDiViDe platform creates a channel to make the research activities in the contributing universities more visible. Industry could also benefit from this platform to promote their brand and products to soon to be engineers. Jochen Vandorpe, Jo Vliegen, Ruben Smeets, Nele Mentens, Milos Drutarovský, Michal Varchola, Kerstin Lemke-Rust, Paul-Gerhard Plöger, Peter Samarin, Dirk Koch, Yngve Hafting, Jim Tørresen |
FPL | 10 |
| 2013 | EasyPR - An easy usable open-source PR systemabstractIn this paper, we present an open source partial reconfiguration (PR) system which is designed for portability and usability serving as a reference for engineers and students interested in using the advanced reconfiguration capabilities available in Xilinx FPGAs. This includes design aspects such as floorplanning and interfacing PR modules as well as fast reconfiguration and online management. The system features relocatable modules which can even contain reconfigurable modules themselves, hence, implementing hierarchical PR. Dirk Koch, Christian Beckhoff, Alexander Wold, Jim Tørresen |
FPT | 1 |
| 2012 | Design techniques for increasing performance and resource utilization of reconfigurable soft CPUsabstractReconfigurable hardware allows application specific customization of soft microprocessors. Techniques such as removing unused instructions, software emulation of instructions, custom instruction set extensions, and run-time reconfigurable instructions have been suggested. However, the techniques have largely been studied separately from each other. The contribution of this paper is a classification method enabling integration of these techniques. This allows for generating an application specific microprocessor based system from a given program. The generated microprocessor is optimized with respect to performance per area. The improvement of our methodology is demonstrated for the CoreBench benchmark. The benefit of combining the removal of unused instructions (ISA subsetting) with software emulation of rarely used instructions is shown to increase performance while at the same time reducing resource requirements. Improvement in both area and performance is accomplished thorough simplifying the design allowing an increase in clock frequency for the synthesized soft CPU. Optimizing only by using custom instructions allowed a 12% increase in performance, but also increased resource usage by 6%. Software emulation combined with ISA subsetting allowed area savings of 7%, but only improved performance by 3%. By combining custom instructions, software emulation and ISA subsetting, we achieved an performance improvement of 15% while at the same time reducing resource requirements. Alexander Wold, Dirk Koch, Jim Tørresen |
DDECS | 2 |
| 2012 | Go Ahead: A Partial Reconfiguration FrameworkabstractExploiting the benefits of partial run-time reconfiguration requires efficient tools. In this paper, we introduce the tool Go Ahead that is able to implement run-time reconfigurable systems for all recent Xilinx FPGAs. This includes in particular support for low cost and low power Spartan-6 FPGAs. Go Ahead assists during floor planning and automates the constraint generation. It interacts with the Xilinx vendor tools and triggers the physical implementation phases all the way down to the final configuration bit streams. Go Ahead enables the building of flexible systems for integrating many reconfigurable modules very efficiently into a system. The tool targets (re)usability, portability to future devices, and migration paths among reconfigurable systems featuring different FPGAs or even FPGA families. Moreover, it provides a scripting interface and all features can be accessed remotely. Christian Beckhoff, Dirk Koch, Jim Tørresen |
FCCM | 2 |
| 2012 | Dynamic Defragmentation of Reconfigurable DevicesabstractWe propose a new method for defragmenting the module layout of a reconfigurable device, enabled by a novel approach for dealing with communication needs between relocated modules and with inhomogeneities found in commonly used FPGAs. Our method is based on dynamic relocation of module positions during runtime, with only very little reconfiguration overhead; the objective is to maximize the length of contiguous free space that is available for new modules. We describe a number of algorithmic aspects of good defragmentation, and present an optimization method based on tabu search. Experimental results indicate that we can improve the quality of module layout by roughly 50% over the static layout. Among other benefits, this improvement avoids unnecessary rejections of modules. Sándor P. Fekete, Tom Kamphans, Nils Schweer, Christopher Tessars, Jan van der Veen, Josef Angermeier, Dirk Koch, Jürgen Teich |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2011 | FPGASort: a high performance sorting architecture exploiting run-time reconfiguration on fpgas for large problem sortingabstractThis paper analyses different hardware sorting architectures in order to implement a highly scaleable sorter for solving huge problems at high performance up to the GB range in linear time complexity. It will be proven that a combination of a FIFO-based merge sorter and a tree-based merge sorter results in the best performance at low cost. Moreover, we will demonstrate how partial run-time reconfiguration can be used for saving almost half the FPGA resources or alternatively for improving the speed. Experiments show a sustainable sorting throughput of 2GB/s for problems fitting into the on-chip FPGA memory and 1 GB/s when using external memory. These values surpass the best published results on large problem sorting implementations on FPGAs, GPUs, and the Cell processor. Dirk Koch, Jim Tørresen |
FPGA | 1 |
| 2011 | A Routing Architecture for Mapping Dataflow Graphs at Run-TimeabstractWhile it is feasible with today's commercial tools to swap from one module to another, many applications demand more advanced configuration schemes, for example, to map different dataflow graphs onto an FPGA. To support this, we will improve an existing on-FPGA communication architecture in order to carry out arbitrary routing among multiple freely placed modules. This results in a circuit switching network that will be efficiently implemented directly within the FPGA routing fabric. Dirk Koch, Jim Tørresen |
FPL | 1 |
| 2010 | Fine-Grained Partial Runtime Reconfiguration on Virtex-5 FPGAsabstractThe architecture of Xilinx FPGAs, has changed remarkable with respect to their ability to implement runtime reconfigurable systems throughout the last generations. This paper will discuss these changes and reveal an on-FPGA communication architecture that is especially tailored to Xilinx Virtex-5 FPGAs. With this architecture, modules can be integrated in a two-dimensional grid with more than a hundred of individual tiles while allowing a throughput of several GB/s to reconfigurable modules. Dirk Koch, Christian Beckhoff, Jim Tørresen |
FCCM | 1 |
| 2010 | Short-Circuits on FPGAs Caused by Partial Runtime ReconfigurationabstractIn this paper, we show how short-circuits on FPGAs can be caused by partial runtime reconfiguration. Short-circuit can even occur on FPGAs that do not offer any tristate resources just by using off the shelf vendor tools without any bitstream manipulation. The duration of the here presented short-circuits ranges from short spikes up to persistent short-circuits that remain active during runtime. Short-circuits will result in increased current consumption and can thus harm the system and must therefore be prevented. An algorithm is derived that detects whether configuration data will cause short-circuits. We implemented this algorithm in a bitstream scanner that can also be used in systems at runtime. Christian Beckhoff, Dirk Koch, Jim Tørresen |
FPL | 2 |
| 2010 | A Bus-Based SoC Architecture for Flexible Module Placement on Reconfigurable FPGAsabstractThis paper proposes an FPGA-based System-on-Chip (SoC) architecture with support for dynamic runtime reconfiguration. The SoC is divided into two parts, the static embedded CPU sub-system and the dynamically reconfigurable part. An additional bus system connects the embedded CPU sub-system with modules within the dynamic area, offering a flexible way to communicate among all SoC components. This makes it possible to implement a reconfigurable design with support for free module placement. An enhanced memory access method is included for high-speed access to an external memory. The dynamic part includes a streaming technology which implements a direct connection between reconfigurable modules. The paper describes the architecture and shows the advantages in a smart camera case study. Andreas Oetken, Stefan Wildermann, Jürgen Teich, Dirk Koch |
FPL | 4 |
| 2010 | Obstacle-free two-dimensional online-routing for run-time reconfigurable FPGA-based systemsabstractBy neatly reserving routing resources of an FPGA at design-time, a circuit switching network can be implemented for integrating reconfigurable modules in a two-dimensional manner at run-time. In this network, paths can be set directly by manipulating fractions of the switch matrix configuration. By utilizing disjoint resources for implementing the network and the modules of the system, the network is capable to route paths to partial modules independent of the present module placement layout. This paper proposes concepts, implementation issues, and a design flow for building reconfigurable systems providing such a network. Furthermore, a timing model will be presented for validating the system at run-time. The applicability of the network will be demonstrated in a prototype system by routing I/O pins to partial modules. Dirk Koch, Christian Beckhoff, Jim Tørresen |
FPT | 1 |
| 2010 | Advanced partial run-time reconfiguration on Spartan-6 FPGAsabstractIn this paper, we demonstrate systems based on Spartan-6 series FPGAs that provide full support for active partial run-time reconfiguration. We will summarize design factors for successfully applying run-time reconfiguration, reveal details on partial reconfiguration on Spartan-6 FPGAs, and introduce our easy to use design flow. In this flow, a module can multiple times be instantiated or even migrated to different systems without the need to physically reimplement such a module. The demo systems can host manifold different partial modules that each are capable to manipulate a video stream. Dirk Koch, Christian Beckhoff, Jim Tørresen |
FPT | 1 |
| 2010 | Routing optimizations for component-based system design and partial run-time reconfiguration on FPGAsabstractIn component-based system design, systems are composed by integrating fully physically implemented hard-IP cores. This design style has the potential to close the design productivity gap that will arise when FPGAs will head over the one million look-up table boundary by minimizing, or even fully removing, costly verification and place&route steps during the system integration phase. This integration can even be carried out at run-time, and consequently, making component-based design the base for implementing partially reconfigurable systems. In this paper, we will discuss requirements on the routing of the components such that modules can be composed to systems without interfering among each other while still being able of providing the top-level component to component communication. We will propose to relax the strict bounding box constraints that have been traditionally applied to implement reconfigurable components. Our results will demonstrate an area and reconfiguration time improvement of up to 33% as compared to the traditional strict bounding box method in a case study using Xilinx Spartan-6 FPGAs. Dirk Koch, Jim Tørresen |
FPT | 1 |
| 2009 | Minimizing Internal Fragmentation by Fine-Grained Two-Dimensional Module Placement for Runtime Reconfiguralble SystemsabstractThis paper analyzes fragmentation issues and proves that the reconfigurable area must be tiled much finer as has been done in existing approaches. The optimal tile grid can typically only be implemented by tiling the reconfigurable area into a two-dimensional grid. This will further increase the utilization of dedicated resources such as block RAMs. In order to provide communication with the reconfigurable modules, the novel ReCoBus communication architecture is enhanced for two-dimensional communication. A case study will demonstrate a system with 248 individual logic tiles that are each less than 200 LUTs in size while still being able of providing a module connection in each particular tile. Dirk Koch, Christian Beckhoff, Jürgen Teich |
FCCM | 1 |
| 2009 | A communication architecture for complex runtime reconfigurable systems and its implementation on spartan-3 FPGAsabstractIn this paper, we present and analyze a sophisticated communication architecture that allows to integrate many different modules into a system by FPGA reconfiguration at runtime. Furthermore, we examine how this architecture can be implemented on low-cost Spartan-3 devices. It will be demonstrated that modules can be exchanged in a system without disturbing the communication architecture. The paper points out, that the capabilities of Spartan-3 FPGAs are sufficient to build complex reconfigurable systems. Dirk Koch, Christian Beckhoff, Jürgen Teich |
FPGA | 1 |
| 2009 | Hardware Decompression Techniques for FPGA-Based Embedded SystemsabstractIn this work, we present hardware decompression accelerators for widening the bottleneck between slow nonvolatile memories on the one side and high-speed FPGA configuration interfaces and fast softcore CPUs on the other side. We discuss different compression algorithms suitable for a hardware accelerated decompression on FPGAs as well as on CPLDs. The algorithms will be investigated with respect to the achievable compression ratio, throughput, and hardware overhead. This leads to various decompressor implementations with one capable to decompress at high data rates of up to 400 megabytes per second under optimal conditions while only requiring slightly more than a hundred lookup tables. We will evaluate how these decompressors perform on configuration bitstreams for different FPGAs as well as for softcore CPU binaries. Dirk Koch, Christian Beckhoff, Jürgen Teich |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2008 | Efficient Reconfigurable On-Chip Buses for FPGAsabstractThis paper presents techniques for generating on-chip buses suitable for dynamically integrating hardware modules into an FPGA-based SoC by partial reconfiguration. The buses permit direct connections of master and slave modules to the bus in combination with a flexible fine-grained module placement and with minimized latency and area overheads. A test system will demonstrate a transfer rate of 800 MB/s while providing an extreme high placement flexibility. Dirk Koch, Christian Haubelt, Jürgen Teich |
FCCM | 1 |
| 2008 | No-break dynamic defragmentation of reconfigurable devicesabstractWe propose a new method for defragmenting the module layout of a reconfigurable device, enabled by a novel approach for dealing with communication needs between relocated modules and with inhomogeneities found in commonly used FPGAs. Our method is based on dynamic relocation of module positions during runtime, with only very little reconfiguration overhead; the objective is to maximize the length of contiguous free space that is available for new modules. We describe a number of algorithmic aspects of good defragmentation, and present an optimization method based on tabu search. Experimental results indicate that we can improve the quality of module layout by roughly 50% over static layout. Among other benefits, this improvement avoids unnecessary rejection of modules. Sándor P. Fekete, Tom Kamphans, Nils Schweer, Christopher Tessars, Jan van der Veen, Josef Angermeier, Dirk Koch, Jürgen Teich |
FPL | 7 |
| 2008 | ReCoBus-Builder - A novel tool and technique to build statically and dynamically reconfigurable systems for FPGASabstractIn this paper, we present the ReCoBus-builder tool chain that simplifies the generation of dynamically reconfigurable systems to almost a push-button process. The generated systems provide one or more resource areas that will be used by different partially reconfigurable modules at runtime. It is possible to integrate multiple partially reconfigurable modules into the same resource area at the same time and these modules can communicate via a fixed bus infrastructure or dedicated point-to-point links with other parts of the system. This allows building encapsulated modules that will be integrated into the system by linking together bitstreams at runtime. We will demonstrate that bitstream linking can further be used to speed up the design process of static only systems by eliminating long synthesis runs or place and route steps, when only small portions of a design are exchanged. Dirk Koch, Christian Beckhoff, Jürgen Teich |
FPL | 1 |
| 2008 | Photo-realistic Rendering of Metallic Car Paint from Image-Based MeasurementsabstractAbstract State‐of‐the‐art car paint shows not only interesting and subtle angular dependency but also significant spatial variation. Especially in sunlight these variations remain visible even for distances up to a few meters and give the coating a strong impression of depth which cannot be reproduced by a single BRDF model and the kind of procedural noise textures typically used. Instead of explicitly modeling the responsible effect particles we propose to use image‐based reflectance measurements of real paint samples and represent their spatial varying part by Bidirectional Texture Functions (BTF). We use classical BRDF models like Cook‐Torrance to represent the reflection behavior of the base paint and the highly specular finish and demonstrate how the parameters of these models can be derived from the BTF measurements. For rendering, the image‐based spatially varying part is compressed and efficiently synthesized. This paper introduces the first hybrid analytical and image‐based representation for car paint and enables the photo‐realistic rendering of all significant effects of highly complex coatings. Martin Rump, Gero Müller, Ralf Sarlette, Dirk Koch, Reinhard Klein |
Comput. Graph. Forum | 4 |
| 2007 | Efficient hardware checkpointing: concepts, overhead analysis, and implementationabstractProgress in reconfigurable hardware technology allows the implementation of complete SoCs in today's FPGAs. In the context design for reliability, software checkpointing is an effective methodology to cope with faults. In this paper, we systematically extend the concept of checkpointing known from software systems to hardware tasks running on reconfigurable devices. We will classify different mechanisms for hardware checkpointing and present formulas for estimating the hardware overhead. Moreover, we will reveal a tool that takes over the burden of modifying hardware modules for checkpointing. Post-synthesis results of applying our methodology to different hardware accelerators will be presented and the results will be compared with the theoretical estimations. Dirk Koch, Christian Haubelt, Jürgen Teich |
FPGA | 1 |
| 2007 | Bitstream Decompression for High Speed FPGA Configuration from Slow MemoriesabstractIn this paper, we present hardware decompression accelerators for bridging the gap between high speed FPGA configuration interfaces and slow configuration memories. We discuss different compression algorithms suitable for a decompression on FPGAs as well as on CPLDs with respect to the achievable compression ratio, throughput, and hardware overhead. This leads to various decompressor implementations with one capable to decompress at high data rates of up to 400 megabytes per second while only requiring slightly more than a hundred look-up tables. Furthermore, we present a sophisticated configuration bitstream benchmark. Dirk Koch, Christian Beckhoff, Jürgen Teich |
FPT | 1 |
| 2007 | Modeling and Synthesis of Hardware-Software MorphingabstractIn state of the art hardware-software-co-design flows for FPGA based systems, the hardware-software partitioning problem is solved offline, thus, omitting the great flexibility provided through partial runtime reconfiguration. The decision which functions are best suitable to be implemented in hardware or software, is typically taken with respect to the expected worst case computational demands and certain objectives like power consumption, throughput or cost. However, if these parameters change at runtime, e.g., due to environmental changes, traditional designed systems lack to adapt to the new conditions, because the hardware-software partitioning is static. This paper systematically presents a new methodology that allows changing the implementation style of tasks at runtime by hardware-software morphing. Based on a formal model, how morphing can be performed without loosing internal states was demonstrated. Moreover, results from applying this methodology were demonstrated to a 16-tap FIR filter. Dirk Koch, Christian Haubelt, Thilo Streichert, Jürgen Teich |
ISCAS | 1 |
| 2004 | A Dynamic NoC Approach for Communication in Reconfigurable Devices
Christophe Bobda, Mateusz Majer, Dirk Koch, Ali Ahmadinia, Jürgen Teich |
FPL | 3 |
| 2004 | Preemptive Hardware Task Management
Dirk Koch |
FPL | 1 |
| 2004 | FPGA architecture extensions for preemptive multitasking and hardware defragmentationabstractThe focus in This work is put onto extensions to typical FPGA hardware architectures in order to support some operating system functions. In particular, we examine the problem of preemptive multitasking and hardware defragmentation on a reconfigurable system based on FPGAs with a dynamically changing set of hardware tasks that can be replaced considering internal states. Dirk Koch, Ali Ahmadinia, Christophe Bobda, Heiko Kalte |
FPT | 1 |