Michael Hübner 0001

dblp:56/2900 · DBLP profile ↗
← Back
82ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-1790-3869ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 77 · 5 first-author · 11 since 2021Software engineering, systems software and programming languages · 13 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1
YearPublicationVenuePosition
2026 Integrating an open-source soft-GPU overlay with RISC-V control and high-bandwidth memory
abstract
Image and signal processing workloads are widely deployed on Graphics Processing Units (GPUs) for high throughput and on Field-Programmable Gate Arrays (FPGAs) for hardware specialization and energy efficiency. Soft GPU overlays on FPGAs aim to combine these advantages, yet existing solutions often depend on fixed hard processors or impose platform constraints that limit portability. This work extends a popular open-source soft GPGPU overlay to integrate a soft RISC-V control plane and enable compatibility with High-Bandwidth Memory (HBM2). The resulting system can be instantiated on FPGA boards without a hard ARM processor, improving portability, simplifying system integration, and broadening deployability. Across representative image and signal processing kernels, the soft GPGPU achieves geometric-mean speedups of 114.60 × over a scalar soft RISC-V core and 19.72 × over a hard ARM core, demonstrating substantial performance benefits while retaining FPGA reconfigurability. HBM2 integration further benefits bandwidth-sensitive workloads by increasing sustained throughput and reducing the performance bottlenecks associated with off-chip memory access. Collectively, these results indicate that GPU-like programmability and performance can be delivered on reconfigurable platforms without reliance on hard CPU subsystems, providing a portable and scalable foundation for embedded vision and DSP acceleration.
Hector Gerardo Muñoz Hernandez, Mahdi Taheri, Muhammad Ali 0010, Keyvan Shahin, Alireza Syavashi, Diana Göhringer, Marc Reichenbach, Christian Herglotz, Michael Hübner 0001
J. Syst. Archit.9
2022 Health Monitoring of Milling Tools under Distinct Operating Conditions by a Deep Convolutional Neural Network model
abstract
One of the most popular manufacturing techniques is milling. It can be used to make a variety of geometric components, such as flat grooves, surfaces, etc. The condition of the milling tool has a major impact on the quality of milling processes. Hence the importance of follow-up. When working on monitoring solutions, it is crucial to take into account different operating variables, such as rotational speed, especially in real world experiences. This work addresses the topic of predictive maintenance by exploiting the fusion of sensor data and the artificial intelligence-based analysis of signals measured by sensors. With a set of data such as vibration and sound reflection from the sensors, we focus on finding solutions for the task of detecting the health condition of machines. A Deep Convolutional Neural Network (DCNN) model is provided with fusion at the sensor data level to detect five consecutive health states of a milling tool; From a healthier state to a state of degradation. In addition, a demonstrator is built with Simulink to simulate and visualize the detection process. To examine the capacity of our model, the signal data was processed individually and subsequently merged. Experiments were carried out on three sets of data recorded during a real milling process. Results using the proposed DCNN architecture with raw data have reached an accuracy of more than 94% for all data sets.
Priscile Suawa Fogou, Michael Hübner 0001
DATE2
2022 STAP: An Architecture and Design Tool for Automata Processing on Memristor TCAMs
abstract
Accelerating finite-state automata benefits several emerging application domains that are built on pattern matching. In-memory architectures, such as the Automata Processor (AP), are efficient to speed them up, at least for outperforming traditional von-Neumann architectures. In spite of the AP’s massive parallelism, current APs suffer from poor memory density, inefficient routing architectures, and limited capabilities. Although these limitations can be lessened by emerging memory technologies, its architecture is still the major source of huge communication demands and lack of scalability. To address these issues, we present STAP , a Scalable TCAM-based architecture for Automata Processing . STAP adopts a reconfigurable array of processing elements, which are based on memristive Ternary CAMs (TCAMs), to efficiently implement Non-deterministic finite automata (NFAs) through proper encoding and mapping methods. The CAD tool for STAP integrates the design flow of automata applications, a specific mapping algorithm, and place and route tools for connecting processing elements by RRAM-based programmable interconnects. Results showed 1.47× higher throughput when processing 16-bit input symbols, and improvements of 3.9× and 25× on state and routing densities over the state-of-the-art AP, while preserving 10 4 programming cycles.
João Paulo C. de Lima, Marcelo Brandalero, Michael Hübner 0001, Luigi Carro
ACM J. Emerg. Technol. Comput. Syst.3
2022 Reduced Precision DWC: An Efficient Hardening Strategy for Mixed-Precision Architectures
abstract
Duplication with Comparison (DWC) is an effective software-level solution to improve the reliability of computing devices. However, it introduces performance and energy consumption overheads that could be unsuitable for high-performance computing or real-time safety-critical applications. In this article, we present Reduced-Precision Duplication with Comparison (RP-DWC) as a means to lower the overhead of DWC by executing the redundant copy in reduced precision. RP-DWC is particularly suitable for modern mixed-precision architectures, such as NVIDIA GPUs, that feature dedicated functional units for computing with programmable accuracy. We discuss the benefits and challenges associated with RP-DWC and show that the intrinsic difference between the mixed-precision copies allows for detecting most, but not all, errors. However, as the undetected faults are the ones that fall into the difference between precisions, they are the ones that produce a much smaller impact on the application output and, thus, might be tolerated. We investigate RP-DWC impact into fault detection, performance, and energy consumption on Volta GPUs. Through fault injection and beam experiment, using three microbenchmarks and four real applications, we show that RP-DWC achieves an excellent coverage (up to 86 percent) with minimal overheads (as low as 0.1 percent time and 24 percent energy consumption overhead).
Fernando Santos 0001, Marcelo Brandalero, Michael B. Sullivan 0001, Pedro Martins Basso, Michael Hübner 0001, Luigi Carro, Paolo Rech
IEEE Trans. Computers5
2021 Artificial Intelligence for Mass Spectrometry and Nuclear Magnetic Resonance Spectroscopy
abstract
Mass Spectrometry (MS) and Nuclear Magnetic Resonance Spectroscopy (NMR) are critical components of every industrial chemical process as they provide information on the concentrations of individual compounds and by-products. These processes are carried out manually and by a specialist, which takes a substantial amount of time and prevents their utilization for real-time closed-loop process control. This paper presents recent advances from two projects that use Artificial Neural Networks (ANNs) to address the challenges of automation and performance-efficient realizations of MS and NMR. In the first part, a complete toolchain has been developed to develop simulated spectra and train ANNs to identify compounds in MS. In the second part, a limited number of experimental NMR spectra have been augmented by simulated spectra to train an ANN with better prediction performance and speed than state-of-the-art analysis. These results suggest that, in the context of the digital transformation of the process industry, we are now on the threshold of a possible strongly simplified use of MS and MRS and the accompanying data evaluation by machine-supported procedures, and can utilize both methods much wider for reaction and process monitoring or quality control.
Florian Fricke, Safdar Mahmood, Javier Hoffmann, Marcelo Brandalero, Sascha Liehr, Simon Kern, Klas Meyer, Stefan Kowarik, Stephan Westerdick, Michael Maiwald, Michael Hübner 0001
DATE11
2021 Design and Implementation Strategy of Adaptive Processor-Based Systems for Error Resilient and Power-Efficient Operation
abstract
The contemporary computing systems are facing two major challenges: excessive power consumption and susceptibility to faults. In order to take advantage of techniques that efficiently address these challenges, the classic ASIC design flow requires some modifications. In this paper, we present a simple and convenient strategy for design and implementation of processor-based systems using highly configurable, cross-layer framework that encompasses techniques such as Adaptive Voltage and Frequency Scaling (AVFS) and Triple Modular Redundancy (TMR). The proposed strategy augments the conventional design flow with two additional steps to integrate the framework's hardware building blocks into the system. Such system is then able to dynamically switch between low power and error resilient operation modes according to the current requirements. By following the proposed strategy, we were able to implement processor-based system that significantly reduces the power consumption / increases the soft error resilience while preserving the performance at negligible area overhead of less than 1%.
Mitko Veleski, Michael Hübner 0001, Milos Krstic, Rolf Kraemer
DDECS2
2021 Towards Machine Learning Support for Embedded System Tests
abstract
The correctness of embedded systems needs to be ensured by a high number of tests. Large amounts of data reflecting the system behavior are collected during these test runs. Automated test evaluations are often limited to checking very specific requirements which can hardly cover all possible kinds of erroneous behaviors. Manual examinations can compensate this by implicit knowledge of experienced test engineers but are very time-consuming and therefore costly.This paper shows how machine learning can support the evaluation of embedded system tests. Assessment of new test runs is based on available data from previous tests and aims at identifying those that deviate from usual behavior. Moreover the paper presents a generic approach that helps to find the most suitable detection algorithm in the given context. A case study proves the effectiveness of this approach. Quantitative comparisons show that our exploration is able to find solutions that outperform state of the art methods.
Stefan Scharoba, Kai-Uwe Basener, Jens Bielefeldt, Hans-Werner Wiesbrock, Michael Hübner 0001
DSD5
2021 AITIA: Embedded AI Techniques for Industrial Applications
abstract
Motivated by an increasing interest from startups in embedded Artificial Intelligence (AI) and by their limited expertise, the AITIA Project targets the development of embedded AI techniques for industrial applications. This extended abstract presents the motivation and the solutions being developed towards four use cases: smart sensors, network intrusion detection, driver-assistance systems, and Industry 4.0.
Marcelo Brandalero, Mitko Veleski, Hector Gerardo Muñoz Hernandez, Muhammad Ali 0010, Laurens Le Jeune, Toon Goedemé, Nele Mentens, Jurgen Vandendriessche, Lancelot Lhoest, Bruno da Silva 0001, Abdellah Touhafi, Diana Göhringer, Michael Hübner 0001
FPL13
2021 Accelerating Fixed-Point Simulations Using Width Reconfigurable Hardware Architectures
abstract
Digital Signal Processing (DSP) systems can be described using either fixed-point or floating-point for their numeric representations. As the general computers use floating-point representation, software-based simulation tools used for modeling and simulation of the arithmetic operations in DSP systems, use floating-point representation as well. However, synthesizing customized hardware for fixed-point arithmetic operations for FPGAs or ASICs is more efficient compared to their floating-point counterparts. Thus it is necessary to convert the representation of a floating-point simulated algorithm on MATLAB for example, to a fixed-point representation which is more suitable for hardware implementation. While former approaches for this conversion step have always been software-based, like on MATLAB itself, this paper presents a new approach to show the possibility of accelerating it by using hardware width reconfigurable designs.
Keyvan Shahin, Michael Hübner 0001
FPL2
2021 The SMART4ALL High Performance Computing Infrastructure: Sharing high-end hardware resources via cloud-based microservices
abstract
The goal of the SMART4ALL EU funded project is to build capacity amongst European stakeholders via the development of self-sustained, cross-border experiments that transfer knowledge and technology between academia and industry. It targets Customized Low Energy Computing (CLEC) in the Cyber-Physical (CPS) and the Internet of Things (IoT) domains. The vision of the project will be mainly realized through funded (via open calls) Pathfinder Application Experiments (PAEs) that will enable the transformation of academic knowledge into products.It is important to mention that SMART4ALL sets forward the concept of marketplace which is offered as a service (Marketplace-as-a-Servtce or MaaS) that acts as one-stopsmart-stop by offering tools, HPC services, and platforms to the PAE grantees. The purpose of this paper is to present the SMART4ALL High-Performance Computing (HPC) infrastructure and services realized according to the Hardware-as-a-Service paradigm. The SMART4A11 HPC infrastructure includes CPU, Memory, GPU, storage as well as FPGA resources connected with high-speed links.
Angelos S. Voros, Christos Panagiotou, Stavros Zogas, Georgios Keramidas, Christos P. Antonopoulos, Michael Hübner 0001, Nikos S. Voros
FPL6
2021 Towards Error Resilient and Power-Efficient Adaptive Multiprocessor System using Highly Configurable and Flexible Cross-Layer Framework
abstract
A typical multiprocessor system often needs to support a wide spectrum of applications. Today, error resilience and low power consumption are two crucial, but non-complementary requirements and meeting both simultaneously is difficult. Thus, adaptivity is becoming increasingly important feature for modern computing systems. In this regard, we integrate a highly-configurable framework with a set of cross-layer techniques efficient in improving error resilience / power consumption into a multiprocessor system. The framework intelligently interchanges methods such as Adaptive Voltage and Frequency Scaling (AVFS), Triple Modular Redundancy (TMR) and clock-gating while the system is on-line. Additionally, flexibility as an inherent multiprocessor feature enables dynamical adaptation of the system to the current requirements. Putting all together, a balanced level between the two key metrics is achieved. We conduct numerous experiments to show the advantages of the proposed approach. Finally, we use the results to confirm the benefits and the effectiveness of the framework.
Mitko Veleski, Michael Hübner 0001, Milos Krstic, Rolf Kraemer
IOLTS2
2020 Deep Learning Utilization in Beamforming Enhancement for Medical Ultrasound
abstract
Ultrasound imaging offers a low cost, noninvasive and portable system, which allowed it to be an invaluable tool for medical imaging. However, the quality of the reconstructed images depends significantly on the beamforming technique utilized. Although advanced data-adaptive methods of reconstruction such as Minimum Variance (MV) beamforming can recover image quality much higher than conventional techniques, their implementation also entails a heavy computational burden. This dichotomy hinders the ultrasound imaging use as a standalone device in some applications such as early breast cancer detection. Deep neural networks (DNNs) have shown a huge potential when applied to many Artificial intelligence (AI) research fields. In this work, the use of Deep learning in improving the quality of the beamforming technique Delay and Sum (DAS) normally used for ultrasound (US) images reconstruction is explored. Three different architectures are implemented: Convolutional AutoEncoder (CAE), Fully Connected network (FC) and U-Net-like architecture. They were trained on datasets simulated using field II. The dataset consists of input-output pairs where the input is Noisy DAS beamformed scan lines and the output is MV beamformed non-noisy scan lines. The networks show a great ability in predicting the beamformed signals along with significantly reducing noise in the reconstructed images. Additionally, the proposed networks improve other image characteristics such as scatterer size and position along with reducing tail characteristic normally found in DAS beamformed ultrasound images. US images constructed by the networks achieved better quality metrics that surpass conventional DAS beamformed images. The CAE, U-net-like architecture, and FC enhanced the signal to noise ratio (SNR) compared to DAS by 218%, 165% and 136% respectively. Additionally, the networks showed higher Contrast to Noise Ratio (CNR) and Contrast Ratio (CR) metrics than DAS beamformed signals. Finally, the proposed approach achieves a 60% enhancement in time consumption of image reconstruction compared to MV technique, which allows higher possible frame rate with a comparable outcome.
Mariam M. Fouad, Yousef Metwally, Georg Schmitz, Michael Hübner 0001, Mohamed Abdelghany
COMPSAC4
2020 Proactive Aging Mitigation in CGRAs through Utilization-Aware Allocation
abstract
Resource balancing has been effectively used to mitigate the long-term aging effects of Negative Bias Temperature Instability (NBTI) in multi-core and Graphics Processing Unit (GPU) architectures. In this work, we investigate this strategy in Coarse-Grained Reconfigurable Arrays (CGRAs) with a novel application-to-CGRA allocation approach. By introducing important extensions to the reconfiguration logic and the datapath, we enable the dynamic movement of configurations throughout the fabric and allow overutilized Functional Units (FUs) to recover from stress-induced NBTI aging. Implementing the approach in a resource-constrained state-of-the-art CGRA reveals 2.2× lifetime improvement with negligible performance overheads and less than 10% increase in area.
Marcelo Brandalero, Bernardo Neuhaus Lignati, Antonio Carlos Schneider Beck, Muhammad Shafique 0001, Michael Hübner 0001
DAC5
2020 RESCUE: Interdependent Challenges of Reliability, Security and Quality in Nanoelectronic Systems
abstract
The recent trends for nanoelectronic computing systems include machine-to-machine communication in the era of Internet-of-Things (IoT) and autonomous systems, complex safety-critical applications, extreme miniaturization of implementation technologies and intensive interaction with the physical world. These set tough requirements on mutually dependent extra-functional design aspects. The H2020 MSCAITN project RESCUE is focused on key challenges for reliability, security and quality, as well as related electronic design automation tools and methodologies. The objectives include both research advancements and cross-sectoral training of a new generation of interdisciplinary researchers. Notable interdisciplinary collaborative research results for the first halfperiod include novel approaches for test generation, soft-error and transient faults vulnerability analysis, cross-layer fault-tolerance and error-resilience, functional safety validation, reliability assessment and run-time management, HW security enhancement and initial implementation of these into holistic EDA tools.
Maksim Jenihhin, Said Hamdioui, Matteo Sonza Reorda, Milos Krstic, Peter Langendörfer, Christian Sauer 0001, Anton Klotz, Michael Hübner 0001, Jörg Nolte, Heinrich Theodor Vierhaus, Georgios N. Selimis, Dan Alexandrescu, Mottaqiallah Taouil, Geert Jan Schrijen, Jaan Raik, Luca Sterpone, Giovanni Squillero, Zoya Dyka
DATE8
2020 Highly Configurable Framework for Adaptive Low Power and Error-Resilient System-On-Chip
abstract
In this paper, a novel, highly configurable framework for low power and error-resilient System-On-Chip is presented. The framework is composed, on the one hand, of the SWIELD configurable flip-flop and on the other hand, of the Chameleon controller. The SWIELD flip-flop is able to operate in three modes. It is driven/configured during runtime via the dedicated controller called Chameleon System Operation Management Unit. The proposed framework is integrated into a complex SoC based on a 32-bit general-purpose processor and the entire system is synthesized using the IHP 130 nm technology library. Numerous simulation experiments have been conducted in order to estimate the system error resilience and power consumption. At expense of negligible area and complexity overhead, the introduced framework shows great potential and excellent results w.r.t. both metrics of interest.
Mitko Veleski, Michael Hübner 0001, Milos Krstic, Rolf Kraemer
DSD2
2020 MCEA: A Resource-Aware Multicore CGRA Architecture for the Edge
abstract
Modern IoT edge devices must address the unpredictability of applications with strict power and temperature constraints. In this scenario, heterogeneous multicore architectures have been driving many solutions due to their high energy efficiency and ability to exploit Task-Level Parallelism. However, while their performance is highly dependent on the quality of the scheduling, their adaptability and generality get restricted when they use fixed-size hardware accelerators. Considering that, this work proposes MCEA, a transparent and power-adaptive multicore reconfigurable architecture. MCEA dynamically adapts the hardware to the workload rather than migrating applications; and predicatively sizes its reconfigurable accelerators without prior knowledge of the applications' behaviors. For that, MCEA uses a synergistic and online profiling system with power gating, achieving performance levels near of homogeneous architectures with fixed and oversized reconfigurable fabric (within 99% on average) while presenting energy efficiency levels similar to heterogeneous architectures statically tuned to a specific workload (within 99% on average). Therefore, MCEA improves Energy-Delay Product in 1.55x and 1.21x when compared to their heterogeneous and homogeneous counterparts, and in 4.72x when compared to a multicore with OoO processors only. We also show that MCEA outperforms a state-of-the-art reconfigurable architecture for the edge under the same power envelope.
Guilherme Korol, Michael G. Jordan, Marcelo Brandalero, Michael Hübner 0001, Mateus B. Rutzig, Antonio Carlos Schneider Beck
FPL4
2020 Reduced-Precision DWC for Mixed-Precision GPUs
abstract
Duplication with Comparison (DWC) is an effective software-level solution to improve the reliability of computing systems, including Graphics Processing Units (GPUs). DWC, however, introduces performance and energy consumption overheads that could be unacceptable for High-Performance Computing (HPC) or real-time safety-critical applications. In this work, we propose Reduced-Precision DWC (RP-DWC): an improvement over the traditional DWC approach that uses mixed-precision GPUs hardware resources to implement fault detection. We investigate, through both fault injection campaigns and accelerated neutron beam experiments, the impact of RPDWC onto performance, energy consumption, and its fault detection capabilites. We show that RP-DWC achieves on average 74% fault coverage (up to 86%) with very small overheads (0.1% time and 24% energy consumption overhead, in the best case).
Fernando Santos 0001, Marcelo Brandalero, Pedro Martins Basso, Michael Hübner 0001, Luigi Carro, Paolo Rech
IOLTS4
2020 Soft Error Hardened Asymmetric 10T SRAM Cell for Aerospace Applications
Ambika Prasad Shah, Santosh Kumar Vishvakarma, Michael Hübner 0001
J. Electron. Test.3
2020 A Machine Learning Methodology for Cache Memory Design Based on Dynamic Instructions
abstract
Cache memories are an essential component of modern processors and consume a large percentage of their power consumption. Its efficacy depends heavily on the memory demands of the software. Thus, finding the optimal cache for a particular program is not a trivial task and usually involves exhaustive simulation. In this article, we propose a machine learning–based methodology that predicts the optimal cache reconfiguration for any given application, based on its dynamic instructions. Our evaluation shows that our methodology reaches 91.1% accuracy. Moreover, an additional experiment shows that only a small portion of the dynamic instructions (10%) suffices to reach 89.71% accuracy.
Osvaldo Navarro, Jones Yudi Mori, Javier Hoffmann, Hector Gerardo Muñoz Hernandez, Michael Hübner 0001
ACM Trans. Embed. Comput. Syst.5
2019 An Integrated on-Silicon Verification Method for FPGA Overlays
Alexandra Kourfali, Florian Fricke, Michael Hübner 0001, Dirk Stroobandt
J. Electron. Test.3
2019 Guest Editorial: Special Issue on Reconfigurable Computing and FPGA Technology
René Cumplido, Maya B. Gokhale, Claudia Feregrino-Uribe, Michael Hübner 0001
J. Parallel Distributed Comput.4
2018 Guest Editorial Circuit and System Design Automation for Internet of Things
abstract
Internet-of-Things (IoT) is the technical backbone of smart cities which are envisioned to cope up with rapid urbanization of human population with limited resources. IoT provides three key features of smart cities such as intelligence, interconnection, and instrumentation. IoT is essentially a system-of-systems which can be considered as a configurable dynamic global network of networks. The main components of IoT include the following: 1) The Things; 2) Internet; 3) LAN; and 4) The Cloud. IoT is built by various diverse components including electronics, sensors, actuators, controllers, networks, firmware, and software. However, the existing electronics, controllers, and processors do not meet IoT requirements, such as multiple sensors, communication protocols, and security requirements. The existing computer-aided design (CAD) or electronic design automation tools are not enough to meet diverse challenges such as time-to-market, complexity, and cost of IoT. The required electronic circuits and systems need to be developed by handling and solving specific requirements. Real-time and ultralow power plays a major role since mobile devices in the IoT have to provide a long availability with a relative small energy budget. At the same time, reliability, availability, real-time constraints, and performance requirements pose significant challenges, and therefore, lead to a high interest in research. In this special issue, different approaches to design novel devices, circuits, and systems for solving the challenges with IoT are targeted. Various novel design automation components including modeling, design flows, simulation methods, and optimizations for designing of modern IoT are targeted, from system level down to device level. The current special issue was envisioned with the above technical considerations. After a rigorous review process, a set of articles were selected for this special issue. These papers are briefly discussed in the rest of the editorial.
Saraju P. Mohanty, Michael Hübner 0001, Chun Jason Xue, Xin Li 0001, Hai Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2018 General-Purpose Computing with Soft GPUs on FPGAs
abstract
Using field-programmable gate arrays (FPGAs) as a substrate to deploy soft graphics processing units (GPUs) would enable offering the FPGA compute power in a very flexible GPU-like tool flow. Application-specific adaptations like selective hardening of floating-point operations and instruction set subsetting would mitigate the high area and power demands of soft GPUs. This work explores the capabilities and limitations of soft General Purpose Computing on GPUs (GPGPU) for both fixed- and floating point arithmetic. For this purpose, we have developed FGPU: a configurable, scalable, and portable GPU architecture designed especially for FPGAs. FGPU is open-source and implemented entirely in RTL. It can be programmed in OpenCL and controlled through a Python API. This article introduces its hardware architecture as well as its tool flow. We evaluated the proposed GPGPU approach against multiple other solutions. In comparison to homogeneous Multi-Processor System-On-Chips (MPSoCs), we found that using a soft GPU is a Pareto-optimal solution regarding throughput per area and energy consumption. On average, FGPU has a 2.9× better compute density and 11.2× less energy consumption than a single MicroBlaze processor when computing in IEEE-754 floating-point format. An average speedup of about 4× over the ARM Cortex-A9 supported with the NEON vector co-processor has been measured for fixed- or floating-point benchmarks. In addition, the biggest FGPU cores we could implement on a Xilinx Zynq-7000 System-On-Chip (SoC) can deliver similar performance to equivalent implementations with High-Level Synthesis (HLS).
Muhammed Al Kadi, Benedikt Janßen, Jones Yudi Mori, Michael Hübner 0001
ACM Trans. Reconfigurable Technol. Syst.4
2017 An open reconfigurable research platform as stepping stone to exascale high-performance computing
abstract
To handle the stringent performance and power requirements of future exascale-class applications, High Performance Computing (HPC) systems need ultra-efficient heterogeneous compute nodes and hardware accelerators with a high degree of specialization. Ideally, dynamic reconfiguration will be an intrinsic feature, so that specific HPC application features can be optimally accelerated, even if they regularly change over time. We create a new and flexible exploration platform for developing reconfigurable architectures, design tools and HPC applications with run-time reconfiguration built-in as a core fundamental feature instead of an add-on. Our project proposes an open research platform that covers the entire stack from architecture up to the application, focusing on the fundamental building blocks for run-time reconfigurable exascale HPC systems: new chip architectures with very low reconfiguration overhead, new tools that truly take reconfiguration as a central design concept, and applications that are tuned to maximally benefit from the proposed run-time reconfiguration techniques. Ultimately, this open platform will enable groundbreaking research towards new exascale computing platforms.
Dirk Stroobandt, Catalin Bogdan Ciobanu, Marco D. Santambrogio, Gabriel Figueiredo, Andreas Brokalakis, Dionisios N. Pnevmatikatos, Michael Hübner 0001, Tobias Becker, Alex J. W. Thom
DATE7
2017 A dynamic partial reconfigurable overlay concept for PYNQ
abstract
Partial reconfiguration is a promising technique in the design of embedded systems since it enables an increase in efficiency and flexibility. However, its usage is still challenging due to the constraints of current FPGAs. In this paper, we present an extension of the Xilinx Python package ‘pynq’ to ease the usage of partial reconfigurable bitstreams. The pynq package belongs to Xilinx's open source project PYNQ. The PYNQ project provides the Python language and its libraries to developers on the Xilinx Zynq. Our pynq package extension enables the usage of partial reconfigurable bitstreams. The concept is based on our Python package ‘pynqpartial’. This package can be integrated into an overlay for PYNQ to enable support for partial reconfigurable bitstreams. Our goal is to enable overlay developers to offer overlays which manage the partial reconfiguration transparent to the Python programmer. We evaluated our approach on the PYNQ-Z1.
Benedikt Janßen, Pascal Zimprich, Michael Hübner 0001
FPL3
2016 Computation and communication challenges to deploy robots in assisted living environments
Georgios Keramidas, Christos P. Antonopoulos, Nikos S. Voros, Fynn Schwiegelshohn, Philipp Wehner, Jens Rettkowski, Diana Göhringer, Michael Hübner 0001, Stasinos Konstantopoulos, Theodoros Giannakopoulos, Vangelis Karkaletsis, Evaggelinos P. Mariatos
DATE8
2016 AutoReloc: Automated Design Flow for Bitstream Relocation on Xilinx FPGAs
abstract
Dynamic and partial reconfiguration of Field Programmable Gate Arrays (FPGA) enable to reuse logic resources for several applications which are scheduled in a sequential order or which are loaded on demand. A fraction of the design on the FPGA is then substituted by another logic function while the rest of the system on the chip stays unaffected. If a design provides several partial reconfigurable areas, the configuration bitstream representing the logic function to be configured in this region has to be adapted to the physical requirements of this chip area. This can be achieved by deploying a repository with all possible configuration bitstreams for all possible regions. It is obvious that storage space can quickly become a limiting parameter in reconfigurable designs. For this purpose, bitstream relocation provides a less storage greedy approach. Only one representation as bitstream of an application needs to be stored. During the configuration process, a relocation algorithm manipulates the bitstream in order to suit it to the respective reconfigurable area. However, reconfigurable regions have to fulfill strong constraints for a relocation to be possible, which makes the selection and placement of reconfigurable regions a complex process. Unfortunately this is not automated by tools so far. In this paper, an approach to automate the development of such relocatable bitstreams is presented along with new algorithms related to relocation specific steps. This approach results in functional designs with minimal intervention from the designer.
André Lalevee, Pierre-Henri Horrein, Matthieu Arzel, Michael Hübner 0001, Sandrine Vaton
DSD4
2016 FGPU: An SIMT-Architecture for FPGAs
abstract
Driven by its high flexibility, good performance and energy efficiency, GPGPU has taken on an increasingly important role in embedded systems. In this paper, we present the basic core of FGPU: a GPU-like, scalable and portable integer soft SIMT-processor implemented in RTL and optimized for FPGA synthesis with a single-level cache system. Compared to a performance-optimized MicroBlaze implementation on the same FPGA, the biggest implemented core of FGPU achieves average wall clock speedups of 49x and a measured power saving of 3.7x with an area overhead of 17.7x. Compared to an ARM CPU with a NEON vector processor, we measured an average speedup of 3.5x over the used benchmark. FGPU is highly parametrizable and it does not contain any manufacturer-specific IP-cores or primitives.
Muhammed Al Kadi, Benedikt Janßen, Michael Hübner 0001
FPGA3
2016 Integer computations with soft GPGPU on FPGAs
abstract
This paper explores the capabilities and limitations of soft GPGPU-based computing on fixed-point arithmetic. The work is based on an existing soft GPU architecture which has been improved and extended to cover broader benchmarks. A generic ALU design for modern FPGA architectures is presented. The enhanced ISA includes conditional instructions and global atomic operations. We extended the tool flow with an LLVM-Backend and used the clang frontend to provide an OpenCL compiler. The improved architecture is evaluated against multiple other solutions: a single MicroBlaze soft processor, a Cortex-A9 ARM with the NEON vector coprocessor and equivalent HLS implementations. We have recorded an average speed up of 10-47× over the MicroBlaze and 0.6-3.4× over the ARM with the NEON engine for the smallest and the biggest soft GPU cores, respectively. Although these cores have an area overhead of 6-22× in comparison to the single soft processor solution, they consumed on average 2.8-7.1× less energy to perform the same tasks. We noticed no performance degradation in comparison to the HLS implementations.
Muhammed Al Kadi, Michael Hübner 0001
FPT2
2016 A Dynamically Reconfigurable Multi-ASIP Architecture for Multistandard and Multimode Turbo Decoding
abstract
The multiplication of wireless communication standards is introducing the need of flexible and reconfigurable multistandard baseband receivers. In this context, multiprocessor turbo decoders have been recently developed in order to support the increasing flexibility and throughput requirements of emerging applications. However, these solutions do not sufficiently address reconfiguration performance issues, which can be a limiting factor in the future. This brief presents the design of a reconfigurable multiprocessor architecture for turbo decoding achieving very fast reconfiguration without compromising the decoding performances.
Vianney Lapotre, Purushotham Murugappa, Guy Gogniat, Amer Baghdadi, Michael Hübner 0001, Jean-Philippe Diguet
IEEE Trans. Very Large Scale Integr. Syst.5
2015 The value of FPGAs as reconfigurable hardware enabling Cyber-Physical Systems
abstract
Industry 4.0 is a reality through the use of intelligent networks capable of gathering and analyzing data and acting on it autonomously. However, important advancements can be made when reconfigurable hardware comes into play. This work presents current trends used in the development of digital circuits, then enumerates a series of challenges faced in the Cyber-Physical Systems environment and, lastly, proposes ways of applying the trends in Cyber-Physical Systems in order to bring their advantages into this growing field.
Tomás Grimm, Benedikt Janßen, Osvaldo Navarro, Michael Hübner 0001
ETFA4
2015 A Framework to the Design and Programming of Many-Core Focal-Plane Vision Processors
abstract
The Focal-Plane Image Processing area aims to bring processing elements as near as possible to the pixels and to the camera's focal-plane. Most of the works reported in the literature uses only simple processing elements, in general analog ones, with few flexibility. With the technology advances, a new generation of Vision Processors is emerging. It is expected that multi/many-core systems will be integrated to the pixel sensors, offering several opportunities for parallelism exploration, resulting in high performance and flexible processing systems. The programmability is one of the main problems in this area, since most programmers are not able to create parallel algorithms and applications. In this work, we propose a methodology to the design and programming of many-core focal-plane vision processors. The application is described using a Domain Specific Language, from which the parallelism characteristics are extracted. Afterwards, a new abstract model is derived using techniques such as Program Slicing (PS) and Task-Graph Clustering (TGC). The abstract model is then transformed in a SystemC/TLM2.0 description, in order to allow for different timing accuracy simulations. The results of the simulations are used together with an ASIP design tool in order to determine both the microarchitecture of processing elements and the communication structure of the new system. Finally, from the model derived before, a new source code is generated and programmed into the new platform. In this context, the main concepts and ideas are described in this work, as well as some partial results.
Jones Yudi Mori, Carlos H. Llanos, Michael Hübner 0001
EUC3
2015 A Holistic Approach for Advancing Robots in Ambient Assisted Living Environments
abstract
Due to the demographic change in western society, new challenges regarding healthcare of the elderly population are at the verge of surfacing. Since young people are not capable of sustaining an adequate healthcare for elderly people, new healthcare fields have to be devised. Recent advances in information and communication technology enable the support of elderly people in their domestic environment. The EU project RADIO will design of an old age compliant smart home environment which specializes in fulfilling the needs of elderly people. This is partially achieved through a mobile robot platform which serves as an assistant to the respective elderly person. Apart from this, the robot also functions as a mobile sensor platform. Under this context, unobtrusiveness is of paramount importance since the robot should be a natural participant of patients' daily life. This paper discusses such a healthcare facility, analyses its requirements and poses the challenges towards this direction.
Fynn Schwiegelshohn, Philipp Wehner, Jens Rettkowski, Diana Göhringer, Michael Hübner 0001, Georgios Keramidas, Christos P. Antonopoulos, Nikos S. Voros
EUC5
2014 Multi-FPGA reconfigurable system for accelerating MATLAB simulations
abstract
The use of reconfigurable FPGA devices to support the execution of computationally intensive software tasks is discussed in this paper. A system architecture consisting of multiple serially-connected FPGAs is developed, where each FPGA holds a pool of reconfigurable regions. An accelerator can be reconfigured into a region, replaced or discarded at runtime. Configurable connection blocks are responsible of directing data between any two accelerators. The whole system is connected via PCIe-interface to a host PC, where a middleware layer hides all hardware management operations, e.g. routing the data sent among the accelerators, and provides the end-user with an API to use the whole system. Recently, the very fast interfaces for reconfiguring parts of the used FPGAs minimize the overhead caused for hardware modifications. In addition, a manual design of hardware accelerators is not more needed with the continuously improving quality of high-level synthesis tools. In this paper, we considered the case where our system is used within MATLAB. We build a small library to compare and improve upon the execution times of some often used functions.
Muhammed Al Kadi, Max Ferger, Volker Stegemann, Michael Hübner 0001
FPL4
2014 Future Trends on Adaptive Processing Systems
abstract
Today ubiquitous computing is steadily growing in daily life, leading to an increasing need of resource awareness especially for devices with limited energy source. The running applications may differ significantly in their requirements and priority and these variations can occur during a single running application as well. Apart from the applications' constraints, there are regularly restrictions regarding the power source, which can vary during runtime. To tackle these issues, the design should focus on systems that can dynamically optimize themselves, according to the current requirements. Adaptive processors offer great advantages, such as the possibility to reconfigure the microarchitecture to specific needs online. Thus, it is possible for the system to be optimized for each stage of the executing applications and the environmental conditions. The design space of adaptive systems is vast if many parameters can be adjusted during runtime. In order to develop and operate these systems, several different approaches must be taken into account. In this work, new design concepts are presented, together with a discussion of future trends in the field of adaptive processing systems.
Benedikt Janßen, Jones Yudi Mori, Osvaldo Navarro, Diana Göhringer, Michael Hübner 0001
ISPA5
2014 A Framework for Supporting Adaptive Fault-Tolerant Solutions
abstract
For decades, computer architects pursued one primary goal: performance. The ever-faster transistors provided by Moore's law were translated into remarkable gains in operation frequency and power consumption. However, the device-level size and architecture complexity impose several new challenges, including a decrease in dependability level due to physical failures. In this article we propose a software-supported methodology based on game theory for adapting the aggressiveness of fault tolerance at runtime. Experimental results prove the efficiency of our solution since it achieves comparable fault masking to relevant solutions, but with significantly lower mitigation cost. More specifically, our framework speeds up the identification of suspicious failure resources on average by 76% as compared to the HotSpot tool. Similarly, the introduced solution leads to average Power×Delay (PDP) savings against an existing TMR approach by 53%.
Kostas Siozios, Dimitrios Soudris, Michael Hübner 0001
ACM Trans. Embed. Comput. Syst.3
2013 Stopping-Free Dynamic Configuration of a Multi-ASIP Turbo Decoder
abstract
The multiplication of wireless standards is introducing the need of flexible and reconfigurable multistandard base band receivers. At the physical layer, multiprocessor turbo decoders have been recently developed in order to provide an answer to the increasing throughput requirement of emerging standards. However these solutions do not sufficiently address reconfiguration performance issues which can be a limiting factor in the future. This work focuses on the design of a reconfigurable multiprocessor architecture for turbo decoding achieving very fast reconfiguration without compromising decoding performances. Dynamic reconfiguration can be performed within a single frame decoding duration opening new perspective for reconfigurable multistandard base band receivers. For that purpose, optimizations at the processing element level and a novel bus-based configuration infrastructure are proposed. Results show that up to 64 processings elements can be dynamically configured in 5.352 μs. This low configuration latency corresponds to a single frame decoding duration when performing 6 decoding iterations for a throughput up to 666 Mbps.
Vianney Lapotre, Purushotham Murugappa, Guy Gogniat, Amer Baghdadi, Michael Hübner 0001, Jean-Philippe Diguet
DSD5
2013 Optimizations for an efficient reconfiguration of an ASIP-based turbo decoder
abstract
The multiplication of wireless standards is introducing the need of flexible multi-standard baseband receivers. A multi-ASIP approach for turbo decoding is an answer to reach high throughput and high flexibility. The increasing demand of throughput for new greedy application on mobile devices and the reduction of latency between two frames create the need of an efficient reconfiguration management of such multi-ASIP platforms. In this paper, we propose to tackle reconfiguration optimization of a multi-standard ASIP for turbo decoding developed during previous work. Results show that for an area overhead of 0.012 mm2in 65 nm CMOS technology, a significant reconfiguration time optimization is achieved thanks to a reduction of the ASIP configuration load of 70%. Moreover, in a multi-ASIP context in which 8 ASIPs are implemented the configuration load is divided by ten thanks to the possibility to use a multicast mechanism for ASIP configuration loading.
Vianney Lapotre, Purushotham Murugappa, Guy Gogniat, Amer Baghdadi, Jean-Philippe Diguet, Jean-Noel Bazin, Michael Hübner 0001
ISCAS7
2013 Introduction to the special section on multiprocessor system-on-chip for cyber-physical systems
abstract
No abstract available.
Michael Hübner 0001, Andreas Herkersdorf
ACM Trans. Embed. Comput. Syst.1
2013 MORPHEUS: A heterogeneous dynamically reconfigurable platform for designing highly complex embedded systems
abstract
Recently, system designers are facing the challenge of developing systems that have diverse features, are more complex and more powerful, with less power consumption and reduced time to market. These contradictory constraints have forced technology providers to pursue design solutions that will allow design teams to meet the above design targets. In that respect, this paper introduces an innovative technology platform, called MORPHEUS, which intents to provide complete design framework for dealing with the aforementioned challenges. MORPHEUS consists of a state of the art architecture that encompasses heterogeneous reconfigurable accelerators for implementing on the same hardware architecture applications with varying characteristics and a tool chain that, through a software oriented approach, eases the implementation of highly complex applications with heterogeneous characteristics. The proposed approach has been tested and evaluated through state of the art cases studies borrowed from complementary application domains.
Nikos S. Voros, Michael Hübner 0001, Jürgen Becker 0001, Matthias Kühnle, Florian Thoma, Arnaud Grasset, Paul Brelet, Philippe Bonnot 0001, Fabio Campi, Eberhard Schüler, Henning Sahlbach, Sean Whitty, Rolf Ernst, Enrico Billich, Claudia Tischendorf, Ulrich Heinkel, Frank Ieromnimon, Dimitrios Kritharidis, Axel Schneider, Joachim Knäblein, Wolfram Putzke-Röming
ACM Trans. Embed. Comput. Syst.2
2013 JITPR: A framework for supporting fast application's implementation onto FPGAs
Harry Sidiropoulos, Kostas Siozios, Peter Figuli, Dimitrios Soudris, Michael Hübner 0001, Jürgen Becker 0001
ACM Trans. Reconfigurable Technol. Syst.5
2012 Invasive manycore architectures
abstract
This paper introduces a scalable hardware and software platform applicable for demonstrating the benefits of the invasive computing paradigm. The hardware architecture consists of a heterogeneous, tile-based manycore structure while the software architecture comprises a multi-agent management layer underpinned by distributed runtime and OS services. The necessity for invasive-specific hardware assist functions is analytically shown and their integration into the overall manycore environment is described.
Jörg Henkel, Andreas Herkersdorf, Lars Bauer, Thomas Wild, Michael Hübner 0001, Ravi Kumar Pujari, Artjom Grudnitsky, Jan Heisswolf, Aurang Zaib, Benjamin Vogel, Vahid Lari, Sebastian Kobbe
ASP-DAC5
2012 Virtualized on-chip distributed computing for heterogeneous reconfigurable multi-core systems
abstract
Efficiently managing the parallel execution of various application tasks onto a heterogeneous multi-core system consisting of a combination of processors and accelerators is a difficult task due to the complex system architecture. The management of reconfigurable multi-core systems which exploit dynamic and partial reconfiguration in order to, e.g. increase the number of processing elements to fulfill the performance demands of the application, is even more complicated. This paper presents a special virtualization layer consisting of one central server and several distributed computing clients to virtualize the complex and adaptive heterogeneous multi-core architecture and to autonomously manage the distribution of the parallel computation tasks onto the different processing elements.
Stephan Werner 0002, Oliver Wolf, Diana Göhringer, Michael Hübner 0001, Jürgen Becker 0001
DATE4
2012 From Scilab to High Performance Embedded Multicore Systems: The ALMA Approach
abstract
The mapping process of high performance embedded applications to today's multiprocessor system on chip devices suffers from a complex tool chain and programming process. The problem here is the expression of parallelism with a pure imperative programming language which is commonly C. This traditional approach limits the mapping, partitioning and the generation of optimized parallel code, and consequently the achievable performance and power consumption of applications from different domains. The Architecture oriented paraLlelization for high performance embedded Multicore systems using scilAb (ALMA) European project aims to bridge these hurdles through the introduction and exploitation of a Scilab-based toolchain which enables the efficient mapping of applications on multiprocessor platforms from high level of abstraction. This holistic solution of the toolchain allows the complexity of both the application and the architecture to be hidden, which leads to a better acceptance, reduced development cost, and shorter time-to-market. Driven by the technology restrictions in chip design, the end of exponential growth of clock speeds, and an unavoidable increasing request of computing performance, ALMA is a fundamental step forward in the necessary introduction of novel computing paradigms and methodologies.
Jürgen Becker 0001, Timo Stripf, Oliver Wolf, Michael Hübner 0001, Steven Derrien, Daniel Ménard, Olivier Sentieys, Gerard K. Rauwerda, Kim Sunesen, Nikolaos Kavvadias, Kostas Masselos, George Goulas, Panayiotis Alefragis, Nikos S. Voros, Dimitrios Kritharidis, Nikolaos Mitas, Diana Göhringer
DSD4
2012 Introduction to the Special Issue on ReCoSoC 2011
abstract
No abstract available.
Michael Hübner 0001
ACM Trans. Reconfigurable Technol. Syst.1
2011 Fast Start-up for Spartan-6 FPGAs using Dynamic Partial Reconfiguration
abstract
This paper introduces the first available tool flow for Dynamic Partial Reconfiguration on the Spartan-6 family. In addition, the paper proposes a new configuration method called Fast Start-up targeting modern FPGA architectures, where the FPGA is configured in two-steps, instead of using a single (monolithic) full device configuration. In this novel approach, only the timing-critical modules are loaded at power-up using the first high-priority bitstream, while the non-timing critical modules are loaded afterwards. This two-step or prioritized FPGA start-up is used in order to meet the extremely tight startup timing specifications found in many modern applications, like PCI-express or automotive applications. Finally, the developed tool flow and methods for Fast Start-up have been used and tested to implement a CAN-based automotive ECU on a Spartan-6 evaluation board (i.e., SP605). By using this novel approach, it was possible to decrease the initial bitstream size and hence, achieve a configuration time speed-up of up to 4.5×, when compared to a standard configuration solution.
Joachim Meyer 0001, Juanjo Noguera, Michael Hübner 0001, Lars Braun, Oliver Sander, R. M. Gil, Rodney Stewart, Jürgen Becker 0001
DATE3
2011 RAMPSoCVM: Runtime Support and Hardware Virtualization for a Runtime Adaptive MPSoC
abstract
Virtualizing complex hardware, such as heterogeneous multiprocessor systems, enables developers to use standard Application Programming Interfaces (APIs) for application integration. Especially, the supply of an Operating System (OS) is well appreciated since many features such as drivers, the runtime environment and scheduling mechanisms are available and well established. For this purpose, Embedded Linux was used as basis OS and extended in order to be able to manage a Runtime Adaptive Multi-Processor System-on-Chip (RAMPSoC) and to provide the standard Message Passing Interface (MPI). This paper describes the adaptation of the Linux kernel supporting MPI with runtime libraries as well as the integration of the software/hardware drivers which supply the message transfer over a reconfigurable and heterogeneous Network-on-Chip (NoC).
Diana Göhringer, Stephan Werner 0002, Michael Hübner 0001, Jürgen Becker 0001
FPL3
2011 Embedded Systems Start-Up under Timing Constraints on Modern FPGAs
abstract
In this paper we present novel techniques, methods and tool flows that enable embedded systems implemented on FPGAs to start-up under tight timing constraints (i.e., hard deadlines). Meeting the application deadline is achieved by exploiting the FPGA programmability in order to implement a two-stage system start-up approach, as well as a suitable memory hierarchy. This reduces the FPGA configuration time as well as the startup time of the embedded software. Thereby the start-up time for timing-critical parts of a design neither dependent on the complexity nor on the start-up time of the complete system. An automotive case study is used to demonstrate the feasibility and quantify the benefits of the proposed approach.
Joachim Meyer 0001, Juanjo Noguera, Michael Hübner 0001, Rodney Stewart, Jürgen Becker 0001
FPL3
2011 A FPGA based fast runtime reconfigurable real-time Multi-Object-Tracker
abstract
This paper presents a real-time Multi-Object-Tracker implemented on a Field Programmable Gate Array (FPGA). This system is able to track three objects simultaneously using different algorithms to get the best result. Each algorithm has its own field of application and the user can decide which algorithm is used for individual objects. Using the dynamic, partial reconfiguration capability of Xilinx FPGAs, the algorithms can be exchanged during runtime without interrupting the object-tracking. To obtain a self reconfigurable system the Internal Configuration Access Port (ICAP) is used. In this application the needed time to exchange the algorithms has to be as short as possible. In this paper we present a design to achieve the theoretical maximum throughput of the ICAP of 400 MB/s.
Matthias Rümmele-Werner, Thomas Perschke, Lars Braun, Michael Hübner 0001, Jürgen Becker 0001
ISCAS4
2010 High Performance Architectures and Compilers
Pedro C. Diniz, Marco Danelutto, Denis Barthou, Marc Gonzales, Michael Hübner 0001
Euro-Par (1)5
2010 A Design Methodology for Application Partitioning and Architecture Development of Reconfigurable Multiprocessor Systems-on-Chip
abstract
Until today, the efficient partitioning and mapping of applications for multiprocessor systems is a challenging task. The deployment of reconfigurable hardware in this domain helps to meet the application requirements more efficiently due to hardware adaptation at design and runtime, which is not applicable in the traditional multiprocessor domain. To exploit this novel degree of freedom in multiprocessor system-on-chip (MPSoC) technology, a novel design methodology is needed, which helps to hide the complexity of the hardware architecture and its realization alternatives from the developer. This paper shows one approach for such a design methodology for the development of the hardware architecture and the application partitioning and mapping. A novel multistep approach based on hierarchical clustering is used for partitioning of the software application and for configuration of a runtime adaptive multiprocessor system. Furthermore, each application module is then partitioned in a Hardware-Software Codesign process in order to achieve a maximum of performance on the local processors and therefore in general for the MPSoC.
Diana Göhringer, Michael Hübner 0001, Michael Benz, Jürgen Becker 0001
FCCM2
2010 A semi-automatic toolchain for reconfigurable multiprocessor systems-on-chip: architecture development and application partitioning (abstract only)
abstract
S.286
Diana Göhringer, Michael Hübner 0001, Michael Benz, Jürgen Becker 0001
FPGA2
2010 Reconfigurable Hardware for Power-over-Fiber Applications
abstract
In this paper we present an optically powered and motorized video camera system. Energy for the camera sensor is supplied by a glass fiber carrying 800 mW of optical power, which the sensor converts back to 320 mW of electrical power. The specific advantage of this arrangement is galvanic isolation and a very high robustness with respect to electromagnetic interference. We demonstrate that sufficient energy can be transmitted for driving an Actel Igloo FPGA, which performs the necessary signal processing. Additionally, with the help of capacitive energy storage, some small actuators can be supplied which move the camera sensor head. The base station of the system, based on a Xilinx Virtex-5 FPGA, holds a LEON-3 based system-on-chip encoding the incoming VGA video stream into Motion-JPEG formatted data in realtime, which may be directly sent to the internet using an Ethernet interface. The prototype has numerous fields of application where it performs much better than stateof-the-art solutions. Most prominent examples are visual sensors in high voltage areas as well as medical endoscopes.
Michael Dreschmann, Michael Hübner 0001, Moritz Röger, Oliver Sander, Christos Klamouris, Jürgen Becker 0001, Wolfgang Freude, Juerg Leuthold
FPL2
2010 Network Bandwidth Optimization of Ethernet-Based Streaming Applications in Automotive Embedded Systems
abstract
Modern cars are equipped with an increasing number of electronic systems which encompass comfort, security, and infotainment related features. The complexity of the physical and logical network topology as well as the bandwidth requirements are steadily increasing. Therefore, a bus standard that can handle upcoming application requirements and is suitable for the automotive environment is needed. One possibility is the use of Ethernet/IP - both standards are widespread and have proven to be very reliable and robust, however they were formerly not designed for embedded systems. Previous work shows that the achievable bandwidth on embedded systems lies far below the theoretical achievable line rate speed of Ethernet. This paper focuses on improving the networking performance of automotive embedded systems which send Ethernet-based data into a network. It highlights several system variants and evaluates the achievable performance with consideration of the necessary hardware and software changes.
Andreas Kern, Christoph Schmutzler, Thilo Streichert, Michael Hübner 0001, Jürgen Teich
ICCCN4
2010 Design Assurance Strategy and Toolset for Partially Reconfigurable FPGA Systems
abstract
The growth of the Reconfigurable Computing (RC) systems community exposes diverse requirements with regard to functionality of Electronic Design Automation (EDA) tools. Low-level design tools are increasingly required for RC bitstream debugging and IP core design assurance, particularly in multiparty Partially Reconfigurable (PR) designs. While tools for low-level analysis of design netlists do exist, there is increasing demand for automated and customisable bitstream analysis tools. This article discusses the need for low-level IP core verification within PR-enabled FPGA systems and reports FDAT (FPGA Design Analysis Tool), a versatile, modular and open tools framework for low-level analysis and verification of FPGA designs. FDAT provides a set of high-level Application Programming Interfaces (APIs) abstracting the Xilinx FPGA fabric, the implemented design (e.g., placed and routed netlist) and the related bitstream. A lightweight graphic front-end allows custom visualisation of the design within the FPGA fabric. The operation of FDAT is governed by “recipe” scripts which support rapid prototyping of the abstract algorithms for system-level design verification. FDAT recipes, being Python scripts, can be ported to embedded FPGA systems, for example, the previously reported Secure Reconfiguration Controller (SeReCon) which enforces an IP core spatial isolation policy in order to provide run-time protection to the PR system. The paper illustrates the application of FDAT for bit-pattern analysis of Virtex-II Pro and Virtex-5 inter-tile routing and verification of the spatial isolation between designs.
Krzysztof Kepa, Fearghal Morgan, Krzysztof Kosciuszkiewicz, Lars Braun, Michael Hübner 0001, Jürgen Becker 0001
ACM Trans. Reconfigurable Technol. Syst.5
2009 Dynamic reconfigurable mixed-signal architecture for safety critical applications
abstract
Current trends show, it is increasingly difficult to manage the constraints of costs, power consumption, size and more than everything else, functional safety, with conventional architectures. This paper presents a new architecture to deal with the current and upcoming requirements in safety critical applications. It proposes the use of diverse redundancy with digital and analog channels, to detect random hardware failures as well as systematic failures. That will increase the functional safety. By exploiting the ability of dynamic and partial hardware reconfiguration of FPGA and FPAA and by using the appropriate failure recovery scenario, the system availability can also be increased. Furthermore, the architecture offers the possibility to combine high accuracy with short response time.
Romuald Girardey, Michael Hübner 0001, Jürgen Becker 0001
FPL2
2009 Star-Wheels Network-on-Chip featuring a self-adaptive mixed topology and a synergy of a circuit - and a packet-switching communication protocol
abstract
Multiprocessor System-on-Chip is a promising realization alternative for the next generation of computing architectures providing the required data processing performance in high performance computing applications. Numerous scientists from industry and academic institutions investigate and develop novel processing elements and accelerators as can be seen in real devices like IBM's Cell or nVIDIA's Tesla GPU. Nevertheless, the on-chip communication of these multiple processor elements has to be optimized tailored to the actual requirement of the data to be processed. Network-on-Chip (NoC), Bus-based or even heterogeneous communication on chip often suffer from the fact of being inflexible due to their fixed physical realization. This paper presents a novel approach for a NoC, exploiting circuit-and packed-switched communication as well as a run-time adaptive and heterogeneous topology. An application scenario from image processing exploiting the implemented NoC on an FPGA delivers results like performance data and hardware costs.
Diana Göhringer, Michael Hübner 0001, Jürgen Becker 0001
FPL3
2008 Design Flows, Communication Based Design and Architectures in Automotive Electronic Systems
abstract
Summary form only given. The complete presentation was not made available for publication as part of the conference proceedings. A steadily increasing number of microprocessors and electronic components with the heavy demand of computation performance in automotive electronic systems affect substantially the design of networked ECUs in today as well as future cars. Novel approaches, based on heterogeneous hardware (Coarse- fine Grained reconfigurable Hardware, Microprocessors) could be a solution to handle the computation intensive tasks, e.g. for driver-assistance systems. The challenge here is to find an optimal trade-off between power consumption, cost, performance and flexibility which leads to the question which technology and which distribution (automotive function centralisation - decentralisation trade-offs!) will be targeted in future car electronics. Introducing novel architecture topologies and corresponding tool flows with standardised specification and verification are here severe challenges. A first approach to meet these challenges is the AUTOSAR development partnership, which aims at a standardisation of automotive software architecture. The purpose of this tutorial is to evaluate and discuss new concepts for communication based design of automotive electronic and car network systems, as well as to discuss and envisage future system design in automotive electronics. Both aspects, hardware / software design and tool-integration will be discussed. The main emphasis in this session is design-flow, tool-development, applications and system design. The tutorial is addressed to hardware and system engineers as well as to researchers. A set of presentations intended to set the stage for the discussion, will be followed by a panel where selected world-wide specialists in the field of automotive electronics will discuss the demands and interests of industry on novel technologies and systems and research activities for future automotive systems.
Jürgen Becker 0001, Michael Hübner 0001, Robert Esser, Andreas Herkersdorf, Walter Stechele, Vera Lauer
DATE2
2008 Design of a HW/SW Communication Infrastructure for a Heterogeneous Reconfigurable Processor
abstract
Reconfigurable architectures and NoC (Network-on-Chip) have introduced new research directions for technology and flexibility issues, which have been largely investigated in the last decades. Exploiting run-time adaptivity opens a new area of research by considering dynamic reconfiguration. In this paper, we present the architecture and associated development tools of an heterogeneous reconfigurable SoC focusing on the chosen communication infrastructure. The SOC integrates units of various sizes of reconfiguration granularity. The included NoC approach demonstrates the mentioned benefits and scalability for actual and future SoC design. On a reference CMOS090 implementation the described interconnect system works at the system reference frequency of 200 MHZ sustaining the required run-time bandwidth on a set of reference applications, at a price ≪ 10% in area in power consumption with respect to the overall system.
Antonio Deledda, Claudio Mucci, Arseni Vitkovski, Philippe Bonnot 0001, Arnaud Grasset, Philippe Millet, Matthias Kühnle, Florian Ries, Michael Hübner 0001, Jürgen Becker 0001, Massimo Coppola, Lorenzo Pieralisi, Riccardo Locatelli, Giuseppe Maruccia, Fabio Campi, Tommaso DeMarco
DATE9
2008 Cost-and Power Optimized FPGA based System Integration: Methodologies and Integration of a Low-Power Capacity-based Measurement Application on Xilinx FPGAs
abstract
The application of field programmable gate arrays (FPGAs) in low power and low cost industrial mass products has become an important issue for designers of electronic systems. The flexibility and performance offered by reconfigurable hardware architectures often stands in the opposite to increased power consumption in comparison to application specific integrated circuit (ASIC) solutions. By exploiting the flexibility of reconfigurable hardware architectures, e.g. the capability of run-time HW reconfiguration of some modern FPGA devices, power consumption of FPGA-based solutions can be further decreased. This paper presents an approach for cost- and power optimized system integration of a low-power capacity-based measurement system by exploiting the dynamic and partial reconfiguration capability of Xilinx FPGAs.
Katarina Paulsson, Michael Hübner 0001, Jürgen Becker 0001
DATE2
2008 Fine grain reconfigurable architectures
abstract
In this booth on fine grain reconfigurable architectures, several research groups demonstrate their joint work on operating concepts for managing dynamic and partial reconfiguration, visualization of bitstreams and routing, presenting an application applying dynamic reconfiguration for video engines as well as work on minimization of reconfiguration data. Unique is that all the above four projects present their work using the same reconfigurable FPGA-based fabric called Erlangen slot machine that has also been built within one project just the purpose of experimenting with dynamic fine grain reconfiguration as an interdisciplinary platform.
Josef Angermeier, Mateusz Majer, Jürgen Teich, Lars Braun, Tobias Schwalb, Philipp Graf, Michael Hübner 0001, Jürgen Becker 0001, Enno Lübbers, Marco Platzner, Christopher Claus, Walter Stechele, Andreas Herkersdorf, Markus Rullmann, Renate Merker
FPL7
2008 Data path driven waveform-like reconfiguration
abstract
The Xilinx Virtex FPGA family provides the capability to perform dynamic partial hardware reconfiguration (DPR). This implies that parts of the system can by dynamically reprogrammed while the rest of the system components continue their execution without being interrupted. Such reconfigurable FPGA systems are becoming more and more common for applications that require a high degree of run-time flexibility. One major research task in this area is to decrease the overhead caused by the reconfiguration duration. This can be done by increasing the reconfiguration rate, which means increasing the system performance when performing the reconfiguration. This paper presents an alternative approach which aims at decreasing the influence of the reconfiguration, by carefully dividing the reconfigurable modules according to the specific data graph and to start processing the data while the following parts of the data graph are still being reconfigured. This prevents data from being stalled and waiting for the reconfiguration to complete. The suggested approach is referred to as waveform-like reconfiguration, since the data processing closely follows the reconfiguration process.
Lars Braun, Katarina Paulsson, Herrmann Krömer, Michael Hübner 0001, Jürgen Becker 0001
FPL4
2008 A multi-platform controller allowing for maximum Dynamic Partial Reconfiguration throughput
abstract
Dynamic and Partial Reconfiguration (DPR) is a special feature offered by Xilinx Field Programmable Gate Arrays (FPGAs), giving the designer the ability to reconfigure a certain portion of the FPGA during run-time without influencing the other parts. This feature allows the hardware to be adaptable to any potential situation. For some applications, such as video-based driver assistance [1], the time needed to exchange a certain portion of the device might be critical. This paper addresses problems, limitations and results of on-chip reconfiguration that enable the user to decide whether DPR is suitable for a certain design prior to its implementation. A method is therefore introduced to calculate the expected reconfiguration throughput and latency. In addition, an IP core is presented that enables fast on-chip DPR close to the maximum achievable speed. Compared to an alternative state-of-the art realization, an increase in speed by a factor of 58 can be obtained.
Christopher Claus, Walter Stechele, Lars Braun, Michael Hübner 0001, Jürgen Becker 0001
FPL5
2008 New dimensions for multiprocessor architectures: Ondemand heterogeneity, infrastructure and performance through reconfigurability - the RAMPSoC approach
abstract
Multiprocessor hardware architectures enable to distribute tasks of an application to several microprocessors, in order to exploit parallelism for accelerating the performance of computation. Especially for the application domain of image data processing, where computation performance is a crucial factor to keep the real-time requirements, this approach is a promising solution for the assembly of high sophisticated algorithms e.g. for object tracking. Changing requirements and the necessary implementation of the tasks in terms of modified algorithms, precision and communication needs to be handled by software and hardware adaptation in state of the art architectures. Field programmable gate arrays (FPGAs) enable to exploit the adaptation of hardware cores and the software running on embedded microprocessor cores on an integrated multiprocessor system.
Diana Göhringer, Michael Hübner 0001, Thomas Perschke, Jürgen Becker 0001
FPL2
2008 Exploitation of dynamic and partial hardware reconfiguration for on-line power/performance optimization
abstract
This paper presents the results from research work done in the field of reconfigurable architectures and systems. Dynamic and partial reconfiguration has mainly been investigated as a way to configure functionalities in hardware on-demand, controlled either by the user or by the system itself. This paper presents work that was aimed at applying hardware reconfiguration even for run-time adaptation of functional implementation in order to enable self-optimization of power and performance according to the run-time specific requirements of the application.
Katarina Paulsson, Michael Hübner 0001, Jürgen Becker 0001
FPL2
2008 Reducing latency times by accelerated routing mechanisms for an FPGA gateway in the automotive domain
abstract
In todays and future automotive electric/electronic architectures the central gateway is one of the key components. The introduction of high performance bus systems like FlexRay and Ethernet, as well as new applications, creates additional requirements for gateway systems. The usage of reconfigurable hardware gives an interesting alternative to existing microcontroller based solutions. A modular gateway prototype based on a field programmable gate array (FPGA) with specialized routing modules allows a significant speed up compared to microcontroller solutions. The architecture of the routing hardware modules as well as the most relevant implementation details are described in this paper. The complete routing functionality was implemented and tested under series constraints. A performance comparison shows significant speedups. Our toolflow for routing table generation is presented in addition. A final version of the gateway has been successfully integrated into a modern mid class vehicle.
Oliver Sander, Michael Hübner 0001, Jürgen Becker 0001, Matthias Traub
FPT2
2008 Runtime adaptive multi-processor system-on-chip: RAMPSoC
abstract
Current trends in high performance computing show, that the usage of multiprocessor systems on chip are one approach for the requirements of computing intensive applications. The multiprocessor system on chip (MPSoC) approaches often provide a static and homogeneous infrastructure of networked microprocessor on the chip die. A novel idea in this research area is to introduce the dynamic adaptivity of reconfigurable hardware in order to provide a flexible heterogeneous set of processing elements during run-time. This extension of the MPSoC idea by introducing run-time reconfiguration delivers a new degree of freedom for system design as well as for the optimized distribution of computing tasks to the adapted processing cells on the architecture related to the changing application requirements. The "computing in time and space"paradigm and the extension with the new degree of freedom for MPSoCs will be presented with the RAMPSoC approach described in this paper.
Diana Göhringer, Michael Hübner 0001, Volker Schatz, Jürgen Becker 0001
IPDPS2
2008 Run-time reconfigurable adaptive multilayer network-on-chip for FPGA-based systems
abstract
Since the 1990s reusable functional blocks, well known as IP-Cores, were integrated on one silicon die. These systems-on-chip (SoC) used a bus-based system for intermodule communication. Technology and flexibility issues forced to introduce a novel communication system called network-on-chip (NoC). Around 1999 this method was introduced and until then it is investigated by several research groups with the aim to connect different IP-Blocks through an effective, flexible and scalable communication network. Exploiting the flexibility of FPGAs, the run-time adaptivity through run-time reconfiguration, opens a new area of research by considering dynamic and partial reconfiguration. This paper presents an approach for exploiting dynamic and partial reconfiguration with Xilinx Virtex-II FPGAs for a multi-layer network-on-chip and the related techniques for adapting the network while run-time to the requirements of an application.
Michael Hübner 0001, Lars Braun, Diana Göhringer, Jürgen Becker 0001
IPDPS1
2008 A framework for dynamic 2D placement on FPGAs
abstract
The presented paper describes an approach of dynamic positioning of functional building blocks on Virtex (Xilinx) FPGAs. The modules can be of a variable rectangular shape. Further, the on-chip location of the area to be reconfigured can be freely chosen, so that any module can be placed anywhere within the defined dynamic region of the FPGA. Thus the utilization of the chip area can be optimized, which in turn reduces e.g. costly area and power consumption. This paper describes a runtime system and the necessary framework, which is able to manage the reconfigurable area. Further it shows how a NoC approach can be applied to shorten wire lengths for communication. This will in turn save routing resources and potentially increases clock speed.
Christian Schuck, Matthias Kühnle, Michael Hübner 0001, Jürgen Becker 0001
IPDPS3
2007 Circuit Switched Run-Time Adaptive Network-on-Chip for Image Processing Applications
abstract
Since the 1990s reusable functional blocks, well known as IP-Cores, have been integrated on one silicon die. These Systems-on-Chip (SoC) used a bus-based system for intermodule communication. Technology, performance and flexibility issues require the introduction of a novel communication system called Network-on-Chip (NoC). Around 1999 this method was introduced and since then has been investigated by several research groups with the aim to connect different IP-Cores through an effective, flexible and scalable communication network. Exploiting the flexibility of FPGAs, the run-time adaptivity through run-time reconfiguration, opens a new area of research by considering dynamic and partial reconfiguration. Since software parts of an electronic system can also be included into reconfigurable hardware by integration of IP-based microcontrollers, the reconfigurable architecture provides a flexible, multi-adaptive heterogeneous platform for HW / SW Co-designs. This paper presents an approach for exploiting dynamic and partial reconfiguration with Xilinx Virtex-II FPGAs for an adaptive circuit switched Network-on-chip and the related techniques for adapting the system during run-time to the requirements of the presented image processing application.
Lars Braun, Michael Hübner 0001, Jürgen Becker 0001, Thomas Perschke, Volker Schatz, Stefan Bach
FPL2
2007 A Graphical Model-Level Debugger for Heterogenous Reconfigurable Architectures
abstract
Graphical modeling languages allow to specify structure and behavior of mixed hardware-and software-systems on high abstraction level and can be automatically rendered into deployable implementations. In this paper we extend a model-based development process by means to debug functionality specified using Matlab Stateflow models in its hardware- and software-implementation on the target system. The user can control and view the system state graphically from the model's level. We introduce an Eclipse based software-tool based on our approach and apply it to a dynamically reconfigurable slot-based FPGA runtime environment.
Philipp Graf, Michael Hübner 0001, Klaus D. Müller-Glaser, Jürgen Becker 0001
FPL2
2007 Implementation of a Virtual Internal Configuration Access Port (JCAP) for Enabling Partial Self-Reconfiguration on Xilinx Spartan III FPGAs
abstract
The exploitation of dynamic and partial hardware reconfiguration on FPGAs is currently being investigated in various research projects, dealing with systems for space applications to automotive and masurement applications. Despite challanges such as a complicated design flow, dynamic reconfigurable systems offer advantages in terms of flexibility and performance. Unfortunately only few kinds of commercial architectures support dynamic and partial reconfiuration, which has lead to Virtex II / IV being main target architectures for this kind of systems. Additionally, the Xilinx Spartan III architecture is dynamically and partially reconfigurable with some limitations, one of them being the lack of an internal configuration port. The Virtex II / IV and V architectures all include the ICAP port, which allows a system to reconfigure itself during run-time without additional external components. Until now, this was not possible on the Spartan III architecture. This paper presents the implementation of a virtual internal configuration port for the Spartan III family of FPGAs. The configuration port was implemented for a hardware reconfigurable measurement system, which is implemented on a Spartan III FPGA due to its cost- and power optimized characteristics.
Katarina Paulsson, Michael Hübner 0001, Günther Auer, Michael Dreschmann, Jürgen Becker 0001
FPL2
2007 On-line Routing of Reconfigurable Functions for Future Self-Adaptive Systems - Investigations within the ÆTHER Project
abstract
The progress in hardware technologies for implementing portable, low power and low cost electronic systems for consumer products has been major the last years. The complexity of embedded systems will further increase at a rate which is not met by the development of advanced CAD tools for managing the large design space. This will likely lead to increased design problems regarding system implementation, test and verification. In the next 15-20 years, it is likely that the consumer products are based on computing devices which are grouped together in networks including thousands or even millions of nodes. The ÆTHER project deals with managing the complexity of such systems based on emerging technologies for future applications. This paper presents how the design complexity can be managed at the hardware level by integrating self-adaptive characteristics, and how the trade-off in performance and flexibility can be optimized to fulfill all application requirements while reducing the design complexity.
Katarina Paulsson, Michael Hübner 0001, Jürgen Becker 0001, Jean-Marc Philippe, Christian Gamrat
FPL2
2007 Communication Architectures for Dynamically Reconfigurable FPGA Designs
abstract
This paper gives a survey of communication architectures which allow for dynamically exchangeable hardware modules. Four different architectures are compared in terms of reconfiguration capabilities, performance, flexibility and hardware requirements. A set of parameters for the classification of the different communication architectures is presented and the pro and cons of each architecture are elaborated. The analysis takes a minimal communication system for connecting four hardware modules as a common basis for the comparison of the diverse data given in the papers on the different architectures.
Thilo Pionteck, Carsten Albrecht, Roman Koch, Erik Maehle, Michael Hübner 0001, Jürgen Becker 0001
IPDPS5
2007 New tool support and architectures in adaptive reconfigurable computing
abstract
Novel methods and reconfigurable architectures provide an increased design space by exploiting the dynamic and partial reconfiguration of hardware. The multi-adaptivity of this heterogeneous reconfigurable architectures reaches from adaptation to performance requirements over adaptation to power consumption in relation to an available amount of energy to adaptation to not predictable requirements from the user. Especially the unpredictable demands and requirements to a computing architecture require a high and filigree adaptivity in order to find an optimized point of operation while run- time. Additional to this issue the increased availability of electronic systems comes by introduction of novel methods for failure redundancy which can be seen as an application of this multi-adaptive system. In this contribution the ideas for a novel system approach will be presented in three parts. First the hardware and methods providing the multi-adaptivity will be presented. This is the basis for higher level design tools and opens a variety of parameters for adaptivity. The mechanisms of reconfigurability will be introduced in detail from basic knowledge to advanced mechanisms and methods. In addition the abstraction levels for manipulation the reconfigurable architecture and points to the tool support for novel reconfigurable FPGA architectures from Xilinx are sketched.
Jürgen Becker 0001, Adam Donlin, Michael Hübner 0001
VLSI-SoC3
2007 Dynamic and Partial FPGA Exploitation
abstract
Today's field programmable gate array (FPGA) architectures, like Xilinx's Virtex-II series, enable partial and dynamic run-time self-reconfiguration. This feature allows the substitution of parts of a hardware design implemented on this reconfigurable hardware, and therefore, a system can be adapted to the actual demands of applications running on the chip. Exploiting this possibility enables the development of adaptive hardware for a huge variety of applications. A novel method for communication interfaces using look up table (LUT)-based communication primitives enables an exact separation of reconfigurable parts and a fast and intelligent bus-system. A new adaptive software/hardware reconfigurable system is presented in this paper, using a real application in the automotive domain implemented on a Xilinx Virtex-II 3000 FPGA to present results.
Jürgen Becker 0001, Michael Hübner 0001, Gerhard Hettich, Rainer Constapel, Joachim Eisenmann, Jürgen Luka
Proc. IEEE2
2006 Elementary block based 2-dimensional dynamic and partial reconfiguration for Virtex-II FPGAs
abstract
The development of field programmable gate arrays (FPGAs) had tremendous improvements in the last few years. They were extended from simple logic circuits to complex systems-on-chip which enable the integration of complete microcontroller systems and their peripheral devices. Virtex-II FPGAs from Xilinx provide the possibility of dynamic and partial reconfiguration. This can be taken advantage of to substitute inactive parts of a hardware system and adapt the complete chip to a different requirement of an application while run-time. Existing approaches allow reconfiguration of slot based systems while run-time. Unfortunately such systems suffer from the fact, that fixed sized reconfigurable slots are not completely utilized by all functional blocks. Therefore a new 2-dimensional approach is necessary to optimize the placement of functions on the reconfiguration area for the FPGA. Benefit is a reduced chip size which leads to a reduction of power dissipation. This paper describes the method and procedure to include a 2-dimensional placement of reconfigurable blocks and the integration to a run-time system.
Michael Hübner 0001, Christian Schuck, Jürgen Becker 0001
IPDPS1
2004 Partial and Dynamically Reconfiguration of Xilinx Virtex-II FPGAs
Brandon Blodget, Christophe Bobda, Michael Hübner 0001, Adronis Niyonkuru
FPL3
2004 Scalable Application-Dependent Network on Chip Adaptivity for Dynamical Reconfigurable Real-Time Systems
Michael Hübner 0001, Michael Ullmann, Lars Braun, A. Klausmann, Jürgen Becker 0001
FPL1
2004 On-Demand FPGA Run-Time System for Dynamical Reconfiguration with Adaptive Priorities
Michael Ullmann, Michael Hübner 0001, Björn Grimm, Jürgen Becker 0001
FPL2
2004 An FPGA Run-Time System for Dynamical On-Demand Reconfiguration
abstract
Summary form only given. The handling of an increasing number of automotive comfort functionalities has become a significant problem for the most automobile manufacturers since communication, power consumption, available space and cost become important issues for a growing number of engine control units. Our contribution presents a first approach for a flexible versatile FPGA-based run-time system supporting a resource saving function multiplex.
Michael Ullmann, Michael Hübner 0001, Björn Grimm, Jürgen Becker 0001
IPDPS2
2003 Real-Time Dynamically Run-Time Reconfiguration for Power-/Cost-optimized Virtex FPGA Realizations
Jürgen Becker 0001, Michael Hübner 0001, Michael Ullmann
VLSI-SOC2