Mario Porrmann

dblp:37/4663 · DBLP profile ↗
← Back
53ranked-venue papers
3as first author
5since 2021 · last 2023
0000-0003-1005-5753ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 43 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 6 · 2 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 VEDLIoT: Next generation accelerated AIoT systems and applications
abstract
The VEDLIoT project aims to develop energy-efficient Deep Learning methodologies for distributed Artificial Intelligence of Things (AIoT) applications. During our project, we propose a holistic approach that focuses on optimizing algorithms while addressing safety and security challenges inherent to AIoT systems. The foundation of this approach lies in a modular and scalable cognitive IoT hardware platform, which leverages microserver technology to enable users to configure the hardware to meet the requirements of a diverse array of applications. Heterogeneous computing is used to boost performance and energy efficiency. In addition, the full spectrum of hardware accelerators is integrated, providing specialized ASICs as well as FPGAs for reconfigurable computing. The project's contributions span across trusted computing, remote attestation, and secure execution environments, with the ultimate goal of facilitating the design and deployment of robust and efficient AIoT systems. The overall architecture is validated on use-cases ranging from Smart Home to Automotive and Industrial IoT appliances. Ten additional use cases are integrated via an open call, broadening the range of application areas.
Kevin Mika, René Griessl, Nils Kucza, Florian Porrmann, Martin Kaiser, Lennart Tigges, Jens Hagemeyer, Pedro Trancoso, Muhammad Waqar Azhar, Fareed Qararyah, Stavroula Zouzoula, Jämes Ménétrey, Marcelo Pasin, Pascal Felber, Carina Marcus, Oliver Brunnegård, Olof Eriksson, Hans Salomonsson, Daniel Ödman, Andreas Ask, António Casimiro, Alysson Neves Bessani, Tiago Carvalho 0002, Karol Gugala, Piotr Zierhoffer, Grzegorz Latosinski, Marco Tassemeier, Mario Porrmann, Hans-Martin Heyn, Eric Knauss, Yufei Mao, Franz Meierhöfer
CF28
2023 Evaluation of heterogeneous AIoT Accelerators within VEDLIoT
abstract
Within VEDLIoT, a project targeting the development of energy-efficient Deep Learning for distributed AIoT applications, several accelerator platforms based on technologies like CPUs, embedded GPUs, FPGAs, or specialized ASICs are evaluated. The VEDLIoT approach is based on modular and scalable cognitive IoT hardware platforms. Modular microserver technology enables the integration of different, heterogeneous accelerators into one platform. Benchmarking of the different accelerators takes into account performance, energy efficiency and accuracy. The results in this paper provide a solid overview regarding available accelerator solutions and provide guidance for hardware selection for AIoT applications from far edge to cloud. VEDLIoT is an H2020 EU project which started in November 2020. It is currently in an intermediate stage. The focus is on the considerations of the performance and energy efficiency of hardware accelerators. Apart from the hardware and accelerator focus presented in this paper, the project also covers toolchain, security and safety aspects. The resulting technology is tested on a wide range of AIoT applications.
René Griessl, Florian Porrmann, Nils Kucza, Kevin Mika, Jens Hagemeyer, Martin Kaiser, Mario Porrmann, Marco Tassemeier, Marcel Flottmann, Fareed Qararyah, Muhammad Waqar Azhar, Pedro Trancoso, Daniel Ödman, Karol Gugala, Grzegorz Latosinski
DATE7
2022 FAQ: A Flexible Accelerator for Q-Learning with Configurable Environment
abstract
Reinforcement Learning is an area of machine learning that is concerned with optimizing the behavior of an agent in an environment by maximizing cumulative rewards. This can be done with classical reinforcement learning algorithms such as Q-Learning and SARSA. This paper presents FAQ, a flexible FPGA-based accelerator for the Q-Learning algorithm. The architecture of the accelerator can be configured in multiple ways, like adjusting the bit width of Q-values or changing the number of pipeline stages. The evaluation shows that FAQ achieves 249% higher throughput than state-of-the-art FPGA implementations while decreasing DSP and BRAM utilization. Additionally, a software-configurable environment was implemented, and the whole system was tested on an Ultra96-V2 development board utilizing the PYNQ framework. Compared to a CPU implementation, FAQ is more than 13 times faster, including communication overhead caused by transferring the environment onto the FPGA and reading the resulting Q-table.
Marc Rothmann, Mario Porrmann
ASAP2
2022 VEDLIoT: Very Efficient Deep Learning in IoT
abstract
The VEDLIoT project targets the development of energy-efficient Deep Learning for distributed AIoT applications. A holistic approach is used to optimize algorithms while also dealing with safety and security challenges. The approach is based on a modular and scalable cognitive IoT hardware platform. Using modular microserver technology enables the user to configure the hardware to satisfy a wide range of applications. VEDLIoT offers a complete design flow for Next-Generation IoT devices required for collaboratively solving complex Deep Learning applications across distributed systems. The methods are tested on various use-cases ranging from Smart Home to Automotive and Industrial IoT appliances. VEDLIoT is an H2020 EU project which started in November 2020. It is currently in an intermediate stage with the first results available.
Martin Kaiser, René Griessl, Nils Kucza, Carola Haumann, Lennart Tigges, Kevin Mika, Jens Hagemeyer, Florian Porrmann, Ulrich Rückert 0001, Micha vor dem Berge, Stefan Krupop, Mario Porrmann, Marco Tassemeier, Pedro Trancoso, Fareed Qararyah, Stavroula Zouzoula, António Casimiro, Alysson Neves Bessani, José Cecílio, Stefan Andersson, Oliver Brunnegård, Olof Eriksson, Roland Weiss 0001, Franz Meierhöfer, Hans Salomonsson, Elaheh Malekzadeh, Daniel Ödman, Anum Khurshid, Pascal Felber, Marcelo Pasin, Valerio Schiavoni, Jämes Ménétrey, Karol Gugala, Piotr Zierhoffer, Eric Knauss, Hans-Martin Heyn
DATE12
2021 Energy-efficient FPGA-accelerated LiDAR-based SLAM for embedded robotics
abstract
Being one of the fundamental problems in autonomous robotics, SLAM (Simultaneous Localization and Mapping) algorithms have gained a lot of attention. Although numerous approaches have been presented for determining 6D poses in 3D environments, one of the main challenges that remains is the required combination of real-time processing and high energy efficiency. In this paper, a combination of CPU and FPGA processing is used to tackle this problem, utilizing a reconfigurable SoC. We present a complete solution for embedded LiDAR-based SLAM that uses a global Truncated Signed Distance Function (TSDF) as map representation. A hardware-in-the-loop environment with ROS integration enables efficient evaluation of new variants of algorithms and implementations. Based on benchmark data sets and real-world environments, we show that our approach compares well to established SLAM algorithms. Compared to a software implementation on a state-of-the-art PC, the proposed implementation achieves a 7-fold speed-up and requires 18 times less energy when using a Xilinx UltraScale+ XCZU15EG.
Marcel Flottmann, Marc Eisoldt, Julian Gaal, Marc Rothmann, Marco Tassemeier, Thomas Wiemann, Mario Porrmann
FPT7
2018 Resource-efficient Reconfigurable Computer-on-Module for Embedded Vision Applications
abstract
The paper proposes a novel architecture for a highly customisable FPGA-SoC-based Computer-on-Module (CoM) targeting embedded vision applications. Apart from a Xilinx Zynq SoC, the module integrates an Adapteva Epiphany floating point accelerator in a Toradex Apalis compliant form factor. The CoM has been successfully integrated into two robot platforms to enhance their vision processing capabilities. For evaluation, visually-guided collision avoidance and navigation has been implemented, mimicking the behaviour of insects. The hardware/software partitioning is presented together with a comparison to an HLS-based solution for the given application. The proposed stream-based FPGA implementation achieves a speedup of 721 and an increase in energy efficiency by a factor of 800 compared to an OpenCV-based implementation on one of the embedded ARM processors of the Zynq SoC.
Daniel Klimeck, Hanno Gerd Meyer, Jens Hagemeyer, Mario Porrmann, Ulrich Rückert 0001
ASAP4
2018 LEGaTO: towards energy-efficient, secure, fault-tolerant toolset for heterogeneous computing
abstract
LEGaTO is a three-year EU H2020 project which started in December 2017. The LEGaTO project will leverage task-based programming models to provide a software ecosystem for Made-in-Europe heterogeneous hardware composed of CPUs, GPUs, FPGAs and dataflow engines. The aim is to attain one order of magnitude energy savings from the edge to the converged cloud/HPC.
Adrián Cristal, Osman S. Unsal, Xavier Martorell, Raúl de la Cruz, Leonardo Arturo Bautista-Gomez, Daniel Jiménez-González, Carlos Álvarez 0001, Behzad Salami 0001, Sergi Madonar, Miquel Pericàs, Pedro Trancoso, Micha vor dem Berge, Gunnar Billung-Meyer, Stefan Krupop, Wolfgang Christmann, Frank Klawonn, Amani Mihklafi, Tobias Becker, Georgi Gaydadjiev, Hans Salomonsson, Devdatt P. Dubhashi, Oron Port, Yoav Etsion, Vesna Nowack, Christof Fetzer, Jens Hagemeyer, Thorsten Jungeblut, Nils Kucza, Martin Kaiser, Mario Porrmann, Marcelo Pasin, Valerio Schiavoni, Isabelly Rocha, Christian Göttel, Pascal Felber
CF31
2018 An Analytical Study of Time of Flight Error Estimation in Two-Way Ranging Methods
abstract
In absence of clock synchronization, Two-Way Ranging (TWR) is the most commonly used technique for measuring the distance between two wireless transceivers. The existing time-of-flight (TOF) error estimation model, the IEEE 802.15.4-2011 standard, is specifically based on clock drift error. However, it is insufficient when an in-depth comparative analysis of different TWR methods is required. In this paper, we propose an extended TOF error estimation model for TWR methods, based on the IEEE 802.15.4 standard. Using the proposed model, we perform an analytical study of TOF error estimation among different TWR methods. The model is validated with numerical simulation results. Moreover, we demonstrate the pitfalls of the symmetric double-sided TWR (SDS-TWR) method, which is commonly used to reduce the TOF error due to clock drifts.
Cung Lian Sang, Michael Adams 0002, Timm Hörmann, Marc Hesse, Mario Porrmann, Ulrich Rückert 0001
IPIN5
2018 CoreVA-MPSoC: A Many-Core Architecture with Tightly Coupled Shared and Local Data Memories
abstract
MPSoCs with hierarchical communication infrastructures are promising architectures for low power embedded systems. Multiple CPU clusters are coupled using an Network-on-Chip (NoC). Our CoreVA-MPSoC targets streaming applications in embedded systems, like signal and video processing. In this work we introduce a tightly coupled shared data memory to each CPU cluster, which can be accessed by all CPUs of a cluster and the NoC with low latency. The main focus is the comparison of different memory architectures and their connection to the NoC. We analyze memory architectures with local data memory only, shared data memory only, and a hybrid architecture integrating both. Implementation results are presented for a 28 nm FD-SOI standard cell technology. A CPU cluster with shared memory shows similar area requirements compared to the local memory architecture. We use post place and route simulations for precise analysis of energy consumption on both cluster and NoC level using the different memory architectures. An architecture with shared data memory shows best performance results in combination with a high resource efficiency. On average, the use of shared memory shows a 17.2 percent higher throughput for a benchmark suite of 10 applications compared to the use of local memory only.
Johannes Ax, Gregor Sievers, Julian Daberkow, Martin Flasskamp, Marten Vohrmann, Thorsten Jungeblut, Wayne Kelly, Mario Porrmann, Ulrich Rückert 0001
IEEE Trans. Parallel Distributed Syst.8
2017 From CPU to FPGA - Acceleration of self-organizing maps for data mining
abstract
Big data and machine learning applications are posing steadily increasing challenges to the used compute platforms in terms of performance and energy efficiency. In this paper we utilize the highly scalable heterogeneous server platform RECS for evaluation of a wide variety of hardware platforms ranging from general purpose CPUs via ARM-based SoCs to GPGPUs and FPGAs. The self-organizing map, a popular neural network model for unsupervised clustering and dimensionality reduction, is used as a typical example for machine learning applications in the big data domain. Optimized implementations of the algorithm have been developed for each of the target architectures. An in-depth analysis of the achieved performance and energy efficiency for a wide variety of application parameters shows that no single architecture performs best in terms of energy efficiency for the complete design space. In our study, ARM-based SoCs achieved the highest efficiency for small network sizes while FPGAs and GPGPUs perform best for large data sets. Compared to an implementation based on the Matlab SOM toolbox, our optimized multi-threaded CPU implementation achieves two orders of magnitude higher performance and energy efficiency. Large simulations especially benefit from the FPGA implementation, which outperforms the optimized CPU implementation by a factor of 220 and provides a 28-times higher energy efficiency.
Jan Lachmair, Thomas Mieth, René Griessl, Jens Hagemeyer, Mario Porrmann
IJCNN5
2017 FPGA-based multi-robot tracking
Arif Irwansyah, Omar W. Ibraheem, Jens Hagemeyer, Mario Porrmann, Ulrich Rückert 0001
J. Parallel Distributed Comput.4
2016 The M2DC Project: Modular Microserver DataCentre
abstract
The Modular Microserver DataCentre (M2DC) project will investigate, develop and demonstrate a modular, highly-efficient, cost-optimized server architecture composed of heterogeneous microserver computing resources, being able to be tailored to meet requirements from various application domains such as image processing, cloud computing or HPC. M2DC will be built on three main pillars: a flexible server architecture that can be easily customised, maintained and updated, advanced management strategies and system efficiency enhancements (SEE), well-defined interfaces to surrounding software data centre ecosystem.
Mariano Cecowski, Giovanni Agosta, Ariel Oleksiak, Michal Kierzynka, Micha vor dem Berge, Wolfgang Christmann, Stefan Krupop, Mario Porrmann, Jens Hagemeyer, René Griessl, Meysam Peykanu, Lennart Tigges, Sven Rosinger, Daniel Schlitt, Christian Pieper, Carlo Brandolese, William Fornaciari, Gerardo Pelosi, Robert Plestenjak, Justin Cinkelj, Loïc Cudennec, Thierry Goubier, Jean-Marc Philippe, Udo Janssen, Chris Adeniyi-Jones
DSD8
2015 Evaluation of interconnect fabrics for an embedded MPSoC in 28 nm FD-SOI
abstract
Embedded many-core architectures contain dozens to hundreds of CPU cores that are connected via a highly scalable NoC interconnect. Our Multiprocessor-System-on-Chip CoreVA-MPSoC combines the advantages of tightly coupled bus-based communication with the scalability of NoC approaches by adding a CPU cluster as an additional level of hierarchy. In this work, we analyze different cluster interconnect implementations with 8 to 32 CPUs and compare them in terms of resource requirements and performance to hierarchical NoCs approaches. Using 28 nm FD-SOI technology the area requirement for 32 CPUs and AXI crossbar is 5.59 mm2including 23.61% for the interconnect at a clock frequency of 830 MHz. In comparison, a hierarchical MPSoC with 4 CPU cluster and 8 CPUs in each cluster requires only 4.83 mm2including 11.61% for the interconnect. To evaluate the performance, we use a compiler for streaming applications to map programs to the different MPSoC configurations. We use this approach for a design-space exploration to find the most efficient architecture and partitioning for an application.
Gregor Sievers, Johannes Ax, Nils Kucza, Martin Flasskamp, Thorsten Jungeblut, Wayne Kelly, Mario Porrmann, Ulrich Rückert 0001
ISCAS7
2014 Reconfigurable high performance architectures: How much are they ready for safety-critical applications?
abstract
Reconfigurable architectures are increasingly employed in a large range of embedded applications, mainly due to their ability to provide high performance and high flexibility, combined with the possibility to be tuned according to the specific task they address. Reconfigurable systems are today used in several application areas, and are also suitable for systems employed in safety-critical environments. The actual development trend in this area is focused on the usage of the reconfigurable features to improve the fault tolerance and the self-test and the self-repair capabilities of the considered systems. The state-of-the-art of the reconfigurable systems is today represented by Very Long Instruction Word (VLIW) processors and reconfigurable systems based on partially reconfigurable SRAM-based FPGAs. In this paper, we present an overview and accurate analysis of these two type of reconfigurable systems. The content of the paper is focused on analyzing design features, fail-safe and reconfigurable features oriented to self-adaptive mitigation and redundancy approaches applied during the design phase. Experimental results reporting a clear status of the test data and fault tolerance robustness are detailed and commented.
Davide Sabena, Luca Sterpone, Mario Schölzel, Tobias Koal, Heinrich Theodor Vierhaus, S. Wong, Robért Glein, Florian Rittner, C. Stender, Mario Porrmann, Jens Hagemeyer
ETS10
2014 A Scalable Server Architecture for Next-Generation Heterogeneous Compute Clusters
abstract
Increasing the energy efficiency of today's high-performance computing systems requires new approaches that go beyond homogeneous architectures, which primarily target maximum performance per node. Heterogeneous architectures that can be tailored towards the specific needs of a particular application are a promising alternative to state-of-the-art server systems. In this paper, we present a novel highly-scalable server architecture that seamlessly integrates variable combinations of general purpose CPUs, embedded CPUs, FPGAs, and GPUs. Embedded CPUs based on the latest ARM Cortex-A15 devices with integrated embedded GPUs are combined with FPGA-based reconfigurable SoCs, which can be used for application-specific hardware acceleration. A dedicated monitoring network enables continuous control and fine-grained observation of all relevant system parameters. Communication between the compute nodes is established by a flexible multi-level interconnect that can be adapted to various Ethernet and Infiniband standards. The communication facilities are further enhanced by direct high-bandwidth, low-latency links between the embedded FPGA-based reconfigurable SoCs.
René Griessl, Meysam Peykanu, Jens Hagemeyer, Mario Porrmann, Stefan Krupop, Micha vor dem Berge, Thomas Kiesel, Wolfgang Christmann
EUC4
2014 CoreVA: A Configurable Resource-Efficient VLIW Processor Architecture
abstract
Mobile signal processing applications have a limited energy budget and require resource-efficient processing elements. General purpose VLIW CPUs offer a high energy efficiency and allow for the execution of a wide range of applications in this domain. In this work we present the configurable 32 bit VLIW processor architecture CoreVA. Besides the number of issue slots, it allows for a fine-grained configuration of the amount and characteristics of the processor's functional units (e.g., ALUs, MACs, or LD/ST units). A design-space exploration is performed to evaluate how these functional units impact area and power consumption. The basic configuration with one ALU, MAC, DIV, and LD/ST unit has a power consumption of 11.796 mW and an area of 0.142 mm2 at a clock frequency of 750 MHz in a 28 nm FD-SOI process. The maximum clock frequency in this process node is 833 MHz. To bear a relation of the hardware requirements to possible performance gains of the application, a signal processing algorithm is used as a benchmark to evaluate the energy consumption of different hardware configurations. The lowest energy consumption is observed with a configuration of 4 issue slots using 4 ALUs, 4 MACs, and 2 LD/ST units. This is an improvement by a factor of 1.68 compared to the single issue slot configuration.
Boris Hübener, Gregor Sievers, Thorsten Jungeblut, Mario Porrmann, Ulrich Rückert 0001
EUC4
2013 On-line testing of permanent radiation effects in reconfigurable systems
abstract
Partially reconfigurable systems are more and more employed in many application fields, including aerospace. SRAM-based FPGAs represent an extremely interesting hardware platform for this kind of systems, because they offer flexibility as well as processing power. In this paper we report about the ongoing development of a software flow for the generation of hard macros for on-line testing and diagnosing of permanent faults due to radiation in SRAM-FPGAs used in space missions. Once faults have been detected and diagnosed the flow allows to generate fine-grained patch hard macros that can be used to mask out the discovered faulty resources, allowing partially faulty regions of the FPGA to be available for further use.
Luca Cassano, Dario Cozzi, Sebastian Korf, Jens Hagemeyer, Mario Porrmann, Luca Sterpone
DATE5
2013 A reconfigurable neuroprocessor for self-organizing feature maps
Jan Lachmair, Erzsébet Merényi, Mario Porrmann, Ulrich Rückert 0001
Neurocomputing3
2013 A Novel Fault Tolerant and Runtime Reconfigurable Platform for Satellite Payload Processing
abstract
Reconfigurable hardware is gaining a steadily growing interest in the domain of space applications. The ability to reconfigure the information processing infrastructure at runtime together with the high computational power of today's FPGA architectures at relatively low power makes these devices interesting candidates for data processing in space applications. Partial dynamic reconfiguration of FPGAs enables maximum flexibility and can be utilized for performance optimization, for improving energy efficiency, and for enhanced fault tolerance. To be able to prove the effectiveness of these novel approaches for satellite payload processing, a highly scalable prototyping environment has been developed, combining dynamically reconfigurable FPGAs with the required interfaces such as SpaceWire, MIL-STD-1553B, and SpaceFibre. The developed systems have been enabled to space harsh environments thanks to an analytical analysis of the radiation effects on its most critical reconfigurable components. Aiming at that scope, a new algorithm for the analysis of critical radiation effects, in particular, related to Single Event Upsets (SEUs) and Multiple Event Upsets (MEUs) has been developed to obtain an effective estimation of the radiation impact and enabling the tuning of the component mapping reducing the routing interaction between the reconfigurable placed modules in their different feasible positions. The experimental performance of the system has been evaluated by a proper dynamic reconfiguration scenario, demonstrating a partial reconfiguration at 400 MByte/s, blind and readback scrubbing is supported and the scrub rate can be adapted individually for different parts of the design. The fault tolerance capability has been proven by means of a new analysis algorithm and by fault injection campaigns of SEUs and MCUs into the FPGA configuration memory.
Luca Sterpone, Mario Porrmann, Jens Hagemeyer
IEEE Trans. Computers2
2013 A systematic approach for optimized bypass configurations for application-specific embedded processors
abstract
The diversity of today's mobile applications requires embedded processor cores with a high resource efficiency, that means, the devices should provide a high performance at low area requirements and power consumption. The fine-grained parallelism supported by multiple functional units of VLIW architectures offers a high throughput at reasonable low clock frequencies compared to single-core RISC processors. To efficiently utilize the processor pipeline, common system architectures have to cope with data hazards due to data dependencies between consecutive operations. On the one hand, such hazards can be resolved by complex forwarding circuits (i.e., a pipeline bypass) which forward intermediate results to a subsequent instruction. On the other hand, the pipeline bypass can strongly affect or even dominate the total resource requirements and degrade the maximum clock frequency. In this work the CoreVA VLIW architecture is used for the development and the analysis of application-specific bypass configurations. It is shown that many paths of a comprehensive bypass system are rarely used and may not be required for certain applications. For this reason, several strategies have been implemented to enhance the efficiency of the total system by introducing application-specific bypass configurations. The configuration can be carried out statically by only implementing required paths or at runtime by dynamically reconfiguring the hardware. An algorithm is proposed which derives an optimized configuration by iteratively disabling single bypass paths. The adaptation of these application-specific bypass configurations allows for a reduction of the critical path by 26%. As a result, the execution time and energy requirements could be reduced by up to 21.5%. Using Dynamic Frequency Scaling (DFS) and dynamic deactivation/reactivation of bypass paths allows for a runtime reconfiguration of the bypass system. This ensures the highest efficiency while processing varying applications.
Thorsten Jungeblut, Boris Hübener, Mario Porrmann, Ulrich Rückert 0001
ACM Trans. Embed. Comput. Syst.3
2012 gNBXe - a Reconfigurable Neuroprocessor for Various Types of Self-Organizing Maps
Jan Lachmair, Erzsébet Merényi, Mario Porrmann, Ulrich Rückert 0001
ESANN3
2012 A TCMS-based architecture for GALS NoCs
abstract
In this work we propose a TCMS (Tightly Coupled Mesochronous Synchronizer)-based architecture of Globally-Asynchronous Locally-Synchronous (GALS) Network-on-Chips (NoC). The NoC is based on the GigaNoC approach, a scalable NoC featuring packet-switched wormhole routing. At a clock frequency of 750MHz a link bandwidth of up to 6 GByte/s is achieved. To provide a high computational performance, the processing engines (PEs) are based on the CoreVA VLIW architecture. The resource efficiency of mesochronous (TMCS-based) and asynchronous (FIFO-based) communication links is analyzed. In addition an asynchronous coupling of the PE to the switch boxes is evaluated. This allows for multi-voltage/multi-frequency scenarios, where the performance of each PE is adapted to the current performance requirements. Analyses have shown, that TCMS-based communication links and asynchronously coupled PEs allow for the high efficiency of GALS-based NoCs with moderate additional resource requirements.
Thorsten Jungeblut, Johannes Ax, Mario Porrmann, Ulrich Rückert 0001
ISCAS3
2011 Automatic HDL-Based Generation of Homogeneous Hard Macros for FPGAs
abstract
The regularity of resources found in FPGAs is a unique feature, which can be utilized in a number of applications, e.g., in timing critical applications or applications with a demand for homogeneous routing. Current synthesis tools do not support an automatic generation of homogeneous FPGA designs, such that a time-consuming hand-crafted design is required. We present a tool flow, which automatically generates homogeneous hard macros for Xilinx FPGAs starting from a high-level description, such as VHDL. Key functionalities of the tool flow are a homogeneous placer and a suitable routing algorithm, which aim at maintaining the homogeneity of the resulting hard macro. The place and route tools use a resource library that is automatically generated for the target FPGA family by extracting relevant information from the vendor tools. The tool chain is demonstrated for the design of hard macros for a time-to-digital converter and a tiled partially reconfigurable region. The resulting designs are evaluated with respect to resource requirements and timing constraints.
Sebastian Korf, Dario Cozzi, Markus Köster, Jens Hagemeyer, Mario Porrmann, Ulrich Rückert 0001, Marco D. Santambrogio
FCCM5
2011 Evaluation of Applied Intra-disk Redundancy Schemes to Improve Single Disk Reliability
abstract
Exponentially growing capacities of disk drives have increased the problem that not only a complete disk can fail, but also individual, small groups of sectors can be erroneous. These sector errors are especially critical during RAID rebuilds because they can only be detected when the corresponding sectors are read. Mechanisms to cope with sector errors, therefore, have become an important way to improve disk reliability. One approach to deal with sector errors is the introduction of intra-disk redundancy, where additional redundancy blocks are calculated and stored for each set of disk sectors. Previous studies have introduced intra-disk redundancy schemes and have evaluated their impact on disk reliability. None of these studies has evaluated the influence on disk drive performance or the underlying energy consumption. The study presented in this paper benchmarks existing schemes concerning these metrics. It shows the surprising result that weaker codes combined with newly introduced scrambling techniques can produce faster layouts with similar reliability properties than previously proposed strong codes.
Matthias Grawinkel, Thorsten Schäfer, André Brinkmann, Jens Hagemeyer, Mario Porrmann
MASCOTS5
2011 Applying dynamic reconfiguration in the mobile robotics domain: A case study on computer vision algorithms
abstract
Mobile robots are widely used in industrial environments and are expected to be widely available in human environments in the near future, for example, in the area of care and service robots. This article proposes an implementation for a highly customizable color recognition module based on Field Programmable Gate Array (FPGA) hardware to accomplish tasks like real-time frame processing for image streams. In comparison to a pure software solution on a CPU, an attached FPGA-based hardware accelerator enables real-time image processing and significantly reduces the required computing power of the CPU. Instead, the CPU can be used for tasks that cannot be efficiently implemented on FPGAs, for example, because of a large control overhead. We concentrate on a multirobot scenario where a group of robots follows a human team member by keeping a specific formation in order to support the human in exploration and object detection. Additionally, the robots provide a communication infrastructure to maintain a stable multihop communication network between the human and a base station recording all actions and evaluating the captured images and transmitted data. Depending on the current operating conditions, the robot system has to be able to execute a wide variety of different tasks. Since only a small number of tasks have to be executed concurrently, dynamic reconfiguration of the FPGA can be used to avoid the parallel implementation of all tasks on the FPGA. Within this context, this article discusses application fields where dynamic reconfiguration of FPGA-based coprocessors significantly reduces the CPU load and presents examples of how dynamic reconfiguration can be used in exploration.
Federico Nava, Donatella Sciuto, Marco D. Santambrogio, Stefan Herbrechtsmeier, Mario Porrmann, Ulf Witkowski, Ulrich Rückert 0001
ACM Trans. Reconfigurable Technol. Syst.5
2011 Design Optimizations for Tiled Partially Reconfigurable Systems
abstract
In partially reconfigurable architectures, system components can be dynamically loaded and unloaded allowing resources to be shared over time. Dynamic system components are represented by partial reconfiguration (PR) modules. In comparison to a static system, the design of a partially reconfigurable system requires additional design steps, such as partitioning the device resources into static and dynamic regions. We present the concept of tiled PR regions, which enables a flexible online-placement of PR modules. Dynamic reconfiguration requires a suitable communication infrastructure to interconnect the static and dynamic system components. We present an embedded communication macro, a communication infrastructure that interconnects PR modules in a tiled PR region. Efficient online-placement of PR modules depends not only on the placement algorithm, but also on design-time aspects such as the chosen synthesis regions of the PR modules. We propose a design method for selecting suitable synthesis regions for the PR modules aiming to optimize their placement at run-time.
Markus Köster, Wayne Luk, Jens Hagemeyer, Mario Porrmann, Ulrich Rückert 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2010 High level specification of embedded listeners for monitoring of Network-on-Chips
abstract
Nowadays, the Network-on-Chip (NoC) paradigm has become more and more popular for building an on-chip communication infrastructure. Like in every traditional network, debugging and performance monitoring are also very important issues in NoC-based systems. Unfortunately, the design process of monitoring hardware is a time consuming activity. The work presented in this paper is based on a high level specification language, called SiLLis (Simplified Language for Listeners), for the convenient development of generic monitoring hardware. SiLLis allows the designer to define complex filter rules on a high abstraction level. In this way, the design time as well as the bandwidth requirements for monitoring data can be drastically reduced. To present the benefits of SiLLis, we define a performance monitor that is integrated into a NoC-based multiprocessor System-on-Chip and can be used both to analyze the performance of the system and to optimize the routing strategy at run-time. By using SiLLis, the performance monitor can be realized with a area overhead of only 0.58 % per NoC node.
Christoph Puttmann, Mario Porrmann, Paolo Roberto Grassi, Marco D. Santambrogio, Ulrich Rückert 0001
ISCAS2
2010 Design Space Exploration for Memory Subsystems of VLIW Architectures
abstract
In this work we present a design space exploration of the memory subsystem of our configurable CoreVA VLIW architecture. The development of resource efficient processor architectures is based on a two-stage tool flow using a high-level processor specification as a reference. We evaluate several memory configurations like one memory port or two memory ports, as well as different write-miss-allocation modes. Applications ranging from LTE protocol stack over baseband processing up to cryptography and multimedia are evaluated in terms of execution time and energy efficiency. Analyses have shown that the application specific configuration of the memory subsystem can improve energy by up to 25%. Our environment allows the rapid profiling and evaluation of algorithms to choose the most efficient configuration.
Thorsten Jungeblut, Gregor Sievers, Mario Porrmann, Ulrich Rückert 0001
NAS3
2010 Runtime Reconfiguration of Multiprocessors Based on Compile-Time Analysis
abstract
In multiprocessors, performance improvement is typically achieved by exploring parallelism with fixed granularities, such as instruction-level, task-level, or data-level parallelism. We introduce a new reconfiguration mechanism that facilitates variations in these granularities in order to optimize resource utilization in addition to performance improvements. Our reconfigurable multiprocessor QuadroCore combines the advantages of reconfigurability and parallel processing. In this article, a unified hardware-software approach for the design of our QuadroCore is presented. This design flow is enabled via compiler-driven reconfiguration which matches application-specific characteristics to a fixed set of architectural variations. A special reconfiguration mechanism has been developed that alters the architecture within a single clock cycle. The QuadroCore has been implemented on Xilinx XC2V6000 for functional validation and on UMC’s 90nm standard cell technology for performance estimation. A diverse set of applications have been mapped onto the reconfigurable multiprocessor to meet orthogonal performance characteristics in terms of time and power. Speedup measurements show a 2--11 times performance increase in comparison to a single processor. Additionally, the reconfiguration scheme has been applied to save power in data-parallel applications. Gate-level simulations have been performed to measure the power-performance trade-offs for two computationally complex applications. The power reports confirm that introducing this scheme of reconfiguration results in power savings in the range of 15--24%.
Madhura Purnaprajna, Mario Porrmann, Ulrich Rückert 0001, Michael Hussmann, Michael Thies, Uwe Kastens
ACM Trans. Reconfigurable Technol. Syst.2
2009 Design optimizations to improve placeability of partial reconfiguration modules
abstract
In partially reconfigurable architectures, system components can be dynamically loaded and unloaded allowing resources to be shared over time. This paper focuses on the relation between the design options of partial reconfiguration modules and their placement at run-time. For a set of dynamic system components, we propose a design method that optimizes the feasible positions of the resulting partial reconfiguration modules to minimize position overlaps. We introduce the concept of subregions, which guarantees the parallel execution of a certain number of partial reconfiguration modules for tiled reconfigurable systems. Experimental results, which are based on a Xilinx Virtex-4 implementation, show that at run-time the average number of available positions can be increased up to 6.4 times and the number of placement violations can be reduced up to 60.6%.
Markus Köster, Wayne Luk, Jens Hagemeyer, Mario Porrmann
DATE4
2008 Power Aware Reconfigurable Multiprocessor for Elliptic Curve Cryptography
abstract
Reconfigurable architectures are being increasingly used for their flexibility and extensive parallelism to achieve accelerations for computationally intensive applications. Although these architectures provide easy adaptability, it is so with an overhead in terms of area, power and timing, as compared to non-reconfigurable ASICs. Here, we propose a low overhead reconfigurable multiprocessor, which provides both parallelism and flexibility. The architecture has been evaluated for its energy efficiency for a computational intensive algorithm used in elliptic curve cryptography (ECC). Typically, algorithms in ECC exhibit task-level parallelism and demand large amount of computational resources for custom implementations to achieve a significant speedup. A finite field multiplication in GF(2233) was chosen as a sample application to evaluate the performance on the QuadroCore reconfigurable multiprocessor architecture. A three-fold performance improvement as compared to a single processor implementation was observed. Further, via reconfiguration to suit the application, power savings of about 24% were noted in UMC's 90 nm standard cell technology.
Madhura Purnaprajna, Christoph Puttmann, Mario Porrmann
DATE3
2008 SelfS - A real-time protocol for virtual ring topologies
abstract
Real-time automation systems have evolved from centrally controlled sensor-actor systems to complex distributed computing systems. Therefore, the communication system becomes a crucial component that strongly influences performance. In this paper we present a simple distributed communication protocol that meets hard real-time constraints without requiring complex synchronization mechanisms. An advantage of the distributed protocol is that network planning can be reduced to a minimum. The protocol is based on virtual rings and can be easily embedded into arbitrary network topologies. Besides a detailed evaluation and analysis of the protocol, the paper includes lower bounds on jitter and performance for arbitrary communication patterns.
Björn Griese, André Brinkmann, Mario Porrmann
IPDPS3
2007 GigaNoC - A Hierarchical Network-on-Chip for Scalable Chip-Multiprocessors
abstract
Due to the technological progress in the semiconductor industry, more and more components can be integrated on a single die forming a complex System-on-Chip. For enabling an efficient interaction between the various building blocks of today's SoCs, efficient communication structures become more and more essential. In this paper, we present the GigaNoC, a hierarchical Network-on-Chip that is especially suitable for scalable Chip-Multiprocessor architectures. The GigaNoC approach features a packet-switched wormhole routing on-chip network that provides the backbone of our multiprocessor architecture. In order to meet bandwidth requirements of different application domains, our Network-on-Chip is easily scalable and parameterizable in various aspects. This work highlights the communication protocol and shows a performance evaluation for different congestion scenarios. Furthermore, we present an FPGA-based prototypical realization and introduce a debugging and verification environment. Finally, implementation results for a standard cell technology are discussed.
Christoph Puttmann, Jörg-Christian Niemann, Mario Porrmann, Ulrich Rückert 0001
DSD3
2007 A Design Methodology for Communication Infrastructures on Partially Reconfigurable FPGAs
abstract
The ability of partial reconfiguration of today's FPGAs allows the exchange of dynamic system components at run-time, which enables the realization of self-reconfigurable systems. To ease the design of a partially reconfigurable system this paper presents an integrated design flow for reconfigurable architectures. The design flow includes tools for system partitioning, floorplanning, and automatic generation of configuration data for the static and the dynamic system components. Furthermore, the design flow comprises the implementation of a homogeneous on-chip communication infrastructure, which is used to interconnect the dynamic system components placed at run-time. For the design of such an on-chip communication infrastructure a layer model is introduced, which divides the communication into five different layers of abstraction. As an example a communication infrastructure is realized on a Xilinx Virtex-2 FPGA based on the Wishbone protocol. A tristate-based and a slice-based implementation are presented and analyzed with respect to efficiency.
Jens Hagemeyer, Boris Kettelhoit, Markus Köster, Mario Porrmann
FPL4
2007 Partial Dynamic Reconfiguration in a Multi-FPGA Clustered Architecture Based on Linux
abstract
Dynamically reconfigurable hardware allows for implementing systems that can be adapted at run-time according to the needs of the user. This paper presents an architecture that is composed of multiple FPGAs that are connected to an embedded processor. Thus, the architecture is referred to as a multi-FPGA clustered architecture (MFCA). All FPGAs can be partially and dynamically reconfigured to integrate user-defined IP-cores into the system at run-time. For the resource management and communication management we have implemented a Linux operating system on the embedded processor that can be used to control the reconfiguration of the FPGAs by means of simple function calls. Furthermore, the Linux OS completely hides the physical infrastructure of the MFCA from user applications, offering a consistent interface to utilize partial reconfiguration.
Vincenzo Rana, Marco D. Santambrogio, Donatella Sciuto, Boris Kettelhoit, Markus Köster, Mario Porrmann, Ulrich Rückert 0001
IPDPS6
2007 A design framework for FPGA-based dynamically reconfigurable digital controllers
abstract
During the past years, it has been shown that dynamic reconfiguration of FPGAs can be used to enhance the resource efficiency and flexibility of digital controllers. The authors have developed a system architecture, which allows the reconfiguration of FPGA-implemented controllers during runtime. Depending on the operating regions of the controlled plant different controllers can be dynamically loaded into the system. In this paper we present a design flow that enables an automated generation of such partial controllers. Furthermore, a high-level design entry allows a comfortable simulation of the controllers with sophisticated tools such as Matlab Simulink.
Carlos Paiz, Boris Kettelhoit, Mario Porrmann
ISCAS3
2007 Resource efficiency of the GigaNetIC chip multiprocessor architecture
Jörg-Christian Niemann, Christoph Puttmann, Mario Porrmann, Ulrich Rückert 0001
J. Syst. Archit.3
2006 A Layer Model for Systematically Designing Dynamically Reconfigurable Systems
abstract
Partial and dynamic reconfiguration significantly enhances the potential of FPGAs, which has been shown in various prototypic implementations in the past. In this paper the authors introduce a new methodology that eases the design of dynamically reconfigurable systems. It is based on a layer model that systematically abstracts from the underlying reconfigurable hardware to the application that wants to use a dynamically loaded hardware module. With six specified layers and well defined interfaces between these layers we reduce the error-proneness of the system design while increasing the reusability of existing system components. The authors demonstrate the benefits of this design methodology with two example designs: a system-on-chip implementation and a multi-FPGA approach.
Boris Kettelhoit, Mario Porrmann
FPL2
2006 Dedicated module access in dynamically reconfigurable systems
abstract
Modern FPGAs, such as the Xilinx Virtex-II series, offer the feature of partial and dynamic reconfiguration, allowing to load various hardware configurations (i.e., HW modules) during run-time. To enable communication with these modules and for controlling purposes, dedicated access to each module as well as dedicated signals to control the global communication are required. This paper discusses several ways of implementing dedicated signals and addresses the impact on dynamically reconfigurable systems. Two new approaches are introduced, which allow a permanent access to the modules and to the communication infrastructure even during reconfiguration
Jens Hagemeyer, Boris Kettelhoit, Mario Porrmann
IPDPS3
2006 Bio-inspired massively parallel architectures for nanotechnologies
abstract
Massively parallel single-chip multiprocessors (CMP) share a number of traits with biological systems such as neural networks. These biological systems have therefore inspired a number of concepts that may help to overcome some of the problems that will come up in future circuit technologies. In this work we present a first comparison of CMPs based on processor cores of different complexity and estimate the efficiency of CMPs with regards to overall performance and energy consumption. The analysis is based on an analytical model of chip multiprocessing that can help to estimate the runtime and energy consumption of different parallel algorithms. As in previous work we will use the GigaNetIC architecture as a basis for the different CMP architectures
Björn Jäger, Mario Porrmann, Ulrich Rückert 0001
ISCAS2
2005 Context Saving and Restoring for Multitasking in Reconfigurable Systems
abstract
Today's Field Programmable Gate Arrays (FPGAs) can be reconfigured partially, which makes it possible to share resources between various functional modules (hardware tasks) over time. This concept is well known in the area of conventional operating systems. However, in order to transfer resource sharing concepts to operating systems on FPGAs, several underlying mechanisms have to be developed. One of these mechanisms is to suspend hardware tasks and restart them at another time and/or another area of the FPGA. Addressing this problem, this paper discusses ways to save and restore the state information of a hardware task. Afterwards, an implementation of a state relocation mechanisms is presented that uses the standard configuration port. In contrast to similar approaches, we significantly reduce the amount of readback data by reading only those configuration frames that contain state information. We finally determine the time overhead for task relocation, which is essential for most multitasking concepts, like defragmentation.
Heiko Kalte, Mario Porrmann
FPL2
2005 Task Placement for Heterogeneous Reconfigurable Architectures
Markus Köster, Mario Porrmann, Heiko Kalte
FPT2
2005 Defragmentation Algorithms for Partially Reconfigurable Hardware
abstract
Dynamic reconfiguration is a promising approach for resource efficient utilization of microelectronic systems. Standard platforms for partial dynamic reconfiguration are field-programmable gate arrays (FPGAs). Multiple hardware tasks can share the same FPGA resources over time, which increases the device utilization in comparison to non-reconfigurable systems. Although, similar resource management is already known in the area of operating systems, there is a requirement to adapt these concepts to the special needs of dynamically reconfigurable systems. Additionally, there is a lack of underlying mechanisms, e.g., to suspend hardware tasks and restart them at a different position within the FPGA. In this article we introduce a mechanism for task relocation that includes saving and restoring of state information of the task. Based on this approach we address the problem of defragmentation. We present defragmentation algorithms that minimize different types of costs. With the help of a detailed simulation model and a benchmark, we finally provide realistic simulation results and compare the different algorithms.
Markus Köster, Heiko Kalte, Mario Porrmann, Ulrich Rückert 0001
VLSI-SoC3
2004 A Mapping Strategy for Resource-Efficient Network Processing on Multiprocessor SoC
abstract
Hardware architectures based on a field of hardware-extended processors can provide flexible computing power for applications where parallelism can be exploited. For multiprocessors, the assignment of functionality to execution units can have a great impact on the performance. Additionally, finding the optimal mapping can be a time-consuming task. We present a multiprocessor architecture along with a suitable design method that includes an automated solution to the mapping problem. Our hardware architecture employs a network-on-chip (NoC) to achieve a high degree of scalability for the application and for the system in respect to future integration technologies. We also show how to reduce the packet buffer requirements with a proper scheduling strategy and present first estimates for the resource consumption of an application targeted for mobile networking.
Matthias Grünewald, Jörg-Christian Niemann, Mario Porrmann, Ulrich Rückert 0001
DATE3
2004 Hardware Support for Dynamic Reconfiguration in Reconfigurable SoC Architectures
Björn Griese, Erik Vonnahme, Mario Porrmann, Ulrich Rückert 0001
FPL3
2004 Study on column wise design compaction for reconfigurable systems
abstract
Some of currently available field programmable gate arrays (FPGAs) can be reconfigured partially, which makes it possible to build up dynamic systems that can be adapted to changing demands during runtime. One basic aspect of such a system is the way the dynamic hardware modules are placed on the FPGA. As most FPGAs offer partial reconfiguration in a column wise manner, a 1D placement of column wise implemented modules seems to be promising. Within This work we present a design study that determines the effects of a column wise module implementation on the resulting frequency and power consumption.
Heiko Kalte, Gareth Lee, Mario Porrmann, Ulrich Rückert 0001
FPT3
2004 gNBX - reconfigurable hardware acceleration of self-organizing maps
abstract
In this work a new FPGA based hardware accelerator (gNBX) for self-organizing maps is introduced. New principles for hardware acceleration of self-organizing maps, which increase the degree of parallelity and therefore the acceleration gain was presented. Our technology independent design description can be mapped on application specific integrated circuits if very high performance is required, as well as on field programmable gate arrays (FPGAs), which offer mid level performance (a speed up factor of up to 70 in comparison with PCs for typical datasets is achieved) at relatively low costs. Additionally, FPGAs offer the flexibility to adapt the hardware to the changing requirements of the application during runtime. Therefore, the hardware can be exploited optimally at all times during the simulation process. Several benchmark scenarios with well known datasets shows the performance of our system.
Christopher Pohl, Marc Franzmeier, Mario Porrmann, Ulrich Rückert 0001
FPT3
2004 System-on-Programmable-Chip Approach Enabling Online Fine-Grained 1D-Placement
abstract
Summary form only given. The increasing logic density of current FPGAs (field programmable gate arrays) enables the integration of whole systems on one programmable chip. Some of these FPGAs provide the additional feature of partial dynamic reconfiguration, which permits to change parts of the device while other parts keep working. Combining the features of system level density and partial dynamic reconfiguration enables the integration of dynamic systems that can be adopted to changing demands during runtime. A lot of theoretical work in this challenging research area has been done on efficiently placing and scheduling modules on the FPGA area. However, there is a lack of applied approaches that can be realized by existing tools and FPGAs. We present a new, realizable approach for the dynamic system integration on Xilinx Virtex FPGAs. In contrast to the existing approaches that consider fixed slots for the module placement, our approach enables the fine-grained placement of modules with variable width along a horizontal communication infrastructure.
Heiko Kalte, Mario Porrmann, Ulrich Rückert 0001
IPDPS2
2003 A holistic methodology for network processor design
abstract
The GigaNetIC project aims to develop high-speed components for networking applications based on massively parallel architectures. A central part of this project is the design, evaluation, and realization of a parameterizable network processing unit. In this paper we present a design methodology for network processors which encompasses the research areas from the application software down to the gate level of the chip. Key components of this holistic approach have been successfully applied to characteristic examples of architecture refinements.
Olaf Bonorden, Nikolaus Brüls, Uwe Kastens, Dinh Khoi Le, Friedhelm Meyer auf der Heide, Jörg-Christian Niemann, Mario Porrmann, Ulrich Rückert 0001, Adrian Slowik, Michael Thies
LCN7
2003 A massively parallel architecture for self-organizing feature maps
abstract
A hardware accelerator for self-organizing feature maps is presented. We have developed a massively parallel architecture that, on the one hand, allows a resource-efficient implementation of small or medium-sized maps for embedded applications, requiring only small areas of silicon. On the other hand, large maps can be simulated with systems that consist of several integrated circuits that work in parallel. Apart from the learning and recall of self-organizing feature maps, the hardware accelerates data pre- and postprocessing. For the verification of our architectural concepts in a real-world environment, we have implemented an ASIC that is integrated into our heterogeneous multiprocessor system for neural applications. The performance of our system is analyzed for various simulation parameters. Additionally, the performance that can be achieved with future microelectronic technologies is estimated.
Mario Porrmann, Ulf Witkowski, Ulrich Rückert 0001
IEEE Trans. Neural Networks1
2002 A reconfigurable SOM hardware accelerator
Mario Porrmann, Marc Franzmeier, Heiko Kalte, Ulf Witkowski, Ulrich Rückert 0001
ESANN1
2002 Dynamically Reconfigurable Hardware - A New Perspective for Neural Network Implementations
Mario Porrmann, Ulf Witkowski, Heiko Kalte, Ulrich Rückert 0001
FPL1
1998 SOM accelerator system
Stefan Rüping 0002, Mario Porrmann, Ulrich Rückert 0001
Neurocomputing2