EDBT 2026 Demo / reviewers in the wild / expert
Marco Platzner
dblp:68/3728
· DBLP profile ↗
81ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0002-6893-063XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 69 · 15 since 2021Software engineering, systems software and programming languages · 8 · 2 since 2021Artificial intelligence and machine learning · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Theory of computation · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DP-MCTS: Deep Playout-Driven MCTS for Approximate Accelerator DesignabstractApproximate computing offers substantial efficiency gains for modern accelerators, yet register-transfer level design space exploration (DSE) remains challenging due to the exponential growth of approximation choices. Although recent use of machine learning (ML)-based accuracy estimators has reduced exploration runtime, it introduces a new challenge: estimator inaccuracy, which becomes increasingly problematic for larger circuits. In this paper, we propose DP-MCTS, an enhanced Monte Carlo Tree Search framework that rethinks estimator usage by treating ML models as lightweight guidance tools rather than decision makers. DP-MCTS employs deep playouts to provide fast, informative look-ahead along candidate paths, coupled with a promotion mechanism that selectively validates promising nodes through simulation-based accuracy evaluation. This synergy preserves broad design space coverage while mitigating estimator-driven drift, significantly improving exploration robustness and solution quality. Experiments across diverse accelerator benchmarks show that DP-MCTS closely matches or even outperforms the solution quality of an MCTS-based search, achieving up to 6.0% additional area reduction and an order of magnitude lower runtime cost, thereby delivering a substantially improved quality-runtime tradeoff over existing search-based frameworks. Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Sayed Morteza Jawadi, Marco Platzner |
DDECS | 4 |
| 2026 | Reliability Assessment in Approximate Accelerator SynthesisabstractWhile optimizing for core hardware performancerelated target metrics, frameworks for approximate accelerators often overlook the reliability aspect. Approximated implementations obtained by these frameworks can potentially differ in terms of reliability and may impact the reliability of the overall system. In particular, approximation changes the data profiles transmitted between system modules, which can trigger crosstalk on interconnect lines and aggravate electromigration. We propose a two-stage process that performs a reliability assessment of the circuit interconnects after the approximate accelerator synthesis. Our approach aims to find the most reliable solutions from the approximate candidate circuits generated by an automated approximation flow. We then leverage Pareto-filtering to strike a balance between area, reliability, and accuracy. Notably, the selected designs achieve up to a 178% improvement in mission time compared to the original accelerator, and a 68% improvement over designs optimized solely for area. In addition, our methodology allows custom priority settings to be adaptable to a user's preference, thereby leading to circuits that meet diverse design constraints. Our experimental results show the effectiveness of our methodology in achieving superior trade-offs between area, reliability, and accuracy, hence uncovering a new dimension for approximate accelerator design methodologies. Somayeh Sadeghi Kohan, Muhammad Awais 0009, Qazi Arbab Ahmed, Marco Platzner, Sybille Hellebrand, Thorsten Jungeblut, Hans-Joachim Wunderlich |
DDECS | 4 |
| 2026 | A Robotics Middleware for FPGAs Supporting Dynamic Function Exchange and Streaming Data DistributionabstractField Programmable Gate Arrays (FPGAs) have the potential to significantly improve the performance and energy efficiency of robotics applications. Several studies propose frameworks that systematically integrate FPGAs into robotic systems, primarily within the Robot Operating System (ROS), the de facto standard middleware for robotics. Recently, two main research directions have emerged: One explores the use of Dynamic Function Exchange (DFX) to time-multiplex hardware-accelerated functions on an FPGA, while the other focuses on intra-FPGA communication, leveraging interconnections between statically mapped hardware-accelerated functions to reduce communication overhead.In this paper, we present a robotics middleware that supports both DFX and dynamic intra-FPGA communication. The middleware employs a static FPGA layout with multiple reconfigurable slots, complemented by router and buffer infrastructure for streaming-based data distribution. Unlike related work, the call-backs executed in reconfigurable slots, the exchanged messages, and the routing configurations do not need to be known at build time, enabling a high degree of runtime flexibility. We describe the hardware architecture, analyze potential deadlock scenarios, and derive constraints required for deadlock-free operation, which our architecture satisfies. We evaluate the middleware using small-scale test cases and a larger autonomous robot simulation. The results quantify the overhead introduced by the flexible design compared to static mappings and demonstrate performance gains over DFX-only and software-only approaches enabled by dynamic intra-FPGA communication. Alexander Nowosad, Christian Lienen, Marco Platzner |
FCCM | 3 |
| 2025 | A Two-Stage Approximation Methodology for Efficient DNN Hardware ImplementationabstractDeploying high-performance Deep Neural Networks (DNNs) on embedded hardware platforms, such as FPGAs, not only promises low latency and high energy-efficiency implementations, but also poses significant challenges due to the resource constraints of such systems. Recent efforts have focused on approximation techniques at the algorithmic level of DNNs to reduce resource requirements, such as network pruning and quantization of weights and activations. Other approximation techniques that approximate register transfer level (RTL) designs by substituting single components with approximated ones are not yet applied to DNNs, mainly because these techniques do not scale well to the typically huge RTL representations of DNNs.This paper introduces the idea of a two-stage approximation methodology for DNNs that combines algorithmic level with RTL approximation to achieve reduced resource requirements at acceptable accuracy. As a novel instance of that idea, we leverage the LogicNets approach that performs a low-bit quantization to map DNN neurons to LUT representations and the CIRCA framework that provides a search-based approximation flow for RTL designs. To balance hardware efficiency and accuracy, and to achieve acceptable runtimes, we limit the RTL approximation to a subset of neurons identified as resilient using a systematic sampling approach. Experimental results demonstrate the potential of the methodology with up to 15.20% further decrease in resource utilization with a mere 2.15% accuracy drop. Amir Hossein Hadipour, Atousa Jafari, Muhammad Awais 0009, Marco Platzner |
DDECS | 4 |
| 2025 | Swift Synthesis of Approximate Hardware Accelerators Using Generative Adversarial NetworksabstractDeploying modern applications with significant resource demands is often challenging, but approximate designs offer a promising alternative by delivering high performance with minimal compromises in output quality. Traditionally, approximate hardware accelerators have been developed through search-based iterative frameworks, which suffer from long runtimes due to the exponential growth of the design space. A significant portion of the runtime is consumed by either invalid nodes or valid nodes that offer minimal improvements in performance metrics, such as runtime or power consumption. This severely limits the thorough exploration of the design space. In this paper, we introduce a novel approach for synthesizing approximate accelerators that leverages sparsity to reduce the complexity of the design space exploration problem. Our method employs a generative adversarial network (GAN) to rapidly generate a diverse set of high-quality design nodes, eliminating the need for costly node evaluations. This enables the swift creation of approximate accelerators generated for any given error threshold in a fraction of time as compared to a simulation-based framework. We conducted experiments on a suite of benchmarks from real-world domains, demonstrating that our methodology can generate approximate hardware designs with significant area and power savings, comparable to state-of-the-art search-based approaches. In a comparative evaluation against two leading methods, our approach achieved equal or better quality results for two out of four benchmarks while reaching up to 55% area savings, thus effectively demonstrating a new avenue for automated generation of approximate accelerators. Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner |
VLSI-SoC | 3 |
| 2025 | Design Space Exploration for Approximate Circuits via Checkpointing and DNN-Based EstimatorsabstractApproximate computing (AC) is a design paradigm that trades in reductions in hardware area, power, and delay of digital circuits for an increased error. Workflows for synthesizing approximate accelerators are often implemented as search-approximate–evaluate cycles that build a search tree by iteratively approximating components of the original circuit. Design space exploration (DSE) for approximate accelerators is challenging due to the sheer size of the design space and the computational costs for validating the error bound during the search. While early work employed rather slow simulation-based methods, more recent approaches use machine learning (ML) to create fast estimators for validating error bounds. For larger accelerators and design spaces, however, the mispredictions of ML-based methods can steer the search process to explore invalid regions of the design space. In this article, we introduce the frameworkDeepApproxfor approximate accelerator synthesis. The novelty of this framework is its combination of fast ML-based quality estimators with occasional simulations to correct possible mispredictions. The simulations are performed at the so-called checkpoints during the search, and we propose two methods: static and dynamic checkpointing (DC). We present theDeepApproxflow and elaborate on the ML quality estimators and the checkpointing mechanisms. Using a set of smaller and larger benchmarks, we compareDeepApproxwith two state-of-the-art flows: a simulation-based method leveraging Monte Carlo tree search (MCTS) and an ML-based method with fast estimators. Our experiments demonstrate thatDeepApproxcan achieve similar reductions in hardware area and power consumption compared to simulation-based methods, albeit at much lower runtimes, and higher reductions in these metrics compared to methods relying exclusively on ML-based estimators. Thus, by combining ML-based techniques with simulations,DeepApproxprovides a scalable DSE for approximate accelerators. Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | FPGADDS: An Intra-FPGA Data Distribution Service for ROS 2 Robotics ApplicationsabstractModern computing platforms for robotics applications comprise a set of heterogeneous elements, e.g., multi-core CPUs, embedded GPUs, and FPGAs. FPGAs are reprogrammable hardware devices that allow for fast and energy-efficient computation of many relevant tasks in robotics. ROS is the de-facto programming standard for robotics and decomposes an application into a set of communicating nodes. ReconROS is a previous approach that can map complete ROS nodes into hardware for acceleration. Since ReconROS relies on standard ROS communication layers, exchanging data between hardware-mapped nodes can lead to a performance bottleneck. This paper presents fpgaDDS, a lean data distribution service for hardware-mapped ROS 2 nodes. fpgaDDS relies on a customized and statically generated streaming-based communication architecture. We detail this communication architecture with its components and outline its benefits. We evaluate fpgaDDS on a test example and a larger autonomous vehicle case study. Compared to a ROS 2 application in software, we achieve speedups of up to 13.34 and reduce jitter by two orders of magnitude. Christian Lienen, Sorel Horst Middeke, Marco Platzner |
IROS | 3 |
| 2022 | Search space characterization for approximate logic synthesisabstractApproximate logic synthesis aims at trading off a circuit's quality to improve a target metric. Corresponding methods explore a search space by approximating circuit components and verifying the resulting quality of the overall circuit, which is costly. Linus Witschen, Tobias Wiersema, Lucas Reuter, Marco Platzner |
DAC | 4 |
| 2022 | MUSCAT: MUS-based Circuit Approximation TechniqueabstractMany applications show an inherent resiliency against inaccuracies and errors in their computations. The design paradigm approximate computing exploits this fact by trading off the application's accuracy against a target metric, e.g., hardware area. This work focuses on approximate computing on the hard-ware level, where approximate logic synthesis seeks to generate approximate circuits under user-defined quality constraints. We propose the novel approximate logic synthesis method MUSCAT to generate approximate circuits which are valid-by-construction. MUSCAT inserts cutpoints into the netlist to employ the commonly-used concept of substituting connections between gates by constant values, which offers potential for subsequent logic minimization. MUSCAT's novelty lies in utilizing formal verification engines to identify minimal unsatisfiable subsets. These subsets determine a maximal number of cutpoints that can be activated together without resulting in a violation against the user-defined quality constraints. As a result, MUSCAT determines an optimal solution w.r.t. the number of activated cutpoints while providing a guarantee on the quality constraints. We present the method and experimentally compare MUS-CAT's open-source implementation to AIG rewriting and components from the EvoApproxLib. We show that our method improves upon these state-of-the-art methods by achieving up to 80 % higher savings in circuit area at typically much lower computation times. Linus Witschen, Tobias Wiersema, Matthias Artmann, Marco Platzner |
DATE | 4 |
| 2022 | Event-Driven Programming of FPGA-accelerated ROS 2 Robotics ApplicationsabstractMany applications from the robotics domain can benefit from FPGA acceleration. A corresponding key question is not only how to integrate hardware accelerators into software-centric robotics programming environments but also how to integrate more advanced approaches like dynamic partial reconfiguration. Recently, several approaches have demonstrated hardware acceleration for the robot operating system (ROS), the dominant programming environment in robotics. ROS is a middleware layer that features the composition of complex robotics applications as a set of nodes that communicate via mechanisms such as publish/subscribe, and distributes them over several compute platforms. In this paper, we present a novel approach for event-based programming of robotics applications that leverages dynamic partial reconfiguration and ReconROS, a framework for flexibly mapping ROS 2 nodes to either software or reconfigurable hardware. The approach bases on the ReconROS executor that schedules callbacks of ROS 2 nodes and utilizes a reconfigurable slot model and partial runtime reconfiguration to load hardware-based callbacks on demand. We describe the ReconROS executor approach, give design examples, and experimentally evaluate its functionality with examples. Christian Lienen, Marco Platzner |
DSD | 2 |
| 2022 | Integrating Safety Guarantees into the Learning Classifier System XCS
Tim Hansmeier, Marco Platzner |
EvoApplications | 2 |
| 2022 | Automated Framework for Fast Synthesis of Approximate Hardware AcceleratorsabstractGenerating approximate accelerators automatically via libraries of functional units faces combinatorial explosion due to extremely large design space. Moreover, long verification times become bottleneck in the process making it nearly impossible to find best suited approximate instances with exhaustive search. Multiple works try to explore the design space with a greedy-based approach that reduces the complexity of the process, albeit overlooking promising combinations and ultimately producing inferior solutions. This paper proposes an automated framework that handles the design space exploration problem with an learning-based search algorithm i.e., MCTS. Furthermore, by leveraging fast machine learning quality estimators, the framework offers up to 47× reduction of runtime when compared to a simulation-based framework. Muhammad Awais 0009, Marco Platzner |
VLSI-SoC | 2 |
| 2022 | Exploiting Hardware-Based Data-Parallel and Multithreading Models for Smart Edge Computing in Reconfigurable FPGAsabstractCurrent edge computing systems are deployed in highly complex application scenarios with dynamically changing requirements. In order to provide the expected performance and energy efficiency values in these situations, the use of heterogeneous hardware/software platforms at the edge has become widespread. However, these computing platforms still suffer from the lack of unified software-driven programming models to efficiently deploy multi-purpose hardware-accelerated solutions. In parallel, edge computing systems also face another huge challenge: operating under multiple conditions that were not taken into account during any of the design stages. Moreover, these conditions may change over time, forcing self-adaptation mechanisms to become a must. This paper presents an integrated architecture to exploit hardware-accelerated data-parallel models and transparent hardware/software multithreading. In particular, the proposed architecture leverages the ARTICo3framework and ReconOS to allow developers to select the most suitable programming model to deploy their edge computing applications onto run-time reconfigurable hardware devices. An evolvable hardware system is used as an additional architectural component during validation, providing support for continuous lifelong learning in smart edge computing scenarios. In particular, the proposed setup exhibits online learning capabilities that include learning by imitation from software-based reference algorithms. Experimental results show the benefits of the proposed approach, exposing different run-time tradeoffs (e.g., computing performance versus functional correctness of the evolved solutions), and highlighting the benefits of using scalable data-parallel models to perform circuit evolution under dynamically changing application scenarios. Alfonso Rodríguez 0002, Andrés Otero, Marco Platzner, Eduardo de la Torre |
IEEE Trans. Computers | 3 |
| 2022 | Design of Distributed Reconfigurable Robotics Systems with ReconROSabstractRobotics applications process large amounts of data in real time and require compute platforms that provide high performance and energy efficiency. FPGAs are well suited for many of these applications, but there is a reluctance in the robotics community to use hardware acceleration due to increased design complexity and a lack of consistent programming models across the software/hardware boundary. In this article, we present ReconROS , a framework that integrates the widely used robot operating system (ROS) with ReconOS, which features multithreaded programming of hardware and software threads for reconfigurable computers. This unique combination gives ROS 2 developers the flexibility to transparently accelerate parts of their robotics applications in hardware. We elaborate on the architecture and the design flow for ReconROS and report on a set of experiments that underline the feasibility and flexibility of our approach. Christian Lienen, Marco Platzner |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2021 | Malicious Routing: Circumventing Bitstream-level Verification for FPGAsabstractThe battle of developing hardware Trojans and corresponding countermeasures has taken adversaries towards ingenious ways of compromising hardware designs by circumventing even advanced testing and verification methods. Besides conventional methods of inserting Trojans into a design by a malicious entity, the design flow for field-programmable gate arrays (FPGAs) can also be surreptitiously compromised to assist the attacker to perform a successful malfunctioning or information leakage attack. The advanced stealthy malicious look-up-table (LUT) attack activates a Trojan only when generating the FPGA bitstream and can thus not be detected by register transfer and gate level testing and verification. However, also this attack was recently revealed by a bitstream-level proof-carrying hardware (PCH) approach. In this paper, we present a novel attack that leverages malicious routing of the inserted Trojan circuit to acquire a dormant state even in the generated and transmitted bitstream. The Trojan's payload is connected to primary inputs/outputs of the FPGA via a programmable interconnect point (PIP). The Trojan is detached from inputs/outputs during place-and-route and re-connected only when the FPGA is being programmed, thus activating the Trojan circuit without any need for a trigger logic. Since the Trojan is injected in a post-synthesis step and remains unconnected in the bitstream, the presented attack can currently neither be prevented by conventional testing and verification methods nor by recent bitstream-level verification techniques. Qazi Arbab Ahmed, Tobias Wiersema, Marco Platzner |
DATE | 3 |
| 2021 | LDAX: A Learning-based Fast Design Space Exploration Framework for Approximate Circuit SynthesisabstractThe majority of existing frameworks for automated synthesis of Approximate Circuits (AxCs) employ a search-based Design Space Exploration (DSE) approach. This includes an iterative process where approximate circuit instances are created and evaluated in terms of quality and performance metrics. The quality evaluation of each instance via either verification or testing results in extremely long computation times, imposing a practical limit on the number of nodes that can be explored in the design space. To overcome this problem, we exploit Random Forests (RFs) to develop extremely fast estimators to evaluate the quality and performance of AxC instances while avoiding time-consuming computations. Utilizing these estimators we build LDAX, an efficient design space exploration framework that attempts to improve the runtime of the AxCs synthesis process. LDAX is based on the fact that, in an iterative search space exploration, a large number of validations can be skipped by leveraging a high-accuracy predictor. To efficiently explore the design space, which is usually represented by a tree and each node denotes an AxC, we propose a learning-based search technique that can quickly analyze a large set of nodes even very deep nodes in the tree. Our experimental results reveal that LDAX can achieve average speed-up of 41 × on a set of practical benchmarks from different application domains in comparison with the most competitive state-of-the-art framework. Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | DeepWind: An Accurate Wind Turbine Condition Monitoring Framework via Deep Learning on Embedded PlatformsabstractCondition monitoring of critical components in wind turbines is of utmost importance to enable low cost and preventive maintenance and to minimize downtime. In this paper we propose DEEPWIND, an end-to-end condition monitoring and fault detection framework that can be implemented on resource-constrained embedded platforms. DEEPWIND exploits multi-channel convolutional neural networks to automatically extract features from sensor data without any need of human feature engineering, and utilizes the features to classify faults occurring in rotor blades of wind turbines. Experiments on a real-world dataset provided by Weidmüller Monitoring Systems GmbH reveal an average F1-score of 0.94 and thus underline the suitability of the approach for fault detection in wind turbines. Hassan Ghasemzadeh Mohammadi, Rahil Arshad, Sneha Rautmare, Suraj Manjunatha, Maurice Kuschel, Felix Paul Jentzsch, Marco Platzner, Alexander Boschmann, Dirk Schollbach |
ETFA | 7 |
| 2020 | A Hybrid Synthesis Methodology for Approximate CircuitsabstractAutomated synthesis of approximate circuits via functional approximations is of prominent importance to provide efficiency in energy, runtime, and chip area required to execute an application. Approximate circuits are usually obtained either through analytical approximation methods leveraging approximate transformations such as bit-width scaling or via iterative search-based optimization methods when a library of approximate components, e.g., approximate adders and multipliers, is available. For the latter, exploring the extremely large design space is challenging in terms of both computations and quality of results. While the combination of both methods can create more room for further approximations, theDesign Space Exploration ~(DSE) becomes a crucial issue. In this paper, we present such a hybrid synthesis methodology that applies a low-cost analytical method followed by parallel stochastic search-based optimization. We address the DSE challenge through efficient pruning of the design space and skipping unnecessary expensive testing and/or verification steps. The experimental results reveal up to 10.57x area savings in comparison with both purely analytical or search-based approaches. Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | Self-aware Cyber-Physical SystemsabstractIn this article, we make the case for the new class of Self-aware Cyber-physical Systems. By bringing together the two established fields of cyber-physical systems and self-aware computing, we aim at creating systems with strongly increased yet managed autonomy, which is a main requirement for many emerging and future applications and technologies. Self-aware cyber-physical systems are situated in a physical environment and constrained in their resources, and they understand their own state and environment and, based on that understanding, are able to make decisions autonomously at runtime in a self-explanatory way. In an attempt to lay out a research agenda, we bring up and elaborate on five key challenges for future self-aware cyber-physical systems: (i) How can we build resource-sensitive yet self-aware systems? (ii) How to acknowledge situatedness and subjectivity? (iii) What are effective infrastructures for implementing self-awareness processes? (iv) How can we verify self-aware cyber-physical systems and, in particular, which guarantees can we give? (v) What novel development processes will be required to engineer self-aware cyber-physical systems? We review each of these challenges in some detail and emphasize that addressing all of them requires the system to make a comprehensive assessment of the situation and a continual introspection of its own state to sensibly balance diverse requirements, constraints, short-term and long-term objectives. Throughout, we draw on three examples of cyber-physical systems that may benefit from self-awareness: a multi-processor system-on-chip, a Mars rover, and an implanted insulin pump. These three very different systems nevertheless have similar characteristics: limited resources, complex unforeseeable environmental dynamics, high expectations on their reliability, and substantial levels of risk associated with malfunctioning. Using these examples, we discuss the potential role of self-awareness in both highly complex and rather more simple systems, and as a main conclusion we highlight the need for research on above listed topics. Kirstie L. Bellman, Christopher Landauer, Nikil Dutt, Lukas Esterle, Andreas Herkersdorf, Axel Jantsch, Nima Taherinejad, Peter R. Lewis 0001, Marco Platzner, Kalle Tammemäe |
ACM Trans. Cyber Phys. Syst. | 9 |
| 2020 | Proof-Carrying Approximate CircuitsabstractApproximate circuits (AxCs) tradeoff computational accuracy against improvements in hardware area, delay, or energy consumption. IP core vendors who wish to create such circuits need to convince consumers of the resulting approximation quality. As a solution, we propose proof-carrying AxCs. The vendor creates an approximate IP core together with a certificate that proves the approximation quality. The proof certificate is bundled with the approximate IP core and sent off to the consumer. The consumer can formally verify the approximation quality of the IP core at a fraction of the typical computational cost for formal verification. In this brief, we first make the case for proof-carrying AxCs and then demonstrate the feasibility of the approach by a set of synthesis experiments using an exemplary approximation framework. Linus Witschen, Tobias Wiersema, Marco Platzner |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | Jump Search: A Fast Technique for the Synthesis of Approximate CircuitsabstractState-of-the-art frameworks for generating approximate circuits automatically explore the search space in an iterative process - often greedily. Synthesis and verification processes are invoked in each iteration to evaluate the found solutions and to guide the search algorithm. As a result, a large number of approximate circuits is subjected to analysis - leading to long runtimes - but only a few approximate circuits might form an acceptable solution. Linus Witschen, Hassan Ghasemzadeh Mohammadi, Matthias Artmann, Marco Platzner |
ACM Great Lakes Symposium on VLSI | 4 |
| 2019 | Zynq-based acceleration of robust high density myoelectric signal processing
Alexander Boschmann, Andreas Agne, Georg Thombansen, Linus Witschen, Florian Kraus, Marco Platzner |
J. Parallel Distributed Comput. | 6 |
| 2018 | A Highly Accurate Energy Model for Task Execution on Heterogeneous Compute NodesabstractHeterogeneous computing with CPUs, GPUs, and FPGAs has strongly gained interest in the last years. While scheduling and optimization problems for runtime have been widely studied, optimizing for energy-related metrics has become an emerging topic only recently due to rising electricity costs and the difficulties of thermal management. Energy-optimizing schedulers need to predict the effect of single task-resource assignment decisions on the consumed energy as well as the energy consumptions for complete schedules. In this paper, we present a highly accurate energy model for heterogeneous compute nodes. Compared to previous work, that differentiated between static and dynamic energy consumption and included idle energies, the new and accurate model looks into the impact of host-side activities of tasks executed on accelerators, covers more performance and power states of the devices, and considers resource-specific features such as reconfiguration processes for FPGA tasks. We present experiments and analyses of task executions on CPU, GPU, and FPGA, which allowed us to derive a set of critical model refinements over previous work. For evaluation, we compare the predictions of our model with previous work and with real measurements gained during the execution of 200 randomly generated schedules. The results show the high accuracy of our new energy model, with prediction errors as low as 0.3% and 0.4%, depending on the task set characteristic. Achim Lösch, Marco Platzner |
ASAP | 2 |
| 2018 | An MCTS-based Framework for Synthesis of Approximate CircuitsabstractApproximate computing has become a very popular design strategy that exploits error resilient computations to achieve higher performance and energy efficiency. Automated synthesis of approximate circuits is performed via functional approximation, in which various parts of the target circuit are extensively examined with a library of approximate components/transformations to trade off the functional accuracy and computational budget (i.e., power). However, as the number of possible approximate transformations increases, traditional search techniques suffer from a combinatorial explosion due to the large branching factor. In this work, we present a comprehensive framework for automated synthesis of approximate circuits from either structural or behavioral descriptions. We adapt the Monte Carlo Tree Search (MCTS), as a stochastic search technique, to deal with the large design space exploration, which enables a broader range of potential possible approximations through lightweight random simulations. The proposed framework is able to recognize the design Pareto set even with low computational budgets. Experimental results highlight the capabilities of the proposed synthesis framework by resulting in up to 61.69% energy saving while maintaining the predefined quality constraints. Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner |
VLSI-SoC | 3 |
| 2017 | reMinMin: A novel static energy-centric list scheduling approach based on real measurementsabstractHeterogeneous compute nodes in form of CPUs with attached GPU and FPGA accelerators have strongly gained interested in the last years. Applications differ in their execution characteristics and can therefore benefit from such heterogeneous resources in terms of performance or energy consumption. While performance optimization has been the only goal for a long time, nowadays research is more and more focusing on techniques to minimize energy consumption due to rising electricity costs. This paper presents reMinMin, a novel static list scheduling approach for optimizing the total energy consumption for a set of tasks executed on a heterogeneous compute node. reMinMin bases on a new energy model that differentiates between static and dynamic energy components and covers effects of accelerator tasks on the host CPU. The required energy values are retrieved by measurements on the real computing system. In order to evaluate reMinMin, we compare it with two reference implementations on three task sets with different degrees of heterogeneity. In our experiments, MinMin is consistently better than a scheduler optimizing for dynamic energy only, which requires up to 19.43% more energy, and very close to optimal schedules. Achim Lösch, Marco Platzner |
ASAP | 2 |
| 2017 | A Zynq-based dynamically reconfigurable high density myoelectric prosthesis controllerabstractThe combination of high-density electromyographic (HD EMG) sensor technology and modern machine learning algorithms allows for intuitive and robust prosthesis control of multiple degrees of freedom. However, HD EMG real-time processing poses a challenge for common microprocessors in an embedded system. With the goal set on an autonomous prosthesis capable of performing training and classification of an amputee's HD EMG signals, the focus of this paper lies in the acceleration of the computationally expensive parts of the embedded signal processing chain: the feature extraction and classification. Using the Xilinx Zynq as a low-cost off-the-shelf system, we present a solution capable of processing 192 HD EMG channels with controller delays below 120 milliseconds, suitable for highly responsive real-world prosthesis control, achieving speed-ups up to 2.8 as compared to a software-only solution. Using dynamic FPGA reconfiguration, the system is able to trade off increased controller delay against improved classification accuracy when signal quality is decreased due to noisy channels. Offloading feature extraction and classification to the FPGA also reduced the system's power consumption, making it more suitable to be used in a battery-powered setup. The system was validated using real-time experiments with online HD EMG data from an amputee to control a state-of-the-art prosthesis. Alexander Boschmann, Georg Thombansen, Linus Witschen, Alex Wiens, Marco Platzner |
DATE | 5 |
| 2017 | Accurate private/shared classification of memory accesses: A run-time analysis system for the LEON3 multi-core processorabstractRelated work has presented simulation-based experiments to classify data accesses in a shared memory multi-core into private and shared. This information can be used to selectively turn on/off cache coherency mechanisms for data blocks, which can save memory bus bandwidth, minimize energy consumption, and reduce application runtimes. In this paper we present an implementation of a private/shared classification mechanism on a LEON3 SPARC multi-core processor running the Linux 2.6 kernel. Our mechanism is paged-based and allows for classifying and counting data accesses at run-time. Compared to previous work, our system provides more accurate, i.e., realistic, data as it includes a real multi-core architecture and an OS. Additionally, our prototype allows us to quantitatively evaluate the overhead for the classification mechanism. We test our system with sequential and parallel benchmarks from the Mibench, ParaMibench, PARSEC, and SPLASH2 application suites. The results show that parallel benchmarks are promising targets for selectively controlling coherency mechanisms and that the run-time overheads induced by our mechanism are rather small. Nam Ho, Ishraq Ibne Ashraf, Paul Kaufmann, Marco Platzner |
DATE | 4 |
| 2017 | Evolvable caches: Optimization of reconfigurable cache mappings for a LEON3/Linux-based multi-core processorabstractReconfigurable hardware technology can help to improve the performance of processor caches by, for instance, tailoring and adapting associativity, block sizes, and replacement strategies to a particular application. A fundamentally different approach in using reconfigurability for better caches, and topic of this work, is cache adaptation by optimizing the memory-to-cache-index mapping function. This idea has been investigated previously but we present and evaluate for the first time a complete hardware implementation of such a system. Using a LEON3 multi-core processor running a standard Linux OS we optimize cache mappings for 12 applications from the MiBench suite with the result that miss rates can be reduced by up to 67% compared to the conventional cache mapping. Nam Ho, Paul Kaufmann, Marco Platzner |
FPT | 3 |
| 2017 | Guest Editorial: IEEE Transactions on Computers and IEEE Transactions on Emerging Topics in Computing Joint Special Section on Innovation in Reconfigurable Computing Fabrics from Devices to ArchitecturesabstractThe papers in this special section focuses on the advancement of the associated performance and reliability objectives via technology and functional heterogeneity, as well as advanced resilience properties in which reconfigurable fabrics are embarking. Emerging device characteristics such as non-volatility and new static versus dynamic energy consumption profiles, as well as novel interconnect mechanisms, and storage blocks within reconfigurable fabrics, in turn innovate architectural advances that enable new applications. Ronald F. DeMara, Marco Platzner, Marco Ottavi |
IEEE Trans. Computers | 2 |
| 2017 | Proof-Carrying Hardware via Inductive InvariantsabstractProof-carrying hardware (PCH) is a principle for achieving safety for dynamically reconfigurable hardware systems. The producer of a hardware module spends huge effort when creating a proof for a safety policy. The proof is then transferred as a certificate together with the configuration bitstream to the consumer of the hardware module, who can quickly verify the given proof. Previous work utilized SAT solvers and resolution traces to set up a PCH technology and corresponding tool flows. In this article, we present a novel technology for PCH based on inductive invariants. For sequential circuits, our approach is fundamentally stronger than the previous SAT-based one since we avoid the limitations of bounded unrolling. We contrast our technology to existing ones and show that it fits into previously proposed tool flows. We conduct experiments with four categories of benchmark circuits and report consumer and producer runtime and peak memory consumption, as well as the size of the certificates and the distribution of the workload between producer and consumer. Experiments clearly show that our new induction-based technology is superior for sequential circuits, whereas the previous SAT-based technology is the better choice for combinational circuits. Tobias Isenberg 0002, Marco Platzner, Heike Wehrheim, Tobias Wiersema |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2017 | The First 25 Years of the FPL Conference: Significant PapersabstractA summary of contributions made by significant papers from the first 25 years of the Field-Programmable Logic and Applications conference (FPL) is presented. The 27 papers chosen represent those which have most strongly influenced theory and practice in the field. Philip H. W. Leong, Hideharu Amano, Jason Helge Anderson, Koen Bertels, João M. P. Cardoso, Oliver Diessel, Guy Gogniat, Mike Hutton, Wayne Luk, Patrick Lysaght, Marco Platzner, Viktor Prasanna 0001, Tero Rissa, Cristina Silvano, Hayden Kwok-Hay So, Yu Wang 0002 |
ACM Trans. Reconfigurable Technol. Syst. | 12 |
| 2016 | Performance-centric scheduling with task migration for a heterogeneous compute node in the data center
Achim Lösch, Tobias Beisel, Tobias Kenter, Christian Plessl, Marco Platzner |
DATE | 5 |
| 2016 | Boolean Difference Based Reliability Evaluation of Fault-Tolerant Circuit Structures on FPGAsabstractThe reliability of FPGA based hardware designs has become an important field of research particularly for space computing. Traditionally, redundancy is utilized in FPGA based designs to achieve reliable or error-tolerant computing. However, the redundant designs vary according to the granularity level and the voter placement algorithms used for the hardware design. The resulting circuit configurations vary in area, latency and power as well as in the achieved reliability. While the evaluation of area, latency and power is done by the FPGA design tools, quantitative data for reliability are usually not derived. Consequently, there is a need for an automated reliability evaluation tool especially considering the huge design space of redundant circuit structures. In this paper we combine the Boolean difference error calculator (BDEC), a probabilistic reliability model for hardware designs, with a reliability model for fault-tolerant circuit structures. As a result, we are able to study the reliability of fault-tolerant circuit structures at the logic layer. We focus on fault-tolerant circuits to be implemented in FPGA and show how to extend our combined model from combinational to sequential circuits. For an automated analysis, we develop a MATLAB-based tool utilizing our extended BDEC model. An application of our automated BDEC tool has been provided with a case study on dynamic reliability management to demonstrate how such a tool providing quantitative reliability data can improve the four-dimensional Pareto optimization for area, latency, power and reliability. Jahanzeb Anwer, Marco Platzner |
DSD | 2 |
| 2016 | Adaptive playouts for online learning of policies during Monte Carlo Tree Search
Tobias Graf, Marco Platzner |
Theor. Comput. Sci. | 2 |
| 2015 | Significant papers from the first 25 years of the FPL conferenceabstractThe list of significant papers from the first 25 years of the Field-Programmable Logic and Applications conference (FPL) is presented in this paper. These 27 papers represent those which have most strongly influenced theory and practice in the field. Philip H. W. Leong, Hideharu Amano, Jason Helge Anderson, Koen Bertels, João M. P. Cardoso, Oliver Diessel, Guy Gogniat, Mike Hutton, Wayne Luk, Patrick Lysaght, Marco Platzner, Viktor Prasanna 0001, Tero Rissa, Cristina Silvano, Hayden Kwok-Hay So, Yu Wang 0002 |
FPL | 12 |
| 2015 | Comparison of thread signatures for error detection in hybrid multi-coresabstractDynamic thread duplication is a known redundancy technique for multi-cores. We have adopted this concept to reconfigurable hardware cores running hardware threads in a hybrid multi-core system. For hardware threads we can compare three different signatures: the sequence of operating system (OS) function calls, the sequence of OS function calls with their parameters, and additionally all memory accesses with their type and data. These signatures allow for an increasing error detection coverage, but also come with increasing performance overheads, which enables an application designer to trade-off the error detection coverage for performance. Our experiments with three benchmark applications show that the best error coverage can be achieved at a performance cost between 2% and 55%, depending on the benchmark, the signature used and the utilization of the CPU. Sebastian Meisner, Marco Platzner |
FPT | 2 |
| 2015 | Distributed Monte Carlo Tree Search: A Novel Technique and its Application to Computer GoabstractMonte Carlo tree search (MCTS) has brought about great success regarding the evaluation of stochastic and deterministic games in recent years. We present and empirically analyze a data-driven parallelization approach for MCTS targeting large HPC clusters with Infiniband interconnect. Our implementation is based on OpenMPI and makes extensive use of its RDMA based asynchronous tiny message communication capabilities for effectively overlapping communication and computation. We integrate our parallel MCTS approach termed UCT-Treesplit in our state-of-the-art Go engine Gomorra and measure its strengths and limitations in a real-world setting. Our extensive experiments show that we can scale up to 128 compute nodes and 2048 cores in self-play experiments and, furthermore, give promising directions for additional improvement. The generality of our parallelization approach advocates its use to significantly improve the search quality of a huge number of current MCTS applications. Lars Schäfers, Marco Platzner |
IEEE Trans. Comput. Intell. AI Games | 2 |
| 2014 | A hardware/software infrastructure for performance monitoring on LEON3 multicore platformsabstractMonitoring applications at run-time and evaluating the recorded statistical data of the underlying micro architecture is one of the key aspects required by many hardware architects and system designers as well as high-performance software developers. To fulfill this requirement, most modern CPUs for High Performance Computing have been equipped with Performance Monitoring Units (PMU) including a set of hardware counters, which can be configured to monitor a rich set of events. Unfortunately, embedded and reconfigurable systems are mostly lacking this feature. Towards rapid exploration of High Performance Embedded Computing in near future, we believe that supporting PMU for these systems is necessary. In this paper, we propose a PMU infrastructure, which supports monitoring of up to seven concurrent events. The PMU infrastructure is implemented on an FPGA and is integrated into a LEON3 platform.We show also the integration of our PMU infrastructure with the perf_event, which is the standard PMU architecture of the Linux kernel. Nam Ho, Paul Kaufmann, Marco Platzner |
FPL | 3 |
| 2014 | Memory security in reconfigurable computers: Combining formal verification with monitoringabstractEnsuring memory access security is a challenge for reconfigurable systems with multiple cores. Previous work introduced access monitors attached to the memory subsystem to ensure that the cores adhere to pre-defined protocols when accessing memory. In this paper, we combine access monitors with a formal runtime verification technique known as proof-carrying hardware to guarantee memory security. We extend previous work on proof-carrying hardware by covering sequential circuits and demonstrate our approach with a prototype leveraging ReconOS/Zynq with an embedded ZUMA virtual FPGA overlay. Experiments show the feasibility of the approach and the capabilities of the prototype, which constitutes the first realization of proof-carrying hardware on real FPGAs. The area overheads for the virtual FPGA are measured as 2x-10x, depending on the resource type. The delay overhead is substantial with almost 100x, but this is an extremely pessimistic estimate that will be lowered once accurate timing analysis for FPGA overlays become available. Finally, reconfiguration time for the virtual FPGA is about one order of magnitude lower than for the native Zynq fabric. Tobias Wiersema, Stephanie Drzevitzky, Marco Platzner |
FPT | 3 |
| 2014 | Integrating Software and Hardware Verification
Marie-Christine Jakobs, Marco Platzner, Heike Wehrheim, Tobias Wiersema |
IFM | 2 |
| 2014 | An FPGA-Based Reconfigurable Mesh Many-CoreabstractThe reconfigurable mesh is a parallel model of computation, which exploits a massive amount of rather simple processing elements connected through a reconfigurable interconnection network. During the last decades, the model received strong interest and many researchers have devised algorithms for it. However, most of this work focuses on theoretical aspects. Due to some idealistic modeling assumptions only a few attempts have been made to implement the model and to study the practical use of reconfigurable meshes. In this paper, we leverage the reconfigurable mesh model to study potential architectures and programming models for future many-cores. We design a reconfigurable mesh in form of a scalable soft core array with a reconfigurable interconnect and implement it on FPGA technology in order to create a prototype platform. We present an overall hardware/software tool flow for generating and programming reconfigurable mesh prototypes. The new language ARMLang and a corresponding compiler facilitate the programming of the massively parallel processor arrays. To our knowledge, this work is the first practical study of word-level reconfigurable meshes. To analyze the performance of our implementation we study four algorithmic kernels from the application domains arithmetic, sorting, graph algorithms and imaging. For each kernel, we devise a reconfigurable mesh program in ARMLang, compile it to our soft core array and measure its runtime depending on the mesh size. Then, we compare the runtimes to two sequential implementations of the algorithms, which are executed on two single core systems. The results show that many-cores leveraging the reconfigurable mesh model can efficiently use a vast number of processing elements and that, for the chosen algorithms, they come close to optimally parallelized programs. Heiner Giefers, Marco Platzner |
IEEE Trans. Computers | 2 |
| 2014 | Self-Awareness as a Model for Designing and Operating Heterogeneous MulticoresabstractSelf-aware computing is a paradigm for structuring and simplifying the design and operation of computing systems that face unprecedented levels of system dynamics and thus require novel forms of adaptivity. The generality of the paradigm makes it applicable to many types of computing systems and, previously, researchers started to introduce concepts of self-awareness to multicore architectures. In our work we build on a recent reference architectural framework as a model for self-aware computing and instantiate it for an FPGA-based heterogeneous multicore running the ReconOS reconfigurable architecture and operating system. After presenting the model for self-aware computing and ReconOS, we demonstrate with a case study how a multicore application built on the principle of self-awareness, autonomously adapts to changes in the workload and system state. Our work shows that the reference architectural framework as a model for self-aware computing can be practically applied and allows us to structure and simplify the design process, which is essential for designing complex future computing systems. Andreas Agne, Markus Happe, Achim Lösch, Christian Plessl, Marco Platzner |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2013 | On-The-Fly Computing: A novel paradigm for individualized IT servicesabstractIn this paper we introduce “On-The-Fly Computing”, our vision of future IT services that will be provided by assembling modular software components available on world-wide markets. After suitable components have been found, they are automatically integrated, configured and brought to execution in an On-The-Fly Compute Center. We envision that these future compute centers will continue to leverage three current trends in large scale computing which are an increasing amount of parallel processing, a trend to use heterogeneous computing resources, and — in the light of rising energy cost — energy-efficiency as a primary goal in the design and operation of computing systems. In this paper, we point out three research challenges and our current work in these areas. Markus Happe, Friedhelm Meyer auf der Heide, Peter Kling, Marco Platzner, Christian Plessl |
ISORC | 4 |
| 2013 | Classification of Electromyographic Signals: Comparing Evolvable Hardware to Conventional ClassifiersabstractEvolvable hardware (EHW) has shown itself to be a promising approach for prosthetic hand controllers. Besides competitive classification performance, EHW classifiers offer self-adaptation, fast training, and a compact implementation. However, EHW classifiers have not yet been sufficiently compared to state-of-the-art conventional classifiers. In this paper, we compare two EHW approaches to four conventional classification techniques:k-nearest-neighbor, decision trees, artificial neural networks, and support vector machines. We provide all classifiers with features extracted from electromyographic signals taken from forearm muscle contractions, and let the algorithms recognize eight to eleven different kinds of hand movements. We investigate classification accuracy on a fixed data set and stability of classification error rates when new data is introduced. For this purpose, we have recorded a short-term data set from three individuals over three consecutive days and a long-term data set from a single individual over three weeks. Experimental results demonstrate that EHW approaches are indeed able to compete with state-of-the-art classifiers in terms of classification performance. Paul Kaufmann, Kyrre Glette, Thiemo Gruber, Marco Platzner, Jim Tørresen, Bernhard Sick |
IEEE Trans. Evol. Comput. | 4 |
| 2011 | Parallel Monte-Carlo Tree Search for HPC Systems
Tobias Graf, Ulf Lorenz, Marco Platzner, Lars Schäfers |
Euro-Par (2) | 3 |
| 2011 | Performance estimation framework for automated exploration of CPU-accelerator architecturesabstractIn this paper we present a fast and fully automated approach for studying the design space when interfacing reconfigurable accelerators with a CPU. Our challenge is, that a reasonable evaluation of architecture parameters requires a hardware/software partitioning that makes best use of each given architecture configuration. Therefore we developed a framework based on the LLVM infrastructure that performs this partitioning with high-level estimation of the runtime on the target architecture utilizing profiling information and code analysis. By making use of program characteristics also during the partitioning process, we improve previous results for various benchmarks and especially for growing interface latencies between CPU and accelerator. Tobias Kenter, Christian Plessl, Marco Platzner, Michael Kauschke |
FPGA | 3 |
| 2011 | Memory Virtualization for Multithreaded Reconfigurable HardwareabstractWith the introduction of multithreaded programming for reconfigurable hardware, it is possible to map both sequential software and parallel hardware to a single CPU/FPGA platform using threads as a unifying development model. At the same time, platform FPGAs are a natural technology for implementing computationally intensive systems in the aerospace, automotive and industrial domains, as they combine high performance and flexibility with lower non-recurring engineering (NRE) costs when compared to low-volume ASIC solutions. The reusability and portability of hardware components in these safety-critical domains could be significantly improved by using multithreaded programming. However, the unique design considerations for memory virtualization, as required in safety-critical systems, are difficult to transfer directly from software to autonomous hardware threads. This paper presents a transparent and efficient way of augmenting current multithreaded and partially reconfigurable hardware runtime environments with dedicated, hardware-thread-aware memory address translation units to provide seamless memory translation for hardware threads. We show an analysis of the overheads, as well as an experimental evaluation of the latencies caused by address translation. Andreas Agne, Marco Platzner, Enno Lübbers |
FPL | 2 |
| 2010 | A novel hybrid evolutionary strategy and its periodization with multi-objective genetic optimizersabstractThis work investigates the effects of the periodization of local and global multi-objective search algorithms. To this, we introduce a model for periodization and define a new multi-objective evolutionary algorithm adopting concepts from Evolutionary Strategies and NSGAII. We show that our method, especially when periodized with standard multi-objective genetic algorithms, excels for the evolution of digital circuits on the Cartesian Genetic Programming model as well as on some standard benchmarks such as the ZDT6. Paul Kaufmann, Tobias Knieper, Marco Platzner |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | A Triple Hybrid Interconnect for Many-Cores: Reconfigurable Mesh, NoC and BarrierabstractNetworks-on-chip (NoC) are very efficient for point-to-point communication but are also known to provide poor broadcast and multicast performance. In this paper, we propose a triple hybrid interconnect for many-cores, consisting of a reconfigurable mesh network and a wormhole routed NoC for data communication, and a barrier network for synchronization. On an FPGA many-core prototype comprising up to 30 Microblaze soft cores we show that the reconfigurable mesh network excels for multicast and broadcast operations, while the NoC performs better for larger messages and more dynamic workloads. Experiments with a parallel Jacobi algorithm demonstrate that the combined use of all three networks delivers the highest performance. Heiner Giefers, Marco Platzner |
FPL | 2 |
| 2009 | IMORC: Application Mapping, Monitoring and Optimization for High-Performance Reconfigurable ComputingabstractMapping applications that consist of a collection of cores to FPGA accelerators and optimizing their performance is a challenging task in high performance reconfigurable computing. We present IMORC, an architectural template and highly versatile on-chip interconnect. IMORC links provide asynchronous FIFOs and bitwidth conversion which allows for flexibly composing accelerators from cores running at full speed within their own clock domains, thus facilitating the re-use of cores and portability. Further, IMORC inserts performance counters for monitoring runtime data. In this paper, we introduce the IMORC architectural template and the on-chip interconnect and demonstrate IMORC on the example of accelerating the 𝑘-th nearest neighbor thinning problem on an XtremeData XD1000 reconfigurable computing system. Tobias Schumacher 0001, Christian Plessl, Marco Platzner |
FCCM | 3 |
| 2009 | Program-driven fine-grained power management for the reconfigurable meshabstractThe reconfigurable mesh model for massively parallel computing has recently been rediscovered and proposed as the basis of a practical many-core architecture. With this paper, we are the first to study the power and energy requirements of reconfigurable meshes. We present two methods to reduce the power consumption, exploiting characteristics of typical reconfigurable mesh algorithms. We extend our previous reconfigurable mesh architecture and the corresponding programming language ARMLang to support program driven power management. On two sample applications, we discuss model-based power estimations and compare the estimates with power measurements on FPGAs. Our program driven power management methods are shown to be effective. For one application we report a reduction in power and energy consumption by 21.09%. Heiner Giefers, Marco Platzner |
FPL | 2 |
| 2009 | Cooperative multithreading in dynamically reconfigurable systemsabstractPreemptive multitasking, a popular technique for timesharing of computational resources in software-based systems, faces considerable difficulties when applied to partially reconfigurable hardware. In this paper, we propose a cooperative scheduling technique for reconfigurable hardware threads as a feasible compromise between computational efficiency and implementation complexity. We have implemented this mechanism for the multithreaded reconfigurable operating system ReconOS and evaluated its overheads and performance on a prototype. Enno Lübbers, Marco Platzner |
FPL | 2 |
| 2009 | An accelerator for K-TH nearest neighbor thinning based on the IMORC infrastructureabstractThe creation and optimization of FPGA accelerators comprising several compute cores and memories are challenging tasks in high performance reconfigurable computing. In this paper, we present the design of such an accelerator for the kth nearest neighbor thinning problem on an XD1000 reconfigurable computing system. The design leverages IMORC, an architectural template and highly versatile on-chip interconnect, to achieve speedups of 74 times over a 2.2 GHz Opteron. Using IMORC with its asynchronous FIFOs and bitwidth conversion in the links between the cores, we are able to quickly create acclerator versions with varying degrees of core-level parallelism and memory mappings. Through the performance monitoring infrastructure of IMORC we gain insight into the data-dependent behavior of the accelerator which facilitates further performance optimizations. Tobias Schumacher 0001, Christian Plessl, Marco Platzner |
FPL | 3 |
| 2009 | ARMLang: A language and compiler for programming reconfigurable mesh many-coresabstractThe reconfigurable mesh serves as a theoretical model for massively parallel computing, but has recently been investigated as a practical architecture for many-cores with light-weight, circuit-switched interconnects. There is a lack of programming environments, including languages, compilers, and debuggers for reconfigurable meshes. In this paper, we present the new language ARMLang for the specification of lockstep programs on regular processor arrays, in particular reconfigurable meshes. Lockstep synchronization is achieved by path equalization and barrier synchronization, both of which are supported by the new language. We further discuss the creation of an ARMLang compiler and a simulation environment that allows for debugging and visualization of the parallel programs. Heiner Giefers, Marco Platzner |
IPDPS | 2 |
| 2009 | ReconOS: Multithreaded programming for reconfigurable computersabstractRising logic densities together with the inclusion of dedicated processor cores push reconfigurable devices from being applied for glue logic and prototyping towards implementing complete reconfigurable systems-on-chip. The mix of fast CPU cores and fine-grained reconfigurable logic allows to map both sequential, control-dominated code and highly parallel data-centric computations onto one platform. However, traditional design techniques that view specialized hardware circuits as passive coprocessors are ill-suited for programming these reconfigurable computers. In particular, the programming models for software—running on an embedded operating system—and digital hardware—synthesized to an FPGA—lack commonalities, which hinders design space exploration and severely impairs the potential for code reuse. In this article, we present ReconOS, an execution environment based on existing embedded operating systems that extends the multithreaded programming model established in the software domain to reconfigurable hardware. Using threads and common synchronization and communication services as an abstraction layer, ReconOS allows for the creation of portable and flexible multithreaded applications targeting CPU/FPGA systems. This article discusses the ReconOS programming model and its execution environment, presents implementations based on modern platform FPGAs and the operating systems eCos and Linux, evaluates time and area overheads of the proposed mechanisms and, finally, demonstrates the feasibility of the multithreading design approach on several case studies. Enno Lübbers, Marco Platzner |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2008 | Fine grain reconfigurable architecturesabstractIn this booth on fine grain reconfigurable architectures, several research groups demonstrate their joint work on operating concepts for managing dynamic and partial reconfiguration, visualization of bitstreams and routing, presenting an application applying dynamic reconfiguration for video engines as well as work on minimization of reconfiguration data. Unique is that all the above four projects present their work using the same reconfigurable FPGA-based fabric called Erlangen slot machine that has also been built within one project just the purpose of experimenting with dynamic fine grain reconfiguration as an interdisciplinary platform. Josef Angermeier, Mateusz Majer, Jürgen Teich, Lars Braun, Tobias Schwalb, Philipp Graf, Michael Hübner 0001, Jürgen Becker 0001, Enno Lübbers, Marco Platzner, Christopher Claus, Walter Stechele, Andreas Herkersdorf, Markus Rullmann, Renate Merker |
FPL | 10 |
| 2008 | A portable abstraction layer for hardware threadsabstractThe multithreaded programming model has been shown to provide a suitable abstraction for reconfigurable computers. Previous implementations of corresponding runtime systems have been limited to a single host operating system, hardware platform, or application domain. This paper presents the implementation of ReconOS, our hardware/software multithreaded programming model, on both eCos and Linux-based host systems as well as on PowerPC and MicroBlaze CPUs. This demonstrates that ReconOS provides a truly portable abstraction layer for programming reconfigurable computers. Further, we quantify the performance of operating system calls and measure the resulting application level performance. Enno Lübbers, Marco Platzner |
FPL | 2 |
| 2008 | Advanced techniques for the creation and propagation of modules in cartesian genetic programmingabstractThe choice of an appropriate hardware representation model is key to successful evolution of digital circuits. One of the most popular models is cartesian genetic programming, which encodes an array of logic gates into a chromosome. While several smaller circuits have been successfully evolved on this model, it lacks scalability. A recent approach towards scalable hardware evolution is based on the automated creation of modules from primitive gates. Paul Kaufmann, Marco Platzner |
GECCO | 2 |
| 2007 | A Many-core Implementation based on the Reconfigurable Mesh ModelabstractThe reconfigurable mesh is a model for massively parallel computing for which many algorithms with very low complexity have been developed. These algorithms execute cycles of bus configuration, communication, and constant-time computation on all processing elements in a lock-step. In this paper, we investigate the use of reconfigurable meshes as coprocessors to accelerate important algorithmic kernels. We discuss the development of a reconfigurable mesh on FPGA technology, including the host integration and the programming tool flow. Then, we present implementation results and a proof-of-concept case study. Heiner Giefers, Marco Platzner |
FPL | 2 |
| 2007 | ReconOS: An RTOS supporting Hard- and Software ThreadsabstractModern platform FPGAs integrate fine-grained reconfigurable logic with processor cores and allow the creation of complete configurable systems-on-chip. However, design methodologies have not kept up with the rise in complexity of the target hardware. In particular, there is little overlap between the programming model for embedded software running on a real-time operating system and the programming model for digital logic. In this paper, we present the operating system ReconOS which supports both software and hardware threads with a single unified programming model. ReconOS is based on eCos, a widely-used real-time operating system (RTOS). We investigate the incurred time and area overheads, especially for inter-thread communication across the hardware/-software boundary, and present a case study demonstrating the feasibility of the RTOS-centric design approach. Enno Lübbers, Marco Platzner |
FPL | 2 |
| 2006 | Executing Hardware Tasks on Dynamically Reconfigurable Devices Under Real-Time ConditionsabstractThis paper presents a prototype system that executes a set of periodic real-time tasks utilizing dynamic hardware reconfiguration. The proposed scheduling technique, MSDL, is not only able to give an offline guarantee for the feasibility of the task set but also minimizes the number of device configurations. After describing this technique, we extend the schedulability analysis to include different runtime system overheads, including the device reconfiguration time. Then we detail a light-weight runtime system that performs the online part of the MSDL scheduling technique. The runtime system is entirely implemented in hardware. Finally, we outline the corresponding synthesis tool flow and report on the overhead posed by the runtime system Klaus Danne, Roland Mühlenbernd, Marco Platzner |
FPL | 3 |
| 2006 | Optimal temporal partitioning based on slowdown and retimingabstractThis paper presents a novel method for optimal temporal partitioning of sequential circuits for time-multiplexed reconfigurable architectures. The method bases on slowdown and retiming and maximizes the circuit's performance during execution while restricting the size of the partitions to respect the resource constraints of the reconfigurable architecture. A mixed integer linear program (MILP) formulation of the problem was provided, which can be solved exactly. In contrast to related work, our approach optimizes performance directly, takes structural modifications of the circuit into account, and is extensible. The application of the new method to temporal partitioning for a coarse-grained reconfigurable architecture was presented Christian Plessl, Marco Platzner, Lothar Thiele |
FPT | 2 |
| 2006 | Partitioned scheduling of periodic real-time tasks onto reconfigurable hardwareabstractReconfigurable hardware devices, such as FPGAs, are increasingly used in embedded systems. To utilize these devices for real-time work loads, scheduling techniques are required that generate predictable task timings. In this paper, we present a partitioning-EDF (earliest deadline first) approach to find such schedules. The FPGA area is partitioned along one dimension into slots. The tasks are partitioned into groups. Then, each group is scheduled to exactly one slot using the EDF rule. We show that the problem of finding an optimal partitioning is related to the well-known 2D level bin-packing problem. We extend a previously reported ILP model to solve our partitioning problem to optimality. By a simulation study we demonstrate that the partitioning-EDF approach is able to find feasible schedules for most task sets with a system utilization of up to 70%. Additionally, we allow a task to be realized in alternative implementations. A simulation study reveals that the scheduling performance increases considerably if three instead of one task variants are considered. Finally, we model and study the impact of the device reconfiguration time on the scheduling performance Klaus Danne, Marco Platzner |
IPDPS | 2 |
| 2006 | An EDF schedulability test for periodic tasks on reconfigurable hardware devicesabstractIn this paper, we consider the scheduling of periodic real-time tasks on reconfigurable hardware devices. Such devices can execute several tasks in parallel. All executing tasks share the hardware resource, which makes the scheduling problem differ from single- and multiprocessor scheduling. We adapt the global EDF multiprocessor scheduling approach to the reconfigurable hardware execution model and define two preemptive scheduling algorithms, EDF-First-k-Fit and EDF-Next-Fit. For these algorithms, we present a novel linear-time schedulability test and give a proof based on a resource augmentation technique. Then, we propose a task placement and relocation scheme utilizing partial device reconfiguration. This scheme allows us to extend the schedulability test to include reconfiguration time overheads. Experiments with synthetic workloads compare the scheduling test with the actual scheduling performance of EDF-First-k-Fit and EDF-Next-Fit. The main evaluation result is that the reconfiguration overhead is acceptable if the task computation times are one order of magnitude higher than the device reconfiguration time. Klaus Danne, Marco Platzner |
LCTES | 2 |
| 2005 | Zippy - A coarse-grained reconfigurable array with support for hardware virtualizationabstractThis paper motivates the use of hardware visualization on coarse-grained reconfigurable architectures. We introduce Zippy, a coarse-grained multi-context hybrid CPU with architectural support for efficient hardware virtualization. The architectural details and the corresponding tool flow are outlined. As a case study, we compare the non-virtualized and the virtualized execution of an ADPCM decoder. Christian Plessl, Marco Platzner |
ASAP | 2 |
| 2005 | A Heuristic Approach to Schedule Periodic Real-Time Tasks on Reconfigurable HardwareabstractThis paper deals with scheduling periodic real-time tasks on reconfigurable hardware devices, such as FPGAs. Reconfigurable hardware devices are increasingly used in embedded systems. To utilize these devices also for systems with real-time constraints, predictable task scheduling is required. We formalize the periodic task scheduling problem and propose two preemptive scheduling algorithms. The first is an adaption of the well-known earliest deadline first (EDF) technique to the FPGA execution model. Although the algorithm reveals good scheduling performance, it lacks an efficient schedulability test and requires a high number of FPGA configurations. The second algorithm uses the concept of servers that reserve area and execution time for other tasks. Tasks are successively merged into servers, which are then scheduled sequentially. While this method is inferior to the EDF-based technique regarding schedulability, it comes with a fast schedulability test and greatly reduces the number of required FPGA configurations. Klaus Danne, Marco Platzner |
FPL | 2 |
| 2004 | Efficient Execution of Process Networks on a Reconfigurable Hardware Virtual MachineabstractIn this paper, we present a novel use of an FPGA as a computing element for streaming based application. We investigate the virtualized execution of dynamic reconfigurable tasks. We use the process networks model as a coordination language which is interpreted on a virtual machine run-time system. We present and discuss the results of a design space exploration, which evaluates the performance of the system architecture for different configurations. Matthias Dyer, Marco Platzner, Lothar Thiele |
FCCM | 2 |
| 2004 | A Runtime Environment for Reconfigurable Hardware Operating Systems
Herbert Walder, Marco Platzner |
FPL | 2 |
| 2004 | Operating Systems for Reconfigurable Embedded Platforms: Online Scheduling of Real-Time TasksabstractToday's reconfigurable hardware devices have huge densities and are partially reconfigurable, allowing for the configuration and execution of hardware tasks in a true multitasking manner. This makes reconfigurable platforms an ideal target for many modern embedded systems that combine high computation demands with dynamic task sets. A rather new line of research is engaged in the construction of operating systems for reconfigurable embedded platforms. Such an operating system provides a minimal programming model and a runtime system. The runtime system performs online task and resource management. In this paper, we first discuss design issues for reconfigurable hardware operating systems. Then, we focus on a runtime system for guarantee-based scheduling of hard real-time tasks. We formulate the scheduling problem for the 1D and 2D resource models and present two heuristics, the horizon and the stuffing technique, to tackle it. Simulation experiments conducted with synthetic workloads evaluate the performance and the runtime efficiency of the proposed schedulers. The scheduling performance for the 1D resource model is strongly dependent on the aspect ratios of the tasks. Compared to the 1D model, the 2D resource model is clearly superior. Finally, the runtime overhead of the scheduling algorithms is shown to be acceptably low. Christoph Steiger, Herbert Walder, Marco Platzner |
IEEE Trans. Computers | 3 |
| 2003 | Online Scheduling for Block-Partitioned Reconfigurable Devices
Herbert Walder, Marco Platzner |
DATE | 2 |
| 2003 | Virtualizing Hardware with Multi-context Reconfigurable Arrays
Rolf Enzler, Christian Plessl, Marco Platzner |
FPL | 3 |
| 2003 | Heuristics for Onine Scheduling Real-Time Tasks to Partially Reconfigurable Devices
Christoph Steiger, Herbert Walder, Marco Platzner |
FPL | 3 |
| 2003 | TKDM - a reconfigurable co-processor in a PC's memory slotabstractThis paper presents TKDM, a PC-based high-performance reconfigurable computing environment. The TKDM hardware consists of an FPGA module that uses the DIMM (dual inline memory module) bus for high-bandwidth and low-latency communication with the host CPU. The system's firmware is integrated with the Linux host operating system and offers functions for data communication and FPGA reconfiguration. The intended use of TKDM is that a dynamically reconfigurable co-processor for data streaming applications. The system's firmware can be customized for specific application domains to facilitate simple and easy-to-use programming interfaces. Christian Plessl, Marco Platzner |
FPT | 2 |
| 2003 | Online Scheduling and Placement of Real-time Tasks to Partially Reconfigurable DevicesabstractThis paper deals with online scheduling of tasks to partially reconfigurable devices. Such devices are able to execute several tasks in parallel. All tasks share the reconfigurable surface as a single resource which leads to highly dynamic allocation situations. To manage such devices at runtime, we propose a reconfigurable operating system that splits into three main modules: scheduler, placer, and loader. The main characteristic of the resulting online scheduling problem is the strong nexus between scheduling and placement. We discuss a fast online placement technique and then focus on scheduling real-time tasks. We devise guarantee-based schedulers for two scenarios, namely tasks with arbitrary and synchronous arrival times. The schedulers exploit the knowledge about task properties to improve the system's performance. The experiments show that the developed schedulers lead to substantial performance gains at an acceptable runtime overhead. Christoph Steiger, Herbert Walder, Marco Platzner, Lothar Thiele |
RTSS | 3 |
| 2003 | The case for reconfigurable hardware in wearable computing
Christian Plessl, Rolf Enzler, Herbert Walder, Jan Beutel, Marco Platzner, Lothar Thiele, Gerhard Tröster |
Pers. Ubiquitous Comput. | 5 |
| 2003 | Instance-Specific Accelerators for Minimum Covering
Christian Plessl, Marco Platzner |
J. Supercomput. | 2 |
| 2002 | Custom Computing Machines for the Set Covering ProblemabstractWe present instance-specific custom computing machines for the set covering problem. Four accelerator architectures are developed that implement branch & bound in 3-valued logic and many of the deduction techniques found in software solvers. We use set covering benchmarks from two-level logic minimization and Steiner triple systems to derive and discuss experimental results. The resulting raw speedups are in the order of four magnitudes on average. Finally, we propose a hybrid solver architecture that combines the raw speed of instance-specific reconfigurable hardware with flexible bounding schemes implemented in software. Christian Plessl, Marco Platzner |
FCCM | 2 |
| 2002 | Partially Reconfigurable Cores for Xilinx Virtex
Matthias Dyer, Christian Plessl, Marco Platzner |
FPL | 3 |
| 2002 | A Framework for Run-time Reconfigurable Systems
Michael Eisenring, Marco Platzner |
J. Supercomput. | 2 |
| 2001 | Object-oriented domain specific compilers for programming FPGAsabstractSimplifying the programming models is paramount to the success of reconfigurable computing with field programmable gate arrays (FPGAs). This paper presents a methodology to combine true object-oriented design of the compiler/CAD tool with an object-oriented hardware design methodology in C++. The resulting system provides all the benefits of object-oriented design to the compiler/CAD tool designer and to the hardware designer/programmer. The two examples for domain-specific compilers presented are BSAT and StReAm. Each domain-specific compiler is targeted at a very specific application domain, such as applications that accelerate Boolean satisfiability problems with BSAT, and applications which lend themselves for implementation as a stream architecture with StReAm. The key benefit of the presented domain specific compilers is a reduction of design time by orders of magnitude while keeping the optimal performance of hand-designed circuits. Oskar Mencer, Marco Platzner, Martin Morf, Michael J. Flynn |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1995 | A Special-purpose Coprocessor for Qualitative Simulation
Gerald Friedl, Marco Platzner, Bernhard Rinner |
Euro-Par | 2 |