Christian Plessl

dblp:76/5938 · DBLP profile ↗
← Back
50ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0001-5728-9982ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 44 · 6 first-author · 16 since 2021Computer networks · 2Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SORCERI: Streaming Overlay Acceleration for Highly Contracted Electron Repulsion Integral Computations in Quantum Chemistry
Philip Stachura, Christian Plessl, Zhenman Fang
FPGA3
2025 Efficient and Distributed Computation of Electron Repulsion Integrals on AMD AI Engines
abstract
Computing electron repulsion integrals (ERIs) is the major computational bottleneck of many quantum mechanical simulation methods, requiring trillions of ERI evaluations per time step. While the computation of independent ERIs is embarrassingly parallel, the efficient computation of individual ERIs on modern processor cores is difficult due to both an insufficient cache size for intermediates of the computation and irregular memory access patterns that are difficult to vectorize. In this paper, we present how our implementation on the AI Engine (AIE) architecture addresses both of these problems. First, we have defined a flexible graph structure, which we call an ERI-Engine, that can be implemented for all 231 canonical ERI quartets from {ss|ss} to {hh| hh} by distributing the computation over 2–14 AIEs. Second, for the larger quartets, we have devised a novel vectorization scheme that leverages the advanced floating-point unit of the AIEs, while also supporting vectorization of independent ERIs for the smaller quartets. Finally, ERI-Engines are horizontally and vertically stackable to fill the entire AIE array, and in particular, the vertically stacked ERI-Engines form a column that uses one or more time-shared channels to stream the results out of the AIE array, almost completely hiding the computational phases of individual ERI-Engines. In terms of absolute performance, we are competitive with recent high-performance implementations of ERI algorithms on FPGAs (SERI) and GPUs (LibintX), as well as well-established highly optimized CPU libraries (Libint, Libcint), while being the unequivocal leader in terms of energy efficiency.
Johannes Menzel, Christian Plessl
FCCM2
2025 Neural Network Inference in High-Performance Computing: Closing the Gap for FINN based Reconfigurable Accelerators
abstract
In recent years, Neural Networks (NNs) have become one of the most prevailing topics in computers science, both in research and in industry. NNs are used for data analysis, natural language processing, autonomous driving and more. As such, NNs also see more application and use in High-Performance Computing (HPC). At the same time, energy efficiency has become an increasingly critical topic. NNs use large amounts of energy for operation, which in return results in large amounts of CO2 emissions. This work presents a comprehensive evaluation of current NN inference soft- and hardware configurations within High-Performance Computing (HPC) environments, with a focus on both performance metrics and energy consumption. NN quantization and accelerators such as FPGAs allow for an increased inference efficiency, both in terms of throughput and energy. Therefore, this work focuses on FINN, an efficient NN inference framework for FPGAs, highlighting its current lack of support for HPC systems. We provide an in-depth analysis of FINN in order to implement extensions to optimize the end-to-end execution for the usage in the HPC environment. We thoroughly evaluate the performance and energy efficiency gains using newly implemented optimizations and compare it against existing NN accelerators for HPC. With our extensions of FINN, we were able to achieve a 1847× higher throughput, while also decreasing the latency on average by 0.9978× and EDP by 0.9979× on an Alveo U55C FPGA. Data flow based NN inference accelerators on an FPGA should be used if the performance and energy footprint of the inference process is crucial, and the batch sizes are small to medium. For extremely large batch sizes and a very limited time for network-to-accelerator (less than a few days), using GPUs is still the way to go. Our results show that with the newly developed driver, we outperform a high-end Nvidia A100 GPU by up to 7.81x in throughput, while having a 0.87x lower latency and 0.88x lower energy delay product.
Linus Jungemann, Bjarne Wintermann, Heinrich Riebler, Christian Plessl
FPGA4
2025 Analyzing performance portability for a SYCL implementation of the 2D shallow water equations
abstract
Abstract SYCL is an open standard for targeting heterogeneous hardware from C++. In this work, we evaluate a SYCL implementation for a discontinuous Galerkin discretization of the 2D shallow water equations targeting CPUs, GPUs, and also FPGAs. The discretization uses polynomial orders zero to two on unstructured triangular meshes. Separating memory accesses from the numerical code allow us to optimize data accesses for the target architecture. A performance analysis shows good portability across x86 and ARM CPUs, GPUs from different vendors, and even two variants of Intel Stratix 10 FPGAs. Measuring the energy to solution shows that GPUs yield an up to 10x higher energy efficiency in terms of degrees of freedom per joule compared to CPUs. With custom designed caches, FPGAs offer a meaningful complement to the other architectures with particularly good computational performance on smaller meshes. FPGAs with High Bandwidth Memory are less affected by bandwidth issues and have similar energy efficiency as latest generation CPUs.
Markus Büttner, Christoph Alt, Tobias Kenter, Harald Köstler, Christian Plessl, Vadym Aizinger
J. Supercomput.5
2024 Reproduction and Extension of Playing Strength Models in Computer Go
abstract
Monte Carlo Tree Search (MCTS) is a well established approach for computer players to the game of Go as well as other combinatorial problems. To improve the playing strength of the algorithm, it can be run for a longer time or be executed in parallel. The traditionally suggested regression model for the achievable playing strength with MCTS uses an exponential decay and nicely matches the observations of other authors as well as our own replication. However, it implies a rather low upper bound, even for non-parallel execution. This appears to disagree with the property of MCTS to asymptotically converge towards perfect play, which should still be far from achievable. Yet, to our knowledge, there has been neither an explicit verification of these findings nor an attempt to explain the observations differently. Here, we look into the real-world playing strength of a state-of-the-art Go program, Katago, in various environments. We find that there is at least one alternative model of very simple nature that does not impose an upper bound while matching the observations to a similar degree.
Sebastian Heuchler, Christian Plessl
CoG2
2024 Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
abstract
Abstract Most FPGA boards in the HPC domain are well-suited for parallel scaling because of the direct integration of versatile and high-throughput network ports. However, the utilization of their network capabilities is often challenging and error-prone because the whole network stack and communication patterns have to be implemented and managed on the FPGAs. Also, this approach conceptually involves a trade-off between the performance potential of improved communication and the impact of resource consumption for communication infrastructure, since the utilized resources on the FPGAs could otherwise be used for computations. In this work, we investigate this trade-off, firstly, by using synthetic benchmarks to evaluate the different configuration options of the communication framework ACCL and their impact on communication latency and throughput. Finally, we use our findings to implement a shallow water simulation whose scalability heavily depends on low-latency communication. With a suitable configuration of ACCL, good scaling behavior can be shown to all 48 FPGAs installed in the system. Overall, the results show that the availability of inter-FPGA communication frameworks as well as the configurability of framework and network stack are crucial to achieve the best application performance with low latency communication.
Marius Meyer, Tobias Kenter, Lucian Petrica, Kenneth O'Brien, Michaela Blott, Christian Plessl
Euro-Par (2)6
2024 HiHiSpMV: Sparse Matrix Vector Multiplication with Hierarchical Row Reductions on FPGAs with High Bandwidth Memory
abstract
The multiplication of a sparse matrix with a dense vector is a vital operation in linear algebra, with applications in numerous contexts. After earlier research on FPGA acceleration of this operation had shown the potential to achieve a high bandwidth efficiency, this workload has received renewed attention with the introduction of high bandwidth memory on FPGA platforms. However, previous designs fell short of scaling to the full bandwidth potential of current FPGA platforms with high bandwidth memory. In this work, we present HiHiSpMV with a novel design approach around hierarchical accumulation, which allows us to overcome several limitations of the related work. Our design is completely implemented with high-level synthesis and compiled with Vitis, and instantiates 16 independent compute units, each processing up to 16 matrix elements per clock cycle. The reduction of these elements is performed without any latency or resource overhead and the subsequent accumulation uses the well-established shift-register design pattern. Due to the independent nature of compute units, our design can connect to all of the 32 high bandwidth memory pseudo-channels in 512-bit interface mode on an Alveo U280 FPGA board. In our tests, we reach up to 86% of the theoretical available bandwidth of this FPGA platform, enabling a computational throughput of up to 98 GFLOPS. This is about 1.5 × faster than the peak throughput of the best related work in that regard, Serpens. On average, the throughput of HiHiSpMV is even 2.7× higher than Serpens.
Abdul Rehman Tareen, Marius Meyer, Christian Plessl, Tobias Kenter
FCCM3
2024 StencilStream: A SYCL-based Stencil Simulation Framework Targeting FPGAs
abstract
Although High-Level Synthesis has dramatically improved the usability of FPGAs in High-Performance Computing applications, it still requires much experience to create well-performing FPGA designs. We introduce the stencil simulation framework StencilStream to separate the concerns of performance engineers and domain scientists. We have formulated our performance-oriented stencil simulation designs for FPGAs along with their management code as C++ templates using SYCL language features. Application developers instantiate these templates with their concrete stencil code and obtain a complete application with little management overhead. In this work, we describe the most interesting API features and the two backend architectures of StencilStream, along with their corresponding performance models. We then evaluate the performance of the architectures using three manually implemented sample applications and highlight the real-world usability of StencilStream further with shallow water simulations obtained from a code-generation flow. Compiled with Intel oneAPI for an Intel Stratix 10 GX 2800 target FPGA, the best-performing sample design exceeds a throughput of 1 TFLOPS.
Jan-Oliver Opdenhövel, Christoph Alt, Christian Plessl, Tobias Kenter
FPL3
2024 SERI: High-Throughput Streaming Acceleration of Electron Repulsion Integral Computation in Quantum Chemistry using HBM-based FPGAs
abstract
The computation of electron repulsion integrals (ERIs) is a key component for quantum chemical methods. The intensive computation and bandwidth demand for ERI evaluation presents a significant challenge for quantum-mechanics-based atomistic simulations with hybrid density functional theory: due to the tens of trillions of ERI computations in each time step, practical applications are usually limited to thousands of atoms. In this work, we propose SERI, a high-throughput streaming accelerator for ERI computation on HBM-based FPGAs. In contrast to prior buffer-based designs, SERI proposes a novel streaming architecture to address the on-chip buffer limitation and the floorplanning challenge, and leverages the high-bandwidth memory to overcome the bandwidth bottleneck in prior designs. Moreover, to meet the varying computation, bandwidth, and floorplanning requirements between the 55 canonical quartet classes in ERI calculation, we design an automation tool, together with an accurate performance model, to automatically customize the architecture and floorplanning strategy for each canonical quartet class to maximize their throughput. Our performance evaluation on the AMD/Xilinx Alveo U280 FPGA board shows that, SERI achieves an average speedup of 9.80 x over the previous best-performing FPGA design, a 3.21x speedup over a 64-core AMD EPYC 7713 CPU, and a 15.64x speedup over an Nvidia A40 GPU. It reaches a peak throughput of 23.8 GERIS ($10^{9}$ ERIs per second) on one Alveo U280 FPGA. SERI will be released soon at https://github.com/SFU-HiAccel/SERI.
Philip Stachura, Christian Plessl, Zhenman Fang
FPL4
2024 A Computation of the Ninth Dedekind Number Using FPGA Supercomputing
abstract
This manuscript makes the claim of having computed the \(9\) th Dedekind number, D(9). This was done by accelerating the core operation of the process with an efficient FPGA design that outperforms an optimized 64-core CPU reference by 95 \(\times\) . The FPGA execution was parallelized on the Noctua 2 supercomputer at Paderborn University. The resulting value for D(9) is 286386577668298411128469151667598498812366. This value can be verified in two steps. We have made the data file containing the 490 M results available, each of which can be verified separately on CPU, and the whole file sums to our proposed value. The paper explains the mathematical approach in the first part, before putting the focus on a deep dive into the FPGA accelerator implementation followed by a performance analysis. The FPGA implementation was done in Register-Transfer Level using a dual-clock architecture and shows how we achieved an impressive FMax of 450 MHz on the targeted Stratix 10 GX 2,800 FPGAs. The total compute time used was 47,000 FPGA hours.
Lennart Van Hirtum, Patrick De Causmaecker, Jens Goemaere, Tobias Kenter, Heinrich Riebler, Michael Lass, Christian Plessl
ACM Trans. Reconfigurable Technol. Syst.7
2023 Computing and Compressing Electron Repulsion Integrals on FPGAs
abstract
The computation of electron repulsion integrals (ERIs) over Gaussian-type orbitals (GTOs) is a challenging problem in quantum-mechanics-based atomistic simulations. In practical simulations, several trillions of ERIs may have to be computed for every time step. In this work, we investigate FPGAs as accelerators for the ERI computation. We use template parameters, here within the Intel oneAPI tool flow, to create customized designs for 256 different ERI quartet classes, based on their orbitals. To maximize data reuse, all intermediates are buffered in FPGA on-chip memory with customized layouts. The pre-calculation of intermediates also helps to overcome data dependencies caused by multi-dimensional recurrence relations. The involved loop structures are partially or even fully unrolled for high throughput of FPGA kernels. Furthermore, a lossy compression algorithm utilizing arbitrary bitwidth integers is integrated in the FPGA kernels. To our best knowledge, this is the first work on ERI computation on FPGAs that supports more than just the single most basic quartet class. Also, the integration of ERI computation and compression is a novelty that is not even covered by CPU or GPU libraries so far. Our evaluation shows that using 16-bit integer for the ERI compression, the fastest FPGA kernels exceed the performance of 10 GERIS ($10\times 10^{9}$ERIs per second) on one Intel Stratix 10 GX 2800 FPGA, with maximum absolute errors around 10−7- 10−5Hartree. The measured throughput can be accurately explained by a performance model. The FPGA kernels deployed on 2 FPGAs outperform similar computations using the widely used libint reference on a two-socket server with 40 Xeon Gold 6148 CPU cores of the same process technology by factors up to 6.0x and on a new two-socket server with 128 EPYC 7713 CPU cores by up to 1.9x.
Tobias Kenter, Robert Schade, Thomas D. Kühne, Christian Plessl
FCCM5
2023 Multi-FPGA Designs and Scaling of HPC Challenge Benchmarks via MPI and Circuit-switched Inter-FPGA Networks
abstract
While FPGA accelerator boards and their respective high-level design tools are maturing, there is still a lack of multi-FPGA applications, libraries, and not least, benchmarks and reference implementations towards sustained HPC usage of these devices. As in the early days of GPUs in HPC, for workloads that can reasonably be decoupled into loosely coupled working sets, multi-accelerator support can be achieved by using standard communication interfaces like MPI on the host side. However, for performance and productivity, some applications can profit from a tighter coupling of the accelerators. FPGAs offer unique opportunities here when extending the dataflow characteristics to their communication interfaces. In this work, we extend the HPCC FPGA benchmark suite by multi-FPGA support and three missing benchmarks that particularly characterize or stress inter-device communication: b_eff, PTRANS, and LINPACK. With all benchmarks implemented for current boards with Intel and Xilinx FPGAs, we established a baseline for multi-FPGA performance. Additionally, for the communication-centric benchmarks, we explored the potential of direct FPGA-to-FPGA communication with a circuit-switched inter-FPGA network that is currently only available for one of the boards. The evaluation with parallel execution on up to 26 FPGA boards makes use of one of the largest academic FPGA installations.
Marius Meyer, Tobias Kenter, Christian Plessl
ACM Trans. Reconfigurable Technol. Syst.3
2022 A High-Fidelity Flow Solver for Unstructured Meshes on Field-Programmable Gate Arrays: Design, Evaluation, and Future Challenges
abstract
The impending termination of Moore’s law motivates the search for new forms of computing to continue the performance scaling we have grown accustomed to. Among the many emerging Post-Moore computing candidates, perhaps none is as salient as the Field-Programmable Gate Array (FPGA), which offers the means of specializing and customizing the hardware to the computation at hand.
Martin Karp, Artur Podobas, Tobias Kenter, Niclas Jansson, Christian Plessl, Philipp Schlatter, Stefano Markidis
HPC Asia5
2022 The HighPerMeshes framework for numerical algorithms on unstructured grids
abstract
Summary Solving partial differential equations (PDEs) on unstructured grids is a cornerstone of engineering and scientific computing. Heterogeneous parallel platforms, including CPUs, GPUs, and FPGAs, enable energy‐efficient and computationally demanding simulations. In this article, we introduce the HighPerMeshes C++‐embedded domain‐specific language (DSL) that bridges the abstraction gap between the mathematical formulation of mesh‐based algorithms for PDE problems on the one hand and an increasing number of heterogeneous platforms with their different programming models on the other hand. Thus, the HighPerMeshes DSL aims at higher productivity in the code development process for multiple target platforms. We introduce the concepts as well as the basic structure of the HighPerMeshes DSL, and demonstrate its usage with three examples. The mapping of the abstract algorithmic description onto parallel hardware, including distributed memory compute clusters, is presented. A code generator and a matching back end allow the acceleration of HighPerMeshes code with GPUs. Finally, the achievable performance and scalability are demonstrated for different example problems.
Samer Alhaddad, Jens Förstner, Stefan Groth, Daniel Grünewald, Yevgen Grynko, Frank Hannig, Tobias Kenter, Franz-Josef Pfreundt, Christian Plessl, Merlind Schotte, Thomas Steinke 0001, Jürgen Teich, Martin Weiser, Florian Wende
Concurr. Comput. Pract. Exp.9
2022 In-depth FPGA accelerator performance evaluation with single node benchmarks from the HPC challenge benchmark suite for Intel and Xilinx FPGAs using OpenCL
Marius Meyer, Tobias Kenter, Christian Plessl
J. Parallel Distributed Comput.3
2022 Towards electronic structure-based ab-initio molecular dynamics simulations with hundreds of millions of atoms
abstract
We push the boundaries of electronic structure-based ab-initio molecular dynamics (AIMD) beyond 100 million atoms. This scale is otherwise barely reachable with classical force-field methods or novel neural network and machine learning potentials. We achieve this breakthrough by combining innovations in linear-scaling AIMD, efficient and approximate sparse linear algebra, low and mixed-precision floating-point computation on GPUs, and a compensation scheme for the errors introduced by numerical approximations. The core of our work is the non-orthogonalized local submatrix method (NOLSM), which scales very favorably to massively parallel computing systems and translates large sparse matrix operations into highly parallel, dense matrix operations that are ideally suited to hardware accelerators. We demonstrate that the NOLSM method, which is at the center point of each AIMD step, is able to achieve a sustained performance of 324 PFLOP/s in mixed FP16/FP32 precision corresponding to an efficiency of 67.7% when running on 1536 NVIDIA A100 GPUs.
Robert Schade, Tobias Kenter, Hossam Elgabarty, Michael Lass, Ole Schütt, Alfio Lazzaro, Hans Pabst, Stephan Mohr, Jürg Hutter, Thomas D. Kühne, Christian Plessl
Parallel Comput.11
2022 The Strong Scaling Advantage of FPGAs in HPC for N-body Simulations
abstract
N-body methods are one of the essential algorithmic building blocks of high-performance and parallel computing. Previous research has shown promising performance for implementing n-body simulations with pairwise force calculations on FPGAs. However, to avoid challenges with accumulation and memory access patterns, the presented designs calculate each pair of forces twice, along with both force sums of the involved particles. Also, they require large problem instances with hundreds of thousands of particles to reach their respective peak performance, limiting the applicability for strong scaling scenarios. This work addresses both issues by presenting a novel FPGA design that uses each calculated force twice and overlaps data transfers and computations in a way that allows to reach peak performance even for small problem instances, outperforming previous single precision results even in double precision, and scaling linearly over multiple interconnected FPGAs. For a comparison across architectures, we provide an equally optimized CPU reference, which for large problems actually achieves higher peak performance per device, however, given the strong scaling advantages of the FPGA design, in parallel setups with few thousand particles per device, the FPGA platform achieves highest performance and power efficiency.
Johannes Menzel, Christian Plessl, Tobias Kenter
ACM Trans. Reconfigurable Technol. Syst.2
2021 High-Performance Spectral Element Methods on Field-Programmable Gate Arrays : Implementation, Evaluation, and Future Projection
abstract
Improvements in computer systems have historically relied on two well-known observations: Moore's law and Dennard's scaling. Today, both these observations are ending, forcing computer users, researchers, and practitioners to abandon the general-purpose architectures' comforts in favor of emerging post-Moore systems. Among the most salient of these post-Moore systems is the Field-Programmable Gate Array (FPGA), which strikes a convenient balance between complexity and performance. In this paper, we study modern FPGAs' applicability in accelerating the Spectral Element Method (SEM) core to many computational fluid dynamics (CFD) applications. We design a custom SEM hardware accelerator operating in double-precision that we empirically evaluate on the latest Stratix 10 GX-series FPGAs and position its performance (and power-efficiency) against state-of-the-art systems such as ARM ThunderX2, NVIDIA Pascal/Volta/Ampere Teslaseries cards, and general-purpose manycore CPUs. Finally, we develop a performance model for our SEM-accelerator, which we use to project future FPGAs' performance and role to accelerate CFD applications, ultimately answering the question: what characteristics would a perfect FPGA for CFD applications have?
Martin Karp, Artur Podobas, Niclas Jansson, Tobias Kenter, Christian Plessl, Philipp Schlatter, Stefano Markidis
IPDPS5
2020 Efficient Ab-Initio Molecular Dynamic Simulations by Offloading Fast Fourier Transformations to FPGAs
abstract
A large share of today's HPC workloads is used for Ab-Initio Molecular Dynamics (AIMD) simulations, where the interatomic forces are computed on-the-fly by means of accurate electronic structure calculations. They are computationally intensive and thus constitute an interesting application class for energy-efficient hardware accelerators such as FPGAs. In this paper, we investigate the potential of offloading 3D Fast Fourier Transformations (FFTs) as a critical routine of plane-wave-based electronic structure calculations to FPGA and in conjunction demonstrate the tolerance of these simulations to lower precision computations.
Arjun Ramaswami, Tobias Kenter, Thomas D. Kühne, Christian Plessl
FPL4
2020 A submatrix-based method for approximate matrix function evaluation in the quantum chemistry code CP2K
abstract
Electronic structure calculations based on density-functional theory (DFT) represent a significant part of today's HPC workloads and pose high demands on high-performance computing resources. To perform these quantum-mechanical DFT calculations on complex large-scale systems, so-called linear scaling methods instead of conventional cubic scaling methods are required. In this work, we take up the idea of the submatrix method and apply it to the DFT computations in the software package CP2K. For that purpose, we transform the underlying numeric operations on distributed, large, sparse matrices into computations on local, much smaller and nearly dense matrices. This allows us to exploit the full floating-point performance of modern CPUs and to make use of dedicated accelerator hardware, where performance has been limited by memory bandwidth before. We demonstrate both functionality and performance of our implementation and show how it can be accelerated with GPUs and FPGAs.
Michael Lass, Robert Schade, Thomas D. Kühne, Christian Plessl
SC4
2019 Transparent Acceleration for Heterogeneous Platforms With Compilation to OpenCL
abstract
Multi-accelerator platforms combine CPUs and different accelerator architectures within a single compute node. Such systems are capable of processing parallel workloads very efficiently while being more energy efficient than regular systems consisting of CPUs only. However, the architectures of such systems are diverse, forcing developers to port applications to each accelerator using different programming languages, models, tools, and compilers. Developers not only require domain-specific knowledge but also need to understand the low-level accelerator details, leading to an increase in the design effort and costs. To tackle this challenge, we propose a compilation approach and a practical realization called HT r OP that is completely transparent to the user. HT r OP is able to automatically analyze a sequential CPU application, detect computational hotspots, and generate parallel OpenCL host and kernel code. The potential of HT r OP is demonstrated by offloading hotspots to different OpenCL-enabled resources (currently the CPU, the general-purpose GPU, and the manycore Intel Xeon Phi) for a broad set of benchmark applications. We present an in-depth evaluation of our approach in terms of performance gains and energy savings, taking into account all static and dynamic overheads. We are able to achieve speedups and energy savings of up to two orders of magnitude, if an application has sufficient computational intensity, when compared to a natively compiled application.
Heinrich Riebler, Gavin Vaz, Tobias Kenter, Christian Plessl
ACM Trans. Archit. Code Optim.4
2018 OpenCL-Based FPGA Design to Accelerate the Nodal Discontinuous Galerkin Method for Unstructured Meshes
abstract
The exploration of FPGAs as accelerators for scientific simulations has so far mostly been focused on small kernels of methods working on regular data structures, for example in the form of stencil computations for finite difference methods. In computational sciences, often more advanced methods are employed that promise better stability, convergence, locality and scaling. Unstructured meshes are shown to be more effective and more accurate, compared to regular grids, in representing computation domains of various shapes. Using unstructured meshes, the discontinuous Galerkin method preserves the ability to perform explicit local update operations for simulations in the time domain. In this work, we investigate FPGAs as target platform for an implementation of the nodal discontinuous Galerkin method to find time-domain solutions of Maxwell's equations in an unstructured mesh. When maximizing data reuse and fitting constant coefficients into suitably partitioned on-chip memory, high computational intensity allows us to implement and feed wide data paths with hundreds of floating point operators. By decoupling off-chip memory accesses from the computations, high memory bandwidth can be sustained, even for the irregular access pattern required by parts of the application. Using the Intel/Altera OpenCL SDK for FPGAs, we present different implementation variants for different polynomial orders of the method. In different phases of the algorithm, either computational or bandwidth limits of the Arria 10 platform are almost reached, thus outperforming a highly multithreaded CPU implementation by around 2x.
Tobias Kenter, Gopinath Mahale, Samer Alhaddad, Yevgen Grynko, Christian Schmitt 0003, Ayesha Afzal, Frank Hannig, Jens Förstner, Christian Plessl
FCCM9
2018 Automated code acceleration targeting heterogeneous openCL devices
abstract
Accelerators can offer exceptional performance advantages. However, programmers need to spend considerable efforts on acceleration, without knowing how sustainable the employed programming models, languages and tools are. To tackle this challenge, we propose and demonstrate a new runtime system called HTrOP that is able to automatically generate and execute OpenCL code from sequential CPU code. HTrOP transforms suitable data-parallel loops into independent OpenCL-typical work-items and handles concrete calls to these devices through a mix of library components and application-specific OpenCL host code. Computational hotspots are identified and can be offloaded to different resources (CPU, GPGPU and Xeon Phi). We demonstrate the potential of HTrOP on a broad set of applications and are able to improve the performance by 4.3X on average.
Heinrich Riebler, Gavin Vaz, Tobias Kenter, Christian Plessl
PPoPP4
2017 Flexible FPGA design for FDTD using OpenCL
abstract
Compared to classical HDL designs, generating FPGA with high-level synthesis from an OpenCL specification promises easier exploration of different design alternatives and, through ready-to-use infrastructure and common abstractions for host and memory interfaces, easier portability between different FPGA families. In this work, we evaluate the extent of this promise. To this end, we present a parameterized FDTD implementation for photonic microcavity simulations. Our design can trade-off different forms of parallelism and works for two independent OpenCL-based FPGA design flows. Hence, we can target FPGAs from different vendors and different FPGA families. We describe how we used pre-processor macros to achieve this flexibility and to work around different shortcomings of the current tools. Choosing the right design configurations, we are able to present two extremely competitive solutions for very different FPGA targets, reaching up to 172 GFLOPS sustained performance. With the portability and flexibility demonstrated, code developers not only avoid vendor lock-in, but can even make best use of real trade-offs between different architectures.
Tobias Kenter, Jens Förstner, Christian Plessl
FPL3
2017 Foreword to the special issue of the 18th IEEE international conference on computational science and engineering (CSE2015)
abstract
The Computational Science and Engineering (CSE) area has earned prominence through advances in electronic and integrated technologies. Advanced computing systems permeate our daily life and have an increasingly importance in many aspects and domains. CSE is shaping future research and development activities in academia and industry, ranging from engineering, science, finance, economics, healthcare, arts, and humanitarian fields. The IEEE International Conference on CSE has been providing a series of highly successful International Conferences on CSE. The 2015 edition of CSE, CSE2015 (http://www.fe.up.pt/cse2015), was held in Porto, Portugal on October 21–23, 2015. It brought together computer scientists, industrial engineers, and researchers to discuss and exchange experimental and theoretical results, work-in-progress, experiences, case studies, and trend-setting ideas, in the areas of advanced computing for solving problems in science and engineering applications. The six extended papers included have been selected from a preliminary set of 12 papers submitted to this special issue and are briefly described as follows. The article ‘Robust resource allocations through performance modeling with stochastic process algebra’ 1 presents a new resource allocation scheme that uses a stochastic process algebra for obtaining resource allocations. These allocations are robust with respect to unpredictable perturbations of the application or system characteristics during runtime. The key idea is to translate performance models into mathematical Markov chain descriptions that can be numerically evaluated without requiring the time consuming simulation process used in competing approaches. A comparison with previous studies shows that the proposed approach achieves competitive results. In addition, process algebras are easier to reproduce because they require neither the effort for learning a simulation framework nor setup or installation cost. Beyond that, the computational efficiency of the approach also allows for embedding the proposed process algebra model into the runtime systems of a model-based framework that re-evaluates the resource assignment whenever a system or application parameter changes at runtime. The article ‘Heterogeneous CPU + GPU Approaches for Mesh Refinement over Lattice-Boltzmann Simulations’ 2 investigates strategies for mapping Lattice-Boltzmann method (LBM) simulations to compute nodes with CPUs and GPUs. The particular challenge addressed is finding an efficient way so that adaptive mesh refinement strategies use both computing resources effectively. While parallelism is abundant in LBM simulations, the challenge is to structure the workload distribution and data access to perform well on both CPU and GPU, which inherently favor different granularities of parallelism. The authors propose two approaches, a multi-domain approach that uses a finer grid in domains where a higher resolution is required; and an irregular grid approach that uses a single Cartesian with non-uniform spacing. Both approaches (multi-domain and irregular grid) are implemented and evaluated for a system comprising a Xeon E5 CPU and an NVidia K20c GPU using either only the GPU or CPU + GPU for the LBM computation. The evaluation shows that the multi-domain approach allows for executing bigger simulations because it requires fewer lattice nodes. The irregular grid approach is easier to implement and delivers a higher throughput (million fluid lattice cell updates per second). For both methods, the CPU + GPU implementation outperforms the homogeneous GPU implementation by 10–30%. The article ‘Methods to Model and Simulate Super Carbon Nanotubes of Higher Order’ 3 presents a new, efficient approach based on graph algebra to simulate the mechanical behavior of super carbon nanotubes (SCNTs). Representing the SCNTs as directed graphs, the authors propose a new data structure that exploits the hierarchy of SCNTs for fast queries. In addition, they propose a novel, iterative solver using the conjugate gradient method. Exploiting the symmetry of SCNTs of order 0, the solver is able to drastically reduce the amount of required calculations and memory for small deformations. Further exploiting structural symmetry and adopting an improved proximity-aware Matrix–vector-Multiplication routine, the performance for SCNTs level 0 can be improved by an additional factor of 2. Up to 4.4 times speedup is achieved when running in parallel on a 16 core SMP system. The authors also explore optimizations for symmetry in SCNTs of order 1. Experimental results show that the new approach outperforms a compressed-row-storage-based reference solver, for SCNTs of order 0 and 1, regardless of deformation, and with much less memory consumption. Because in practice memory consumption is oftentimes the limiting factor for scaling, this approach can significantly expand the realm of feasible simulations for SCNTs. In the article, the readers can also find introductions to the basic mathematical formulations and algorithms for simulating SCNTs as well as a brief summary of previous results. The article ‘Combinatorial Optimization of DNA Sequence Analysis on Heterogeneous Systems’ 4 presents an experimental study of counting the occurrences of query patterns in DNA sequences in parallel. DNA sequence analysis has many important practical applications, and is both data and computation intensive. The authors parallelize the Aho Corasick pattern matching algorithm, and explore its execution on heterogeneous systems, such as Intel Xeon E5 with Xeon Phi as co-processor, for acceleration. To achieve maximal performance on heterogeneous systems, the authors employ simulated annealing to determine the number of threads, thread affinities, and data placement on the host and the accelerator as such configuration is critical to performance and system utilization. Using real-world DNA sequences, the authors evaluate the efficiency of their approach. They show that the average speedup achieved is 1.6 times compared against the host-only parallelization and 2 times against device-only parallelization. The article ‘Using Adaptive Runtime Filtering to Support an Event-based Performance Analysis’ 5 presents an approach to filter tracing data in the context of event-based monitoring for performance improvements. The approach is based on self-guided filters that automatically adapt to an application's runtime behavior and are able to reduce performance data to manageable sizes for large-scale parallel applications and long execution programs. The article presents four runtime filters, each one targeting a specific type of data redundancy. They evaluate their approach with five real-world applications from different scientific domains. Compared to the default settings, their filters achieve a data size reduction of two orders of magnitude while increasing execution time regarding the overhead of tracing by less than one percent on average. The article also examines the influence of filtering on the performance analysis and identifies its limitations and presents three schemes to help performance analysts to correctly interpret filtered traces or even reconstruct parts of a filtered trace. The article ‘Automatic source-to-source error compensation of floating-point programs: code synthesis to optimize accuracy and time’ 6 presents a source-to-source C compiler approach for automatically improving the numerical accuracy of floating-point programs without significantly increasing execution time. The approach is based on the automatic compensation of floating-point operations by applying error-free transformations, and on the synthesis of code for both accuracy and execution time criteria. The use of partial compensation is proposed in order to trade-off performance and accuracy. The article also presents a number of code transformations to increase accuracy and to tune the impact on execution time. In addition, the authors present a method to find the best transformation satisfying execution time or accuracy constraints. The approach presented is evaluated with a number of case studies using two target computing environments and is able to produce some compensated algorithms as accurate and efficient as the ones derived by hand. We would like to acknowledge the authors of the articles included in this special issue for the hard work on preparing high-quality papers, the work of the anonymous reviewers on providing very important insights and suggestions that undoubtedly helped authors to improve their papers, and the support of the CCPE editors, Geoffrey C. Fox and David W. Walker.
Christian Plessl, Guojing Cong, João M. P. Cardoso
Concurr. Comput. Pract. Exp.1
2017 Efficient Branch and Bound on FPGAs Using Work Stealing and Instance-Specific Designs
abstract
Branch and bound (B8B) algorithms structure the search space as a tree and eliminate infeasible solutions early by pruning subtrees that cannot lead to a valid or optimal solution. Custom hardware designs significantly accelerate the execution of these algorithms. In this article, we demonstrate a high-performance B8B implementation on FPGAs. First, we identify general elements of B8B algorithms and describe their implementation as a finite state machine. Then, we introduce workers that autonomously cooperate using work stealing to allow parallel execution and full utilization of the target FPGA. Finally, we explore advantages of instance-specific designs that target a specific problem instance to improve performance. We evaluate our concepts by applying them to a branch and bound problem, the reconstruction of corrupted AES keys obtained from cold-boot attacks. The evaluation shows that our work stealing approach is scalable with the available resources and provides speedups proportional to the number of workers. Instance-specific designs allow us to achieve an overall speedup of 47 × compared to the fastest implementation of AES key reconstruction so far. Finally, we demonstrate how instance-specific designs can be generated just-in-time such that the provided speedups outweigh the additional time required for design synthesis.
Heinrich Riebler, Michael Lass, Robert Mittendorf, Thomas Löcke, Christian Plessl
ACM Trans. Reconfigurable Technol. Syst.5
2016 Performance-centric scheduling with task migration for a heterogeneous compute node in the data center
Achim Lösch, Tobias Beisel, Tobias Kenter, Christian Plessl, Marco Platzner
DATE4
2015 Transparent offloading of computational hotspots from binary code to Xeon Phi
Marvin Damschen, Heinrich Riebler, Gavin Vaz, Christian Plessl
DATE4
2014 Reconstructing AES Key Schedules from Decayed Memory with FPGAs
abstract
In this paper, we study how AES key schedules can be reconstructed from decayed memory. This operation is a crucial and time consuming operation when trying to break encryption systems with cold-boot attacks. In software, the reconstruction of the AES master key can be performed using a recursive, branch-and-bound tree-search algorithm that exploits redundancies in the key schedule for constraining the search space. In this work, we investigate how this branch-and-bound algorithm can be accelerated with FPGAs. We translate the recursive search procedure to a state machine with an explicit stack for each recursion level and create optimized datapaths to accelerate in particular the processing of the most frequently accessed tree levels. We support two different decay models, of which especially the more realistic non-idealized asymmetric decay model causes very high runtimes in software. Our implementation on a Maxeler dataflow computing system outperforms a software implementation for this model by up to 27x, which makes cold-boot attacks against AES practical even for high error rates.
Heinrich Riebler, Tobias Kenter, Christian Plessl, Christoph Sorge
FCCM3
2014 Runtime Resource Management in Heterogeneous System Architectures: The SAVE Approach
abstract
Modern computing systems featuring different kinds of processing elements have proven to be efficient in terms of performance/energy trade-offs. Furthermore these systems usually have to execute multiple concurrent tasks without any apriori knowledge on expected arrival times, in an unpredictable and very dynamic environment. This scenario has propelled an interest towards self-adaptive systems that dynamically reorganize the use of system resources to optimize for a given goal. The SAVE project will develop a Heterogeneous System Architecture that will decide at runtime to execute task on the appropriate kind of resources, based on the current requirements. This paper presents a first implementation of a resource allocation policy that dynamically shares heterogeneous resources between multiple running applications. Resource allocation mechanisms are discussed and evaluated in an experimental campaign, showing how the policy helps in attaining users' applications goals.
Gianluca Durelli, Marcello Pogliani, Antonio Miele, Christian Plessl, Heinrich Riebler, Marco D. Santambrogio, Gavin Vaz, Cristiana Bolchini
ISPA4
2014 Self-Awareness as a Model for Designing and Operating Heterogeneous Multicores
abstract
Self-aware computing is a paradigm for structuring and simplifying the design and operation of computing systems that face unprecedented levels of system dynamics and thus require novel forms of adaptivity. The generality of the paradigm makes it applicable to many types of computing systems and, previously, researchers started to introduce concepts of self-awareness to multicore architectures. In our work we build on a recent reference architectural framework as a model for self-aware computing and instantiate it for an FPGA-based heterogeneous multicore running the ReconOS reconfigurable architecture and operating system. After presenting the model for self-aware computing and ReconOS, we demonstrate with a case study how a multicore application built on the principle of self-awareness, autonomously adapts to changes in the workload and system state. Our work shows that the reference architectural framework as a model for self-aware computing can be practically applied and allows us to structure and simplify the design process, which is essential for designing complex future computing systems.
Andreas Agne, Markus Happe, Achim Lösch, Christian Plessl, Marco Platzner
ACM Trans. Reconfigurable Technol. Syst.4
2013 FPGA-accelerated key search for cold-boot attacks against AES
abstract
Cold-boot attacks exploit the fact that DRAM contents are not immediately lost when a PC is powered off. Instead the contents decay rather slowly, in particular if the DRAM chips are cooled to low temperatures. This effect opens an attack vector on cryptographic applications that keep decrypted keys in DRAM. An attacker with access to the target computer can reboot it or remove the RAM modules and quickly copy the RAM contents to non-volatile memory. By exploiting the known cryptographic structure of the cipher and layout of the key data in memory, in our application an AES key schedule with redundancy, the resulting memory image can be searched for sections that could correspond to decayed cryptographic keys; then, the attacker can attempt to reconstruct the original key. However, the runtime of these algorithms grows rapidly with increasing memory image size, error rate and complexity of the bit error model, which limits the practicability of the approach. In this work, we study how the algorithm for key search can be accelerated with custom computing machines. We present an FPGA-based architecture on a Maxeler dataflow computing system that outperforms a software implementation up to 205x, which significantly improves the practicability of cold-attacks against AES.
Heinrich Riebler, Tobias Kenter, Christoph Sorge, Christian Plessl
FPT4
2013 On-The-Fly Computing: A novel paradigm for individualized IT services
abstract
In this paper we introduce “On-The-Fly Computing”, our vision of future IT services that will be provided by assembling modular software components available on world-wide markets. After suitable components have been found, they are automatically integrated, configured and brought to execution in an On-The-Fly Compute Center. We envision that these future compute centers will continue to leverage three current trends in large scale computing which are an increasing amount of parallel processing, a trend to use heterogeneous computing resources, and — in the light of rising energy cost — energy-efficiency as a primary goal in the design and operation of computing systems. In this paper, we point out three research challenges and our current work in these areas.
Markus Happe, Friedhelm Meyer auf der Heide, Peter Kling, Marco Platzner, Christian Plessl
ISORC5
2012 Convey vector personalities - FPGA acceleration with an openmp-like programming effort?
abstract
Although the benefits of FPGAs for accelerating scientific codes are widely acknowledged, the use of FPGA accelerators in scientific computing is not widespread because reaping these benefits requires knowledge of hardware design methods and tools that is typically not available with domain scientists. A promising but hardly investigated approach is to develop tool flows that keep the common languages for scientific code (C,C++, and Fortran) and allow the developer to augment the source code with OpenMP-like directives for instructing the compiler which parts of the application shall be offloaded the FPGA accelerator. In this work we study whether the promise of effective FPGA acceleration with an OpenMP-like programming effort can actually be held. Our target system is the Convey HC-1 reconfigurable computer for which an OpenMP-like programming environment exists. As case study we use an application from computational nanophotonics. Our results show that a developer without previous FPGA experience could create an FPGA-accelerated application that is competitive to an optimized OpenMP-parallelized CPU version running on a two socket quad-core server. Finally, we discuss our experiences with this tool flow and the Convey HC-1 from a productivity and economic point of view.
Björn Meyer, Jörn Schumacher, Christian Plessl, Jens Förstner
FPL3
2012 Exploration of ring oscillator design space for temperature measurements on FPGAs
abstract
While numerous publications have presented ring oscillator designs for temperature measurements a detailed study of the ring oscillator's design space is still missing. In this work, we introduce metrics for comparing the performance and area efficiency of ring oscillators and a methodology for determining these metrics. As a result, we present a systematic study of the design space for ring oscillators for a Xilinx Virtex-5 platform FPGA.
Christoph Ruething, Andreas Agne, Markus Happe, Christian Plessl
FPL4
2011 Cooperative multitasking for heterogeneous accelerators in the Linux Completely Fair Scheduler
abstract
This paper presents an extension of the Completely Fair Scheduler (CFS) to support cooperative multitasking with time-sharing for heterogeneous processing elements in Linux. We extend the kernel to be aware of accelerators, hold different run queues for these components and perform scheduling decisions using application provided meta information and a fairness measure. Our additional programming model allows the integration of checkpoints into applications, which permits the preemption and subsequent migration of applications between accelerators. We show that cooperative multitasking is possible on heterogeneous systems and that it increases application performance and system utilization.
Tobias Beisel, Tobias Wiersema, Christian Plessl, André Brinkmann
ASAP3
2011 Performance estimation framework for automated exploration of CPU-accelerator architectures
abstract
In this paper we present a fast and fully automated approach for studying the design space when interfacing reconfigurable accelerators with a CPU. Our challenge is, that a reasonable evaluation of architecture parameters requires a hardware/software partitioning that makes best use of each given architecture configuration. Therefore we developed a framework based on the LLVM infrastructure that performs this partitioning with high-level estimation of the runtime on the target architecture utilizing profiling information and code analysis. By making use of program characteristics also during the partitioning process, we improve previous results for various benchmarks and especially for growing interface latencies between CPU and accelerator.
Tobias Kenter, Christian Plessl, Marco Platzner, Michael Kauschke
FPGA2
2010 Using shared library interposing for transparent application acceleration in systems with heterogeneous hardware accelerators
abstract
Todays computer systems increasingly comprise heterogeneous computing elements like multi-core processors, graphics processing units, and specialized co-processors, which allow parallel processing. Programming applications to utilize such systems is a complex process and needs good knowledge about the hardware architecture. Automatic and transparent use of these resources is a major concern of domain specific software developers and users. We present a new approach of using shared library interposing to replace libraries in binary applications with highly optimized accelerated versions. A plugin-based framework was developed, which allows interposing shared library calls, delegating them to accelerator specific libraries and adapting them to the library specific interface. Accelerator specific plugins can be added with a high degree of automatism. First steps were taken to develop a fast and intelligent selection component, choosing the best possible accelerator for a shared library call. It was shown, that such a framework may be efficiently used to apply shared library interposing to transparently speedup existing applications. The BLAS library for linear algebra was used as an example to develop plugins for an acceleratable library. Runtimes of BLAS functions were measured on different architectures and expose significant differences depending on the used implementation and hardware, showing the potentially high speedups of the approach.
Tobias Beisel, Manuel Niekamp, Christian Plessl
ASAP3
2009 IMORC: Application Mapping, Monitoring and Optimization for High-Performance Reconfigurable Computing
abstract
Mapping applications that consist of a collection of cores to FPGA accelerators and optimizing their performance is a challenging task in high performance reconfigurable computing. We present IMORC, an architectural template and highly versatile on-chip interconnect. IMORC links provide asynchronous FIFOs and bitwidth conversion which allows for flexibly composing accelerators from cores running at full speed within their own clock domains, thus facilitating the re-use of cores and portability. Further, IMORC inserts performance counters for monitoring runtime data. In this paper, we introduce the IMORC architectural template and the on-chip interconnect and demonstrate IMORC on the example of accelerating the 𝑘-th nearest neighbor thinning problem on an XtremeData XD1000 reconfigurable computing system.
Tobias Schumacher 0001, Christian Plessl, Marco Platzner
FCCM2
2009 An accelerator for K-TH nearest neighbor thinning based on the IMORC infrastructure
abstract
The creation and optimization of FPGA accelerators comprising several compute cores and memories are challenging tasks in high performance reconfigurable computing. In this paper, we present the design of such an accelerator for the kth nearest neighbor thinning problem on an XD1000 reconfigurable computing system. The design leverages IMORC, an architectural template and highly versatile on-chip interconnect, to achieve speedups of 74 times over a 2.2 GHz Opteron. Using IMORC with its asynchronous FIFOs and bitwidth conversion in the links between the cores, we are able to quickly create acclerator versions with varying degrees of core-level parallelism and memory mappings. Through the performance monitoring infrastructure of IMORC we gain insight into the data-dependent behavior of the accelerator which facilitates further performance optimizations.
Tobias Schumacher 0001, Christian Plessl, Marco Platzner
FPL2
2009 PermaDAQ: A scientific instrument for precision sensing and data recovery in environmental extremes
Jan Beutel, Stephan Gruber, Andreas Hasler, Roman Lim, Andreas Meier 0003, Christian Plessl, Igor Talzi, Lothar Thiele, Christian F. Tschudin, Matthias Woehrle, Mustafa Yuecel
IPSN6
2009 Demo abstract: Operating a sensor network at 3500 m above sea level
Jan Beutel, Stephan Gruber, Andreas Hasler, Roman Lim, Andreas Meier 0003, Christian Plessl, Igor Talzi, Lothar Thiele, Christian F. Tschudin, Matthias Woehrle, Mustafa Yuecel
IPSN6
2006 Optimal temporal partitioning based on slowdown and retiming
abstract
This paper presents a novel method for optimal temporal partitioning of sequential circuits for time-multiplexed reconfigurable architectures. The method bases on slowdown and retiming and maximizes the circuit's performance during execution while restricting the size of the partitions to respect the resource constraints of the reconfigurable architecture. A mixed integer linear program (MILP) formulation of the problem was provided, which can be solved exactly. In contrast to related work, our approach optimizes performance directly, takes structural modifications of the circuit into account, and is extensible. The application of the new method to temporal partitioning for a coarse-grained reconfigurable architecture was presented
Christian Plessl, Marco Platzner, Lothar Thiele
FPT1
2005 Zippy - A coarse-grained reconfigurable array with support for hardware virtualization
abstract
This paper motivates the use of hardware visualization on coarse-grained reconfigurable architectures. We introduce Zippy, a coarse-grained multi-context hybrid CPU with architectural support for efficient hardware virtualization. The architectural details and the corresponding tool flow are outlined. As a case study, we compare the non-virtualized and the virtualized execution of an ADPCM decoder.
Christian Plessl, Marco Platzner
ASAP1
2003 Virtualizing Hardware with Multi-context Reconfigurable Arrays
Rolf Enzler, Christian Plessl, Marco Platzner
FPL2
2003 TKDM - a reconfigurable co-processor in a PC's memory slot
abstract
This paper presents TKDM, a PC-based high-performance reconfigurable computing environment. The TKDM hardware consists of an FPGA module that uses the DIMM (dual inline memory module) bus for high-bandwidth and low-latency communication with the host CPU. The system's firmware is integrated with the Linux host operating system and offers functions for data communication and FPGA reconfiguration. The intended use of TKDM is that a dynamically reconfigurable co-processor for data streaming applications. The system's firmware can be customized for specific application domains to facilitate simple and easy-to-use programming interfaces.
Christian Plessl, Marco Platzner
FPT1
2003 The case for reconfigurable hardware in wearable computing
Christian Plessl, Rolf Enzler, Herbert Walder, Jan Beutel, Marco Platzner, Lothar Thiele, Gerhard Tröster
Pers. Ubiquitous Comput.1
2003 Instance-Specific Accelerators for Minimum Covering
Christian Plessl, Marco Platzner
J. Supercomput.1
2002 Custom Computing Machines for the Set Covering Problem
abstract
We present instance-specific custom computing machines for the set covering problem. Four accelerator architectures are developed that implement branch & bound in 3-valued logic and many of the deduction techniques found in software solvers. We use set covering benchmarks from two-level logic minimization and Steiner triple systems to derive and discuss experimental results. The resulting raw speedups are in the order of four magnitudes on average. Finally, we propose a hybrid solver architecture that combines the raw speed of instance-specific reconfigurable hardware with flexible bounding schemes implemented in software.
Christian Plessl, Marco Platzner
FCCM1
2002 Partially Reconfigurable Cores for Xilinx Virtex
Matthias Dyer, Christian Plessl, Marco Platzner
FPL2