EDBT 2026 Demo / reviewers in the wild / expert
Peter Zipf
dblp:z/PeterZipf
· DBLP profile ↗
61ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-4725-4246ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 58 · 6 first-author · 11 since 2021Theory of computation · 3Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discovering Optimal Constant Matrix Multiplication Circuits With Boolean SatisfiabilityabstractWhen designing arithmetic circuits, the multiplication of a variable by a constant number can be performed by a series of add/subtract, and shift operations instead of a multiplication (e.g., 7x = 23x – x = (x << 3) – x). This allows for significant reductions in complexity of multiplication circuits, as adders are typically much more hardware-efficient than a generic multiplier, and shift operations can be performed without any overhead by appropriately connecting the concerned signals. This concept extends to the multiplication of a constant matrix by a vector of variables, known as constant matrix multiplication (CMM). So far, there exists no way of discovering CMM algorithms with provably minimal add/subtract count. Here we show how the CMM problem can be reduced to a series of Boolean Satisfiability (SAT) problems, enabling the use of powerful SAT solvers to determine optimal solutions for arbitrary matrices. Modeling the problem in a closed mathematical framework allows us to straightforwardly extend our algorithm towards secondary objectives, namely word-size reductions of the add/subtract operations and pipelining for throughput maximization. Compared to state-of-the-art heuristic methods we are consistently able to achieve improvements regarding the resulting implementation complexity, even for small but practical matrix sizes such as 2 x 2 or 3 x 3. These results enable a reduction in implementation costs for a wide range of practical applications such as digital filters, convolutional cores in artificial neural networks, or the multiplication by complex constants in discrete transforms like the Fast Fourier Transform. Nicolai Fiege, Peter Zipf |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Multiplexer Optimizations for Virtex FPGAsabstractMultiplexers (MUX) are essential elements in FieldProgrammable Gate Arrays (FPGA), widely used in practical applications. Due to the LUT-based architecture of FPGAs, multiplexers that switch among many signals or operate on large word sizes incur significant resource costs, as these costs scale linearly with the data word size. Vivado's automatic synthesis flow often produces sub-optimal MUX implementations, necessitating hand-crafted solutions to minimize resource overhead. Here, we present three MUX implementation schemes that reduce resource usage for various input signal counts. These optimizations enable enhanced resource efficiency in applications ranging from circuits generated by High-Level Synthesis (HLS) tools to optimized digital filters and artificial neural networks. Nicolai Fiege, Martin Hardieck, Peter Zipf |
FPL | 3 |
| 2025 | Improving Boolean Satisfiability-Based Modulo SchedulingabstractModulo scheduling is a highly effective approach for maximizing throughput in loops with static memory dependencies, interleaving computations across consecutive loop iterations. Despite substantial advancements in scheduling procedures, it remains the most computationally intensive phase for high-level synthesis flows. A recent approach encodes the modulo scheduling problem as a series of Boolean satisfiability (SAT) instances, capitalizing on the efficiency of modern SAT solvers. This approach significantly reduces solving time and increases the availability of throughput-optimal solutions compared to integer linear programming-based algorithms. This work introduces two enhancements for SAT-based modulo scheduling: (i) an algorithm to rapidly calculate a lower bound for the schedule length, improving the identification of latency-optimal schedules; and (ii) a streamlined SAT formulation with fewer clauses, facilitating quicker solver decisions. Extensive experimental evaluations show that these improvements lead to an increased number of throughput-optimal and latency-optimal solutions. Nicolai Fiege, Peter Zipf |
FPL | 2 |
| 2025 | Fantastic Circuits and Where to Find Them - A Holistic ILP Formulation for Model-Based Hardware DesignabstractThe end of Moore’s law and Dennard scaling emphasizes the need for application-specific computing architectures to achieve high resource and energy efficiency and real-time performance. The concept of a silicon compiler remains an enduring aspiration for design time reduction. In order to generate hardware implementations at register transfer level from behavioral descriptions, design automation tools must address challenging and interdependent problems, including allocation, scheduling, and binding. Additionally, manual intervention by the user is necessary to balance the resources vs. performance tradeoff via, for example, function inlining or loop unrolling/pipelining. Existing approaches typically solve these problems sequentially, compromising optimality in favor of simplicity and runtime. Here we show how to model the whole model-based design flow as one holistic integer linear programming (ILP) formulation aiming at consistently deriving the optimal microarchitecture for any given application. Incorporating clock gating minimizes the number of useless operations with negligible resource overhead (if any), while always guaranteeing optimal throughput. The unified nature of the proposed ILP model enables implementations unmatched by state-of-the-art approaches in terms of resource efficiency and measured power consumption. These results facilitate a streamlined design flow for highly optimized embedded systems in the context of model-based design. Nicolai Fiege, Peter Zipf |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2024 | Bit-Level Optimized Constant Multiplication Using Boolean SatisfiabilityabstractMultiplierless constant multiplication using bit-shifts, additions and subtractions has been an active research topic in the last decades. The multiplication with multiple constants, known as the multiple constant multiplication (MCM) problem, is of special interest because of its practical relevance, notably for digital filter implementation. In this work we propose to use the speed of modern Boolean satisfiability (SAT) solvers to find fast and optimal solutions. The solutions are optimal either with respect to the adder count or the bit level cost. In contrast to previous approaches, we also consider negative fundamentals that are sometimes cheaper to realize than their positive counterparts leading to more compact hardware implementations. Our experiments show that our approach is able to find optimal single constant multiplication (SCM) and MCM circuits for practically relevant test instances in reasonable time. We also prove the necessity for the post-add right shift operation for SCM. Using our SAT formulation to enumerate all possible implementations for some of our test instances we show the importance of considering bit-level costs and negative fundamentals when solving MCM problems. Nicolai Fiege, Martin Kumm, Peter Zipf |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | More AddNet: A deeper insight into DNNs using FPGA-optimized multipliersabstractWe present a training tool flow for deep neural networks (DNN) optimized for a hardware-efficient FPGA-implementation based on reconfigurable constant-coefficient multipliers (RCCMs). RCCMs replace the costly generic multipliers by shift-and-add operations. In previous work, it was shown that RCCMs offer a better alternative for saving FPGA area than utilizing low-precision arithmetic. This work proposes an improved tool flow that enables layer-wise weight quantization, a larger search space by additional RCCM coefficient sets and an optimized retraining. This leads to an improved accuracy compared to the previous method. In addition, hardware requirements are lower as only 1 to 3 adders per multiplication are used. This reduces the overall complexity and the required memory bandwidth simultaneously. We evaluate our tool flow using multiple networks (ResNets) on the ImageNet data set. Martin Hardieck, Tobias Habermann, Fabian Wagner, Michael Mecik, Martin Kumm, Peter Zipf |
ISCAS | 6 |
| 2023 | BLOOP: Boolean Satisfiability-based Optimized Loop PipeliningabstractModulo scheduling is the premier technique for throughput maximization of loops in high-level synthesis by interleaving consecutive loop iterations. The number of clock cycles between data insertions is called the initiation interval (II). For throughput maximization, this value should be as low as possible; therefore, its minimization is the main optimization goal. Despite its long historical existence, modulo scheduling always remained a relevant research topic over the years with many exact and heuristic algorithms available in the literature. Nevertheless, we are able to leverage the scalability of modern Boolean Satisfiability (SAT) solvers to outperform state-of-the-art ILP-based algorithms for latency-optimal modulo scheduling for both integer and rational IIs. Our algorithm is able to compute valid modulo schedules for the whole CHStone and MachSuite benchmark suites, with 99% of the solutions being proven to be throughput optimal for a timeout of only 10 minutes per candidate II. For various time limits, not a single tested scheduler from the state of the art is able to compute more verified optimal solutions or even a single schedule with a higher throughput than our proposed approach. Using an HLS toolflow, we show that our algorithm can be effectively used to generate Pareto-optimal FPGA implementations regarding throughput and resource usage. Nicolai Fiege, Peter Zipf |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | Improving Energy Efficiency in Loop Pipelining by Rational-II Modulo SchedulingabstractModulo scheduling is a commonly used high-level synthesis (HLS) technique to maximize throughput by overlapping the computation of consecutive loop iterations [1] – [5] . For maximum throughput, the number of cycles to wait between successive sample insertions (called initiation interval, II) should be as low as possible. Nicolai Fiege, Patrick Sittel, Peter Zipf |
FCCM | 3 |
| 2022 | Optimal Binding and Port Assignment for Loop Pipelining in High-Level SynthesisabstractIn order to provide high throughput for custom hardware implementations, academic and commercial high-level synthesis (HLS) tools use loop pipelining by modulo scheduling. When provided a resource allocation and a schedule, the binding algorithm can be used to reduce the number of required lifetime registers (LR) and multiplexers (MUX). Contrary to non-modulo schedules, optimal solutions to the binding problem for implementing modulo schedules with respect to minimizing required LRs and MUXs have not been published. To address this topic, we propose a novel optimal binding algorithm to simultaneously minimize MUX and LR costs for loop pipelining using Integer Linear Programming. We evaluated our algorithm on a set of commonly used benchmark instances from digital signal processing and report that all encountered problems could be solved, with 36.53% of the solutions being optimal within a time limit of only five minutes. Compared to worst case evaluations, we report MUX and LR savings of up to 42.74% and 26.62%, respectively. To evaluate the impact on the resulting circuit after place and route, we studied FPGA implementations of several benchmark instances and recorded look-up table and flip-flop reductions of up to 13.70% and 5.24%, respectively, compared to previous work and to an extensive set of randomly generated bindings when state-of-the-art algorithms fail to find a feasible solution. Nicolai Fiege, Patrick Sittel, Peter Zipf |
FPL | 3 |
| 2022 | Speeding Up Optimal Modulo Scheduling with Rational Initiation IntervalsabstractCompared to integer initiation intervals (II), rational IIs improve throughput achieved by loop pipelining in many cases. This comes at the expense of a higher need for data path elements (i.e., multiplexers and registers) and the need for solving more complex scheduling problems. To optimally solve these problems, we improved an existing ILP formulation for latency-optimal modulo scheduling with rational IIs that now finds 6.08x more solutions and 6.10x as many optimal ones within the same time budget. Compared to the best alternative from previous work, our improved algorithm finds 1.15x more solutions and 2.97x as many optimal ones. Nicolai Fiege, Patrick Sittel, Peter Zipf |
FPL | 3 |
| 2022 | Optimal and Heuristic Approaches to Modulo Scheduling With Rational Initiation Intervals in Hardware SynthesisabstractA well-known approach for generating custom hardware with high throughput and low resource usage ismodulo scheduling, in which the number of clock cycles between successive inputs [the initiation interval (II)] can be lower than the latency of the computation. The II is traditionally aninteger, but in this article, we explore the benefits of allowing it to be arationalnumber. A rational II can be interpreted as theaveragenumber of clock cycles between successive inputs. Since the minimum rational II can be less than the minimum integer II, higher throughput is possible; moreover, allowing rational IIs gives more options in a design-space exploration. We formulate rational-II modulo scheduling as an integer linear programming (ILP) problem that is able to find latency-optimal schedules for a fixed rational II. We also propose two heuristic approaches that make rational-II scheduling more feasible: one based on identifying strongly connected components in the data-flow graph, and one based on iteratively relaxing the target II until a solution is found. We have applied our methods to a standard benchmark of hardware designs, and our results demonstrate an average speedup with respect to II of$1.24\times $in 35% of the encountered scheduling problems compared to state-of-the-art formulations. Patrick Sittel, Nicolai Fiege, John Wickerson, Peter Zipf |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Modulo Scheduling with Rational Initiation Intervals in Custom Hardware DesignabstractIn modulo scheduling, the number of clock cycles between successive inputs (the initiation interval, II) is traditionally an integer, but in this paper, we explore the benefits of allowing it to be a rational number. This rational II can be interpreted as the average number of clock cycles between successive inputs. As the minimum rational II can be less than the minimum integer II, this translates to higher throughput. We formulate rational-II modulo scheduling as an integer linear programming (ILP) problem that is able to find latency-optimal schedules for a fixed rational II. We have applied our scheduler to a standard benchmark of hardware designs, and our results demonstrate a significant speedup compared to state-of-the-art integer-II and rational-II formulations. Patrick Sittel, John Wickerson, Martin Kumm, Peter Zipf |
ASP-DAC | 4 |
| 2020 | AddNet: Deep Neural Networks Using FPGA-Optimized MultipliersabstractLow-precision arithmetic operations to accelerate deep-learning applications on field-programmable gate arrays (FPGAs) have been studied extensively, because they offer the potential to save silicon area or increase throughput. However, these benefits come at the cost of a decrease in accuracy. In this article, we demonstrate that reconfigurable constant coefficient multipliers (RCCMs) offer a better alternative for saving the silicon area than utilizing low-precision arithmetic. RCCMs multiply input values by a restricted choice of coefficients using only adders, subtractors, bit shifts, and multiplexers (MUXes), meaning that they can be heavily optimized for FPGAs. We propose a family of RCCMs tailored to FPGA logic elements to ensure their efficient utilization. To minimize information loss from quantization, we then develop novel training techniques that map the possible coefficient representations of the RCCMs to neural network weight parameter distributions. This enables the usage of the RCCMs in hardware, while maintaining high accuracy. We demonstrate the benefits of these techniques using AlexNet, ResNet-18, and ResNet-50 networks. The resulting implementations achieve up to 50% resource savings over traditional 8-bit quantized networks, translating to significant speedups and power savings. Our RCCM with the lowest resource requirements exceeds 6-bit fixed point accuracy, while all other implementations with RCCMs achieve at least similar accuracy to an 8-bit uniformly quantized design, while achieving significant resource savings. Julian Faraone, Martin Kumm, Martin Hardieck, Peter Zipf, Xueyuan Liu 0002, David Boland, Philip H. W. Leong |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | Reconfigurable Convolutional Kernels for Neural Networks on FPGAsabstractConvolutional neural networks (CNNs) gained great success in machine learning applications and much attention was paid to their acceleration on field programmable gate arrays (FPGAs). The most demanding computational complexity of CNNs is found in the convolutional layers, which account for 90% of the total operations. The fact that parameters in convolutional layers do not change over a long time interval in weight stationary CNNs allows the use of reconfiguration to reduce the resource requirements. This work proposes several alternative reconfiguration schemes that significantly reduce the complexity of sum-of-products operations. The proposed direct configuration schemes provide the least resource requirements and fast reconfiguration times of 32 clock cycles but require additional memory for the pre-computed configurations. The proposed online reconfiguration scheme uses an online computation of the LUT contents to avoid this memory overhead. Finally, a scheme that duplicates the reconfigurable LUTs is proposed for which the reconfiguration time can be completely hidden in the computation time. Combined with a few online reconfiguration circuits, this provides the same configuration memory and configuration time as a conventional parallel kernel but offers large resource reductions of up to 80% of the LUTs. Martin Hardieck, Martin Kumm, Konrad Möller, Peter Zipf |
FPGA | 4 |
| 2019 | Unrolling Ternary Neural NetworksabstractThe computational complexity of neural networks for large-scale or real-time applications necessitates hardware acceleration. Most approaches assume that the network architecture and parameters are unknown at design time, permitting usage in a large number of applications. This article demonstrates, for the case where the neural network architecture and ternary weight values are known a priori , that extremely high throughput implementations of neural network inference can be made by customising the datapath and routing to remove unnecessary computations and data movement. This approach is ideally suited to FPGA implementations as a specialized implementation of a trained network improves efficiency while still retaining generality with the reconfigurability of an FPGA. A VGG-style network with ternary weights and fixed point activations is implemented for the CIFAR10 dataset on Amazon’s AWS F1 instance. This article demonstrates how to remove 90% of the operations in convolutional layers by exploiting sparsity and compile-time optimizations. The implementation in hardware achieves 90.9 ± 0.1% accuracy and 122k frames per second, with a latency of only 29µs, which is the fastest CNN inference implementation reported so far on an FPGA. Stephen Tridgell, Martin Kumm, Martin Hardieck, David Boland, Duncan J. M. Moss, Peter Zipf, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2018 | Karatsuba with Rectangular Multipliers for FPGAsabstractThis work presents an extension of Karatsuba's method to efficiently use rectangular multipliers as a base for larger multipliers. The rectangular multipliers that motivate this work are the embedded 18 × 25-bit signed multipliers found in the DSP blocks of recent Xilinx FPGAs: The traditional Karatsuba approach must under-use them as square 18 × 18 ones. This work shows that rectangular multipliers can be efficiently exploited in a modified Karatsuba method if their input word sizes have a large greatest common divider. In the Xilinx FPG A case, this can be obtained by using the embedded multipliers as 16 × 24 unsigned and as 17 × 25 signed ones. The obtained architectures are implemented with due detail to architectural features such as the pre-adders and post-adders available in Xilinx DSP blocks. They are synthesized and compared with traditional Karatsuba, but also with (non-Karatsuba) state-of-the-art tiling techniques that make use of the full rectangular multipliers. The proposed technique improves resource consumption and performance for multipliers of numbers larger than 64 bits. Martin Kumm, Oscar Gustafsson, Florent de Dinechin, Johannes Kappauf, Peter Zipf |
ARITH | 5 |
| 2018 | ILP-Based Modulo Scheduling and Binding for Register MinimizationabstractA key element for achieving high throughput, e.g. circuits generated with high-level synthesis (HLS) methods and model-based hardware design, is the use of modulo scheduling. Integer linear programming (ILP)-based modulo schedulers are capable of computing schedules that are optimal regarding throughput and latency, while keeping run times to practically usable lengths. However, the generated schedules may lead to an excessive number of registers for storing intermediate values. We propose extensions for ILP-based modulo scheduling that minimizes these registers. The ILP formulation incorporates the elimination of redundant registers by post binding optimization. Extensive experiments on different benchmark sets show average register reductions of 30.4% compared to commonly used minimum lifetime approaches that reduce register requirements. This comes without any loss in throughput or latency and with less than 4% additional scheduling run time compared to state-of-the-art ILP-based modulo schedulers. Patrick Sittel, Martin Kumm, Julian Oppermann, Konrad Möller, Peter Zipf, Andreas Koch 0001 |
FPL | 5 |
| 2018 | Optimal Shift Reassignment in Reconfigurable Constant Multiplication CircuitsabstractThis paper presents a new method called optimal shift reassignment (OSR), used for reconfigurable multiplication circuits. These circuits consist of adders, subtractors, shifts, and multiplexers (MUXs). They calculate the multiplication of an input number by one out of several constants which can be selected dynamically during run-time. The OSR method is based on the idea that shifts can be placed at different positions along the circuit, while the calculated output constant stays the same. This differs from previous approaches, which were limited by the fact that all constants within the constant multiplier were forced to be odd. The OSR method subsequently releases this restriction. As a result, the number of required MUXs in the circuit can be reduced. This happens when the shift reassignment aligns the shift values of different inputs of an MUX. Experimental results show MUX savings of up to 50% and average savings between 11% and 16% using the OSR method compared to previous approaches. Konrad Möller, Martin Kumm, Mario Garrido, Peter Zipf |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Resource Optimal Design of Large Multipliers for FPGAsabstractThis work presents a resource optimal approach for the design of large multipliers for FPGAs. These are composed of smaller multipliers which can be DSP blocks or logic-based multipliers. A previously proposed multiplier tiling methodology is used to describe feasible solutions of the problem. The problem is then formulated as an integer linear programming (ILP) problem which can be solved by standard ILP solvers. It can be used to minimize the total implementation cost or to trade the LUT cost against the DSP cost. It is demonstrated that although the problem is NP-complete, optimal solutions can be found for most practical multiplier sizes up to 64x64. Synthesis experiments on relevant multiplier sizes show slice reductions of up to 47.5% compared to state-of-the-art heuristic approaches. Martin Kumm, Johannes Kappauf, Matei Istoan, Peter Zipf |
ARITH | 4 |
| 2017 | Model-based hardware design based on compatible sets of isomorphic subgraphsabstractHardware applications in an industrial context often have tight area, latency and throughput requirements or a specific combination thereof. This paper presents a method to improve area and throughput figures for folded circuits generated during a model-based hardware design process. The method targets FPGA implementations and is based on the automatic combination of isomorphic subgraphs and the detailed consideration of pipelined primitive operations for folding core scheduling. In the course of a design space exploration, the user is provided with fine-grain control over the area/throughput trade-off. Patrick Sittel, Konrad Möller, Martin Kumm, Peter Zipf, Bogdan Pasca 0001, Mark Jervis |
FPT | 4 |
| 2017 | Optimization of Constant Matrix Multiplication with Low Power and High ThroughputabstractConstant matrix multiplication (CMM), i.e., the multiplication of a constant matrix with a vector, is a common operation in digital signal processing. It is a generalization of multiple constant multiplication (MCM) where a single variable is multiplied by a constant vector. Like MCM, CMM can be reduced to additions/subtractions and bit shifts. Finding a circuit with minimal number of add/subtract operations is known as the CMM problem. While this leads to a reduction in circuit area it may be less efficient for power consumption or throughput. It is well studied for the MCM problem that a) reducing the adder depth (AD) leads to a reduced power consumption and b) pipeline resources have to be considered during optimization to enhance throughput without wasting area. This paper addresses the optimization of CMM circuits which considers both adder depth and pipelining for the first time. For that, a heuristic is proposed which evaluates the most attractive graph topologies. It is shown that the proposed method requires 12.5% less adders with min. AD and 38.5% less pipelined operations. Synthesis results for recent FPGAs show that these reductions also translate to superior results in terms of delay and power consumption compared to the state-of-the-art. Martin Kumm, Martin Hardieck, Peter Zipf |
IEEE Trans. Computers | 3 |
| 2017 | Reconfigurable Constant Multiplication for FPGAsabstractThis paper introduces a new heuristic to generate pipelined run-time reconfigurable constant multipliers for field-programmable gate arrays (FPGAs). It produces results close to the optimum. It is based on an optimal algorithm which fuses already optimized pipelined constant multipliers generated by an existing heuristic called reduced pipelined adder graph (RPAG). Switching between different single or multiple constant outputs is realized by the insertion of multiplexers. The heuristic searches for a solution that results in minimal multiplexer overhead. Using the proposed heuristic reduces the run-time of the fusion process, which raises the usability and application domain of the proposed method of run-time reconfiguration. An extensive evaluation of the proposed method confirms a 9%-26% FPGA resource reduction on average compared to previous work. For reconfigurable multiple constant multiplication, resource savings of up to 75% can be shown compared to a standard generic lookup table based multiplier. Two low level optimizations are presented, which further reduce resource consumption and are included into an automatic VHDL code generation based on the FloPoCo library. Konrad Möller, Martin Kumm, Marco Kleinlein, Peter Zipf |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Efficient sum of absolute difference computation on FPGAsabstractAn improved architecture for efficiently computing the sum of absolute differences (SAD) on FPGAs is proposed in this work. It is based on a configurable adder/subtractor implementation in which each adder input can be negated at runtime. The negation of both inputs at the same time is explicitly allowed and used to compute the sum of absolute values in a single adder stage. The architecture can be mapped to modern FPGAs from Xilinx and Altera. An analytic complexity model as well as synthesis experiments yield an average look-up table (LUT) reduction of 17.4% for an input word size of 8 bit compared to state-of-the-art. As the SAD computation is a resource demanding part in image processing applications, the proposed circuit can be used to replace the SAD core of many applications to enhance their efficiency. Martin Kumm, Marco Kleinlein, Peter Zipf |
FPL | 3 |
| 2015 | An Efficient Softcore Multiplier Architecture for Xilinx FPGAsabstractThis work presents an efficient implementation of a softcore multiplier, i.e., a multiplier architecture which can be efficiently mapped to the slice resources of modern Xilinx FPGAs. Instead of dividing the multiplication into the generation of partial products and the summation using a compressor tree, as done in modern multipliers, an array-like architecture is proposed. Each row of the array generates a partial product which is directly added to results of previous rows using the fast carry chain. A radix-4 Booth encoding/decoding is used to reduce the I/O count of the partial product generation which makes it possible to map both, the Booth encoder and decoder, into a single 6-input look up table (LUT). Like a conventional Booth multiplier, this nearly halves the number of rows compared to a ripple carry array multiplier. In addition, the compressor tree is completely avoided and an efficient and regular structure retains that uses up to 50% less slice resources compared to previous approaches and offers a multiply accumulate (MAC) operation without extra resources. Martin Kumm, Shahid Abbas, Peter Zipf |
ARITH | 3 |
| 2014 | Pipelined compressor tree optimization using integer linear programmingabstractCompressor trees offer an effective realization of the multiple input addition needed by many arithmetic operations. However, mapping the commonly used carry save adders (CSA) of classical compressor trees to FPGAs suffers from a poor resource utilization. This can be enhanced by using generalized performance counters (GPCs). Prior work has shown that high efficient GPCs can be constructed by exploiting the low-level structure of the FPGA. However, due to their irregular shape, the selection of those is not straight forward. Furthermore, the compressor tree has to be pipelined to achieve the potential FPGA performance. Then, a selection between registered GPCs or flip-flops has to be done to balance the pipeline. This work defines the pipelined compressor tree synthesis as an optimization problem and proposes a (resource) optimal method using integer linear programming (ILP). Besides that, two new GPC mappings with high efficiency are proposed for Xilinx FPGAs. Martin Kumm, Peter Zipf |
FPL | 2 |
| 2014 | An FPGA-optimized architecture of horn and schunck optical flow algorithm for real-time applicationsabstractOptical flow estimation of image sequences is one of the key elements for motion detection. However, processing the optical flow in real-time is still an open task due to its computationally expensive nature. In this paper we present an FPGA-optimized architecture for optical flow estimation based on the algorithm of Horn and Schunck. While existing FPGA-realizations are only partly real-time capable, on a Stratix IV our architecture enables the computation of the optical flow for each pixel of a frame with 640 × 512 pixels at a framerate of 30 fps in iterative and up to 4k resolution (4,096 × 2,304 pixels) at a framerate of 20 fps in full-pipelined form. Michael Kunz, Alexander Ostrowski, Peter Zipf |
FPL | 3 |
| 2014 | Pipelined reconfigurable multiplication with constants on FPGAsabstractThis paper presents a new algorithm to automatically create pipelined run-time reconfigurable constant multipliers. Reconfiguration between several constants is achieved by merging optimized pipelined adder graphs using multiplexers. The adder graphs perform the required multiplications by using additions, subtractions and bit-shifts only. They are generated by an existing heuristic called RPAG. The resulting reconfigurable pipelined single and multiple constant multipliers can be used for time-multiplexed multiplication reducing the required FPGA logic resources. In contrast to earlier approaches aiming at application-specific integrated circuits (ASICs) we introduce pipelining and take special care of the size of the added multiplexers to obtain FPGA-optimized solutions. We can show that the achieved pipelined run-time reconfigurable constant multipliers on average only need about 77% of the slices compared to the best solutions based on the merging of adder graphs published so far. Konrad Möller, Martin Kumm, Marco Kleinlein, Peter Zipf |
FPL | 4 |
| 2013 | Multiple constant multiplication with ternary addersabstractThe scaling operation, i. e., the multiplication with a single constant is a frequently used operation in many kinds of numeric algorithms. The multiple constant multiplication (MCM) is a generalization where a variable is multiplied by several constants. This kind of operation is heavily used, e. g., in digital filters or discrete transforms. It was shown in recent work that small, fast and power efficient MCM implementations can be realized by using the fast carry chains of FPGAs rather than wasting specialized embedded multipliers. However, in the work so far, only common two-input adders were used. As FPGAs today support ternary adders, i. e., adders with three inputs, this work investigates the optimization of pipelined MCM circuits which include ternary adders. It is shown experimentally that 27% less operations are needed on average by using ternary adders, resulting in 15% slice (Xilinx) and 10% ALM (Altera) reductions, respectively. Martin Kumm, Martin Hardieck, Jens Willkomm, Peter Zipf, Uwe Meyer-Bäse |
FPL | 4 |
| 2013 | Partial LUT size analysis in distributed arithmetic FIR Filters on FPGAsabstractDistributed arithmetic is a popular method for implementing digital FIR filters on FPGAs. One essential optimization method is the division of large look-up tables (LUTs) into smaller partial LUTs by using additional adders. Previous work indicates, that the size of these partial LUTs should be chosen to the LUT input size of the FPGA which was 4 for a long time. Nowadays, modern FPGAs offer 6-input LUTs which can be configured to two 5-input LUTs with shared inputs. This paper investigates the optimal input size of partial LUTs on FPGAs with 4-input and 5/6-input LUTs. On FPGAs with 4-input LUTs, it turnes out that only in 62% of the cases (out of 220), a LUT input size of 4 leads to the best implementation. However, the slice overhead is 6.3% on average for the other cases. On FPGAs with 5/6-input LUTs, the least slice overhead (10% on average) is paid when the LUT input size is chosen to 6. However, it was shown that a resource reduction of up to 32% can be achieved when all input sizes in the range 4...7 are evaluated. Using the best partial LUT size, slice reductions of over 50% on average compared to Xilinx Coregen could be achieved for Virtex 6 FPGAs. Martin Kumm, Konrad Möller, Peter Zipf |
ISCAS | 3 |
| 2013 | Reconfigurable FIR filter using distributed arithmetic on FPGAsabstractAn architecture for a dynamically run-time reconfigurable finite impulse response (FIR) filter is presented in this work. It is based on distributed arithmetic (DA) combined with a look-up table (LUT) reduction technique which allows the direct mapping to reconfigurable LUTs (CFGLUT) of the latest Xilinx FPGAs. The resulting FIR filter can be reconfigured with arbitrary coefficients which are only limited by their length and word size. The number of filter instances for reconfiguration is only limited by the block memory of the FPGA which typically allows hundreds of different configurations. The proposed reconfigurable architecture consumes 16% less slices on average than a fixed coefficient DA filter generated by Xilinx Coregen. As the direct mapping to CFGLUTs leads to invalid filter output during reconfiguration, an alternative architecture is proposed which avoids this limitation at the cost of 19% more slice resources on average. Using a parallel reconfiguration scheme, reconfiguration times of about 100ns could be achieved. Martin Kumm, Konrad Möller, Peter Zipf |
ISCAS | 3 |
| 2012 | Reduced complexity single and multiple constant multiplication in floating point precisionabstractThis paper addresses the automatic generation and optimization of single and multiple constant multipliers in IEEE 754 floating point precision for FPGAs. It is shown that sharing of partial results in multiplication and exponent addition as well as the handling of special input values can greatly reduce the overall hardware complexity. Two methods are used to reduce the complexity of the integer multiplier block: An existing method using adder arithmetic only and a novel optimization method using a reduced amount of embedded multipliers to compute products with large coefficient values. Superior results are shown compared to previous methods. Using adder arithmetic, a slice reduction of 18% for single and 38% for multiple constant floating point multiplication could be achieved. Using embedded multipliers, nearly half of the multipliers could be saved for multiple constants compared to the conventional approach. Martin Kumm, Katharina Liebisch, Peter Zipf |
FPL | 3 |
| 2012 | Area estimation of look-up table based fixed-point computations on the example of a real-time high dynamic range imaging systemabstractIn many FPGA-based designs, fixed-point computations can be efficiently implemented by look-up tables. However, the precision of these computations has a great influence on the hardware costs. We present two simple models for fast but precise area estimation without synthesis, that can be used for word length optimization. As an example we analyze a lookup table based implementation of a circuit for processing image combination and tone-mapping of high dynamic range (HDR) images. The results can be generalized and be applied to any problem with soft accuracy requirements. Michael Kunz, Martin Kumm, Martin Heide, Peter Zipf |
FPL | 4 |
| 2012 | Pipelined adder graph optimization for high speed multiple constant multiplicationabstractThis paper addresses the direct optimization of pipelined adder graphs (PAGs) for high speed multiple constant multiplication (MCM). The optimization opportunities are described and a definition of the pipelined multiple constant multiplication (PMCM) problem is given. It is shown that the PMCM problem is a generalization of the MCM problem with limited adder depth (AD). A novel algorithm to solve the PMCM problem heuristically, called RPAG, is presented. RPAG outperforms previous methods which are based on pipelining the solutions of conventional MCM algorithms. A flexible cost evaluation is used which enables the optimization for FPGA or ASIC targets on high or low abstraction levels. Results for both technologies are given and compared with the most recent methods. Even for the special case of limited AD it is shown that RPAG often produces better results compared to the prominent Hcubalgorithm with minimal total AD constraint. Martin Kumm, Peter Zipf, Mathias Faust, Chip-Hong Chang |
ISCAS | 2 |
| 2011 | High speed low complexity FPGA-based FIR filters using pipelined adder graphsabstractA method for generating high speed FIR filters with low complexity for FPGAs is presented. The realization is split into two parts. First, an adder graph is obtained using an existing multiple constant multiplication (MCM) algorithm. This adder graph describes the required multiplier block of the FIR filter using only additions/subtractions and shifts. Secondly, a novel FPGA-specific combined schedule and pipeline optimization is performed to gain the maximum speed while using a minimal performance penalty. FPGA-specific characteristics are exploited during optimization including the reduction of pipeline registers by duplicating adders in later stages. The optimization is formulated as binary integer linear programming (BILP) problem. It is shown that the generated number of pipelined operations based on the HcubMCM algorithm is reduced up to 29.1% on average compared to an as-soon-as-possible (ASAP) scheduling using cut-set retiming. Synthesis results are obtained by generating VHDL code, showing that the proposed method outperforms the recently proposed Add/Shift method in resource complexity (54.1% reduction on average) while a competitive performance is achieved (88.2% speed of Add/Shift on average). Martin Kumm, Peter Zipf |
FPT | 2 |
| 2009 | Design and evaluation of an energy-efficient dynamically reconfigurable architecture for wireless sensor nodesabstractWe explore the design of a coarse-grained reconfigurable architecture for wireless sensor network nodes, which combines high energy efficiency with programmability and hence meets the requirements of small energy-constraint embedded systems. Its energy consumption, area, and performance are evaluated and compared to processor and ASIC architectures. Our case study particularly focuses on the question if the architecture concept of frequent dynamic reconfiguration of a small heterogeneous data path can lead to suitable system solutions for the target domain. To answer this, the effect of the reconfiguration overhead on total system efficiency is examined closely. As important result, our experiments show the low energy consumption achieved, the low reconfiguration overhead, and the specific region of the architecture in the design space between processors and ASICs. In particular, large energy-savings of factor 2 to 6 and speed-ups of factor 6 to 14 compared to processors are obtained on average. Our work shows the high suitability of frequent runtime reconfiguration of small coarse-grain data paths for the design of very efficient but yet programmable embedded systems platforms. Heiko Hinkelmann, Peter Zipf, Manfred Glesner |
FPL | 2 |
| 2009 | An integrated tool flow to realize runtime-reconfigurable applications on a new class of partial multi-context FPGAsabstractEfficient cycle-based reconfiguration of datapaths can be realized on current FPGAs by designing merged datapaths, which can execute different tasks depending on the datapath control. In our previous work, we provided a synthesis tool for the automated generation of such datapaths. The objective in this paper is a reduction of the resource requirements for implementing the reconfiguration control on the FPGA. First, we extend our existing tool flow with a novel mechanism for efficient partial multi-context reconfiguration and provide tool support to generate according circuitry, control, and configuration data. Second, we propose an extension of current FPGA architectures to efficiently support cycle-based runtime control. We employ a special content-addressable multi-context memory resource for controlling the datapath functionality, which on average requires only 15% memory for storing the control data compared to a BlockRAM based approach, and only 53% compared to regular multi-context switching. Markus Rullmann, Renate Merker, Heiko Hinkelmann, Peter Zipf, Manfred Glesner |
FPL | 4 |
| 2009 | Towards a unique FPGA-based identification circuit using process variationsabstractA compact chip identification (ID) circuit with improved reliability is presented. Ring oscillators are used to measure the spatial process variation and the ID is based on their relative speeds. A novel averaging and postprocessing scheme is employed to accurately determine the faster of two similar-frequency ring oscillators in the presence of noise. Using this scheme, the average number of unstable bits i.e. bits which can change in value between readings, measured on an FPGA is shown to be reduced from 5.3% to 0.9% at 20degC. Within the range 20-60degC, the percentage of unstable bits is within 2.8%. An analysis of the effectiveness of the scheme and the distribution of the errors is given over different temperature ranges and FPGA chips. Haile Yu, Philip H. W. Leong, Heiko Hinkelmann, Leandro Möller, Manfred Glesner, Peter Zipf |
FPL | 6 |
| 2009 | Generation of Synthetic Floating-Point benchmark circuitsabstractSynthetic Floating-Point (SFP), a synthetic benchmark generator program for floating-point circuits is presented. SFP consists of two independent modules for characterisation and generation. The characterisation module extracts key dataflow statistics of an arbitrary software program. Generation involves producing randomised circuits with desired statistics which are either the output of the characterisation module or directly generated by the user. Using the basic linear algebra subprograms (BLAS) library, Whetstone benchmark and LINPACK benchmark, it is demonstrated that SFP can be used to generate floating-point benchmarks with different user-specified properties as well as benchmarks that mimic real computational programs. Thomas C. P. Chau, S. Man Ho Ho, Philip H. W. Leong, Peter Zipf, Manfred Glesner |
IPDPS | 4 |
| 2008 | Coarse-grained reconfigurationabstractIn the last years, aside from fine-grained reconfigurable architectures such as FPGAs, coarse-grained reconfigurable architectures (CGRAs), which typically have building blocks of a fixed bit-width (8 bit, 16 bit, etc.), have gained in importance in academia as well as in industry. CGRAs are usually used for domain-specific computations and have advantages over traditional FPGAs in terms of area and power cost, performance, and reconfiguration time. Thus, architectures with coarse-grained reconfiguration features have also been studied in projects (Sec. 1, 2, 4) within the priority program Reconfigurable Computing Systems and the project CoMap (Sec. 3), which are all sponsored by the German science foundation. Sven Eisenhardt, Thomas Schweizer, Julio de Oliveira Filho, Tobias Oppold, Wolfgang Rosenstiel, Alexander Thomas, Jürgen Becker 0001, Frank Hannig, Dmitrij Kissler, Hritam Dutta, Jürgen Teich, Heiko Hinkelmann, Peter Zipf, Manfred Glesner |
FPL | 13 |
| 2008 | Application-specific reconfigurable processorsabstractApplication-specific reconfigurable processor architectures provide a remarkable potential for systems which achieve concurrently high performance, area efficiency, energy efficiency, run-time adaptivity, and sufficient flexibility. Thus, they represent competitive design alternatives that provide significant improvements in some of these figures of merit in comparison to non-reconfigurable architectures. Research results of three projects on the analysis of architecture concepts, design, and evaluation of application-specific reconfigurable processors in different domains are presented. Heiko Hinkelmann, Peter Zipf, Manfred Glesner, Matthias Alles, Timo Vogt, Norbert Wehn, Götz Kappen, Tobias G. Noll |
FPL | 2 |
| 2008 | A scalable reconfiguration mechanism for fast dynamic reconfigurationabstractHardware reconfiguration during run-time provides attractive features like fast adaptivity, high hardware utilisation, and low area consumption due to efficient reuse of hardware components. In this paper, a novel multi-layered reconfiguration mechanism is proposed that allows frequent dynamic reconfiguration at very low latencies. It combines successful existing techniques such as multi-context and partial reconfiguration with new ideas like tag-matching and reconfiguration profiles to one uniform approach. As an important feature, the proposed reconfiguration mechanism is well scalable and can be adapted to given hardware structures easily, thus being applicable to virtually any reconfigurable fabric. In contrast to many existing techniques, it also supports even very heterogeneous architectures found for instance in custom reconfigurable systems. By experimental results, we show that our reconfiguration mechanism provides significantly lower reconfiguration latencies compared to some common existing techniques. Heiko Hinkelmann, Peter Zipf, Manfred Glesner |
FPT | 2 |
| 2008 | An area-efficient FPGA realisation of a codebook-based image compression methodabstractWe present a hardware implementation of an efficient image compression method optimised for small FPGAs. The compression method is based on a codebook of reference patterns to support multiplication-free quantisation of the image data. Based on specific features of a low-cost FPGA architecture, a pipelined implementation is developed and evaluated. The implemented hardware benefits from the simple structure of the compression method and is optimised for area and performance. The realised hardware as well as the underlying compression mechanism are described and the synthesis results for different model variants are compared. The results show that a high compression rate is possible at extremely low hardware costs. Also, a high frame rate can be obtained even on a low-cost FPGA. Peter Zipf, Heiko Hinkelmann, Radu Dogaru, Manfred Glesner |
FPT | 1 |
| 2008 | Applying Dynamic Reconfiguration for Fault Tolerance in Fine-Grained Logic ArraysabstractThis paper presents the realization of a fault tolerance technique for a dynamically reconfigurable array of programmable cells. The three parts of the technique, fault detection, fault reconfiguration, and fault recovery, are implemented completely in hardware and form a self-contained system. Each of the parts can be exchanged by an alternative implementation without affecting the remaining parts too much, thus making the concept adaptable to different reconfigurable circuits. A hardware realization for the core mechanism is discussed and a prototypical design of a field-programmable gate array implementing the complete system is described. The technological development towards nanoscale feature sizes and the growing influence of deep-submicrometer effects will result in an inherent unreliability of the individual components of future circuit implementations and a higher vulnerability towards external influences. The technique discussed can be used to exploit dynamic reconfiguration capabilities of programmable arrays to alleviate system vulnerability towards these effects and thus to enhance their overall reliability. Peter Zipf |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2007 | A Power Estimation Model for an FPGA-based Softcore ProcessorabstractWe describe the application of a hybrid functional level power analysis (FLPA) and instruction level power analysis (ILPA) approach to a processor model implemented on an FPGA. This technique enables the estimation of the task specific power consumption of the modeled processor, in our case a LEON2, very early during a system design flow, based on the software which will run on it. The FLPA/ILPA model used during our work as well as the test scenarios and the measured results are described. Later, the function block separation and the power consumption modeling are discussed. Finally, the model is validated by benchmarking. The obtained model is promising in the sense that a) its estimations are close (4% on average) to the measured data, and b) the model structure is similar to that of hardcore processors which is not a trivial result. Peter Zipf, Heiko Hinkelmann, Manfred Glesner, Holger Blume, Tobias G. Noll |
FPL | 1 |
| 2007 | A Domain-Specific Dynamically Reconfigurable Hardware Platform for Wireless Sensor NetworksabstractIn this paper, a new generic sensor node platform for wireless sensor networks (WSN) is presented, demonstrating that high energy efficiency, flexibility and performance can be achieved by the use of dynamically reconfigurable hardware in the WSN domain. The core of the presented platform is formed by the combination of a RISC processor and a dynamically reconfigurable function unit optimized for efficient data processing in WSN applications. A novel reconfiguration mechanism is applied enabling rapid dynamic reconfiguration with very short latencies. Thereby, we can show that the overhead on performance and energy consumption caused by dynamic reconfiguration can be reduced to a moderate, non-critical value and is clearly outweighed by the significantly improved performance and energy consumption for data processing on the reconfigurable function unit. The evaluation of the platform and its comparison to a standard processor-based platform finally demonstrate high gains in energy-efficiency of one to two orders of magnitude. Heiko Hinkelmann, Peter Zipf, Manfred Glesner |
FPT | 2 |
| 2006 | A signal theory based approach to the statistical analysis of combinatorial nanoelectronic circuitsabstractIn this paper we present a method which allows the statistical analysis of nanoelectronic Boolean networks with respect to timing uncertainty and noise. All signals are considered to be instationary random processes which is the most general signal representation. As one cannot deal with random processes per se, we focus on certain statistical properties which are propagated through networks of Boolean gates yielding the instationary probability density function (pdf) of each signal in the network. Finally, several values of interest as the error probability, the average path delay or the average signal trace over time can be extracted from these pdf Oliver Soffke, Peter Zipf, Tudor Murgan, Manfred Glesner |
DATE | 2 |
| 2006 | Multitasking Support for Dynamically Reconfig Urable SystemsabstractReconfigurable hardware is often used in an attempt to boost performance of an embedded system while minimizing the cost penalty of additional hardware. The key method applied to achieve this goal is the re-use of hardware resources for different tasks. However, reconfigurable systems tend to have a large internal state, which complicates rapid task switching, as it is required for multitasking in real-time systems, significantly. We present a control technique enabling a fast and efficient preemption of dynamically reconfigurable systems which is largely performed by hardware. The technique handles the problem of internal state preservation and reduces the preemption effort by applying the concept of preemption points. The implementation of the system control and an evaluation of the performance is given, showing the high efficiency of the concept. Heiko Hinkelmann, Andreas Gunberg, Peter Zipf, Leandro Soares Indrusiak, Manfred Glesner |
FPL | 3 |
| 2005 | CONAN - A Design Exploration Framework for Reliable Nano-ElectronicsabstractIn this paper we introduce a design methodology that allows the system/circuit designer to build reliable systems out of unreliable nano-scale components. The central point of our approach is a generic (parametrical) architectural template. Configurable nanostructures for reliable nano electronics (CONAN), which embeds support for reliability at various levels of abstractions. Some of the main reliability sources are regular and decentralized structures based on simple basic computation cells designed to be robust against disturbances and noise, fault tolerance based on hardware, time and information redundancy applied at the basic cell level as well as at higher levels, self diagnosis assisted by the dynamic reconfiguration of basic computation cells and interconnect rerouting. Within the CONAN template, both technology dependent and independent models co-exists such that the more abstract layers are technology independent while the lower levels can be retargeted to various fabrication technologies. Our proposal is application-oriented and allows the designers to deal with unpredictability, and low reliability, which are unavoidable characteristics of future emerging nano-devices. When combined with the underlying software, the tools supporting the CONAN approach allow the designer to check whether the design constraints are fulfilled before performing a detailed implementation and provides means to trade area, delay, and power consumptions for reliability. As such, this proposal is a call-to-arms to mobilize the efforts of systems designers in order to achieve a systematic design methodology for reliable systems. Sorin Cotofana, Alexandre Schmid, Yusuf Leblebici, Adrian M. Ionescu, Oliver Soffke, Peter Zipf, Manfred Glesner, Antonio Rubio 0001 |
ASAP | 6 |
| 2005 | Programmable and Reconfigurable Hardware Architectures for the Rapid Prototyping of Cellular AutomataabstractIn this paper we describe and compare several architectures of cellular automata to be used as hardware accelerators in the evaluation loop of a genetic algorithm. In addition, two dynamically reconfigurable cell interconnection networks are presented capable to realize nonregular lattices. Cellular automata are basic computational structures of interacting units which may expose self-organization and emergent behaviour. To investigate this behaviour subject to different cell interconnect patterns, an automated inspection flow is needed, including a very fast evaluation of single specimen. For that purpose, a flexible FPGA-based accelerator for cellular automata evaluation is used, which can be accessed transparently by a Java client running the genetic algorithm. Several architectures have been developed for that: a straightforward implementation of the cellular automaton, an area-reduced architecture, a dynamically reconfigurable interconnection network, which allows the interconnection of single cells under certain constraints and, finally, a dynamically reconfigurable interconnection network, which allows to connect cells arbitrarily. Peter Zipf, Oliver Soffke, Andre Schumacher, Radu Dogaru, Manfred Glesner |
FPL | 1 |
| 2005 | A Hardware-in-the-Loop System to Evaluate the Performance of Small-World Cellular AutomataabstractThis paper presents the realisation of a hardware-in-the-loop system to investigate the performance of different cellular automata (CA) structures. The system is applied to regular lattice CAs and to small-world CAs, which are expected to expose better characteristics than lattice automata due to their nature-inspired structure. CA functionality is evolved using a genetic algorithm (GA) implemented as a distributed Java program running on a host computer. The performance evaluation of whole generations of individual automata is transferred to a specialised hardware architecture on an FPGA-board in order to speed up this process. For this, a customisable version of an automaton is residing on the board and personalisation data can be downloaded to it. The objective of the approach is to gather qualitative and quantitative data on the differences between the two types of CAs. We discuss two CA implementations, one of a lattice CA and one of a small-world CA. Their properties are characterised and their integration into the overall evaluation system is described. Peter Zipf, Oliver Soffke, Andre Schumacher, Clemens Schlachta, Radu Dogaru, Manfred Glesner |
FPL | 1 |
| 2004 | IP Generation for an FPGA-Based Audio DAC Sigma-Delta Converter
Ralf Ludewig, Oliver Soffke, Peter Zipf, Manfred Glesner, Kong-Pang Pun, Kuen Hung Tsoi, Kin-Hong Lee, Philip H. W. Leong |
FPL | 3 |
| 2004 | The XPP Architecture and Its Co-simulation Within the Simulink Environment
Mihail Petrov, Tudor Murgan, Frank May, Martin Vorbach, Peter Zipf, Manfred Glesner |
FPL | 5 |
| 2003 | A granularity-based classification model for systems-on-a-chipabstractField-programmable logic has become an increasingly important technology for the design of digital circuits. One interesting point in the field of reconfigurable logic is its classification within the implementation space of other technologies. Such a classification gains importance if FPGA technology becomes an integral part of Systems-on-a-Chip (SoC). The poster discusses an approach to classify technologies based on their granularity. Therefore, a new distinction into homogeneous and heterogeneous granularity is proposed. This leads to a theoretic classification method for the combination of two or more technologies. Using this method the result of a combination of two or more technologies can be determined qualitatively. The poster further shows opportunities to improve standard CAD-software based on the model proposed. Stephan Bingemer, Peter Zipf, Manfred Glesner |
FPGA | 2 |
| 2003 | Evaluation and Run-Time Optimization of On-chip Communication Structures in Reconfigurable Architectures
Tudor Murgan, Mihail Petrov, Alberto García Ortiz, Ralf Ludewig, Peter Zipf, Thomas Hollstein, Manfred Glesner, Bernard Ölkrug, Jörg Brakensiek |
FPL | 5 |
| 2003 | Arbitrary function approximation in HDLs with application to the N-body problemabstractA module generator is described that allows for the generation of synthesizable VHDL modules which implement arbitrary functions in fixed point precision using the Symmetric Table Addition Method (STAM). This module generator was interfaced to a high level synthesis tool "fly" which automatically generates fully-pipelined circuits from a Perl-like language. The resulting system was applied to the N-body problem and results are presented. It was found that a function generator module is a very useful addition to a hardware description language. Chun Hok Ho, Kuen Hung Tsoi, Jackson H. C. Yeung, Yuet Ming Lam, Kin-Hong Lee, Philip H. W. Leong, Ralf Ludewig, Peter Zipf, Alberto García Ortiz, Manfred Glesner |
FPT | 8 |
| 2003 | An Integrated Model Bridging the Gap between Technology and Economy
Stephan Bingemer, Peter Zipf, Manfred Glesner |
VLSI-SOC | 2 |
| 2003 | A hierarchical generic approach for on-chip communication, testing and debugging of SoCs
Thomas Hollstein, Ralf Ludewig, Christoph Mager, Peter Zipf, Manfred Glesner |
VLSI-SOC | 4 |
| 2003 | An Adaptive Trace-Back Solution for State-Parallel Viterbi Decoders
Mihail Petrov, Abdulfattah Mohammad Obeid, Tudor Murgan, Peter Zipf, Jörg Brakensiek, Bernard Ölkrug, Manfred Glesner |
VLSI-SOC | 4 |
| 2002 | Fly - A Modifiable Hardware Compiler
Chun Hok Ho, Philip H. W. Leong, Kuen Hung Tsoi, Ralf Ludewig, Peter Zipf, Alberto García Ortiz, Manfred Glesner |
FPL | 5 |
| 2002 | A Framework for Teaching (Re)Configurable Architectures in Student Projects
Thilo Pionteck, Peter Zipf, Lukusa D. Kabulepa, Manfred Glesner |
FPL | 2 |
| 2002 | Handling FPGA Faults and Configuration Sequencing Using a Hardware Extension
Peter Zipf, Manfred Glesner, Christine Bauer 0002, Hans Wojtkowiak |
FPL | 1 |