VLDB 2026 Research / reviewers in the wild / expert
Oskar Mencer
dblp:08/325
· DBLP profile ↗
55ranked-venue papers
9as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 51 · 8 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
Electronic design automation · 34% Reconfigurable computing and FPGAs · 29% High-performance computing · 15% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 22 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
high-level synthesis |
0.6 | 4 | 2020 | Performance Portable FPGA Design · FPGA 2020 CHIPS: Custom Hardware Instruction Processor Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 Accuracy-Guaranteed Bit-Width Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Electronic design automation
design methodology |
0.5 | 1 | 2021 | On Predictable Reconfigurable System Design · ACM Trans. Archit. Code Optim. 2021 |
Reconfigurable computing and FPGAs › reconfigurable computing
reconfigurable system design |
0.5 | 1 | 2021 | On Predictable Reconfigurable System Design · ACM Trans. Archit. Code Optim. 2021 |
High-performance computing › performance engineering
performance portability |
0.4 | 1 | 2020 | Performance Portable FPGA Design · FPGA 2020 |
Parallel and multicore computing
dataflow computing |
0.2 | 1 | 2013 | Finite-Difference Wave Propagation Modeling on Special-Purpose Dataflow Machines · IEEE Trans. Parallel Distributed Syst. 2013 |
Integrated circuit design › digital circuit design
arithmetic circuit design |
0.2 | 2 | 2010 | FPGA Designs with Optimized Logarithmic Arithmetic · IEEE Trans. Computers 2010 Optimizing Hardware Function Evaluation · IEEE Trans. Computers 2005 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator |
0.1 | 1 | 2021 | On Predictable Reconfigurable System Design · ACM Trans. Archit. Code Optim. 2021 |
High-performance computing › scientific computing
scientific computing application |
0.1 | 1 | 2021 | On Predictable Reconfigurable System Design · ACM Trans. Archit. Code Optim. 2021 |
Performance modeling and evaluation
analytical modeling |
0.1 | 1 | 2020 | Performance Portable FPGA Design · FPGA 2020 |
Electronic design automation › high-level synthesis › arithmetic-level optimization
bit-width optimization |
0.1 | 2 | 2006 | Accuracy-Guaranteed Bit-Width Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 MiniBit: bit-width optimization via affine arithmetic · DAC 2005 |
Reconfigurable computing and FPGAs
FPGA arithmetic |
0.1 | 1 | 2010 | FPGA Designs with Optimized Logarithmic Arithmetic · IEEE Trans. Computers 2010 |
Integrated circuit design › digital circuit design › arithmetic circuit design
logarithmic number system |
0.1 | 1 | 2010 | FPGA Designs with Optimized Logarithmic Arithmetic · IEEE Trans. Computers 2010 |
Processor architecture and microarchitecture › instruction set architecture › instruction set extension
custom instruction |
0.1 | 1 | 2008 | CHIPS: Custom Hardware Instruction Processor Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Electronic design automation › hardware/software co-design
custom instruction identification |
0.1 | 1 | 2008 | CHIPS: Custom Hardware Instruction Processor Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Processor architecture and microarchitecture › instruction set architecture
instruction set extension |
0.1 | 1 | 2008 | CHIPS: Custom Hardware Instruction Processor Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Reconfigurable computing and FPGAs › FPGA resource optimization
FPGA design optimization |
0.1 | 2 | 2006 | Accuracy-Guaranteed Bit-Width Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 MiniBit: bit-width optimization via affine arithmetic · DAC 2005 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.1 | 1 | 2006 | ASC: a stream compiler for computing with FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Electronic design automation › design automation tools › FPGA CAD
FPGA design tools |
0.1 | 1 | 2006 | ASC: a stream compiler for computing with FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Electronic design automation
design space exploration |
0.1 | 1 | 2005 | Optimizing Hardware Function Evaluation · IEEE Trans. Computers 2005 |
Reconfigurable computing and FPGAs
FPGA implementation |
0.1 | 1 | 2005 | Optimizing Hardware Function Evaluation · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture › computer arithmetic
function evaluation unit |
0.1 | 1 | 2005 | Optimizing Hardware Function Evaluation · IEEE Trans. Computers 2005 |
High-performance computing › supercomputing
exascale computing |
0.0 | 1 | 2013 | Finite-Difference Wave Propagation Modeling on Special-Purpose Dataflow Machines · IEEE Trans. Parallel Distributed Syst. 2013 |
Methods — techniques the papers use, named apart from their topics
analytical modeling · 0.5high-level synthesis · 0.4analytical performance modeling · 0.4java · 0.2domain-specific language · 0.2integer linear programming · 0.1polynomial approximation · 0.1bit-accurate simulation · 0.1if-conversion · 0.1design space exploration · 0.1adaptive simulated annealing · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Forecasting Prescription Efficacy
Hao-Ren Yao, Oskar Mencer, Han-Sun Chiang, Der-Chen Chang, Ophir Frieder |
ECIR (5) | 2 |
| 2021 | On Predictable Reconfigurable System DesignabstractWe propose a design methodology to facilitate rigorous development of complex applications targeting reconfigurable hardware. Our methodology relies on analytical estimation of system performance and area utilisation for a given specific application and a particular system instance consisting of a controlflow machine working in conjunction with one or more reconfigurable dataflow accelerators. The targeted application is carefully analyzed, and the parts identified for hardware acceleration are reimplemented as a set of representative software models. Next, with the results of the application analysis, a suitable system architecture is devised and its performance is evaluated to determine bottlenecks, allowing predictable design. The architecture is iteratively refined, until the final version satisfying the specification requirements in terms of performance and required hardware area is obtained. We validate the presented methodology using a widely accepted convolutional neural network (VGG-16) and an important HPC application (BQCD). In both cases, our methodology relieved and alleviated all system bottlenecks before the hardware implementation was started. As a result the architectures were implemented first time right, achieving state-of-the-art performance within 15% of our modelling estimations. Nils Voss, Bastiaan Kwaadgras, Oskar Mencer, Wayne Luk, Georgi Gaydadjiev |
ACM Trans. Archit. Code Optim. | 3 |
| 2020 | Performance Portable FPGA DesignabstractFPGA platforms are widely used for application acceleration. Although a number of high-level design frameworks exist, application and performance portability across different platforms remain challenging. To address the above problem, we propose an API design for high-level development tools to separate platform-dependent code from the remaining application design. Additionally, we propose design guidelines to assist with performance portability. To demonstrate our techniques, a large-scale application, originally developed for an Intel Stratix-V FPGA is ported to several new Xilinx Virtex UltraScale+ systems. The accelerated application, developed in a high-level framework, is rapidly moved onto the new platforms with minimal changes. The original, unmodified kernel code delivers a 1.74x speedup due to increased clock frequency on the new platform. Subsequently, the application is further optimised to make use of the additional resources available on the larger Ultrascale+ FPGAs, guided by a simple analytical performance model. This results in an additional performance increase of up to 7.4x. Using the presented framework, we demonstrate rapid deployment of the same application across a number of different platforms that leverage the same FPGA family but differ in their low-level implementation details and the available peripherals. As a result, the same application code supports five different platforms: Maxeler MAX5C DFE, Amazon EC2 F1, Xilinx Alveo U200, U250 and the original Intel Stratix-V accelerator card, with performance close to what is theoretically achievable for each of these platforms. Nils Voss, Tobias Becker, Simon Tilbury, Georgi Gaydadjiev, Oskar Mencer, Anna Maria Nestorov, Enrico Reggiani, Wayne Luk |
FPGA | 5 |
| 2019 | Towards Real Time Radiotherapy SimulationabstractWe propose a novel reconfigurable hardware architecture to implement Monte Carlo based simulation of physical dose accumulation for intensity-modulated adaptive radiotherapy. The long term goal of our effort is to provide accurate online dose calculation in real-time during patient treatment. This will allow wider adoption of personalised patient therapies which has the potential to significantly reduce dose exposure to the patient as well as shorten treatment and greatly reduce costs. The proposed architecture exploits the inherent parallelism of Monte Carlo simulations to perform domain decomposition and provide high resolution simulation without being limited by on-chip memory capacity. We present our architecture in detail and provide a performance model to estimate execution time, hardware area and bandwidth utilisation. Finally, we evaluate our architecture on a Xilinx VU9P platform and show that three cards are sufficient to meet our real time target of 100 million randomly generated particle histories per second. Nils Voss, Peter Ziegenhein, Lukas Vermond, Joost Hoozemans, Oskar Mencer, Uwe Oelfke, Wayne Luk, Georgi Gaydadjiev |
ASAP | 5 |
| 2019 | Memory Mapping for Multi-die FPGAsabstractThis paper proposes an algorithm for mapping logical to physical memory resources on FPGAs. Our greedy strategy based algorithm is specifically designed to facilitate timing closure on modern multi-die FPGAs for static-dataflow accelerators utilising most of the on-chip resources. The main objective of the proposed algorithm is to ensure that specific sub-parts of the design under consideration can fully reside within a single die to limit inter-die communication. The above is achieved by performing the memory mapping for each sub-part of the design separately while keeping allocation of the available physical resources balanced. As a result the number of inter-die connections is reduced on average by 50% compared to an algorithm targeting minimal area usage for real, complex applications using most of the on-chip's resources. Additionally, our algorithm is the only one out of the four evaluated approaches which successfully produces place and route results for all 33 applications and benchmarks. Nils Voss, Pablo Quintana, Oskar Mencer, Wayne Luk, Georgi Gaydadjiev |
FCCM | 3 |
| 2017 | Convolutional Neural Networks on Dataflow EnginesabstractIn this paper we discuss a high performance implementation for Convolutional Neural Networks (CNNs) inference on the latest generation of Dataflow Engines (DFEs). We discuss the architectural choices made during the design phase taking into account the DFE chip properties. We then perform design space exploration, considering the memory bandwidth and resources utilisation constraints derived from the used DFE and the chosen architecture. Finally, we discuss the high performance implementation and compare the obtained performance against other implementations, showing that our proposed design reaches 2,450 GOPS when running VGG16 as a test case. Nils Voss, Marco Bacis, Oskar Mencer, Georgi Gaydadjiev, Wayne Luk |
ICCD | 3 |
| 2016 | Dataflow design for optimal incremental SVM trainingabstractThis paper proposes a new parallel architecture for incremental training of a Support Vector Machine (SVM), which produces an optimal solution based on manipulating the Karush-Kuhn-Tucker (KKT) conditions. Compared to batch training methods, our approach avoids re-training from scratch when training dataset changes. The proposed architecture is the first to adopt an efficient dataflow organisation. The main novelty is a parametric description of the parallel dataflow architecture, which deploys customisable arithmetic units for dense linear algebraic operations involved in updating the KKT conditions. The proposed architecture targets on-line SVM training applications. Experimental evaluation with real world financial data shows that our architecture implemented on Stratix-V FPGA achieved significant speedup against LIBSVM on Core i7-4770 CPU. Shengjia Shao, Oskar Mencer, Wayne Luk |
FPT | 2 |
| 2014 | A highly-efficient and green data flow engine for solving euler atmospheric equationsabstractAtmospheric modeling is an essential issue in the study of climate change. However, due to the complicated algorithmic and communication models, scientists and researchers are facing tough challenges in finding efficient solutions to solve the atmospheric equations. In this paper, we accelerate a solver for the three-dimensional Euler atmospheric equations through reconfigurable data flow engines. We first propose a hybrid design that achieves efficient resource allocation and data reuse. Furthermore, through algorithmic offsetting, fast memory table, and customizable-precision arithmetic, we map a complex Euler kernel into a single FPGA chip, which can perform 956 floating point operations per cycle. In a 1U-chassis, our CPU-DFE unit with 8 FPGA chips is 18.5 times faster and 8.3 times more power efficient than a multicore system based on two 12-core Intel E5-2697 (Ivy Bridge) CPUs, and is 6.2 times faster and 5.2 times more power efficient than a hybrid unit equipped with two 12-core Intel E5-2697 (Ivy Bridge) CPUs and three Intel Xeon Phi 5120d (MIC) cards. Lin Gan 0001, Haohuan Fu, Chao Yang 0002, Wayne Luk, Wei Xue 0003, Oskar Mencer, Xiaomeng Huang, Guangwen Yang 0002 |
FPL | 6 |
| 2013 | Going to the wire: The next generation financial risk management platform
Ari Studnitzer, Oskar Mencer |
Hot Chips Symposium | 2 |
| 2013 | Finite-Difference Wave Propagation Modeling on Special-Purpose Dataflow MachinesabstractModeling wave propagation through the earth is an important application in geoscience. We present a framework for wave propagation modeling on special-purpose hardware, which dramatically improves the application performance compared to conventional CPUs. We utilize custom hardware platforms consisting of a mix of x86 CPUs and dataflow engines connected by high-bandwidth communication links. Application programmers describe their algorithms in a domain specific language using Java syntax, with special dataflow semantics overlayed on top of the Java language. The application-specific dataflow engines run at hundreds of MHz with massive parallelism and deliver high performance/Watt, up to 30 times more energy efficient than conventional CPUs. The power efficiency of this approach suggests that dataflow computing may have a key role to play in the improvements in power efficiency necessary to reach exascale computing. Oliver Pell, Jacob A. Bower, Robert G. Dimond, Oskar Mencer, Michael J. Flynn |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2012 | Dataflow supercomputingabstractOver the past decades parallel processor speedup has been an elusive quantity for a broad class of applications. Yet with the end of performance scaling for single processors the need for speedup has never been greater. The problem is not technology but programming models. One answer to this speedup problem is to create an idealized data flow machine that exactly corresponds to the application and stream data through the resulting machine. This approach can be emulated with FPGAs, providing more than an order of magnitude speedup even as executed as an emulation of the data flow machine. Michael J. Flynn, Oliver Pell, Oskar Mencer |
FPL | 3 |
| 2012 | Rapid computation of value and risk for derivatives portfoliosabstractSUMMARY We report new results from an on‐going project to accelerate derivatives computations. Our earlier work was focused on accelerating the valuation of credit derivatives. In this paper, we extend our work in two ways: by applying the same techniques, first, to accelerate the computation of portfolio level risk for credit derivatives and, second, to different asset classes using a different type of mathematical model, which together present challenges that are quite different to those dealt with in our earlier work. Specifically, we report acceleration over 270 times faster than a single Intel Core for a multi‐asset Monte Carlo model. We also explore the implications for risk. Copyright © 2011 John Wiley & Sons, Ltd. Stephen Weston, James Spooner, Sébastien Racanière, Oskar Mencer |
Concurr. Comput. Pract. Exp. | 4 |
| 2012 | FISH: Fast Instruction SyntHesis for Custom ProcessorsabstractThis paper presents Fast Instruction SyntHesis (FISH), a system that supports automatic generation of custom instruction processors from high-level application descriptions to enable fast design space exploration. FISH is based on novel methods for automatically adapting the instruction set to match an application in a high-level language such as C or C++. FISH identifies custom instruction candidates using two approaches: 1) by enumerating maximal convex subgraphs of application data flow graphs and 2) by integer linear programming (ILP). The experiments, involving ten multimedia and cryptography benchmarks, show that our contributed algorithms are the fastest among the state-of-the-art techniques. In most cases, enumeration takes only milliseconds to execute. The longest enumeration run-time observed is less than six seconds. ILP is usually slower than enumeration, but provides us with a complementary solution technique. Both enumeration and ILP allow the use of multiple different merit functions in the evaluation of data-flow subgraphs. The experiments demonstrate that, using only modest additional hardware resources, up to 30-fold performance improvement can be obtained with respect to a single-issue base processor. Kubilay Atasu, Wayne Luk, Oskar Mencer, Can C. Özturan, Günhan Dündar |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | Surviving the end of scaling of traditional micro processors in HPC
Olav Lindtjorn, Robert G. Clapp, Oliver Pell, Oskar Mencer, Michael J. Flynn |
Hot Chips Symposium | 4 |
| 2010 | FPGA Designs with Optimized Logarithmic ArithmeticabstractUsing a general polynomial approximation approach, we present an arithmetic library generator for the logarithmic number system (LNS). The generator produces optimized LNS arithmetic libraries that improve significantly over previous LNS designs on area and latency. We also provide area cost estimation and bit-accurate simulation tools that facilitate comparison between LNS and floating-point designs. Haohuan Fu, Oskar Mencer, Wayne Luk |
IEEE Trans. Computers | 2 |
| 2008 | Fast custom instruction identification by convex subgraph enumerationabstractAutomatic generation of custom instruction processors from high-level application descriptions enables fast design space exploration, while offering very favorable performance and silicon area combinations. This work introduces a novel method for adapting the instruction set to match an application captured in a high-level language. A simplified model is used to find the optimal instructions via enumeration of maximal convex subgraphs of application data flow graphs (DFGs). Our experiments involving a set of multimedia and cryptography benchmarks show that an order of magnitude performance improvement can be achieved using only a limited amount of hardware resources. In most cases, our algorithm takes less than a second to execute. Kubilay Atasu, Oskar Mencer, Wayne Luk, Can C. Özturan, Günhan Dündar |
ASAP | 2 |
| 2008 | An Approach to Graph and Netlist CompressionabstractWe introduce an EDIF netlist graph-based compression algorithm which is lossy with respect to the original byte stream but lossless in terms of the circuit information it contains. The algorithm builds on the graph mining tool SUBDUE. Our algorithm, CEDIF (compressed EDIF), compresses the EDIF file to a size about 39% of the size of the compressed file resulting from the state-of-the-art PAQ text compression algorithm and to about 85% of the size of a GRAPHITOUR-like graph compression algorithm. Jeehong Yang, Serap A. Savari, Oskar Mencer |
DCC | 3 |
| 2008 | Power-Aware and Branch-Aware Word-Length OptimizationabstractPower reduction is becoming more important as circuit size increases. This paper presents a tool called PowerCutter which employs accuracy-guaranteed word-length optimization to reduce power consumption of circuits. We adapt circuit word-lengths at run time to decrease power consumption, with optimizations based on branch statistics. Our tool uses a technique based on Automatic Differentiation to analyze library cores specified as black box functions, which do not include implementation information. We use this technique to analyze benchmarks containing library functions such as square root. Our approach shows that power savings of up to 32% can be achieved on benchmarks which cannot be analyzed by previous approaches, because library cores with an unknown implementation are used. William George Osborne, José Gabriel F. Coutinho, Wayne Luk, Oskar Mencer |
FCCM | 4 |
| 2008 | Optimizing residue arithmetic on FPGAsabstractResidue Number System (RNS), which originates from the Chinese Remainder Theorem, is regarded as a promising number representation in the domain of Digital Signal Processing (DSP). This paper describes our work on optimizing residue arithmetic units on the platform of reconfigurable devices, such as FPGAs. First, we provide improved designs for residue arithmetic units. For reverse converters from RNS to binary numbers, we propose a novel design that uses only n-bit additions. Compared to previous work, the design consumes up to 14.3% less area and provides lower latency. Second, we develop a reconfigurable RNS arithmetic library generator for the moduli set {2n−1, 2n, 2n+1}. The generator supports a wide range of RNS numbers, and enables us to perform an extensive comparison between RNS and other number representations at both the arithmetic unit level and the application level. The comparison shows that, for applications involving a large number of multiplications, the RNS designs can reduce up to 1/2 DSP48s for large bit-width settings. Haohuan Fu, Oskar Mencer, Wayne Luk |
FPT | 2 |
| 2008 | Finding Speedup in Parallel ProcessorsabstractWhile recently the focus of architects and programmers has been on multi core, the alternative of processor node plus array oriented accelerator has some significant advantages especially in compute intensive static applications. We propose an acceleration methodology based on FPGA arrays (but, in principle it could be GPU or Cell based). The methodology uses a comprehensive application analysis supported by high performance FPGA hardware. The analysis provides a dataflow graph of the application which is replicated in SIMD for multiple data strips until limited by the pin bandwidth, then pipelined (MISD) until circuit limited. An oil exploration application shows the possibility of speedup of over 300x over an Intel Xeon. Michael J. Flynn, Robert G. Dimond, Oskar Mencer, Oliver Pell |
ISPDC | 3 |
| 2008 | CHIPS: Custom Hardware Instruction Processor SynthesisabstractThis paper describes an integer-linear-programming (ILP)-based system called custom hardware instruction processor synthesis (CHIPS) that identifies custom instructions for critical code segments, given the available data bandwidth and transfer latencies between custom logic and a baseline processor with architecturally visible state registers. Our approach enables designers to optionally constrain the number of input and output operands for custom instructions. We describe a design flow to identify promising area, performance, and code-size tradeoffs. We study the effect of input/output constraints, register-file ports, and compiler transformations such as if-conversion. Our experiments show that, in most cases, the solutions with the highest performance are identified when the input/output constraints are removed. However, input/output constraints help our algorithms identify frequently used code segments, reducing the overall area overhead. Results for 11 benchmarks covering cryptography and multimedia are shown, with speed-ups between 1.7 and 6.6 times, code-size reductions between 6% and 72%, and area costs ranging between 12 and 256 adders for maximum speed-up. Our ILP-based approach scales well: benchmarks with basic blocks consisting of more than 1000 instructions can be optimally solved, most of the time within a few seconds. Kubilay Atasu, Can C. Özturan, Günhan Dündar, Oskar Mencer, Wayne Luk |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2007 | Optimizing instruction-set extensible processors under data bandwidth constraintsabstractThe authors present a methodology for generating optimized architectures for data bandwidth constrained extensible processors. The authors describe a scalable integer linear programming (ILP) formulation, that extracts the most profitable set of instruction-set extensions given the available data bandwidth and transfer latency. Unlike previous approaches, the authors differentiate between number of inputs and outputs for instruction-set extensions and the number of register file ports. This differentiation makes the approach applicable to architectures that include architecturally visible state registers and dedicated data transfer channels. The authors support a comprehensive design space exploration to characterize the area/performance trade-offs for various applications. The authors evaluate our approach using actual ASIC implementations to demonstrate that our automatically customized processors meet timing within the target silicon area. For an embedded processor with only two register read ports and one register write port, the authors obtain up to 4.3times speed-up with extensions incurring only a 35% area overhead Kubilay Atasu, Robert G. Dimond, Oskar Mencer, Wayne Luk, Can C. Özturan, Günhan Dündar |
DATE | 3 |
| 2007 | Optimizing Logarithmic Arithmetic on FPGAsabstractThis paper proposes optimizations of the methods and parameters used in both mathematical approximation and hardware design for logarithmic number system (LNS) arithmetic. First, we introduce a general polynomial approximation approach with an adaptive divide-in-halves segmentation method for evaluation of LNS arithmetic functions. Second, we develop a library generator that automatically generates optimized LNS arithmetic units with a wide bit-width range from 21 to 64 bits, to support LNS application development and design exploration. The basic arithmetic units are tested on practical FPGA boards as well as software simulation. When compared with existing LNS designs, our generated units provide in most cases 6% to 37% reduction in area and 20% to 50% reduction in latency. The key challenge for LNS remains on the application level. We show the performance of LNS versus floating-point for realistic applications: digital sine/cosine waveform generator, matrix multiplication and radiative Monte Carlo simulation. Our infrastructure for fast prototyping LNS FPGA applications allows us to efficiently study LNS number representation and its tradeoffs in speed and size when compared with floating-point designs. Haohuan Fu, Oskar Mencer, Wayne Luk |
FCCM | 2 |
| 2007 | Automatic Accuracy-Guaranteed Bit-Width Optimization for Fixed and Floating-Point SystemsabstractIn this paper we present Minibit+, an approach that optimizes the bit-widths of fixed-point and floating-point designs, while guaranteeing accuracy. Our approach adopts different levels of analysis giving the designer the opportunity to terminate it at any stage to obtain a result. Range analysis is achieved using a combined affine and interval arithmetic approach to reduce the number of bits. Precision analysis involves a coarse-grain and fine-grain analysis. The best representation, in fixed-point or floating-point, for the numbers is then chosen based on the range, precision and latency. Three case studies are used: discrete cosine transform, B-Splines and RGB to YCbCr color conversion. Our analysis can run over 200 times faster than current approaches to this problem while producing more accurate results, on average within 2-3% of an exhaustive search. William George Osborne, Ray C. C. Cheung, José Gabriel F. Coutinho, Wayne Luk, Oskar Mencer |
FPL | 5 |
| 2007 | Instrumented Multi-Stage Word-Length OptimizationabstractIn this paper we present a tool, LengthFinder, for optimizing word-lengths of hardware designs with fixed-point arithmetic based on analytical error models that guarantee accuracy. LengthFinder adopts a multi-stage approach, with four novel features. First, the code analysis stage selects loops to instrument, such that information about the number of iterations can be extracted to generate more accurate results. Second, aggressive heuristics are used to produce non-uniform word-lengths rapidly while meeting requirements from the guaranteed error functions. Third, a method capable of reducing the search space has been developed for data-partitioning with a variable word-length reduction. Fourth, a genetic algorithm with selective-crossover and high mutation probability is applied to obtain near-optimal results. The benefits of LengthFinder are illustrated with various case studies. We show that LengthFinder can run over 200 times faster than previous techniques (Lee et al., 2006), while producing more accurate results, relative to values obtained from integer linear programming. William George Osborne, José Gabriel F. Coutinho, Ray C. C. Cheung, Wayne Luk, Oskar Mencer |
FPT | 5 |
| 2007 | Improving Bounds for FPGA Logic MinimizationabstractWe present a methodology for improving the bounds of combinational designs implemented on networks of lookup tables, moving them closer to the theoretical minimum. Our work effectively extends optimality to span logic minimization and technology mapping. We obtain a proof of optimality by restricting ourselves to 4-input look-up tables (LUTs) and generating all possible circuits up to a certain area or latency depending on the optimization mode. Since simple-minded generation would take a long time, we develop levels of abstraction (steps) and techniques to restrict and order the search space, and produce results in practical time. We use logic decomposition to break up large designs, using the resulting trees to guide our search and prune the search space. The price of this optimality is that we are limited to small blocks; however, such blocks can be used to build larger designs. Tim Todman, Haofan Fu, Oskar Mencer, Wayne Luk |
FPT | 3 |
| 2006 | Automating processor customisation: optimised memory access and resource sharingabstractWe propose a novel methodology to generate application specific instruction processors (ASIPs) including custom instructions. Our implementation balances performance and area requirements by making custom instructions reusable across similar pieces of code. In addition to arithmetic and logic operations, table look-ups within custom instructions reduce costly accesses to global memory. We present synthesis and cycle-accurate simulation results for six embedded benchmarks running on customised processors. Reusable custom instructions achieve an average 319% speedup with only 5% additional area. The maximum speedup of 501% for the advanced encryption standard (AES) requires only 3.6% additional area Robert G. Dimond, Oskar Mencer, Wayne Luk |
DATE | 2 |
| 2006 | Combining Instruction Coding and Scheduling to Optimize Energy in System-on-FPGAabstractIn this paper, we investigate a combination of two techniquesnstruction coding and instruction re-ordering - for optimizing energy in embedded processor control. We present the first practical, hardware implementation incorporating both approaches as part of a novel flow for automatic power-optimization of an FPGA soft processor. Our infrastructure generates customized processors and associated software, to enable power optimizations to be evaluated on multiple architectures and FPGA platforms. We evaluate using both software estimates of power and actual measurements from both low-cost and high-performance FPGAs. We generate over 150 optimized processor designs for two FPGA platforms, two processor architectures and six different benchmarks at four different clock rates and achieve consistent measured dynamic power reduction of up to 74%, without performance cost. Our results are applicable beyond processor optimization, quantifying the benefits of practical switching reduction and highlighting non-obvious pitfalls and complexities in dynamic power optimization Robert G. Dimond, Oskar Mencer, Wayne Luk |
FCCM | 2 |
| 2006 | ASC-Based Acceleration in an FPGA with a Processor Core Using Software-Only SkillsabstractA stream compiler (ASC) generates net lists for hardware (FPGA) accelerators from C-like descriptions, obviating the need for hardware skills. The authors present a backend adapter that enables integration of such accelerators with a processor core in the same FPGA. Development of hybrid ASC-accelerated applications using software-only skills is thus made possible, as illustrated by a hybrid power-conscious iDCT implementation Evgeny Fiksman, Yitzhak Birk, Oskar Mencer |
FCCM | 3 |
| 2006 | FPGAs, GPUs and the PS2 - A Single Programming MethodologyabstractField programmable gate arrays (FPGAs), graphics processing units (GPUs) and Sony's Playstation 2 vector units offer scope for hardware acceleration of applications. Implementing algorithms on multiple architectures can be a long and complicated process. We demonstrate an approach to compiling for FPGAs, GPUs and PS2 vector units using a unified description based on A Stream Compiler (ASC) for FPGAs. As an example of its use we implement a Monte Carlo simulation using ASC. The unified description allows us to evaluate optimisations for specific architectures on top of a single base description, saving time and effort Lee W. Howes, Paul Price, Oskar Mencer, Olav Beckmann |
FCCM | 3 |
| 2006 | Comparing FPGAs to Graphics Accelerators and the Playstation 2 Using a Unified Source DescriptionabstractField programmable gate arrays (FPGAs), graphics processing units (GPUs) and Sony's Playstation 2 vector units offer scope for hardware acceleration of applications. We compare the performance of these architectures using a unified description based onA Stream Compiler(ASC) for FPGAs, which has been extended to target GPUs and PS2 vector units. Programming these architectures from a single description enables us to reason about optimizations for the different architectures. Using the ASC description we implement a Monte Carlo simulation, a fast Fourier transform (FFT) and a weighted sum algorithm. Our results show that without much optimization the GPU is suited to the Monte Carlo simulation, while the weighted sum is better suited to PS2 vector units. FPGA implementations benefit particularly from architecture specific optimizations which ASC allows us to easily implement by adding simple annotations to the shared code. Lee W. Howes, Paul Price, Oskar Mencer, Olav Beckmann, Oliver Pell |
FPL | 3 |
| 2006 | Comparing floating-point and logarithmic number representations for reconfigurable accelerationabstractThe paper investigates floating-point and logarithmic number representations for computing with FPGAs. The key issue is to select the best number format for an application to improve performance and accuracy. Using A Stream Compiler, ASC as the hardware design and compilation tool, a convenient scheme to compare the designs of both floating-point and logarithmic numbers and select the solution with the best performance and accuracy, was developed. Its contributions are: (1) optimized function evaluations for conversions between logarithmic and floating-point numbers; (2) design and implementation of logarithmic arithmetic, with optimized segmentation and polynomial degree; (3) a practical comparison case study of Monte Carlo radiative heat transfer simulation. Compared to prior work, our design supports two to six times more LNS conversion and LNS arithmetic units on one FPGA. For Monte Carlo simulation, our designs of both number systems produce 39-80% higher throughput with either a smaller area or a higher accuracy Haohuan Fu, Oskar Mencer, Wayne Luk |
FPT | 2 |
| 2006 | Towards optimal custom instruction processors
Wayne Luk, Kubilay Atasu, Robert G. Dimond, Oskar Mencer |
Hot Chips Symposium | 4 |
| 2006 | Accuracy-Guaranteed Bit-Width OptimizationabstractAn automated static approach for optimizing bit widths of fixed-point feedforward designs with guaranteed accuracy, called MiniBit, is presented. Methods to minimize both the integer and fraction parts of fixed-point signals with the aim of minimizing the circuit area are described. For range analysis, the technique in this paper identifies the number of integer bits necessary to meet range requirements. For precision analysis, a semianalytical approach with analytical error models in conjunction with adaptive simulated annealing is employed to optimize the number of fraction bits. The analytical models make it possible to guarantee overflow/underflow protection and numerical accuracy for all inputs over the user-specified input intervals. Using a stream compiler for field-programmable gate arrays (FPGAs), the approach in this paper is demonstrated with polynomial approximation, RGB-to-YCbCr conversion, matrix multiplication, B-splines, and discrete cosine transform placed and routed on a Xilinx Virtex-4 FPGA. Improvements for a given design reduce the area and the latency by up to 26% and 12%, respectively, over a design using optimum uniform fraction bit widths. Studies show that MiniBit-optimized designs are within 1% of the area produced from the integer linear programming approach Dong-U Lee, Altaf Abdul Gaffar, Ray C. C. Cheung, Oskar Mencer, Wayne Luk, George A. Constantinides |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2006 | ASC: a stream compiler for computing with FPGAsabstractA stream compiler (ASC) for computing with field programmable gate arrays (FPGAs) emerges from the ambition to bridge the hardware-design productivity gap where the number of available transistors grows more rapidly than the productivity of very large scale integration (VLSI) and FPGA computer-aided-design (CAD) tools. ASC addresses this problem with a softwarelike programming interface to hardware design (FPGAs) while keeping the performance of hand-designed circuits at the same time. ASC improves productivity by letting the programmer optimize the implementation on the algorithm level, the architecture level, the arithmetic level, and the gate level, all within the same C++ program. The increased productivity of ASC is applied to the hardware acceleration of a wide range of applications. Traditionally, hardware accelerators are tediously handcrafted to achieve top performance. ASC simplifies design-space exploration of hardware accelerators by transforming the hardware-design task into a software-design process, using only "GNU compiler collection (GCC)" and "make" to obtain a hardware netlist. From experience, the hardware-design productivity and ease of use are close to pure software development. This paper presents results and case studies with optimizations that are: 1) on the gate level-Kasumi and International Data Encryption Algorithm (IDEA) encryptions; 2) on the arithmetic level-redundant addition and multiplication function evaluation for two-dimensional (2-D) rotation; and 3) on the architecture level-Wavelet and Lempel-Ziv (LZ)-like compression Oskar Mencer |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Automating custom-precision function evaluation for embedded processorsabstractDue to resource and power constraints, embedded processors often cannot afford dedicated floating-point units. For instance, the IBM PowerPC processor embedded in Xilinx Virtex-II Pro FPGAs only supports emulated floating-point arithmetic, which leads to slow operation when floating-point arithmetic is desired. This paper presents a customizable mathematical library using fixed-point arithmetic for elementary function evaluation. We approximate functions via polynomial or rational approximations depending on the user-defined accuracy requirements. The data representation for the inputs and outputs are compatible with IEEE single-precision and double-precision floating-point formats. Results show that our 32-bit polynomial method achieves over 80 times speedup over the single-precision mathematical library from Xilinx, while our 64-bit polynomial method achieves over 30 times speedup. Ray C. C. Cheung, Dong-U Lee, Oskar Mencer, Wayne Luk, Peter Y. K. Cheung |
CASES | 3 |
| 2005 | MiniBit: bit-width optimization via affine arithmeticabstractMiniBit, our automated approach for optimizing bit-widths of fixed-point designs is based on static analysis via affine arithmetic. We describe methods to minimize both the integer and fraction parts of fixed-point signals with the aim of minimizing circuit area. Our range analysis technique identifies the number of integer bits required. For precision analysis, we employ a semi-analytical approach with analytical error models in conjunction with adaptive simulated annealing to find the optimum number of fraction bits. Improvements for a given design reduce area and latency by up to 20% and 12% respectively, over optimum uniform fraction bit-widths on a Xilinx Virtex-4 FPGA. Dong-U Lee, Altaf Abdul Gaffar, Oskar Mencer, Wayne Luk |
DAC | 3 |
| 2005 | CUSTARD - A Customisable Threaded FPGA Soft Processor and ToolsabstractWe propose CUSTARD - customisable threaded architecture - a soft processor design space that combines support for multiple hardware threads and automatically generated custom instructions. Multiple threads incur low additional hardware cost and allow fine-grained concurrency without multiple processor cores or software overhead. Custom instructions, generated for a specific application, accelerate frequently performed computations by implementing them as dedicated hardware. In this paper we present a flexible processor and compiler generation system, FPGA implementations of CUSTARD and performance/area results for media and cryptography benchmarks. Robert G. Dimond, Oskar Mencer, Wayne Luk |
FPL | 2 |
| 2005 | Custom Hardware Architectures for Posture Analysis
M. P. T. Juvonen, José Gabriel F. Coutinho, J. L. Wang, Benny P. L. Lo, Wayne Luk, Oskar Mencer, Guang-Zhong Yang |
FPT | 6 |
| 2005 | Optimizing Hardware Function EvaluationabstractWe present a methodology and an automated system for function evaluation unit generation. Our system selects the best function evaluation hardware for a given function, accuracy requirements, technology mapping, and optimization metrics, such as area, throughput, and latency. Function evaluation f(x) typically consists of range reduction and the actual evaluation on a small convenient interval such as [0, /spl pi//2) for sin(x). We investigate the impact of hardware function evaluation with range reduction for a given range and precision of x and f(x) on area and speed. An automated bit-width optimization technique for minimizing the sizes of the operators in the data paths is also proposed. We explore a vast design space for fixed-point sin(x), log(x), and /spl radic/x accurate to one unit in the last place using MATLAB and ASC, a stream compiler for field-programmable gate arrays (FPGAs). In this study, we implement over 2,000 placed-and-routed FPGA designs, resulting in over 100 million application-specific integrated circuit (ASIC) equivalent gates. We provide optimal function evaluation results for range and precision combinations between 8 and 48 bits. Dong-U Lee, Altaf Abdul Gaffar, Oskar Mencer, Wayne Luk |
IEEE Trans. Computers | 3 |
| 2004 | Unifying Bit-Width Optimisation for Fixed-Point and Floating-Point DesignsabstractThis paper presents a method that offers a uniform treatment for bit-width optimisation of both fixed-point and floating-point designs. Our work utilises automatic differentiation to compute the sensitivities of outputs to the bit-width of the various operands in the design. This sensitivity analysis enables us to explore and compare fixed-point and floating-point implementation for a particular design. As a result, we can automate the selection of the optimal number representation for each variable in a design to optimize area and performance. We implement our method in the BitSize tool targeting reconfigurable architectures, which takes user-defined constraints to direct the optimisation procedure. We illustrate our approach using applications such as ray-tracing and function approximation. Altaf Abdul Gaffar, Oskar Mencer, Wayne Luk, Peter Y. K. Cheung |
FCCM | 2 |
| 2004 | Automating Optimized Table-with-Polynomial Function Evaluation for FPGAs
Dong-U Lee, Oskar Mencer, David J. Pearce 0001, Wayne Luk |
FPL | 2 |
| 2004 | Adaptive range reduction for hardware function evaluationabstractFunction evaluation f(x) typically consists of range reduction and the actual function evaluation on a small interval. We investigate optimization of range reduction given the range and precision of x and f(x). For every function evaluation there exists a convenient interval such as [0, /spl pi//2) for sin(x). The adaptive range reduction method, which we propose in this work, involves deciding whether range reduction can be used effectively for a particular design. The decision depends on the function being evaluated, precision, and optimization metrics such as area, latency and throughput. In addition, the input and output range has an impact on the preferable function evaluation method such as polynomial, table-based, or combinations of the two. We explore this vast design space of adaptive range reduction for fixed-point sin(x), log(x) and /spl radic/(x) accurate to one unit in the last place using MATLAB and ASC, A Stream Compiler. These tools enable us to study over 1000 designs resulting in over 40 million Xilinx equivalent circuit gates, in a few hours' time. The final objective is to progress towards a fully automated library that provides optimal function evaluation hardware units given input/output range and precision. Dong-U Lee, Altaf Abdul Gaffar, Oskar Mencer, Wayne Luk |
FPT | 3 |
| 2003 | Floating Point Unit Generation and Evaluation for FPGAsabstractMost commercial and academic floating point libraries for FPGAs (field programmable gate arrays) provide only a small fraction of all possible floating point units. In contrast, the floating point unit generation approach outlined in this paper allows for the creation of a vast collection of floating point units with differing throughput, latency, and area characteristics. Given performance requirements, our generation tool automatically chooses the proper implementation algorithm and architecture to create a compliant floating point unit. Our approach is fully integrated into standard C++ using ASC, a stream compiler for FPGAs, and the PAM-Blox II module generation environment. The floating point units created by our approach exhibit a factor of two latency improvement versus commercial FPGA floating point units, while consuming only half of the FPGA logic area. Russell Tessier, Oskar Mencer |
FCCM | 3 |
| 2003 | Hardware Design with a Scripting Language
Per Haglund, Oskar Mencer, Wayne Luk, Benjamin Tai |
FPL | 2 |
| 2003 | Design space exploration with A Stream CompilerabstractWe consider speeding up general-purpose applications with hardware accelerators. Traditionally hardware accelerators are tediously hand-crafted to achieve top performance ASC (A Stream Complier) simplifies exploration of hardware accelerators by transforming the hardware design task into a software design process using only 'gcc' and 'make' to obtain a hardware netlist. ASC enables programmers to customize hardware accelarators at three levels of abstraction: the architecture level, the functional block level, and the bit level. All three customizations are based on one uniform representation: a single C++ program with custom types and operators for each level of abstraction. This representation allows ASC users to express and reason about the design space, extract parallelism at each level and quickly evaluate different design choices. In addition, since the user has full control over each gate-level resource in the entire design. ASC accelerator performance can always be equal to or better than hand-crafted designs, usually with much less effort. We present several ASC bench marks, including wavelet compression and Kasumi encryption. Oskar Mencer, David J. Pearce 0001, Lee W. Howes, Wayne Luk |
FPT | 1 |
| 2002 | PAM-Blox II: Design and Evaluation of C++ Module Generation for Computing with FPGAsabstractThis paper explores the implications of integrating flexible module generation into a compiler for FPGAs. The objective is to improve the programmability of FPGAs, or in other words, the productivity of the FPGA programmer. We describe (1) the module generation library PAM-Blox II, the second generation of object-oriented module generators in C++, targeted at computing with FPGAs, and (2) examples of design tradeoffs and performance results using redundant representations for addition and multiplication, and technology mapping of comparison and elementary function evaluation. PAM-Blox II is built on top of a set of extensions to the gate level FPGA design library PamDC to provide a more efficient, portable, scalable, and maintainable module generator library. Using PAM-Blox II we demonstrate a simplified interface to bit-level programability. The simplification results from the bottom-up approach and a close coupling of architecture generation, module generation and gate level CAD. The tradeoffs for the module generators are based on trading area for speed and hand-optimizing technology mapping to the specific FPGA technology. As an example, we show that redundant number representations hold one key to unleashing the full potential of reconfigurability on the bit-level. The presented module generators are applied to encryption and compression to show the impact of the bit-level optimizations on application performance. Oskar Mencer |
FCCM | 1 |
| 2002 | HAGAR: Efficient Multi-context Graph Processors
Oskar Mencer, Zhining Huang, Lorenz Huelsbergen |
FPL | 1 |
| 2002 | Floating-point bitwidth analysis via automatic differentiationabstractAutomatic bitwidth analysis is a key ingredient for highlevel programming of FPGAs and high-level synthesis of VLSI circuits. The objective is to find the minimal number of bits to represent a value in order to minimise the circuit area and to improve efficiency of the respective arithmetic operations, while satisfying user-defined numerical constraints. We present a novel approach to bitwidth- or precision-analysis for floating-point designs. The approach involves analysing the dataflow graph representation of a design to see how sensitive the output of a node is to changes in the outputs of other nodes: higher sensitivity requires higher precision and hence more output bits. We automate such sensitivity analysis by a mathematical method called automatic differentiation, which involves differentiating variables in a design with respect to other variables. We illustrate our approach by optimising the bitwidth for two examples, a discrete Fourier transform (DFT) implementation and a Finite Impulse Response (FIR) filter implementation. Altaf Abdul Gaffar, Oskar Mencer, Wayne Luk, Peter Y. K. Cheung, Nabeel Shirazi |
FPT | 2 |
| 2001 | Pipelined Function Evaluation on FPGAs
Nicolas Boullis, Oskar Mencer, Wayne Luk, Henry Styles |
FCCM | 2 |
| 2001 | Parameterized Function Evaluation for FPGAs
Oskar Mencer, Nicolas Boullis, Wayne Luk, Henry Styles |
FPL | 1 |
| 2001 | Object-oriented domain specific compilers for programming FPGAsabstractSimplifying the programming models is paramount to the success of reconfigurable computing with field programmable gate arrays (FPGAs). This paper presents a methodology to combine true object-oriented design of the compiler/CAD tool with an object-oriented hardware design methodology in C++. The resulting system provides all the benefits of object-oriented design to the compiler/CAD tool designer and to the hardware designer/programmer. The two examples for domain-specific compilers presented are BSAT and StReAm. Each domain-specific compiler is targeted at a very specific application domain, such as applications that accelerate Boolean satisfiability problems with BSAT, and applications which lend themselves for implementation as a stream architecture with StReAm. The key benefit of the presented domain specific compilers is a reduction of design time by orders of magnitude while keeping the optimal performance of hand-designed circuits. Oskar Mencer, Marco Platzner, Martin Morf, Michael J. Flynn |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2000 | StReAm: Object-Oriented Programming of Stream Architectures Using PAM-BloxabstractGeneral-purpose microprocessors are designed to deliver low-latency computation with maximal clock frequencies, leading to high power consumption. In terms of performance and power consumption, latency-tolerant applications can be efficiently implemented with architectures that provide high throughputs, such as Imagine, Score, RaPiD, and the PCI-PipeRench. Stream architectures are fully-pipelined, high-throughput microarchitectures. Algorithms are executed by mapping dataflow graphs to hardware and streaming the data through the architecture. In the best case, when all the loops are fully unrolled, the dataflow graph is acyclic and the clock frequency of the design is the data rate. In this paper, we develop a domain-specific compiler for programming FPGAs, called StReAm. StReAm is build on top of the module generation environment PAM-Blox. Oskar Mencer, Heiko Hübert, Martin Morf, Michael J. Flynn |
FCCM | 1 |
| 1998 | PAM-Blox: High Performance FPGA Design for Adaptive ComputingabstractPAM-Blox are object-oriented circuit generators on top of the PCI Pamette design environment, PamDC. High-performance FPGA design for adaptive computing is simplified by using a hierarchy of optimized hardware objects described in C++. PAM-Blox consist of two major layers of abstraction. First, PamBlox are parameterizable simple elements such as counters and adders. Automatic placement of carry chains and flexible shapes are supported. PaModules are more complex elements possibly instantiating PamBlox. PaModules have fixed shapes and are usually optimized for a specific data-width. Examples for PaModules are multipliers, Coordinate Rotations (CORDICs), and special arithmetic units for encryption. The key difference of our approach to most other design tools for FPGAs is that the designer has total control over placement at each level of the design hierarchy, which is the key to high-performance FPGA design. Second, the object interface was chosen carefully to encourage code-reuse and simplify code-sharing between designers. PAM-Blox are intended to be part of an open library that allows design sharing between members of the adaptive computing community. Oskar Mencer, Martin Morf, Michael J. Flynn |
FCCM | 1 |
| 1998 | Hardware software tri-design of encryption for mobile communication unitsabstractWe explore the design space of field programmable gate arrays (FPGAs), processors and ASICs-hardware-software tri-design-in the framework of encryption for hand-held communication units. The IDEA (International Data Encryption Algorithm) is used to show the tradeoffs for the suggested technologies. The measures for comparing different options are: performance, programmability and power (P/sup 3/). More specifically we use the performance to power, or operations to energy ratio MOPS/Watt and Mbits/s/Watt to compare processors, FPGAs and ASICs. We compare the latest digital signal processor (DSP) from Texas Instruments to Xilinx XC4000 series FPGAs. Many DSP-like applications perform very well on FPGAs. We show the benefits and limitations of FPGA technology for IDEA. Oskar Mencer, Martin Morf, Michael J. Flynn |
ICASSP | 1 |