Peter Y. K. Cheung

dblp:54/1029 · DBLP profile ↗
← Back
197ranked-venue papers
4as first author
5since 2021 · last 2023
0000-0002-8236-1816ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 185 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 11Graphics, computer vision, multimedia, augmented reality and games · 8Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2
YearPublicationVenuePosition
2023 Enabling Binary Neural Network Training on the Edge
abstract
The ever-growing computational demands of increasingly complex machine learning models frequently necessitate the use of powerful cloud-based infrastructure for their training. Binary neural networks are known to be promising candidates for on-device inference due to their extreme compute and memory savings over higher-precision alternatives. However, their existing training methods require the concurrent storage of high-precision activations for all layers, generally making learning on memory-constrained devices infeasible. In this article, we demonstrate that the backward propagation operations needed for binary neural network training are strongly robust to quantization, thereby making on-the-edge learning with modern models a practical proposition. We introduce a low-cost binary neural network training strategy exhibiting sizable memory footprint reductions while inducing little to no accuracy loss vs Courbariaux & Bengio’s standard approach. These decreases are primarily enabled through the retention of activations exclusively in binary format. Against the latter algorithm, our drop-in replacement sees memory requirement reductions of 3–5×, while reaching similar test accuracy (± 2 pp) in comparable time, across a range of small-scale models trained to classify popular datasets. We also demonstrate from-scratch ImageNet training of binarized ResNet-18, achieving a 3.78× memory reduction. Our work is open-source, and includes the Raspberry Pi-targeted prototype we used to verify our modeled memory decreases and capture the associated energy drops. Such savings will allow for unnecessary cloud offloading to be avoided, reducing latency, increasing energy efficiency, and safeguarding end-user privacy.
Erwei Wang, James J. Davis 0001, Daniele Moro, Jia Jie Lim, Claudionor José Nunes Coelho Jr., Satrajit Chatterjee, Peter Y. K. Cheung, George A. Constantinides
ACM Trans. Embed. Comput. Syst.8
2023 Logic Shrinkage: Learned Connectivity Sparsification for LUT-Based Neural Networks
abstract
Field-programmable gate array (FPGA)–specific deep neural network (DNN) architectures using native lookup tables (LUTs) as independently trainable inference operators have been shown to achieve favorable area-accuracy and energy-accuracy trade-offs. The first work in this area, LUTNet, exhibited state-of-the-art performance for standard DNN benchmarks. In this article, we propose the learned optimization of such LUT-based topologies, resulting in higher-efficiency designs than via the direct use of off-the-shelf, hand-designed networks. Existing implementations of this class of architecture require the manual specification of the number of inputs per LUT, K . Choosing appropriate K a priori is challenging. Doing so at even high granularity, for example, per layer, is a time-consuming and error-prone process that leaves FPGAs’ spatial flexibility underexploited. Furthermore, prior works see LUT inputs connected randomly, which does not guarantee a good choice of network topology. To address these issues, we propose logic shrinkage , a fine-grained netlist pruning methodology enabling K to be automatically learned for every LUT in a neural network targeted for FPGA inference. By removing LUT inputs determined to be of low importance, our method increases the efficiency of the resultant accelerators. Our GPU-friendly solution to LUT input removal is capable of processing large topologies during their training with negligible slowdown. With logic shrinkage, we improve the area and energy efficiency of the best-performing LUTNet implementation of the CNV network classifying CIFAR-10 by 1.54× and 1.31×, respectively, while matching its accuracy. This implementation also reaches 2.71× the area efficiency of an equally accurate, heavily pruned binary neural network (BNN). On ImageNet, with the Bi-Real Net architecture, employment of logic shrinkage results in a post-synthesis area reduction of 2.67× vs. LUTNet, allowing for implementation that was previously impossible on today’s largest FPGAs. We validate the benefits of logic shrinkage in the context of real application deployment by implementing a face mask detection DNN using a BNN, LUTNet, and logic-shrunk layers. Our results show that logic shrinkage results in area gains versus LUTNet (up to 1.20×) and equally pruned BNNs (up to 1.08×), along with accuracy improvements.
Erwei Wang, Marie Auffret, Georgios-Ilias Stavrou, Peter Y. K. Cheung, George A. Constantinides, Mohamed S. Abdelfattah, James J. Davis 0001
ACM Trans. Reconfigurable Technol. Syst.4
2022 Logic Shrinkage: Learned FPGA Netlist Sparsity for Efficient Neural Network Inference
abstract
FPGA-specific DNN architectures using the native LUTs as independently trainable inference operators have been shown to achieve favorable area-accuracy and energy-accuracy tradeoffs. The first work in this area, LUTNet, exhibited state-of-the-art performance for standard DNN benchmarks. In this paper, we propose the learned optimization of such LUT-based topologies, resulting in higher-efficiency designs than via the direct use of off-the-shelf, hand-designed networks. Existing implementations of this class of architecture require the manual specification of the number of inputs per LUT, K. Choosing appropriate K a priori is challenging, and doing so at even high granularity, e.g. per layer, is a time-consuming and error-prone process that leaves FPGAs' spatial flexibility underexploited. Furthermore, prior works see LUT inputs connected randomly, which does not guarantee a good choice of network topology. To address these issues, we propose logic shrinkage, a fine-grained netlist pruning methodology enabling K to be automatically learned for every LUT in a neural network targeted for FPGA inference. By removing LUT inputs determined to be of low importance, our method increases the efficiency of the resultant accelerators. Our GPU-friendly solution to LUT input removal is capable of processing large topologies during their training with negligible slowdown. With logic shrinkage, we better the area and energy efficiency of the best-performing LUTNet implementation of the CNV network classifying CIFAR-10 by 1.54x and 1.31x, respectively, while matching its accuracy. This implementation also reaches 2.71x the area efficiency of an equally accurate, heavily pruned BNN. On ImageNet with the Bi-Real Net architecture, employment of logic shrinkage results in a post-synthesis area reduction of 2.67x vs LUTNet, allowing for implementation that was previously impossible on today's largest FPGAs.
Erwei Wang, James J. Davis 0001, Georgios-Ilias Stavrou, Peter Y. K. Cheung, George A. Constantinides, Mohamed S. Abdelfattah
FPGA4
2021 Accelerating Recurrent Neural Networks for Gravitational Wave Experiments
abstract
This paper presents novel reconfigurable architectures for reducing the latency of recurrent neural networks (RNNs) that are used for detecting gravitational waves. Gravitational interferometers such as the LIGO detectors capture cosmic events such as black hole mergers which happen at unknown times and of varying durations, producing time-series data. We have developed a new architecture capable of accelerating RNN inference for analyzing time-series data from LIGO detectors. This architecture is based on optimizing the initiation intervals (II) in a multi-layer LSTM (Long Short-Term Memory) network, by identifying appropriate reuse factors for each layer. A customizable template for this architecture has been designed, which enables the generation of low-latency FPGA designs with efficient resource utilization using high-level synthesis tools. The proposed approach has been evaluated based on two LSTM models, targeting a ZYNQ 7045 FPGA and a U250 FPGA. Experimental results show that with balanced II, the number of DSPs can be reduced up to 42% while achieving the same IIs. When compared to other FPGA-based LSTM designs, our design can achieve about 4.92 to 12.4 times lower latency.
Zhiqiang Que, Erwei Wang, Umar Marikar, Eric A. Moreno, Jennifer Ngadiuba, Hamza Javed, Bartlomiej Borzyszkowski, Thea Aarrestad, Vladimir Loncar, Sioni Summers, Maurizio Pierini, Peter Y. K. Cheung, Wayne Luk
ASAP12
2021 Post-lockdown abatement of COVID-19 by fast periodic switching
abstract
COVID-19 abatement strategies have risks and uncertainties which could lead to repeating waves of infection. We show-as proof of concept grounded on rigorous mathematical evidence-that periodic, high-frequency alternation of into, and out-of, lockdown effectively mitigates second-wave effects, while allowing continued, albeit reduced, economic activity. Periodicity confers (i) predictability, which is essential for economic sustainability, and (ii) robustness, since lockdown periods are not activated by uncertain measurements over short time scales. In turn-while not eliminating the virus-this fast switching policy is sustainable over time, and it mitigates the infection until a vaccine or treatment becomes available, while alleviating the social costs associated with long lockdowns. Typically, the policy might be in the form of 1-day of work followed by 6-days of lockdown every week (or perhaps 2 days working, 5 days off) and it can be modified at a slow-rate based on measurements filtered over longer time scales. Our results highlight the potential efficacy of high frequency switching interventions in post lockdown mitigation. All code is available on Github at https://github.com/V4p1d/FPSP_Covid19. A software tool has also been developed so that interested parties can explore the proof-of-concept system.
Michelangelo Bin, Peter Y. K. Cheung, Emanuele Crisostomi, Pietro Ferraro, Hugo Lhachemi, Roderick Murray-Smith, Connor W. Myant, Thomas Parisini, Robert Shorten, Sebastian Stein 0003, Lewi Stone
PLoS Comput. Biol.2
2020 LUTNet: Learning FPGA Configurations for Highly Efficient Neural Network Inference
abstract
Research has shown that deep neural networks contain significant redundancy, and thus that high classification accuracy can be achieved even when weights and activations are quantized down to binary values. Network binarization on FPGAs greatly increases area efficiency by replacing resource-hungry multipliers with lightweight XNOR gates. However, an FPGA's fundamental building block, the K-LUT, is capable of implementing far more than an XNOR: it can perform any K-input Boolean operation. Inspired by this observation, we propose LUTNet, an end-to-end hardware-software framework for the construction of area-efficient FPGA-based neural network accelerators using the native LUTs as inference operators. We describe the realization of both unrolled and tiled LUTNet architectures, with the latter facilitating smaller, less power-hungry deployment over the former while sacrificing area and energy efficiency along with throughput. For both varieties, we demonstrate that the exploitation of LUT flexibility allows for far heavier pruning than possible in prior works, resulting in significant area savings while achieving comparable accuracy. Against the state-of-the-art binarized neural network implementation, we achieve up to twice the area efficiency for several standard network models when inferencing popular datasets. We also demonstrate that even greater energy efficiency improvements are obtainable.
Erwei Wang, James J. Davis 0001, Peter Y. K. Cheung, George A. Constantinides
IEEE Trans. Computers3
2019 LUTNet: Rethinking Inference in FPGA Soft Logic
abstract
Research has shown that deep neural networks contain significant redundancy, and that high classification accuracies can be achieved even when weights and activations are quantised down to binary values. Network binarisation on FPGAs greatly increases area efficiency by replacing resource-hungry multipliers with lightweight XNOR gates. However, an FPGA's fundamental building block, the K-LUT, is capable of implementing far more than an XNOR: it can perform any K-input Boolean operation. Inspired by this observation, we propose LUTNet, an end-to-end hardware-software framework for the construction of area-efficient FPGA-based neural network accelerators using the native LUTs as inference operators. We demonstrate that the exploitation of LUT flexibility allows for far heavier pruning than possible in prior works, resulting in significant area savings while achieving comparable accuracy. Against the state-of-the-art binarised neural network implementation, we achieve twice the area efficiency for several standard network models when inferencing popular datasets. We also demonstrate that even greater energy efficiency improvements are obtainable.
Erwei Wang, James J. Davis 0001, Peter Y. K. Cheung, George A. Constantinides
FCCM3
2019 Accelerating Position-Aware Top-k ListNet for Ranking Under Custom Precision Regimes
abstract
Document ranking is used to order query results by relevance with ranking models. ListNet is a well-know ranking approach for constructing and training learning to rank models. Compared with traditional learning approaches, ListNet delivers better accuracy, but is computationally too expensive to learn models with large datasets due to the large number of permutations involved in computing the gradients. This paper introduces a position-aware sampling approach, which takes the importance of ranking positions into account and shows better accuracy than previous sampling methods. We also propose an effective quantisation method based on FPGA devices for the ListNet algorithm, which organises the gradient values to several batches, and associates each batch with a different fractional precision. We implemented our approach on a Xilinx Ultrascale+ board and applied it to the MQ 2008 benchmark dataset for ranking. The experiment results show a 4.42x speedup over an Nvidia GTX 1080T GPU implementation with 2% accuracy loss.
Erwei Wang, Shane T. Fleming, David B. Thomas, Peter Y. K. Cheung
FPL5
2018 Hardware Compilation of Deep Neural Networks: An Overview
abstract
Deploying a deep neural network model on a reconfigurable platform, such as an FPGA, is challenging due to the enormous design spaces of both network models and hardware design. A neural network model has various layer types, connection patterns and data representations, and the corresponding implementation can be customised with different architectural and modular parameters. Rather than manually exploring this design space, it is more effective to automate optimisation throughout an end-to-end compilation process. This paper provides an overview of recent literature proposing novel approaches to achieve this aim. We organise materials to mirror a typical compilation flow: front end, platform-independent optimisation and back end. Design templates for neural network accelerators are studied with a specific focus on their derivation methodologies. We also review previous work on network compilation and optimisation for other hardware platforms to gain inspiration regarding FPGA implementation. Finally, we propose some future directions for related research.
Ruizhe Zhao, Shuanglong Liu, Ho-Cheung Ng, Erwei Wang, James J. Davis 0001, Xinyu Niu, Huifeng Shi, George A. Constantinides, Peter Y. K. Cheung, Wayne Luk
ASAP10
2018 A PYNQ-Based Framework for Rapid CNN Prototyping
abstract
This work presents a self-contained and modifiable framework for fast and easy convolutional neural network prototyping on the Xilinx PYNQ platform. With a Python-based programming interface, the framework combines the convenience of high-level abstraction with the speed of optimised FPGA implementation. Our work is freely available on GitHub for the community to use and build upon.
Erwei Wang, James J. Davis 0001, Peter Y. K. Cheung
FCCM3
2018 Accelerating Top-k ListNet Training for Ranking Using FPGA
abstract
Document ranking is used to order query results by relevance, with different document ranking models providing trade-offs between ranking accuracy and training speed. ListNet is a well-known ranking approach which achieves high accuracy, but is infeasible in practice because training time is quadratic in the number of training documents. This paper considers the acceleration of ListNet training using FPGAs, and improves training speed by using hardware-oriented algorithmic optimisations, and by transforming algorithm structures to remove dependencies and expose parallelism. We implemented our approach on a Xilinx ultrascale FPGA board and applied it to the MQ 2008 benchmark dataset for ranking. Compared to existing ranking approaches ours shows an improvement from 0.29 to 0.33 in ranking accuracy on the same dataset using the NDCG@10 metric. Taking into account the communication between software and hardware, we are able to achieve a 3.21x speedup over an Intel Xeon1.6 GHz CPU implementation.
Shane T. Fleming, David B. Thomas, Peter Y. K. Cheung
FPT4
2018 CypherDB: A Novel Architecture for Outsourcing Secure Database Processing
abstract
CypherDB addresses the problem of protecting the confidentiality of database stored externally in a cloud and enabling efficient computation over it to thwart any curious-but-honest cloud computing service provider. It works by encrypting the entire outsourced database and executing queries over the encrypted data using our novel CypherDB secure processor architecture. To optimize computational efficiency, our proposed processor architecture provides tightly-coupled datapaths that avoid information leakage during database access and query execution. Our simulation using a well-known database benchmark TPC-H over a commercial grade Database Management System (SQLite) demonstrates that our proposed architecture incurs an average of about 10 percent overhead when compared with the same set of operations without secure database processing.
Bony H. K. Chen, Paul Y. S. Cheung, Peter Y. K. Cheung, Yu-Kwong Kwok
IEEE Trans. Cloud Comput.3
2018 KAPow: High-Accuracy, Low-Overhead Online Per-Module Power Estimation for FPGA Designs
abstract
In an FPGA system-on-chip design, it is often insufficient to merely assess the power consumption of the entire circuit by compile-time estimation or runtime power measurement. Instead, to make better decisions, one must understand the power consumed by each module in the system. In this work, we combine measurements of register-level switching activity and system-level power to build an adaptive online model that produces live breakdowns of power consumption within the design. Online model refinement avoids time-consuming characterization while also allowing the model to track long-term operating condition changes. Central to our method is an automated flow that selects signals predicted to be indicative of high power consumption, instrumenting them for monitoring. We named this technique KAPow, for ‘K’ounting Activity for Power estimation, which we show to be accurate and to have low overheads across a range of representative benchmarks. We also propose a strategy allowing for the identification and subsequent elimination of counters found to be of low significance at runtime, reducing algorithmic complexity without sacrificing significant accuracy. Finally, we demonstrate an application example in which a module-level power breakdown can be used to determine an efficient mapping of tasks to modules and reduce system-wide power consumption by up to 7%.
James J. Davis 0001, Eddie Hung, Joshua M. Levine, Edward A. Stott, Peter Y. K. Cheung, George A. Constantinides
ACM Trans. Reconfigurable Technol. Syst.5
2017 STRIPE: Signal selection for runtime power estimation
abstract
Knowledge of power consumption at a subsystem level can facilitate adaptive energy-saving techniques such as power gating, runtime task mapping and dynamic voltage and/or frequency scahng. While we have the ability to attribute power to an arbitrary hardware system's modules in real time, the selection of the particular signals to monitor for the purpose of power estimation within any given module has yet to be treated as a primary concern. In this paper, we show how the automatic analysis of circuit structure and behaviour inferred through vectored simulation can be used to produce high-quality rankings of signals' importance, with the resulting selections able to achieve lower power estimation error than those of prior work coupled with decreases in area, power and modelling complexity. In particular, by monitoring just eight signals per module (~0.3% of the total) across the 15 we examined, we demonstrate how to achieve runtime module-level estimation errors 1.5-6.9× lower than when rehant on the signal selections made in accordance with a more straightforward, previously published metric.
James J. Davis 0001, Joshua M. Levine, Edward A. Stott, Eddie Hung, Peter Y. K. Cheung, George A. Constantinides
FPL5
2016 KAPow: A System Identification Approach to Online Per-Module Power Estimation in FPGA Designs
abstract
In a modern FPGA system-on-chip design, it is often insufficient to simply assess the total power consumption of the entire circuit by design-time estimation or runtime power rail measurement. Instead, to make better runtime decisions, it is desirable to understand the power consumed by each individual module in the system. In this work, we combine board-level power measurements with register-level activity counting to build an online model that produces a breakdown of power consumption within the design. Online model refinement avoids the need for a time-consuming characterisation stage and also allows the model to track long-term changes to operating conditions. Our flow is named KAPow, a (loose) acronym for 'K'ounting Activity for Power estimation, which we show to be accurate, with per-module power estimates as close to ±5mW of true measurements, and to have low overheads. We also demonstrate an application example in which a per-module power breakdown can be used to determine an efficient mapping of tasks to modules and reduce system-wide power consumption by over 8%.
Eddie Hung, James J. Davis 0001, Joshua M. Levine, Edward A. Stott, Peter Y. K. Cheung, George A. Constantinides
FCCM5
2016 Increasing Network Size and Training Throughput of FPGA Restricted Boltzmann Machines Using Dropout
abstract
Restricted Boltzmann Machines (RBMs) are widely used in modern machine learning tasks. Existing implementations are limited in network size and training throughput by available DSP resources. In this work we propose a new algorithm and architecture for FPGAs called dropout-RBM (dRBM) system. Compared to the state-of-art design methods on the same FPGA, dRBM with a dropout rate 0.5 doubles the maximum affordable network size using only half of DSP and BRAM resources. This is achieved by an application of a technique called dropout, which is a relatively new method used to avoid overfitting of data. Here we instead apply dropout as a technique for reducing the required DSPs and BRAM resources, while also having the side-effect of increasing robustness of training. Also to improve the processing throughput, we propose a multi-mode matrix multiplication module that maximizes the DSP efficiency. For the MNIST classificationbenchmark, a Stratix IV EP4SGX530 FPGA running dRBM is 34x faster than a single-precision Matlab implementation running on Intel i7 2.9GHz CPU.
Jiang Su, David B. Thomas, Peter Y. K. Cheung
FCCM3
2016 Knowledge is Power: Module-level Sensing for Runtime Optimisation (Abstact Only)
abstract
We propose the compile-time instrumentation of coexisting modules?IP blocks, accelerators, etc.?implemented in FPGAs. The efficient mapping of tasks to execution units can then be achieved, for power and/or timing performance, by tracking dynamic power consumption and/or timing slack online at module-level granularity. Our proposed instrumentation is transparent, thereby not affecting circuit functionality. Power and timing overheads have proven to be small and tend to be outweighed by the exposed runtime benefits.
James J. Davis 0001, Eddie Hung, Joshua M. Levine, Edward A. Stott, Peter Y. K. Cheung, George A. Constantinides
FPGA5
2015 Preface
abstract
The first International Conference on Field-Programmable Logic and Applications (FPL) was held in 1991 at Oxford University. In the ensuing years, it has grown to become the largest conference covering the rapidly growing area of field-programmable logic. Many of the advances achieved in reconfigurable system architectures, applications, embedded processors, design automation methods and tools were first published in the proceedings of the FPL conference series. The objective of FPL is to bring together researchers and practitioners from both academia and industry around the world.
Peter Y. K. Cheung, Wayne Luk, Cristina Silvano
FPL1
2015 An efficient architecture for zero overhead data en-/decryption using reconfigurable cryptographic engine
abstract
Many applications use encryption to protect data confidentiality, which require decryption before any data processing. Integrating ASIC design of encryption engines and general-purpose processor can yield the best overall performance in program execution as it benefits from low latency hardware engine and high processor memory bandwidth. However, ASIC design is fixed once manufactured, which cannot afford any changes in the implemented cryptographic algorithm. FPGA implementation is attractive in terms of its re-configurability but it is generally much slower than ASIC design. In this demo, we present a novel scheme that can offload the latency of reconfigurable cryptographic engine from the overall execution and define an en-/decryption data interface, which is independent of the underlying encryption algorithms. To verify our proposed scheme, we implemented a FPGA prototype, which integrated our design with OpenRISC on ALTERA DE2i-150 evaluation board. We prove that our proposed architecture can flexibly and efficiently en-/decrypt the data with zero overheads towards overall program execution with careful design. Our case study on SQLite shows that the query execution over a 1GB encrypted database on our implemented system introduces performance overhead ranging from 0% to 14%.
Bony H. K. Chen, Paul Y. S. Cheung, Peter Y. K. Cheung, Yu-Kwong Kwok
FPT3
2015 Mapping Adaptive Particle Filters to Heterogeneous Reconfigurable Systems
abstract
This article presents an approach for mapping real-time applications based on particle filters (PFs) to heterogeneous reconfigurable systems, which typically consist of multiple FPGAs and CPUs. A method is proposed to adapt the number of particles dynamically and to utilise runtime reconfigurability of FPGAs for reduced power and energy consumption. A data compression scheme is employed to reduce communication overhead between FPGAs and CPUs. A mobile robot localisation and tracking application is developed to illustrate our approach. Experimental results show that the proposed adaptive PF can reduce up to 99% of computation time. Using runtime reconfiguration, we achieve a 25% to 34% reduction in idle power. A 1U system with four FPGAs is up to 169 times faster than a single-core CPU and 41 times faster than a 1U CPU server with 12 cores. It is also estimated to be 3 times faster than a system with four GPUs.
Thomas C. P. Chau, Xinyu Niu, Alison Eele, Jan M. Maciejowski, Peter Y. K. Cheung, Wayne Luk
ACM Trans. Reconfigurable Technol. Syst.5
2014 Image progressive acquisition for hardware systems
abstract
As the resolution of digital images increases, accessing raw image data from memory has become a major consideration during the design of image/video processing systems. This is due to the fact that the bandwidth requirement and energy consumption of such image accessing process has increased. Inspired by the successful application of progressive image sampling techniques in many image processing tasks, this work proposes to apply similar concept within hardware systems to efficiently trade image quality for reduced memory bandwidth requirement and lower energy consumption. Based on this idea, a hardware system is proposed that is placed between the memory subsystem and the processing core of the design. The proposed system alters the conventional memory access pattern to progressively and adaptively access pixels from a target memory external to the system. The sampled pixels are used to reconstruct an approximation to the ground truth, which is stored in an internal image buffer for further processing. The system is prototyped on FPGA and its performance evaluation shows that a saving of up to 85% of memory accessing time and 33%/45% of image acquisition time/energy is achieved on the benchmark image “lena” while maintaining a PSNR of about 30 dB.
Jianxiong Liu, Christos-Savvas Bouganis, Peter Y. K. Cheung
DATE3
2014 SMCGen: Generating Reconfigurable Design for Sequential Monte Carlo Applications
abstract
The Sequential Monte Carlo (SMC) method is a simulation-based approach to compute posterior distributions. SMC methods often work well on applications considered intractable by other methods due to high dimensionality, but they are computationally demanding. While SMC has been implemented efficiently on FPGAs, design productivity remains a challenge. This paper introduces a design flow for generating efficient implementation of reconfigurable SMC designs. Through templating the SMC structure, the design flow enables efficient mapping of SMC applications to multiple FPGAs. The proposed design flow consists of a parametrisable SMC computation engine, and an open-source software template which enables efficient mapping of a variety of SMC designs to reconfigurable hardware. Design parameters that are critical to the performance and to the solution quality are tuned using a machine learning algorithm based on surrogate modelling. Experimental results for three case studies show that design performance is substantially improved after parameter optimisation. The proposed design flow demonstrates its capability of producing reconfigurable implementations for a range of SMC applications that have significant improvement in speed and in energy efficiency over optimised CPU and GPU implementations.
Thomas C. P. Chau, Maciej Kurek, James Stanley Targett, Jake Humphrey, Georgios Skouroupathis, Alison Eele, Jan M. Maciejowski, Benjamin Cope, Kathryn Cobden, Philip H. W. Leong, Peter Y. K. Cheung, Wayne Luk
FCCM11
2014 Reducing Overheads for Fault-Tolerant Datapaths with Dynamic Partial Reconfiguration
abstract
As process scaling and transistor count inflation continue, silicon chips are becoming increasingly susceptible to faults. Although FPGAs are particularly vulnerable to these effects, their runtime reconfigurability offers unique opportunities for fault tolerance. This work presents an application combining algorithmic-level error detection with dynamic partial reconfiguration (DPR) to allow faults manifested within its datapath at runtime to be circumvented at low cost.
James J. Davis 0001, Peter Y. K. Cheung
FCCM2
2014 Timing Fault Detection in FPGA-Based Circuits
abstract
The operation of FPGA systems, like most VLSI technology, is traditionally governed by static timing analysis, whereby safety margins for operating and manufacturing uncertainty are factored in at design-time. If we operate FPGA designs beyond these conservative margins we can obtain substantial energy and performance improvements. However, doing this carelessly would cause unacceptable impacts to reliability, lifespan and yield - issues which are growing more severe with continuing process scaling. Fortunately, the flexibility of FPGA architecture allows us to monitor and control reliability problems with a variety of runtime instrumentation and adaptation techniques. In this paper we develop a system for detecting timing faults in arbitrary FPGA circuits based on Razor-like shadow register insertion. Through a combination of calibration, timing constraint and adaptation of the CAD flow, we deliver low-overhead, trustworthy fault detection for FPGA-based circuits.
Edward A. Stott, Joshua M. Levine, Peter Y. K. Cheung, Nachiket Kapre
FCCM3
2014 Dynamic voltage & frequency scaling with online slack measurement
abstract
Timing margins in FPGAs are already significant and as process scaling continues they will have to grow to guarantee operation under increased variation. Margins enforce worst-case operation even in typical conditions and result in devices operating more slowly and consuming more energy than necessary. This paper presents a method of dynamic voltage and frequency scaling that uses online slack measurement to determine timing headroom in a circuit while it is operating and scale the voltage and/or frequency in response. Doing so can significantly reduce power consumption or increase throughput with a minimal overhead. The method is demonstrated on a number of benchmark circuits under a range of operating conditions, constraints and optimisation targets.
Joshua M. Levine, Edward A. Stott, Peter Y. K. Cheung
FPGA3
2014 Achieving low-overhead fault tolerance for parallel accelerators with dynamic partial reconfiguration
abstract
While allowing for the fabrication of increasingly complex and efficient circuitry, transistor shrinkage and count-per-device expansion have major downsides: chiefly increased variation, degradation and fault susceptibility. For this reason, design-time consideration of fault tolerance will have to be given to increasing numbers of electronic systems in the future to ensure yields, reliabilities and lifetimes remain acceptably high. Many commonly implemented operators are suited to modification resulting in datapath error detection capabilities with low area overheads. FPGAs are uniquely placed to allow further area savings to be made when incorporating fault avoidance mechanisms thanks to their dynamic reconfigurability. In this paper, we examine the practicalities and costs involved in implementing hardware-software fault tolerance on a test platform: a parallel matrix multiplication accelerator in hardware, with controller in software, running on a Xilinx Zynq system-on-chip. A combination of `bolt-on' error detection logic and software-triggered routing reconfiguration serve to provide low-overhead datapath fault tolerance at runtime. Rapid yet accurate fault diagnoses along with low hardware (area), software (configuration storage) and performance penalties are achieved.
James J. Davis 0001, Peter Y. K. Cheung
FPL2
2013 A variation-adaptive retiming method exploiting reconfigurability
abstract
In this article we present a variation-aware post placement and routing (P&R) retiming method to counteract process variation in FPGAs. Variation-aware retiming takes into account exact variation maps (measured on FPGAs) as opposed to statistical static timing analysis (SSTA) which models process variation with statistical distributions. Experiments are conducted using variation maps measured from 100 Cyclone III FPGAs, and the retiming algorithm is applied using MATLAB. We have shown that for circuits with several retiming choices of equivalent logic depth, up to 30% delay improvement can be achieved for a given variation coefficient of σ/μ = 0.3.
Justin S. J. Wong, Sumanta Chaudhuri, George A. Constantinides, Peter Y. K. Cheung
FPL5
2013 SMI: Slack Measurement Insertion for online timing monitoring in FPGAs
abstract
Shadow registers, driven by a variable-phase clock, can be used to extract useful timing information from a circuit during operation. This paper presents Slack Measurement Insertion (SMI), an automated tool flow for inserting shadow registers into an FPGA design to enable measurement of timing slack. The flow provides a parameterised level of circuit coverage and results in minimal timing and area overheads. We demonstrate the process through its application to three complex benchmark designs.
Joshua M. Levine, Edward A. Stott, George A. Constantinides, Peter Y. K. Cheung
FPL4
2013 Acceleration of real-time Proximity Query for dynamic active constraints
abstract
Proximity Query (PQ) is a process to calculate the relative placement of objects. It is a critical task for many applications such as robot motion planning, but it is often too computationally demanding for real-time applications, particularly those involving human-robot collaborative control. This paper derives a PQ formulation which can support non-convex objects represented by meshes or cloud points. We optimise the proposed PQ for reconfigurable hardware by function transformation and reduced precision, resulting in a novel data structure and memory architecture for data streaming while maintaining the accuracy of results. Run-time reconfiguration is adopted for dynamic precision optimisation. Experimental results show that our optimised PQ implementation on a reconfigurable platform with four FPGAs is 58 times faster than an optimised CPU implementation with 12 cores, 9 times faster than a GPU, and 3 times faster than a double precision implementation with four FPGAs.
Thomas C. P. Chau, Ka-Wai Kwok, Gary C. T. Chow, Kuen Hung Tsoi, Kit-Hang Lee, Zion Tsz Ho Tse, Peter Y. K. Cheung, Wayne Luk
FPT7
2013 Datapath fault tolerance for parallel accelerators
abstract
While we reap the benefits of process scaling in terms of transistor density and switching speed, consideration must be given to the negative effects it causes: increased variation, degradation and fault susceptibility. Above device level, such phenomena and the faults they induce can lead to reduced yield, decreased system reliability and, in extreme cases, total failure after a period of successful operation. Although error detection and correction are almost always considered for highly sensitive and susceptible applications such as those in space, for other, more general-purpose applications they are often overlooked. In this paper, we present a parallel matrix multiplication accelerator running in hardware on the Xilinx Zynq system-on-chip platform, along with ‘bolt-on’ logic for detecting, locating and avoiding faults within its datapath. Designs of various sizes are compared with respect to resource overhead and performance impact. Our largest-implemented fault-tolerant accelerator was found to consume 17.3% more area, run at a 3.95% lower frequency and incur an 18.8% execution time penalty over its equivalent fault-susceptible design during fault-free operation.
James J. Davis 0001, Peter Y. K. Cheung
FPT2
2013 Exploiting stochastic delay variability on FPGAs with adaptive partial rerouting
abstract
Aggressive transistor scaling will soon lead us to the physical upper-bound of process technology, where stochastic process variability dominates the timing performance of FPGA components. In this paper, a variation-aware partial-rerouting method is proposed to mitigate and take advantage of the effect of delay variability due to process variation. The variation in logic delay across each FPGA (variation map) is measured on commercial FPGAs and is used to assess the effectiveness and potential gain of the proposed method on current FPGA architectures. Our partial-rerouting method achieved 5.25% improvement in critical path delay under a delay variability of σ/μ = 0.3, and is considerably less time consuming than using variation-aware full chipwise routing, which gave a slightly better timing gain of 6.41% but requires 8x more execution time when optimising for 100 target FPGAs with unique variation maps.
Justin S. J. Wong, Sumanta Chaudhuri, George A. Constantinides, Peter Y. K. Cheung
FPT5
2013 High-level power and performance estimation of FPGA-based soft processors and its application to design space exploration
Adam Powell, Christos-Savvas Bouganis, Peter Y. K. Cheung
J. Syst. Archit.3
2013 Timing Measurement Platform for Arbitrary Black-Box Circuits Based on Transition Probability
abstract
The key aspects of a good on-chip timing measurement platform are high measurement resolution, accuracy, and low area overhead. A measurement method based on transition probability (TP) has shown promising characteristics in all these areas. In this paper, the TP measurement method is examined through simulation to understand its apparent effectiveness and accuracy in measuring complex circuits. Timing uncertainties and logic glitch activities are considered in detail, and the effect of varying input vectors' probability distributions is analyzed to enable further accuracy improvements. Using a field-programmable gate array, the method is implemented and demonstrated as a modular on-chip test platform for testing complex arbitrary circuits. Practical circuits found in typical modular designs, including fixed/floating-point arithmetic and filter circuits, are chosen to evaluate the test platform. The resolution of the timing measurements ranges from 0.3 to 8.0 ps, and the measurement errors against reference measurements are found to be within 3.6%. The test platform can be applied to VLSI designs with minor area overhead, and provides designers with precise and accurate physical timing information of circuits.
Justin S. J. Wong, Peter Y. K. Cheung
IEEE Trans. Very Large Scale Integr. Syst.2
2012 Online Measurement of Timing in Circuits: For Health Monitoring and Dynamic Voltage & Frequency Scaling
abstract
Reliability, power consumption and timing performance are key considerations for the utilisation of field-programmable gate arrays. Online measurement techniques can determine the timing characteristics of an FPGA application while it is operating, and facilitate a range of benefits. Degradation can be monitored by tracking changes in timing performance, while power consumption can be reduced through dynamic voltage scaling (DVS) of the power supply to exploit any spare timing headroom. If higher performance is the objective, dynamic frequency scaling (DFS) can be used to maximise operating frequency. In both cases, online timing measurement of the application circuit is used to exploit favourable operating conditions. This work demonstrates a method of online measurement, achieved by sweeping the phase of a secondary clock signal, driving additional shadowing registers strategically added to the application design. The measurement technique and initial voltage and frequency scaling experiments are demonstrated on an Alter a Cyclone III FPGA. Timing performance can be measured with a best case resolution of 96ps. The additional circuitry results in minimal overhead in terms of area and performance. Power savings of 23% dynamic and 13% static in an example circuit are achieved through DVS, or performance improvements of 21% through DFS, when compared with operating at nominal core voltage, or timing model FMax.
Joshua M. Levine, Edward A. Stott, George A. Constantinides, Peter Y. K. Cheung
FCCM4
2012 Adaptive Sequential Monte Carlo approach for real-time applications
abstract
This paper presents an adaptive Sequential Monte Carlo approach for real-time applications. Sequential Monte Carlo method is employed to estimate the states of dynamic systems using weighted particles. The proposed approach reduces the run-time computation complexity by adapting the size of the particle set. Multiple processing elements on FPGAs are dynamically allocated for improved energy efficiency without violating real-time constraints. A robot localisation application is developed based on the proposed approach. Compared to a non-adaptive implementation, the dynamic energy consumption is reduced by up to 70% without affecting the quality of solutions.
Thomas C. P. Chau, Wayne Luk, Peter Y. K. Cheung, Alison Eele, Jan M. Maciejowski
FPL3
2012 A two-stage variation-aware placement method for FPGAS exploiting variation maps classification
abstract
Technology scaling causes increasing and unavoidable delay variability in FPGAs. This paper proposes a 2-stage variation-aware placement method that benefits from the optimality of a full-chipwise (chip-by-chip) placement but only requires a fraction of total execution time for a large number of FPGAs with different variation patterns. By classifying variation maps into finite number of classes, variation-aware placement only need to be executed based on the median map of each class to produce the placement for the other FPGAs (variation maps) in that class to save execution time. Our proposed method is implemented in a modified version of VPR 5.0 and verified using variation maps measured from 129 DE0 boards equipped with Cyclone III FPGAs. The mean timing gain of 7.36% is observed in 20 MCNC benchmarks with 16 clusters, while reducing execution time by a factor of 8 compared to full-chipwise placement.
Justin S. J. Wong, Sumanta Chaudhuri, George A. Constantinides, Peter Y. K. Cheung
FPL5
2012 Early performance estimation of image compression methods on soft processors
abstract
This paper presents a power and execution time estimation framework for an FPGA-based soft processor when considering the implementation of image compression techniques. Using the proposed framework, a quick power consumption and execution time estimate can be obtained early in the design phase allowing system designers to estimate these performance metrics without the need of implementing the algorithm or generating all possible soft processor architectures. This estimate is performed using both high-level algorithm parameters and soft processor architecture parameters. For system designers this can result in fast design space exploration. The model can predict the execution time of an algorithm with an average of 139% less relative error than predictions using only architecture parameters with the same framework.
Adam Powell, Christos-Savvas Bouganis, Peter Y. K. Cheung
FPL3
2011 Health monitoring of live circuits in FPGAs based on time delay measurement (abstract only)
abstract
Literature suggests that timing performance degradation in VLSI could be a major concern in future process technologies. FPGAs are well suited to cope with this challenge, due to their flexibility at design-, manufacture- and run-time.
Joshua M. Levine, Edward A. Stott, George A. Constantinides, Peter Y. K. Cheung
FPGA4
2011 Improved delay measurement method in FPGA based on transition probability
abstract
The ability to measure delay of arbitrary circuits on FPGA offers many opportunities for on-chip characterisation and optimisation. This paper describes an improved delay measurement method by monitoring the transition probability at the output nodes as the operating frequency is swept.
Justin S. J. Wong, Peter Y. K. Cheung
FPGA2
2011 Improving FPGA Reliability with Wear-Levelling
abstract
As VLSI circuits achieve smaller geometries, reliability is becoming an growing problem. The flexibility of FPGAs enables novel techniques for meeting this challenge, and one such technique is wear-levelling: periodic reconfiguration to eliminate electrical stress hotspots. In this work we have have carried out accelerated-life experiments in FPGAs to assess the feasibility of three wear-levelling strategies for reducing timing degradation. All three techniques resulted in significant improvements to robustness compared with a static configuration, and we have demonstrated that wear-levelling is a promising tool for improving FPGA reliability.
Edward A. Stott, Peter Y. K. Cheung
FPL2
2011 Timing speculation in FPGAs: Probabilistic inference of data dependent failure rates
abstract
The goal of this work is to model and predict timing failure rates of digital blocks in FPGAs due to delay variation. The timing failures are modelled as transient faults depending on the input vectors to the block, as opposed to conventional critical path failures. Firstly, we present transition tables derived from the truth tables of the logic gates which are apt for our purpose. Next we present a model of transient circuit behaviour based on Bayesian Networks and a method to infer timing failure rates. We implemented these methods with the Bayes Network Toolbox in MATLAB, and compared our results with Monte-Carlo simulations based on SAE-J2748 package for VHDL, and actual implementations and measurements on a Cyclone III FPGA. The test cases are simple arithmetic circuits, such as adders and multipliers. We show that, with this method comparison of error rates is possible among various implementations. We conclude with emphasis on the need for such a method for advancing probabilistic computing paradigms.
Sumanta Chaudhuri, Justin S. J. Wong, Peter Y. K. Cheung
FPT3
2011 Compiling C-like Languages to FPGA Hardware: Some Novel Approaches Targeting Data Memory Organization
abstract
This paper describes our approaches to raise the level of abstraction at which hardware suitable for accelerating computationally intensive applications can be specified. Field-programmable gate arrays are becoming adopted as a computational platform by the high-performance computing community, but there are challenges to extract maximum performance from these devices. Unlike other approaches, our focus is on data memory organization and input–output bandwidth considerations, which are the typical stumbling block of existing hardware compilation schemes. We describe our approaches, which are based on formal optimization techniques, and present some results showing the advantage of exposing the interaction between data memory system design and parallelism extraction to the compiler.
Qiang Liu 0011, George A. Constantinides, Kostas Masselos, Peter Y. K. Cheung
Comput. J.4
2011 Introduction to special section FPGA 2009
abstract
No abstract available.
Peter Y. K. Cheung
ACM Trans. Reconfigurable Technol. Syst.1
2010 Exploration of hardware sharing for image encoders
abstract
Hardware sharing can be used to reduce the area and the power dissipation of a design. This is of particular interest in the field of image and video compression, where an encoder must deal with different design tradeoffs depending on the characteristics of the signal to be encoded and the constraints imposed by the users. This paper introduces a novel methodology for exploring the design space based on the amount of hardware sharing between different functional blocks, giving as a result a set of feasible solutions which are broad in terms of hardware cost and throughput capabilities. The proposed approach, inspired by the notion of a partition in set theory, has been applied to optimize and to evaluate the sharing alternatives of a group of image and video compression key computational kernels when mapped onto a Xilinx Virtex-5 FPGA.
Sebastián López, Roberto Sarmiento, Philip G. Potter, Wayne Luk, Peter Y. K. Cheung
DATE5
2010 Energy-Aware Optimisation for Run-Time Reconfiguration
abstract
Run-time reconfiguration has been shown to produce power and energy efficient designs. However, it is important to take into account the energy overhead of the reconfiguration process itself. This paper presents an analytical model that covers the effects of power consumption and configuration speed of the reconfiguration process. Based on this model, a method is introduced that establishes the optimal degree of parallelism for designs supporting partial run-time reconfiguration. Our energy-aware approach is illustrated by optimising designs for software-defined radio: a reconfigurable FIR filter is shown to be up to 49% more energy efficient and up to 87% more area efficient than a non-reconfigurable design.
Tobias Becker, Wayne Luk, Peter Y. K. Cheung
FCCM3
2010 Degradation in FPGAs: measurement and modelling
abstract
Progress in VLSI technology is driven by increasing circuit density through process scaling, but with shrinking geometry comes an increasing threat to reliability. FPGAs are uniquely placed to tackle degradation and faults due to their regular structure and ability to reconfigure, giving them the potential to implement system-level reliability enhancements. To assess the scale of the challenge, a method for measuring and monitoring degradation in an FPGA was developed and used to conduct an accelerated life test on a modern device. This revealed a clear, gradual degradation in timing performance that matches the expected effects of Negative-Bias Temperature Instability and Hot Carrier Injection, two of the most important VLSI degradation mechanisms. Further insight into ageing phenomena was gained using modelling -- showing how degradation in a typical LUT would be affected by different usage conditions, and predicting in detail the effects on circuit behaviour.
Edward A. Stott, Justin S. J. Wong, N. Pete Sedcole, Peter Y. K. Cheung
FPGA4
2010 GPU Versus FPGA for High Productivity Computing
abstract
Heterogeneous or co-processor architectures are becoming an important component of high productivity computing systems (HPCS). In this work the performance of a GPU based HPCS is compared with the performance of a commercially available FPGA based HPC. Contrary to previous approaches that focussed on specific examples, a broader analysis is performed by considering processes at an architectural level. A set of benchmarks is employed that use different process architectures in order to exploit the benefits of each technology. These include the asynchronous pipelines common to map tasks, a partially synchronous tree common to reduce tasks and a fully synchronous, fully connected mesh. We show that the GPU is more productive than the FPGA architecture for most of the benchmarks and conclude that FPGA-based HPCS is being marginalised by GPUs.
David Huw Jones, Adam Powell, Christos-Savvas Bouganis, Peter Y. K. Cheung
FPL4
2010 Degradation Analysis and Mitigation in FPGAs
abstract
FPGAs are powerful platforms for investigating impending challenges associated with process scaling, such as variation and degradation. Their versatility allows us to gather empirical data and evaluate novel solutions. We carried out accelerated-life tests on modern FPGA devices and obtained a useful characterisation of the ageing processes that afflict them. We also quantified the potential benefits of three degradation mitigation strategies based on exploiting spare logic and interconnect resources. The work helps cement the role of reconfigurable logic as a vitally-important technology in the face of the uncertainties of future process scaling.
Edward A. Stott, Justin S. J. Wong, Peter Y. K. Cheung
FPL3
2010 A Salient Region Detector for GPU Using a Cellular Automata Architecture
David Huw Jones, Adam Powell, Christos-Savvas Bouganis, Peter Y. K. Cheung
ICONIP (2)4
2010 Wave-pipelined intra-chip signaling for on-FPGA communications
Terrence S. T. Mak, N. Pete Sedcole, Peter Y. K. Cheung, Wayne Luk
Integr.3
2010 Performance Comparison of Graphics Processors to Reconfigurable Logic: A Case Study
abstract
A systematic approach to the comparison of the graphics processor (GPU) and reconfigurable logic is defined in terms of three throughput drivers. The approach is applied to five case study algorithms, characterized by their arithmetic complexity, memory access requirements, and data dependence, and two target devices: the nVidia GeForce 7900 GTX GPU and a Xilinx Virtex-4 field programmable gate array (FPGA). Two orders of magnitude speedup, over a general-purpose processor, is observed for each device for arithmetic intensive algorithms. An FPGA is superior, over a GPU, for algorithms requiring large numbers of regular memory accesses, while the GPU is superior for algorithms with variable data reuse. In the presence of data dependence, the implementation of a customized data path in an FPGA exceeds GPU performance by up to eight times. The trends of the analysis to newer and future technologies are analyzed.
Benjamin Cope, Peter Y. K. Cheung, Wayne Luk, Lee W. Howes
IEEE Trans. Computers2
2010 FPGA Architecture Optimization Using Geometric Programming
abstract
This paper is concerned with the application of geometric programming to the design of homogeneous field programmable gate array (FPGA) architectures. The paper builds on an increasing body of work concerned with modeling reconfigurable architectures, and presents a full area and delay model of an FPGA. We use a geometric programming framework to show how transistor sizing and high-level architecture parameter selection can now be solved as a concurrent optimization problem. We validate the model through the use of simulation program with integrated circuit emphasis (SPICE) models and the versatile place and route (VPR) FPGA architecture simulation tool. Not only does the optimization framework allow architectures to be optimized orders of magnitude faster than previous work, but the combined optimization can lead to different architectural conclusions compared to conventional methods by exploring the coupling between the two sets of optimization variables. Specifically, we show that as delay takes more significance in the objective of the optimization, there should be more lookup tables in a logic block, whereas conventional techniques suggest that there should be fewer lookup tables in an FPGA logic block.
Alastair M. Smith, George A. Constantinides, Peter Y. K. Cheung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2010 Benchmarking and evaluating reconfigurable architectures targeting the mobile domain
abstract
We present the GroundHog 2009 benchmarking suite that evaluates the power consumption of reconfigurable technology for applications targeting the mobile computing domain. This benchmark suite includes seven designs; one design targets fine-grained FPGA fabrics allowing for quick state-of-the-art evaluation, and six designs are specified at a high level allowing them to target a range of existing and future reconfigurable technologies. Each of the six designs can be stimulated with the help of synthetically generated input stimuli created by an open-source tool included in the downloadable suite. Another tool is included to help verify the correctness of each implemented design. To demonstrate the potential of this benchmark suite, we evaluate the power consumption of two modern industrial FPGAs targeting the mobile domain. Also, we show how an academic FPGA framework, VPR 5.0, that has been updated for power estimates can be used to estimates the power consumption of different FPGA architectures and an open-source CAD flow mapping to these architectures.
Peter Jamieson, Tobias Becker, Peter Y. K. Cheung, Wayne Luk, Tero Rissa, Teemu Pitkänen
ACM Trans. Design Autom. Electr. Syst.3
2010 Efficient Heterogeneous Architecture Floorplan Optimization using Analytical Methods
abstract
This paper argues the case for the use of analytical models in FPGA architecture exploration. We show that the problem, when simplified, is amenable to formal optimization techniques such as integer linear programming. However, the simplification process may lead to inaccurate models. To test the overall methodology, we feed the resulting architectures to VPR 5.0 and quantify their performance in comparison with traditional design methodologies. Our results show that the resulting architectures are better than those found using parameter sweep techniques. In addition, we show that these architectures can be further improved by combining the accuracy of VPR 5.0 with the efficiency of analytical techniques. This is achieved using a closed loop framework which iteratively refines the analytical model using the place and route outputs from VPR.
Asma Kahoul, Alastair M. Smith, George A. Constantinides, Peter Y. K. Cheung
ACM Trans. Reconfigurable Technol. Syst.4
2010 An Automated Flow for Arithmetic Component Generation in Field-Programmable Gate Arrays
abstract
State-of-the-art configurable logic platforms, such as Field-Programmable Gate Arrays (FPGAs), consist of a heterogeneous mixture of different component types. Compared to traditional homogeneous configurable platforms, heterogeneity provides speed and density advantages. This is due to the replacement of inefficient programmable logic and routing with specialized logic and fixed interconnect in components such as memories, embedded processor units, and fused arithmetic units. Given the increasing complexity of these components, this article introduces a method to automatically propose and explore the benefits of different types of fused arithmetic units. The methods are based on common subgraph extraction techniques, meaning that it is possible to explore different subcircuits that occur frequently across a set of benchmarks. A quantitative analysis is performed of the various fused arithmetic circuits identified by our tool, which are then automatically synthesized to an ASIC process, providing a study of the speed and area benefits of the components. The results of this study provide bounds on the performance of heterogeneous FPGAs: by incorporating coarse-grain components which match the specific needs of a set of benchmarks we show that significant improvements in circuit speed and area can be made.
Alastair M. Smith, George A. Constantinides, Peter Y. K. Cheung
ACM Trans. Reconfigurable Technol. Syst.3
2010 Exploration of Heterogeneous FPGAs for Mapping Linear Projection Designs
abstract
In many applications, a reduction of the amount of the original data or a representation of the original data by a small set of variables is often required. Among many techniques, the linear projection is often chosen due to its computational attractiveness and good performance. For applications where real-time performance and flexibility to accommodate new data are required, the linear projection is implemented in field-programmable gate arrays (FPGAs) due to their fine-grain parallelism and reconfigurability properties. Currently, the optimization of such a design is considered as a separate problem from the basis calculation leading to suboptimal solutions. In this paper, we propose a novel approach that couples the calculation of the linear projection basis, the area optimization problem, and the heterogeneity exploration of modern FPGAs. The power of the proposed framework is based on the flexibility to insert information regarding the implementation requirements of the linear basis by assigning a proper prior distribution to the basis matrix. Results from real-life examples on modern FPGA devices demonstrate the effectiveness of our approach, where up to 48% reduction in the required area is achieved compared to the current approach, without any loss in the accuracy or throughput of the design.
Christos-Savvas Bouganis, Iosifina Pournara, Peter Y. K. Cheung
IEEE Trans. Very Large Scale Integr. Syst.3
2009 Partition-based exploration for reconfigurable JPEG designs
abstract
This paper proposes a novel approach for design space exploration by characterizing hardware sharing based on the notion of a partition in set theory. Related designs with different degrees of hardware sharing can be captured concisely by a Hasse diagram, highlighting designs with shared building blocks. Hardware sharing can be implemented in various ways, such as component multiplexing, instruction-set processors, or run-time reconfiguration. We illustrate how the proposed approach can be applied to exploring the design space for FPGA implementations of JPEG image compression.
Philip G. Potter, Wayne Luk, Peter Y. K. Cheung
DATE3
2009 Benchmarking Reconfigurable Architectures in the Mobile Domain
abstract
In this paper, we introduce GroundHog 2009 benchmarking suite that can be used to evaluate the power consumption of reconfigurable technology implementing applications targeting the mobile computing domain. This benchmark suite includes seven designs; one design targets fine-grained FPGA fabrics, and six designs are specified at a high level, which allows them to target a range of reconfigurable technologies. Each of the six designs can be stimulated with synthetically generated input stimuli created by a tool included in the suite. Additionally, another tool can help verify the correctness of each implemented design. Finally,we use our benchmark suite to evaluate the power consumption of two modern FPGAs targeting the mobile domain.
Peter Jamieson, Tobias Becker, Wayne Luk, Peter Y. K. Cheung, Tero Rissa, Teemu Pitkänen
FCCM4
2009 Compensating for variability in FPGAs by re-mapping and re-placement
abstract
Two complementary techniques for reducing the effect of within-die variability on the critical path delay in FPGA circuits are reported. The first technique selects the best LUT mapping from a set of alternative mappings of a logic function for each LUT cluster in the FPGA. The second selects the best assignment of LUTs to physical locations within a cluster. The techniques can be used together, and are shown in Monte Carlo experiments to reduce both the mean and standard deviation of critical path delay.
N. Pete Sedcole, Edward A. Stott, Peter Y. K. Cheung
FPL3
2009 Area estimation and optimisation of FPGA routing fabrics
abstract
This paper presents a methodology for estimating and optimising FPGA routing fabrics using high-level modelling and convex optimisation techniques. Experimental methods for exploring design spaces suffer from expensive computation time, which is exacerbated by increased dimensionality due to the larger number of architectural parameters. In this paper we build on previously published work to describe a model of FPGA routing area. This model is used in conjunction with a form of optimisation known as geometric programming, in order to analytically derive optimised FPGA architectural parameters, demonstrating the power and accuracy of model-based approaches in configurable architecture design. We show that routing parameters such as connection and switch box flexibilities can be architected to save around 6% of area instead of using traditional ldquorules of thumbrdquo.
Alastair M. Smith, George A. Constantinides, Peter Y. K. Cheung
FPL3
2009 Concurrently optimizing FPGA architecture parameters and transistor sizing: Implications for FPGA design
abstract
This paper presents a method that combines high-level and low-level architecture parameter exploration. The paper builds on an increasing body of work concerned with modeling reconfigurable architectures, and presents a full area and delay model of an FPGA. The optimization of this model is based on the use of geometric programming, and allows high-level architecture parameter selection and transistor sizing to be done concurrently. We use the framework to demonstrate that concurrent optimization of both high and low-level parameters can lead to significantly different architectural conclusions.
Alastair M. Smith, George A. Constantinides, Steve Wilton, Peter Y. K. Cheung
FPT4
2009 A sensor-based approach to linear blur identification for real-time video enhancement
abstract
Super-resolution (SR) methods are largely affected by the accurate evaluation of the Point Spread Function (PSF) that is related to the input frames. When the frames are degraded by heavy motion blur, the PSFs are highly non-isotropic, which further complicates their estimation. The ill-posed nature of blur identification is usually addressed using the assumption of linear and uniform motion. However, in real-life systems, this may deviate significantly from the actual motion blur. To resolve the above, this work proposes combining a scheme that validates the initial motion assumption with the real-time reconfiguration property of an adaptive image sensor. If the linearity and uniformity assumption is invalid for a given motion region, the sensor is locally reconfigured to larger pixels that produce higher frame-rate samples with reduced blur. Once the appropriate configuration that gives rise to a valid motion assumption is applied, highly accurate PSFs are estimated, resulting to an improved SR reconstruction quality.
Maria E. Angelopoulou, Christos-Savvas Bouganis, Peter Y. K. Cheung
ICIP3
2009 Throughput Maximization for Wave-pipelined Interconnects using Cascaded Buffers and Transistor Sizing
abstract
This paper presents two new design methodologies for throughput-centric wave-pipelined interconnects: cascaded buffers insertion and transistor sizing. Experimental results show that up to 185% throughput improvement can be achieved by applying the new proposed approaches compared with conventional interconnect optimization techniques, such as buffer insertion. Moreover, with the combination of cascaded buffers insertion and adequate techniques in supply voltage scaling, up to 60% dynamic power reduction can be gained compared to the conventional design.
Terrence S. T. Mak, N. Pete Sedcole, Peter Y. K. Cheung
ISCAS4
2009 Combining Data Reuse With Data-Level Parallelization for FPGA-Targeted Hardware Compilation: A Geometric Programming Framework
abstract
A nonlinear optimization framework is proposed in this paper to automate exploration of the design space consisting of data-reuse (buffering) decisions and loop-level parallelization, in the context of field-programmable-gate-array-targeted hardware compilation. Buffering frequently accessed data in on-chip memories can reduce off-chip memory accesses and open avenues for parallelization. However, the exploitation of both data reuse and parallelization is limited by the memory resources available on-chip. As a result, considering these two problems separately, e.g., first exploring data reuse and then exploring data-level parallelization, based on the data-reuse options determined in the first step, may not yield the performance-optimal designs for limited on-chip memory resources. We consider both problems at the same time, exposing the dependence between the two. We show that this combined problem can be formulated as a nonlinear program and further show that efficient solution techniques exist for this problem, based on recent advances in optimization of so-calledgeometricprogrammingproblems. The results from applying this framework to several real benchmarks implemented on a Xilinx device demonstrate that given different constraints on on-chip memory utilization, the corresponding performance-optimal designs are automatically determined by the framework. We have also implemented designs determined by a two-stage optimization method that first explores data reuse and then explores parallelization on the same platform, and by comparison, the performance-optimal designs proposed by our framework are faster than the designs determined by the two-stage method by up to 5.7 times.
Qiang Liu 0011, George A. Constantinides, Kostas Masselos, Peter Y. K. Cheung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2009 Word-length selection for power minimization via nonlinear optimization
abstract
This article describes the first method for minimizing the dynamic power consumption of a Digital Signal Processing (DSP) algorithm implemented on reconfigurable hardware via word-length optimization. Fast models for estimating the power consumption of the arithmetic components and the routing power of these algorithm implementations are used within a constrained nonlinear optimization formulation that solves a relaxed version of word-length optimization. Tight lower and upper bounds on the cost of the integer word-length problem can be obtained using the proposed solution, with typical upper bounds being 2.9% and 5.1% larger than the lower bounds for area and power consumption, respectively. Heuristics can then use the upper bound as a starting point from which to get even closer to the known lower bound. Results show that power consumption can be improved by up to 40% compared to that achieved when using simple word-length selection techniques, and further comparisons are made between the minimization of different cost functions that give insight into the advantages offered by multiple word-length optimization.
Jonathan A. Clarke, George A. Constantinides, Peter Y. K. Cheung
ACM Trans. Design Autom. Electr. Syst.3
2009 Robust Real-Time Super-Resolution on FPGA and an Application to Video Enhancement
abstract
The high density image sensors of state-of-the-art imaging systems provide outputs with high spatial resolution, but require long exposure times. This limits their applicability, due to the motion blur effect. Recent technological advances have lead to adaptive image sensors that can combine several pixels together in real time to form a larger pixel. Larger pixels require shorter exposure times and produce high-frame-rate samples with reduced motion blur. This work proposes combining an FPGA with an adaptive image sensor to produce an output of high resolution both in space and time. The FPGA is responsible for the spatial resolution enhancement of the high-frame-rate samples using super-resolution (SR) techniques in real time. To achieve it, this article proposes utilizing the Iterative Back Projection (IBP) SR algorithm. The original IBP method is modified to account for the presence of noise, leading to an algorithm more robust to noise. An FPGA implementation of this algorithm is presented. The proposed architecture can serve as a general purpose real-time resolution enhancement system, and its performance is evaluated under various noise levels.
Maria E. Angelopoulou, Christos-Savvas Bouganis, Peter Y. K. Cheung, George A. Constantinides
ACM Trans. Reconfigurable Technol. Syst.3
2009 Synthesis and Optimization of 2D Filter Designs for Heterogeneous FPGAs
abstract
Many image processing applications require fast convolution of an image with one or more 2D filters. Field-Programmable Gate Arrays (FPGAs) are often used to achieve this goal due to their fine grain parallelism and reconfigurability. However, the heterogeneous nature of modern reconfigurable devices is not usually considered during design optimization. This article proposes an algorithm that explores the space of possible implementation architectures of 2D filters, targeting the minimization of the required area, by optimizing the usage of the different components in a heterogeneous device. This is achieved by exploring the heterogeneous nature of modern reconfigurable devices using a Singular Value Decomposition based algorithm, which provides an efficient mapping of filter's implementation requirements to the heterogeneous components of modern FPGAs. In the case of multiple 2D filters, the proposed algorithm also exploits any redundancy that exists within each filter and between different filters in the set, leading to designs with minimized area. Experiments with real filter sets from computer vision applications demonstrate an average of up to 38% reduction in the required area.
Christos-Savvas Bouganis, Sung-Boem Park, George A. Constantinides, Peter Y. K. Cheung
ACM Trans. Reconfigurable Technol. Syst.4
2009 Self-Measurement of Combinatorial Circuit Delays in FPGAs
abstract
This article proposes a Built-In Self-Test (BIST) method to accurately measure the combinatorial circuit delays on an FPGA. The flexibility of the on-chip clock generation capability found in modern FPGAs is employed to step through a range of frequencies until timing failure in the combinatorial circuit is detected. In this way, the delay of any combinatorial circuit can be determined with a timing resolution of the order of picoseconds. Parallel and optimized implementations of the method for self-characterization of the delay of all the LUTs on an FPGA are also proposed. The method was applied to Altera Cyclone II and III FPGAs . A complete self-characterization of LUTs on a Cyclone II was achieved in 2.5 seconds, utilizing only 13kbit of block RAM to store the results. More extensive tests were carried out on the Cyclone III and the delays of adder circuits and embedded multiplier blocks were successfully measured. This self-measurement method paves the way for matching timing requirements in designs to FPGAs as a means of combating the problem of process variations.
Justin S. J. Wong, N. Pete Sedcole, Peter Y. K. Cheung
ACM Trans. Reconfigurable Technol. Syst.3
2008 Using Reconfigurable Logic to Optimise GPU Memory Accesses
abstract
Memory access patterns common in video processing algorithms, which are unsuited to the GPU (Graphics Processing Unit) memory system, are identified. We develop REDA (Reconfigurable Engine for Data Access) to improve GPU performance for such access patterns, by employing reconfigurable logic for address mapping. It is shown that a sixty times reduction in number of video memory accesses can be achieved for previously unsuited access patterns, with no detriment to well suited patterns. Surprisingly, memory access locality is also improved.
Benjamin Cope, Peter Y. K. Cheung, Wayne Luk
DATE2
2008 High-throughput interconnect wave-pipelining for global communication in FPGAs
abstract
Global interconnection is fundamental to high bandwidth links for inter-module communication in FPGAs. The long range global interconnections at high clock frequencies are becoming more problematic. This is due to the large circuit delay and leakage power caused by interconnect switches along the line. The delay is worsen by the global interconnect deterioration in technology scaling. In this paper, we address this problem by presenting an interconnect wave-pipelining strategy by using the existing programmable interconnects fabrics to provide high-throughput global communication in FPGA. A novel global interconnection circuit model is presented and, from the model, interconnection throughput can be derived. The model has been verified using SPICE simulation and delay results from the Xilinx FPGA Editor. We demonstrate the feasibility of our proposal by implementing a wave-pipelined interconnect circuit in a Xilinx Virtex-5 FPGA device. The circuit is able to achieve a throughput that is 3 times faster than a conventional synchronous approach. We conclude this paper by having a discussion about two strategies to further enhance the wave-pipelining throughput
Terrence S. T. Mak, N. Pete Sedcole, Peter Y. K. Cheung, Wayne Luk
FPGA3
2008 Measuring and modeling FPGA clock variability
abstract
As integrated circuits are scaled down it becomes difficult to maintain uniformity in process parameters across each individual die. To avoid significant performance loss through pessimistic over-design new design strategies are required that are cognizant of within-die performance variability. This paper examines the effect of variability on the clock resources in FPGA devices. Techniques and circuits for measuring clock skew variations are presented, with initial results indicating that skew variability is at least as significant as signal path variability in modern FPGAs. A model of variation in FPGA clock networks is proposed, and used to suggest strategies for reducing the impact of variations on the performance of implemented designs
N. Pete Sedcole, Justin S. J. Wong, Peter Y. K. Cheung
FPGA3
2008 Towards benchmarking energy efficiency of reconfigurable architectures
abstract
Energy research in reconfigurable architectures often involves legacy benchmarks such as the MCNC benchmarks. These benchmarks, however, are not well-suited for assessing energy consumption of reconfigurable technology, since they lack realistic input stimuli. This paper reviews and categorises a range of computation system benchmarks, and shows that there are no comprehensive benchmarks targeting reconfigurable architectures that would stimulate energy or power research. We review existing energy research in the field which involves microbenchmarks, in-house designs, or legacy benchmark suites used to evaluate power optimisations.
Tobias Becker, Peter Jamieson, Wayne Luk, Peter Y. K. Cheung, Tero Rissa
FPL4
2008 Combining data reuse exploitationwith data-level parallelization for FPGA targeted hardware compilation: A geometric programming framework
abstract
A geometric programming framework is proposed in this paper to automate exploration of the design space consisting of data reuse (buffering) exploitation and loop-level parallelization, in the context of FPGA-targeted hardware compilation. We expose the dependence between data reuse and data-level parallelization and explore both problems under the on-chip memory constraint for performance-optimal designs within a single optimization step. Results from applying this framework to several real benchmarks demonstrate that given different constraints on on-chip memory utilization, the corresponding performance-optimal designs are automatically determined by the framework, and performance improvements up to 4.7 times have been achieved compared with the method that first explores data reuse and then performs parallelization.
Qiang Liu 0011, George A. Constantinides, Kostas Masselos, Peter Y. K. Cheung
FPL4
2008 Fault tolerant methods for reliability in FPGAs
abstract
Reliability and process variability are serious issues for FPGAs in the future. Fortunately FPGAs have the ability to reconfigure in the field and at runtime, thus providing opportunities to overcome some of these issues. This paper provides the first comprehensive survey of fault detection methods and fault tolerance schemes specifically for FPGAs, with the goal of laying a strong foundation for future research in this field. All methods and schemes are qualitatively compared and some particularly promising approaches highlighted.
Edward A. Stott, N. Pete Sedcole, Peter Y. K. Cheung
FPL3
2008 Combating process variation on FPGAS with a precise at-speed delay measurement method
abstract
The goal of this PhD project is to devise a way to combat the effect of process variation on propagation delays in modern FPGAs. Through our research, we have devised a novel measurement method that is capable of measuring the delays of components on FPGAs with picosecond timing resolution and fine spatial granularity. The method avoids the use of external test equipment and able to measure stochastic delay variability, which is becoming increasingly significant. The aim is to exhaustively test FPGA components based on this method and use the results to optimise the placement and routing of circuits in FPGAs to maximise performance under the negative influence of process variation.
Justin S. J. Wong, Peter Y. K. Cheung, N. Pete Sedcole
FPL2
2008 Wave-pipelined signaling for on-FPGA communication
abstract
On-FPGA communication is becoming more problematic as the long interconnection performance is deteriorating in technology scaling. In this paper, we address this issue by presenting a new wave-pipelined signaling scheme to achieve high-throughput communication in FPGA. The throughput and power consumption of a wave-pipelined link have been derived analytically and compared to the conventional synchronous link. Two circuit designs are proposed to realize wave-pipelined link using FPGA fabrics. The proposed approaches are also compared with conventional synchronous and asynchronous pipelining techniques. It is shown that, the wave-pipelined approach can achieve up to 5.66 times improvement in throughput versus the synchronous link and 13% improvement in power consumption and 35% improvement in delay versus the synchronous register-pipelining. Also, trade-offs between power, speed and area between the proposed and conventional designs are studied.
Terrence S. T. Mak, N. Pete Sedcole, Peter Y. K. Cheung, Wayne Luk
FPT3
2008 Modelling and compensating for clock skew variability in FPGAs
abstract
As integrated circuits are scaled down it becomes difficult to maintain uniformity in process parameters across each individual die. To avoid significant performance loss through pessimistic over-design new design strategies are required that are cognisant of within-die performance variability. This paper examines the effect of process variability on the clock resources in FPGA devices. A model of variation in clock skew in FPGA clock networks is presented. Techniques for reducing the impact of variations on the performance of implemented designs are proposed and analysed, demonstrating that skew variation can be reduced by 70% or more through a combination of phase adjustment and clock rerouting. Measurements on a Virtex-5 FPGA validate the feasibility and benefits of the proposed compensation strategies.
N. Pete Sedcole, Justin S. J. Wong, Peter Y. K. Cheung
FPT3
2008 Co-optimisation of datapath and memory in outer loop pipelining
abstract
When targeting algorithms to FPGAs both the array to memory assignment and the selection of data reuse structures should be considered to maximise performance. In this work we present an integer linear programming formulation for the combined problem of array to memory assignment and data reuse selection. We include a number of cost functions to minimise during memory optimisation and show how these optimisations can be integrated into a loop pipelining framework to iteratively update the memory subsystem during scheduling. By co-optimising the datapath and memory subsystem we are able to produce near optimal (fastest) solutions, with an upper bound on the distance from the optimal. Our results show an average speedup of up to 4x over a non-optimised memory subsystem when integrated into an existing outer loop pipelining framework.
Kieron Turkington, George A. Constantinides, Peter Y. K. Cheung, Kostas Masselos
FPT3
2008 A transition probability based delay measurement method for arbitrary circuits on FPGAs
abstract
This paper proposes a novel test method for measuring the worst case path delay of any circuit on an FPGA, combinatorial or sequential, where little prior knowledge of the circuitpsilas internal structure is required. The method is based on detecting changes in the transition probability profile on the circuitpsilas output nodes while a range of test clock frequencies is stepped through. The method is applied to three classes of circuits, all implemented on an Altera Cyclone III FPGA: an adder carry chain, an embedded multiplier and a linear-feedback shift-register. The measured delays are compared to that found by a previously published, but much more time consuming, method and their results match to within 12%.
Justin S. J. Wong, N. Pete Sedcole, Peter Y. K. Cheung
FPT3
2008 Video enhancement on an adaptive image sensor
abstract
The high density pixel sensors of the latest imaging systems provide images with high resolution, but require long exposure times, which limit their applicability due to the motion blur effect. Recent technological advances have lead to image sensors that can combine in real-time several pixels together to form a larger pixel. Larger pixels require shorter exposure times and produce high-frame-rate samples with reduced motion blur. This work proposes ways of configuring such a sensor to maximize the raw information collected from the environment, and methods to process that information and enhance the final output. In particular, a super-resolution and a deconvolution-based approach, for motion deblurring on an adaptive image sensor, are proposed, compared and evaluated.
Maria E. Angelopoulou, Christos-Savvas Bouganis, Peter Y. K. Cheung
ICIP3
2008 Glitch-aware output switching activity from word-level statistics
abstract
This paper presents models for estimating the transition activity of signals at the output of adders in Field Programmable Gate Arrays (FPGAs), given only word-level measures of the correlation and variance of the input signals to these components. This will allow the power consumed in the output wires of these components to be estimated from a high-level description before RTL-synthesis, without resorting to time-consuming low-level simulation. The proposed model combines knowledge of the internal construction of adders on FPGAs with the Transition Density model for activity propagation [1] and typical activity profiles for signals within Digital Signal Processing (DSP) systems according to the DBT model [2], and is characterized using device-level power measurements. The resulting closed form expression allows power consumption estimates to be quickly made in order to drive design exploration decisions during power-aware synthesis. The model has been verified by comparing it to power estimates generated by the low-level power estimation tool XPower, achieving a mean relative error in total activity of 2.1%, whilst being several orders of magnitude times faster than XPower.
Jonathan A. Clarke, George A. Constantinides, Peter Y. K. Cheung, Alastair M. Smith
ISCAS3
2008 Implementation of Wave-Pipelined Interconnects in FPGAs
Terrence S. T. Mak, Crescenzo D'Alessandro, N. Pete Sedcole, Peter Y. K. Cheung, Alexandre Yakovlev, Wayne Luk
NOCS4
2008 Comments on the BCS Lecture "The Future of Computer Technology and its Implications for the Computer Industry" by Professor Steve Furber
abstract
Department of Electrical and Electronic Engineering, Imperial College London, UK Email: [email protected] Professor Furber has, in a clear and succinct manner, provided us with a pragmatic and honest overview of the challenges facing the computer industry in the future. We have had a brief history of its development. Most of us in the audience, I am sure, are encouraged by achievements made in the last 60 years. Some may even feel proud knowing that they have made their personal contributions. We have been warned of the imminent danger facing the industry, on reliability (or more accurately, unreliability), on escalating costs and on challenging business models that the industry operates under. We are illuminated with some light at the end of the tunnel. In particular, we learn about the UK's efforts in setting up the Microelectronics design Grand Challenges, and the interesting paradigm in computing by learning from nature through the working of the brain.
Peter Y. K. Cheung, Alexandre Yakovlev
Comput. J.1
2008 Affective Level Video Segmentation by Utilizing the Pleasure-Arousal-Dominance Information
abstract
In this paper, we offer an entirely new view to the problem of high level video parsing. We developed a novel computation method for affective level video segmentation. Its function was to extract emotional segments from videos. Its design was based on the pleasure-arousal-dominance (P-A-D) model of affect representation , which in principle can represent a large number of emotions. Our method had two stages. The first P-A-D estimation stage was defined within framework of the dynamic Bayesian networks (DBNs). A spectral clustering algorithm was applied in the final stage to determine the emotional segments of the video. The performance of our method was compared with the time adaptive clustering (TAC) algorithm and an accelerated version of it which we had developed. According to Vendrig , the TAC algorithm was the best segmentation method. Experiment results will show the feasibility of our method.
Sutjipto Arifin, Peter Y. K. Cheung
IEEE Trans. Multim.2
2008 Parametric Yield Modeling and Simulations of FPGA Circuits Considering Within-Die Delay Variations
abstract
Variations in the semiconductor fabrication process results in differences in parameters between transistors on the same die, a problem exacerbated by lithographic scaling. Field-Programmable Gate Arrays may be able to compensate for within-die delay variability, by judicious use of reconfigurability. This article presents two strategies for compensating within-die stochastic delay variability by using reconfiguration: reconfiguring the entire FPGA, and relocating subcircuits within an FPGA. Analytical models for the theoretical bounds on the achievable gains are derived for both strategies and compared to models for worst-case design as well as statistical static timing analysis (SSTA). All models are validated by comparison to circuit-level Monte Carlo simulations. It is demonstrated that significant improvements in circuit yield and timing are possible using SSTA alone, and these improvements can be enhanced by employing reconfiguration-based techniques.
N. Pete Sedcole, Peter Y. K. Cheung
ACM Trans. Reconfigurable Technol. Syst.2
2008 Integrated Floorplanning, Module-Selection, and Architecture Generationfor Reconfigurable Devices
abstract
This paper is concerned with the application of formal optimization methods to the design of mixed-granularity field-programmable gate arrays (FPGAs). In particular, we investigate the appropriate mix and floorplan of heterogeneous elements: multipliers, RAMs, and lookup table (LUT)-based logic, in order to maximize the performance of a set of digital signal processing (DSP) benchmark applications, given a fixed silicon budget. A mathematical programming framework is introduced, along with a set of heuristics, capable of providing upper-bounds on the achievable reconfigurable-to-fixed-logic performance ratio. Moreover, we use linear-programming bounding procedures from the operations research community to provide lower-bounds on the same quantity. Our results provide, for the first time, quantifications of the optimal performance/area-enhancing capability of multipliers and RAM blocks within a system context. The approach detailed provides a formal mechanism to explore future technology nodes.
Alastair M. Smith, George A. Constantinides, Peter Y. K. Cheung
IEEE Trans. Very Large Scale Integr. Syst.3
2008 Outer Loop Pipelining for Application Specific Datapaths in FPGAs
abstract
Most hardware compilers apply loop pipelining to increase the parallelism achieved, but pipelining is restricted to the only innermost level in a nested loop. In this work we extend and adapt an existing outer loop pipelining approach known as single dimension software pipelining to generate schedules for field-programmable gate-array (FPGA) hardware coprocessors. Each loop level in nine test loops is pipelined and the resulting schedules are implemented in VHDL and targeted to an Altera Stratix II FPGA. The results show that the fastest solution for all but one of the loops occurs when pipelining is applied one to three levels above the innermost loop. Across the nine test loops we achieve an acceleration over the innermost loop solution of up to seven times, with a mean speedup of 3.2 times. The results suggest that inclusion of outer loop pipelining in future hardware compilers may be worthwhile as it can allow significantly improved results to be achieved at the cost of a small increase in compile time.
Kieron Turkington, Turkington A. Constantinides, Kostas Masselos, Peter Y. K. Cheung
IEEE Trans. Very Large Scale Integr. Syst.4
2007 Bridging the Gap between FPGAs and Multi-Processor Architectures: A Video Processing Perspective
abstract
This work explores how the graphics processing unit (GPU) pipeline model can influence future multi-core architectures which include reconfigurable logic cores. The design challenges of implementing five algorithms on two field programmable gate arrays (FPGAs) and two GPUs are explained and performance results contrasted. Explored algorithm features include data dependence, flexible data reuse patterns and histogram generation. A customisable systemC model, which demonstrates that features of the GPU pipeline can be transferred to a general multi-core architecture, is implemented. The customisations are: choice of processing unit (PU); processing pattern; and on-chip memory organisation. Example tradeoffs are: the choice of processing pattern for histogram equalisation; choice of number of PUs; and memory sizing for motion vector estimation. It is shown that a multi-core architecture can be optimised for video processing by combining a GPU pipeline with cores that support reconfigurable datapath operations.
Benjamin Cope, Peter Y. K. Cheung, Wayne Luk
ASAP2
2007 A Hybrid Memory Sub-system for Video Coding Applications
abstract
This paper introduces a parameterisable, application and platform-independent, hybrid memory sub-system for custom hardware. This memory sub-system consists of a scratchpad memory (SPM) and a custom parallel cache, which exploits data re-use effectively in spite of data dependence. The cache is capable of exploiting spatial locality of memory accesses in two dimensions, making it ideal for video applications. Further, we present a case study involving the Quad-tree Structured Pulse Code Modulation (QSDPCM) algorithm, commonly used in MPEG applications. Specifically, the data dependent nature of memory accesses is demonstrated. Using the memory sub-system, performance improvements of up to 1.7times and 1.4times are obtained when the application is implemented on an Altera Stratix 2 chip and a Xilinx Virtex 2 chip respectively, compared to a SPM implementation. In addition, memory savings of up to 3.2times are achieved. These results emphasize the importance of developing dynamic memory sub-systems for custom hardware applications.
Su-Shin Ang, George A. Constantinides, Wayne Luk, Peter Y. K. Cheung
FCCM4
2007 Enhancing Relocatability of Partial Bitstreams for Run-Time Reconfiguration
abstract
This paper introduces a method that enhances the relocatability of partial bitstreams for FPGA run-time reconfiguration. Reconfigurable applications usually employ partial bitstreams which are specific to one target region on the FPGA. Previously, techniques have been proposed that allow relocation between identical regions on the FPGA. However, as FPGAs are becoming increasingly heterogeneous, this approach is often too restrictive. We introduce a method that circumvents the problem of having to find fully identical regions based on compatible subsets of resources, enabling flexible placement of relocatable modules. In a software defined radio prototype with two reconfigurable regions, the number of partial bitstreams is reduced by 50% and the compile time is shortened by 43%.
Tobias Becker, Wayne Luk, Peter Y. K. Cheung
FCCM3
2007 Efficient Mapping of Dimensionality Reduction Designs onto Heterogeneous FPGAs
abstract
Dimensionality reduction or feature extraction has been widely used in applications that require to reduce the amount of original data, like in image compression, or to represent the original data by a small set of variables that capture the main modes of data variation, as in face recognition and detection applications. A linear projection is often chosen due to its computational attractiveness. The calculation of the linear basis that best explains the data is usually addressed using the Karhunen-Loeve transform (KLT). Moreover, for applications where real-time performance and flexibility to accommodate new data are required, the linear projection is implemented in FPGAs due to their fine-grain parallelism and reconfigurability properties. Currently, the optimization of such a design, in terms of area usage and efficient allocation of the embedded multipliers that exist in modern FPGAs, is considered as a separate problem to the basis calculation. In this paper, we propose a novel approach that couples the calculation of the linear projection basis, the area optimization problem, and the heterogeneity exploration of modern FPGAs under a probabilistic Bayesian framework. The power of the proposed framework is based on the flexibility to insert information regarding the implementation requirements of the linear basis by assigning a proper prior distribution. Results using real-life examples demonstrate the effectiveness of our approach.
Christos-Savvas Bouganis, Iosifina Pournara, Peter Y. K. Cheung
FCCM3
2007 Automatic On-chip Memory Minimization for Data Reuse
abstract
FPGA-based computing engines have become a promising option for the implementation of computationally intensive applications due to high flexibility and parallelism. However, one of the main obstacles to overcome when trying to accelerate an application on an FPGA is the bottleneck in off-chip communication, typically to large memories. Often it is known at compile-time that the same data item is accessed many times, and as a result can be loaded once from large off-chip RAM onto scarce on-chip RAM, alleviating this bottleneck. This paper addresses how to automatically derive an address mapping that reduces the size of the required on-chip memory for a given memory access pattern. Experimental results demonstrate that, in practice, our approach reduces on-chip storage requirements to the minimum, corresponding to a reduction in on-chip memory size of up to 40times (average 10times) for some benchmarks compared to a naive approach. At the same time, no clock period penalty or increase in control logic area compared to this approach is observed for these benchmarks.
Qiang Liu 0011, George A. Constantinides, Kostas Masselos, Peter Y. K. Cheung
FCCM4
2007 Parametric yield in FPGAs due to within-die delay variations: a quantitative analysis
abstract
Variations in the semiconductor fabrication process results in variability in parameters between transistors on the same die, a problem exacerbated by lithographic scaling. The re-configurability of Field-Programmable Gate Arrays presents the opportunity to compensate for within-die delay variability. This paper presents three reconfiguration-based strategies for compensating within-die stochastic delay variability in FPGAs: reconfiguring the entire FPGA, relocating subcircuits within an FPGA, and reconfiguring signal paths within a design. The yield of each strategy is analysed and compared with worst-case design and statistical static timing analysis (SSTA). It is demonstrated that significant im-provements in circuit yield and timing are possible using SSTA alone, and these improvements can be enhanced by employing reconfiguration-based techniques.
N. Pete Sedcole, Peter Y. K. Cheung
FPGA2
2007 On the feasibility of early routing capacitance estimation for FPGAs
abstract
Knowing the capacitance of circuit nets in an FPGA design is essential when computing the dynamic power consumed by switching these nets. Before a circuit is placed, however, there is little information available to allow the capacitance of routing wires to be estimated. In this paper we study the feasibility of estimating routing capacitance before RTL-synthesis to allow high-level power consumption optimization algorithms to be able to target routing power. We propose a novel method for estimating the capacitance of nets before RTL-synthesis and show that this method improves the accuracy and the rank ordering of the net-by-net estimates made over existing fan-out based techniques.
Jonathan A. Clarke, George A. Constantinides, Peter Y. K. Cheung
FPL3
2007 Efficient mapping of a Kalman filter into an FPGA using Taylor Expansion
abstract
The Kalman filter is widely used as an estimator in many modern applications. In the case where its implementation in hardware is required, the computational complexity of the algorithm dictates the use of many resources. This paper presents an approximation of the conventional Kalman filter by using Taylor expansion and matrix calculus in order to remove the hardware expensive part of the algorithm. The Bierman-Thornton algorithm, as the exact counterpart of our proposed Approximate Kalman filter algorithm, is also implemented for comparison purposes. Comparing to the Bierman-Thornton algorithm, the FPGA implementation results demonstrate that our proposed Approximate Kalman filter implementation achieves one order of magnitude higher throughput using less hardware resources, obtaining similar convergence rate and accuracy.
Christos-Savvas Bouganis, Peter Y. K. Cheung
FPL3
2007 Fused-Arithmetic Unit Generation for Reconfigurable Devices using Common Subgraph Extraction
abstract
To complement the flexible, fine-grain logic in field programmable gate arrays (FPGAs), configurable hardware devices now incorporate more complex coarse-grain components such as memories, embedded processing units and fused-arithmetic units. These components provide speed and density advantages due to the specialised logic and fixed interconnect. In this paper, a methodology is presented to automatically propose and explore the benefits of different types of fused arithmetic units for configurable devices. The methods are based on common subgraph extraction techniques, meaning that it is possible to explore different subcircuits that occur frequently across a set of benchmarks. A quantitative analysis is performed of the various fused-arithmetic circuits identified by our tool, which are then automatically synthesised to an ASIC process, providing a study of the speed and area benefits of the components. We report improvements of up to 3.3times in speed and 19.7times in area for the average improvement of particular silicon cores identified by our approach when compared to implementation of the same sub-circuits implemented in a commercial mixed-granularity FPGA in a comparable 90nm technology. The average improvements across all embedded cores identified by our approach are 1.67times and 5.55times when designing the ASIC cores for fastest speed performance.
Alastair M. Smith, George A. Constantinides, Peter Y. K. Cheung
FPT3
2007 Self-characterization of Combinatorial Circuit Delays in FPGAs
abstract
This paper proposes a built-in self-test (BIST) method to measure accurately the combinatorial circuit delays on an FPGA. The flexibility of the on-chip clock generation capability found in modern FPGAs is employed to step through a range of frequencies until timing failure in the combinatorial circuit is detected. In this way, the delay of any combinatorial circuit can be determined with a timing resolution of 1 ps or lower. A parallel implementation of the method for self-characterization of the delay of all the LUTs on an FPGA is also proposed. The method was applied to an Altera Cyclone-II FPGA (EP2C35). A complete self-characterization was achieved in 3 seconds, utilizing only 13 kbit of block RAM to store the results. This self-characterization method paves the way for matching timing requirements in designs to FPGAs as a means of combating the problem of process variations.
Justin S. J. Wong, N. Pete Sedcole, Peter Y. K. Cheung
FPT3
2007 A Novel Video Parsing Algorithm Utilizing the Pleasure-Arousal-Dominance Emotional Information
abstract
One of the major problems faced when designing a high level video parsing system is that viewers usually have doubts about the exact boundaries of an episode. Moreover, due to the different emotional states that viewers have while viewing a video, it is very difficult to improve the performance of these algorithms using convention methods. To solve this problem, this paper presents a novel spectral clustering based high level video parsing algorithm that utilizes the Pleasure-Arousal-Dominance (P-A-D) (Mehrabian, 1996 and Valdez and Mehrabian, 1994) emotional content of the video.
Sutjipto Arifin, Peter Y. K. Cheung
ICIP (6)2
2007 A computation method for video segmentation utilizing the pleasure-arousal-dominance emotional information
abstract
Extracting video structures is important for video indexing and navigation in large digital video archives. It is usually achieved by video segmentation algorithms. Little research efforts has been invested on segmentation solutions that utilize the video's emotional content. These solutions not only have the potential of providing better performances than existing segmentation methods, but are also able to provide a more natural video segmentation with which viewers can associate with. The development of an affect-based segmentation solution faces many challenges, such as the dynamic and time evolving nature of a video's emotional content. This paper introduces a novel computation method for affect-based video segmentation. It is designed based on the Pleasure-Arousal-Dominance (P-A-D) emotion model[18], which in principle can represent a large number of emotions. This method consists of a P-A-D estimation stage and a segmentation stage. A P-A-D estimator based on the Dynamic Bayesian Networks (DBNs) is proposed for the first stage. A clustering-based algorithm that utilizes the video's P-A-D information is proposed for the second stage. Experimental results demonstrate the feasibility of the method.
Sutjipto Arifin, Peter Y. K. Cheung
ACM Multimedia2
2007 A Hybrid Analog-Digital Routing Network for NoC Dynamic Routing
abstract
Dynamic routing can substantially enhance the quality of service for multiprocessor communication, and can provide intelligent adaptation of faulty links during run time. Implementing dynamic routing on a network-on-chip (NoC) platform requires a design that provides highly efficient optimal path computation coupled with reduced area and power consumption. In this paper, we present a hybrid analog-digital routing network design that enables efficient dynamic routing on an NoC architecture. The digital part provides accurate real-time traffic estimation using a temporal cost evaluation and adaptation scheme. The analog network, which is distributed within the digital communication network, provides an efficient implementation for the optimal routing algorithm with extremely low power consumption. Our results demonstrate the effectiveness of the hybrid analog-digital design, with a significant improvement in latency over the static routing for random hot spot traffics
Terrence S. T. Mak, N. Pete Sedcole, Peter Y. K. Cheung, Wayne Luk, Kai-Pui Lam
NOCS3
2007 Run-Time Integration of Reconfigurable Video Processing Systems
abstract
Embedded systems in field-programmable gate arrays (FPGAs) can be customized and adaptive if assembled from modular components at run time. This paper examines realizing run-time system assembly by extension of platform-based design. Two major challenges are addressed in this paper. First, the design of a reconfigurable platform architecture suitable for run-time system assembly is described. Different systems are constructed by integrating the platform architecture with different modular components, which employ the communication infrastructure supplied by the platform in order to interact. Second, where on-chip communications channels use shared media, we propose techniques for modeling the intermodule communication behavior based on statistical time-division multiplexing. The proposed techniques enable system designers to guarantee that logical communication requirements between the adjunct modules can be satisfied by the infrastructure. An in-depth analysis is presented and then verified with cycle-accurate simulations for an example reconfigurable platform for real-time video applications.
N. Pete Sedcole, Peter Y. K. Cheung, George A. Constantinides, Wayne Luk
IEEE Trans. Very Large Scale Integr. Syst.2
2006 Hardware efficient architectures for Eigenvalue computation
abstract
Eigenvalue computation is essential in many fields of science and engineering. For high performance and real-time applications, this may need to be done in hardware. This paper focuses on the exploration of hardware architectures which compute eigenvalues of symmetric matrices. We propose to use the approximate Jacobi method for general case symmetric matrix eigenvalue problem. The paper illustrates that the proposed architecture is more efficient than previous architectures reported in the literature. Moreover, for the special case of 3times3 symmetric matrices, we propose to use an algebraic method. It is shown that the pipelined architecture based on the algebraic method has a significant advantage in terms of area
Christos-Savvas Bouganis, Peter Y. K. Cheung, Philip H. W. Leong, Stephen J. Motley
DATE3
2006 A Novel Hueristic and Provable Bounds for Reconfigurable Architecture Design
abstract
This paper is concerned with the application of formal optimisation methods to the design of mixed-granularity FPGAs. In particular, we investigate the appropriate mix and floorplan of heterogeneous elements: multipliers, RAMs, and LUT-based logic, in order to maximise the performance of a set of DSP benchmark applications, given a fixed silicon budget. We extend our previous mathematical programming framework by proposing a novel set of heuristics, capable of providing upper bounds on the achievable reconfigurable-tofixed- logic performance ratio. Our results provide, for the first time, quantifications of the optimal performance/areaenhancing capability of multipliers and RAM blocks within a system context, and indicate that only a minimal performance benefit can be achieved over Virtex II by re-organising the device floorplan, when using optimal technology mapping.
Alastair M. Smith, George A. Constantinides, Peter Y. K. Cheung
FCCM3
2006 Yield enhancements of design-specific FPGAs
abstract
The high unit cost of FPGA devices often deters their use beyond the prototyping stage. Efforts have been made to reduce the part-cost of FPGA devices, resulting in the development of Design-Specific FPGAs. These parts offer cost reductions by limiting manufacturing tests and improving the number of working devices in a wafer. This paper addresses the issue of yield enhancement in Design-Specific FPGAs. In this paper, an analytical model predicting the probability of mapping a specific design onto potentially defective FPGAs is developed. When combined with existing yield modelling techniques, a quantitative measure of the potential yield improvements of the Design-Specific FPGA approach is reported for current and future technology nodes. It is found that this approach, while beneficial with current manufacturing technology, may not be suitable for 22nm technology or beyond.
Nicola Campregher, Peter Y. K. Cheung, George A. Constantinides, Milan Vasilko
FPGA2
2006 Towards Affective Level Video Applications: A Novel FPGA-Based Video Arousal Content Modeling System
abstract
The affective content of a video is defined as the expected amount and type of emotion that are contained in a video. Utilizing this affective content will extend the current scope of application possibilities. The dimensional approach to representing emotion can play an important role in the development of an affective video content analyzer. The three basic affect dimensions are defined as valence, arousal and control [Hanjalic, A, et al. 2005]. This paper presents a novel FPGA-based system for modeling the arousal content of a video based on user saliency and film grammar. The design is implemented on a Xilinx Virtex-Il xc2v6000 on board a RC300 board and it runs 25 times faster than a Pentium 4-based PC at 3.4 Ghz.
Sutjipto Arifin, Peter Y. K. Cheung
FPL2
2006 FPGA-Accelerated Pre-Attentive Segmentation in Primary Visual Cortex
abstract
Visual attention systems inspired by the behavior of neural architectures have attracted the attention of many researchers in the computer vision field. Of special interest is the model proposed by Li where the bottom-up saliency features of an image are detected through a mechanism that simulates the operation of the primary visual cortex (V1). Beyond its biological nature, the specific model is also of interest because it performs texture segmentation and contour enhancement using the same circuitry. The main drawback of the proposed model is its computational complexity, making it time consuming to simulate the model in software to, e.g., explore the model parameters, and also limits its applicability in real-time scenarios. In this work, we explore the inherent parallelism that exists in the model and propose a flexible hardware architecture that can accelerate the model. Moreover, the flexibility of the proposed architecture to adapt to similar models of the brain is of significant concern. Performance evaluation shows that the proposed architecture gives results close to the software model, achieving at the same time a speed up of one order of magnitude.
Christos-Savvas Bouganis, Peter Y. K. Cheung, Zhaoping Li 0001
FPL2
2006 Reconfiguration and Fine-Grained Redundancy for Fault Tolerance in FPGAs
abstract
As manufacturing technology enters the ultra-deep submicron era, wafer yields are destined to drop due to higher occurrence of physical defects on the die. This paper proposes a yield enhancement scheme based on the use of spare interconnect resources in each routing channel to tolerate functional faults. By using a node-covering technique and integer-linear programming (ILP) methods, the scheme is shown to provide minimal area and timing overheads. Significant yield improvements can thus be achieved.
Nicola Campregher, Peter Y. K. Cheung, George A. Constantinides, Milan Vasilko
FPL2
2006 Efficient Realtime FPGA Implementation of the Trace Transform
abstract
The trace transform is a novel image transform that is able to exhibit useful properties such as scale and rotation invariance and occlusion robustness. As a result, it is particularly suited to a variety of classification and recognition tasks including image database search, token registration, activity monitoring, character recognition and face authentication. The main obstacle to the widespread use of the transform is its high computational complexity. This has precluded a detailed investigation of transform parameters. This paper presents an architecture and implementation of a trace transform engine on a Virtex-II FPGA. By exploiting the inherent parallelism in the algorithm and the use of optimised functional blocks, a huge performance gain is achieved, exceeding realtime video processing requirements for a 256 times 256 image
Suhaib A. Fahmy, Christos-Savvas Bouganis, Peter Y. K. Cheung, Wayne Luk
FPL3
2006 On-FPGA Communication Architectures and Design Factors
abstract
The recent development of Platform-FPGA or Field-Programmable System-on-Chip architectures, with immersed coarse-grain processors, embedded memories and IP cores, offers the potential for immense computing power as well as opportunities for rapid system prototyping. These platforms require high-performance on-chip communication architectures for efficient and reliable inter-processor communication. However, as the number of embedded processors increases, communication bandwidth between embedded components becomes a limiting factor to overall system performance. In this paper, we survey the state-of-the-art on-FPGA communication architectures and methodologies. Salient factors, which include quantitative performance metrics and qualitative factors, relevant to design are identified and used to analyze and classify the on-FPGA communication architectures. This survey aims to facilitate innovation in and development of future on-FPGA communication architectures.
Terrence S. T. Mak, N. Pete Sedcole, Peter Y. K. Cheung, Wayne Luk
FPL3
2006 A Novel Heuristic and Provable Bounds for Reconfigurable Architecture Design
abstract
This paper is concerned with the application of formal optimisation methods to the design of mixed-granularity FP-GAs. In particular, we investigate the appropriate mix and floorplan of heterogeneous elements: multipliers, RAMs, and LUT-based logic, in order to maximise the performance of a set of DSP benchmark applications, given a fixed silicon budget. We extend our previous mathematical programming framework by proposing a novel set of heuristics, capable of providing upper-bounds on the achievable reconfigurable-to-fixed-logic performance ratio. Moreover, we use linear-programming bounding procedures from the operations research community to provide lower-bounds on the same quantity. Our results provide, for the first time, quantifications of the optimal performance/area-enhancing capability of multipliers and RAM blocks within a system context, and indicate that only a minimal performance benefit can be achieved over Virtex II by re-organising the device floorplan, when using optimal technology mapping.
Alastair M. Smith, George A. Constantinides, Peter Y. K. Cheung
FPL3
2006 The cost of data dependence in motion vector estimation for reconfigurable platforms
abstract
Motion vector estimation is frequently performed as a prelude to the exploitation of temporal redundancies in video applications. As a result, a large volume of work has been done to develop techniques to avoid the heavy memory access requirements of full search motion vector estimation. Often, these approaches introduce data dependence to the algorithm, leading to memory accesses which cannot be determined at design time. Consequently, this complicates the exploitation of data reuse in hardware. In this work, the cost of data dependence is quantified. Experiments indicate that a data dependent fast motion vector estimation approach is faster than full search by up to 47% in the absence of data re-use optimisation. However, full search is approximately 16 times faster than the `fast' motion vector estimation algorithm when a static line buffering scheme and a parallel caching scheme are used respectively to exploit data re-use. Therefore, it is established that data dependence in motion vector estimation is very expensive in terms of hardware performance
Su-Shin Ang, George A. Constantinides, Wayne Luk, Peter Y. K. Cheung
FPT4
2006 A comparison of 2-D discrete wavelet transform computation schedules on FPGAs
abstract
When it comes to the computation of the 2D discrete wavelet transform (DWT), three major computation schedules have been proposed, namely the row-column, the line-based and the block-based. In this work, the lifting-based designs of these schedules are implemented on FPGA-based platforms to execute the forward 2D DWT, and their comparison is presented. Our implementations are optimized in terms of throughput and memory requirements, in accordance with the specifications of each one of the three computation schedules and the lifting decomposition. All implementations are parameterized with respect to the image size and the number of decomposition levels. Experimental results prove that the suitability of each implementation for a particular application depends on the given specifications, concerning the throughput and the hardware cost
Maria E. Angelopoulou, Kostas Masselos, Peter Y. K. Cheung, Yiannis Andreopoulos
FPT3
2006 A statistical framework for dimensionality reduction implementation in FPGAs
abstract
Dimensionality reduction or feature extraction has been widely used in applications that require a set of data to be represented by a small set of variables. A linear projection is often chosen due to its computational attractiveness. The calculation of the linear basis that best explains the data is usually addressed using the Karhunen-Loeve transform (KLT). Moreover, for applications where real-time performance and flexibility to accommodate new data are required, the linear projection is implemented in FPGAs due to their fine-grain parallelism and reconfigurability properties. Currently, the optimization of such a design in terms of area usage is considered as a separate problem to the basis calculation. In this paper, we propose a novel approach that couples the calculation of the linear projection basis and the area optimization problems under a probabilistic Bayesian framework. The power of the proposed framework is based on the flexibility to insert information regarding the implementation requirements of the linear basis by assigning a proper prior distribution. Results using real-life examples demonstrate the effectiveness of our approach
Christos-Savvas Bouganis, Iosifina Pournara, Peter Y. K. Cheung
FPT3
2006 Within-die delay variability in 90nm FPGAs and beyond
abstract
Semiconductor scaling causes increasing and unavoidable within-die parametric variability. This paper describes accurate measurement techniques for characterising both systematic and stochastic delay variability in FPGAs. Results and analysis are presented from measurements made on a sample of 90nm devices, showing that delay per logic element varies stochastically by plusmn3.54% on average over the set. The delay also varies by up to 3.66% across a single die from correlated sources of variability. The results are extrapolated to determine the impact at future technology nodes. The predicted significant performance degradation that variability will cause demonstrates the importance of new circuit or system design techniques to cope with variations in future FPGAs
N. Pete Sedcole, Peter Y. K. Cheung
FPT2
2006 User Attention Based Arousal Content Modeling
abstract
The affective content of a video is defined as the expected amount and type of emotion that are contained in a video. Utilizing this affective content will extend the current scope of application possibilities. The dimensional approach to representing emotion can play an important role in the development of an affective video content analyzer. The three basic affect dimensions are defined as valence, arousal and control. This paper presents a novel FPGA-based system for modeling the arousal content of a video based on user saliency and film grammar. The design is implemented on a Xilinx Virtex-II xc2v6000 on board a RC300 board.
Sutjipto Arifin, Peter Y. K. Cheung
ICIP2
2006 A Spatiotemporal Saliency Framework
abstract
This paper presents a novel bio-inspired spatiotemporal saliency framework. The framework incorporates spatial feature detection, feature tracking and motion prediction in order to generate a spatiotemporal saliency map. Experimental results demonstrate its ability and robustness to produce saliency responses to motion pop-up phenomena that are in line with humans responses. Moreover, the limited storage requirements permit real-time implementations of the proposed framework.
Christos-Savvas Bouganis, Peter Y. K. Cheung
ICIP3
2006 Fast word-level power models for synthesis of FPGA-based arithmetic
abstract
This paper presents power models for multiplication and addition components on FPGAs which can be used at a high-level design description stage to estimate their logic and intra-component routing power consumption. The models presented are parameterized by the word-length of the component and the word-level statistics of its input signals. A key feature of these power models is the ability to handle both zero mean and non-zero mean signals. A method for measuring intra-component routing power consumption is presented, enabling the power models to account for both logic and routing power in components. The resulting models are equations which can be used to estimate the power consumed in an arithmetic component in a fraction of a second at the pre-placement stage of the design flow. The models have a mean relative error of 7.2% compared to bit-level power simulation of the placed-and-routed design
Jonathan A. Clarke, Altaf Abdul Gaffar, George A. Constantinides, Peter Y. K. Cheung
ISCAS4
2005 Automating custom-precision function evaluation for embedded processors
abstract
Due to resource and power constraints, embedded processors often cannot afford dedicated floating-point units. For instance, the IBM PowerPC processor embedded in Xilinx Virtex-II Pro FPGAs only supports emulated floating-point arithmetic, which leads to slow operation when floating-point arithmetic is desired. This paper presents a customizable mathematical library using fixed-point arithmetic for elementary function evaluation. We approximate functions via polynomial or rational approximations depending on the user-defined accuracy requirements. The data representation for the inputs and outputs are compatible with IEEE single-precision and double-precision floating-point formats. Results show that our 32-bit polynomial method achieves over 80 times speedup over the single-precision mathematical library from Xilinx, while our 64-bit polynomial method achieves over 30 times speedup.
Ray C. C. Cheung, Dong-U Lee, Oskar Mencer, Wayne Luk, Peter Y. K. Cheung
CASES5
2005 Reconfigurable Elliptic Curve Cryptosystems on a Chip
abstract
The paper presents a system-on-a-chip (SoC) architecture, which targets reconfigurable hardware, for elliptic curve cryptosystems (ECC). A four-level partitioning scheme is described for exploring the area and speed tradeoffs. A design generator is used to generate parameterisable building blocks for the configurable SoC architecture. A secure Web server, which runs on a reconfigurable soft-processor and an embedded hard-processor, shows over 2000 times speedup when computationally-intensive operations run on the customised building blocks. The embedded on-chip timer block gives accurate performance information. The design factors of configurable SoC architectures are also discussed and evaluated.
Ray C. C. Cheung, Wayne Luk, Peter Y. K. Cheung
DATE3
2005 Hardware Acceleration of Hidden Markov Model Decoding for Person Detection
abstract
This paper explores methods for hardware acceleration of hidden Markov model (HMM) decoding for the detection of persons in still images. Our architecture exploits the inherent structure of the HMM trellis to optimise a Viterbi decoder for extracting the state sequence front observation features. Further performance enhancement is obtained by computing the HMM trellis states in parallel. The resulting hardware decoder architecture is mapped onto a field programmable gate array (FPGA). The performance and resource usage of our design is investigated for different levels of parallelism. Performance advantages over software are evaluated. We show how this work contributes to a real-time system for person-tracking in video-sequences.
Suhaib A. Fahmy, Peter Y. K. Cheung, Wayne Luk
DATE2
2005 A Novel 2D Filter Design Methodology for Heterogeneous Devices
abstract
In many image processing applications, fast convolution of an image with a large 2D filter is required. Field programable gate arrays (FPGAs) are often used to achieve this goal due to their fine grain parallelism and reconfigurability. However, the heterogeneous nature of modern reconfigurable devices is not usually considered during design optimization. This paper proposes an algorithm that explores the implementation architecture of 2D filters, targeting the minimization of the required area, by optimizing the usage of the different components in a heterogeneous device. Experiments show that the proposed algorithm can achieve a reduction in the required area in a range o to 70% when compared to current techniques.
Christos-Savvas Bouganis, George A. Constantinides, Peter Y. K. Cheung
FCCM3
2005 Analysis of yield loss due to random photolithographic defects in the interconnect structure of FPGAs
abstract
This paper presents an analysis of the potential yield loss in FPGA due to random defects in metal layers. A proven yield model is adapted to target the FPGA interconnect layers in order to predict the manufacturing yield. Defect parameters from the 2003 SIA roadmap are used to investigate the trend in yield loss due to defects in interconnect layers in the future. It is shown that the low yield predicted for the 45nm technology node and beyond is a cause for concern. The potential impact on yield using two different approaches, namely redundant circuits and fault tolerant design, is also presented.
Nicola Campregher, Peter Y. K. Cheung, George A. Constantinides, Milan Vasilko
FPGA2
2005 Exploration of heterogeneous reconfigurable architectures (abstract only)
abstract
The purpose of this paper is to detail the method and findings of an architectural exploration of mixed granularity field programmable gate arrays (FPGAs). The work carried out for the purposes of this study involves the creation of an analytical framework within which a set of benchmark circuits can be studied. The idea is to maximise the performance over all benchmark circuits by choosing an optimal set of silicon cores to be placed within a given area constraint. When connected with flexible configurable routing, these cores should together be capable of performing any one of the benchmark circuits. In this paper the problem is cast as a formal optimisation, and solved using existing optimisation tools. Any multiplication or memory operation is allowed to be implemented either by configuring fine-grain resources, or by using specialised functional units such as those found in a Xilinx Virtex 2 FPGA. The design space is explored by examining the tradeoffs between area, speed and flexibility. The architectures generated are contrasted to commercial architectures with fixed ratios of functional units and, in addition, a sensitivity analysis is performed to see how the results are affected by the archtectural parameters of the problem.
Alastair M. Smith, George A. Constantinides, Peter Y. K. Cheung
FPGA3
2005 Heterogeneity Exploration for Multiple 2D Filter Designs
abstract
Many image processing applications require fast convolution of an image with a set of large 2D filters. Field-programmable gate arrays (FPGAs) are often used to achieve this goal due to their fine grain parallelism and reconfigurability. This paper presents a novel algorithm for the class of designs that implement a convolution with a set of 2D filters. Firstly, it explores the heterogeneous nature of modern reconfigurable devices using a singular value decomposition based algorithm, which orders the coefficients according to their impact to the filters' approximation. Secondly, it exploits any redundancy that exists within each filter and between different filters in the set, leading to designs with minimized area. Experiments with real filter sets from computer vision applications demonstrate up to 60% reduction in the required area.
Christos-Savvas Bouganis, Peter Y. K. Cheung, George A. Constantinides
FPL2
2005 Yield modelling and Yield Enhancement for FPGAs using Fault Tolerance Schemes
abstract
This paper presents a revised model for the yield analysis of FPGA interconnect layers. Based on proven yield models, this work improves the predictions and assumptions of previously reported analysis. The model is then applied to three well known yield improvement schemes to quantify the enhancement offered by these schemes.
Nicola Campregher, Peter Y. K. Cheung, George A. Constantinides, Milan Vasilko
FPL2
2005 Error Modelling of Dual FiXed-point Arithmetic and its Application in Field Programmable Logic
abstract
Dual FiXed-point (DFX) is a new data representation which is an efficient compromise between fixed-point and floating-point representations. DFX has an implementation complexity similar to that of a fixed-point system with the improved dynamic range capability of a floating-point system. Automating the process of DFX scaling optimisation requires the knowledge of its truncation/rounding noise properties. This paper presents truncation and rounding error models for DFX arithmetic as traditional error models do not apply to DFX. The models were tested on a 159-tap FIR filter and the benefits of using DFX over floating-point are demonstrated with implementations on a Xilinx Virtex II Pro.
Chun Te Ewe, Peter Y. K. Cheung, George A. Constantinides
FPL2
2005 Novel FPGA-Based Implementation of Median and Weighted Median Filters for Image Processing
abstract
An efficient hardware implementation of a median filter is presented. Input samples are used to construct a cumulative histogram, which is then used to find the median. The resource usage of the design is independent of window size, but rather, dependent on the number of bits in each input sample. This offers a realisable way of efficiently implementing large-windowed median filtering, as required by transforms such as the Trace Transform. The method is then extended to weighted median filtering. The designs are synthesised for a Xilinx Virtex II FPGA and the performance and area compared to another implementation for different sized windows. Intentional use of the heterogeneous resources on the FPGA in the design allows for a reduction in slice usage and high throughput.
Suhaib A. Fahmy, Peter Y. K. Cheung, Wayne Luk
FPL2
2005 Using DSP Blocks For ROM Replacement: A Novel Synthesis Flow
abstract
This paper describes a method based on polynomial approximation for transferring ROM resources used in FPGA designs to multiplication and addition operations. The technique can be applied to any FPGA architecture containing embedded multiplication, however this paper focuses on using the DSP blocks of Altera Stratix and Stratix II architectures. The transformation is combined with other resource transfers and integrated in a synthesis flow targeting designs implemented on heterogeneous FPGAs. The main advantage of such a system is in handling user constraints on each type of resource: DSP block, LUT and ROM, in addition to timing-related constraints. The flow is based on an extension to the Altera Quartus II synthesis software and Quartus University Interface Program (QUIP) framework. Results are provided for implementations of benchmark algorithms and it is shown through a design-space exploration that the set of achievable designs for the algorithms has been extended by the use of the proposed methods.
Gareth W. Morris, George A. Constantinides, Peter Y. K. Cheung
FPL3
2005 Power and Area Optimization for Multiple Restricted Multiplication
abstract
This paper presents a design and optimization technique for the multiple restricted multiplication problem [N. Sidahao, G. A. Constantinides, and F. Y. Cheung (2004)]. This refers to a situation where a single variable is multiplied by several coefficients which, while not constant, are drawn from a finite set of constants that change with time. The approach exploits dedicated registers in FPGA architecture for further time-step based optimization over previous approaches [N. Sidahao, G. A. Constantinides, and F. Y. Cheung. S. S. Demirsoy, A. G. Dempster, and I. Kale (2003)]. It is also combined with an effective technique, based on high-level power modelling, for power optimization. The problem is formulated into an integer linear program for finding solutions to the minimum-costs. The new approach results up to 22% area saving compared to the optimal non-register approach in [N. Sidahao, G. A. Constantinides, and F. Y. Cheung (2004)], and 80% of all results also show 21%-48% power savings.
Nalin Sidahao, George A. Constantinides, Peter Y. K. Cheung
FPL3
2005 An Analytical Approach to Generation and Exploration of Reconfigurable Architectures
abstract
The purpose of this paper is to detail a high-level analytical modelling and optimisation environment for mixed-granularity field programmable gate arrays (FPGAs). The work carried out for the purposes of this study involves the creation of an analytical framework that can be used to optimise the design of a reconfigurable device for a set of benchmarks. The strengths of this approach are the simultaneous placement, module selection and architecture generation. In this paper, the problem is cast as a formal optimisation, and may be solved using existing optimisation tools. In addition, the approach is adapted into an heuristic for larger benchmark sets. The design space is explored by examining the tradeoffs between area, speed and flexibility, and some comparisons to commercial architectures are drawn.
Alastair M. Smith, George A. Constantinides, Peter Y. K. Cheung
FPL3
2005 Have GPUs Made FPGAs Redundant in the Field of Video Processing?
Benjamin Cope, Peter Y. K. Cheung, Wayne Luk, Sarah Witt
FPT2
2005 FPGA Based Router for Cognitive Packet Networks
Laurence A. Hey, Peter Y. K. Cheung, Michael Gellman
FPT2
2005 Customizable elliptic curve cryptosystems
abstract
This paper presents a method for producing hardware designs for elliptic curve cryptography (ECC) systems over the finite field GF(2/sup m/), using the optimal normal basis for the representation of numbers. Our field multiplier design is based on a parallel architecture containing multiple m-bit serial multipliers; by changing the number of such serial multipliers, designers can obtain implementations with different tradeoffs in speed, size and level of security. A design generator has been developed which can automatically produce a customised ECC hardware design that meets user-defined requirements. To facilitate performance characterization, we have developed a parametric model for estimating the number of cycles for our generic ECC architecture. The resulting hardware implementations are among the fastest reported: for a key size of 270 bits, a point multiplication in a Xilinx XC2V6000 FPGA at 35 MHz can run over 1000 times faster than a software implementation on a Xeon computer at 2.6 GHz.
Ray C. C. Cheung, N. J. Telle, Wayne Luk, Peter Y. K. Cheung
IEEE Trans. Very Large Scale Integr. Syst.4
2005 Optimum and heuristic synthesis of multiple word-length architectures
abstract
This paper explores the problem of architectural synthesis (scheduling, allocation, and binding) for multiple word-length systems. It is demonstrated that the resource allocation and binding problem, and the interaction between scheduling, allocation, and binding, are complicated by the existence of multiple word-length operators. Both optimum and heuristic approaches to the combined problem are formulated. The optimum solution involves modeling as an integer linear program, while the heuristic solution considers intertwined scheduling, binding, and resource word-length selection. Techniques are introduced to perform scheduling with incomplete word-length information, to combine binding and word-length selection, and to refine word-length information based on critical path analysis. Results are presented for several benchmark and artificial examples, demonstrating significant resource savings of up to 46% are possible by considering these problems within the proposed unified framework.
George A. Constantinides, Peter Y. K. Cheung, Wayne Luk
IEEE Trans. Very Large Scale Integr. Syst.2
2004 A Novel Implementation of Tile-Based Address Mapping
abstract
Tile-based data layout has been applied to achieve various objectives such as minimizing cache conflicts and memory row switching activity. In some applications of tile-based mapping, the size of the tile can be assumed to be a power of two. In this paper, this 'power of two' assumption has been used to drastically simplify the tile-based address mapping functions. Once optimized, the implementation of the non-linear tile-based mapping consumes 60% less power than the implementation of the linear row-major mapping. This result is very interesting because one would normally expect a power penalty in the address generation stage of the more sophisticated tile-based mapping. Moreover, on average tile-based mapping implementation takes 10% less area and incurs virtually no additional delay over row-major mapping implementation.
Sambuddhi Hettiaratchi, Peter Y. K. Cheung
DATE2
2004 Unifying Bit-Width Optimisation for Fixed-Point and Floating-Point Designs
abstract
This paper presents a method that offers a uniform treatment for bit-width optimisation of both fixed-point and floating-point designs. Our work utilises automatic differentiation to compute the sensitivities of outputs to the bit-width of the various operands in the design. This sensitivity analysis enables us to explore and compare fixed-point and floating-point implementation for a particular design. As a result, we can automate the selection of the optimal number representation for each variable in a design to optimize area and performance. We implement our method in the BitSize tool targeting reconfigurable architectures, which takes user-defined constraints to direct the optimisation procedure. We illustrate our approach using applications such as ray-tracing and function approximation.
Altaf Abdul Gaffar, Oskar Mencer, Wayne Luk, Peter Y. K. Cheung
FCCM4
2004 Migrating Functionality from ROMS to Embedded Multipliers
abstract
This poster proposes a technique, based on polynomial approximation, which can be applied to convert ROMs into a combination of arithmetic operations and smaller ROMs. We show that this technique highlights new areas of the multiplier/4LUT design space over existing methods.
Gareth W. Morris, George A. Constantinides, Peter Y. K. Cheung
FCCM3
2004 A Structured System Methodology for FPGA Based System-on-A-Chip Design
abstract
The ever increasing quantities of logic resources combined with heterogeneous integrated performance enhancing primitives in high-end FPGAs creates a design complexity challenge that requires new methodologies to address. We present a structured system based design methodology which aims to increase productivity and exploit reconfigurability in large scale FPGAs. The methodology is exemplified by sonic-on-a-chip, a video image processing system.
N. Pete Sedcole, Peter Y. K. Cheung, George A. Constantinides, Wayne Luk
FCCM2
2004 A Steerable Complex Wavelet Construction and Its Implementation on FPGA
Christos-Savvas Bouganis, Peter Y. K. Cheung, Jeffrey Ng, Anil A. Bharath
FPL2
2004 BIST Based Interconnect Fault Location for FPGAs
Nicola Campregher, Peter Y. K. Cheung, Milan Vasilko
FPL2
2004 Dual Fixed-Point: An Efficient Alternative to Floating-Point Computation
Chun Te Ewe, Peter Y. K. Cheung, George A. Constantinides
FPL2
2004 SoftSONIC: A Customisable Modular Platform for Video Applications
Tero Rissa, Peter Y. K. Cheung, Wayne Luk
FPL2
2004 A Structured Methodology for System-on-an-FPGA Design
N. Pete Sedcole, Peter Y. K. Cheung, George A. Constantinides, Wayne Luk
FPL2
2004 Multiple Restricted Multiplication
Nalin Sidahao, George A. Constantinides, Peter Y. K. Cheung
FPL3
2004 A scalable hardware architecture for prime number validation
abstract
This work presents a scalable architecture for prime number validation which targets reconfigurable hardware. The primality test is crucial for security systems, especially for most public-key schemes. The Rabin-Miller Strong Pseudoprime Test has been mapped into hardware, which makes use of a circuit for computing Montgomery modular exponentiation to further speed up the validation and to reduce the hardware cost. A design generator has been developed to generate a variety of scalable and non-scalable Montgomery multipliers based on user-defined parameters. The performance and resource usage of our designs, implemented in Xilinx reconfigurable devices, have been explored using very large prime numbers. Our work demonstrates the flexibility and trade-offs in using reconfigurable platform for prototyping cryptographic hardware in embedded systems. It is shown that, for instance, a 1024-bit primality test can be completed in less than a second, and a low cost XC3S2000 FPGA chip can accommodate a 32k-bit scalable primality test with 64 parallel processing elements.
Ray C. C. Cheung, Ashley Brown, Wayne Luk, Peter Y. K. Cheung
FPT4
2004 Scalable structured data access by combining autonomous memory blocks
abstract
Many hardware designs, especially those for signal and image processing, involve structured data access such as queues, stacks and stripes. This work presents parametric descriptions as abstractions for such structured data access, and explains how these abstractions can be supported either as FPGA libraries targeting existing reconfigurable hardware devices, or as dedicated logic implementations forming autonomous memory blocks (AMBs). Scalable architectures combining the address generation logic in AMBs together to provide larger storage with parallel data access, are also examined. The effectiveness of this approach is illustrated with size and performance estimates for our FPGA libraries and dedicated logic implementations of AMBs. It is shown that for two-dimensional filtering, the dedicated AMBs can be 7 times smaller and 5 times faster than the FPGA libraries performing the same function.
Wim J. C. Melis, Peter Y. K. Cheung, Wayne Luk
FPT2
2004 Guest Editors' Introduction: Field Programmable Logic and Applications
abstract
THE impact of Field Programmable Logic on the computing community has been growing for more than a decade. Field Programmable Logic devices are no longer just a prototyping vehicle for Application-Specific Integrated Circuits (ASICs), but are increasingly found in computer systems where the user configurable logic and interconnects offer unique advantages. This special section contains seven papers reporting on a number of interesting advances in the architectures, compilation techniques, and applications of configurable computer systems, all chosen from the 13th International Conference on Field Programmable Logic and Its Applications, held on 1-3 September 2003 in Lisbon, Portugal. Two papers are selected to reflect the diversity that reconfigurable architecture can offer. The paper “The MOLEN Polymorphic Processor” by S. Vassiliadis, S. Wong, G. Gaydadjiev, K. Bertels, G. Kuzmanov, and E. Moscu Panainte presents a mixed general purpose and custom computing machine, proposing their own computing paradigm, instruction set, and compiler methodology. This paper attempts to combine both by extending a general purpose instruction set with eight special instructions to implement reconfigurable functions. This paper illustrates that the traditional barrier between the software and hardware worlds is fast diminishing. Many of the modern Field Programmable Gate Array (FPGA) architectures which include embedded processors also illustrate this new reality. The second paper, “An Asynchronous Dataflow FPGA Architecture” by J. Teifel and R. Manohar, presents an FPGA architecture able to implement highperformance asynchronous logic using the dataflow paradigm. These asynchronous circuits do not need a global clock to ensure that computation proceeds in the right sequence. Instead, all cells compute concurrently and are connected by specific communication channels which guarantee the necessary data dependencies according to a dataflow scheme. They developed a specific asynchronous FPGA device instead of utilizing conventional clocked FPGA architectures, as has been done by others in the past. Software environments and tools for reconfigurable computers can be very different from those found in conventional computers. Three papers are selected to demonstrate such differences. The first, “Operating Systems for Reconfigurable Embedded Platforms” by C. Steiger, H. Walder, and M. Platzner, addresses some issues in the design of an operating system for a reconfigurable system, focusing on the runtime environment that guarantees proper scheduling of real-time tasks. Unlike conventional software-only scheduling, this operating system requires a strong connection between the scheduling and placement of hardware modules. The second paper, “Exploiting Program Branch Probabilities in Hardware Compilation” by H. Styles and W. Luk, addresses a very interesting topic relating to the optimization of circuits implementing behaviors with branching constructs. For many years, software compilation has taken advantage of branch probabilities to optimize the average-case performance of algorithms. This paper extends the approach to compilation for reconfigurable hardware. It is demonstrated that an approach based on queuing theory can provide insights into an appropriate trade off between circuit area and circuit performance for each component in a design. As a result, the overall design has significantly improved performance under the same area constraint, compared to commonapproaches that do not consider load balancing issues. The last compilation paper by K. Shayee, J. Park, and P. Diniz considers the impact of compiler loop transformations on hardware designs implemented in reconfigurable logic. It has long been accepted that loop transformations offer a useful way to formalize and automate design space exploration. However, the impact of loop transformations on circuit performance is not always well-understood due to the simplifying architectural models often employed. This paper studies the impact of loop transformations on both area and performance measures and focuses on the particularly interesting area of loop transformations within architectures containing a limited number of memory channels. Computer systems based on Field Programmable Logic only compete favorably against conventional computer systems in specific applications. Two such applications are chosen for the last two papers. The first by C. Ebeling, C. Fisher, G. Xing, M. Shen, and H. Liu presents the design and implementation of an OFDM receiver in the RaPiD reconfigurable architecture as a case study for comparing the relative cost and performance of ASIC, programmable, FPGA, and domain-specific reconfigurable systems. The last paper, by I. Skliarova and A. Ferrari, gives a IEEE TRANSACTIONS ON COMPUTERS, VOL. 53, NO. 11, NOVEMBER 2004 1361
Peter Y. K. Cheung, George A. Constantinides, José T. de Sousa
IEEE Trans. Computers1
2004 A Gaussian Noise Generator for Hardware-Based Simulations
abstract
Hardware simulation offers the potential of improving code evaluation speed by orders of magnitude over workstation or PC-based simulation. We describe a hardware-based Gaussian noise generator used as a key component in a hardware simulation system, for exploring channel code behavior at very low bit error rates (BERs) in the range of 10/sup -9/ to 10/sup -10/. The main novelty is the design and use of nonuniform piecewise linear approximations in computing trigonometric and logarithmic functions. The parameters of the approximation are chosen carefully to enable rapid computation of coefficients from the inputs while still retaining high fidelity to the modeled functions. The output of the noise generator accurately models a true Gaussian Probability Density Function (PDF) even at very high /spl sigma/ values. Its properties are explored using: 1) several different statistical tests, including the chi-square test and the Anderson-Darling test, and 2) an application for decoding of low-density parity-check (LDPC) codes. An implementation at 133 MHz on a Xilinx Virtex-II XC2V4000-6 FPGA produces 133 million samples per second, which is seven times faster than a 2.6 GHz Pentium-IV PC; another implementation on a Xilinx Spartan-IIE XC2S300E-7 FPGA at 62 MHz is capable of a three times speedup. The performance can be improved by exploiting parallelism: an XC2V4000-6 FPGA with nine parallel instances of the noise generator at 105 MHz can run 50 times faster than a 2.6 GHz Pentium-IV PC. We illustrate the deterioration of clock speed with the increase in the number of instances.
Dong-U Lee, Wayne Luk, John D. Villasenor, Peter Y. K. Cheung
IEEE Trans. Computers4
2003 Mesh Partitioning Approach to Energy Efficient Data Layout
Sambuddhi Hettiaratchi, Peter Y. K. Cheung
DATE2
2003 A Hardware Gaussian Noise Generator for Channel Code Evaluation
abstract
Hardware simulation of channel codes offers the potential of improving code evaluation speed by orders of magnitude over workstation of PC-based simulation. We describe a hardware-based Gaussian noise generator used as a key component in a hardware simulation system, for exploring channel code behavior at very low bit error rates (BERs) in the range of 10/sup -9/ to 10/sup -10/. The main novelty is the design and use of nonuniform piecewise linear approximations in computing trigonometric and logarithmic functions. The parameters of the approximation are chosen carefully to enable rapid computation of coefficients from the inputs, while still retaining extremely high fidelity to the modeled functions. The output of the noise generator accurately models a true Gaussian PDF even at very high /spl sigma/ values. Its properties are explored using: (a) several different statistical tests, including the chi-square test and the Kolmogorov-Smirnov test, and (b) an application for decoding of low density parity check (LDPC) codes. An implementation at 133MHz on a Xilinx Virtex-II XC2V4000-6 FPGA produces 133 million samples per second, which is 40 times faster than a 2.13GHz PC; another implementation on a Xilinx Spartan-IIE XC2S300E-7 FPGA at 62MHz is capable of a 20 times speedup. The performance can be improved by exploiting parallelism: an XC2V4000-6 FPGA with three parallel instances of the noise generator at 126 MHz can run 100 times faster than a 2.13GHz PC. We illustrate the deterioration of clock speed with the increase in the number of instances.
Dong-U Lee, Wayne Luk, John D. Villasenor, Peter Y. K. Cheung
FCCM4
2003 Non-uniform Segmentation for Hardware Function Evaluation
Dong-U Lee, Wayne Luk, John D. Villasenor, Peter Y. K. Cheung
FPL4
2003 Globally Asynchronous Locally Synchronous FPGA Architectures
Andrew Royal, Peter Y. K. Cheung
FPL2
2003 A Reconfigurable Platform for Real-Time Embedded Video Image Processing
N. Pete Sedcole, Peter Y. K. Cheung, George A. Constantinides, Wayne Luk
FPL2
2003 A Unified Codesign Run-Time Environment for the UltraSONIC Reconfigurable Computer
Theerayod Wiangtong, Peter Y. K. Cheung, Wayne Luk
FPL2
2003 Cluster-Driven Hardware/Software Partitioning and Scheduling Approach for a Reconfigurable Computer System
Theerayod Wiangtong, Peter Y. K. Cheung, Wayne Luk
FPL2
2003 High-level language extensions for run-time reconfigurable systems
abstract
This paper presents high-level language extensions for designs that can be reconfigured at run time. Such extensions provide a unified framework for instantiating and controlling reconfigurable hardware blocks. Our frame-work involves capturing functional blocks at the task level, with language constructs for describing run-time reconfigurable tasks, and dynamic datatypes for describing run-time parametrisable designs. Two compilation paths, one involving the Handel-C system and the other involving the RT Pebble tools, have been developed. The effectiveness of our approach has been evaluated using designs for shape-adaptive template matching and network firewall.
Arran Derbyshire, Wayne Luk, Peter Y. K. Cheung
FPT4
2003 Hierarchical segmentation schemes for function evaluation
abstract
This paper presents a method for evaluating functions based on piecewise polynomial approximation with a novel hierarchical segmentation scheme. The use of a novel hierarchy scheme of uniform segments and segments with size varying by powers of two enables us to approximate non-linear regions of a function particularly well. This partitioning is automated: efficient look-up tables and their coefficients are generated for a given function, input range, order of the polynomials, desired accuracy and finite precision constraints. We describe an algorithm to find the optimum number of segments and the placement of their boundaries, which is used to analyze the properties of a function and to benchmark out approach. Our method is illustrated using three non-linear compound functions, /spl radic/-log(x), x log(x) and a high order rational function. We present results for various operand sizes between 8 and 24 bits for first and second order polynomial approximations.
Dong-U Lee, Wayne Luk, John D. Villasenor, Peter Y. K. Cheung
FPT4
2003 Wordlength optimization for linear digital signal processing
abstract
This paper presents an approach to the wordlength allocation and optimization problem for linear digital signal processing systems implemented as custom parallel processing units. Two techniques are proposed, one which guarantees an optimum set of wordlengths for each internal variable, and one which is a heuristic approach. Both techniques allow the user to tradeoff implementation area for arithmetic error at system outputs. Optimality (with respect to the area and error estimates) is guaranteed through modeling as a mixed integer linear program. It is demonstrated that the proposed heuristic leads to area improvements of 6% to 45% combined with speed increases compared to the optimum uniform wordlength design. In addition, the heuristic reaches within 0.7% of the optimum multiple wordlength area over a range of benchmark problems.
George A. Constantinides, Peter Y. K. Cheung, Wayne Luk
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2003 Synthesis of saturation arithmetic architectures
abstract
This paper describes a synthesis technique for automating the design of linear Digital Signal Processing (DSP) systems such as digital filters. The proposed methodology makes optimized use of saturation arithmetic to produce a small design implemented directly in hardware. An analytical technique is proposed to estimate the saturation error resulting from a particular implementation, and an optimization procedure is introduced to aim for the smallest implementation satisfying user-specified bounds on saturation and roundoff error. Results are presented illustrating significant speedup and area reduction compared with standard DSP design techniques: up to 22% improvement in area and 28% improvement in speed have been obtained on Field Programmable Gate Array (FPGA) implementations.
George A. Constantinides, Peter Y. K. Cheung, Wayne Luk
ACM Trans. Design Autom. Electr. Syst.2
2002 Performance-Area Trade-Off of Address Generators for Address Decoder-Decoupled Memory
abstract
Multimedia applications are characterized by a large, number of data accesses and complex array index manipulations. The built-in address decoder in the RAM memory model commonly used by most memory synthesis tools, unnecessarily restricts the freedom of address generator synthesis. Therefore a memory model in which the address decoder is decoupled from the memory cell array is proposed. In order to demonstrate the benefits and limitations of this alternative memory model, synthesis results for a Shift Register based Address Generator that does not require address decoding are compared to those for a counter-based address generator that requires address decoding. Results show that delay can be nearly halved at the expense of increased area.
Sambuddhi Hettiaratchi, Peter Y. K. Cheung, Thomas J. W. Clarke
DATE2
2002 Optimum Wordlength Allocation
abstract
This paper presents an approach to the wordlength allocation and optimization problem for linear digital signal processing systems implemented in Field-Programmable Gate Arrays. The proposed technique guarantees an optimum set of wordlengths for each internal variable, allowing the user to trade-off implementation area for error at system outputs. Optimality is guaranteed through modelling as a mixed integer linear program, constructed through novel techniques for the linearization of error and area constraints. Optimum results in this field are valuable since they can be used to assess the effectiveness of heuristic wordlength optimization techniques. It is demonstrated that one such previously published heuristic reaches within 0.7% of the optimum area over a range of benchmark problems.
George A. Constantinides, Peter Y. K. Cheung, Wayne Luk
FCCM2
2002 Customising Floating-Point Designs
abstract
This paper describes a method for customising the representation of floating-point numbers that exploits the flexibility of reconfigurable hardware. The method determines the appropriate size of mantissa and exponent for each operation in a design, so that a cost function with a given error specification for the output relative to a reference representation can be satisfied. Currently our tool, which adopts an iterative implementation of this method, supports single- or double-precision floating-point representation as the reference representation. It produces customised floating-point formats with arbitrary-sized mantissa and exponent. Results show that, for calculations involving large dynamic ranges, our method can achieve significant hardware reduction and speed improvement with respect to a design adopting the reference representation.
Altaf Abdul Gaffar, Wayne Luk, Peter Y. K. Cheung, Nabeel Shirazi
FCCM3
2002 Reconfigurable Shape-Adaptive Template Matching Architectures
abstract
This paper presents reconfigurable computing strategies for a Shape-Adaptive Template Matching (SA-TM) method to retrieve arbitrarily shaped objects within images or video frames. A generic systolic array architecture is proposed as the basis for comparing three designs: a static design where the configuration does not change after compilation, a partially-dynamic design where a static circuit can be reconfigured to use different on-chip data, and a dynamic design which completely, adapts to a particular computation. While the logic resources required to implement the static and partially-dynamic designs are constant and depend only on the size of the search frame, the dynamic design is adapted to the size and shape of the template object, and hence requires much less area. The execution time of the matching process greatly depends on the number of frames the same object is matched at. For a small number of frames, the dynamic and partially-dynamic designs suffer from high reconfiguration overheads. This overhead is significantly reduced if the matching process is repeated on a large number of consecutive frames. We find that the dynamic SA-TM design in a 50 MHz Virtex 1000E device, including reconfiguration time, can perform almost 7,000 times faster than a 1.4 GHz Pentium 4 PC when processing a 100/spl times/100 template on 300 consecutive video frames in HDTV format.
Jörn Gause, Peter Y. K. Cheung, Wayne Luk
FCCM2
2002 Image Registration of Real-Time Video Data Using the SONIC Reconfigurable Computer Platform
abstract
This paper is concerned with the image registration problem as applied to video sequences that have been subjected to geometric distortions. This work involves the development of a computationally efficient algorithm to restore the video sequence using image registration techniques. An approach based on motion vectors is proposed and is found to be successful in restoring the video sequence for any affine transform based distortion. The algorithm is implemented in FPGA hardware targeted for a reconfigurable computing platform called SONIC It is shown that the algorithm can efficiently restore the video data in realtime.
Wim J. C. Melis, Peter Y. K. Cheung, Wayne Luk
FCCM2
2002 Tabu Search with Intensification Strategy for Functional Partitioning in Hardware-Software Codesign
abstract
This paper presents tabu search (TS) method with intensification strategy for hardware-software partitioning. The algorithm operates on functional blocks for designs represented as directed acyclic graphs (DAG), with the objective of minimising processing time under various hardware area constraints. Results are compared to two other heuristic search algorithms: genetic algorithm (GA) and simulated annealing (SA). The comparison involves a scheduling model based on list scheduling for calculating processing time used as a system cost, assuming that shared resource conflicts do not occur. The results show that TS, which rarely appears for solving this kind of problem, is superior to SA and GA in terms of both search time and the quality of solutions. In addition, we have implemented intensification strategy in TS called penalty reward, which can further improve the quality of results.
Theerayod Wiangtong, Peter Y. K. Cheung, Wayne Luk
FCCM2
2002 Automating Customisation of Floating-Point Designs
Altaf Abdul Gaffar, Wayne Luk, Peter Y. K. Cheung, Nabeel Shirazi
FPL3
2002 Image Registration of Real-Time Broadcast Video Using the UltraSONIC Reconfigurable Computer
Wim J. C. Melis, Peter Y. K. Cheung, Wayne Luk
FPL2
2002 Run-Time Adaptive Flexible Instruction Processors
Shay Ping Seng, Wayne Luk, Peter Y. K. Cheung
FPL3
2002 Floating-point bitwidth analysis via automatic differentiation
abstract
Automatic bitwidth analysis is a key ingredient for highlevel programming of FPGAs and high-level synthesis of VLSI circuits. The objective is to find the minimal number of bits to represent a value in order to minimise the circuit area and to improve efficiency of the respective arithmetic operations, while satisfying user-defined numerical constraints. We present a novel approach to bitwidth- or precision-analysis for floating-point designs. The approach involves analysing the dataflow graph representation of a design to see how sensitive the output of a node is to changes in the outputs of other nodes: higher sensitivity requires higher precision and hence more output bits. We automate such sensitivity analysis by a mathematical method called automatic differentiation, which involves differentiating variables in a design with respect to other variables. We illustrate our approach by optimising the bitwidth for two examples, a discrete Fourier transform (DFT) implementation and a Finite Impulse Response (FIR) filter implementation.
Altaf Abdul Gaffar, Oskar Mencer, Wayne Luk, Peter Y. K. Cheung, Nabeel Shirazi
FPT4
2002 Strassen's matrix multiplication for customisable processors
abstract
Strassen's algorithm is an efficient method for multiplying large matrices. We explore various ways of mapping Strassen's algorithm into reconfigurable hardware that contains one or more customisable instruction processors. Our approach has been implemented using Nios processors with custom instructions and with custom-designed coprocessors, taking advantage of the additional logic and memory blocks available on a reconfigurable platform.
Henry M. D. Ip, James D. Low, Peter Y. K. Cheung, George A. Constantinides, Wayne Luk, Shay Ping Seng, Paul Metzgen
FPT3
2002 Incremental programming for reconfigurable engines
abstract
We present an incremental approach to developing programs for reconfigurable engines, systems which contain both instruction processors and reconfigurable hardware. The purpose is to support rapid production of prototypes, as well as their further systematic refinement and adaptation when required. The key elements of our approach include abstractions and tools based on high-level descriptions, and facilities for optimizations such as domain-specific data partitioning and run-time reconfiguration. The application of our approach is illustrated using the SONIC reconfigurable engine, which contains a multi-FPGA card in a PC system designed for video image processing.
Dong-U Lee, Wayne Luk, Peter Y. K. Cheung
FPT4
2002 PD-XML: extensible markup language for processor description
abstract
This paper introduces PD-XML, a meta-language for describing instruction processors in general and with an emphasis on embedded processors, with the specific aim of enabling their rapid prototyping, evaluation and eventual design and implementation. PD-XML is not specific to any one architecture, compiler or simulation environment and hence provides greater flexibility than related machine description methodologies. We demonstrate how PD-XML can be interfaced to existing description methodologies and tool-flows. In particular we show how PD-XML specifications can be translated into appropriate machine descriptions for the parametric HPL-PD VLIW processor, and for the Flexible Instruction Processor (FIP) approach targeting reconfigurable implementations.
Shay Ping Seng, Krishna V. Palem, Rodric M. Rabbah, Weng-Fai Wong, Wayne Luk, Peter Y. K. Cheung
FPT6
2002 Energy efficient address assignment through minimized memory row switching
abstract
Data transfer intensive applications consume a significant amount of energy in memory access. The selection of a memory location from a memory array involves driving row and column select lines. A signal transition on a row select line often consumes significantly more energy than a transition on a column select line. In order to exploit this difference in energy consumption of row and column select lines, we propose a novel address assignment methodology that aims to minimize high energy row transitions by assigning spatially and temporally local data items to the same row. The problem of energy efficient address assignment has been formulated as a multi-way graph partitioning problem and solved with a heuristic. Our experiments demonstrate that our methodology achieves row transition counts very close to the optimum and that the methodology can, for some examples, reduce row transition count by 40--70% over row major mapping. Moreover, we also demonstrate that our methodology is capable of handling access sequences with over 15 million accesses in moderate time.
Sambuddhi Hettiaratchi, Peter Y. K. Cheung, Thomas J. W. Clarke
ICCAD2
2001 Heuristic datapath allocation for multiple wordlength systems
abstract
This paper introduces a heuristic to solve the combined scheduling, resource building, and wordlength selection problem for multiple wordlength systems. The algorithm involves an iterative refinement of operator wordlength information, leading to a scheduled and bound data-flow graph. Scheduling is performed with incomplete wordlength information during the intermediate stages of this refinement process. Results show significant area savings over known alternative approaches.
George A. Constantinides, Peter Y. K. Cheung, Wayne Luk
DATE2
2001 The Multiple Wordlength Paradigm
George A. Constantinides, Peter Y. K. Cheung, Wayne Luk
FCCM2
2001 The Effect of FPGA Granularity on Video Codec Implementations
Jörn Gause, Carsten Reuter, Holger Kropp, Peter Y. K. Cheung, Wayne Luk
FCCM4
2001 A Digit-Serial Structure for Reconfigurable Multipliers
Chakkapas Visavakul, Peter Y. K. Cheung, Wayne Luk
FPL2
2000 Flexible instruction processors
abstract
This paper introduces the notion of a Flexible Instruction Processor (FIP) for systematic customisation of instruction processor design and implementation. The features of our approach include: (a) a modular framework based on “processor templates” that capture various instruction processor styles, such as stack-based or register-based styles; (b) enhancements of this framework to improve functionality and performance, such as hybrid processor templates and superscalar operation; (c) compilation strategies involving standard compilers and FIP-specific compilers, and the associated design flow; (d) technology-independent and technology-specific optimisations, such as techniques for ecient resource sharing in FPGA implementations. Our current implementation of the FIP framework is based on a highlevel parallel language called Handel-C, which can be compiled into hardware. Various customised Java Virtual Machines and MIPS style processors have been developed using existing FPGAs to evaluate the eectiveness and promise of this approach.
Shay Ping Seng, Wayne Luk, Peter Y. K. Cheung
CASES3
2000 Multiple Precision for Resource Minimization
abstract
Presents the Synoptix high-level synthesis and precision optimization system for FPGAs. Given abstract specifications in the form of infinite-precision signal flow graphs and a set of error constraints, Synoptix creates hardware descriptions of fixed-point arithmetic implementations. The width of each signal is individually optimized in order to achieve the minimal resource utilization while satisfying user-specified constraints such as signal-to-noise ratio. A heuristic for solving the optimization problem is introduced, and the results of implementations on an Altera Flex10k-based reconfigurable computing platform are reported. It is demonstrated that significant area reductions can be obtained by optimizing signal widths individually, compared to the use of a single uniform signal width.
George A. Constantinides, Peter Y. K. Cheung, Wayne Luk
FCCM2
2000 Roundoff-noise shaping in filter design
abstract
This paper presents a technique for the spectral shaping of roundoff noise in fixed-point implementations of digital filters. An automated feasibility test is introduced, in order to decide whether a given filter realisation meets user-specified constraints on the roundoff noise power spectrum. This feasibility test is used by an algorithm for optimization of individual signal widths within a filter structure. Some results are presented, illustrating how the optimization produces filters closely meeting the specification, leading to significant improvements in implementation area.
George A. Constantinides, Peter Y. K. Cheung, Wayne Luk
ISCAS2
1999 SONIC - A Plug-In Architecture for Video Processing
abstract
This paper presents the SONIC reconfigurable computing architecture and the first implementation, SONIC-I. SONIC is designed to support a software plug-in methodology to accelerate video image processing applications. SONIC differs from other architectures through the use of Plug-In Processing Elements (PIPEs) and the Application Programmer's Interface (API). Each PIPE contains a reconfigurable processor, a scalable router that also formats video data, and a frame-buffer memory. The SONIC architecture integrates multiple PIPEs together using a specialised bus structure which enables flexible and optimal pipelined processing. SONIC-I communicates with the host PC through the PCI bus and has 8 PIPEs. We have developed an easy to use API which allows SONIC-I to he used by multiple applications simultaneously. Preliminary results show that a 19 tap separable 2-D FIR filter implemented on a single PIPE achieves processing rates of more than 15 frames per second operating on 512/spl times/512 video transferred over the PCI bus. We estimate that using all 8 PIPEs, we could obtain real-time processing rates for complex operations such as image warping.
Simon D. Haynes, Peter Y. K. Cheung, Wayne Luk, John Stone
FCCM2
1999 Reconfigurable Computing for Augmented Reality
abstract
Augmented reality involves combining three-dimensional real and synthetic objects for real-time user interaction. We describe a framework for supporting augmented reality applications by appropriate hardware and software. The benefits of reconfigurable computing, which allows optimised video analysis and synthesis to adapt to environmental changes, are explained using this framework. Our approach is illustrated by video mixing, image extraction, and object tracking. Prototype designs have been implemented using an FPGA-based platform and run at full video frame rate for images up to size 640 by 980 pixels.
Wayne Luk, J. Rice, Nabeel Shirazi, Peter Y. K. Cheung
FCCM5
1998 A Reconfigurable Multiplier Array For Video Image Processing Tasks, Suitable For Embedding In An FPGA Structure
abstract
This paper presents a design for a reconfigurable multiplier array. The multiplier is constructed using an array of 4 bit Flexible Array Blocks (FABs), which could be embedded within a conventional FPGA structure. The array can be configured to perform a number of 4n/spl times/4m bit signed/unsigned binary multiplications. We have estimated that the FABs are about 25 times more efficient in area than the equivalent multiplier implemented using a conventional FPGA structure alone.
Simon D. Haynes, Peter Y. K. Cheung
FCCM2
1998 Automating Production of Run-Time Reconfigurable Designs
abstract
This paper describes a method that automates a key step in producing run-time reconfigurable designs: the identification and mapping of reconfigurable regions. In this method, two successive circuit configurations are matched to locate the components common to them, so that reconfiguration time can be minimized. The circuit configurations are represented as a weighted bipartite graph, to which an efficient matching algorithm is applied. Our method, which supports hierarchical and library-based design, is device-independent and has been tested using Xilinx 6200 FPGAs. A number of examples in arithmetic, pattern matching and image processing are selected to illustrate our approach.
Nabeel Shirazi, Wayne Luk, Peter Y. K. Cheung
FCCM3
1997 Compilation tools for run-time reconfigurable designs
abstract
This paper describes a framework and tools for automating the production of designs which can be partially reconfigured at run time. The tools include: a partial evaluator, which produces configuration files for a given design, where the number of configurations can be minimised by a process, known as compile-time sequencing; an incremental configuration calculator, which takes the output of the partial evaluator and generates an initial configuration file and incremental configuration files that partially update preceding configurations; and a tool which further optimises designs for FPGAs supporting simultaneous configuration of multiple cells. While many of our techniques are independent of the design language and device used, our tools currently target Xilinx 6200 devices. Simultaneous configuration, for example, can be used to reduce the time for reconfiguring an adder to a subtractor from time linear with respect to its size to constant time at best and logarithmic time at worst.
Wayne Luk, Nabeel Shirazi, Peter Y. K. Cheung
FCCM3
1997 Asnchronous Wrapper for Heterogeneous Systems
abstract
We propose a new method for creating globally asynchronous locally synchronous (GALS) circuits. Each locally synchronous module is surrounded by an "asynchronous wrapper" which provides an asynchronous interface to an otherwise synchronous circuit. Every locally synchronous (LS) region operates independently, minimising problems of clock skew and enabling regions to run at different clock speeds if desired. Metastability can never cause the system to fail because an asynchronous handshake "stretches" or "pauses" the local clock until data has stabilised. When new data is not available for processing, the local clock stretches, automatically preventing the LS block from consuming power. Once new data does arrive, the block responds directly in phase with the handshake without wasted synchronisation time. The LS modules can be designed using typical synchronous techniques. However, since the external interface to each LS block uses asynchronous handshaking, we can now freely mix synchronous and asynchronous circuits.
David S. Bormann, Peter Y. K. Cheung
ICCD2
1997 Diagnosis of Boards for Realistic Interconnect Shorts
José T. de Sousa, Peter Y. K. Cheung
J. Electron. Test.2
1996 On the viability of FPGA-based integrated coprocessors
abstract
The paper examines the viability of using integrated programmable logic as a coprocessor to support a host CPU core. This adaptive coprocessor is compared to a VLIW machine in term of both die area occupied and performance. The parametric bounds necessary to justify the adoption of an FPGA-based coprocessor are established. An abstract field programmable gate array model is used to investigate the area and delay characteristics of arithmetic circuits implemented on FPGA architectures to determine the potential speedup of FPGA-based coprocessors. Analysis shows that integrated FPGA arrays are suitable as coprocessor platforms for realising algorithms that require only limited numbers of multiplication instructions. Inherent FPGA characteristics limit the data-path widths that can be supported efficiently for these applications. An FPGA-based adaptive coprocessor requires a large minimum die area before any advantage over a VLIW machine of a comparable size can be realised.
Osama T. Albaharna, Peter Y. K. Cheung, Thomas J. Clarke
FCCM2
1996 Modelling and optimising run-time reconfigurable systems
abstract
We present a simple model for specifying and optimising designs which contain elements that can be reconfigured at run-time. In this model the control mechanism for reconfiguration can be implemented in many ways: by the user using multiplexers or other logic blocks, or by FPGAs which support dynamic partial reconfiguration. The model can be used for assessing trade-offs in run-time reconfigurable systems such as operation speed, design size, reconfiguration time and complexity of reconfiguration controllers; current work includes expressing the model in a framework which also captures layout information. Our approach is illustrated by various reconfigurable implementations for filtering and locating edges in images. The design tradeoffs of these implementations are being evaluated on a PCI platform, which contains a Xilinx 6216 device.
Wayne Luk, Nabeel Shirazi, Peter Y. K. Cheung
FCCM3
1996 Adaptive Automatic Facial Feature Segmentation
abstract
Automatic facial feature detection is typically solved by using manually segmented images to train a feature detector In this paper we investigate whether it is possible to improve the detection performance of such a feature detector by using additional unsegmented images. We propose a new adaptive automatic facial feature segmentation algorithm which aims to do this. The experimental results using this algorithm demonstrate that it is possible to improve the detection performance obtained from a small segmented training set by using a larger number of additional unsegmented images.
Hasan Demirel, Thomas J. Clarke, Peter Y. K. Cheung
FG3
1996 Hierarchical tolerance analysis using statistical behavioral models
abstract
A methodology is presented for deriving statistical models of analog and digital circuit cells at the behavioral level. These models can be combined in a single simulation environment for efficient yield estimation of large circuits. The motivation is the growing importance of mixed analog/digital ASICs and the impracticality of traditional approaches to tolerance analysis based on computationally intensive device-level simulation. An efficient method of mapping from device-level space to behavioral space which requires no a priori assumption about the analytical mapping is presented. The method is demonstrated using an operational amplifier example. By combining the mapping with statistical methods, tolerance information is included in the behavioral model. A statistical model giving the mean, standard deviation, and correlation of behavioral parameters is obtained. Hence the tolerance analysis problem can be defined at the behavioral level of simulation and the statistical behavioral models combined to estimate the variation of system level performance. This hierarchical methodology is demonstrated using a two-stage flash analog-to-digital converter circuit. Compared to device-level simulation, a fifteen-fold gain in efficiency, and accuracy to within 2% were achieved in the yield estimate using a static performance specification.
Timo Koskinen, Peter Y. K. Cheung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1994 Area & Time Limitations of FPGA-based Virtual Hardware
abstract
This paper examines the limitations of integrating programmable logic with a powerful core processor on the same die. An abstract model to investigate the area and delay of field programmable gate array architectures is presented. The model is used to show that a system implemented on FPGAs will require as much as 100 times more die area than its custom VLSI implementation and would be about 10 times slower. Our analysis shows that this high cost, inherent to the current FPGA-based architectures, is a severe limitation to virtual hardware development. A new approach to cell architecture and array organization is needed to deliver high computational speed-ups comparable to multiple processor systems with the same total die area.>
Osama T. Albaharna, Peter Y. K. Cheung, Thomas J. Clarke
ICCD2
1994 Analog Fault Diagnosis - A Practical Approach
abstract
This paper describes a new implementation of model-based diagnosis for analog circuits that uses techniques from circuit optimisation to perform both hypothesis testing and component modelling. The approach explicitly takes account of tolerance deviations on components and of errors in measurements; it can diagnosis both catastrophic faults due to open/short circuiting and parametric faults due to component tolerances. The method can be used to locate faults within medium to large circuits within several minutes without an excessive number of measurements.>
Peter Y. K. Cheung
ISCAS2
1994 Virtual Hardware and the Limits of Computational Speed-up
abstract
This paper investigates the limits of achievable computational speed-up using available FPGA-based virtual hardware platforms as examples. It is shown that even if the additional hardware area is limited to only a 100% increment over the size of a realistic future general purpose processor, the virtual platform would still be able to exploit enough of the available algorithmic concurrency to push the overall task speed-up sufficiently close to its theoretical maximum.>
Osama T. Albaharna, Peter Y. K. Cheung, Thomas J. Clarke
ISCAS2
1994 A Method of Representative Fault Selection in Digital Circuits for ATPG
abstract
A new method of representative fault selection in digital circuits based on the concepts of test equivalent and test implied faults is introduced in order to minimise the number of target faults to be considered in ATPG and fault simulation. Experimental results on a set of ISCAS benchmark combinational and sequential circuits have shown a significant reduction in the number of target faults when compared with other recently published techniques.>
Akachai Sang-In, Peter Y. K. Cheung
ISCAS2
1993 A New Schematic-driven Floorplanning Algorithm for Analog Cell Layout
Nasir-ud-Din Gohar, Peter Y. K. Cheung
ISCAS2
1991 A Tag Coprocessor Architecture for Symbolic Languages
abstract
A novel architecture is presented for the efficient execution of symbolic languages on conventional von Neumann, register-based machines. Unlike other symbolic processing architectures, this is based on a tag coprocessor (TC) which is designed to work in parallel with a conventional RISC CPU such as the MIPS R3000. The TC performs almost all the tag manipulation operations independently of the CPU. It can also perform stack height checking, range checking and loop control. This design significantly enhances the execution speed of symbolic languages such as Lisp and Prolog on a RISC processor, yet all existing software for the CPU without the TC will work with minimal modification. The simplicity of the TC architecture provides a cost-effective way of designing systems specifically for artificial intelligence applications.>
Vicente Fuentes-Sánchez, Peter Y. K. Cheung
ICCD2