EDBT 2026 Demo / reviewers in the wild / expert
Philip H. W. Leong
dblp:54/222 · also Philip Heng Wai Leong
· DBLP profile ↗
142ranked-venue papers
9as first author
23since 2021 · last 2025
0000-0002-3923-3499ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 119 · 7 first-author · 17 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 1 since 2021Computer networks · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AMD Versal Implementations of FAM and SSCA EstimatorsabstractCyclostationary analysis is widely used in signal processing, particularly in the analysis of human-made signals, and spectral correlation density (SCD) is often used to characterise cyclostationarity. Unfortunately, for real-time applications, even utilising the fast Fourier transform (FFT), the high computational complexity associated with estimating the SCD limits its applicability. In this work, we present optimised, high-speed fieldprogrammable gate array (FPGA) implementations of two SCD estimation techniques. Specifically, we present an implementation of the FFT accumulation method (FAM) running entirely on the AMD Versal AI engine (AIE) array. We also introduce an efficient implementation of the strip spectral correlation analyser (SSCA) that can be used for window sizes up to 220. For both techniques, a generalised methodology is presented to parallelise the computation while respecting memory size and data bandwidth constraints. Compared to an NVIDIA GeForce RTX 3090 graphics processing unit (GPU) which uses a similar 7 nm technology to our FPGA, for the same accuracy, our FAM/SSCA implementations achieve speedups of$4.43 \times / 1.90 \times$and a$30.5 \mathrm{x} / 24.5 \mathrm{x}$improvement in energy efficiency. Carol Jingyi Li, Ruilin Wu, Philip H. W. Leong |
FPL | 3 |
| 2025 | Highly Parallel CNN Accelerator for RepVGG-Like Network Training on FPGAsabstractIn this article, we propose a generic FPGA-based training accelerator tailored for RepVGG-like networks, which strikes a balance between maximizing training-time accuracy and minimizing inference-time latency. The proposed accelerator leverages fine-grain channel-level parallelism within computational units specially designed for multiple branches of the basic building block within the RepVGG-like network. Specifically, we employ a Conv block for forward Conv and backward deConv, along with a dilated Conv block, including a weight kernel partition scheme for efficient weight gradient calculation. Furthermore, we aggressively exploit a 2-stage coarse-grain task-level parallelism for low-latency CNN training: 1) parallelism among multiple branches of the basic building block of RepVGG and 2) parallelism between error back-propagation and weight gradient calculation in the backward path. Through experiments on the CIFAR-10 dataset using 16-bit fixed-point arithmetic, we demonstrate state-of-the-art batch 1 throughput of 150 GOPs for training and 183 GOPs for inference. Chuliang Guo, Binglei Lou, David Boland, Philip H. W. Leong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | FPGA-based Block Minifloat Training Accelerator for a Time Series Prediction NetworkabstractTime series forecasting is the problem of predicting future data samples from historical information and recent deep neural network (DNNs) based techniques have achieved excellent results compared with conventional statistical approaches. Many applications at the edge can utilize this technology and most implementations have focused on inference, an ability to train at the edge would enable the DNN to adapt to changing conditions. Unfortunately, training requires approximately three times more memory and computation than inference. Moreover, edge applications are often constrained by energy efficiency. In this work, we implement a block minifloat (BM) training accelerator for a time series prediction network, N-BEATS. Our architecture involves a mixed-precision GEMM accelerator that utilizes BM arithmetic. We use a 4-bit DSP packing scheme to optimize the implementation further, achieving a throughput of 779 Gops. The resulting power efficiency is 42.4 Gops/W, 3.1 \(\times\) better than a graphics processing unit in a similar technology. Haoyan Qi, David Boland, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2025 | Toward High-Performance Network Coding: FPGA Acceleration With Bounded-Value GeneratorsabstractThe network coding enhances performance in network communications and distributed storage by increasing throughput and robustness while reducing latency. Batched sparse (BATS) codes are a class of capacity-achieving network codes, but their practical applications are hindered by their structure, computational intensity, and power demands of finite field (FF) operations. Most literature focuses on algorithmic-level techniques to improve the coding efficiency. Optimization with an algorithm/hardware co-designing approach has long been neglected. Leveraging the unique structure of BATS codes, we first present cyclic-shift BATS (CS-BATS), a hardware-friendly variant. Next, we propose a simple but effective bounded-value (BV) generator, to reduce the size of a finite field multiplier by up to 70%. Finally, we report on a scalable and resource-efficient field-programmable gate array (FPGA)-based network coding accelerator that achieves a throughput of 27 Gb/s, a speedup of more than 300 over software. Jiaxin Qing, Philip H. W. Leong, Kin-Hong Lee, Raymond W. Yeung |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | PolyLUT-Add: FPGA-based LUT Inference with Wide InputsabstractFPGAs have distinct advantages as a technology for deploying deep neural networks (DNNs) at the edge. Lookup Table (LUT) based networks, where neurons are directly modeled using LUTs, help maximize this promise of offering ultra-low latency and high area efficiency on FPGAs. Unfortunately, LUT resource usage scales exponentially with the number of inputs to the LUT, restricting PolyLUT to small LUT sizes. This work introduces PolyLUT-Add, a technique that enhances neuron connectivity by combining A PolyLUT sub-neurons via addition to improve accuracy. Moreover, we describe a novel architecture to improve its scalability. We evaluated our implementation over the MNIST, Jet Substructure classification, and Network Intrusion Detection benchmark and found that for similar accuracy, PolyLUT-Add achieves a LUT reduction of $2.0-13.9 \times$ with a $1.2-1.6 \times$ decrease in latency. Binglei Lou, Richard Rademacher, David Boland, Philip H. W. Leong |
FPL | 4 |
| 2024 | S$^{3}$CA: A Sparse Strip Spectral Correlation AnalyzerabstractThe spectral correlation density (SCD) is widely used to characterize cyclostationary signals and the strip spectral correlation analyzer (SSCA) is commonly used to estimate the SCD. Although the SSCA utilizes the fast Fourier transform (FFT) for computational efficiency, its real-time implementation still poses challenges as large input sizes are often involved. In this work, we present a sparse strip spectral correlation analyzer (S3CA) based on the sparse fast Fourier transform (SFFT). The S3CA approach involves computing a sparse, downsampled channel-data product (CDP) which is then passed to a modified SFFT implementation to obtain the spectral density. For an input of length 2 million samples, the S3CA is 30× faster than the conventional SSCA. Carol Jingyi Li, Richard Rademacher, David Boland, Craig T. Jin, Chad M. Spooner, Philip H. W. Leong |
IEEE Signal Process. Lett. | 6 |
| 2024 | Efficient Radius Search for Adaptive Foveal Sizing Mechanism in Collaborative Foveated Rendering FrameworkabstractCollaborative Foveated Rendering (CFR) is the latest collaborative rendering framework proposed to enable high frame rate VR applications on mobile devices. Compared with the strategies adopted in conventional collaborative rendering, the pixel-based Adaptive Foveal Sizing (AFS) mechanism in CFR offers a more flexible and intelligent workload trade-off by predicting the radius. However, the performance of the AFS mechanism in actual deployment depends on its adaptability to two factors, including the Sudden Environmental Variations (SEV) and the Random Discrete Latency (RDL). Guaranteeing the performance of the AFS mechanism by adapting to these two factors is of great significance to guaranteeing users' immersive experience.This paper identifies the existence of the SEV and RDL phenomenon in the AFS mechanism for the first time, and contributes the first method that offers the effective and real-time AFS mechanism implementation for the practical deployment, namely the Efficient Radius Search (ERS).The ERS method efficiently searches the largest radius online that controls the rendering workload within the foveated layer just below the offline baked threshold, thereby achieving the immediate response to SEV and reducing the oscillating frame rendering latency led by RDL.Through the experiments on 3 VR applications and 4 mobile devices, the resulting 2.44× to 9.07× higher frame rate precision compared with the state-of-the-art method demonstrate the superiority of the ERS method. Chenhao Xie 0001, Liansheng Liu, Philip H. W. Leong, Shuaiwen Song |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | FedOrbit: Energy Efficient Federated Learning for Orbital Edge Computing Using Block Minifloat ArithmeticabstractLow Earth Orbit (LEO) satellite constellations have diverse applications, including earth observation, communication services, navigation, and positioning. These constellations have evolved into a valuable data source; however, their use in a ground station (GS) for analysis via machine learning algorithms presents challenges due to constraints on power consumption, communication bandwidth, and onboard computing capabilities. While the combination of Federated Learning (FL) and Orbital Edge Computing has been employed to address these challenges, its heavy reliance on the GS for model aggregation and edge resource limitations remains a research challenge. This article presents FedOrbit, a novel energy-efficient and decentralised FL method to optimise communication with the GS and reduce power consumption. FedOrbit utilises reinforcement learning for cluster formation, satellite visiting patterns for master satellite selection, and block minifloat arithmetic for power reduction. Extensive performance evaluation under Walker Delta-based LEO constellation configurations and different datasets reveals that FedOrbit can maintain high accuracy while significantly reduce communication demand, power consumption and training time in comparison to state-of-the-art FL approaches. The proposed technique can also reduce the training time by 5× compared with the centralised FL approaches. In addition, the utilisation of block minifloat representation as low-precision arithmetic enhanced the energy consumption by 3.5× compared with the single-precision (FP32) format. Mohammad Reza Jabbarpour, Bahman Javadi, Philip H. W. Leong, Rodrigo N. Calheiros, David Boland |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Single-Batch CNN Training using Block Minifloats on FPGAsabstractTraining convolutional neural networks remains a challenge on resource-limited edge devices due to its intensive computations, large storage requirements, and high bandwidth. Error back-propagation, gradient generation, and weight update usually require high precision to guarantee model accuracy, which places a further burden on computation and bandwidth. This paper presents the first parallel FPGA CNN training accelerator with block minifloat datatypes. We first propose a heuristic bit-width allocation technique to derive a unified 8-bit block minifloat format with a sign bit, 2 exponent bits, and 5 mantissa bits. In contrast to previous techniques, the same data format is used for weights, activations, errors, and gradients. Using this format, accuracy similar to 32-bit single precision floating point is achieved and thus simplifies the FPGA-based designs of computational units such as multiply-and-add. In addition, we propose a unified Conv block to deal with Conv and transposed Conv in the forward and backward paths respectively; and a dilated Conv block with a weight kernel partition scheme for gradient generation. Both Conv blocks support non-unit stride, this being crucial for the residual connections that appear in modern CNNs. For training of ResNet20 on the CIFAR-10 dataset with a batch size of 1, our accelerator on a Xilinx Ultrascale+ ZCU102 FPGA achieves state-of-the-art single-batch throughput of 144.64 and 192.68 GOPs with and without batch normalisation layers respectively. Chuliang Guo, Binglei Lou, Xueyuan Liu 0002, David Boland, Philip H. W. Leong |
FPGA | 5 |
| 2023 | The Wyner Variational Autoencoder for Unsupervised Multi-Layer Wireless FingerprintingabstractWireless fingerprinting is a device identification approach which leverages hardware imperfections and wireless channel variations as unique user-centric signatures. Recent studies have also demonstrated that user behavior can be used as a signature by collecting network traffic data, e.g., packet length, without the need to decode/decrypt the payload. Inspired by these results, we propose a multi-layer fingerprinting framework that jointly combines the multi-layer signatures for improved identification performance. In contrast to previous works in the area, our multi-view learning approach is rooted in the common information framework developed by Wyner [1] and is able to exploit data with multiple forms to enable the extraction of the user-centric signatures shared among the multi-layer features without the need for labels (i.e., unsupervised learning setup). We further use variational inference to obtain a computationally efficient algorithm based on a tight surrogate bound on the loss function. Our evaluation framework is based on a dataset obtained by combining real-world video traffic with simulated physical layer characteristics. Finally, our empirical results show that our Wyner Variational Autoencoder significantly outper-forms the state-of-the-art baseline in the unsupervised wireless fingerprinting setting. Teng-Hui Huang, Thilini Dahanayaka, Kanchana Thilakarathna, Philip H. W. Leong, Hesham El Gamal |
GLOBECOM | 4 |
| 2023 | BOOST: Block Minifloat-Based On-Device CNN Training Accelerator with Transfer LearningabstractAdapting CNNs to changing problems is challenging on resource-limited edge devices due to intensive computations, high precision requirements, large storage needs, and high bandwidth. This paper presents BOOST, a novel block minifloat (BM)-based parallel CNN training accelerator on memory- and computation-constrained FPGAs for transfer learning (TL). By updating a small number of layers online, BOOST enables adaptation to changing problems. Our approach utilizes a unified 8-bit BM datatype (bm(2,5) ), i.e., with a sign bit, 2 exponent bits, and 5 mantissa bits, and proposes unified Conv and dilated Conv blocks that support non-unit stride and enable task-level parallelism during back-propagation to minimize latency. For ResNet20 and VGG-like training on CIFAR-10 and SVHN datasets, BOOST achieves near 32-bit floating point accuracy, reducing latency by 21%-43% and BRAM usage by 63%-66% compared to back-propagation training without TL. Notably, BOOST outperforms the prior SOTA works to achieve perbatch throughput of 131 and 209 GOPs for ResNet20 and VGG-like respectively. Chuliang Guo, Binglei Lou, Xueyuan Liu 0002, David Boland, Philip H. W. Leong, Cheng Zhuo |
ICCAD | 5 |
| 2023 | On-Board Federated Learning in Orbital Edge ComputingabstractLow Earth Orbit (LEO) satellite constellations are used for a wide range of applications including earth observation, communication services, navigation, and positioning. They have emerged as a new source of data but transferring this data to a ground station (GS) for analysis and machine learning requires extensive bandwidth and incurs high latency. Limited battery capacity, communication and computing capabilities are other factors affecting the training process. Federated Learning (FL) is being used to address these challenges, although it heavily relies on the GS for model aggregation. In this paper, we consider Orbital Edge Computing (OEC) as an architecture for LEO satellite constellations and propose an on-board Federated Learning to reduce communication with the GS. We present a novel decentralised FL algorithm, called FedOrbit, based on reinforcement learning cluster formation and satellite visiting patterns to utilise intra and inter-satellite communications for model aggregation. Extensive performance evaluation under Walker Delta-based LEO constellation configurations and different datasets including MNIST, CIFAR-10, and EuroSat revealed that FedOrbit can significantly reduce communication rounds, power consumption and training time in comparison to state-of-the-art FL approaches while maintaining a high accuracy. FedOrbit demonstrates a significant decrease in power consumption, specifically by 8.8% and 79.1% for the MNIST dataset, when compared to decentralised and centralised FL approaches, respectively. The proposed technique can also reduce the training time by 5× and 48× compared with the decentralised and centralised FL approaches, respectively. Mohammad Reza Jabbarpour, Bahman Javadi, Philip H. W. Leong, Rodrigo N. Calheiros, David Boland, Chris Butler |
ICPADS | 3 |
| 2023 | Fixed-point FPGA Implementation of the FFT Accumulation Method for Real-time Cyclostationary AnalysisabstractThe spectral correlation density (SCD) is an important tool in cyclostationary signal detection and classification. Even using efficient techniques based on the fast Fourier transform (FFT), real-time implementations are challenging because of the high computational complexity. A key dimension for computational optimization lies in minimizing the wordlength employed. In this article, we analyze the relationship between wordlength and signal-to-quantization noise in fixed-point implementations of the SCD function. A canonical SCD estimation algorithm, the FFT accumulation method (FAM) using fixed-point arithmetic, is studied. We derive closed-form expressions for SQNR and compare them at wordlengths ranging from 14 to 26 bits. The differences between the calculated SQNR and bit-exact simulations are less than 1 dB. Furthermore, an HLS-based FPGA design is implemented on a Xilinx Zynq UltraScale+ XCZU28DR-2FFVG1517E RFSoC. Using less than 25% of the logic fabric on the device, it consumes 7.7 W total on-chip power and has a power efficiency of 12.4 GOPS/W, which is an order of magnitude improvement over an Nvidia Tesla K40 graphics processing unit (GPU) implementation. In terms of throughput, it achieves 50 MS/sec, which is a speedup of 1.6 over a recent optimized FPGA implementation. Carol Jingyi Li, Xiangwei Li, Binglei Lou, Craig T. Jin, David Boland, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2023 | A Scalable Systolic Accelerator for Estimation of the Spectral Correlation Density Function and Its FPGA ImplementationabstractThe spectral correlation density (SCD) function is the time-averaged correlation of two spectral components used for analyzing periodic signals with time-varying spectral content. Although the analysis is extremely powerful, it has not been widely adopted in real-time applications due to its high computational complexity. In this article, we present an efficient FPGA implementation of the FFT accumulation method (FAM) for estimating the SCD function and its alpha profile. The implementation uses a linear systolic array with a bi-directional datapath consisting of DSP-based processing elements (PEs) with a dedicated instruction schedule, achieving a PE utilization of 88.2%. The 128-PE implementation achieves a clock frequency in excess of 530 MHz and consumes 151K LUTs, 151K FFs, 264 BRAMs, 4 URAMs, and 1,054 DSPs, which is less than 36% of the logic fabric on a Zynq UltraScale+ XCZU28DR-2FFVG1517E RFSoC device. It has a modest 12.5W power consumption and an energy efficiency of 4,832 MOPS/W, which is 20.6× better than the published state-of-the-art GPU implementation. In terms of throughput, it achieves 15,340 windows/s (15,340 windows/s × 2,048 samples/window = 31.4 MS/s), which is a 4.65× improvement compared to the above-mentioned GPU implementation and 807× compared to an existing hybrid FPGA-GPU implementation. Xiangwei Li, Douglas L. Maskell, Carol Jingyi Li, Philip H. W. Leong, David Boland |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2023 | fSEAD: A Composable FPGA-based Streaming Ensemble Anomaly Detection LibraryabstractMachine learning ensembles combine multiple base models to produce a more accurate output. They can be applied to a range of machine learning problems, including anomaly detection. In this article, we investigate how to maximize the composability and scalability of an FPGA-based streaming ensemble anomaly detector (fSEAD). To achieve this, we propose a flexible computing architecture consisting of multiple partially reconfigurable regions, pblocks, which each implement anomaly detectors. Our proof-of-concept design supports three state-of-the-art anomaly detection algorithms: Loda, RS-Hash, and xStream. Each algorithm is scalable, meaning multiple instances can be placed within a pblock to improve performance. Moreover, fSEAD is implemented using High-level synthesis (HLS), meaning further custom anomaly detectors can be supported. Pblocks are interconnected via an AXI-switch, enabling them to be composed in an arbitrary fashion before combining and merging results at runtime to create an ensemble that maximizes the use of FPGA resources and accuracy. Through utilizing reconfigurable Dynamic Function eXchange (DFX), the detector can be modified at runtime to adapt to changing environmental conditions. We compare fSEAD to an equivalent central processing unit (CPU) implementation using four standard datasets, with speedups ranging from 3× to 8×. Binglei Lou, David Boland, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2022 | On-Device Saliency Prediction Based on Pseudoknowledge DistillationabstractSaliency prediction models aim to mimic the human visual system’s attention process, and the research has made significant progress due to recent advancements in deep convolution neural networks. However, the high memory requirements and intensive computational demands make these approaches less suitable for Internet-of-Things (IoT) devices, and there is a need for an improved computational efficiency and reduced memory footprint to facilitate distributed IoT intelligence. This article proposes a pseudoknowledge distillation (PKD) training method for creating a compact real-time saliency prediction model. The proposed method can effectively transfer knowledge from computationally expensive once-for-all (OFA-595) as a single teacher model and a combination of OFA-595 and EfficientNet-B7 as a multiteacher model to an early exit evolutionary algorithm network student model by utilizing knowledge distillation and pseudolabeling. Five saliency benchmark datasets are used to demonstrate PKD’s improved prediction performance and its reduced inference time without modifying the original student model. Ayaz Umer, Chakkrit Termritthikun, Tie Qiu 0001, Philip H. W. Leong, Ivan Lee 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | NITI: Training Integer Neural Networks Using Integer-Only ArithmeticabstractLow bitwidth integer arithmetic has been widely adopted in hardware implementations of deep neural network inference applications. However, despite the promised energy-efficiency improvements demanding edge applications, the use of low bitwidth integer arithmetic for neural network training remains limited. Unlike inference, training demands high dynamic range and numerical accuracy for high quality results, making the use of low-bitwidth integer arithmetic particularly challenging. To address this challenge, we present a novel neural network training framework called NITI that exclusively utilizes low bitwidth integer arithmetic. NITI stores all parameters and accumulates intermediate values as 8-bit integers while using no more than 5 bits for gradients. To provide the necessary dynamic range during the training process, a per-layer block scaling exponentiation scheme is utilized. By deeply integrating with the rounding procedures and integer entropy loss calculation, the proposed scaling scheme incurs only minimal overhead in terms of storage and additional computation. Furthermore, a hardware-efficient pseudo-stochastic rounding scheme that eliminates the need for external random number generation is proposed to facilitate conversion from wider intermediate arithmetic results to lower precision for storage. Since NITI operates only with standard 8-bit integer arithmetic and storage, it is possible to accelerate it using existing low bitwidth operators originally developed for inference in commodity accelerators. To demonstrate this, an open-source software implementation of end-to-end training, using native 8-bit integer operations in modern GPUs is presented. In addition, experiments have been conducted on an FPGA-based training accelerator to evaluate the hardware advantage of NITI. When compared with an equivalent training setup implemented with floating point storage and arithmetic, NITI has no accuracy degradation on the MNIST and CIFAR10 datasets. On ImageNet, NITI achieves similar accuracy as state-of-the-art integer training frameworks without relying on full-precision floating-point first and last layers. Maolin Wang 0002, Seyedramin Rasoulinezhad, Philip H. W. Leong, Hayden Kwok-Hay So |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Introduction to Special Section on FPGA 2021abstractNo abstract available. Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2022 | Rethinking Embedded Blocks for Machine Learning ApplicationsabstractThe underlying goal of FPGA architecture research is to devise flexible substrates that implement a wide variety of circuits efficiently. Contemporary FPGA architectures have been optimized to support networking, signal processing, and image processing applications through high-precision digital signal processing (DSP) blocks. The recent emergence of machine learning has created a new set of demands characterized by: (1) higher computational density and (2) low precision arithmetic requirements. With the goal of exploring this new design space in a methodical manner, we first propose a problem formulation involving computing nested loops over multiply-accumulate (MAC) operations, which covers many basic linear algebra primitives and standard deep neural network (DNN) kernels. A quantitative methodology for deriving efficient coarse-grained compute block architectures from benchmarks is then proposed together with a family of new embedded blocks, called MLBlocks. An MLBlock instance includes several multiply-accumulate units connected via a flexible routing, where each configuration performs a few parallel dot-products in a systolic array fashion. This architecture is parameterized with support for different data movements, reuse, and precisions, utilizing a columnar arrangement that is compatible with existing FPGA architectures. On synthetic benchmarks, we demonstrate that for 8-bit arithmetic, MLBlocks offer 6× improved performance over the commercial Xilinx DSP48E2 architecture with smaller area and delay; and for time-multiplexed 16-bit arithmetic, achieves 2× higher performance per area with the same area and frequency. All source codes and data, along with documents to reproduce all the results in this article, are available at http://github.com/raminrasoulinezhad/MLBlocks . Seyedramin Rasoulinezhad, Esther Roorda, Steve Wilton, Philip H. W. Leong, David Boland |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2022 | FPGA Architecture Exploration for DNN AccelerationabstractRecent years have seen an explosion of machine learning applications implemented on Field-Programmable Gate Arrays (FPGAs) . FPGA vendors and researchers have responded by updating their fabrics to more efficiently implement machine learning accelerators, including innovations such as enhanced Digital Signal Processing (DSP) blocks and hardened systolic arrays. Evaluating architectural proposals is difficult, however, due to the lack of publicly available benchmark circuits. This paper addresses this problem by presenting an open-source benchmark circuit generator that creates realistic DNN-oriented circuits for use in FPGA architecture studies. Unlike previous generators, which create circuits that are agnostic of the underlying FPGA, our circuits explicitly instantiate embedded blocks, allowing for meaningful comparison of recent architectural proposals without the need for a complete inference computer-aided design (CAD) flow. Our circuits are compatible with the VTR CAD suite, allowing for architecture studies that investigate routing congestion and other low-level architectural implications. In addition to addressing the lack of machine learning benchmark circuits, the architecture exploration flow that we propose allows for a more comprehensive evaluation of FPGA architectures than traditional static benchmark suites. We demonstrate this through three case studies which illustrate how realistic benchmark circuits can be generated to target different heterogeneous FPGAs. Esther Roorda, Seyedramin Rasoulinezhad, Philip H. W. Leong, Steve Wilton |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2021 | MLBlocks: FPGA Blocks for Machine Learning ApplicationsabstractThe underlying goal of FPGA architecture research is to devise flexible substrates which implement a wide variety of circuits efficiently. Contemporary FPGA architectures have been optimized to support networking, signal processing and image processing applications through high precision digital signal processing (DSP) blocks. The recent emergence of machine learning has created a new set of demands characterized by: 1) higher computational density and 2) low precision arithmetic requirements. With the goal of exploring this new design space in a methodical manner, we first propose a problem formulation involving computing nested loops over multiply-accumulate (MAC) operations, which covers many basic linear algebra primitives and standard deep neural network (DNN) layers. A quantitative methodology for deriving efficient coarse-grained compute block architectures from benchmarks is then proposed together with a family of new compute units, called MLBlocks. These blocks are flexible mesh-based systolic array units parameterized with different data movements, data reuse, and multi-precision support. They utilize a columnar arrangement which is compatible with existing FPGA architectures. Finally, using synthetic benchmarks, we demonstrate that MLBlocks offer significantly improved performance over the commercial Xilinx DSP48E2, while maintaining similar area and timing requirements to current DSPs. Seyedramin Rasoulinezhad, David Boland, Philip H. W. Leong |
FPGA | 3 |
| 2021 | APIR-DSP: An approximate PIR-DSP architecture for error-tolerant applicationsabstractIn error-tolerant applications such as low-precision DNNs and digital filters, approximate arithmetic circuits can significantly reduce hardware resource utilization. In this work we propose an embedded block for field-programmable gate arrays, called APIR-DSP, which incorporates an approximate 9×9 hard multiplier based on the PIR-DSP architecture to improve speed and reduce area. In addition, a DSP unit evaluation platform based on Yosys and VPR which packs multiply accumulate operations into DSP blocks is developed. Using this tool we synthesis designs from Verilog implementations of matrix multiplication in DeepBench and the DoReFaNet low-precision neural network and show that APIR-DSP significantly reduces DSP resources and improves hardware utilization and performance compared with the Xilinx DSP48E2 embedded block. Compared with exact multiplication, it is shown that accuracy loss is optimized with the SNR of an FIR filter being reduced by 1.03 dB. For DNNs, accuracy loss for AlexNet is 0.31% on CIFAR10 dataset and no accuracy loss for LeNet on MNIST dataset is observed. Synthesis results show that the APIR-DSP enjoys an area reduction of 21.60%, critical path reduction of 4.85% and power consumption is reduced by 2.80%, compared with PIR-DSP. Yuan Dai, Hao Zhou 0008, Seyedramin Rasoulinezhad, Philip H. W. Leong, Lingli Wang |
FPT | 6 |
| 2021 | A Block Minifloat Representation for Training Deep Neural Networks
Sean Fox, Seyedramin Rasoulinezhad, Julian Faraone, David Boland, Philip H. W. Leong |
ICLR | 5 |
| 2020 | LUXOR: An FPGA Logic Cell Architecture for Efficient Compressor Tree ImplementationsabstractWe propose two tiers of modifications to FPGA logic cell architecture to deliver a variety of performance and utilization benefits with only minor area overheads. In the first tier, we augment existing commercial logic cell datapaths with a 6-input XOR gate in order to improve the expressiveness of each element, while maintaining backward compatibility. This new architecture is vendor-agnostic, and we refer to it as LUXOR. We also consider a secondary tier of vendor-specific modifications to both Xilinx and Intel FPGAs, which we refer to as X-LUXOR+ and I-LUXOR+ respectively. We demonstrate that compressor tree synthesis using generalized parallel counters (GPCs) is further improved with the proposed modifications. Using both the Intel adaptive logic module and the Xilinx slice at the 65nm technology node for a comparative study, it is shown that the silicon area overhead is less than 0.5% for LUXOR and 5-6% for LUXOR+, while the delay increments are 1-6% and 3-9% respectively. We demonstrate that LUXOR can deliver an average reduction of 13-19% in logic utilization on micro-benchmarks from a variety of domains. BNN benchmarks benefit the most with an average reduction of 37-47% in logic utilization, which is due to the highly-efficient mapping of the XnorPopcount operation on our proposed LUXOR+ logic cells. Seyedramin Rasoulinezhad, Siddhartha 0003, Hao Zhou 0008, Lingli Wang, David Boland, Philip H. W. Leong |
FPGA | 6 |
| 2020 | Vision Guided Crop Detection in Field Robots using FPGA-Based Reconfigurable ComputersabstractA case study in applying modern FPGAs as a platform to accelerate intelligent vision-guided crop detection in agricultural field robots is presented. A state-of-the-art YOLOv3 object detection neural network was adapted to detect broccoli and cauliflower in image dataset obtained from autonomous agricultural robots. A baseline floating point implementation achieved 96% mAP, and an efficient, quantized implementation suitable for FPGA implementation 92% mAP. The proposed FPGA solution has 136.86 ms inference latency while consuming 12.43W in a low latency configuration, and 28.48 frames per second while consuming 17.78W in a high throughput one. Compared to an embedded GPU implementation of the same task, the FPGA solution was 4.12 times more power-efficient and offers 6.85 times higher throughput, translating to faster and longer operation of a battery-powered field robot. Cyrus Wing-Hei Chan, Philip H. W. Leong, Hayden Kwok-Hay So |
ISCAS | 2 |
| 2020 | Kernel Normalised Least Mean Squares with Delayed Model AdaptationabstractKernel adaptive filters (KAFs) are non-linear filters which can adapt temporally and have the additional benefit of being computationally efficient through use of the “kernel trick”. In a number of real-world applications, such as channel equalisation, the non-linear mapping provides significant improvements over conventional linear techniques such as the least mean squares (LMS) and recursive least squares (RLS) algorithms. Prior works have focused mainly on the theory and accuracy of KAFs, with little research on their implementations. This article proposes several variants of algorithms based on the kernel normalised least mean squares (KNLMS) algorithm which utilise a delayed model update to minimise dependencies. Subsequently, this work proposes corresponding hardware architectures which utilise this delayed model update to achieve high sample rates and low latency while also providing high modelling accuracy. The resultant delayed KNLMS (DKNLMS) algorithms can achieve clock rates up to 12× higher than the standard KNLMS algorithm, with minimal impact on accuracy and stability. A system implementation achieves 250 GOps/s and a throughput of 187.4 MHz on an Ultra96 board with 1.8× higher throughput than previous state of the art. Nicholas J. Fraser, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2020 | AddNet: Deep Neural Networks Using FPGA-Optimized MultipliersabstractLow-precision arithmetic operations to accelerate deep-learning applications on field-programmable gate arrays (FPGAs) have been studied extensively, because they offer the potential to save silicon area or increase throughput. However, these benefits come at the cost of a decrease in accuracy. In this article, we demonstrate that reconfigurable constant coefficient multipliers (RCCMs) offer a better alternative for saving the silicon area than utilizing low-precision arithmetic. RCCMs multiply input values by a restricted choice of coefficients using only adders, subtractors, bit shifts, and multiplexers (MUXes), meaning that they can be heavily optimized for FPGAs. We propose a family of RCCMs tailored to FPGA logic elements to ensure their efficient utilization. To minimize information loss from quantization, we then develop novel training techniques that map the possible coefficient representations of the RCCMs to neural network weight parameter distributions. This enables the usage of the RCCMs in hardware, while maintaining high accuracy. We demonstrate the benefits of these techniques using AlexNet, ResNet-18, and ResNet-50 networks. The resulting implementations achieve up to 50% resource savings over traditional 8-bit quantized networks, translating to significant speedups and power savings. Our RCCM with the lowest resource requirements exceeds 6-bit fixed point accuracy, while all other implementations with RCCMs achieve at least similar accuracy to an 8-bit uniformly quantized design, while achieving significant resource savings. Julian Faraone, Martin Kumm, Martin Hardieck, Peter Zipf, Xueyuan Liu 0002, David Boland, Philip H. W. Leong |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2019 | PIR-DSP: An FPGA DSP Block Architecture for Multi-precision Deep Neural NetworksabstractQuantisation is a key optimisation strategy to improve the performance of floating-point deep neural network (DNN) accelerators. Digital signal processing (DSP) blocks on field-programmable gate arrays are not efficiently utilised when the accelerator precision is much lower than the DSP precision. Through three modifications to Xilinx DSP48E2 DSP blocks, we address this issue for important computations in embedded DNN accelerators, namely the standard, depth-wise, and pointwise convolutional layers. First, we propose a flexible precision, run-time decomposable multiplier architecture for CNN implementations. Second, we propose a significant upgrade to DSPDSP interconnect, providing a semi-2D low precision chaining capability which supports our low-precision multiplier. Finally, we improve data reuse via a register file which can also be configured as FIFO. Compared with the 27 × 18-bit mode in the Xilinx DSP48E2, our Precision, Interconnect, and Reuseoptimised DSP (PIR-DSP) offers a 6× improvement in multiplyaccumulate operations per DSP in the 9 × 9-bit case, 12× for 4 × 4 bits, and 24× for 2 × 2 bits. We estimate that PIR-DSP decreases the run time energy to 31/19/13% of the original value in a 9/4/2-bit MobileNet-v2 DNN implementation. Seyedramin Rasoulinezhad, Hao Zhou 0008, Lingli Wang, Philip H. W. Leong |
FCCM | 4 |
| 2019 | Unrolling Ternary Neural NetworksabstractThe computational complexity of neural networks for large-scale or real-time applications necessitates hardware acceleration. Most approaches assume that the network architecture and parameters are unknown at design time, permitting usage in a large number of applications. This article demonstrates, for the case where the neural network architecture and ternary weight values are known a priori , that extremely high throughput implementations of neural network inference can be made by customising the datapath and routing to remove unnecessary computations and data movement. This approach is ideally suited to FPGA implementations as a specialized implementation of a trained network improves efficiency while still retaining generality with the reconfigurability of an FPGA. A VGG-style network with ternary weights and fixed point activations is implemented for the CIFAR10 dataset on Amazon’s AWS F1 instance. This article demonstrates how to remove 90% of the operations in convolutional layers by exploiting sparsity and compile-time optimizations. The implementation in hardware achieves 90.9 ± 0.1% accuracy and 122k frames per second, with a latency of only 29µs, which is the fastest CNN inference implementation reported so far on an FPGA. Stephen Tridgell, Martin Kumm, Martin Hardieck, David Boland, Duncan J. M. Moss, Peter Zipf, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2019 | A Two-Speed, Radix-4, Serial-Parallel MultiplierabstractIn this paper, we present a two-speed, radix-4, serial-parallel multiplier for accelerating applications such as digital filters, artificial neural networks, and other machine learning algorithms. Our multiplier is a variant of the serial-parallel (SP) modified radix-4 Booth multiplier that adds only the nonzero Booth encodings and skips over the zero operations, making the latency dependent on the multiplier value. Two subcircuits with different critical paths are utilized so that throughput and latency are improved for a subset of multiplier values. The multiplier is evaluated on an Intel Cyclone V field-programmable gate array against standard parallel-parallel and SP multipliers across four different process-voltage-temperature corners. We show that for bit widths of 32 and 64, our optimizations can result in a 1.42×-$3.36× improvement over the standard parallel Booth multiplier in terms of area-time depending on the input set. Duncan J. M. Moss, David Boland, Philip H. W. Leong |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | SYQ: Learning Symmetric Quantization for Efficient Deep Neural NetworksabstractInference for state-of-the-art deep neural networks is computationally expensive, making them difficult to deploy on constrained hardware environments. An efficient way to reduce this complexity is to quantize the weight parameters and/or activations during training by approximating their distributions with a limited entry codebook. For very low-precisions, such as binary or ternary networks with 1-8-bit activations, the information loss from quantization leads to significant accuracy degradation due to large gradient mismatches between the forward and backward functions. In this paper, we introduce a quantization method to reduce this loss by learning a symmetric codebook for particular weight subgroups. These subgroups are determined based on their locality in the weight matrix, such that the hardware simplicity of the low-precision representations is preserved. Empirically, we show that symmetric quantization can substantially improve accuracy for networks with extremely low-precision weights and activations. We also demonstrate that this representation imposes minimal or no hardware implications to more coarse-grained approaches. Source code is available at https://www.github.com/julianfaraone/SYQ. Julian Faraone, Nicholas J. Fraser, Michaela Blott, Philip H. W. Leong |
CVPR | 4 |
| 2018 | FPGA Fastfood - A High Speed Systolic Implementation of a Large Scale Online Kernel MethodabstractIn this paper, we describe a systolic Field Programmable Gate Array (FPGA) implementation of the Fastfood algorithm that is optimised to run at a high frequency. The Fastfood algorithm supports online learning for large scale kernel methods. Empirical results show that 500 MHz clock rates can be sustained for an architecture that can solve problems with input dimensions that are $10^3$ times larger than previously reported. Unlike many recent deep learning publications, this design implements both training and prediction. This enables the use of kernel methods in applications requiring a rare combination of capacity, adaption and speed. Sean Fox, David Boland, Philip H. W. Leong |
FPGA | 3 |
| 2018 | A Customizable Matrix Multiplication Framework for the Intel HARPv2 Xeon+FPGA Platform: A Deep Learning Case StudyabstractGeneral Matrix to Matrix multiplication (GEMM) is the cornerstone for a wide gamut of applications in high performance computing (HPC), scientific computing (SC) and more recently, deep learning. In this work, we present a customizable matrix multiplication framework for the Intel HARPv2 CPU+FPGA platform that includes support for both traditional single precision floating point and reduced precision workloads. Our framework supports arbitrary size GEMMs and consists of two parts: (1) a simple application programming interface (API) for easy configuration and integration into existing software and (2) a highly customizable hardware template. The API provides both compile and runtime options for controlling key aspects of the hardware template including dynamic precision switching; interleaving and block size control; and fused deep learning specific operations. The framework currently supports single precision floating point (FP32), 16, 8, 4 and 2 bit Integer and Fixed Point (INT16, INT8, INT4, INT2) and more exotic data types for deep learning workloads: INT16xTernary, INT8xTernary, BinaryxBinary. Duncan J. M. Moss, Krishnan Srivatsan, Eriko Nurvitadhi, Piotr Ratuszniak, Jaewoong Sim, Asit K. Mishra, Debbie Marr, Suchit Subhaschandra, Philip H. W. Leong |
FPGA | 10 |
| 2018 | Customizing Low-Precision Deep Neural Networks for FPGAsabstractIn this paper, we argue that instead of solely focusing on developing efficient architectures to accelerate well-known low-precision CNNs, we should also seek to modify the network to suit the FPGA. We develop a fully automative toolflow which focuses on modifying the network through filter pruning, such that it efficiently utilizes the FPGA hardware whilst satisfying a predefined accuracy threshold. Although fewer weights are re-moved in comparison to traditional pruning techniques designed for software implementations, the overall model complexity and feature map storage is greatly reduced. We implement the AlexNet and TinyYolo networks on the large-scale ImageNet and PascalVOC datasets, to demonstrate up to roughly 2× speedup in frames per second and 2× reduction in resource requirements over the original network, with equal or improved accuracy. Julian Faraone, Giulio Gambardella, Nicholas J. Fraser, Michaela Blott, Philip H. W. Leong, David Boland |
FPL | 5 |
| 2018 | RNA: An Accurate Residual Network Accelerator for Quantized and Reconstructed Deep Neural NetworksabstractWith the continuous refinement of Deep Neural Networks (DNNs), a series of deep and complex networks such as Residual Networks (ResNets) show impressive prediction accuracy in image classification tasks. Unfortunately, the structural complexity and computational cost of residual networks make hardware implementation difficult. In this paper, we present the quantized and reconstructed deep neural network (QR-DNN) technique, which first inserts batch normalization (BN) layers in the network during training, and later removes them to facilitate efficient hardware implementation. Moreover, an accurate and efficient residual network accelerator (RNA) is presented based on QR-DNN with batch-normalization-free structures and weights represented in a logarithmic number system. RNA employs a systolic array architecture to perform shift-and-accumulate operations instead of multiplication operations. QR-DNN is shown to achieve a 1% ~ 2% improvement in accuracy over existing techniques, and RNA over previous best fixed point accelerators. An FPGA implementation on a Xilinx Zynq XC7Z045 device achieves 804.03 GOPS, 104.15 FPS and 91.41% top-5accuracyfortheResNet-50benchmark, andstate-of-the-art results are also reported for AlexNet and VGG. Wei Cao 0002, Philip H. W. Leong, Lingli Wang |
FPL | 4 |
| 2018 | Simultaneous Inference and Training Using On-FPGA Weight Perturbation TechniquesabstractWe present an FPGA-optimized implementation of online neural network training based on weight perturbation (WP) techniques. When compared to the classic backpropagation (BP) algorithm, WP is capable of delivering competitive performance while occupying minimal area resources. Perturbation-based methods have been demonstrated as viable training techniques and are suitable for on-line learning applications which adapt to changing conditions. The viability of applying WP-based on-chip training for low-precision fixed-point hardware is demonstrated on two distinct MLP benchmarks: the Iris dataset classification network and an RF anomaly detector. When synthesized to a Xilinx Kintex-7 XC7K410T FPGA, WP offers a 3-10x area savings with <;1% degradation in accuracy compared with backpropagation. Compared with an inference-only implementation the overhead of introducing on-chip learning is approximately 30%. Siddhartha 0003, Steve Wilton, David Boland, Barry Flower, Perry Blackmore, Philip H. W. Leong |
FPT | 6 |
| 2018 | Real-time FPGA-based Anomaly Detection for Radio Frequency SignalsabstractWe describe an open source, FPGA accelerated neural network-based anomaly detector. The detector derives its training set from observed exemplar data and continuous learning in software can be undertaken in an unsupervised manner. Trained network weights are passed to the FPGA, which performs continuous high-speed anomaly detection, combining parallelism reduced precision, and a single-chip design to maximise performance and energy efficiency. Our design can process continuous 200 MS/s complex inputs, producing anomaly classifications at the same rate, with a latency of 105 ns, an improvement of at least 4 orders of magnitude over a software radio such as GNU Radio. Duncan J. M. Moss, David Boland, Peyam Pourbeik, Philip H. W. Leong |
ISCAS | 4 |
| 2017 | FINN: A Framework for Fast, Scalable Binarized Neural Network Inference
Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip H. W. Leong, Magnus Jahre, Kees A. Vissers |
FPGA | 5 |
| 2017 | High performance binary neural networks on the Xeon+FPGA™ platformabstractConvolutional neural networks (CNNs) are deployed in a wide range of image recognition, scene segmentation and object detection applications. Achieving state of the art accuracy in CNNs often results in large models and complex topologies that require significant compute resources to complete in a timely manner. Binarised neural networks (BNNs) have been proposed as an optimised variant of CNNs, which constrain the weights and activations to +1 or -1 and thus offer compact models and lower computational complexity per operation. This paper presents a high performance BNN accelerator on the Intel®Xeon+FPGA™ platform. The proposed accelerator is designed to take advantage of the Xeon+FPGA system in a way that a specialised FPGA architecture can be targeted for the most compute intensive parts of the BNN whilst other parts of the topology can be handled by the Xeon™ CPU. The implementation is evaluated by comparing the raw compute performance and energy efficiency for key layers in standard CNN topologies against an Nvidia Titan X Pascal GPU and other published FPGA BNN accelerators. The results show that our single-package integrated Arria™ 10 FPGA accelerator coupled with a high-end Xeon CPU can offer comparable performance and better energy efficiency than a high-end discrete Titan X GPU card. In addition, our solution delivers the best performance compared to previous BNN FPGA implementations. Duncan J. M. Moss, Eriko Nurvitadhi, Jaewoong Sim, Asit K. Mishra, Debbie Marr, Suchit Subhaschandra, Philip H. W. Leong |
FPL | 7 |
| 2017 | Wearable healthcare systems: A single channel accelerometer based anomaly detector for studies of gait freezing in Parkinson's diseaseabstractThe causality of gait freezing in patients with advanced Parkinson's disease is still not fully understood. Clinicians are interested in investigating the freezing of gait (FoG) histogram of patients in their daily life. To that end, one needs a real-time signal processing platform that can help record freezing information (e.g., timing and the duration of every gait freezing occurrences). Wearable wireless sensors have been proposed to monitor FoG epochs. Existing automated methods using accelerometers have been introduced with high accuracy performance only for subject-dependent settings (e.g., an individual offline training process). This is a troublesome for large scale out-of-lab deployment and time-consuming. In this work, we used spectral coherence analysis for accelerometer data to apply an anomaly detection approach. Conventional features such as energy and freezing index are introduced to help refine normal epochs while the anomaly scores from spectral coherence measures define FoG epochs. Using this new set of features, our new FoG detector for subject-independent settings achieves the mean ±SD sensitivity (specificity) of 89.2±0.3% (95.6 ± 0.3%). To our best knowledge, this is the best performance for automated subject-independent approaches in literature of freezing of gait detection. Thuy T. Pham, Diep N. Nguyen, Eryk Dutkiewicz, Alistair Lee McEwan, Philip H. W. Leong |
ICC | 5 |
| 2017 | Compressing Low Precision Deep Neural Networks Using Sparsity-Induced Regularization in Ternary Networks
Julian Faraone, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip H. W. Leong |
ICONIP (2) | 5 |
| 2017 | Forecasting Financial Time Series with Grammar-Guided Feature GenerationabstractThe application of machine learning techniques to forecast financial time series is not a recent development, yet it continues to attract considerable attention because of the difficulty of the problem that is compounded by the nonlinear and nonstationary nature of the time series. The choice of an appropriate set of features is crucial to improve forecasting accuracy of machine learning techniques. In this article, we propose a systematic way for generating rich features using context‐free grammars. Our proposed methodology identifies potential candidates for new technical indicators that consistently improve forecasts compared with some well‐known indicators. The notion of grammar families as a compact representation to generate a rich class of features is exploited, and implementation issues are discussed in detail. The proposed methodology is tested on closing price data of major stock market indices, and the forecasting performance is compared with some standard techniques. A comparison with the conventional approach using standard technical indicators and naive approaches is shown. Anthony Mihirana De Silva, Richard I. A. Davis, Syed Ahmed Pasha, Philip H. W. Leong |
Comput. Intell. | 4 |
| 2017 | FPGA Implementations of Kernel Normalised Least Mean Squares ProcessorsabstractKernel adaptive filters (KAFs) are online machine learning algorithms which are amenable to highly efficient streaming implementations. They require only a single pass through the data and can act as universal approximators, i.e. approximate any continuous function with arbitrary accuracy. KAFs are members of a family of kernel methods which apply an implicit non-linear mapping of input data to a high dimensional feature space, permitting learning algorithms to be expressed entirely as inner products. Such an approach avoids explicit projection into the feature space, enabling computational efficiency. In this paper, we propose the first fully pipelined implementation of the kernel normalised least mean squares algorithm for regression. Independent training tasks necessary for hyperparameter optimisation fill pipeline stages, so no stall cycles to resolve dependencies are required. Together with other optimisations to reduce resource utilisation and latency, our core achieves 161 GFLOPS on a Virtex 7 XC7VX485T FPGA for a floating point implementation and 211 GOPS for fixed point. Our PCI Express based floating-point system implementation achieves 80% of the core’s speed, this being a speedup of 10× over an optimised implementation on a desktop processor and 2.66× over a GPU. Nicholas J. Fraser, Duncan J. M. Moss, Julian Faraone, Stephen Tridgell, Craig T. Jin, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2017 | The First 25 Years of the FPL Conference: Significant PapersabstractA summary of contributions made by significant papers from the first 25 years of the Field-Programmable Logic and Applications conference (FPL) is presented. The 27 papers chosen represent those which have most strongly influenced theory and practice in the field. Philip H. W. Leong, Hideharu Amano, Jason Helge Anderson, Koen Bertels, João M. P. Cardoso, Oliver Diessel, Guy Gogniat, Mike Hutton, Wayne Luk, Patrick Lysaght, Marco Platzner, Viktor Prasanna 0001, Tero Rissa, Cristina Silvano, Hayden Kwok-Hay So, Yu Wang 0002 |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2017 | Lossless Compression Decoders for Bitstreams and Software Binaries Based on High-Level SynthesisabstractAs the density of field-programmable gate arrays continues to increase, the size of configuration bitstreams grows accordingly. Compression techniques can reduce memory size and save external memory bandwidth. To accelerate the configuration process and reduce the software startup time, four open-source lossless compression decoders developed using high-level synthesis techniques are presented. Moreover, in order to balance the objectives of compression ratio, decompression throughput, and hardware resource overhead, various improvements and optimizations are proposed. Full bitstreams and software binaries have been collected as a benchmark, and 33 partial bitstreams have also been developed and integrated into the benchmark. Evaluations of the synthesizable compression decoders are demonstrated on a Xilinx ZC706 board, showing higher decompression throughput than those of the existing lossless compression decoders using our benchmark. The proposed decoders can reduce software startup time by up to 31.23% in embedded systems and 69.83% reduction of reconfiguration time for partial reconfigurable systems. Jian Yan 0002, Junqi Yuan, Philip H. W. Leong, Wayne Luk, Lingli Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Random projections for scaling machine learning on FPGAsabstractRandom projections have recently emerged as a powerful technique for large scale dimensionality reduction in machine learning applications. Crucially, the projection can be obtained from sparse probability distributions, enabling hardware implementations with little overhead. In this paper, we describe a Field-Programmable Gate Array (FPGA) implementation alongside a kernel adaptive filter (KAF) that is capable of reducing computational resources by introducing a controlled error term, achieving higher modelling capacity for given hardware resources. Empirical results involving classification, regression and novelty detection show that a 40% net increase in available resources and improvements in prediction accuracy is achievable for projections which halve the input vector length, enabling us to scale-up hardware implementations of KAF learning algorithms by at least a factor of 2. An implementation on a FPGA-based network card allows novelty detection of an 8× 24-bit input vector with latency of 404 ns, this being a 26-fold reduction compared to an Intel Core i5-2400 processor. Sean Fox, Stephen Tridgell, Craig T. Jin, Philip H. W. Leong |
FPT | 4 |
| 2016 | Feature Engineering and Supervised Learning Classifiers for Respiratory Artefact Removal in Lung Function TestsabstractA critical task in forced oscillation technique (FOT), a promising lung function test, is to remove respiratory artefacts. Manual removal by specialists is widely used but time- consuming and subjective. Most existing automated techniques have involved simple thresholding methods in an unsupervised manner. Breath cycles can be classified by a binary classification model (classes: artefactual and accepted). While attempting to use off-the-shelf sorting algorithms (e.g., one-class support vector machine, knearest neighbours, and adaptive boosting ensemble), we noticed their poor detection performance. This may result from the dependence of samples as found in physiological studies of the lung function that challenges the learning process. Specifically, statistics of breaths that we recorded may change from one to another patient and even within the same recording of a patient. We introduce an additional feature engineering step that is an intermediate module to decorrelate samples, called feature learning (using Wilcoxon signed rank tests). To that end, we collected FOT recordings from various groups of patients (paediatric and adult including healthy and asthmatics). Artefacts in this work were recorded naturally and processed in a complete-breath approach. Performance metrics include evaluations on preservation of "accepted" breaths in the filtered output (including F1- score, throughput, and approval rate). Our experiment found that our feature engineering steps significantly improve the artefact removal performance of all implemented classifiers especially with feature inputs selected by mutual information criterion. Thuy T. Pham, Diep N. Nguyen, Eryk Dutkiewicz, Alistair Lee McEwan, Cindy Thamrin, Paul D. Robinson, Philip H. W. Leong |
GLOBECOM | 7 |
| 2016 | A Microcoded Kernel Recursive Least Squares Processor Using FPGA TechnologyabstractKernel methods utilize linear methods in a nonlinear feature space and combine the advantages of both. Online kernel methods, such as kernel recursive least squares (KRLS) and kernel normalized least mean squares (KNLMS), perform nonlinear regression in a recursive manner, with similar computational requirements to linear techniques. In this article, an architecture for a microcoded kernel method accelerator is described, and high-performance implementations of sliding-window KRLS, fixed-budget KRLS, and KNLMS are presented. The architecture utilizes pipelining and vectorization for performance, and microcoding for reusability. The design can be scaled to allow tradeoffs between capacity, performance, and area. The design is compared with a central processing unit (CPU), digital signal processor (DSP), and Altera OpenCL implementations. In different configurations on an Altera Arria 10 device, our SW-KRLS implementation delivers floating-point throughput of approximately 16 GFLOPs, latency of 5.5μ S , and energy consumption of 10 − 4 J, these being improvements over a CPU by factors of 12, 17, and 24, respectively. Yeyong Pang, Yu Peng 0002, Xiyuan Peng, Nicholas J. Fraser, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2015 | Technology Scaling in FPGAs: Trends in Applications and ArchitecturesabstractSince the release of the first commercial field programmable gate array (FPGA) in 1985, devices have enjoyed continuous improvements in all metrics due to technology scaling, architectural advances and the addition of features. In this paper, we explore performance and utilization trends associated with research designs as a function of FPGA technology progression. The data used is a subset of designs presented at the IEEE International Symposium on Field-Programmable Custom Computing Machines (FCCM) over the past 20 years. These are compared to trends from theoretical and vendor sources, models generated and comparisons made. Finally, we compare operating frequency trends from our analysis to the trends exhibited by a set of vendor IP cores mapped to four generations of devices. The results of this investigation suggest that design implementations are generally following the theoretical trends and that the inclusion of embedded hard IP blocks has provided designers with additional performance benefits. Lesley Shannon, Veronica Cojocaru, Cong Nguyen Dao, Philip H. W. Leong |
FCCM | 4 |
| 2015 | A fully pipelined kernel normalised least mean squares processor for accelerated parameter optimisationabstractKernel adaptive filters (KAFs) are online machine learning algorithms which are amenable to highly efficient streaming implementations. They require only a single pass through the data during training and can act as universal approximators, i.e. approximate any continuous function with arbitrary accuracy. KAFs are members of a family of kernel methods which apply an implicit nonlinear mapping of input data to a high dimensional feature space, permitting learning algorithms to be expressed entirely as inner products. Such an approach avoids explicit projection into the feature space, enabling computational efficiency. In this paper, we propose the first fully pipelined floating point implementation of the kernel normalised least mean squares algorithm for regression. Independent training tasks necessary for parameter optimisation fill L cycles of latency ensuring the pipeline does not stall. Together with other optimisations to reduce resource utilisation and latency, our core achieves 160 GFLOPS on a Virtex 7 XC7VX485T FPGA, and the PCI-based system implementation is 70× faster than an optimised software implementation on a desktop processor. Nicholas J. Fraser, Duncan J. M. Moss, Stephen Tridgell, Craig T. Jin, Philip H. W. Leong |
FPL | 6 |
| 2015 | Significant papers from the first 25 years of the FPL conferenceabstractThe list of significant papers from the first 25 years of the Field-Programmable Logic and Applications conference (FPL) is presented in this paper. These 27 papers represent those which have most strongly influenced theory and practice in the field. Philip H. W. Leong, Hideharu Amano, Jason Helge Anderson, Koen Bertels, João M. P. Cardoso, Oliver Diessel, Guy Gogniat, Mike Hutton, Wayne Luk, Patrick Lysaght, Marco Platzner, Viktor Prasanna 0001, Tero Rissa, Cristina Silvano, Hayden Kwok-Hay So, Yu Wang 0002 |
FPL | 1 |
| 2015 | UniStream: A unified stream architecture combining configuration and data processingabstractThis paper proposes UniStream, a unified stream architecture based on point-to-point stream channels combining both bitstream configuration and data stream processing. In addition, unified APIs are provided to support bitstream configuration and data stream processing, as well as the stream interconnect. A cost model is also presented for the overhead on the stream interconnect, hardware task configuration and data stream processing at system level, which can be used during the early stage of development. The flexibility and high efficiency of UniStream are demonstrated on Xilinx Virtex-5 and Virtex-6 FPGAs. Experimental results on bitstream configuration/ read-back, data encryption/decryption and Discrete Cosine Transformation show that performance can be significantly improved with different stream modes. Jian Yan 0002, Jifang Jin, Ying Wang 0032, Xuegong Zhou, Philip H. W. Leong, Lingli Wang |
FPL | 5 |
| 2015 | Braiding: A scheme for resolving hazards in kernel adaptive filtersabstractComputational cost presents a barrier in the application of machine learning algorithms to large-scale real-time learning problems. Kernel adaptive filters (KAFs) have low computational cost with the ability to learn online and are hence favoured for such applications. Unfortunately, dependencies of the outputs on the weight updates prohibit pipelining. This paper introduces a combination of parallel execution and conditional forwarding, called braiding, which overcomes dependencies by expressing the output as a combination of the earlier state and other examples in the pipeline. To demonstrate its utility, braiding is applied to the implementation of classification, regression and novelty detection algorithms based on the Naive Online regularised Risk Minimization Algorithm (NORMA). Fixed point, open source implementations are described which can achieve data rates of around 130 MSamples/s with a latency of 10 to 13 clock cycles. This constitutes a two orders of magnitude increase in throughput and one order of magnitude decrease in latency compared to a single core CPU implementation. Stephen Tridgell, Duncan J. M. Moss, Nicholas J. Fraser, Philip H. W. Leong |
FPT | 4 |
| 2015 | Distributed kernel learning using Kernel Recursive Least SquaresabstractConstructing accurate models that represent the underlying structure of Big Data is a costly process that usually constitutes a compromise between computation time and model accuracy. Methods addressing these issues often employ parallelisation to handle processing. Many of these methods target the Support Vector Machine (SVM) and provide a significant speed up over batch approaches. However, the convergence of these methods often rely on multiple passes through the data. In this paper, we present a parallelised algorithm that constructs a model equivalent to a serial approach, whilst requiring only a single pass of the data. We first employ the Kernel Recursive Least Squares (KRLS) algorithm to construct several models from subsets of the overall data. We then show that these models can be combined using KRLS to create a single compact model. Our parallelised KRLS methodology significantly improves execution time and demonstrates comparable accuracy when compared to the parallel and serial SVM approaches. Nicholas J. Fraser, Duncan J. M. Moss, Nicolas Epain, Philip H. W. Leong |
ICASSP | 4 |
| 2015 | Phase recovery for time of arrival estimation in the presence of interferenceabstractTime of arrival is a commonly used mechanism for geolocation. In many environments, multipath propagation of RF signals limits accuracy. Wide bandwidths reduce the impact of multipath propagation, but frequently results in interference from other devices. When the bandwidth is assembled from several measured sub-bands, there is an additional difficulty of a random phase offset between bands. In this paper, we introduce a compressive sensing scheme which recovers both corrupted samples and phase offsets between sub-bands. For interferers with bandwidths up to 12 MHz, we further show that the proposed scheme leads to improved localisation accuracy compared to previously published techniques. David Humphrey, Mark Hedley, Philip H. W. Leong |
ICASSP | 4 |
| 2015 | MCALIB: Measuring Sensitivity to Rounding Error with Monte Carlo ProgrammingabstractRuntime analysis provides an effective method for measuring the sensitivity of programs to rounding errors. To date, implementations have required significant changes to source code, detracting from their widespread application. In this work, we present an open source system that automates the quantitative analysis of floating point rounding errors through the use of C-based source-to-source compilation and a Monte Carlo arithmetic library. We demonstrate its application to the comparison of algorithms, detection of catastrophic cancellation, and determination of whether single precision floating point provides sufficient accuracy for a given application. Methods for obtaining quantifiable measurements of sensitivity to rounding error are also detailed. Michael Frechtling, Philip H. W. Leong |
ACM Trans. Program. Lang. Syst. | 2 |
| 2014 | Dynamic hedging of foreign exchange risk using stochastic model predictive controlabstractA risk management system for foreign exchange (FX) brokers is described. Stochastic model predictive control (SMPC) is used to reduce positions in foreign holdings over a receding horizon, while minimising a mean-variance cost function. Computation of the broker's position incorporates elements which model client flow, transaction costs, market impact, and exchange rate. Using both synthetic and historical data, the technique is shown to outperform two simple hedging strategies on a risk-cost Pareto frontier. Prediction of client and market behaviour are shown to further enhance the hedging outcome. Farzad Noorian, Philip H. W. Leong |
CIFEr | 2 |
| 2014 | SMCGen: Generating Reconfigurable Design for Sequential Monte Carlo ApplicationsabstractThe Sequential Monte Carlo (SMC) method is a simulation-based approach to compute posterior distributions. SMC methods often work well on applications considered intractable by other methods due to high dimensionality, but they are computationally demanding. While SMC has been implemented efficiently on FPGAs, design productivity remains a challenge. This paper introduces a design flow for generating efficient implementation of reconfigurable SMC designs. Through templating the SMC structure, the design flow enables efficient mapping of SMC applications to multiple FPGAs. The proposed design flow consists of a parametrisable SMC computation engine, and an open-source software template which enables efficient mapping of a variety of SMC designs to reconfigurable hardware. Design parameters that are critical to the performance and to the solution quality are tuned using a machine learning algorithm based on surrogate modelling. Experimental results for three case studies show that design performance is substantially improved after parameter optimisation. The proposed design flow demonstrates its capability of producing reconfigurable implementations for a range of SMC applications that have significant improvement in speed and in energy efficiency over optimised CPU and GPU implementations. Thomas C. P. Chau, Maciej Kurek, James Stanley Targett, Jake Humphrey, Georgios Skouroupathis, Alison Eele, Jan M. Maciejowski, Benjamin Cope, Kathryn Cobden, Philip H. W. Leong, Peter Y. K. Cheung, Wayne Luk |
FCCM | 10 |
| 2014 | An FPGA-based spectral anomaly detection systemabstractAnomaly detection based on spectral features is applicable to a diverse range of problems including prognostic and health management, vibration analysis, astronomy, biomedicai engineering and computational finance. The input data could be regularly sampled, as in the case of a standard analogue to digital converter sampling a bandlimited signal at above the Nyquist rate, or irregularly sampled, as in the case of stock quotes or astronomical data. In this paper, we present new online algorithms for the computation of power spectra for regularly or irregularly sampled data, and performing anomaly detection on time series data. Both algorithms allow hardware implementations with O(l) time complexity, this being the minimum for any system that considers all the samples. We combine the two algorithms to form a power Spectrum-based Anomaly Detector (SAD). We also describe an implementation of SAD which has minimal hardware requirements, and achieves one to two orders of magnitude improvement in speed, latency, power and energy over a traditional processor-based design. Duncan J. M. Moss, Nicholas J. Fraser, Philip H. W. Leong |
FPT | 4 |
| 2014 | Design space exploration for FPGA-based hybrid multicore architectureabstractThis paper presents a parameterized system-level design framework, which enables rapid and powerful research for hybrid multicore architecture exploration and hardware/software co-design. The framework comprises the component-based hardware design and application compiler, which make it easy for a designer to build stream-oriented applications with FPGA-based hybrid multicore architectures. The high modularity and parameterization of the framework supports fast multicore architecture exploration of different topologies, routing schemes, processor types, customized hardware processing units and memory system organizations. The compiler tool chain is used to map C/C++ based applications onto the soft processing units. Experimental results targeting the JPEG encoding application demonstrate the feasibility and performance improvement of this framework. Jian Yan 0002, Junqi Yuan, Ying Wang 0032, Philip H. W. Leong, Lingli Wang |
FPT | 4 |
| 2013 | Cluster analysis of high-dimensional high-frequency financial time seriesabstractRecently the availability of tick data is driving renewed interest in statistical tools for the analysis of high-dimensional irregularly spaced time series. Since the standard tools require that the data are evenly spaced, the traditional multivariate time series analysis techniques are inadequate for the analysis of tick data. We develop for perhaps the first time a proper procedure that performs cluster analysis of tick data using the joint information of the temporal process and the continuous-valued data at the actual sampling times. A simulation example studies the problem with the standard approach and demonstrates the reliability of our proposed method. Data analyses of major stock market indices and currencies are provided. Syed Ahmed Pasha, Philip H. W. Leong |
CIFEr | 2 |
| 2013 | A low latency kernel recursive least squares processor using FPGA technologyabstractThe kernel recursive least squares (KRLS) algorithm performs non-linear regression in an online manner, with similar computational requirements to linear techniques. In this paper, an implementation of the KRLS algorithm utilising pipelining and vectorisation for performance; and microcoding for reusability is described. The design can be scaled to allow tradeoffs between capacity, performance and area. Compared with a central processing unit (CPU) and digital signal processor (DSP), the processor improves on execution time, latency and energy consumption by factors of 5, 5 and 12 respectively. Yeyong Pang, Yu Peng 0002, Nicholas J. Fraser, Philip H. W. Leong |
FPT | 5 |
| 2013 | A Hybrid Feature Selection and Generation Algorithm for Electricity Load Prediction Using Grammatical EvolutionabstractAccurate load prediction plays a major role in devising effective power system control strategies. Successful prediction systems often use machine learning (ML) methods. The success of ML methods, among other things, depends on a suitable choice of input features which are usually selected by domain-experts. In this paper, we propose a novel systematic way of generating and selecting better features for daily peak electricity load prediction using kernel methods. Grammatical evolution is used to evolve an initial population of well performing individuals, which are subsequently mapped to feature subsets derived from wavelets and technical indicator type formulae used in finance. It is shown that the generated features can improve results, while requiring no domain-specific knowledge. The proposed method is focused on feature generation and can be applied to a wide range of ML architectures and applications. Anthony Mihirana De Silva, Farzad Noorian, Richard I. A. Davis, Philip H. W. Leong |
ICMLA (2) | 4 |
| 2013 | Architecture and Design Flow for a Highly Efficient Structured ASICabstractAs fabrication process technology continues to advance, mask set costs have become prohibitively expensive. Structured application specific integrated circuits (sASICs) offer a middle ground in price and performance between ASICs and field-programmable gate arrays (FPGAs) by sharing masks across different designs. In this paper, two sASIC architectures are proposed, the first being based on three-input lookup-tables, and the second on AOI22 gates. The sASICs are programmed using a standard-cell compatible design flow. They are customized using a minimum of three masks, i.e., two metals and one via. The area and delay of the sASIC are compared with ASICs and FPGAs. Results over a set of benchmark circuits show that our AOI22-based sASIC had an average of 1.76x/1.41x increase in area/delay compared to ASICs, a considerable improvement compared with the 26.56x/5.09x increase for FPGAs. This is, to the best of our knowledge, the best performance reported in the literature for a practical sASIC. A prototype using the sASIC was fabricated using a universal machine control 0.13-μm mixed-mode/RF process. It was fully verified using scan and functional tests, and used in a demonstration system. S. Man Ho Ho, Yanqing Ai, Thomas C. P. Chau, Steve C. L. Yuen, Oliver Chiu-sing Choy, Philip H. W. Leong, Kong-Pang Pun |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2012 | A Mixed Precision Methodology for Mathematical OptimisationabstractThis paper introduces a novel mixed precision methodology for mathematical optimisation. It involves the use of reduced precision FPGA optimisers for searching potential regions containing the global optimum, and double precision optimisers on a general purpose processor (GPP) for verifying the results. An empirical method is proposed to determine parameters of the mixed precision methodology running on a reconfigurable accelerator consisting of FPGA and GPP. The effectiveness of our approach is evaluated using a set of optimisation benchmarks. Using our mixed precision methodology and a modern reconfigurable accelerator, we can locate the global optima 1.7 to 6 times faster compared with quad-core optimiser. The mixed precision optimisations search up to 40.3 times more starting vector per unit time compared with quad core optimisers and only 0.7% to 2.7% of these searches are refined using GPP double precision optimisers. The proposed methodology also allows us to accelerate problems with more complicated functions or to solve problems involving higher dimensions. Gary C. T. Chow, Wayne Luk, Philip H. W. Leong |
FCCM | 3 |
| 2012 | A mixed precision Monte Carlo methodology for reconfigurable accelerator systemsabstractThis paper introduces a novel mixed precision methodology applicable to any Monte Carlo (MC) simulation. It involves the use of data-paths with reduced precision, and the resulting errors are corrected by auxiliary sampling. An analytical model is developed for a reconfigurable accelerator system with a field-programmable gate array (FPGA) and a general purpose processor (GPP). Optimisation based on mixed integer geometric programming is employed for determining the optimal reduced precision and optimal resource allocation among the MC data-paths and correction datapaths. Experiments show that the proposed mixed precision methodology requires up to 11 % additional evaluations while less than 4 % of all the evaluations are computed in the reference precision; the resulting designs are up to 7.1 times faster and 3.1 times more energy efficient than baseline double precision FPGA designs, and up to 163 times faster and 170 times more energy efficient than quad-core software designs optimised with the Intel compiler and Math Kernel Library. Our methodology also produces designs for pricing Asian options which are 4.6 times faster and 5.5 times more energy efficient than NVIDIA Tesla C2070 GPU implementations. Gary Chun Tak Chow, Anson H. T. Tse, Qiwei Jin, Wayne Luk, Philip H. W. Leong, David B. Thomas |
FPGA | 5 |
| 2012 | Hardware efficient parallel particle filter for tracking in wireless networksabstractLocation tracking is being increasingly used across many applications. While GPS is the most widely used location tracking technology, it is unavailable in many environments such as indoors and underground. Local positioning systems (LPS) that use time of arrival based ranging can provide high accuracy location tracking for many applications. Tracking location using range measurements is a non-linear state estimation problem and the measurement noise is often non-Gaussian in environments where LPS are typically used. Hence a particle filter is an appropriate state estimator for location tracking in LPS. Particle filters are computationally complex and have a serial bottleneck that prevents straightforward parallel implementation. In this paper we present a parallel architecture for the particle filter that can be efficiently implemented in a field programmable gate array or a fixed-point digital signal processor. We show that processing can be divided into up to twenty parallel particle filters to massively increase the processing speed. Mixing between the filters is essential and we present a new algorithm for this that minimises computational complexity and memory bandwidth. Finally, for efficient hardware implementation fixed point arithmetic should be used and we empirically determine the required precision. Yi Qiao Zhang, Thuraiappah Sathyan, Mark Hedley, Philip H. W. Leong, Syed Ahmed Pasha |
PIMRC | 4 |
| 2012 | An FPGA Chip Identification Generator Using Configurable Ring OscillatorsabstractPhysically unclonable functions (PUF) are commonly used in applications such as hardware security and intellectual property protection. Various PUF implementation techniques have been proposed to translate chip-specific variations into a unique binary string. It is difficult to maintain repeatability of chip ID generation, especially over a wide range of operating conditions. To address this problem, we propose utilizing configurable ring oscillators and an orthogonal re-initialization scheme to improve repeatability. An implementation on a Xilinx Spartan-3e field-programmable gate array was tested on nine different chips. Experimental results show that the bit flip rate is reduced from 1.5% to approximately 0 at a fixed supply voltage and room temperature. Over a 20°C-80°C temperature range and 25% variation in supply voltage, the bit flip rate is reduced from 1.56% to 3.125×10-7. Haile Yu, Philip H. W. Leong, Qiang Xu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Optimizing Floating Point Units in Hybrid FPGAsabstractThis paper introduces a methodology to optimize coarse-grained floating point units (FPUs) in a hybrid field-programmable gate array (FPGA), where the FPU consists of a number of interconnected floating point adders/subtracters (FAs), multipliers (FMs), and wordblocks (WBs). The wordblocks include registers and lookup tables (LUTs) which can implement fixed point operations efficiently. We employ common subgraph extraction to determine the best mix of blocks within an FPU and study the area, speed and utilization tradeoff over a set of floating point benchmark circuits. We then explore the system impact of FPU density and flexibility in terms of area, speed, and routing resources. Finally, we derive an optimized coarse-grained FPU by considering both architectural and system-level issues. This proposed methodology can be used to evaluate a variety of FPU architecture optimizations. The results for the selected FPU architecture optimization show that although high density FPUs are slower, they have the advantages of improved area, area-delay product, and throughput. Chi Wai Yu, Alastair M. Smith, Wayne Luk, Philip H. W. Leong, Steve Wilton |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | Mixed Precision Processing in Reconfigurable SystemsabstractCustomisable data formats provide an opportunity for exploring trade-offs in accuracy and performance of reconfigurable systems. This paper introduces a novel methodology for mixed-precision comparison, which improves comparison performance by using reduced-precision data paths while maintaining accuracy by using high-precision data paths. Our methodology adopts reduced-precision data-paths for preliminary comparison, and high-precision data-paths when the accuracy for preliminary comparison is insufficient. We develop an analytical model for performance estimation of the proposed mixed-precision methodology. Optimisation based on integer linear programming is employed for determining the optimal precision and resource allocation for each of the data paths. The effectiveness of our approach is evaluated using a common collision detection problem. Performance gains of 4 to 7.3 times are obtained over baseline fixed-precision designs for the same FPGAs. With the help of the proposed mixed-precision methodology, our FPGA designs are 15.4 to 16.7 times faster than software running on multi-core CPUs with the same technology. Gary Chun Tak Chow, K. W. Kwok, Wayne Luk, Philip H. W. Leong |
FCCM | 4 |
| 2011 | A Model for Peak Matrix Performance on FPGAs
Colin Yu Lin, Hayden Kwok-Hay So, Philip H. W. Leong |
FCCM | 3 |
| 2011 | A monte-carlo floating-point unit for self-validating arithmeticabstractMonte-Carlo arithmetic is a form of self-validating arithmetic that accounts for the effect of rounding errors. We have implemented a floating point unit that can perform either IEEE 754 or Monte-Carlo floating point computation, allowing hardware accelerated validation of results during execution. Experiments show that our approach has a modest hardware overhead and allows the propagation of rounding error to be accurately estimated. Jackson H. C. Yeung, Evangeline F. Y. Young, Philip H. W. Leong |
FPGA | 3 |
| 2011 | On timing yield improvement for FPGA designs using architectural symmetry (abstract only)abstractAs semiconductor manufacturing technology continues towards reduced feature sizes, timing yield will degrade due to increased process variation. Traditional variation aware design (VAD) methodologies address this problem by using chipwise placement and routing optimizations given the variation distribution is obtained. However, it is very time-consuming to do chipwise variation characterization and optimization. Therefore, this work proposes the use of symmetry in FPGA architectures so that a large range of timing-equivalent configurations can be derived from a single initial implementation by configuration rotation and flipping, allowing the application of post-silicon tuning to mitigate the effects of process variation. Additionally, logic element swaps further improve timing performance. An FPGA design methodology is presented which combines configuration-level redundancy and fine-grained design tuning. The proposed methodology does not need variation characterization and customized placement and routing for each individual FPGA. Compared to other variation aware design methods, it is more cost-efficient in terms of run-time, especially for design implementation on a large amount of FPGAs. Twenty MCNC benchmark circuits in different process technologies were used to show that the proposed method is effective in improving yield and timing in the presence of process variation. Haile Yu, Qiang Xu 0001, Philip H. W. Leong |
FPGA | 3 |
| 2011 | A Model for Matrix Multiplication Performance on FPGAsabstractComputations involving matrices form the kernel of a large spectrum of computationally demanding applications for which FPGAs have been utilized as accelerators. Their performance is related to their underlying architectural and system parameters such as computational resources, memory and I/O bandwidth. A simple analytic model that gives an estimate of the performance of FPGA-based sparse matrix-vector and matrix-matrix multiplication is presented, dense matrix multiplication being a special case. The efficiency of existing implementations are compared to the model and performance trends for future technologies examined. Colin Yu Lin, Hayden Kwok-Hay So, Philip H. W. Leong |
FPL | 3 |
| 2011 | On Timing Yield Improvement for FPGA Designs Using Architectural SymmetryabstractAs semiconductor manufacturing technology continues towards reduced feature sizes, timing yield will degrade due to increased process variation. This work proposes the use of architectural symmetry in FPGA so that multiple timing-equivalent configurations can be derived from a single initial implementation, allowing the application of post-silicon tuning to mitigate process variation effects. Experimental results on twenty MCNC benchmark circuits for various process technologies demonstrate timing yield improvement using the proposed method. Haile Yu, Qiang Xu 0001, Philip H. W. Leong |
FPL | 3 |
| 2011 | Spiking neural network-based auto-associative memory using FPGA interconnect delaysabstractThis paper describes the design of an auto-associative memory based on a spiking neural network (SNN). The architecture is able to effectively utilize the massive interconnect resources available in FPGA architectures as a good match to the axons in biological neural networks. A complete implementation of the memory on a single FPGA is presented. The signal processing circuitry is composed from simple, parallel building blocks and the training logic is implemented using an on-chip soft processor. Chong H. Ang, Craig T. Jin, Philip H. W. Leong, André van Schaik |
FPT | 3 |
| 2011 | An Analytical Model Relating FPGA Architecture to Logic Density and DepthabstractThis paper presents an analytical model that relates FPGA architectural parameters to the logic size and depth of an FPGA implementation. In particular, the model relates the lookup-table size, the cluster size, and the number of inputs per cluster to the amount of logic that can be packed into each lookup-table and cluster, the number of used inputs per cluster, and the depth of the circuit after technology mapping and clustering. Comparison to experimental results shows that our model has good accuracy. We illustrate how the model can be used in FPGA architectural investigations to complement the experimental approach. The model's accuracy, combined with the simple form of the equations, make them a powerful tool for FPGA architects to better understand and guide the development of future FPGA architectures. Joydip Das, Andrew Lam, Steve Wilton, Philip H. W. Leong, Wayne Luk |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2010 | Rapid prototyping on a structured ASIC fabricabstractWe describe the architecture of a structured ASIC fabric in which the logic and routing can be customized using three masks. A standard Cadence based design flow is employed, and using an active dynamic backlight controller as an example, performance is compared to that of an ASIC implementation in the same technology. Steve C. L. Yuen, Yanqing Ai, Brian P. W. Chan, Thomas C. P. Chau, Sam M. H. Ho, Oscar K. L. Lau, Kong-Pang Pun, Philip H. W. Leong, Oliver Chiu-sing Choy |
ASP-DAC | 8 |
| 2010 | Design of a single layer programmable Structured ASIC libraryabstractA Structured Application-specific Integrated Circuit (SASIC) is a programmable fabric in which a small set of masks are customized for a particular application, serving to reduce the associated non-recurring engineering cost (NRE). In this paper we describe the implementation of a SASIC logic cell which is programmable via a single metal layer. A SASIC fabric prototype is fabricated and all implemented functions are verified on silicon. Experimental measurement verifies correct operation of our SASIC with a clock frequency of over 250 MHz. Thomas C. P. Chau, David W. L. Wu, Yanqing Ai, Brian P. W. Chan, Sam M. H. Ho, Oscar K. L. Lau, Steve C. L. Yuen, Kong-Pang Pun, Oliver Chiu-sing Choy, Philip H. W. Leong |
DDECS | 10 |
| 2010 | A Karatsuba-Based Montgomery MultiplierabstractModular multiplication of long integers is an important building block for cryptographic algorithms. Although several FPGA accelerators have been proposed for large modular multiplication, previous systems have been based on O(N2) algorithms. In this paper, we present a Montgomery multiplier that incorporates the more efficient Karatsuba algorithm which is O(N(log 3/log 2)). This system is parameterizable to different bitwidths and makes excellent use of both embedded multipliers and fine-grained logic. The design has significantly lower LUT-delay product and multiplier-delay product compared with previous designs. Initial testing on a Virtex-6 FPGA showed that it is 60-190 times faster than an optimized multi-threaded software implementation running on an Intel Xeon 2.5 GHz CPU. The proposed multiplier system is also estimated to be 95-189 times more energy efficient than the software-based implementation. This high performance and energy efficiency makes it suitable for server-side applications running in a datacenter environment. Gary Chun Tak Chow, Kenneth Eguro, Wayne Luk, Philip H. W. Leong |
FPL | 4 |
| 2010 | Structured ASIC: Methodology and comparisonabstractAs fabrication process technology continues to advance, mask set costs have become prohibitively expensive. Structured ASICs can offer price and performance between ASICs and FPGAs. They are attractive for mid-volume production and offer good intellectual property security. In this paper, a structured ASIC methodology, where 2 metal- and 1 via-mask are customised, is described. The CAD tools are fully compatible with conventional ASIC design flows and a comparison of area and delay performance with ASICs and FPGAs is given. A prototype structured ASIC implementing an LED-backlit LCD controller was fabricated in a 0.13 μm CMOS process. It was verified and power consumption compared with an ASIC design. Sam M. H. Ho, Steve C. L. Yuen, Hiu Ching Poon, Thomas C. P. Chau, Yanqing Ai, Philip H. W. Leong, Oliver Chiu-sing Choy, Kong-Pang Pun |
FPT | 6 |
| 2010 | An FPGA chip identification generator using configurable ring oscillatorabstractAn improved chip identification (ID) generator, otherwise known as a physically unclonable function (PUF) is described. Similar to previous designs, a cell, i, is used to obtain a measure of the difference in period of four ring oscillators and obtain the residue Ri, a random variable. Experiments show it is normally distributed with a mean of 0. A binary output value of 0 or 1 assigned depending on the sign of Ri. When |E(Ri)| is large, this scheme consistently gives the same output. Unfortunately, when it is small, the repeatability is compromised, particularly when variations in operating conditions such as supply voltage and temperature are also taken into account, which is a common problem for all previous works. To address this problem, we propose a cell with configurable ring oscillators together with an orthogonal re-initialisation scheme. Together, these two techniques maximise repeatability by causing the distribution of the mean of different Ri's to change from normal to bimodal. We implement this design in the Xilinx Spartan-3e FPGA. Nine FPGA chips are tested, and experimental results show that the new method significantly enhances reliability of ID generation and tolerance to environmental changes. Bit flip rate is reduced from 1.5% to approximately 0 at a fixed supply voltage and room temperature. Over the 20 - 80°C temperature range, and a 25% variation in supply voltage, the bit flip rate is reduced from 1.56% to 3.125 × 10-7, which is a 50000x improvement. Haile Yu, Philip H. W. Leong, Qiang Xu 0001 |
FPT | 2 |
| 2010 | Fine-grained characterization of process variation in FPGAsabstractAs semiconductor manufacturing continues towards reduced feature sizes, yield loss due to process variation becomes increasingly important. To address this issue on FPGA platforms, several variation aware design (VAD) methodologies have been proposed. In this work we present a practical method of process variation characterization (PVC) to facilitate VAD using only intrinsic FPGA resources. The scheme is based on measuring the difference between ring oscillator (RO) delay at different locations within a die, and can be used to perform process variation characterization for LE delays and interconnect delays including direct connection, double wire and hex wires. The difference in loop delays can also be estimated from equations using parameters extracted from primitives and compared with direct measurements. On a Xilinx Spartan-3e device, it was found that the error between the estimated and measured values was on average less than 10%. Haile Yu, Qiang Xu 0001, Philip H. W. Leong |
FPT | 3 |
| 2009 | A comparison of via-programmable gate array logic cell circuitsabstractVia-programmable gate arrays (VPGAs) offer a middle ground between application specific integrated circuits and field programmable gate arrays in terms of flexibility, manufactuing cost, speed, power and area. In this paper, we present a novel VPGA logic cell, the complementary universal logic gate (CULG) which can be used to implement both sequential and combinatorial elements. Its performance is compared with a number of other designs including transmission gate, differential cascode voltage switch with pass gate, and standard cell. The CULG is found to have comparable power-delay product and process variation sensitivity to the other designs while offering the lowest power consumption. Thomas C. P. Chau, Philip H. W. Leong, Sam M. H. Ho, Brian P. W. Chan, Steve C. L. Yuen, Kong-Pang Pun, Oliver Chiu-sing Choy, Xinan Wang |
FPGA | 2 |
| 2009 | Modeling post-techmapping and post-clustering FPGA circuit depthabstractThis paper presents an analytical model that relates FPGA architectural parameters to the expected speed of FPGA implementation. More precisely, the model relates the lookup-table size, cluster size, and number of inputs per cluster to the depth of the circuit after technology mapping and after clustering. Comparison to experimental results with large MCNC circuits shows that our models are accurate. We show how the models can be used in FPGA architectural investigations to complement the more usual experimental approach. Joydip Das, Steve Wilton, Philip H. W. Leong, Wayne Luk |
FPL | 3 |
| 2009 | Towards a unique FPGA-based identification circuit using process variationsabstractA compact chip identification (ID) circuit with improved reliability is presented. Ring oscillators are used to measure the spatial process variation and the ID is based on their relative speeds. A novel averaging and postprocessing scheme is employed to accurately determine the faster of two similar-frequency ring oscillators in the presence of noise. Using this scheme, the average number of unstable bits i.e. bits which can change in value between readings, measured on an FPGA is shown to be reduced from 5.3% to 0.9% at 20degC. Within the range 20-60degC, the percentage of unstable bits is within 2.8%. An analysis of the effectiveness of the scheme and the distribution of the errors is given over different temperature ranges and FPGA chips. Haile Yu, Philip H. W. Leong, Heiko Hinkelmann, Leandro Möller, Manfred Glesner, Peter Zipf |
FPL | 2 |
| 2009 | A detailed delay path model for FPGAsabstractA complete circuit-level description of a representative FPGA is presented in this paper, from which a simple RC delay model as a function of architectural and technology parameters is derived. Using this model, the expression for the optimal delay of any path through the FPGA can be formulated. We distill our model into being purely architecture dependent, and use it to capture new insight into how FPGA parameters can directly affect its delay. Several applications of this model are: (1) to gain better intuition of how architecture and process parameters affect the delay path in an FPGA, (2) for initial studies into new circuit designs and integrated circuit technologies, (3) in CAD tools for optimisation and sensitivity analysis. The technique described can be applied to arbitrary circuits, and simulations show that our closed form equations give delay values that are accurate to approximately 10% when compared to HSPICE simulation. Eddie Hung, Steve Wilton, Haile Yu, Thomas C. P. Chau, Philip H. W. Leong |
FPT | 5 |
| 2009 | Generation of Synthetic Floating-Point benchmark circuitsabstractSynthetic Floating-Point (SFP), a synthetic benchmark generator program for floating-point circuits is presented. SFP consists of two independent modules for characterisation and generation. The characterisation module extracts key dataflow statistics of an arbitrary software program. Generation involves producing randomised circuits with desired statistics which are either the output of the characterisation module or directly generated by the user. Using the basic linear algebra subprograms (BLAS) library, Whetstone benchmark and LINPACK benchmark, it is demonstrated that SFP can be used to generate floating-point benchmarks with different user-specified properties as well as benchmarks that mimic real computational programs. Thomas C. P. Chau, S. Man Ho Ho, Philip H. W. Leong, Peter Zipf, Manfred Glesner |
IPDPS | 3 |
| 2009 | Floating-Point FPGA: Architecture and ModelingabstractThis paper presents an architecture for a reconfigurable device that is specifically optimized for floating-point applications. Fine-grained units are used for implementing control logic and bit-oriented operations, while parameterized and reconfigurable word-based coarse-grained units incorporating word-oriented lookup tables and floating-point operations are used to implement datapaths. In order to facilitate comparison with existing FPGA devices, the virtual embedded block scheme is proposed to model embedded blocks using existing field-programmable gate array (FPGA) tools. This methodology involves adopting existing FPGA resources to model the size, position, and delay of the embedded elements. The standard design flow offered by FPGA and computer-aided design vendors is then applied and static timing analysis can be used to estimate the performance of the FPGA with the embedded blocks. On selected floating-point benchmark circuits, our results indicate that the proposed architecture can achieve four times improvement in speed and 25 times reduction in area compared with a traditional FPGA device. Chun Hok Ho, Chi Wai Yu, Philip H. W. Leong, Wayne Luk, Steve Wilton |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | Map-reduce as a Programming Model for Custom Computing MachinesabstractThe map-reduce model requires users to express their problem in terms of a map function that processes single records in a stream, and a reduce function that merges all mapped outputs to produce a final result. By exposing structural similarity in this way, a number of key issues associated with the design of custom computing machines including parallelisation; design complexity; software-hardware partitioning; hardware-dependency, portability and scalability can be easily addressed. We present an implementation of a map-reduce library supporting parallel field programmable gate arrays (FPGAs) and graphics processing units (GPUs). Parallelisation due to pipelining, multiple data paths and concurrent execution of FPGA/GPU hardware is automatically achieved. Users first specify the map and reduce steps for the problem in ANSI Cand no knowledge of the underlying hardware or parallelisation is needed. The source code is then manually translated into a pipelined data path which, along with the map-reduce library, is compiled into appropriate binary configurations for the processing units. We describe our experience in developing a number of benchmark problems in signal processing, Monte Carlo simulation and scientific computing as well as report on the performance of FPGA, GPU and heterogeneous systems. Jackson H. C. Yeung, C. C. Tsang, Kuen Hung Tsoi, Bill S. H. Kwan, Chris C. C. Cheung, Anthony P. C. Chan, Philip H. W. Leong |
FCCM | 7 |
| 2008 | FPGA interconnect design using logical effortabstractLogical effort (LE) is a linear technique for modeling the delay of a circuit in a technology independent manner. It offers the potential to simplify delay models for FPGAs and gain more insight into how the parameters affect the result. In this paper, the LE model will be introduced and an application to FPGA interconnect driver sizing described. Simple closed form equations are given for delay, sensitivity of delay to driver size and optimal delay. The results are shown to closely agree with Spice simulation Haile Yu, Yuk Hei Chan, Philip H. W. Leong |
FPGA | 3 |
| 2008 | Rapid estimation of power consumption for hybrid FPGAsabstractA hybrid FPGA consists of island-style fine-grained units and domain-specific coarse-grained units. This paper describes an approach to estimate the power consumption of a set of hybrid FPGA architectures. The dynamic power consumption of the fine-grained units is obtained using standard FPGA tools, and the coarse-grained units using standard ASIC tools. Based on this approach, the dynamic power consumption of different hybrid FPGA architectures can be studied and we report on results over a set of floating point benchmark circuits. Chun Hok Ho, Philip H. W. Leong, Wayne Luk, Steve Wilton |
FPL | 2 |
| 2008 | Mapping and scheduling with task clustering for heterogeneous computing systemsabstractThis paper presents a new approach for mapping task graphs to heterogeneous hardware/software computing systems using heuristic search techniques. Two techniques: (1) integration of clustering, mapping, and scheduling in a single step and (2) multiple neighborhood functions strategy are proposed to enhance quality of mapping/scheduling solutions. Our approach is demonstrated by case studies involving 40 randomly generated task graphs, as well as four real applications including signal processing and pattern recognition. Experimental results show that the proposed integrated approach outperforms a separate approach in terms of quality of the mapping/scheduling solution by up to 18.3% for a heterogeneous system which includes a microprocessor, a floating-point digital signal processor, and an FPGA. Yuet Ming Lam, José Gabriel F. Coutinho, Wayne Luk, Philip H. W. Leong |
FPL | 4 |
| 2008 | An analytical model describing the relationships between logic architecture and FPGA densityabstractThis paper describes an analytical model, based principally on Rentpsilas Rule, that relates logic architectural parameters to the area efficiency of an FPGA. In particular, the model relates the lookup-table size, the cluster size, and the number of inputs per cluster to the amount of logic that can be packed into each lookup-table and cluster, and the number of used inputs per cluster. Comparison to experimental results show that our models are accurate. This accuracy combined with the simple form of the equations make them a powerful tool for FPGA architects to better understand and guide the development of future FPGA architectures. Andrew Lam, Steve Wilton, Philip H. W. Leong, Wayne Luk |
FPL | 3 |
| 2008 | FPGA interconnect design using logical effortabstractLogical effort (LE) is a linear technique for modelling the delay of a circuit in a technology independent manner. It offers the potential to simplify delay models for FPGAs and gain more insight into how the parameters affect the result. In this paper, the LE model will be introduced and an application to FPGA interconnect driver sizing described. Simple closed form equations are given for delay, sensitivity of delay to driver size and optimal delay. The results are shown to closely agree with Spice simulations. Haile Yu, Yuk Hei Chan, Philip H. W. Leong |
FPL | 3 |
| 2008 | Unrolling-based loop mapping and schedulingabstractThis paper presents an loop unrolling based mapping and scheduling strategy to maximum the parallelism of an application described as task graph targeting on a heterogeneous computing systems. Loops are statically unrolled using compile-time parameters and dynamic tasks are generated to handle run-time conditions, such that the closer the match of run-time conditions and compile-time parameters, the higher the performance. Experimental results obtained using a speech recognition system show the proposed method outperforms an approach without unrolling by 2.1 times, and using the processing time of a 2.6 GHz microprocessor as a reference, a speed up of 10 times can be achieved when compile-time and run-time parameters are matched, while the performance drops gradually when they are different. Yuet Ming Lam, José Gabriel F. Coutinho, Wayne Luk, Philip H. W. Leong |
FPT | 4 |
| 2008 | Optimizing coarse-grained units in floating point hybrid FPGAabstractThis paper introduces a novel methodology to optimize coarse-grained floating point units (FPUs) in a hybrid FPGA. We employ common subgraph extraction to determine the number of floating point adders/subtracters (FAs), multipliers (FMs) and wordblocks (WBs) in the FPUs. We flrst study the area, speed and utilization trade-off of the selected FPU subgraphs in a set of floating point benchmark circuits. We then explore the impact of density and flexibility of FPUs on the system in terms of area, speed and routing resources. We derive an optimized coarse-grained FPU by considering both architectural and system level issues. The results show that: (1) embedding more types of coarse-grained FPU in the system causes at most 21.3% increase in delay, (2) the area of the system can be reduced by 27.4% by embedding high density subgraphs, (3) the high density subgraphs requires 14.8% fewer routing resources. Chi Wai Yu, Alastair M. Smith, Wayne Luk, Philip H. W. Leong, Steve Wilton |
FPT | 4 |
| 2008 | A Synthesizable Datapath-Oriented Embedded FPGA Fabric for Silicon Debug ApplicationsabstractWe present an architecture for a synthesizable datapath-oriented FPGA core that can be used to provide post-fabrication flexibility to an SoC. Our architecture is optimized for bus-based operations and employs a directional routing architecture, which allows it to be synthesized using standard ASIC design tools and flows. The primary motivation for this architecture is to provide an efficient mechanism to support on-chip debugging. The fabric can also be used to implement other datapath-oriented circuits such as those needed in signal processing and computation-intensive applications. We evaluate our architecture using a set of benchmark circuits and compare it to previous fabrics in terms of area, speed, and power. Steve Wilton, Chun Hok Ho, Bradley R. Quinton, Philip H. W. Leong, Wayne Luk |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2007 | A synthesizable datapath-oriented embedded FPGA fabricabstractWe present an architecture for a synthesizable datapath-oriented Field Programmable Gate Array (FPGA) core which can be used to provide post-fabrication flexibility to a System-on-Chip (SoC). Our architecture is optimized for bus-based operations that are common in signal processing and computation intensive applications. It employs a directional routing architecture, which allows it to be synthesized using standard ASIC design tools and flows. We also describe a proof-of-concept layout of our core. It is shown that the proposed architecture is significantly more area efficient than the best previously reported synthesizable programmable logic core. Steve Wilton, Chun Hok Ho, Philip H. W. Leong, Wayne Luk, Bradley R. Quinton |
FPGA | 3 |
| 2007 | Domain-Specific Hybrid FPGA: Architecture and Floating Point ApplicationsabstractThis paper presents a novel architecture for domain-specific FPGA devices. This architecture can be optimised for both speed and density by exploiting domain-specific information to produce efficient reconfigurable logic with multiple granularity. In the reconfigurable logic, general-purpose finegrained units are used for implementing control logic and bit-oriented operations, while domain-specific coarse-grained units and heterogeneous blocks are used for implementing datapaths; the precise amount of each type of resources can be customised to suit specific application domains. Issues and challenges associated with the design flow and the architecture modelling are addressed. Examples of the proposed architecture for speeding up floating point applications are illustrated. Current results indicate that the proposed architecture can achieve 2.5 times improvement in speed and 18 times reduction in area on average, when compared with traditional FPGA devices on selected floating point benchmark circuits. Chun Hok Ho, Chi Wai Yu, Philip H. W. Leong, Wayne Luk, Steve Wilton |
FPL | 3 |
| 2006 | Hardware efficient architectures for Eigenvalue computationabstractEigenvalue computation is essential in many fields of science and engineering. For high performance and real-time applications, this may need to be done in hardware. This paper focuses on the exploration of hardware architectures which compute eigenvalues of symmetric matrices. We propose to use the approximate Jacobi method for general case symmetric matrix eigenvalue problem. The paper illustrates that the proposed architecture is more efficient than previous architectures reported in the literature. Moreover, for the special case of 3times3 symmetric matrices, we propose to use an algebraic method. It is shown that the pipelined architecture based on the algebraic method has a significant advantage in terms of area Christos-Savvas Bouganis, Peter Y. K. Cheung, Philip H. W. Leong, Stephen J. Motley |
DATE | 4 |
| 2006 | Virtual Embedded Blocks: A Methodology for Evaluating Embedded Elements in FPGAsabstractEmbedded elements, such as block multipliers, are increasingly used in advanced field programmable gate array (FPGA) devices to improve efficiency in speed, area and power consumption. A methodology is described for assessing the impact of such embedded elements on efficiency. The methodology involves creating dummy elements, called virtual embedded blocks (VEBs), in the FPGA to model the size, position and delay of the embedded elements. The standard design flow offered by FPGA and CAD vendors can be used for mapping, placement, routing and retiming of designs with VEBs. The speed and resource utilisation of the resulting designs can then be inferred using the FPGA vendor's timing analysis tools. We illustrate the application of this methodology to the evaluation of various schemes of involving embedded elements that support floating-point computations Chun Hok Ho, Philip H. W. Leong, Wayne Luk, Steve Wilton, Sergio López-Buedo |
FCCM | 2 |
| 2006 | FPGA Based Acceleration of the Linpack Benchmark: A High Level Code Transformation ApproachabstractDue to their increasing resource densities, field programmable gate arrays (FPGAs) have become capable of efficiently implementing large scale scientific applications involving floating point computations. In this paper FPGAs are compared to a high end microprocessor with respect to sustained performance for a popular floating point CPU performance benchmark, namely LINPACK 1000. A set of translation and optimization steps have been applied to transform a sequential C description of the LINPACK benchmark, based on a monolithic memory model, into a parallel Handel-C description that utilizes the plurality of memory resources available on a realistic reconfigurable computing platform. The experimental results show that the latest generation of FPGAs, programmed using Handel-C, can achieve a sustained floating point performance up to 6 times greater than the microprocessor while operating at a clock frequency that is 60 times lower. The transformations are applied in a way that could be generalized, allowing efficient compilation approaches for the mapping of high level descriptions onto FPGAs. Kieron Turkington, Kostas Masselos, George A. Constantinides, Philip H. W. Leong |
FPL | 4 |
| 2006 | An FPGA-Based Electronic Cochlea with Dual Fixed-Point ArithmeticabstractAn improved FPGA implementation of an electronic cochlea filter is presented. We show that by using decimation, the computations of the electronic cochlea can be reduced. Furthermore, employing dual fixed-point arithmetic, gives a significant improvement in signal to noise ratio. A sequential architecture is described which employs pipelined infinite impulse response filter stages. The accuracy, performance and resource utilisation of a number of different implementations are compared. Chak-Kuen Wong, Philip H. W. Leong |
FPL | 2 |
| 2006 | FPGA-based MSB-first bit-serial variable block size motion estimation processorabstractH.264/AVC is the latest video coding standard adopting variable block size, quarter-pixel accuracy, motion vector prediction and multi-reference frames for motion estimation. These new features result in much higher computation requirements than previous coding standards. In this paper we propose a novel most significant bit (MSB) first bit-serial architecture for full-search block matching (FSBM) variable block size motion estimation. Since the nature of MSB-first processing enables early termination of the sum of absolute difference (SAD) calculation, the average hardware performance can be enhanced. The architecture has been simulated, synthesized and implemented on a Xilinx Virtex-II XC2V6000 FPGA. The maximum frequency achieved is 340 MHz and the throughput rate is around 18674 macroblocks per second within a -16 to 15 search range. The resource utilization is 3345 LUTs and it can encode CIF resolution video in real time Brian M. H. Li, Philip H. W. Leong |
FPT | 2 |
| 2006 | A Scalable FPGA Implementation of Cellular Neural Networks for Gabor-type FilteringabstractWe describe an implementation of Gabor-type filters on field programmable gate arrays using the cellular neural network (CNN) architecture. The CNN template depends upon the parameters (e.g., orientation, bandwidth) of the Gabor-type filter and can be modified at runtime so that the functionality of Gabor-type filter can be changed dynamically. Our implementation uses the Euler method to solve the ordinary differential equation describing the CNN. The design is scalable to allow for different pixel array sizes, as well as simultaneous computation of multiple filter outputs tuned to different orientations and bandwidths. For 1024 pixel frames, an implementation on a Xilinx Virtex XC2V1000-4 device uses 1842 slices, operates at 120 MHz and achieves 23,000 Euler iterations over one frame per second. Ocean Y. H. Cheung, Philip H. W. Leong, Eric K. C. Tsang, Bertram E. Shi |
IJCNN | 2 |
| 2006 | Development of a Human Airbag System for Fall Protection Using MEMS Motion Sensing TechnologyabstractThis paper describes the development of a human airbag system which is designed to reduce the impact force from falls. A micro inertial measurement unit (muIMU), based on MEMS accelerometers and gyro sensors is developed as the motion sensing part of the system. A recognition algorithm is used for real-time fall determination. With the algorithm, a microcontroller integrated with the muIMU can discriminate falling-down motion from normal human motions and trigger an airbag system when a fall occurs. Our airbag system is designed to have fast response with moderate input pressure, i.e., the experimental response time is less than 0.3 second under 0.4 MPa. In addition, we present our progress on using support vector machine (SVM) training together with the muIMU to better distinguish falling and normal motions. Experimental results show that selected eigenvector sets generated from 200 experimental data sets can be accurately separated into falling and other motions Guangyi Shi, Cheung-Shing Chan, Yilun Luo, Guanglie Zhang, Wen Jung Li, Philip H. W. Leong, Kwok-Sui Leung |
IROS | 6 |
| 2006 | A Hardware Gaussian Noise Generator Using the Box-Muller Method and Its Error AnalysisabstractWe present a hardware Gaussian noise generator based on the Box-Muller method that provides highly accurate noise samples. The noise generator can be used as a key component in a hardware-based simulation system, such as for exploring channel code behavior at very low bit error rates, as low as 10-12to 10-13. The main novelties of this work are accurate analytical error analysis and bit-width optimization for the elementary functions involved in the Box-Muller method. Two 16-bit noise samples are generated every clock cycle and, due to the accurate error analysis, every sample is analytically guaranteed to be accurate to one unit in the last place. An implementation on a Xilinx Virtex-4 XC4VLX100-12 FPGA occupies 1,452 slices, three block RAMs, and 12 DSP slices, and is capable of generating 750 million samples per second at a clock speed of 375 MHz. The performance can be improved by exploiting concurrent execution: 37 parallel instances of the noise generator at 95 MHz on a Xilinx Virtex-II Pro XC2VP100-7 FPGA generate seven billion samples per second and can run over 200 times faster than the output produced by software running on an Intel Pentium-4 3 GHz PC. The noise generator is currently being used at the Jet Propulsion Laboratory, NASA to evaluate the performance of low-density parity-check codes for deep-space communications Dong-U Lee, John D. Villasenor, Wayne Luk, Philip H. W. Leong |
IEEE Trans. Computers | 4 |
| 2005 | Mullet - A Parallel Multiplier GeneratorabstractA module generator called Mullet for producing near-optimal parallel multipliers in a technology independent manner is presented. Using this tool, a large number of candidate designs can be generated in order to find combinations of primitive elements which produce the best multiplier. The process of multiplication is broken down into a partial product generator (PPG) and a partial product summer (PPS). Both of these tasks can be done in a number of different ways and the best solution depends on the size of the required multiplier as well as the technology used. Mullet can be combined with a searching algorithm to find the best multiplier based on some objective. It can generate high quality multipliers for irregular architectures such as FPGAs and use features such as the dedicated multipliers. The tool can also be used to explore tradeoffs between architectures and to calibrate timing models of the primitive components. Synthesized examples using Xilinx FPGA devices are comparisons are made with those produced by the Xilinx CoreGenerator and XST tools. Kuen Hung Tsoi, Philip H. W. Leong |
FPL | 2 |
| 2005 | Ziggurat-based Hardware Gaussian Random Number GeneratorabstractAn architecture and implementation of a high performance Gaussian random number generator (GRNG) is described. The GRNG uses the Ziggurat algorithm which divides the area under the probability density function into three regions (rectangular, wedge and tail). The rejection method is then used and this amounts to determining whether a random point falls into one of the three regions. The vast majority of points lie in the rectangular region and are accepted to directly produce a random variate. For the nonrectangular regions, which occur 1.5% of the time, the exponential or logarithm functions must be computed and an iterative fixed point operation unit is used. Computation of the rectangular region is heavily pipelined and a buffering scheme is used to allow the processing of rectangular regions to continue to operate in parallel with evaluation of the wedge and tail computation. The resulting system can generate 169 million normally distributed random numbers per second on a Xilinx XC2VP3O-6 device. Guanglie Zhang, Philip H. W. Leong, Dong-U Lee, John D. Villasenor, Ray C. C. Cheung, Wayne Luk |
FPL | 2 |
| 2005 | Implementation of Gabor-Type Filters on Field Programmable Gate Arrays
Ocean Y. H. Cheung, Philip H. W. Leong, Eric K. C. Tsang, Bertram E. Shi |
FPT | 2 |
| 2005 | Dynamic Voltage Scaling for Commercial FPGAs
Gary Chun Tak Chow, L. S. M. Tsui, Philip H. W. Leong, Wayne Luk, Steve Wilton |
FPT | 3 |
| 2005 | Reconfigurable Acceleration for Monte Carlo Based Financial Simulation
Guanglie Zhang, Philip H. W. Leong, Chun Hok Ho, Kuen Hung Tsoi, Chris C. C. Cheung, Dong-U Lee, Ray C. C. Cheung, Wayne Luk |
FPT | 2 |
| 2005 | A hardware Gaussian noise generator using the Wallace methodabstractWe describe a hardware Gaussian noise generator based on the Wallace method used for a hardware simulation system. Our noise generator accurately models a true Gaussian probability density function even at high /spl sigma/ values. We evaluate its properties using: 1) several different statistical tests, including the chi-square test and the Anderson-Darling test and 2) an application for decoding of low-density parity-check (LDPC) codes. Our design is implemented on a Xilinx Virtex-II XC2V4000-6 field-programmable gate array (FPGA) at 155 MHz; it takes up 3% of the device and produces 155 million samples per second, which is three times faster than a 2.6-GHz Pentium-IV PC. Another implementation on a Xilinx Spartan-III XC3S200E-5 FPGA at 106 MHz is two times faster than the software version. Further improvement in performance can be obtained by concurrent execution: 20 parallel instances of the noise generator on an XC2V4000-6 FPGA at 115 MHz can run 51 times faster than software on a 2.6-GHz Pentium-IV PC. Dong-U Lee, Wayne Luk, John D. Villasenor, Guanglie Zhang, Philip H. W. Leong |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2004 | An Arithmetic Library and Its Application to the N-body ProblemabstractComputer arithmetic is a specialist field of study, and it is very difficult for designers to choose the most efficient method for implementing a given algorithm due to the large number of design choices available. In this paper, an object oriented arithmetic library is presented which can be used to simulate and generate designs which use fixed, floating, logarithmic and hybrid number representations. The advantages of this approach are that a user can explore tradeoffs concerning precision, accuracy and speed from single high level description. Furthermore, users need not be intimately familiar with the implementation details of the underlying libraries, thus allowing users to develop systems employing advanced computer arithmetic without detailed knowledge of its implementation. The application of this library to a coprocessor which implements the force pipeline for an N-body solver is described. Kuen Hung Tsoi, Chun Hok Ho, Jackson H. C. Yeung, Philip H. W. Leong |
FCCM | 4 |
| 2004 | IP Generation for an FPGA-Based Audio DAC Sigma-Delta Converter
Ralf Ludewig, Oliver Soffke, Peter Zipf, Manfred Glesner, Kong-Pang Pun, Kuen Hung Tsoi, Kin-Hong Lee, Philip H. W. Leong |
FPL | 8 |
| 2004 | An FPGA-based Othello endgame solverabstractA single chip FPGA-based Othello endgame solver is presented in This work. The solver includes all the hardware for move checking, disc flipping, move selection, board evaluation and alpha-beta pruning. On a Xilinx Virtex XCVW00E-6 device operating at 50 MHz, the chip can search 3.14 million Othello positions per second. The endgame chip achieves a speedup of 3.5 over an 800 MHz Pentium III machine, showing that performance similar to that of a high end microprocessor can be achieved using modest FPGA resources. By using a larger FPGA, a more sophisticated search algorithm and an improved datapath, we believe that a single FPGA based endgame solver with at least two orders of magnitude better performance can be developed. Chak-Kuen Wong, K. K. Lo, Philip H. W. Leong |
FPT | 3 |
| 2003 | FPGA-based SIMD ProcessorabstractA massively parallel single instruction multiple data stream (SIMD) processor designed specifically for cryptographic key search applications is presented. This design exploits fine grain parallelism and the high memory bandwidth available in an FPGA (field programmable gate array) by integrating 95 simple processors and memory on a single FPGA chip. Performance is compared with a previously reported hardwired design on a RC4 key search application. Stanley Y. C. Li, Gap C. K. Cheuk, Kin-Hong Lee, Philip H. W. Leong |
FCCM | 4 |
| 2003 | Compact FPGA-based True and Pseudo Random Number GeneratorsabstractTwo FPGA-based (field programmable gate array) implementations of random number generators intended for embedded cryptographic applications are presented. The first is a true random number generator (TRNG) which employs oscillator phase noise, and the second is a bit serial implementation of a Blum Blum Shub (BBS) pseudorandom number generator (PRNG). Both designs are extremely compact and can be implemented on any FPGA of PLD device. They were designed specifically for use as FPGA-based cryptographic hardware cores. The TRNG and PRNG were tested using the NIST and Diehard random number test suites. Kuen Hung Tsoi, Ka Hei Leung, Philip H. W. Leong |
FCCM | 3 |
| 2003 | A Smith-Waterman Systolic Cell
Chi Wai Yu, K. H. Kwong, Kin-Hong Lee, Philip H. W. Leong |
FPL | 4 |
| 2003 | An FPGA-based re-configurable 24-bit 96kHz sigma-delta audio DACabstractThis paper presents a reconfigurable sigma-delta audio Digital-to-Analog Converter (DAC) which is suitable for embedded FPGA applications. The Sigma-Delta Modulator (SDM) design can be configured as a 3rd or 5th order SDM and allows different input word lengths. Different input sampling rates are also entertained by employing a programmable interpolator. The DAC accepts 16-/18-/20-/24-bit PCM data at sampling rates of 32/44.1/48/88.2/96 kHz for applications in CD, SACD and DVD audio. Ray C. C. Cheung, Kong-Pang Pun, Steve C. L. Yuen, Kuen Hung Tsoi, Philip H. W. Leong |
FPT | 5 |
| 2003 | Arbitrary function approximation in HDLs with application to the N-body problemabstractA module generator is described that allows for the generation of synthesizable VHDL modules which implement arbitrary functions in fixed point precision using the Symmetric Table Addition Method (STAM). This module generator was interfaced to a high level synthesis tool "fly" which automatically generates fully-pipelined circuits from a Perl-like language. The resulting system was applied to the N-body problem and results are presented. It was found that a function generator module is a very useful addition to a hardware description language. Chun Hok Ho, Kuen Hung Tsoi, Jackson H. C. Yeung, Yuet Ming Lam, Kin-Hong Lee, Philip H. W. Leong, Ralf Ludewig, Peter Zipf, Alberto García Ortiz, Manfred Glesner |
FPT | 6 |
| 2003 | Modular exponentiation using parallel multipliersabstractA field programmable gate array (FPGA) semi-systolic implementation of a modular exponentiation unit, suitable for use in implementing the RSA public key cryptosystem is presented. The design is carefully matched with features of the FPGA architecture, utilizing embedded 18/spl times/18-bit multipliers on the FPGA and employing a carry save addition scheme. Using this architecture, a 1024-bit modular exponentiation can operate at 90 MHz on a Xilinx XC2V3000-6 device and perform a 1024-bit RSA decryption in 0.66 ms with the Chinese Remainder Theorem. S. H. Tang, K. S. Tsui, Philip H. W. Leong |
FPT | 3 |
| 2003 | A variable-radix digit-serial design methodology and its application to the discrete cosine transformabstractA variable-radix digit-serial design methodology and its application to the implementation of a systolic structure for computing the discrete cosine transform is presented. Based on the parameters supplied by a user, different fixed-point designs can be derived from a single floating-point description where tradeoffs among quantization effects, throughput, latency, and area can be addressed. The resulting hardware implementations have variables of different wordlengths and operators of different radices. This design methodology enables efficient exploration of a complex design space to determine the most suitable implementation for a particular application. Monk-Ping Leong, Philip H. W. Leong |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | A Massively Parallel RC4 Key Search EngineabstractA massively parallel implementation of an RC4 key search engine on an FPGA is described. The design employs parallelism at the logic level to perform many operations per cycle, uses on-chip memories to achieve very high memory bandwidth, floorplanning to reduce routing delays and multiple decryption units to achieve further parallelism. A total of 96 RC4 decryption engines were integrated on a single Xilinx Virtex XCV1000-E field programmable gate array (FPGA). The resulting design operates at a 50 MHz clock rate and achieves a search speed of 6.06 /spl times/ 10/sup 6/ keys/second, which is a speedup of 58 over a 1.5 GHz Pentium 4 PC. Kuen Hung Tsoi, Kin-Hong Lee, Philip H. W. Leong |
FCCM | 3 |
| 2002 | Fly - A Modifiable Hardware Compiler
Chun Hok Ho, Philip H. W. Leong, Kuen Hung Tsoi, Ralf Ludewig, Peter Zipf, Alberto García Ortiz, Manfred Glesner |
FPL | 2 |
| 2002 | An FPGA Based SHA-256 Processor
Kurt K. Ting, Steve C. L. Yuen, Kin-Hong Lee, Philip H. W. Leong |
FPL | 4 |
| 2002 | Implementation of an FPGA based accelerator for virtual private networksabstractVirtual Private Networks (VPN) are becoming increasingly popular network architectures for corporate networks. As VPNs are built on the Internet infrastructure, the data exchange among different local area networks will be passed through the Internet and thus can be easily eavesdropped, masqueraded, etc. Therefore, certain security measures must be used to deal with these privacy issues. The Internet Protocol Security (IPSec) by the Internet Engineering Task Force (IETF) addresses the abovementioned security issues and the Free Secure Wide Area Network (FreeS/WAN) is an open source software implementation of IPSec for Linux which uses triple-DES as the default encryption mode. As shown in this paper, the performance of FreeS/WAN with IPSec is 50% of that without encryption. In order to improve its performance, a field programmable gate array (FPGA) based triple-DES accelerator was built on a reconfigurable computing development platform called Pilchard and achieved a throughput of more than 120 Mb/sec for triple-DES in cipher-block chaining mode, a speedup of 3 over a software implementation, Measurements show that an FPGA-accelerated FreeS/WAN offers a 30% speedup for the TCP protocol over the original software library. Ocean Y. H. Cheung, Philip H. W. Leong |
FPT | 2 |
| 2002 | A system level implementation of Rijndael on a memory-slot based FPGA cardabstractThis paper describes system level issues encountered in a high performance implementation of a Rijndael encryption core on a memory-slot based reconfigurable computing platform called Pilchard. The Rijndael algorithm was adopted in 2000 by the US National Institute of Standards and Technology (NIST) as the Advanced Encryption Standard (AES). In the implementation of Rijndael, changing the number of unrolled rounds in the encryption core can affect the performance of the system. It is shown that for the design presented, the highest performance of 755 Mbit/sec was achieved by implementing a core with a single round. Although it is relatively easy to implement a high performance core on an FPGA, due to I/O bottlenecks, achieving high system level performance is more difficult. In order to optimize the performance of the host/FPGA interface, special instructions from the Intel Pentium III streaming SIMD extensions (SSE) along with write-combining memory operations were used. These features enabled the measured throughput of the AES core to reach 445 Mbit/sec which, although still slower than the AES core, was double that of an unoptimized interface. Dennis K. Y. Tong, Pui Sze Lo, Kin-Hong Lee, Philip H. W. Leong |
FPT | 4 |
| 2002 | A microcoded elliptic curve processor using FPGA technologyabstractThe implementation of a microcoded elliptic curve processor using field-programmable gate array technology is described. This processor implements optimal normal basis field operations in F(2/sup n/). The design is synthesized by a parameterized module generator, which can accommodate arbitrary n and also produce field multipliers with different speed/area tradeoffs. The control part of the processor is microcoded, enabling curve operations to be incorporated into the processor and hence reducing the chip's I/O requirements. The microcoded approach also facilitates rapid development and algorithmic optimization: for example, projective and affine coordinates were supported using different microcode. The design was successfully tested on a Xilinx Virtex XCV1000-6 device and could perform an elliptic curve multiplication over the field F(2/sup n/) using affine and projective coordinates for n=113,155, and 173. Philip H. W. Leong, Ivan K. H. Leung |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | Tradeoffs in Parallel and Serial Implementations of the International Data Encryption Algorithm IDEA
Ocean Y. H. Cheung, Kuen Hung Tsoi, Philip H. W. Leong, Monk-Ping Leong |
CHES | 3 |
| 2001 | Parameterized Module Generator for an FPGA-Based Electronic Cochlea
Monk-Ping Leong, Craig T. Jin, Philip H. W. Leong |
FCCM | 3 |
| 2001 | Pilchard - A Reconfigurable Computing Platform with Memory Slot Interface
Philip H. W. Leong, Monk-Ping Leong, Ocean Y. H. Cheung, T. Tung, C. M. Kwok, Ming Yiu Wong, Kin-Hong Lee |
FCCM | 1 |
| 2001 | A bitstream reconfigurable FPGA implementation of the WSAT algorithmabstractA field programmable gate array (FPGA) implementation of a coprocessor which uses the WSAT algorithm to solve Boolean satisfiability problems is presented. The input is a SAT problem description file from which a software program directly generates a problem-specific circuit design which can be downloaded to a Xilinx Virtex FPGA device and executed to find a solution. On an XCV300, problems of 50 variables and 170 clauses can be solved. Compared with previous approaches, it avoids the need for resynthesis, placement, and routing for different constraints. Our coprocessor is eminently suitable for embedded applications where energy, weight and real-time response are of concern. Philip H. W. Leong, Chiu-Wing Sham, H. Y. Wong, Wing Seung Yuen, Monk-Ping Leong |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2000 | A Bit-Serial Implementation of the International Data Encryption Algorithm IDEAabstractA high-performance implementation of the International Data Encryption Algorithm (IDEA) is presented in this paper. Using a novel bit-serial architecture to perform multiplication modulo 2/sup 16/+1, the implementation occupies a minimal amount of hardware. The bit-serial architecture enabled the algorithm to be deeply pipelined to achieve a system clock rate of 125 MHz on a Xilinx Virtex XCV300-6, delivering a throughput of 500 Mb/sec. With a XCV1000-6 device, the estimated performance is 2 Gb/sec, three orders of magnitude faster than a software implementation on a 450 MHz Intel Pentium II. This design is suitable for applications in on-line encryption for high-speed networks. Monk-Ping Leong, Ocean Y. H. Cheung, Kuen Hung Tsoi, Philip H. W. Leong |
FCCM | 4 |
| 2000 | FPGA Implementation of a Microcoded Elliptic Curve Cryptographic ProcessorabstractElliptic curve cryptography (ECC) has been the focus of much recent attention since it offers the highest security per bit of any known public key cryptosystem. This benefit of smaller key sizes makes ECC particularly attractive for embedded applications since its implementation requires less memory and processing power. In this paper a microcoded Xilinx Virtex based elliptic curve processor is described. In contrast to previous implementations, it implements curve operations as well as optimal normal basis field operations in F(2/sup n/); the design is parameterized for arbitrary n; and it is microcoded to allow for rapid development of the control part of the processor. The design was successfully tested on a Xilinx Virtex XCV300-4 and, for n=113 bits, utilized 1290 slices at a maximum frequency of 45 MHz and achieved a thirty-fold speedup over an optimized software implementation. Ka Hei Leung, K. W. Ma, Wai Keung Wong, Philip H. W. Leong |
FCCM | 4 |
| 1999 | Automatic Floating to Fixed Point Translation and its Application to Post-Rendering 3D WarpingabstractThe automatic conversion of floating point software implementations of algorithms to a equivalent fixed point implementation which can be efficiently implemented in an FCCM remains an obstacle in the rapid systems prototyping design flow. Floating point to fixed point conversion is tedious, error prone and requires a good knowledge of fixed point computer arithmetic. This paper describes a software system called fp designed to automate the process. It consists of a fixed point C++ class; a profiler which is used to determine the number of bits of precision required for each signal in the hardware implementation; an optimiser which finds the minimal number of bits required for a specified degree of accuracy in the implementation and finally and a compiler which takes the information collected by the system and outputs synthesisable VHDL code. A post-rendering 3D image warping application designed using this system is used as an example. Monk-Ping Leong, M. Y. Yeung, C. K. Yeung, Chi-Wing Fu, Pheng-Ann Heng, Philip H. W. Leong |
FCCM | 6 |
| 1998 | An FPGA Implementation of GENET for Solving Graph Coloring Problems abstractConstraint satisfaction problems (CSPs) can be used to model problems in a wide variety of application areas, such as time-table scheduling, bandwidth allocation, and car-sequencing. To solve a CSP means finding appropriate values for its set of variables such that all of the specified constraints are satisfied. Almost all CSPs have exponential time complexity and instances of them may require a prohibitively large amount of time to solve. Consequently, much research has been done in developing efficient methods to solve CSPs. In particular, a generic neural network (GENET) model, developed by C.J. Wang and E.P.K. Tsang (1991), has been demonstrated to work extremely well in solving many CSPs, often finding solutions where other methods fail. Philip H. W. Leong, K. T. Chan, Siew Kok Hui, H. K. Yeung, M. F. Lo, Jimmy Ho-Man Lee |
FCCM | 2 |
| 1998 | A FPGA Based Forth MicroprocessorabstractSystems which employ a microprocessor together with an application specific FPGA based coprocessor are common today. These applications can reduce power consumption and system costs by incorporating the microprocessor in the FPGA. For such applications, a microprocessor which has good performance, occupies a minimal amount of FPGA resources, has a good high level language software development environment and good code density is desirable. In this paper a 16 bit FPGA based microprocessor, called MSL16, optimised for such applications is described. MSL16 utilises a stack architecture with each instruction occupying only 4 bits, leading to a small instruction set, simple datapath and control, and high code density. MSL16 was specifically designed to efficiently execute the programming language "Forth". The Forth language has the desirable features of portability and high code density, and it is well suited to control, DSP, real-time and embedded applications. Philip H. W. Leong, P. K. Tsang |
FCCM | 1 |
| 1995 | A low-power VLSI arrhythmia classifierabstractThe design, implementation, and operation of a low-power multilayer perceptron chip (Kakadu) in the framework of a cardiac arrhythmia classification system is presented in this paper. This classifier, called MATIC, makes timing decisions using a decision tree, and a neural network is used to identify heartbeats with abnormal morphologies. This classifier was designed to be suitable for use in implantable devices and a VLSI (very large scale integration) neural-network chip (Kakadu) was designed so that the computationally expensive neural-network algorithm can be implemented with low power consumption. Kakadu implements a (10,6,4) perceptron and has a typical power consumption of tens of microwatts. When used with the arrhythmia classification system, the chip can operate with an average power consumption of less than 25 nW. Philip H. W. Leong, Marwan A. Jabri |
IEEE Trans. Neural Networks | 1 |
| 1993 | Kakadu - A Low Power Analogue Neural Network ClassifierabstractAn analogue neural network VLSI chip designed for low power operation is presented. This chip consists of 84 synapse elements arranged as arrays of size 10 x 6 and 6 x 4 and was fabricated using a standard 1.2 micron double metal single poly CMOS process. The synapses are digitally programmable and static weight storage is provided. The chip has a typical power consumption of tens of microwatts. It has been successfully trained and tested on a range of classification problems including 4-bit parity, character recognition and morphological-based classification of intracardiac electrogram signals. Philip H. W. Leong, Marwan A. Jabri |
Int. J. Neural Syst. | 1 |
| 1991 | ANN Board Classification for Heart Defibrillators
Marwan A. Jabri, Stephen Pickard, Philip H. W. Leong, Z. Chi, Barry Flower |
NIPS | 3 |