EDBT 2026 Demo / reviewers in the wild / expert
David Boland
dblp:23/5076
· DBLP profile ↗
44ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0001-5370-4464ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 8 first-author · 17 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CroSatFL: Energy-Efficient Federated Learning with Cross-Aggregation for Satellite Edge Computing
Bahman Javadi, Rodrigo N. Calheiros, David Boland, Philip Leong |
CCGrid | 4 |
| 2026 | TurboFuzz: FPGA Accelerated Hardware Fuzzing for Processor Agile VerificationabstractVerification is a critical process for ensuring the correctness of modern processors. The increasing complexity of processor designs and the emergence of new instruction set architectures (ISAs) like RISC-V have created demands for more agile and efficient verification methodologies, particularly regarding verification efficiency and faster coverage convergence. While simulation-based approaches now attempt to incorporate advanced software testing techniques such as fuzzing to improve coverage, they face significant limitations when applied to processor verification, notably poor performance and inadequate test case quality. Hardware-accelerated solutions using FPGA or ASIC platforms have tried to address these issues, yet they struggle with challenges including host-FPGA communication overhead, inefficient test pattern generation, and suboptimal implementation of the entire multi-step verification process. In this paper, we present TurboFuzz, an end-to-end hardwareaccelerated verification framework that implements the entire Test Generation-Simulation-Coverage Feedback loop on a single FPGA for modern processor verification. TurboFuzz enhances test quality through optimized test case (seed) control flow, efficient inter-seed scheduling, and hybrid fuzzer integration, thereby improving coverage and execution efficiency. Additionally, it employs a feedback-driven generation mechanism to accelerate coverage convergence. Experimental results show that TurboFuzz achieves up to 2.23× more coverage collection than software-based fuzzers within the same time budget, and up to 571× performance speedup when detecting real-world issues, while maintaining full visibility and debugging capabilities with moderate area overhead. Xueqi Li 0001, Sa Wang, David Boland, Yungang Bao, Kan Shi |
HPCA | 5 |
| 2025 | Corvus: Efficient HW/SW Co-Verification Framework for RISC-V Instruction Extensions with FPGA AccelerationabstractThe RISC-V instruction set architecture (ISA) offers flexibility for domain-specific custom instruction extensions. While the basic RISC-V ISA contains common instructions, the extended accelerators provide additional computing power to meet diverse needs. High-level synthesis (HLS) is often used to agilely create custom extension accelerators, allowing engineers to design complex digital circuits using high-level languages such as C/C++, further improving development efficiency. However, verifying a design that includes RISC-V cores and custom extensions is rarely studied and can be challenging. Traditional approaches for verifying HLS-generated designs use C-RTL co-simulation, primarily focusing on the unit level. This method can be extremely time-consuming and often makes impractical assumptions about interactions between HLS-generated circuits and the processor. Therefore, system-level verification is essential to extensively exercise the RISC-V cores, the custom extensions, and their interconnections. Zijian Jiang, Keran Zheng, David Boland, Yungang Bao, Kan Shi |
ASP-DAC | 3 |
| 2025 | Hercules: Efficient Verification of High-Level Synthesis Designs with FPGA AccelerationabstractHigh-Level Synthesis (HLS) enables software engineers to create intricate digital circuit designs using high-level languages like C/C++. While HLS tools can perform functional verification using C/C++ simulation, it is harder to verify that the generated RTL is also correct. This problem is exacerbated for designs which include HLS-generated IPs, such as PCIe or DDR interfaces, or hand-written RTL. While it is possible to perform cycle-accurate verification using C/RTL co-simulation, conventional methods are both slow and typically only focus on unit-level verification which can make it harder to identify the root cause of a bug. Shuoxiang Xu, Zijian Jiang, David Boland, Yungang Bao, Kan Shi |
FPGA | 4 |
| 2025 | Highly Parallel CNN Accelerator for RepVGG-Like Network Training on FPGAsabstractIn this article, we propose a generic FPGA-based training accelerator tailored for RepVGG-like networks, which strikes a balance between maximizing training-time accuracy and minimizing inference-time latency. The proposed accelerator leverages fine-grain channel-level parallelism within computational units specially designed for multiple branches of the basic building block within the RepVGG-like network. Specifically, we employ a Conv block for forward Conv and backward deConv, along with a dilated Conv block, including a weight kernel partition scheme for efficient weight gradient calculation. Furthermore, we aggressively exploit a 2-stage coarse-grain task-level parallelism for low-latency CNN training: 1) parallelism among multiple branches of the basic building block of RepVGG and 2) parallelism between error back-propagation and weight gradient calculation in the backward path. Through experiments on the CIFAR-10 dataset using 16-bit fixed-point arithmetic, we demonstrate state-of-the-art batch 1 throughput of 150 GOPs for training and 183 GOPs for inference. Chuliang Guo, Binglei Lou, David Boland, Philip H. W. Leong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | FPGA-based Block Minifloat Training Accelerator for a Time Series Prediction NetworkabstractTime series forecasting is the problem of predicting future data samples from historical information and recent deep neural network (DNNs) based techniques have achieved excellent results compared with conventional statistical approaches. Many applications at the edge can utilize this technology and most implementations have focused on inference, an ability to train at the edge would enable the DNN to adapt to changing conditions. Unfortunately, training requires approximately three times more memory and computation than inference. Moreover, edge applications are often constrained by energy efficiency. In this work, we implement a block minifloat (BM) training accelerator for a time series prediction network, N-BEATS. Our architecture involves a mixed-precision GEMM accelerator that utilizes BM arithmetic. We use a 4-bit DSP packing scheme to optimize the implementation further, achieving a throughput of 779 Gops. The resulting power efficiency is 42.4 Gops/W, 3.1 \(\times\) better than a graphics processing unit in a similar technology. Haoyan Qi, David Boland, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2024 | Hassert: Hardware Assertion-Based Verification Framework with FPGA AccelerationabstractHardware verification is typically the bottleneck of the chip development cycle, mainly due to the time-consuming simulation and debugging process using software simulators. Assertion-Based Verification (ABV) has been widely adopted to provide better visibility into microarchitecture details and automatically detect unexpected behaviors. While ABV significantly improves verification efficiency, checking assertions using software simulators requires extremely long times for large benchmarks. Prototyping designs on an FPGA is a potential alternative to verify hardware, but it lacks fine-grained debugging capabilities for when errors occur. Weijie Weng, Lijia Cai, David Boland, Yungang Bao, Kan Shi |
ASPLOS (4) | 6 |
| 2024 | PolyLUT-Add: FPGA-based LUT Inference with Wide InputsabstractFPGAs have distinct advantages as a technology for deploying deep neural networks (DNNs) at the edge. Lookup Table (LUT) based networks, where neurons are directly modeled using LUTs, help maximize this promise of offering ultra-low latency and high area efficiency on FPGAs. Unfortunately, LUT resource usage scales exponentially with the number of inputs to the LUT, restricting PolyLUT to small LUT sizes. This work introduces PolyLUT-Add, a technique that enhances neuron connectivity by combining A PolyLUT sub-neurons via addition to improve accuracy. Moreover, we describe a novel architecture to improve its scalability. We evaluated our implementation over the MNIST, Jet Substructure classification, and Network Intrusion Detection benchmark and found that for similar accuracy, PolyLUT-Add achieves a LUT reduction of $2.0-13.9 \times$ with a $1.2-1.6 \times$ decrease in latency. Binglei Lou, Richard Rademacher, David Boland, Philip H. W. Leong |
FPL | 3 |
| 2024 | S$^{3}$CA: A Sparse Strip Spectral Correlation AnalyzerabstractThe spectral correlation density (SCD) is widely used to characterize cyclostationary signals and the strip spectral correlation analyzer (SSCA) is commonly used to estimate the SCD. Although the SSCA utilizes the fast Fourier transform (FFT) for computational efficiency, its real-time implementation still poses challenges as large input sizes are often involved. In this work, we present a sparse strip spectral correlation analyzer (S3CA) based on the sparse fast Fourier transform (SFFT). The S3CA approach involves computing a sparse, downsampled channel-data product (CDP) which is then passed to a modified SFFT implementation to obtain the spectral density. For an input of length 2 million samples, the S3CA is 30× faster than the conventional SSCA. Carol Jingyi Li, Richard Rademacher, David Boland, Craig T. Jin, Chad M. Spooner, Philip H. W. Leong |
IEEE Signal Process. Lett. | 3 |
| 2024 | FedOrbit: Energy Efficient Federated Learning for Orbital Edge Computing Using Block Minifloat ArithmeticabstractLow Earth Orbit (LEO) satellite constellations have diverse applications, including earth observation, communication services, navigation, and positioning. These constellations have evolved into a valuable data source; however, their use in a ground station (GS) for analysis via machine learning algorithms presents challenges due to constraints on power consumption, communication bandwidth, and onboard computing capabilities. While the combination of Federated Learning (FL) and Orbital Edge Computing has been employed to address these challenges, its heavy reliance on the GS for model aggregation and edge resource limitations remains a research challenge. This article presents FedOrbit, a novel energy-efficient and decentralised FL method to optimise communication with the GS and reduce power consumption. FedOrbit utilises reinforcement learning for cluster formation, satellite visiting patterns for master satellite selection, and block minifloat arithmetic for power reduction. Extensive performance evaluation under Walker Delta-based LEO constellation configurations and different datasets reveals that FedOrbit can maintain high accuracy while significantly reduce communication demand, power consumption and training time in comparison to state-of-the-art FL approaches. The proposed technique can also reduce the training time by 5× compared with the centralised FL approaches. In addition, the utilisation of block minifloat representation as low-precision arithmetic enhanced the energy consumption by 3.5× compared with the single-precision (FP32) format. Mohammad Reza Jabbarpour, Bahman Javadi, Philip H. W. Leong, Rodrigo N. Calheiros, David Boland |
IEEE Trans. Serv. Comput. | 5 |
| 2023 | Single-Batch CNN Training using Block Minifloats on FPGAsabstractTraining convolutional neural networks remains a challenge on resource-limited edge devices due to its intensive computations, large storage requirements, and high bandwidth. Error back-propagation, gradient generation, and weight update usually require high precision to guarantee model accuracy, which places a further burden on computation and bandwidth. This paper presents the first parallel FPGA CNN training accelerator with block minifloat datatypes. We first propose a heuristic bit-width allocation technique to derive a unified 8-bit block minifloat format with a sign bit, 2 exponent bits, and 5 mantissa bits. In contrast to previous techniques, the same data format is used for weights, activations, errors, and gradients. Using this format, accuracy similar to 32-bit single precision floating point is achieved and thus simplifies the FPGA-based designs of computational units such as multiply-and-add. In addition, we propose a unified Conv block to deal with Conv and transposed Conv in the forward and backward paths respectively; and a dilated Conv block with a weight kernel partition scheme for gradient generation. Both Conv blocks support non-unit stride, this being crucial for the residual connections that appear in modern CNNs. For training of ResNet20 on the CIFAR-10 dataset with a batch size of 1, our accelerator on a Xilinx Ultrascale+ ZCU102 FPGA achieves state-of-the-art single-batch throughput of 144.64 and 192.68 GOPs with and without batch normalisation layers respectively. Chuliang Guo, Binglei Lou, Xueyuan Liu 0002, David Boland, Philip H. W. Leong |
FPGA | 4 |
| 2023 | ENCORE: Efficient Architecture Verification Framework with FPGA AccelerationabstractVerification typically consumes the majority of the time in the hardware development cycle. Primarily this is because multiple iterations to debug hardware using software simulation is extremely time-consuming. While FPGAs can be utilised to accelerate the simulation, existing methods either provide limited visibility of design details, or are expensive to check against a reference model dynamically at the system level. Kan Shi, Shuoxiang Xu, Yuhan Diao, David Boland, Yungang Bao |
FPGA | 4 |
| 2023 | BOOST: Block Minifloat-Based On-Device CNN Training Accelerator with Transfer LearningabstractAdapting CNNs to changing problems is challenging on resource-limited edge devices due to intensive computations, high precision requirements, large storage needs, and high bandwidth. This paper presents BOOST, a novel block minifloat (BM)-based parallel CNN training accelerator on memory- and computation-constrained FPGAs for transfer learning (TL). By updating a small number of layers online, BOOST enables adaptation to changing problems. Our approach utilizes a unified 8-bit BM datatype (bm(2,5) ), i.e., with a sign bit, 2 exponent bits, and 5 mantissa bits, and proposes unified Conv and dilated Conv blocks that support non-unit stride and enable task-level parallelism during back-propagation to minimize latency. For ResNet20 and VGG-like training on CIFAR-10 and SVHN datasets, BOOST achieves near 32-bit floating point accuracy, reducing latency by 21%-43% and BRAM usage by 63%-66% compared to back-propagation training without TL. Notably, BOOST outperforms the prior SOTA works to achieve perbatch throughput of 131 and 209 GOPs for ResNet20 and VGG-like respectively. Chuliang Guo, Binglei Lou, Xueyuan Liu 0002, David Boland, Philip H. W. Leong, Cheng Zhuo |
ICCAD | 4 |
| 2023 | On-Board Federated Learning in Orbital Edge ComputingabstractLow Earth Orbit (LEO) satellite constellations are used for a wide range of applications including earth observation, communication services, navigation, and positioning. They have emerged as a new source of data but transferring this data to a ground station (GS) for analysis and machine learning requires extensive bandwidth and incurs high latency. Limited battery capacity, communication and computing capabilities are other factors affecting the training process. Federated Learning (FL) is being used to address these challenges, although it heavily relies on the GS for model aggregation. In this paper, we consider Orbital Edge Computing (OEC) as an architecture for LEO satellite constellations and propose an on-board Federated Learning to reduce communication with the GS. We present a novel decentralised FL algorithm, called FedOrbit, based on reinforcement learning cluster formation and satellite visiting patterns to utilise intra and inter-satellite communications for model aggregation. Extensive performance evaluation under Walker Delta-based LEO constellation configurations and different datasets including MNIST, CIFAR-10, and EuroSat revealed that FedOrbit can significantly reduce communication rounds, power consumption and training time in comparison to state-of-the-art FL approaches while maintaining a high accuracy. FedOrbit demonstrates a significant decrease in power consumption, specifically by 8.8% and 79.1% for the MNIST dataset, when compared to decentralised and centralised FL approaches, respectively. The proposed technique can also reduce the training time by 5× and 48× compared with the decentralised and centralised FL approaches, respectively. Mohammad Reza Jabbarpour, Bahman Javadi, Philip H. W. Leong, Rodrigo N. Calheiros, David Boland, Chris Butler |
ICPADS | 5 |
| 2023 | Fixed-point FPGA Implementation of the FFT Accumulation Method for Real-time Cyclostationary AnalysisabstractThe spectral correlation density (SCD) is an important tool in cyclostationary signal detection and classification. Even using efficient techniques based on the fast Fourier transform (FFT), real-time implementations are challenging because of the high computational complexity. A key dimension for computational optimization lies in minimizing the wordlength employed. In this article, we analyze the relationship between wordlength and signal-to-quantization noise in fixed-point implementations of the SCD function. A canonical SCD estimation algorithm, the FFT accumulation method (FAM) using fixed-point arithmetic, is studied. We derive closed-form expressions for SQNR and compare them at wordlengths ranging from 14 to 26 bits. The differences between the calculated SQNR and bit-exact simulations are less than 1 dB. Furthermore, an HLS-based FPGA design is implemented on a Xilinx Zynq UltraScale+ XCZU28DR-2FFVG1517E RFSoC. Using less than 25% of the logic fabric on the device, it consumes 7.7 W total on-chip power and has a power efficiency of 12.4 GOPS/W, which is an order of magnitude improvement over an Nvidia Tesla K40 graphics processing unit (GPU) implementation. In terms of throughput, it achieves 50 MS/sec, which is a speedup of 1.6 over a recent optimized FPGA implementation. Carol Jingyi Li, Xiangwei Li, Binglei Lou, Craig T. Jin, David Boland, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2023 | A Scalable Systolic Accelerator for Estimation of the Spectral Correlation Density Function and Its FPGA ImplementationabstractThe spectral correlation density (SCD) function is the time-averaged correlation of two spectral components used for analyzing periodic signals with time-varying spectral content. Although the analysis is extremely powerful, it has not been widely adopted in real-time applications due to its high computational complexity. In this article, we present an efficient FPGA implementation of the FFT accumulation method (FAM) for estimating the SCD function and its alpha profile. The implementation uses a linear systolic array with a bi-directional datapath consisting of DSP-based processing elements (PEs) with a dedicated instruction schedule, achieving a PE utilization of 88.2%. The 128-PE implementation achieves a clock frequency in excess of 530 MHz and consumes 151K LUTs, 151K FFs, 264 BRAMs, 4 URAMs, and 1,054 DSPs, which is less than 36% of the logic fabric on a Zynq UltraScale+ XCZU28DR-2FFVG1517E RFSoC device. It has a modest 12.5W power consumption and an energy efficiency of 4,832 MOPS/W, which is 20.6× better than the published state-of-the-art GPU implementation. In terms of throughput, it achieves 15,340 windows/s (15,340 windows/s × 2,048 samples/window = 31.4 MS/s), which is a 4.65× improvement compared to the above-mentioned GPU implementation and 807× compared to an existing hybrid FPGA-GPU implementation. Xiangwei Li, Douglas L. Maskell, Carol Jingyi Li, Philip H. W. Leong, David Boland |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2023 | fSEAD: A Composable FPGA-based Streaming Ensemble Anomaly Detection LibraryabstractMachine learning ensembles combine multiple base models to produce a more accurate output. They can be applied to a range of machine learning problems, including anomaly detection. In this article, we investigate how to maximize the composability and scalability of an FPGA-based streaming ensemble anomaly detector (fSEAD). To achieve this, we propose a flexible computing architecture consisting of multiple partially reconfigurable regions, pblocks, which each implement anomaly detectors. Our proof-of-concept design supports three state-of-the-art anomaly detection algorithms: Loda, RS-Hash, and xStream. Each algorithm is scalable, meaning multiple instances can be placed within a pblock to improve performance. Moreover, fSEAD is implemented using High-level synthesis (HLS), meaning further custom anomaly detectors can be supported. Pblocks are interconnected via an AXI-switch, enabling them to be composed in an arbitrary fashion before combining and merging results at runtime to create an ensemble that maximizes the use of FPGA resources and accuracy. Through utilizing reconfigurable Dynamic Function eXchange (DFX), the detector can be modified at runtime to adapt to changing environmental conditions. We compare fSEAD to an equivalent central processing unit (CPU) implementation using four standard datasets, with speedups ranging from 3× to 8×. Binglei Lou, David Boland, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | Rethinking Embedded Blocks for Machine Learning ApplicationsabstractThe underlying goal of FPGA architecture research is to devise flexible substrates that implement a wide variety of circuits efficiently. Contemporary FPGA architectures have been optimized to support networking, signal processing, and image processing applications through high-precision digital signal processing (DSP) blocks. The recent emergence of machine learning has created a new set of demands characterized by: (1) higher computational density and (2) low precision arithmetic requirements. With the goal of exploring this new design space in a methodical manner, we first propose a problem formulation involving computing nested loops over multiply-accumulate (MAC) operations, which covers many basic linear algebra primitives and standard deep neural network (DNN) kernels. A quantitative methodology for deriving efficient coarse-grained compute block architectures from benchmarks is then proposed together with a family of new embedded blocks, called MLBlocks. An MLBlock instance includes several multiply-accumulate units connected via a flexible routing, where each configuration performs a few parallel dot-products in a systolic array fashion. This architecture is parameterized with support for different data movements, reuse, and precisions, utilizing a columnar arrangement that is compatible with existing FPGA architectures. On synthetic benchmarks, we demonstrate that for 8-bit arithmetic, MLBlocks offer 6× improved performance over the commercial Xilinx DSP48E2 architecture with smaller area and delay; and for time-multiplexed 16-bit arithmetic, achieves 2× higher performance per area with the same area and frequency. All source codes and data, along with documents to reproduce all the results in this article, are available at http://github.com/raminrasoulinezhad/MLBlocks . Seyedramin Rasoulinezhad, Esther Roorda, Steve Wilton, Philip H. W. Leong, David Boland |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2021 | MLBlocks: FPGA Blocks for Machine Learning ApplicationsabstractThe underlying goal of FPGA architecture research is to devise flexible substrates which implement a wide variety of circuits efficiently. Contemporary FPGA architectures have been optimized to support networking, signal processing and image processing applications through high precision digital signal processing (DSP) blocks. The recent emergence of machine learning has created a new set of demands characterized by: 1) higher computational density and 2) low precision arithmetic requirements. With the goal of exploring this new design space in a methodical manner, we first propose a problem formulation involving computing nested loops over multiply-accumulate (MAC) operations, which covers many basic linear algebra primitives and standard deep neural network (DNN) layers. A quantitative methodology for deriving efficient coarse-grained compute block architectures from benchmarks is then proposed together with a family of new compute units, called MLBlocks. These blocks are flexible mesh-based systolic array units parameterized with different data movements, data reuse, and multi-precision support. They utilize a columnar arrangement which is compatible with existing FPGA architectures. Finally, using synthetic benchmarks, we demonstrate that MLBlocks offer significantly improved performance over the commercial Xilinx DSP48E2, while maintaining similar area and timing requirements to current DSPs. Seyedramin Rasoulinezhad, David Boland, Philip H. W. Leong |
FPGA | 2 |
| 2021 | A Block Minifloat Representation for Training Deep Neural Networks
Sean Fox, Seyedramin Rasoulinezhad, Julian Faraone, David Boland, Philip H. W. Leong |
ICLR | 4 |
| 2020 | LUXOR: An FPGA Logic Cell Architecture for Efficient Compressor Tree ImplementationsabstractWe propose two tiers of modifications to FPGA logic cell architecture to deliver a variety of performance and utilization benefits with only minor area overheads. In the first tier, we augment existing commercial logic cell datapaths with a 6-input XOR gate in order to improve the expressiveness of each element, while maintaining backward compatibility. This new architecture is vendor-agnostic, and we refer to it as LUXOR. We also consider a secondary tier of vendor-specific modifications to both Xilinx and Intel FPGAs, which we refer to as X-LUXOR+ and I-LUXOR+ respectively. We demonstrate that compressor tree synthesis using generalized parallel counters (GPCs) is further improved with the proposed modifications. Using both the Intel adaptive logic module and the Xilinx slice at the 65nm technology node for a comparative study, it is shown that the silicon area overhead is less than 0.5% for LUXOR and 5-6% for LUXOR+, while the delay increments are 1-6% and 3-9% respectively. We demonstrate that LUXOR can deliver an average reduction of 13-19% in logic utilization on micro-benchmarks from a variety of domains. BNN benchmarks benefit the most with an average reduction of 37-47% in logic utilization, which is due to the highly-efficient mapping of the XnorPopcount operation on our proposed LUXOR+ logic cells. Seyedramin Rasoulinezhad, Siddhartha 0003, Hao Zhou 0008, Lingli Wang, David Boland, Philip H. W. Leong |
FPGA | 5 |
| 2020 | AddNet: Deep Neural Networks Using FPGA-Optimized MultipliersabstractLow-precision arithmetic operations to accelerate deep-learning applications on field-programmable gate arrays (FPGAs) have been studied extensively, because they offer the potential to save silicon area or increase throughput. However, these benefits come at the cost of a decrease in accuracy. In this article, we demonstrate that reconfigurable constant coefficient multipliers (RCCMs) offer a better alternative for saving the silicon area than utilizing low-precision arithmetic. RCCMs multiply input values by a restricted choice of coefficients using only adders, subtractors, bit shifts, and multiplexers (MUXes), meaning that they can be heavily optimized for FPGAs. We propose a family of RCCMs tailored to FPGA logic elements to ensure their efficient utilization. To minimize information loss from quantization, we then develop novel training techniques that map the possible coefficient representations of the RCCMs to neural network weight parameter distributions. This enables the usage of the RCCMs in hardware, while maintaining high accuracy. We demonstrate the benefits of these techniques using AlexNet, ResNet-18, and ResNet-50 networks. The resulting implementations achieve up to 50% resource savings over traditional 8-bit quantized networks, translating to significant speedups and power savings. Our RCCM with the lowest resource requirements exceeds 6-bit fixed point accuracy, while all other implementations with RCCMs achieve at least similar accuracy to an 8-bit uniformly quantized design, while achieving significant resource savings. Julian Faraone, Martin Kumm, Martin Hardieck, Peter Zipf, Xueyuan Liu 0002, David Boland, Philip H. W. Leong |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2019 | Unrolling Ternary Neural NetworksabstractThe computational complexity of neural networks for large-scale or real-time applications necessitates hardware acceleration. Most approaches assume that the network architecture and parameters are unknown at design time, permitting usage in a large number of applications. This article demonstrates, for the case where the neural network architecture and ternary weight values are known a priori , that extremely high throughput implementations of neural network inference can be made by customising the datapath and routing to remove unnecessary computations and data movement. This approach is ideally suited to FPGA implementations as a specialized implementation of a trained network improves efficiency while still retaining generality with the reconfigurability of an FPGA. A VGG-style network with ternary weights and fixed point activations is implemented for the CIFAR10 dataset on Amazon’s AWS F1 instance. This article demonstrates how to remove 90% of the operations in convolutional layers by exploiting sparsity and compile-time optimizations. The implementation in hardware achieves 90.9 ± 0.1% accuracy and 122k frames per second, with a latency of only 29µs, which is the fastest CNN inference implementation reported so far on an FPGA. Stephen Tridgell, Martin Kumm, Martin Hardieck, David Boland, Duncan J. M. Moss, Peter Zipf, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2019 | A Two-Speed, Radix-4, Serial-Parallel MultiplierabstractIn this paper, we present a two-speed, radix-4, serial-parallel multiplier for accelerating applications such as digital filters, artificial neural networks, and other machine learning algorithms. Our multiplier is a variant of the serial-parallel (SP) modified radix-4 Booth multiplier that adds only the nonzero Booth encodings and skips over the zero operations, making the latency dependent on the multiplier value. Two subcircuits with different critical paths are utilized so that throughput and latency are improved for a subset of multiplier values. The multiplier is evaluated on an Intel Cyclone V field-programmable gate array against standard parallel-parallel and SP multipliers across four different process-voltage-temperature corners. We show that for bit widths of 32 and 64, our optimizations can result in a 1.42×-$3.36× improvement over the standard parallel Booth multiplier in terms of area-time depending on the input set. Duncan J. M. Moss, David Boland, Philip H. W. Leong |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | FPGA Fastfood - A High Speed Systolic Implementation of a Large Scale Online Kernel MethodabstractIn this paper, we describe a systolic Field Programmable Gate Array (FPGA) implementation of the Fastfood algorithm that is optimised to run at a high frequency. The Fastfood algorithm supports online learning for large scale kernel methods. Empirical results show that 500 MHz clock rates can be sustained for an architecture that can solve problems with input dimensions that are $10^3$ times larger than previously reported. Unlike many recent deep learning publications, this design implements both training and prediction. This enables the use of kernel methods in applications requiring a rare combination of capacity, adaption and speed. Sean Fox, David Boland, Philip H. W. Leong |
FPGA | 2 |
| 2018 | Customizing Low-Precision Deep Neural Networks for FPGAsabstractIn this paper, we argue that instead of solely focusing on developing efficient architectures to accelerate well-known low-precision CNNs, we should also seek to modify the network to suit the FPGA. We develop a fully automative toolflow which focuses on modifying the network through filter pruning, such that it efficiently utilizes the FPGA hardware whilst satisfying a predefined accuracy threshold. Although fewer weights are re-moved in comparison to traditional pruning techniques designed for software implementations, the overall model complexity and feature map storage is greatly reduced. We implement the AlexNet and TinyYolo networks on the large-scale ImageNet and PascalVOC datasets, to demonstrate up to roughly 2× speedup in frames per second and 2× reduction in resource requirements over the original network, with equal or improved accuracy. Julian Faraone, Giulio Gambardella, Nicholas J. Fraser, Michaela Blott, Philip H. W. Leong, David Boland |
FPL | 6 |
| 2018 | Simultaneous Inference and Training Using On-FPGA Weight Perturbation TechniquesabstractWe present an FPGA-optimized implementation of online neural network training based on weight perturbation (WP) techniques. When compared to the classic backpropagation (BP) algorithm, WP is capable of delivering competitive performance while occupying minimal area resources. Perturbation-based methods have been demonstrated as viable training techniques and are suitable for on-line learning applications which adapt to changing conditions. The viability of applying WP-based on-chip training for low-precision fixed-point hardware is demonstrated on two distinct MLP benchmarks: the Iris dataset classification network and an RF anomaly detector. When synthesized to a Xilinx Kintex-7 XC7K410T FPGA, WP offers a 3-10x area savings with <;1% degradation in accuracy compared with backpropagation. Compared with an inference-only implementation the overhead of introducing on-chip learning is approximately 30%. Siddhartha 0003, Steve Wilton, David Boland, Barry Flower, Perry Blackmore, Philip H. W. Leong |
FPT | 3 |
| 2018 | Real-time FPGA-based Anomaly Detection for Radio Frequency SignalsabstractWe describe an open source, FPGA accelerated neural network-based anomaly detector. The detector derives its training set from observed exemplar data and continuous learning in software can be undertaken in an unsupervised manner. Trained network weights are passed to the FPGA, which performs continuous high-speed anomaly detection, combining parallelism reduced precision, and a single-chip design to maximise performance and energy efficiency. Our design can process continuous 200 MS/s complex inputs, producing anomaly classifications at the same rate, with a latency of 105 ns, an improvement of at least 4 orders of magnitude over a software radio such as GNU Radio. Duncan J. M. Moss, David Boland, Peyam Pourbeik, Philip H. W. Leong |
ISCAS | 2 |
| 2017 | Dynamic bitwidth assignment for efficient dot productsabstractThe benefits of customising the precision throughout an FPGA design according to a design tolerance are well known. However, customising the precision of a design at runtime has the potential for an even greater performance impact. In this paper, we add the ability to dynamically choose the internal precision of a datapath. This enables a result that is at least as accurate as the worst-case under standard precisions, whilst internally operating at a lower precision. We demonstrate this technique on fused floating-point dot-product circuits. We show that for circuits with inputs that have a wide dynamic range, we can see substantial resource savings. We provide examples with savings of up to 75% of the DSPs and 16% of the ALMs over an optimised fused dot-product design. Simon Joel Schmidt, David Boland |
FPL | 2 |
| 2017 | FPGA acceleration of multilevel ORB feature extraction for computer visionabstractIn this paper, we present the first multilevel implementation of the Harris-Stephens corner detector and the ORB feature extractor running on FPGA hardware, for computer vision and robotics applications. ORB is a fundamental component of many robotics applications, and requires significant computation. The design has been validated both in behavioural simulation and in implementation on an Arria V FPGA connected to a desktop PC via PCI-Express. A Linux kernel-mode driver and userspace library allow integration of the acceleration hardware into C++ programs. The device has significantly higher throughput than a CPU implementation (150 MPixel/s vs 27 MPixel/s) and a GPU implementation (40 MPixel/s), with much lower power draw (5.3 W vs 145 W). This throughput is equivalent to 72 fps at 1920 × 1080 or 488 fps at 640 × 480. Josh Weberruss, Lindsay Kleeman, David Boland, Tom Drummond |
FPL | 3 |
| 2016 | Reducing Memory Requirements for High-Performance and Numerically Stable Gaussian EliminationabstractGaussian elimination is a well-known technique to compute the solution to a system of linear equations and boosting its performance is highly desirable. While straightforward parallel techniques are limited either by I/O or on-chip memory bandwidth, block-based algorithms offer the potential to bridge this gap by interleaving I/O with computation. However, these algorithms require the amount of on-chip memory to be at least the square of the number of processing elements available. Using the latest generation Altera FPGAs with hardened floating-point units, this is no longer the case. It follows that the amount of on-chip memory limits performance, a problem that is only likely to increase unless on-chip memory dominates FPGA architecture. In addition to this limitation, existing FPGA implementations of block-based Gaussian elimination either sacrifice numerical stability or efficiency. The former limits the usefulness of these implementations to a small class of matrices, the latter limits its performance. David Boland |
FPGA | 1 |
| 2015 | Imprecise Datapath Design: An Overclocking ApproachabstractIn this article, we describe an alternative circuit design methodology when considering trade-offs between accuracy, performance, and silicon area. We compare two different approaches that could trade accuracy for performance. One is the traditional approach where the precision used in the datapath is limited to meet a target latency. The other is a proposed new approach which simply allows the datapath to operate without timing closure. We demonstrate analytically and experimentally that on average our approach obtains either smaller errors or equivalent faster operating frequencies in comparison to the traditional approach. This is because the worst case caused by timing violations only happens rarely, while precision loss results in errors to most data. We also show that for basic arithmetic operations such as addition, applying our approach to the simple building block of ripple carry adders can achieve better accuracy or performance than using faster adder designs to achieve similar latency. Kan Shi, David Boland, George A. Constantinides |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2014 | Datapath Synthesis for Overclocking: Online Arithmetic for Latency-Accuracy Trade-offsabstractDigital circuits are currently designed to ensure timing closure. Releasing this constraint by allowing timing violations could lead to significant performance improvements, but conventional forms of computer arithmetic do not fail gracefully when pushed beyond deterministic operation. In this paper we take a fresh look at Online Arithmetic, originally proposed for digit serial operation, and synthesize unrolled digit parallel online operators to allow for graceful degradation. We quantify the impact of timing violation on key arithmetic primitives, and show that substantial performance benefits can be obtained in comparison to binary arithmetic. Since timing errors are caused by long carry chains, these result in errors in least significant digits with online arithmetic, causing less impact than conventional implementations. Using analytical models and empirical FPGA results from an image processing application, we demonstrate an error reduction over 89% and an improvement in SNR of over 20dB for the same clock rate. Kan Shi, David Boland, Edward A. Stott, Samuel Bayliss, George A. Constantinides |
DAC | 2 |
| 2014 | Efficient FPGA implementation of digit parallel online arithmetic operatorsabstractOnline arithmetic has been widely studied for ASIC implementation. Online components were originally designed to perform computations in digit serial with most significant digit (MSD) first, resulting in the ability to chain arithmetic operators together for low latency. More recently, research has shown that digit parallel online operators can fail more gracefully when operating beyond the deterministic clocking region in comparison to operators with conventional arithmetic. Unfortunately, the utilization of online arithmetic operators in the past has required a large area overhead for FPGA implementation. In this paper, we propose novel approaches to implement the key primitives of online arithmetic, adders and multipliers, efficiently on modern Xilinx FPGAs with 6-input LUTs and carry resources. We demonstrate experimentally that in comparison to a direct RTL synthesis, the proposed architectures achieve slice savings of over 67% and 69%, and speed-ups of over 1.2x and 1.5x for adders and multipliers, respectively. As a result, the area overheads of using online adders and multipliers in place of traditional arithmetic primitives is reduced from 8.41 x and 8.11 x to 1.88x and 1.84x respectively. Finally, because an online multiplier generates MSDs first, we also demonstrate the method to create an online multiplier with a reduced precision output that is smaller than a traditional multiplier producing the same result. We show that this can lead to silicon area savings of up to 56%. Kan Shi, David Boland, George A. Constantinides |
FPT | 2 |
| 2013 | Accuracy-Performance Tradeoffs on an FPGA through OverclockingabstractEmbedded applications can often demand stringent latency requirements. While high degrees of parallelism within custom FPGA-based accelerators may help to some extent, it may also be necessary to limit the precision used in the datapath to boost the operating frequency of the implementation. However, by reducing the precision, the engineer introduces quantization error into the design. In this paper, we demonstrate that for many applications it would be preferable to simply overclock the design and accept that timing violations may arise. Since the errors introduced by timing violations occur rarely, they will cause less noise than quantization errors. Through the use of analytical models and empirical results on a Xilinx Virtex-6 FPGA, we show that a geometric mean reduction of 67.9% to 98.8% in error expectation or a geometric mean improvement of 3.1% to 27.6% in operating frequency can be obtained using this alternative design methodology. Kan Shi, David Boland, George A. Constantinides |
FCCM | 2 |
| 2013 | Word-length optimization beyond straight line codeabstractThe silicon area benefits that result from word-length optimization have been widely reported by the FPGA community. However, to date, most approaches are restricted to straight line code, or code that can be converted into straight line code using techniques such as loop-unrolling. In this paper, we take the first steps towards creating analytical techniques to optimize the precision used throughout custom FPGA accelerators for algorithms that contain loops with data dependent exit conditions. To achieve this, we build on ideas emanating from the software verification community to prove program termination. Our idea is to apply word-length optimization techniques to find the minimum precision required to guarantee that a loop with data dependent exit conditions will terminate. Without techniques to analyze algorithms containing these types of loops, a hardware designer may elect to implement every arithmetic operator throughout a custom FPGA-based accelerator using IEEE-754 standard single or double precision arithmetic. With this approach, the FPGA accelerator would have comparable accuracy to a software implementation. However, we show that using our new technique to create custom fixed and floating point designs, we can obtain silicon area savings of up to 50% over IEEE standard single precision arithmetic, or 80% over IEEE standard double precision arithmetic, at the same time as providing guarantees that the created hardware designs will work in practice. David Boland, George A. Constantinides |
FPGA | 1 |
| 2013 | Revisiting the reduction circuit: A case study for simultaneous architecture and precision optimisationabstractWord-length optimisation techniques have traditionally been used to minimise the precision in a fixed hardware datapath subject to a given error tolerance. In this paper, we discuss how using word-length optimisation techniques to structure a hardware datapath can result in designs achieving the same functionality with even less silicon area. To demonstrate this, we revisit the addition reduction circuit and its use within matrix-vector multiplication. Our results show that given freedom over how to parallelise this circuit, for a fixed error and latency budget we can obtain mean silicon area savings of 58% a typical fixed-point design. We achieve this by creating a more numerically stable parallel architecture instead of replicating the initial design. Since freedom over datapath design is common for high-level synthesis tools, we hope this will inspire word-length optimisation techniques to be applied at the same time as making structural decisions within the design flow of these tools. David Boland, George A. Constantinides |
FPT | 1 |
| 2013 | Overclocking datapath for latency-error tradeoffabstractRelaxing constraints of 100% accuracy in datapath can provide the freedom to create designs with better performance or energy efficiency. This paper develops probabilistic models, which enable us to explore these trade-offs for key arithmetic primitives. We show that because specific input patterns are required to cause timing violations and that these patterns arise rarely, a lower expected error can be attained by allowing some timing variations to occur, instead of reducing the precision of a circuit to meet a target latency. Experiments show that a mean reduction of 5.6× ~ 36.7× in error expectation and an improvement of 7.2dB ~ 19.7dB in signal-to-noise ratio can be obtained for practical applications. Kan Shi, David Boland, George A. Constantinides |
ISCAS | 2 |
| 2013 | A Scalable Precision Analysis FrameworkabstractIn embedded computing, typically some form of silicon area or power budget restricts the potential performance achievable. For algorithms with limited dynamic range, custom hardware accelerators manage to extract significant additional performance for such a budget via mapping operations in the algorithm to fixed-point. However, for complex applications requiring floating-point computation, the potential performance improvement over software is reduced. Nonetheless, custom hardware can still customize the precision of floating-point operators, unlike software which is restricted to IEEE standard single or double precision, to increase the overall performance at the cost of increasing the error observed in the final computational result. Unfortunately, because it is difficult to determine if this error increase is tolerable, this task is rarely performed. We present a new analytical technique to calculate bounds on the range or relative error of output variables, enabling custom hardware accelerators to be tolerant of floating point errors by design. In contrast to existing tools that perform this task, our approach scales to larger examples and obtains tighter bounds, within a smaller execution time. Furthermore, it allows a user to trade the quality of bounds with execution time of the procedure, making it suitable for both small and large-scale algorithms. David Boland, George A. Constantinides |
IEEE Trans. Multim. | 1 |
| 2012 | A scalable approach for automated precision analysisabstractThe freedom over the choice of numerical precision is one of the key factors that can only be exploited throughout the datapath of an FPGA accelerator, providing the ability to trade the accuracy of the final computational result with the silicon area, power, operating frequency, and latency. However, in order to tune the precision used throughout hardware accelerators automatically, a tool is required to verify that the hardware will meet an error or range specification for a given precision. Existing tools to perform this task typically suffer either from a lack of tightness of bounds or require a large execution time when applied to large scale algorithms; in this work, we propose an approach that can both scale to larger examples and obtain tighter bounds, within a smaller execution time, than the existing methods. The approach we describe also provides a user with the ability to trade the quality of bounds with execution time of the procedure, making it suitable within a word-length optimization framework for both small and large-scale algorithms. David Boland, George A. Constantinides |
FPGA | 1 |
| 2011 | Bounding Variable Values and Round-Off Effects Using Handelman RepresentationsabstractThe precision used in an algorithm affects the error and performance of individual computations, the memory usage, and the potential parallelism for a fixed hardware budget. This paper describes a new method to determine the minimum precision required to meet a given error specification for an algorithm consisting of the basic algebraic operations. Using this approach, it is possible to significantly reduce the computational word-length in comparison to existing methods, and this can lead to superior hardware designs. We demonstrate the proposed procedure on an iteration of the conjugate gradient algorithm, achieving proofs of bounds that can translate to global word-length savings ranging from a few bits to proving the existence of ranges that must otherwise be assumed to be unbounded when using competing approaches. We also achieve comparable bounds to recent literature in a small fraction of the execution time, with greater scalability. David Boland, George A. Constantinides |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2011 | Optimizing memory bandwidth use and performance for matrix-vector multiplication in iterative methodsabstractComputing the solution to a system of linear equations is a fundamental problem in scientific computing, and its acceleration has drawn wide interest in the FPGA community [Morris et al. 2006; Zhang et al. 2008; Zhuo and Prasanna 2006]. One class of algorithms to solve these systems, iterative methods, has drawn particular interest, with recent literature showing large performance improvements over General-Purpose Processors (GPPs) [Lopes and Constantinides 2008]. In several iterative methods, this performance gain is largely a result of parallelization of the matrix-vector multiplication, an operation that occurs in many applications and hence has also been widely studied on FPGAs [Zhuo and Prasanna 2005; El-Kurdi et al. 2006]. However, whilst the performance of matrix-vector multiplication on FPGAs is generally I/O bound [Zhuo and Prasanna 2005], the nature of iterative methods allows the use of on-chip memory buffers to increase the bandwidth, providing the potential for significantly more parallelism [deLorimier and DeHon 2005]. Unfortunately, existing approaches have generally only either been capable of solving large matrices with limited improvement over GPPs [Zhuo and Prasanna 2005; El-Kurdi et al. 2006; deLorimier and DeHon 2005], or achieve high performance for relatively small matrices [Lopes and Constantinides 2008; Boland and Constantinides 2008]. This article proposes hardware designs to take advantage of symmetrical and banded matrix structure, as well as methods to optimize the RAM use, in order to both increase the performance and retain this performance for larger-order matrices. David Boland, George A. Constantinides |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2010 | Automated Precision Analysis: A Polynomial Algebraic ApproachabstractWhen migrating an algorithm onto hardware, the potential saving that can be obtained by tuning the precision used in the algorithm to meet a range or error specification is often overlooked; the major reason is that it is hard to choose a number system which can guarantee any such specification can be met. Instead, the problem is mitigated by opting to use IEEE standard single or double precision so as to be `no worse' than a software implementation. However, the flexibility in the number representation is one of the key factors that can only be exploited on FPGAs, unlike GPUs and general purpose processors, and hence ignoring this potential significantly limits the performance achievable on an FPGA. To this end, this paper describes a tool which analyses algorithms with given input ranges under a finite precision to provide information that could be used to tune the hardware to the algorithm specifications. We demonstrate the proposed procedure on an iteration of the conjugate gradient algorithm, achieving a reduction in slices of over 40% when meeting the same error specification found by traditional methods. We also show it achieves comparable bounds to recent literature in a small fraction of the execution time, with greater scalability. David Boland, George A. Constantinides |
FCCM | 1 |
| 2008 | An FPGA-based implementation of the MINRES algorithmabstractDue to continuous improvements in the resources available on FPGAs, it is becoming increasingly possible to accelerate floating point algorithms. The solution of a system of linear equations forms the basis of many problems in engineering and science, but its calculation is highly time consuming. The minimum residual algorithm (MINRES) is one method to solve this problem, and is highly effective provided the matrix exhibits certain characteristics. This paper examines an IEEE 754 single precision floating point implementation of the MINRES algorithm on an FPGA. It demonstrates that through parallelisation and heavy pipelining of all floating point components it is possible to achieve a sustained performance of up to 53 GFLOPS on the Virtex5-330T. This compares favourably to other hardware implementations of floating point matrix inversion algorithms, and corresponds to an improvement of nearly an order of magnitude compared to a software implementation. David Boland, George A. Constantinides |
FPL | 1 |