EDBT 2026 Demo / reviewers in the wild / expert
Suhaib A. Fahmy
dblp:83/2660 · also Suhaib Ahmed Fahmy
· DBLP profile ↗
96ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0003-0568-5048ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 87 · 8 first-author · 19 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Configurable p-Neurons Using Modular p-Bits
Saleh Bunaiyan, Mohammad Alsharif, Abdelrahman S. Abdelrahman, Hesham ElSawy, Suraj S. Cheema, Suhaib A. Fahmy, Kerem Yunus Çamsari, Feras Al-Dirini |
ISCAS | 6 |
| 2025 | Automatic Code-Generation for Accelerating Structured-Mesh-Based Explicit Numerical Solvers on FPGAsabstractStructured-mesh-based stencil computations are a common motif in many numerical algorithms, such as for solving PDEs. Recent work has shown promising runtime performance and energy efficiency when mapping these applications to FPGAs. However, this requires significant manual effort with hardwarespecific optimizations and customization. We present a new codegeneration framework that automates the mapping of stencil applications to FPGAs. From a domain-specific declaration, the framework applies radical, ad-hoc, and dynamic optimizations, including a novel window-buffer chaining scheme, and an on-chip loopback pipeline approach that minimizes memory transaction overhead. We show that the generated code is able to match or exceed the performance of hand-tuned state-of-the-art designs. A range of stencil solvers benchmarked, including non-trivial multi-stencil solvers, on two FPGAs, demonstrating performanceportability. Comparisons to best-in-class solvers on an Nvidia H100 GPU indicate competitive performance with 2-19× energy savings. Beniel Thileepan, Suhaib A. Fahmy, Gihan R. Mudalige |
PACT | 2 |
| 2025 | High Throughput Low Latency Network Intrusion Detection on FPGAs: A Raw Packet Approach
Abid Rafique, Suhaib A. Fahmy, Aman Arora 0001 |
FPGA | 3 |
| 2025 | Exploring Microscaling MX Minifloat Systolic Arrays on FPGAsabstractThe rapid advancement of generative AI and exponential growth of model parameters has driven the pursuit of alternative arithmetic formats to enhance efficiency while maintaining inference accuracy. While earlier efforts to implement small floating point formats (minifloats) have been somewhat ad-hoc, the Microscaling MX formats proposed by the Open Compute Project offer a standard around which hardware designers can converge. This article explores the design space of systolic array architectures for Microscaling MX formats on FPGAs. It provides a detailed analysis of the area and timing characteristics of systolic arrays implemented on AMD UltraScale+ FPGAs for 6-bit and 8-bit MX formats, exploring different accumulation and pipelining strategies while maintaining high computational throughput. Our most optimized design reaches a peak throughput of 568 GOPS using 4-stage and 3-stage pipelined Exact accumulation for MXFP6 E3M2 and E2M3, respectively. Abdurauf Abdurakhmanov, Suhaib A. Fahmy |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2025 | ResiLogic: Leveraging Composability and Diversity to Design Fault and Intrusion Resilient ChipsabstractA long-standing challenge is the design of chips resilient to faults and glitches. Both fine-grained gate diversity and coarse-grained modular redundancy have been used in the past. However, these approaches have not been well-studied under other threat models where some stakeholders in the supply chain are untrusted. Increasing digital sovereignty tensions raise concerns regarding the use of foreign off-the-shelf tools and intellectual property (IP), or off-sourcing fabrication, driving research into the design of resilient chips under this threat model. This article addresses a threat model considering three pertinent attacks to resilience: distribution, zonal, and compound attacks. To mitigate these attacks, we introduce theResiLogicframework that exploitsDiversity by Composability: constructing diverse circuits composed of smaller diverse ones by design. This approach enables designers to develop circuits in the early stages of design without the need for additional redundancy in terms of space or cost. To generate diverse circuits, we propose a technique using E-Graphs with new rewrite definitions for diversity. Using this approach at different levels of granularity is shown to improve the resilience of circuit design inResiLogicup to$\times 5$against the three considered attacks. Ahmad T. Sheikh, Ali Shoker, Suhaib A. Fahmy, Paulo Veríssimo |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | Leveraging MLIR for Efficient Irregular-Shaped CGRA Overlay Design: (PhD Forum Paper)abstractCoarse-grained reconfigurable arrays (CGRAs) are reconfigurable architectures that combine the efficiency of custom datapaths with the flexibility of FPGAs. Optimized word-level functional units with reconfigurable interconnect enable a wide range of applications. Due to the irregular data movement patterns and the diversity of real-world HPC applications, a traditional homogeneous regular-shaped CGRA architecture can be sub-optimal. A more application-specific arrangement of resources could, however, result in poor hardware reuse and increased deployment effort for a CGRA. Using CGRA overlays can increase hardware reuse while maintaining high performance of a more optimized architecture. This proposal uses and builds upon the built-in tools of the MLIR infrastructure to target a class of applications that can be analyzed, rewritten, and optimized based on the available resources of an underlying CGRA architecture, thereby better exploiting hardware resources. Mohamed Bouaziz 0002, Suhaib A. Fahmy |
ASAP | 2 |
| 2024 | High Throughput Massive MIMO Signal Decoding Using Multi-Level Tree Search on FPGAsabstractSupporting the evolution of wireless communication beyond 5G using high-performance networks requires massive device connectivity. Massive Multiple-Input Multiple-Output (MIMO) systems have been used and proven to increase the data throughput of wireless links. However, scaling such systems to a large number of antennas and a high modulation factor entails a significant computational cost for signal decoding using conventional non-linear decoders. Heuristic tree-search-based approaches have been proposed to address this challenge as a means to achieve real-time decoding requirements for large MIMO configurations. FPGAs represent an ideal platform for accelerating MIMO signal decoding on account of their low latency and potential for integration within the signal processing chain while consuming low power. This paper presents soft-ware/hardware co-design for a multi-level tree search approach that integrates the computation of multiple tree levels. The proposed heuristic transforms the tree search process into a streaming operation well suited for the FPGA's architecture. We show a series of hardware design and algorithmic optimizations that significantly improve scalability and decoding time, resulting in a design capable of decoding 64×64 64-QAM MIMO within 10ms real-time requirements. Mohamed W. Hassan, Hatem Ltaief, Suhaib A. Fahmy |
FCCM | 3 |
| 2024 | Resilient and Secure Programmable System-on-Chip Accelerator OffloadabstractComputational offload to hardware accelerators is gaining traction due to increasing computational demands and efficiency challenges. Programmable hardware, like FPGAs, offers a promising platform in rapidly evolving application areas, with the benefits of hardware acceleration and software programmability. Unfortunately, such systems composed of multiple hardware components must consider integrity in the case of malicious components. In this work, we propose Samsara, the first secure and resilient platform that derives, from Byzantine Fault Tolerance (BFT), protocols to enhance the computing resilience of programmable hardware. Samsara uses a novel lightweight hardware-based BFT protocol for Systems-on-Chip, called H-Quorum, that implements the theoretical-minimum latency between applications and replicated compute nodes. To withstand malicious behaviors, Samsara supports hardware rejuvenation, which is used to replace, relocate, or diversify faulty compute nodes. Samsara's architecture ensures the security of the entire workflow while keeping the latency overhead, of both computation and rejuvenation, close to the non-replicated counterpart. Inês Pinto Gouveia, Ahmad T. Sheikh, Ali Shoker, Suhaib A. Fahmy, Paulo Veríssimo |
SRDS | 4 |
| 2024 | Introduction to the Special Section on FPGA 2023
Suhaib A. Fahmy, Jason D. Bakos |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2023 | REFL: Resource-Efficient Federated LearningabstractFederated Learning (FL) enables distributed training by learners using local data, thereby enhancing privacy and reducing communication. However, it presents numerous challenges relating to the heterogeneity of the data distribution, device capabilities, and participant availability as deployments scale, which can impact both model convergence and bias. Existing FL schemes use random participant selection to improve the fairness of the selection process; however, this can result in inefficient use of resources and lower quality training. In this work, we systematically address the question of resource efficiency in FL, showing the benefits of intelligent participant selection, and incorporation of updates from straggling participants. We demonstrate how these factors enable resource efficiency while also improving trained model quality. Ahmed M. Abdelmoniem, Atal Narayan Sahu, Marco Canini, Suhaib A. Fahmy |
EuroSys | 4 |
| 2023 | Exploring FPGA Acceleration for Distributed Serverless ComputingabstractServerless computing has become a popular cloud computing paradigm. However, its deployment abstraction entails significant performance overheads. We explore the potential for enabling serverless computing on FPGAs and present some early results that show the concurrency and scalability benefits on a stream processing workload. The FPGA enables flows of network data to flow directly into accelerators, which is beneficial for scenarios involving large packets and multiple request streams. Ziyi Yang 0001, Suhaib A. Fahmy |
FPL | 2 |
| 2023 | Signal Detection for Large MIMO Systems Using Sphere Decoding on FPGAsabstractWireless communication systems rely on aggressive spatial multiplexing Multiple-Input Multiple-Output (MIMO) access points to enhance network throughput. A significant computational hurdle for large MIMO systems is signal detection and decoding, which has exponentially increasing computational complexity as the number of antennas increases. Hence, the feasibility of large MIMO systems depends on suitable implementations of signal decoding schemes.This paper presents an FPGA-based Sphere Decoder (SD) architecture that provides high-performance signal decoding for large MIMO systems, supporting up to 16-QAM modulation. The SD algorithm is refactored to map well to the FPGA architecture using a GEMM-based approach to exploit the parallel computational power of FPGAs. We implement FPGA-specific optimization techniques to improve computational complexity. We show significant improvement in time to decode the received signal with under 10–2BER. The design is deployed on a Xilinx Alveo U280 FPGA and shows up to a 9× speedup compared to optimized multi-core CPU execution, achieving real-time requirements. Our proposed design reduces power consumption by a geo-mean of 38.1× compared to CPU implementation, which is important in real-world deployments. We also evaluate our design against alternative approaches on GPU. Mohamed W. Hassan, Adel Dabah, Hatem Ltaief, Suhaib A. Fahmy |
IPDPS | 4 |
| 2023 | ZyPR: End-to-end Build Tool and Runtime Manager for Partial Reconfiguration of FPGA SoCs at the EdgeabstractPartial reconfiguration (PR) is a key enabler to the design and development of adaptive systems on modern Field Programmable Gate Array (FPGA) Systems-on-Chip (SoCs), allowing hardware to be adapted dynamically at runtime. Vendor-supported PR infrastructure is performance-limited and blocking, drivers entail complex memory management, and software/hardware design requires bespoke knowledge of the underlying hardware. This article presents ZyPR: a complete end-to-end framework that provides high-performance reconfiguration of hardware from within a software abstraction in the Linux userspace, automating the process of building PR applications with support for the Xilinx Zynq and Zynq UltraScale+ architectures, aimed at enabling non-expert application designers to leverage PR for edge applications. We compare ZyPR against traditional vendor tooling for PR management as well as recent open source tools that support PR under Linux. The framework provides a high-performance runtime along with low overhead for its provided abstractions. We introduce improvements to our previous work, increasing the provisioning throughput for PR bitstreams on the Zynq Ultrascale+ by 2× and 5.4× compared to Xilinx’s FPGA Manager. Alex R. Bucknall, Suhaib A. Fahmy |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2023 | Streaming Overlay Architecture for Lightweight LSTM Computation on FPGA SoCsabstractLong-Short Term Memory (LSTM) networks, and Recurrent Neural Networks (RNNs) in general, have demonstrated their suitability in many time series data applications, especially in Natural Language Processing (NLP) . Computationally, LSTMs introduce dependencies on previous outputs in each layer that complicate their computation and the design of custom computing architectures, compared to traditional feed-forward networks. Most neural network acceleration work has focused on optimising the core matrix-vector operations on highly capable FPGAs in server environments. Research that considers the embedded domain has often been unsuitable for streaming inference, relying heavily on batch processing to achieve high throughput. Moreover, many existing accelerator architectures have not focused on fully exploiting the underlying FPGA architecture, resulting in designs that achieve lower operating frequencies than the theoretical maximum. This paper presents a flexible overlay architecture for LSTMs on FPGA SoCs that is built around a streaming dataflow arrangement, uses DSP block capabilities directly, and is tailored to keep parameters within the architecture while moving input data serially to mitigate external memory access overheads. The architecture is designed as an overlay that can be configured to implement alternative models or update model parameters at runtime. It achieves higher operating frequency and demonstrates higher performance than other lightweight LSTM accelerators, as demonstrated in an FPGA SoC implementation. Lenos Ioannou, Suhaib A. Fahmy |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | High throughput multidimensional tridiagonal system solvers on FPGAsabstractWe present a high performance tridiagonal solver library for Xilinx FPGAs optimized for multiple multi-dimensional systems common in real-world applications. An analytical performance model is developed and used to explore the design space and obtain rapid performance estimates that are over 85% accurate. This library achieves an order of magnitude better performance when solving large batches of systems than previous FPGA work. A detailed comparison with a current state-of-the-art GPU library for multi-dimensional tridiagonal systems on an Nvidia V100 GPU shows the FPGA achieving competitive or better runtime and significant energy savings of over 30%. Through this design, we learn lessons about the types of applications where FPGAs can challenge the current dominance of GPUs. Kamalakkannan Kamalavasan, Gihan R. Mudalige, István Z. Reguly, Suhaib A. Fahmy |
ICS | 4 |
| 2022 | Estimation of missing air pollutant data using a spatiotemporal convolutional autoencoderabstractAbstract A key challenge in building machine learning models for time series prediction is the incompleteness of the datasets. Missing data can arise for a variety of reasons, including sensor failure and network outages, resulting in datasets that can be missing significant periods of measurements. Models built using these datasets can therefore be biased. Although various methods have been proposed to handle missing data in many application areas, more air quality missing data prediction requires additional investigation. This study proposes an autoencoder model with spatiotemporal considerations to estimate missing values in air quality data. The model consists of one-dimensional convolution layers, making it flexible to cover spatial and temporal behaviours of air contaminants. This model exploits data from nearby stations to enhance predictions at the target station with missing data. This method does not require additional external features, such as weather and climate data. The results show that the proposed method effectively imputes missing data for discontinuous and long-interval interrupted datasets. Compared to univariate imputation techniques (most frequent, median and mean imputations), our model achieves up to 65% RMSE improvement and 20–40% against multivariate imputation techniques (decision tree, extra-trees, k -nearest neighbours and Bayesian ridge regressors). Imputation performance degrades when neighbouring stations are negatively correlated or weakly correlated. I Nyoman Kusuma Wardana, Julian W. Gardner, Suhaib A. Fahmy |
Neural Comput. Appl. | 3 |
| 2022 | Correction to: Estimation of missing air pollutant data using a spatiotemporal convolutional autoencoder
I Nyoman Kusuma Wardana, Julian W. Gardner, Suhaib A. Fahmy |
Neural Comput. Appl. | 3 |
| 2022 | Coarse Grained FPGA Overlay for Rapid Just-In-Time Accelerator CompilationabstractCoarse-grained FPGA overlays built around the runtime programmable DSP blocks in modern FPGAs can achieve high throughput and improved scalability compared to traditional overlays built without detailed consideration of FPGA architecture. These overlays can be mapped to using higher level compilers, achieving fast compilation, software-like programmability and run-time management, and high-level design abstraction. OpenCL allows programs running on a host computer to launch accelerator kernels which can be compiled at run-time for a specific architecture, thus enabling portability. However, prohibitive hardware compilation times in traditional design flows mean that the tools cannot effectively use just-in-time (JIT) compilation or runtime performance scaling on FPGAs. We present a methodology for runtime compilation of dataflow graphs expressed as OpenCL kernels onto coarse-grained overlays. The methodology benefits from the high level of abstraction afforded by using the OpenCL programming model, while the mapping to the overlay significantly reduces compilation and load times. Key characteristics of this work include highly performant DSP-optimized functional units that scale to large overlays on modern devices and the ability to perform automatic resource-aware kernel replication up to the size of the overlay. We demonstrate place and route times orders of magnitude better than traditional HLS flows, even when running on an embedded processor in the Xilinx Zynq. Abhishek Kumar Jain, Douglas L. Maskell, Suhaib A. Fahmy |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Sit Here: Placing Virtual Machines Securely in Cloud EnvironmentsabstractA Cloud Computing Environment (CCE) leverages the advantages offered by virtualisation to enable virtual machines (VMs) within the same physical machine (PM) to share physical resources. Cloud service providers (CSPs) accommodate the fluctuating resource demands of cloud users dynamically, through elastic resource provisioning. CSPs use VM allocation techniques such as VM placement and VM migration to optimise the use of shared physical resources in the CCE. However, these techniques are exposed to potential security threats that can lead to the problem of malicious co-residency between VMs. This threat happens when a malicious VM is co-located with a critical (or target) VM on the same PM. Hence, the VM allocation techniques need to be made secure. While earlier works propose specific solutions to address this malicious co-residency problem, our work here proposes to investigate the allocation patterns that are more likely to lead to a secure allocation. Furthermore, we introduce a s ecurity-aware VM allocation algorithm (SRS) that aims to allocate the VMs securely, to reduce the potential for co-residency between malicious and target VMs. Our study shows: (i) our SRS algorithm outperforms all state-of-the-art allocation algorithms and (ii) algorithms that adopt stacking-based behaviours are more likely to return secure allocations than those with spreading or random behaviours. Mansour Aldawood, Arshad Jhumka, Suhaib A. Fahmy |
CLOSER | 3 |
| 2021 | Runtime Abstraction for Autonomous Adaptive Systems on Reconfigurable HardwareabstractAutonomous systems increasingly rely on on-board computation to avoid the latency overheads of offloading to more powerful remote computing. This requires the integration of hardware accelerators to handle the complex computations demanded by data-intensive sensors. FPGAs offer hardware acceleration with ample flexibility and interfacing capabilities when paired with general purpose processors, with the ability to reconfigure at runtime using partial reconfiguration (PR). Managing dynamic hardware is complex and has been left to designers to address in an ad-hoc manner, without first-class integration in autonomous software frameworks. This paper presents an abstracted runtime for managing adaptation of FPGA accelerators, including PR and parametric changes, that presents as a typical interface used in autonomous software systems. We present a demonstration using the Robot Operating System (ROS), showing negligible latency overhead as a result of the abstraction. Alex R. Bucknall, Suhaib A. Fahmy |
DATE | 2 |
| 2021 | Heterogeneous Communication Virtualization for Distributed Embedded ApplicationsabstractDistributed embedded applications (DEAs) are typically implemented on diverse embedded nodes interconnected through communication network(s) to exchange data and control information to achieve the desired functionality. Conventional approaches of utilising a single large-bandwidth link in a distributed system are not efficient in large DEAs owing to diverse requirements and factors like cost, reliability, scalability and criticality, among others. Heterogeneous communication is a promising approach in DEAs, where the diverse nature of underlying protocols (wired/wireless, synchronous/asynchronous, multiple access modes and others) can be leveraged to meet such requirements, in addition to the benefits like aggregated bandwidth and robustness. However, utilising them ‘directly’ places significant complexity on the application as it needs to dynamically evaluate the channels and utilise different protocol structures for each case. Virtualising the communication channels would present a unified interface to the application by abstracting away low-level details, similar to virtualisation applied in compute architectures. However, unlike architecture virtualisation, virtualising heterogeneous communication particularly for resource-constrained device networks involves unique challenges imposed by the physical (wired/wireless) and logical domains (limited-bandwidth, small payload, protocols, channel access schemes, etc.), which needs to be concurrently evaluated to optimise the communication system. This paper presents a model and an optimal transmission strategy as the proof of concept for deploying heterogeneous communication in DEAs. The model is described at an abstracted level while capturing transmission parameters of multiple channels, which are then optimised to meet the application’s communication requirements. The model and the optimisation method are validated through simulation and a practical case study. Thinh Hung Pham, Shanker Shreejith, Sebastian Steinhorst, Suhaib A. Fahmy, Samarjit Chakraborty |
DSD | 4 |
| 2021 | High-Level FPGA Accelerator Design for Structured-Mesh-Based Explicit Numerical SolversabstractThis paper presents a workflow for synthesizing near-optimal FPGA implementations of structured-mesh based stencil applications for explicit solvers. It leverages key characteristics of the application class and its computation-communication pattern and the architectural capabilities of the FPGA to accelerate solvers for high-performance computing applications. Key new features of the workflow are (1) the unification of standard state-of-the-art techniques with a number of high-gain optimizations such as batching and spatial blocking/tiling, motivated by increasing throughput for real-world workloads and (2) the development and use of a predictive analytical model to explore the design space, and obtain resource and performance estimates. Three representative applications are implemented using the design workflow on a Xilinx Alveo U280 FPGA, demonstrating near-optimal performance and over 85% predictive model accuracy. These are compared with equivalent highly-optimized implementations of the same applications on modern HPC-grade GPUs (Nvidia V100), analyzing time to solution, bandwidth, and energy consumption. Performance results indicate comparable runtimes with the V100 GPU, with over 2× energy savings for the largest non-trivial application on the FPGA. Our investigation shows the challenges of achieving high performance on current generation FPGAs compared to traditional architectures. We discuss determinants for a given stencil code to be amenable to FPGA implementation, providing insights into the feasibility and profitability of a design and its resulting performance. Kamalakkannan Kamalavasan, Gihan R. Mudalige, István Z. Reguly, Suhaib A. Fahmy |
IPDPS | 4 |
| 2021 | Power-Efficient Mapping of Large Applications on Modern Heterogeneous FPGAsabstractThe increasing size of modern field programmable gate arrays (FPGAs) allows for ever more complex applications to be mapped onto them. However, long design implementation times for large designs can severely affect design productivity. A modular design methodology can improve design productivity in a divide and conqueror fashion but at the expense of degraded performance and power consumption of the resulting implementation. To reduce the dominant power dissipation component in FPGAs, the routing power, methodologies have been proposed that consider data communication between modules during module formation and placement on the FPGA. Selecting proper mapping region on target FPGAs, on the other hand, is becoming a critical process because of the heterogeneous resources and column arrangements in modern FPGAs. Selecting inappropriate FPGA regions for mapping could lead to degraded performance. Hence, we propose a methodology that uses communication-aware module placement, such that modules are mapped by selecting the best shape and region on the FPGA factoring the columnar resource arrangements. Additionally, techniques for module locking and splitting have been proposed for deterministic convergence of the algorithm and for improved module placement. This methodology exhibits nearly 19% routing power reduction with respect to commercial CAD flows without any degradation in achievable performance. Kalindu Herath, Alok Prakash, Suhaib A. Fahmy, Thambipillai Srikanthan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Characterizing Latency Overheads in the Deployment of FPGA AcceleratorsabstractFPGA hardware accelerators have recently enjoyed significant attention as platforms for further accelerating computation in the datacenter but they potentially add additional layers of hardware and software interfacing that can further increase communication latency. In this paper, we characterize these overheads for streaming applications where latency can be an important consideration. We examine the latency and throughput characteristics of traditional server-based PCIe connected accelerators, and the more recent approach of network attached FPGA accelerators. We additionally quantify the additional overhead introduced by virtualising accelerators on FPGAs. Ryan A. Cooke, Suhaib A. Fahmy |
FPL | 2 |
| 2020 | High Throughput Accelerator Interface Framework for a Linear Time-Multiplexed FPGA OverlayabstractCoarse-grained FPGA overlays improve design productivity through software-like programmability and fast compilation. However, the effectiveness of overlays as accelerators is dependent on suitable interface and programming integration into a typically processor-based computing system, an aspect which has often been neglected in evaluations of overlays. We explore the integration of a time-multiplexed FPGA overlay over a server-class PCI Express interface. We show how this integration can be optimised to maximise performance, and evaluate the area overhead. We also propose a user-friendly programming model for such an overlay accelerator system. Xiangwei Li, Kizheppatt Vipin, Douglas L. Maskell, Suhaib A. Fahmy, Abhishek Kumar Jain |
ISCAS | 4 |
| 2020 | A model for distributed in-network and near-edge computing with heterogeneous hardware
Ryan A. Cooke, Suhaib A. Fahmy |
Future Gener. Comput. Syst. | 2 |
| 2020 | High Throughput Spatial Convolution Filters on FPGAsabstractDigital signal processing (DSP) on field-programmable gate arrays (FPGAs) has long been appealing because of the inherent parallelism in these computations that can be easily exploited to accelerate such algorithms. FPGAs have evolved significantly to further enhance the mapping of these algorithms, included additional hard blocks, such as the DSP blocks found in modern FPGAs. Although these DSP blocks can offer more efficient mapping of DSP computations, they are primarily designed for 1-D filter structures. We present a study on spatial convolutional filter implementations on FPGAs, optimizing around the structure of the DSP blocks to offer high throughput while maintaining the coefficient flexibility that other published architectures usually sacrifice. We show that it is possible to implement large filters for large 4K resolution image frames at frame rates of 30-60 FPS, while maintaining functional flexibility. Lenos Ioannou, Abdullah Al-Dujaili, Suhaib A. Fahmy |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | Network Intrusion Detection Using Neural Networks on FPGA SoCsabstractNetwork security is increasing in importance as systems become more interconnected. Much research has been conducted on large appliances for network security, but these do not scale well to lightweight systems such as those used in the Internet of Things (IoT). Meanwhile, the low power processors used in IoT devices do not have the required performance for detailed packet analysis. We present an approach for network intrusion detection using neural networks, implemented on FPGA SoC devices that can achieve the required performance on embedded systems. The design is flexible, allowing model updates in order to adapt to emerging attacks. Lenos Ioannou, Suhaib A. Fahmy |
FPL | 2 |
| 2019 | Neural Network Overlay Using FPGA DSP BlocksabstractWith the increasing wider application of neural networks, there has been significant focus on accelerating this class of computations. Larger, more complex networks are being proposed in a variety of domains, requiring more powerful computation platforms. The inherent parallelism and regularity of neural network structures means custom architectures can be adopted for this purpose. FPGAs have been widely used to implement such accelerators because of their flexibility, achievable performance, efficiency, and abundant peripherals. While platforms that utilize multicore CPUs and GPUs are also competitive, FPGAs offer superior energy efficiency, and a wider space of optimisations to enhance performance and efficiency. FPGAs are also more suitable for performing such computations at the edge, where multicore CPUs and GPUs are are less likely to be used and energy efficiency is paramount. Lenos Ioannou, Suhaib A. Fahmy |
FPL | 2 |
| 2019 | The Power-optimised Software EnvelopeabstractAdvances in processor design have delivered performance improvements for decades. As physical limits are reached, refinements to the same basic technologies are beginning to yield diminishing returns. Unsustainable increases in energy consumption are forcing hardware manufacturers to prioritise energy efficiency in their designs. Research suggests that software modifications may be needed to exploit the resulting improvements in current and future hardware. New tools are required to capitalise on this new class of optimisation. In this article, we present the Power Optimised Software Envelope (POSE) model, which allows developers to assess the potential benefits of power optimisation for their applications. The POSE model is metric agnostic and in this article, we provide derivations using the established Energy-Delay Product metric and the novel Energy-Delay Sum and Energy-Delay Distance metrics that we believe are more appropriate for energy-aware optimisation efforts. We demonstrate POSE on three platforms by studying the optimisation characteristics of applications from the Mantevo benchmark suite. Our results show that the Pathfinder application has very little scope for power optimisation while TeaLeaf has the most, with all other applications in the benchmark suite falling between the two. Finally, we extend our POSE model with a formulation known as System Summary POSE—a meta-heuristic that allows developers to assess the scope a system has for energy-aware software optimisation independent of the code being run. Stephen I. Roberts, Steven A. Wright 0001, Suhaib A. Fahmy, Stephen A. Jarvis |
ACM Trans. Archit. Code Optim. | 3 |
| 2018 | A time-multiplexed FPGA overlay with linear interconnectabstractCoarse-grained overlays improve FPGA design productivity by providing fast compilation and software like programmability. Soft processor based overlays with well-defined ISAs are attractive to application developers due to their ease of use. However, these overlays have significant FPGA resource overheads. Time multiplexed (TM) CGRA-like overlays represent an interesting alternative as they are able to change their behavior on a cycle by cycle basis while the compute kernel executes. This reduces the FPGA resource needed, but at the cost of a higher initiation interval (II) and hence reduced throughput. The fully flexible routing network of current CGRA-like overlays results in high FPGA resource usage. However, many application kernels are acyclic and can be implemented using a much simpler linear feed-forward routing network. This paper examines a DSP block based TM overlay with linear interconnect where the overlay architecture takes account of the application kernels' characteristics and the underlying FPGA architecture, so as to minimize the II and the FPGA resource usage. We examine a number of architectural extensions to the DSP block based functional unit to improve the II, throughput and latency. The results show an average 70% reduction in II, with corresponding improvements in throughput and latency. Xiangwei Li, Abhishek Kumar Jain, Douglas L. Maskell, Suhaib A. Fahmy |
DATE | 4 |
| 2018 | A Smart Network Interface Approach for Distributed Applications on Xilinx Zynq SoCsabstractNetworked embedded systems have seen tremendous growth with many more complex critical and non-critical systems exchanging information over networks of various types. At each node, information is processed by the network stack before the application sees the data. Large portions of the stack are in software, resulting in significant and non-deterministic delays. While hybrid compute platforms like the Xilinx Zynq can accelerate processing tasks through offloading to programmable logic, the delays incurred due to connectivity can significantly impact overall application latency. In this paper, we present a smart network interface approach for the Xilinx Zynq platform based on datapath extensions within the otherwise standard Ethernet interface. We show that this approach improves computation offload latency by 24–27% and throughput by 37% for a complex computational kernel. Shanker Shreejith, Ryan A. Cooke, Suhaib A. Fahmy |
FPL | 3 |
| 2018 | Efficient Spectrum Sensing for Aeronautical LDACS Using Low-Power Correlators
Shanker Shreejith, Libin K. Mathew, A. Prasad Vinod 0001, Suhaib A. Fahmy |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | In-network online data analytics with FPGAsabstractThe growth of sensor technology, communication systems and computation have led to vast quantities of data being available for relevant parties to utilise. Applications such as the monitoring and analysis of industrial equipment, smart surveillance, and fraud detection rely on the `real-time' analysis of time sensitive data gathered from distributed sources. A variety of processing tasks, such as filtering, aggregation, machine learning algorithms, or other transformations to be carried out on this data in order to extract value from it. Centralised computation strategies are often deployed in these scenarios, with the majority of the data being forwarded though the network to a datacenter environment, typically due to the lack of required computational or storage resources at the leaves of the network, and data from other sources or historical data being required. This approach has also traditionally been viewed as more scalable, as resources can be augmented through the addition of extra compute hardware and cloud services. Ryan A. Cooke, Suhaib A. Fahmy |
FPL | 2 |
| 2017 | VEGa: A High Performance Vehicular Ethernet Gateway on Hybrid FPGAabstractModern vehicles employ a large amount of distributed computation and require the underlying communication scheme to provide high bandwidth and low latency. Existing communication protocols like Controller Area Network (CAN) and FlexRay do not provide the required bandwidth, paving the way for adoption of Ethernet as the next generation network backbone for in-vehicle systems. Ethernet would co-exist with safety-critical communication on legacy networks, providing a scalable platform for evolving vehicular systems. This requires a high-performance network gateway that can simultaneously handle high bandwidth, low latency, and isolation; features that are not achievable with traditional processor based gateway implementations. We present VEGa, a configurable vehicular Ethernet gateway architecture utilising a hybrid FPGA to closely couple software control on a processor with dedicated switching circuit on the reconfigurable fabric. The fabric implements isolated interface ports and an accelerated routing mechanism, which can be controlled and monitored from software. Further, reconfigurability enables the switching behaviour to be altered at runtime under software control, while the configurable architecture allows easy adaptation to different vehicular architectures using highlevel parameter settings. We demonstrate the architecture on the Xilinx Zynq platform and evaluate the bandwidth, latency, and isolation using extensive tests in hardware. Shanker Shreejith, Philipp Mundhenk, Andreas Ettner, Suhaib A. Fahmy, Sebastian Steinhorst, Martin Lukasiewycz, Samarjit Chakraborty |
IEEE Trans. Computers | 4 |
| 2017 | Multipumping Flexible DSP Blocks for Resource Reduction on Xilinx FPGAsabstractFor complex datapaths, resource sharing can help reduce area consumption. Traditionally, resource sharing is applied when the same resource can be scheduled for different uses in different cycles, often resulting in a longer schedule. Multipumping is a method whereby a resource is clocked at a frequency that is a multiple of the surrounding circuit, thereby offering multiple executions per global clock cycle. This allows a single resource to be shared among multiple uses in the same cycle. This concept maps well to modern field-programmable gate arrays (FPGAs), where hard macro blocks are typically capable of running at higher frequencies than most designs implemented in the logic fabric. While this technique has been demonstrated for static resources, modern digital signal processing (DSP) blocks are flexible, supporting varied operations at runtime. In this paper, we demonstrate multipumping for resource sharing of the flexible DSP48E1 macros in Xilinx FPGAs. We exploit their dynamic programmability to enable resource sharing for the full set of supported DSP block operations, and compare this to multipumping only multipliers and DSP blocks with fixed configurations. The proposed approach saves on average 48% DSP blocks at a cost of 74% more LUTs, effectively saving 30% equivalent LUT area and is feasible for the majority of designs, in which clock frequency is typically below half the maximum supported by the DSP blocks. Bajaj Ronak, Suhaib A. Fahmy |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Security in Automotive Networks: Lightweight Authentication and AuthorizationabstractWith the increasing amount of interconnections between vehicles, the attack surface of internal vehicle networks is rising steeply. Although these networks are shielded against external attacks, they often do not have any internal security to protect against malicious components or adversaries who can breach the network perimeter. To secure the in-vehicle network, all communicating components must be authenticated, and only authorized components should be allowed to send and receive messages. This is achieved through the use of an authentication framework. Cryptography is widely used to authenticate communicating parties and provide secure communication channels (e.g., Internet communication). However, the real-time performance requirements of in-vehicle networks restrict the types of cryptographic algorithms and protocols that may be used. In particular, asymmetric cryptography is computationally infeasible during vehicle operation. In this work, we address the challenges of designing authentication protocols for automotive systems. We present Lightweight Authentication for Secure Automotive Networks (LASAN), a full lifecycle authentication approach. We describe the core LASAN protocols and show how they protect the internal vehicle network while complying with the real-time constraints and low computational resources of this domain. By leveraging the fixed structure of automotive networks, we minimize bandwidth and computation requirements. Unlike previous work, we also explain how this framework can be integrated into all aspects of the automotive product lifecycle, including manufacturing, vehicle maintenance, and software updates. We evaluate LASAN in two different ways: First, we analyze the security properties of the protocols using established protocol verification techniques based on formal methods. Second, we evaluate the timing requirements of LASAN and compare these to other frameworks using a new highly modular discrete event simulator for in-vehicle networks, which we have developed for this evaluation. Philipp Mundhenk, Andrew Paverd, Artur Mrowca, Sebastian Steinhorst, Martin Lukasiewycz, Suhaib A. Fahmy, Samarjit Chakraborty |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2016 | Message from the ASAP 2016 chairsabstractWe welcome you to the 27th IEEE International Conference on Application-specific Systems, Architectures and Processors (ASAP 2016). This year's event takes place in London, United Kingdom on the campus of the Imperial College London. Prior to this year's visit to London, the conference has been held in many places around the globe including Oxford (1986), San Diego (1988), Killarney (1989), Princeton (1990), Barcelona (1991), Berkeley (1992), Venice (1993), San Francisco (1994), Strasbourg (1995), Chicago (1996), Zurich (1997), Boston (2000), San Jose (2002), The Hague (2003), Galveston (2004), Samos (2005), Steamboat Springs (2006), Montreal (2007), Leuven (2008), Boston (2009), Rennes (2010), Santa Monica (2011), Delft (2012), and Washington, D.C (2013), Zurich (2014), and Toronto (2015). Though this is the 27th iteration of ASAP, it is actually the 30 year anniversary of the first conference in Oxford. Suhaib A. Fahmy |
ASAP | 2 |
| 2016 | Throughput oriented FPGA overlays using DSP blocks
Abhishek Kumar Jain, Douglas L. Maskell, Suhaib A. Fahmy |
DATE | 3 |
| 2016 | Accelerated Artificial Neural Networks on FPGA for fault detection in automotive systems
Shanker Shreejith, Bezborah Anshuman, Suhaib A. Fahmy |
DATE | 3 |
| 2016 | DeCO: A DSP Block Based FPGA Accelerator Overlay with Low Overhead InterconnectabstractCoarse-grained FPGA overlay architectures paired with general purpose processors offer a number of advantages for general purpose hardware acceleration because of software-like programmability, fast compilation, application portability, and improved design productivity. However, the area overheads of these overlays, and in particular architectures with island-style interconnect, negate many of these advantages, preventing their use in practical FPGA-based systems. Crucially, the interconnect flexibility provided by these overlay architectures is normally over-provisioned for accelerators based on feed-forward pipelined datapaths, which in many cases have the general shape of inverted cones. We propose DeCO, a cone shaped cluster of FUs utilizing a simple linear interconnect between them. This reduces the area overheads for implementing compute kernels extracted from compute-intensive applications represented as directed acyclic dataflow graphs, while still allowing high data throughput. We perform design space exploration by modeling programmability overhead as a function of overlay design parameters, and compare to the programmability overhead of island-style overlays. We observe 87% savings in LUT requirements using the proposed approach compared to DSP block based island-style overlays. Our experimental evaluation shows that the proposed overlay exhibits an achievable frequency of 395 MHz, close to the DSP theoretical limit on the Xilinx Zynq. We also present an automated tool flow that provides a rapid and vendor-independent mapping of the high level compute kernel code to the proposed overlay. Abhishek Kumar Jain, Xiangwei Li, Pranjul Singhai, Douglas L. Maskell, Suhaib A. Fahmy |
FCCM | 5 |
| 2016 | Initiation Interval Aware Resource Sharing for FPGA DSP BlocksabstractResource sharing attempts to minimise usage of hardware blocks by mapping multiple operations onto same block at the cost of an increase in schedule length and initiation interval (II). Sharing multi-cycle high-throughput DSP blocks using traditional approaches results in significantly high II, determined by structure of dataflow graph of the design, thus limiting achievable throughput. We have developed a resource sharing technique that minimises the number of DSP blocks and schedule length given an II constraint. Bajaj Ronak, Suhaib A. Fahmy |
FCCM | 2 |
| 2016 | Automated Verification Code Generation in HLS Using Software Execution Traces (Abstract Only)abstractImproved quality of results from high level synthesis (HLS) tools has led to their increased adoption. Despite the automated translation from high level descriptions to register-transfer level (RTL) implementations, functional verification remains a major challenge. Verification can take significantly more time than the design process; if there is a functional mismatch, developers must back-trace thousands of signals and cycles to determine underlying cause. The challenge is further exacerbated with HLS-produced RTL, which is often not human readable. To overcome these challenges, we present a verification technique that uses software-execution traces and automated insertion of verification code into the HLS-generated RTL to assist in debugging. The verification code helps pinpoint the earliest instance of RTL simulation mismatch, either caused by HLS engine bugs or design bugs, and related instructions. We also integrate a watchdog timer to examine the execution of control-flow and perform source-to-source transformation on benchmarks to take advantage of our proposed instrumentation. We also create a framework to insert various types of bugs, e.g. data-flow, control-flow and operational bugs, to evaluate our technique. We use the CHStone benchmark suite and demonstrate that our verification detects over 90% of the inserted bugs, with over 70% of them detected within 10 cycles. In addition, the proposed flow can detect real-life bugs existing in previously released versions of CHStone suite as well. Swathi T. Gurumani, Suhaib A. Fahmy, Deming Chen, Kyle Rupnow |
FPGA | 3 |
| 2016 | Designing a virtual runtime for FPGA accelerators in the cloudabstractFPGAs can provide high performance and energy efficiency to many applications; therefore, they are attractive computing platforms in a cloud environment. However, FPGA application development requires extensive hardware design knowledge which significantly limits the potential user base. Moreover, in a cloud setting, allocating a whole FPGA to a user is often wasteful and not cost effective due to low device utilization. To make FPGA application development easier, firstly, we propose a methodology that provides clean abstractions with high-level APIs and a simple execution model that supports both software and hardware execution. Secondly, to improve device utilization and share the FPGA among multiple users, we developed a lightweight runtime system that provides hardware-assisted memory virtualization and memory protection, enabling multiple applications to simultaneously execute on the device. Mikhail Asiatici, Nithin George, Kizheppatt Vipin, Suhaib A. Fahmy, Paolo Ienne |
FPL | 4 |
| 2016 | Improved resource sharing for FPGA DSP blocksabstractSharing multi-cycle hardware blocks like the DSP48E1 primitive in Xilinx FPGAs can result in significant resource savings, but complicates scheduling. For high-throughput, DSP blocks must be pipelined, which results in a high initiation interval (II) for resource shared implementations. In this paper, we propose a resource reduction technique that minimises DSP block usage while also offering improved II over traditional approaches. This is integrated in a high-level tool which takes datapath descriptions in C and generates synthesisable Verilog RTL with different levels of resource sharing. We demonstrate significantly improved throughput compared to traditional resource sharing while achieving resource reduction compared to resource unconstrained and HLS implementations. The approach explores an otherwise infeasible design space between resource unconstrained and traditional resource sharing methods. Bajaj Ronak, Suhaib A. Fahmy |
FPL | 2 |
| 2016 | JetStream: An open-source high-performance PCI Express 3 streaming library for FPGA-to-Host and FPGA-to-FPGA communicationabstractMany FPGA-based accelerators are constrained by the available resources and multi-FPGA solutions can be necessary for building more capable systems. Available PCIe solutions provide only FPGA-to-Host communication. In this paper we present JetStream, an open-source1modular PCIe 3 library, supporting not only fast FPGA-to-Host communication, but also allowing direct FPGA-to-FPGA communication which fully bypasses the memory subsystem. The direct mode saves memory bandwidth for multicast modes and permits to connect multiple FPGAs in various software defined topologies. We show the benefits of JetStream with a large FIR filter spanning four FPGA boards, achieving throughputs of up to 7.09 GB/s per link. Utilizing direct FPGA-to-FPGA transfers reduces the required memory bandwidth by up to 75%. Malte Vesper, Dirk Koch, Kizheppatt Vipin, Suhaib A. Fahmy |
FPL | 4 |
| 2016 | Mapping for Maximum Performance on FPGA DSP BlocksabstractThe digital signal processing (DSP) blocks on modern field programmable gate arrays (FPGAs) are highly capable and support a variety of different datapath configurations. Unfortunately, inference in synthesis tools can fail to result in circuits that reach maximum DSP block throughput. We have developed a tool that maps graphs of add/sub/mult nodes to DSP blocks on Xilinx FPGAs, ensuring maximum throughput. This is done by delaying scheduling until after the graph has been partitioned onto DSP blocks and scheduled based on their pipeline structure, resulting in a throughput optimized implementation. Our tool prepares equivalent implementations in a variety of other methods, including high-level synthesis (HLS) for comparison. We show that the proposed approach offers an improvement in frequency of 100% over standard pipelined code, and 23% over Vivado HLS synthesis implementation, while retaining code portability, at the cost of a modest increase in logic resource usage. Bajaj Ronak, Suhaib A. Fahmy |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Efficient Integer Frequency Offset Estimation Architecture for Enhanced OFDM SynchronizationabstractIn orthogonal frequency-division multiplexing (OFDM) systems, integer frequency offset (IFO) causes a circular shift of the subcarrier indices in the frequency domain. The IFO can be mitigated through strict RF front-end design, which tends to be expensive, or by strictly limiting mobility and channel agility, which constrains operating scenarios. The IFO is, therefore, often estimated and removed at baseband, allowing implementations to benefit from the relaxed RF front-end specifications and to be tolerant to both Doppler shift and multistandard channel selection. This paper proposes a novel architecture for the IFO estimation which achieves reduced power consumption and lower computational cost than contemporary methods, while achieving excellent estimation performance, close to theoretically achievable bounds. A pilot subsampling technique enables fourfold resource sharing to reduce the computational cost, while multiplierless computation yields further power reduction. Performance exceeds that of the conventional techniques, while being much more efficient. When implemented on field-programmable gate array for IEEE 802.16-2009, the dynamic power reductions of 78% are achieved. The architecture and method is applicable to other OFDM standards including IEEE 802.11 and IEEE 802.22. Thinh Hung Pham, Suhaib A. Fahmy, Ian McLoughlin 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Virtualized FPGA Accelerators for Efficient Cloud ComputingabstractHardware accelerators implement custom architectures to significantly speed up computations in a wide range of domains. As performance scaling in server-class CPUs slows, we propose the integration of hardware accelerators in the cloud as a way to maintain a positive performance trend. Field programmable gate arrays (FPGAs) represent the ideal way to integrate accelerators in the cloud, since they can be reprogrammed as needs change and allow multiple accelerators to share optimised communication infrastructure. We discuss a framework that integrates reconfigurable accelerators in a standard server with virtualised resource management and communication. We then present a case study that quantifies the efficiency benefits and break-even point for integrating FPGAs in the cloud. Suhaib A. Fahmy, Kizheppatt Vipin, Shanker Shreejith |
CloudCom | 1 |
| 2015 | Security analysis of automotive architectures using probabilistic model checkingabstractThis paper proposes a novel approach to security analysis of automotive architectures at the system-level. With an increasing amount of software and connectedness of cars, security challenges are emerging in the automotive domain. Our proposed approach enables assessment of the security of architecture variants and can be used by decision makers in the design process. First, the automotive Electronic Control Units (ECUs) and networks are modelled at the system-level using parameters per component, including an exploitability score and patching rates that are derived from an automated or manual assessment. For any specific architecture variant, a Continuous-Time Markov Chain (CTMC) model is determined and analyzed in terms of confidentiality, integrity and availability, using probabilistic model checking. The introduced case study demonstrates the applicability of our approach, enabling, for instance, the exploration of parameters like patch rate targets for ECU manufacturers. Philipp Mundhenk, Sebastian Steinhorst, Martin Lukasiewycz, Suhaib A. Fahmy, Samarjit Chakraborty |
DAC | 4 |
| 2015 | Security aware network controllers for next generation automotive embedded systemsabstractModern cars incorporate complex distributed computing systems that manage all aspects of vehicle dynamics, comfort, and safety. Increased automation has demanded more complex networking in vehicles, that now contain a hundred or more compute units. As these networks were developed as silos, little attention was given to security early on. However, this has become a key challenge in the automotive domain, as these systems have been shown to be susceptible to various attacks, with sometimes catastrophic consequences. Addressing security in such systems requires consideration of the network and compute units, both hardware and software, and complex real-time constraints. We present an approach that integrates application authentication, message encryption and network access control into a smart network interface, without compromising network determinism. A custom interface with partial reconfiguration support on FPGAs enables seamless integration of security at the interface, offering a level of security not possible with standard layered approaches. Shanker Shreejith, Suhaib A. Fahmy |
DAC | 2 |
| 2015 | Lightweight authentication for secure automotive networks
Philipp Mundhenk, Sebastian Steinhorst, Martin Lukasiewycz, Suhaib A. Fahmy, Samarjit Chakraborty |
DATE | 4 |
| 2015 | Efficient Overlay Architecture Based on DSP BlocksabstractDesign productivity and long compilation times are major issues preventing the mainstream adoption of FPGAs in general purpose computing. Several overlay architectures have emerged to tackle these challenges, but at the cost of increased area and performance overheads. This paper examines a coarse grained overlay architecture designed using the flexible DSP48E1 primitive on Xilinx FPGAs. This allows pipelined execution at significantly higher throughput without adding significant area overheads to the PE. We map several benchmarks, using our custom mapping tool, and show that the proposed overlay architecture delivers a throughput of up to 21.6 GOPS and provides an 11 -- 52% improvement in throughput compared to Vivado HLS implementations. Abhishek Kumar Jain, Suhaib A. Fahmy, Douglas L. Maskell |
FCCM | 2 |
| 2015 | On Data Forwarding in Deeply Pipelined Soft ProcessorsabstractWe can design high-frequency soft-processors on FPGAs that exploit deep pipelining of DSP primitives, supported by selective data forwarding, to deliver up to 25% performance improvements across a range of benchmarks. Pipelined, in-order, scalar processors can be small and lightweight but suffer from a large number of idle cycles due to dependency chains in the instruction sequence. Data forwarding allows us to more deeply pipeline the processor stages while avoiding an associated increase in the NOP cycles between dependent instructions. Full forwarding can be prohibitively complex for a lean soft processor, so we explore two approaches: an external forwarding path around the DSP block execution unit in FPGA logic and using the intrinsic loopback path within the DSP block primitive. We show that internal loopback improves performance by 5% compared to external forwarding, and up to 25% over no data forwarding. The result is a processor that runs at a frequency close to the fabric limit of 500 MHz, but without the significant dependency overheads typical of such processors. Hui Yan Cheah, Suhaib A. Fahmy, Nachiket Kapre |
FPGA | 2 |
| 2015 | Minimizing DSP block usage through multi-pumpingabstractResource sharing in the mapping of an algorithm to an architecture allows the same resource to be scheduled for different uses in different cycles, generally at the cost of increased schedule length. Multi-pumping is a method whereby a resource is clocked at a frequency that is a multiple of the surrounding circuit, thereby offering multiple executions per global clock, and therefore sharing in the same clock cycle. This concept maps well to FPGA architectures, where hard macro blocks are typically capable of running at higher frequencies than standard logic. While this technique has been demonstrated for multipliers, modern DSP blocks are more complex with multiple computational nodes. In this paper, we apply multi-pumping to minimise DSP block usage, while taking advantage of the multiple nodes they support. The proposed approach uses, on average, 39% fewer DSP blocks, at a cost of 19% more LUTs and 7% more registers. Bajaj Ronak, Suhaib A. Fahmy |
FPT | 2 |
| 2015 | JIT trace-based verification for high-level synthesisabstractHigh level synthesis (HLS) tools are increasingly adopted for hardware design as the quality of tools consistently improves. Concerted development effort on HLS tools represents significant software development effort, and debugging and validation represents a significant portion of that effort. However, HLS tools are different from typical large-scale software systems; HLS tool output must be subsequently verified through functional verification of the generated RTL implementation. Debugging machine-generated functionally incorrect RTL is time-consuming and cumbersome requiring back-tracing through hundreds of signals and simulation cycles to determine the underlying error. This challenging process requires support framework in the HLS flow to enable fast and efficient pinpointing of the incorrectness in the tool. In this paper, we present a debug framework that uses just-in-time (JIT) traces and automated insertion of verification code into the generated RTL to assist in debugging an HLS tool. This framework aids the user by quickly pinpointing the earliest instance of execution mismatch, paired with detailed information on the faulty signal, and the corresponding instruction from the application source. Using CHStone benchmarks, we demonstrate that this technique can significantly reduce bug detection latency: often with zero cycle detection. Magzhan Ikram, Swathi T. Gurumani, Suhaib A. Fahmy, Deming Chen, Kyle Rupnow |
FPT | 4 |
| 2014 | A scalable and compact systolic architecture for linear solversabstractWe present a scalable design for accelerating the problem of solving a dense linear system of equations using LU Decomposition. A novel systolic array architecture that can be used as a building block in scientific applications is described and prototyped on a Xilinx Virtex 6 FPGA. This solver has a throughput of around 3.2 million linear systems per second for matrices of size N=4 and around 80 thousand linear systems per second for matrices of size N=16. In comparison with similar work, our design offers up to a 12-fold improvement in speed whilst requiring up to 50% less hardware resources. As a result, a linear system of size N=64 can be implemented on a single FPGA, whereas previous work was limited to a size of N=12 and resorted to complex multi-FPGA architectures to scale. Finally, the scalable design can be adapted to different sized problems with minimum effort. Kevin Shen-Hoong Ong, Suhaib A. Fahmy, Keck Voon Ling |
ASAP | 2 |
| 2014 | Experiments in Mapping Expressions to DSP BlocksabstractMapping complex mathematical expressions to DSP blocks by relying on synthesis from pipelined code is inefficient and results in significantly reduced throughput. We have developed a tool to demonstrate the benefit of considering the structure and pipeline arrangement of the DSP block in mapping of functions. Implementations where the structure of the DSP block is considered during pipelining achieve double the throughput of other methods, demonstrating that the structure of the DSP block must be considered when scheduling complex expressions. Bajaj Ronak, Suhaib A. Fahmy |
FCCM | 2 |
| 2014 | Automated Partial Reconfiguration Design for Adaptive Systems with CoPR for ZynqabstractDynamically adaptive systems (DAS) respond to environmental conditions, by modifying their processing at runtime and selecting alternative configurations of computation. Field programmable gate arrays, with their support for partial reconfiguration (PR) represent an ideal platform for implementing such systems. Designing partially reconfigurable systems has traditionally been a difficult task requiring FPGA expertise. This paper presents a fully automated framework for implementing PR based adaptive systems. The designer specifies a set of valid configurations containing instances of modules from a standard library. The tool automates partitioning of modules into regions, floorplanning regions on the FPGA fabric, and generation of bitstreams. A runtime system manages the loading of bitstreams automatically through API calls. Kizheppatt Vipin, Suhaib A. Fahmy |
FCCM | 2 |
| 2014 | Square-rich fixed point polynomial evaluation on FPGAsabstractPolynomial evaluation is important across a wide range of application domains, so significant work has been done on accelerating its computation. The conventional algorithm, referred to as Horner's rule, involves the least number of steps but can lead to increased latency due to serial computation. Parallel evaluation algorithms such as Estrin's method have shorter latency than Horner's rule, but achieve this at the expense of large hardware overhead. This paper presents an efficient polynomial evaluation algorithm, which reforms the evaluation process to include an increased number of squaring steps. By using a squarer design that is more efficient than general multiplication, this can result in polynomial evaluation with a 57.9% latency reduction over Horner's rule and 14.6% over Estrin's method, while consuming less area than Horner's rule, when implemented on a Xilinx Virtex 6 FPGA. When applied in fixed point function evaluation, where precision requirements limit the rounding of operands, it still achieves a 52.4% performance gain compared to Horner's rule with only a 4% area overhead in evaluating 5th degree polynomials. Simin Xu, Suhaib A. Fahmy, Ian McLoughlin 0001 |
FPGA | 2 |
| 2014 | Efficient multi-standard cognitive radios on FPGAsabstractCognitive radios that support multiple standards and modify operation depending on environmental conditions are becoming more important as the demand for higher bandwidth and efficient spectrum use increases. Traditional implementations in custom ASICs cannot support such flexibility, with standards changing at a faster pace, while software baseband implementations fail to achieve the performance required. Hence, FPGAs offer an ideal platform bringing together flexibility, performance, and efficiency. This work explores the possible techniques for designing multi-standard radios on FPGAs, and explores how partial reconfiguration can be leveraged in a way that is amenable for domain experts with minimal FPGA knowledge. Thinh Hung Pham, Suhaib A. Fahmy, Ian McLoughlin 0001 |
FPL | 2 |
| 2014 | Efficient mapping of mathematical expressions into DSP blocksabstractMapping complex mathematical expressions to DSP blocks through standard inference from pipelined code is inefficient and results in significantly reduced throughput. In this paper, we demonstrate the benefit of considering the structure and pipeline arrangement of DSP blocks during mapping. We have developed a tool that can map mathematical expressions using RTL inference, through high level synthesis with Vivado HLS, and through a custom approach that incorporates DSP block structure. We can show that the proposed method results in circuits that run at around double the frequency of other methods, demonstrating that the structure of the DSP block must be considered when scheduling complex expressions. Bajaj Ronak, Suhaib A. Fahmy |
FPL | 2 |
| 2014 | DyRACT: A partial reconfiguration enabled accelerator and test platformabstractIntegrating FPGAs with a general purpose computer remains difficult, but recent efforts have resulted in open frameworks that offer a software API and hardware interface to allow easier integration. However, such systems only support static FPGA designs. With the addition of partial reconfiguration (PR) support, such frameworks can enable more effective use of FPGAs. Now, designers can incorporate hardware accelerators within their software applications, and these can be loaded dynamically as required. We present a PR-enabled FPGA platform that allows user modules to be loaded onto the FPGA, inputs to be applied, results obtained, and functions to be swapped at runtime. The interface and PR management logic are part of the static region, while multiple accelerators can be loaded using high level functions provided by the API. Reconfiguration and data transfer are both managed over the PCIe interface from the host PC, with communication throughput of more than 1.5 GB/s (75% of peak PCIe bandwidth) and reconfiguration of a large accelerator in 20 ms. Kizheppatt Vipin, Suhaib A. Fahmy |
FPL | 2 |
| 2014 | Analysis and optimization of a deeply pipelined FPGA soft processorabstractFPGA soft processors have been shown to achieve high frequency when designed around the specific capabilities of heterogenous resources on modern FPGAs. However, such performance comes at a cost of deep pipelines, which can result in a larger number of idle cycles when executing programs with long dependency chains in the instruction sequence. We perform a full design-space exploration of a DSP block based soft processor to examine the effect of pipeline depth on frequency, area, and program runtime, noting the significant number of NOPs required to resolve dependencies. We then explore the potential of a restricted data forwarding approach in improving runtime by significantly reducing NOP padding. The result is a processor that runs close to the fabric limit of 500MHz with a case for simple data forwarding. Hui Yan Cheah, Suhaib A. Fahmy, Nachiket Kapre |
FPT | 2 |
| 2014 | Zero latency encryption with FPGAs for secure time-triggered automotive networksabstractSecurity has emerged as a key concern in increasingly complex embedded automotive networks. The distributed architecture and broadcast transmission characteristics mean they are vulnerable and provide little resistance to intrusive and non-intrusive attack mechanisms. Incorporating data security using traditional approaches introduces significant latency which can be problematic in the presence of real-time deadlines. We demonstrate how a security layer can be added within the network communication controller in modern time-triggered systems, without introducing additional latency or processing overheads. This allows critical communications to be secured in a manner that is transparent to the processors in the electronic control units (ECUs), while also safeguarding network communication properties. Shanker Shreejith, Suhaib A. Fahmy |
FPT | 2 |
| 2014 | A case for leveraging 802.11p for direct phone-to-phone communicationsabstractWiFi cannot effectively handle the demands of device-to-device communication between phones, due to insufficient range and poor reliability. We make the case for using IEEE 802.11p DSRC instead, which has been adopted for vehicle-to-vehicle communications, providing lower latency and longer range. We demonstrate a prototype motivated by a novel fabrication process that deposits both III-V and CMOS devices on the same die. In our system prototype, the designed RF front-end is interfaced with a baseband processor on an FPGA, connected to Android phones. It consumes 0.02uJ/bit across 100m assuming free space. Application-level power control dramatically reduces power consumption by 47-56%. Pilsoon Choi, Jason H. Gao 0001, Nadesh Ramanathan, Mengda Mao, Shipeng Xu, Chirn Chye Boon, Suhaib A. Fahmy, Li-Shiuan Peh |
ISLPED | 7 |
| 2014 | Shaping Spectral Leakage for IEEE 802.11p Vehicular CommunicationsabstractIEEE 802.11p is a recently defined standard for the physical (PHY) and medium access control (MAC) layers for Dedicated Short-Range Communications. Four Spectrum Emission Masks (SEMs) are specified in 802.11p that are much more stringent than those for current 802.11 systems. In addition, the guard interval in 802.11p has been lengthened by reducing the bandwidth to support vehicular communication (VC) channels, and this results in a narrowing of the frequency guard. This raises a significant challenge for filtering the spectrum of 802.11p signals to meet the specifications of the SEMs. We investigate state of the art pulse shaping and filtering techniques for 802.11p, before proposing a new method of shaping the 802.11p spectral leakage to meet the most stringent, class D, SEM specification. The proposed method, performed at baseband to relax the strict constraints of the radio frequency (RF) front-end, allows 802.11p systems to be implemented using commercial off-the-shelf (COTS) 802.11a RF hardware, resulting in reduced total system cost. Thinh Hung Pham, Ian McLoughlin 0001, Suhaib A. Fahmy |
VTC Spring | 3 |
| 2014 | The iDEA DSP Block-Based Soft Processor for FPGAsabstractDSP blocks in modern FPGAs can be used for a wide range of arithmetic functions, offering increased performance while saving logic resources for other uses. They have evolved to better support a plethora of signal processing tasks, meaning that in other application domains they may be underutilised. The DSP48E1 primitives in new Xilinx devices support dynamic programmability that can help extend their usefulness; the specific function of a DSP block can be modified on a cycle-by-cycle basis. However, the standard synthesis flow does not leverage this flexibility in the vast majority of cases. The lean DSP Extension Architecture (iDEA) presented in this article builds around the dynamic programmability of a single DSP48E1 primitive, with minimal additional logic to create a general-purpose processor supporting a full instruction-set architecture. The result is a very compact, fast processor that can execute a full gamut of general machine instructions. We show a number of simple applications compiled using an MIPS compiler and translated to the iDEA instruction set, comparing with a Xilinx MicroBlaze to show estimated performance figures. Being based on the DSP48E1, this processor can be deployed across next-generation Xilinx Artix-7, Kintex-7, Virtex-7, and Zynq families. Hui Yan Cheah, Fredrik Brosser, Suhaib A. Fahmy, Douglas L. Maskell |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2013 | Microkernel hypervisor for a hybrid ARM-FPGA platformabstractReconfigurable architectures have found use in a wide range of application domains, but mostly as static accelerators for computationally intensive functions. Commodity computing adoption has not taken off due primarily to design complexity challenges. Yet reconfigurable architectures offer significant advantages in terms of sharing hardware between distinct isolated tasks, under tight time constraints. Trends towards amalgamation of computing resources in the automotive and aviation domains have so far been limited to non-critical systems, because processor approaches suffer from a lack of predictability and isolation. Hybrid reconfigurable platforms may provide a promising solution to this, by allowing physically isolated access to hardware resources, and support for computationally demanding applications, but with improved programmability and management. We propose virtualized execution and management of software and hardware tasks using a microkernel-based hypervisor running on a commercial hybrid computing platform (the Xilinx Zynq). We demonstrate a framework based on the CODEZERO hypervisor, which has been modified to leverage the capabilities of the FPGA fabric. It supports discrete hardware accelerators, dynamically reconfigurable regions, and regions of virtual fabric, allowing for application isolation and simpler use of hardware resources. A case study demonstrating multiple independent (and isolated) software and hardware tasks is presented. Khoa Dang Pham, Abhishek Kumar Jain, Suhaib A. Fahmy, Douglas L. Maskell |
ASAP | 4 |
| 2013 | System architecture and software design for electric vehiclesabstractThis paper gives an overview of the system architecture and software design challenges for Electric Vehicles (EVs). First, we introduce the EV-specific components and their control, considering the battery, electric motor, and electric powertrain. Moreover, technologies that will help to advance safety and energy efficiency of EVs such as drive-by-wire and information systems are discussed. Regarding the system architecture, we present challenges in the domain of communication and computation platforms. A paradigm shift towards time-triggered in-vehicle communication systems becomes inevitable for the sake of determinism, making the introduction of new bus systems and protocols necessary. At the same time, novel computational devices promise high processing power at low cost which will make a reduction in the number of Electronic Control Units (ECUs) possible. As a result, the software design has to be performed in a holistic manner, considering the controlled component while transparently abstracting the underlying hardware architecture. For this purpose, we show how middleware and verification techniques can help to reduce the design and test complexity. At the same time, with the growing connectivity of EVs, security has to become a major design objective, considering possible threats and a security-aware design as discussed in this paper. Martin Lukasiewycz, Sebastian Steinhorst, Sidharta Andalam, Florian Sagstetter, Peter Waszecki, Wanli Chang 0001, Matthias Kauer, Philipp Mundhenk, Shanker Shreejith, Suhaib A. Fahmy, Samarjit Chakraborty |
DAC | 10 |
| 2013 | An approach for redundancy in FlexRay networks using FPGA partial reconfigurationabstractSafety-critical in-vehicle electronic control units (ECUs) demand high levels of determinism and isolation, since they directly influence vehicle behaviour and passenger safety. As modern vehicles incorporate more complex computational systems, ensuring the safety of critical systems becomes paramount. One-to-one redundant units have been previously proposed as measures for evolving critical functions like x-by-wire. However, these may not be viable solutions for power-constrained systems like next generation electric vehicles. Reconfigurable architectures offer alternative approaches to implementing reliable safety critical systems using more efficient hardware. In this paper, we present an approach for implementing redundancy in safety-critical in-car systems, that uses FPGA partial reconfiguration and a customised bus controller to offer fast recovery from faults. Results show that such an integrated design is better than alternatives that use discrete bus interface modules. Shanker Shreejith, Kizheppatt Vipin, Suhaib A. Fahmy, Martin Lukasiewycz |
DATE | 3 |
| 2013 | An Approach to a Fully Automated Partial Reconfiguration Design FlowabstractAdoption of partial reconfiguration (PR) in mainstream FPGA system design remains underwhelming primarily due the significant FPGA design expertise that is required. We present an approach to fully automating a design flow that accepts a high level description of a dynamically adaptive application and generates a fully functional, optimised PR design. This tool can determine the most suitable FPGA for a design to meet a given reconfiguration time constraint and makes full use of available resources. The flow targets adaptive systems, where the dynamic behaviour and switching order are not known up front. Kizheppatt Vipin, Suhaib A. Fahmy |
FCCM | 2 |
| 2013 | Efficient Large Integer Squarers on FPGAabstractThis paper presents an optimised high throughput architecture for integer squaring on FPGAs. The approach reduces the number of DSP blocks required compared to a standard multiplier. Previous work has proposed the tiling method for double precision squaring, using the least number of DSP blocks so far. However that approach incurs a large overhead in terms of look-up table (LUT) consumption and has a complex and irregular structure that is not suitable for higher word size. The architecture proposed in this paper can reduce DSP block usage by an equivalent amount to the tiling method while incurring a much lower LUT overhead: 21.8% fewer LUTs for a 53-bit squarer. The architecture is mapped to a Xilinx Virtex 6 FPGA and evaluated for a wide range of operand word sizes, demonstrating its scalability and efficiency. Simin Xu, Suhaib A. Fahmy, Ian McLoughlin 0001 |
FCCM | 2 |
| 2013 | Iterative floating point computation using FPGA DSP blocksabstractThis paper presents a single precision floating point unit design for multiplication and addition/subtraction using FPGA DSP blocks. The design is based around the DSP48E1 primitive found in Virtex-6 and all 7-series FPGAs from Xilinx. Since the DSP48E1 can be dynamically configured and used for many of the sub-operations involved in IEEE 754-2008 binary32 floating point multiplication and addition, we demonstrate an iterative combined operator that uses a single DSP block and minimal logic. Logic-only and fixed-configuration DSP block designs, and other state-of-the-art implementations, including the Xilinx CoreGen operators are compared to this approach. Since FPGA based systems typically run at a fraction of the maximum possible FPGA speed, and in some cases, floating point computations may not be required in every cycle, the iterative approach represents an efficient way to leverage DSP resources for what can otherwise be costly operations. Fredrik Brosser, Hui Yan Cheah, Suhaib A. Fahmy |
FPL | 3 |
| 2013 | Enhancing communication on automotive networks using data layer extensionsabstractAutomotive systems function within a distributed computing paradigm consisting of networks of sensors, actuators and processing units. More advanced functions are finding their way into the automotive domain, making computing and networking more complex. While safety, security and determinism are primary concerns for many systems, the networking protocols do not provide extensions to address them and it is left to application designers to tackle these issues. Standardising simple features like time-stamping of messages and health status flags can help improve robustness and mitigate risks associated with replay attacks. It is also possible to integrate further protection, like encryption of messages, addressing the increasing security concerns in this domain. In this paper, we demonstrate a systematic way of accommodating such enhancements within a standard automotive network that retains interoperability with existing systems. We show how such enhancements can be made possible in both software and hardware to help add functionality above the core network specification. We also show that the enhancements incorporated at the hardware layer offers 20× better performance than the software-based approach. Shanker Shreejith, Suhaib A. Fahmy |
FPT | 2 |
| 2013 | Accelerating validation of time-triggered automotive systems on FPGAsabstractAutomotive systems comprise a high number of networked safety-critical functions. Any design changes or addition of new functionality must be rigorously tested to ensure that no performance or safety issues are introduced, and this consumes a significant amount of time. Validation should be conducted using a faithful representation of the system, and so typically, a full subsystem is built for validation. We present a scalable scheme for emulating a complete cluster of automotive embedded compute units on an FPGA, with accelerated network communication using custom physical level interfaces. With these interfaces, we can achieve acceleration of system emulation by 8× or more, with a systematic way of exploring real-world issues like jitter, network delays, and data corruption, among others. By using the same communication infrastructure as in a real deployed system, this validation is closer to the requirements of standards compliance. This approach also enables hardware-in-the-loop (HIL) validation, allowing rapid prototyping of distributed functions, including changes in network topology and parameters, and modification of time-triggered schedules without physical hardware modification. We present an implementation of this framework on the Xilinx ML605 evaluation board that integrates six FlexRay automotive functions to demonstrate the potential of the framework. Shanker Shreejith, Suhaib A. Fahmy, Martin Lukasiewycz |
FPT | 2 |
| 2013 | System-level FPGA device driver with high-level synthesis supportabstractWe can exploit the standardization of communication abstractions provided by modern high-level synthesis tools like Vivado HLS, Bluespec and SCORE to provide stable system interfaces between the host and PCIe-based FPGA accelerator platforms. At a high level, our FPGA driver attempts to provide CUDA-like driver behavior, and more, to FPGA programmers. On the FPGA fabric, we develop an AXI-compliant, lightweight interface switch coupled to multiple physical interfaces (PCIe, Ethernet, DRAM) to provide programmable, portable routing capability between the host and user logic on the FPGA. On the host, we adapt the RIFFA 1.0 driver to provide enhanced communication APIs along with bitstream configuration capability allowing low-latency, high-throughput communication and safe, reliable programming of user logic on the FPGA. Our driver only consumes 21% BRAMs and 14% logic overhead on a Xilinx ML605 platform or 9% BRAMs and 8% logic overhead on a Xilinx V707 board. We are able to sustain DMA transfer throughput (to DRAM) of 1.47GB/s (74% peak) of the PCIe (x4 Gen2) bandwidth, 120.2MB/s (96%) of the Ethernet (1G) bandwidth and 5.93GB/s (92.5%) of DRAM bandwidth. Kizheppatt Vipin, Shanker Shreejith, Dulitha Gunasekera, Suhaib A. Fahmy, Nachiket Kapre |
FPT | 4 |
| 2013 | Optimization of the HEFT Algorithm for a CPU-GPU EnvironmentabstractScheduling applications efficiently on a network of computing systems is crucial for high performance. This problem is known to be NP-Hard and is further complicated when applied to a CPU-GPU heterogeneous environment. Heuristic algorithms like Heterogeneous Earliest Finish Time (HEFT) have shown to produce good results for other heterogeneous environments like Grids and Clusters. In this paper, we propose a novel optimization of this algorithm that takes advantage of dissimilar execution times of the processors in the chosen environment. We optimize both the task ranking as well as the processor selection steps of the HEFT algorithm. By balancing the locally optimal result with the globally optimal result, we show that performance can be improved significantly without any change in the complexity of the algorithm (as compared to HEFT). Using randomly generated Directed A cyclic Graphs (DAGs), the new algorithm HEFT-NC (No-Cross) is compared with HEFT both in terms of speedup and schedule length. We show that the HEFT-NC outperforms HEFT algorithm and is consistent across different graph shapes and task sizes. Karan R. Shetti, Suhaib A. Fahmy, Timo Bretschneider |
PDCAT | 2 |
| 2013 | Architecture for Real-Time Nonparametric Probability Density Function EstimationabstractAdaptive systems are increasing in importance across a range of application domains. They rely on the ability to respond to environmental conditions, and hence real-time monitoring of statistics is a key enabler for such systems. Probability density function (PDF) estimation has been applied in numerous domains; computational limitations, however, have meant that proxies are often used. Parametric estimators attempt to approximate PDFs based on fitting data to an expected underlying distribution, but this is not always ideal. The density function can be estimated by rescaling a histogram of sampled data, but this requires many samples for a smooth curve. Kernel-based density estimation can provide a smoother curve from fewer data samples. We present a general architecture for nonparametric PDF estimation, using both histogram-based and kernel-based methods, which is designed for integration into streaming applications on field-programmable gate array (FPGAs). The architecture employs heterogeneous resources available on modern FPGAs within a highly parallelized and pipelined design, and is able to perform real-time computation on sampled data at speeds of over 250 million samples per second, while extracting a variety of statistical properties. Suhaib A. Fahmy, A. R. Mohan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Low-Power Correlation for IEEE 802.16 OFDM Synchronization on FPGAabstractThis brief compares the use of multiplierless and DSP slice-based cross-correlation for IEEE 802.16d orthogonal frequency division multiplexing (OFDM) timing synchronization on Xilinx Virtex-6 and Spartan-6 field programmable gate arrays (FPGAs). The natural approach, given the availability of embedded DSP blocks on these FPGAs, would be to implement standard multiplier-based cross-correlation. However, this can consume a significant number of DSP blocks, which may not fit on low-power devices. Hence, we compare a DSP48E1 slice-based design to four different quantizations of multiplierless correlation in terms of resource utilization and power consumption. OFDM timing synchronization accuracy is evaluated for each system at different signal-to-noise ratios. Results show that even relatively coarse multiplierless coefficient quantization can yield accurate timing synchronization, and does so at high clock speeds. Multiplierless designs enjoy reduced power consumption over the DSP48E1 Slice-based design, and can be used where DSP Slice resources are insufficient, such as on low-power FPGA devices. Thinh Hung Pham, Suhaib A. Fahmy, Ian McLoughlin 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Embedded systems and software challenges in electric vehiclesabstractThe design of electric vehicles require a complete paradigm shift in terms of embedded systems architectures and software design techniques that are followed within the conventional automotive systems domain. It is increasingly being realized that the evolutionary approach of replacing the engine of a car by an electric engine will not be able to address issues like acceptable vehicle range, battery lifetime performance, battery management techniques, costs and weight, which are the core issues for the success of electric vehicles. While battery technology has crucial importance in the domain of electric vehicles, how these batteries are used and managed pose new problems in the area of embedded systems architecture and software for electric vehicles. At the same time, the communication and computation design challenges in electric vehicles also have to be addressed appropriately. This paper discusses some of these research challenges. Samarjit Chakraborty, Martin Lukasiewycz, Christian Buckl, Suhaib A. Fahmy, Naehyuck Chang, Sangyoung Park, Younghyun Kim 0001, Patrick Leteinturier, Hans Adlkofer |
DATE | 4 |
| 2012 | A lean FPGA soft processor built using a DSP blockabstractAs Field Programmable Gate Arrays (FPGAs) have advanced, the capabilities and variety of embedded resources have increased. In the last decade, signal processing has become one of the main driving applications for FPGA adoption, so FPGA vendors tailored their architectures to such applications. The resulting embedded digital signal processing (DSP) blocks have now advanced to the point of supporting a wide range of operations. In this paper, we explore how these DSP blocks can be applied to general computation. We show that the DSP48E1 blocks in Xilinx Virtex-6 devices support a wide range of standard processor instructions which can be designed into the core of a basic processor with minimal additional logic usage. Hui Yan Cheah, Suhaib A. Fahmy, Douglas L. Maskell, Chidamber Kulkarni |
FPGA | 2 |
| 2012 | Evaluating the efficiency of DSP Block synthesis inference from flow graphsabstractThe embedded DSP Blocks in FPGAs have become significantly more capable in recent generations of devices. While vendor synthesis tools can infer the use of these resources, the efficiency of this inference is not guaranteed. Specific language structures are suggested for implementing standard applications but others that do not fit these standard designs can suffer from inefficient synthesis inference. In this paper, we demonstrate this effect by synthesising a number of arithmetic circuits, showing that standard code results in a significant resource and timing overhead compared to considered use of DSP Blocks and their plethora of configuration options through custom instantiation. Bajaj Ronak, Suhaib A. Fahmy |
FPL | 2 |
| 2012 | iDEA: A DSP block based FPGA soft processorabstractThis paper presents a very lean DSP Extension Architecture (iDEA) soft processor for Field Programmable Gate Arrays (FPGAs). iDEA has been built to be as lightweight as possible, utilising the run-time flexibility of the DSP48E1 primitive in Xilinx FPGAs to serve as many processor functions as possible. We show how the primitive's flexibility can be leveraged within a general-purpose processor, what additional circuitry is needed, and present a full instruction-set architecture. The result is a very compact processor that can run at high speed, while executing a full gamut of general machine instructions. We provide results for a number of simple applications, and show how the processor's resource requirements and frequency compare to a Xilinx MicroBlaze soft core. Based on the DSP48E1, this processor can be deployed across next-generation Xilinx Artix-7, Kintex-7, and Virtex-7 families. Hui Yan Cheah, Suhaib A. Fahmy, Douglas L. Maskell |
FPT | 2 |
| 2012 | A high speed open source controller for FPGA Partial ReconfigurationabstractPartial Reconfiguration (PR) is an advanced technique, which improves the flexibility of FPGAs by allowing portions of a design to be reconfigured at runtime by overwriting parts of the configuration memory. PR is an important enabler for implementing adaptive systems. However, the design of such systems can be challenging, and this is especially true of the configuration controller. The generally supported methods and IP have low throughput, resulting in long configuration time that precludes PR from systems where this operation needs to be fast. In this paper, we present a high-speed configuration controller that provides several features useful in adaptive systems. The design has been released for use by the wider research community. Kizheppatt Vipin, Suhaib A. Fahmy |
FPT | 2 |
| 2011 | Enabling high level design of adaptive systems with partial reconfigurationabstractAdaptive systems have the ability to respond to environmental conditions, by modifying their processing at runtime. While this is easy to do software systems, modern algorithms can be computationally expensive, requiring powerful processors. At the same time hardware is not as flexible. Field programmable gate arrays (FPGAs) are recognised as being suitable for adaptive systems implementation, due to their flexibility and high performance. The use of partial reconfiguration on FPGAs to implement adaptive systems has been proposed many times in the literature. However the design process for partially reconfigurable systems is complex and requires specialist knowledge on behalf of the application designer. Hence, it has remained a rarely used capability outside of academic circles. We propose a new approach to leverage partial reconfiguration within adaptive systems, by integrating with, rather than circumventing, supported vendor tool flows, while automating many of the steps that have made such designs more difficult in the past. This makes it possible for system designers with less FPGA expertise to use partial reconfiguration when designing adaptive systems. Kizheppatt Vipin, Suhaib A. Fahmy |
FPT | 2 |
| 2011 | A threat-based Connect6 implementation on FPGAabstractConnect6 is a new generation k-in-a-row game, which has drawn great interest not only from game enthusiasts, but also from researchers, due to its characteristics such as fairness and high state-space complexity. In this paper we describe the design and implementation of an FPGA-based Connect6 player that can compete against other computer-based opponents, communicating through a serial interface. Our algorithmic implementation utilises only basic FPGA building blocks such as LUTs and flip-flops and does not include any IP cores or hardware macros, making it portable across different FPGA platforms without design modifications. The design has been implemented and validated both on a Xilinx Spartan-3A, and a Xilinx Spartan-6 FPGA boards. The algorithm uses a powerful threat-based placement strategy that maximises the FPGA's winning opportunity while reducing the opponent's options. Extended simulation and evaluation based on software and human players confirms that our FPGA-based implementation performs well, and the algorithm used in the design leads to a high probability of success. Kizheppatt Vipin, Suhaib A. Fahmy |
FPT | 2 |
| 2011 | Efficient region allocation for adaptive partial reconfigurationabstractWhile partial reconfiguration on FPGAs has attracted significant research interest in recent years, designing systems that leverage it remains a specialist skill. Systems with a large number of reconfigurable modules can be challenging to design. Deciding on how many reconfigurable regions to use is not always straightforward, yet this choice impacts area efficiency and configuration latency. Current FPGA partial reconfiguration tool flows require the designer to have detailed knowledge of the physical architecture of the FPGA. It is the responsibility of the designer to decide on the location and size of regions. In this paper we introduce a formulation for determining the tradeoff between the number of reconfigurable regions, the allocation of modules to regions, and the reconfiguration overhead, represented by area and reconfiguration time. Throughout the investigation, we consider the heterogeneous nature of modern FPGAs as well as limitations imposed by current tools. We then show that adopting an optimal allocation can result in both area savings and a reduction in reconfiguration time over a standard approach to allocation. Kizheppatt Vipin, Suhaib A. Fahmy |
FPT | 2 |
| 2011 | A Model-Based Approach to Cognitive Radio DesignabstractCognitive radio is a promising technology for fulfilling the spectrum and service requirements of future wireless communication systems. Real experimentation is a key factor for driving research forward. However, the experimentation testbeds available today are cumbersome to use, require detailed platform knowledge, and often lack high level design methods and tools. In this paper we propose a novel cognitive radio design technique, based on a high-level model which is implementation independent, supports design-time correctness checks, and clearly defines the underlying execution semantics. A radio designed using this technique can be synthesised to various real radio platforms automatically; detailed knowledge of the target platform is not required. The proposed technique therefore simplifies cognitive radio design and implementation significantly, allowing researchers to validate ideas in experiments without extensive engineering effort. One example target platform is proposed, comprising software and reconfigurable hardware. The design technique is demonstrated for this platform through the development of two realistic cognitive radio applications. Jörg Lotze, Suhaib A. Fahmy, Juanjo Noguera, Linda Doyle |
IEEE J. Sel. Areas Commun. | 2 |
| 2010 | Histogram-based probability density function estimation on FPGAsabstractProbability density functions (PDFs) have a wide range of uses across an array of application domains. Since computing the PDF of real-time data is typically expensive, various estimations have been devised that attempt to approximate the real PDFs based on fitting data to an expected underlying distribution. As we move to more adaptive systems, real-time monitoring of signal statistics increases in importance. In this paper, we present a technique that leverages the heterogeneous resources on modern FPGAs to enable real time computation of PDFs of sampled data at speeds of over 200 Msamples per second. We detail a flexible architecture that can be used to extract statistical information in real time while consuming a moderate amount of area, allowing it to be incorporated into existing FPGA-based applications. Suhaib A. Fahmy |
FPT | 1 |
| 2009 | Generic Software Framework for Adaptive Applications on FPGAsabstractAdaptive systems are set to become more mainstream, as numerous practical applications in the communications domain emerge. FPGAs offer an ideal implementation platform, combining high performance with flexibility. While significant research has been undertaken in the area of FPGA partial reconfiguration, it has focussed primarily on low-level architecture-specific implementations. Building upon this previous work, we present a system model and software architecture for implementing runtime adaptive applications on FPGAs, separating the control and processing planes and abstracting away the details of hardware reconfiguration from the system designer. Hardware processing components appear as software components in the runtime system, enabling their inclusion in adaptive applications. We present an adaptive wireless application, demonstrating the use of the model and software architecture. Suhaib A. Fahmy, Jörg Lotze, Juanjo Noguera, Linda Doyle, Robert Esser |
FCCM | 1 |
| 2009 | Development Framework for Implementing FPGA-Based Cognitive Network NodesabstractThis paper identifies important features a cognitive radio framework should provide, namely a virtual architecture for hardware abstraction, an adaptive run-time system for managing cognition, and high level design tools for cognitive radio development. We evaluate a range of existing frameworks with respect to these, and propose a novel FPGA-based framework that provides all these features. By abstracting away the details of hardware reconfiguration, radio designers can implement FPGA-based cognitive nodes much as they would do for a software implementation. We apply the proposed framework to the design and implementation of a receiver node that works in two modes: discovery, where it uses spectrum sensing to find a radio transmission, and communication, in which it receives and demodulates the said transmission. We show how the whole design process does not require any hardware experience on behalf of the radio designer. Jörg Lotze, Suhaib A. Fahmy, Juanjo Noguera, Baris Özgül, Linda Doyle, Robert Esser |
GLOBECOM | 2 |
| 2006 | Investigating Trace Transform Architectures for Face AuthenticationabstractThe overall aim of this research is to thoroughly investigate which functionals are able to provide the best performance for the trace transform. The resultant framework can be applied generally to other domains too. This is an example of how the rapid prototyping afforded by FPGAs and the performance advantages serve as the ideal platform for algorithmic exploration. Furthermore, the power of runtime-reconfigurability assists further in the investigation Suhaib A. Fahmy |
FPL | 1 |
| 2006 | Efficient Realtime FPGA Implementation of the Trace TransformabstractThe trace transform is a novel image transform that is able to exhibit useful properties such as scale and rotation invariance and occlusion robustness. As a result, it is particularly suited to a variety of classification and recognition tasks including image database search, token registration, activity monitoring, character recognition and face authentication. The main obstacle to the widespread use of the transform is its high computational complexity. This has precluded a detailed investigation of transform parameters. This paper presents an architecture and implementation of a trace transform engine on a Virtex-II FPGA. By exploiting the inherent parallelism in the algorithm and the use of optimised functional blocks, a huge performance gain is achieved, exceeding realtime video processing requirements for a 256 times 256 image Suhaib A. Fahmy, Christos-Savvas Bouganis, Peter Y. K. Cheung, Wayne Luk |
FPL | 1 |
| 2005 | Hardware Acceleration of Hidden Markov Model Decoding for Person DetectionabstractThis paper explores methods for hardware acceleration of hidden Markov model (HMM) decoding for the detection of persons in still images. Our architecture exploits the inherent structure of the HMM trellis to optimise a Viterbi decoder for extracting the state sequence front observation features. Further performance enhancement is obtained by computing the HMM trellis states in parallel. The resulting hardware decoder architecture is mapped onto a field programmable gate array (FPGA). The performance and resource usage of our design is investigated for different levels of parallelism. Performance advantages over software are evaluated. We show how this work contributes to a real-time system for person-tracking in video-sequences. Suhaib A. Fahmy, Peter Y. K. Cheung, Wayne Luk |
DATE | 1 |
| 2005 | Novel FPGA-Based Implementation of Median and Weighted Median Filters for Image ProcessingabstractAn efficient hardware implementation of a median filter is presented. Input samples are used to construct a cumulative histogram, which is then used to find the median. The resource usage of the design is independent of window size, but rather, dependent on the number of bits in each input sample. This offers a realisable way of efficiently implementing large-windowed median filtering, as required by transforms such as the Trace Transform. The method is then extended to weighted median filtering. The designs are synthesised for a Xilinx Virtex II FPGA and the performance and area compared to another implementation for different sized windows. Intentional use of the heterogeneous resources on the FPGA in the design allows for a reduction in slice usage and high throughput. Suhaib A. Fahmy, Peter Y. K. Cheung, Wayne Luk |
FPL | 1 |