Khaled Benkrid

dblp:50/3989 · DBLP profile ↗
← Back
43ranked-venue papers
16as first author
0since 2021 · last 2015
0009-0009-2673-2668ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 37 · 13 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Reconfigurable computing and FPGAs · 60% Embedded and real-time systems · 11% Hardware accelerators and domain-specific architectures · 11%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
2 papers
Image and video processing · 100%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs › FPGA virtualization
FPGA operating system
0.212013
R3TOS: A Novel Reliable Reconfigurable Real-Time Operating System for Highly Adaptive, Efficient, and Dependable Computing on FPGAs · IEEE Trans. Computers 2013
Reconfigurable computing and FPGAs
FPGA implementation
0.132005
A single-FPGA implementation of image connected component labelling · FPGA 2003
Design framework for the implementation of the 2-D orthogonal discrete wavelet transform on FPGA · FPGA 2003
An integrated framework for the high level design of high performance signal processing circuits on FPGAs (abstract only) · FPGA 2005
Reconfigurable computing and FPGAs
FPGA accelerator
0.112009
A high performance fpga-based implementation of position specific iterated blast · FPGA 2009
Hardware accelerators and domain-specific architectures › bioinformatics accelerator
sequence alignment accelerator
0.112009
A high performance fpga-based implementation of position specific iterated blast · FPGA 2009
Integrated circuit design › digital signal processing circuits
digital signal processor design
0.112005
An integrated framework for the high level design of high performance signal processing circuits on FPGAs (abstract only) · FPGA 2005
Embedded and real-time systems › real-time scheduling
hardware task scheduling
0.012013
R3TOS: A Novel Reliable Reconfigurable Real-Time Operating System for Highly Adaptive, Efficient, and Dependable Computing on FPGAs · IEEE Trans. Computers 2013
Embedded and real-time systems
real-time scheduling
0.012013
R3TOS: A Novel Reliable Reconfigurable Real-Time Operating System for Highly Adaptive, Efficient, and Dependable Computing on FPGAs · IEEE Trans. Computers 2013
Parallel and multicore computing › parallel graph algorithms
connected components
0.012003
A single-FPGA implementation of image connected component labelling · FPGA 2003
Bioinformatics and computational biology › sequence alignment › heuristic alignment
PSI-BLAST
0.012009
A high performance fpga-based implementation of position specific iterated blast · FPGA 2009
Bioinformatics and computational biology
sequence alignment
0.012009
A high performance fpga-based implementation of position specific iterated blast · FPGA 2009
Image and video processing › binary image processing
connected component labeling
0.012003
A single-FPGA implementation of image connected component labelling · FPGA 2003
Image and video processing
image segmentation
0.012003
A single-FPGA implementation of image connected component labelling · FPGA 2003
Image and video processing
wavelet transform
0.012003
Design framework for the implementation of the 2-D orthogonal discrete wavelet transform on FPGA · FPGA 2003
Electronic design automation
high-level synthesis
0.012003
Design framework for the implementation of the 2-D orthogonal discrete wavelet transform on FPGA · FPGA 2003

Methods — techniques the papers use, named apart from their topics

online scheduling · 0.2dynamic partial reconfiguration · 0.2handel-c · 0.1symmetric extension · 0.1structural hardware description · 0.1HIDE · 0.1wordlength rounding analysis · 0.0serial iterative algorithm · 0.0pyramid algorithms · 0.0pyramid algorithm · 0.0prolog · 0.0logic programming · 0.0EDIF generation · 0.0
YearPublicationVenuePosition
2015 Microkernel Architecture and Hardware Abstraction Layer of a Reliable Reconfigurable Real-Time Operating System (R3TOS)
abstract
This article presents a new solution for easing the development of reconfigurable applications using Field-Programable Gate Arrays (FPGAs). Namely, our Reliable Reconfigurable Real-Time Operating System (R3TOS) provides OS-like support for partially reconfigurable FPGAs. Unlike related works, R3TOS is founded on the basis of resource reusability and computation ephemerality. It makes intensive use of reconfiguration at very fine FPGA granularity, keeping the logic resources used only while performing computation and releasing them as soon as it is completed. To achieve this goal, R3TOS goes beyond the traditional approach of using reconfigurable slots with fixed boundaries interconnected by means of a static communication infrastructure. Instead, R3TOS approaches a static route-free system where nearly everything is reconfigurable. The tasks are concatenated to form a computation chain through which partial results naturally flow, and data are exchanged among remotely located tasks using FPGA’s reconfiguration mechanism or by means of “removable” routing circuits. In this article, we describe the R3TOS microkernel architecture as well as its hardware abstraction services and programming interface. Notably, the article presents a set of novel circuits and mechanisms to overcome the limitations and exploit the opportunities of Xilinx reconfigurable technology in the scope of hardware multitasking and dependability.
Xabier Iturbe, Khaled Benkrid, Chuan Hong, Ali Ebrahim, Raul Torrego, Tughrul Arslan
ACM Trans. Reconfigurable Technol. Syst.2
2013 Reconfigurable feeding network for GSM/GPS/3G/WiFi and global LTE applications
abstract
This paper presents a novel miniaturized reconfigurable and switchable feeding network to cover GSM, GPS, 3G, WiFi and global LET standards. The feeding network consists of four conventional Wilkinson power dividers which can be individually reconfigured in length using PIN diodes switches. By controlling the bias voltages of these PIN diodes, the operating frequency of the proposed design can be converted between four different bands: 600MHz-900MHz, 1.2GHz-1.6GHz, 1.8GHz-2.2GHz and 2.4GHz-2.6GHz. The first frequency band (600MHz-900MHz) is applied to satisfy the applications of LTE US (700MHz), LTE UK (800MHz) and GSM (850MHz, 900MHz). The second band (1.2GHz-1.6GHz) targets GPS L1 (1.575GHz) and GPS L2 (1.227GHz). Different GSM (1800MHz, 1900MHz) and 3G standards (UMTS, W-CDMA, TD-SCDMA and CDMA2000) are located in the third frequency band (1.8GHz-2.2GHz). The last band (2.4GHz-2.6GHz) is used to cover WiFi (2.45GHz) and LTE Europe (2.6GHz). The miniaturized and optimized feeding network exhibits good performance for S-Parameters in each band, which includes low return loss, equal power splitting and suitable insertion loss. Within the simulation environment, three types of PIN diode models were constructed and investigated in order to improve accuracy. The feeding network is implemented on an FR4 substrate. Fabrication and measurement results closely correlate with those obtained during design simulations. The reconfigurable feeding network can be particularly applied to commercial multiband communication systems.
Tughrul Arslan, Khaled Benkrid, Ahmed O. El-Rayis, Nakul Haridas
ISCAS3
2013 Guest Editors' introduction: Special section on adaptive hardware and systems
abstract
This special section of IEEE Transactions on Computers presents some of the latest research developments in the field of adaptive hardware and systems. The creation of this section was motivated by lively discussions held at the annual NASA/ESA Adaptive Hardware and Systems (AHS) conference, which showed a need for such special section at a top ranked journal. At the end of a rigorous review process, ten papers were selected for publication from a set of high quality submissions consisting of regular papers and extended papers from the AHS 2012 conference proceedings. The articles are then briefly described.
Khaled Benkrid, Didier Keymeulen, Umeshkumar D. Patel, David Merodio Codinachs
IEEE Trans. Computers1
2013 R3TOS: A Novel Reliable Reconfigurable Real-Time Operating System for Highly Adaptive, Efficient, and Dependable Computing on FPGAs
abstract
Despite the clear potential of FPGAs to push the current power wall beyond what is possible with general-purpose processors, as well as to meet ever more exigent reliability requirements, the lack of standard tools and interfaces to develop reconfigurable applications limits FPGAs' user base and makes their programming not productive. R3TOS is our contribution to tackle this problem. It provides systematic OS support for FPGAs, allowing the exploitation of some of the most advanced capabilities of FPGA technology by inexperienced users. What makes R3TOS special is its nonconventional way of exploiting on-chip resources: These are used indistinguishably for carrying out either computation or communication tasks at different times. Indeed, R3TOS does not rely on any static infrastructure apart from its own core circuitry, which is constrained to a specific region within the FPGA where it is implemented. Thus, the rest of the device is kept free of obstacles, with the spare resources ready to be used as and whenever needed. At runtime, the hardware tasks are scheduled and allocated with the dual objective of improving computation density and circumventing damaged resources on the FPGA.
Xabier Iturbe, Khaled Benkrid, Chuan Hong, Ali Ebrahim, Raul Torrego, Imanol Martinez, Tughrul Arslan, Jon Pérez 0001
IEEE Trans. Computers2
2012 Design and implementation of fault-tolerant soft processors on FPGAs
abstract
This paper presents a novel hardware mechanism to facilitate the design and implementation of soft processors on FPGAs using the Error-correcting code (ECC)-protected memory and Triple Modular Redundancy (TMR). Such techniques highly harden the fault tolerance of soft processors, especially their memories, which are the most radiation susceptible resources on FPGAs. This is demonstrated in the implementation of a fault-tolerant PicoBlaze processor on Xilinx FPGAs, in which we used an additional LookAhead technique to synchronize the processor with ECC-protected Block RAM (ECC BRAM). The resulting fault-tolerant PicoBlaze processor has the benefit of having a self-recoverable program memory in the presence of Single Error Upsets (SEUs), without halting the processor. Our techniques can be applied to other soft processors e.g. Xilinx MicroBlaze or Altera Nios.
Chuan Hong, Khaled Benkrid, Xabier Iturbe, Ali Ebrahim
FPL2
2012 An adaptive FPGA implementation of multi-core K-nearest neighbour ensemble classifier using dynamic partial reconfiguration
abstract
Classification of highly dimensional Microarray data using K-nearest neighbour (K-NN) is a time-consuming task when implemented on general purpose processors (GPPs), and such it can benefit greatly from a parallel hardware implementation. In this work, an FPGA implementation of the K-NN classifier is presented and compared with an equivalent implementation running on GPP. Then, a novel FPGA-based multi-core implementation of the K-NN ensemble classifier, which exploits dynamic partial reconfiguration (DPR) is presented. The FPGA implementation of the single core K-NN classifier was found to be 92× faster than a GPP implementation, and the ensemble implementation was found to offer ~5× speed-up of the FPGA reconfiguration time. In addition, the paper investigates the effect of data dimensionality on classification time on both FPGAs and GPPs, showing that FPGAs scale up better than GPPs with higher data dimensionality.
Hanaa M. Hussain, Khaled Benkrid, Chuan Hong, Huseyin Seker 0001
FPL2
2012 IP-XACT extensions for IP interoperability guarantees and software model generation
abstract
This paper presents a set of novel metadata extensions that are used to specify the interfaces on Xilinx IP cores and their software models under a uniform data model which allows enhanced design rule checking in the system design process. We also present a suite of tools which can be used to generate executable software simulation models of complete systems from their specifications under that data model. These tools may be used stand-alone, or may be used to extend the capabilities of the Vivado IP Integrator tool that has recently been released by Xilinx. Our tool flow has been used successfully to generate software simulation models for two 3GPP Long Term Evolution (LTE) physical layer systems: uplink receive and downlink transmit.
Thomas P. Perry, Richard L. Walke, Rob Payne, Stefan Petko, Khaled Benkrid
FPL5
2012 Data coding functions for Software Defined Radios implemented on R3TOS
abstract
This paper presents the implementation of several data coding functions used in Software Defined Radios, on R3TOS: a Reliable, Reconfigurable and Real-Time Operating System. The latter offers efficient high performance computing on FPGAs as well as protection against emerging faults, hence making it a perfect candidate for the implementation of SDRs. In particular, R3TOS' ICAP-based Inter-task Communication Infrastructure (I2CI) has been used for data feeding and collection from coding functions, while a task context saving and restoration procedure, and a fast function parameterization system have been developed in order to improve system performance. The design of the data coding functions has been carried out using Xilinx's rapid prototyping tool System Generator in order to ease their development.
Raul Torrego, Inaki Val, Eñaut Muxika, Xabier Iturbe, Khaled Benkrid
FPL5
2012 Digital Hardware Design Teaching: An Alternative Approach
abstract
This article presents the design and implementation of a complete review of undergraduate digital hardware design teaching in the School of Engineering at the University of Edinburgh. Four guiding principles have been used in this exercise: learning-outcome driven teaching, deep learning, affordability, and flexibility. This has identified discrete electronics as key components in the early stages of the curriculum and FPGAs as an economical platform for the teaching of various digital hardware design concepts and techniques in later stages of the curriculum. In particular, the article presents the detailed design and implementation of one digital hardware design laboratory, called Gateway, which introduces students to synchronous digital circuit development from high level functional specifications, uses Verilog for hardware description and FPGAs as an implementation platform. Biggs’ theory of constructive alignment was applied in the design of this lab’s learning outcomes, lab content, teaching and learning methods, and assessment methods. The lab makes extensive use of multimedia in both lab content delivery and demonstration applications developed by students. Student feedback following the deployment of this lab was overwhelmingly positive and an evaluation of the lab results compared to previous lab offerings’ shows the merit of the approach taken.
Khaled Benkrid, Thomas Clayton
ACM Trans. Comput. Educ.1
2011 Methods and Mechanisms for Hardware Multitasking: Executing and Synchronizing Fully Relocatable Hardware Tasks in Xilinx FPGAs
abstract
This paper presents the details of a novel technique which allows for the implementation and execution of completely relocatable hardware tasks onto dynamically reconfigurable FPGAs. Our novel technique harnesses the internal configuration access port (ICAP) for inter-task communication and synchronization, leading to very little logic overheads. The advantages of this technique include fault-tolerance, as tasks could be relocated freely on the fabric to circumvent damaged resources, and high performance, due to better exploitation of the logic fabric. The work is part of a larger effort in our group which aims to build a fully operational dynamically reconfigurable computer which would satisfy the often conflicting requirements of high performance, fault-tolerance and high level programming.
Xabier Iturbe, Khaled Benkrid, Tughrul Arslan, Raul Torrego, Imanol Martinez
FPL2
2011 High Performance Phylogenetic Analysis With Maximum Parsimony on Reconfigurable Hardware
abstract
We present in this paper the detailed field-programmable gate-array (FPGA) design of the Maximum Parsimony method for molecular-based phylogenetic analysis and its implementation on the nodes of an FPGA supercomputer called Maxwell. This is the first FPGA implementation of this method for nucleotide sequence data reported in the literature. The hardware architecture consists in a linear systolic array composed of 20 processing elements each of which performing Sankoff's algorithm for a different tree topology in parallel. This array computes the scores of all theoretically possible trees for a given number of taxa in several iterations. The currently supported maximum number of taxa is 12 but this number can be easily increased. Furthermore, the resulting implementation outperforms an equivalent desktop-based software implementation (using phylogenetic analysis using parsimony software) by several orders of magnitude. The speed-up values achieved by the hardware implementation on a single node of the Maxwell machine can reach up to four orders of magnitude for the 12-taxa case while implementations on several Maxwell nodes can yield even higher speed-ups. This is achieved through harnessing both coarse-grain and fine-grain parallelism available in the algorithm and corresponding hardware implementation platform.
Server Kasap, Khaled Benkrid
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Highly efficient mapping of the Smith-Waterman algorithm on CUDA-compatible GPUs
abstract
This paper describes a multi-threaded parallel design and implementation of the Smith-Waterman (SW) algorithm on graphic processing units (GPUs) with NVIDIA corporation's Compute Unified Device Architecture (CUDA). Central to this is a divide and conquer approach which divides the computation of a whole pairwise sequence alignment matrix into multiple sub-matrices (or parallelograms) each running efficiently on the available hardware resources of the GPU in hand, with temporary intermediate data stored in global memory. Moreover, we use thread warps and padding techniques in order to decrease the cost of thread synchronization, as well as loop unrolling in order to reduce the cost of conditional branches. While intermediate data is stored in global memory for large queries, the most inner loop in our implementation will only access shared memory and registers. As a result of these optimizations, our implementation of the SW algorithm achieves a throughput ranging between 9.09 GCUPS (Giga Cell Update per Second) and 12.71 GCUPS on a single-GPU version, and a throughput between 29.46 GCUPS and 43.05 GCUPS on a quad-GPU platform. Compared with the best GPU implementation of the SW algorithm reported to date, our implementation achieves up to 46 % improvement in speed. The source code of our implementation is available in the public domain for Bioinformaticians to benefit from its performance.
Keisuke Dohi, Khaled Benkrid, Cheng Ling, Tsuyoshi Hamada, Yuichiro Shibata
ASAP2
2010 ATB: Area-Time response Balancing algorithm for scheduling real-time hardware tasks
abstract
This paper describes a novel scheduling algorithm for the execution of hardware tasks with real-time constraints onto partially and dynamically reconfigurable FPGAs. The Area-Time response Balancing scheduling algorithm (ATB) is inspired by the well-known Earliest Deadline First (EDF) algorithm, which is extended with a technique for reducing the fragmentation on FPGA's reconfigurable area. This technique promotes the reuse of the resources that are released when great area tasks finish their execution by smaller area tasks as long as the real-time constraints permit to do so. Providing an exclusively time-based algorithm, such as EDF, with support for dealing with area-related issues ensures the best results. Simulation results reported in this paper show that ATB misses 23% less deadlines than EDF. Moreover, since FPGA's damaged resources provoke unpredictable fragmentation on the device, ATB is currently the best scheduling option to be used in a Reliable Reconfigurable Real-Time Operating System (R3TOS).
Xabier Iturbe, Khaled Benkrid, Tughrul Arslan, Imanol Martinez, Mikel Azkarate-askatsua
FPT2
2010 High-Performance Quasi-Monte Carlo Financial Simulation: FPGA vs. GPP vs. GPU
abstract
Quasi-Monte Carlo simulation is a special Monte Carlo simulation method that uses quasi-random or low-discrepancy numbers as random sample sets. In many applications, this method has proved advantageous compared to the traditional Monte Carlo simulation method, which uses pseudo-random numbers, thanks to its faster convergence and higher level of accuracy. This article presents the design and implementation of a massively parallelized Quasi-Monte Carlo simulation engine on an FPGA-based supercomputer, called Maxwell. It also compares this implementation with equivalent graphics processing units (GPUs) and general purpose processors (GPP)-based implementations. The detailed comparison between these three implementations (FPGA vs. GPP vs. GPU) is done in the context of financial derivatives pricing based on our Quasi-Monte Carlo simulation engine. Real hardware implementations on the Maxwell machine show that FPGAs outperform equivalent GPP-based software implementations by 2 orders of magnitude, with the speed-up figure scaling linearly with the number of processing nodes used (FPGAs/GPPs). The same implementations show that FPGAs achieve a ~ 3x speedup compared to equivalent GPU-based implementations. Power consumption measurements also show FPGAs to be 336x more energy efficient than CPUs, and 16x more energy efficient than GPUs.
Xiang Tian 0001, Khaled Benkrid
ACM Trans. Reconfigurable Technol. Syst.2
2009 A high performance fpga-based implementation of position specific iterated blast
abstract
We present in this paper the first reported FPGA implementation of the Position Specific Iterated BLAST (PSI-BLAST) algorithm. The latter is a heuristic biological sequence alignment algorithm that is widely used in the bioinformatics and computational biology world in order to detect weak homologs. The architecture of our FPGA implementation is parameterized in terms of sequence lengths, scoring matrix, gap penalties and cut-off and threshold values. It is composed of various blmocks each of which performs one step of the algorithm in parallel. This results in high performance implementations, which easily outperform equivalent software implementations by one order of magnitude or more. Furthermore, the core was captured in an FPGA-platform-independent language, namely the Handel-C language, to which no specific resource inference or placement constraints were applied. This makes our core portable across different FPGA families and architectures.
Server Kasap, Khaled Benkrid, Ying Liu 0003
FPGA2
2009 Novel Area-Efficient FPGA Architectures for FIR Filtering With Symmetric Signal Extension
abstract
This paper presents four novel area-efficient field-programmable gate-array (FPGA) bit-parallel architectures of finite impulse response (FIR) filters that smartly support the technique of symmetric signal extension while processing finite length signals at their boundaries. The key to this is a clever use of variable-depth shift registers which are efficiently implemented in Xilinx FPGAs in the form of shift register logic (SRL) components. Comparisons with the conventional architecture of FIR filter with symmetric boundary processing show considerable area saving especially with long-tap filters. For instance, our architecture implementation of the 8-tap low Daubechies-8 FIR filter achieves ~ 30% reduction in the area requirement (in terms of slices) compared to the conventional architecture while maintaining the same throughput. Two of the above-cited novel architectures are dedicated to the special case of symmetric FIR filters. The first architecture is highly area-efficient but requires a clock frequency doubler. While this reduces the overall processing speed (to a maximum of 2), it does maintain a high throughput. Moreover, this speed penalty is cancelled in bi-phase filters which are widely used in multirate architectures (e.g., wavelets). Our second symmetric FIR filter architecture saves less logic than the first architecture (e.g., 10% with the 9-tap low Biorthogonal 9&7 symmetric filter instead of 37% with the first architecture) but overcomes its speed penalty as it matches the throughput of the conventional architecture.
Abdsamad Benkrid, Khaled Benkrid
IEEE Trans. Very Large Scale Integr. Syst.2
2009 A Highly Parameterized and Efficient FPGA-Based Skeleton for Pairwise Biological Sequence Alignment
abstract
This paper presents the design and implementation of the most parameterisable field-programmable gate array (FPGA)-based skeleton for pairwise biological sequence alignment reported in the literature. The skeleton is parameterised in terms of the sequence symbol type, i.e., DNA, RNA, or protein sequences, the sequence lengths, the match score, i.e., the score attributed to a symbol match, mismatch or gap, and the matching task, i.e., the algorithm used to match sequences, which includes global alignment, local alignment, and overlapped matching. Instances of the skeleton implement the Smith-Waterman and the Needleman-Wunsch algorithms. The skeleton has the advantage of being captured in the Handel-C language, which makes it FPGA platform-independent. Hence, the same code could be ported across a variety of FPGA families. It implements the sequence alignment algorithm in hand using a pipeline of basic processing elements, which are tailored to the algorithm parameters. This paper presents a number of optimizations built into the skeleton and applied at compile-time depending on the user-supplied parameters. These result in high performance FPGA implementations tailored to the algorithm in hand. For instance, actual hardware implementations of the Smith-Waterman algorithm for Protein sequence alignment achieve speedups of two orders of magnitude compared to equivalent standard desktop software implementations.
Khaled Benkrid, Ying Liu 0003, Abdsamad Benkrid
IEEE Trans. Very Large Scale Integr. Syst.1
2008 High performance FPGA-based core for BLAST sequence alignment with the two-hit method
abstract
This paper presents the design and implementation of a high performance FPGA-based core for BLAST sequence alignment with the two-hit method. BLAST with two-hit is a very widely used heuristic biological sequence alignment algorithm, and this paper is the first reported FPGA implementation of it, to our knowledge. The architecture of our core is parameterized in terms of the sequence lengths, match scores, gap penalties, and cut-off and threshold values. It is composed of various blocks each of which performs one step of the algorithm in parallel with the others. This results in a high performance and efficient FPGA implementation, which outperforms equivalent software implementations by one order of magnitude or more. Real hardware implementations show that our core is 52 times faster than equivalent software implementations, on average. Furthermore, the core was captured in an FPGA-platform-independent language, namely the Handel-C language, to which no specific resource inference or placement constraints were applied. Hence, the same code can be easily ported to different FPGA families and architectures.
Server Kasap, Khaled Benkrid, Ying Liu 0003
BIBE2
2008 Design and implementation of a high performance financial Monte-Carlo simulation engine on an FPGA supercomputer
abstract
Monte-Carlo simulation is a very widely used technique in scientific computations in general with huge computation benefits in solving problems where closed form solutions are impossible to derive. This technique is also characterized by a high degree of parallelism as a large number of different simulation paths need to be calculated, which makes it ideal for a parallel hardware implementation. This paper illustrates the benefits of such implementation in the context of financial computing as it implements a financial Monte-Carlo simulation engine on an FPGA-based supercomputer, called Maxwell, developed at the University of Edinburgh. The latter consists of a 32 CPU cluster augmented with 64 Virtex-4 Xilinx FPGAs connected in a 2D torus. Our engine can implement various Monte-Carlo simulations on the Maxwell machine with speed-ups in the 3-order magnitude compared to equivalent software implementations. This is illustrated in this paper in the context of an implementation of the Black-Scholes option pricing model. Real hardware implementation shows that our FPGA-based implementation of the Black-Scholes model outperforms an equivalent software implementation running on a workstation cluster with the same number of computing nodes (CPU/FPGA) by a factor of 750, which is the fastest ever reported FPGA implementation of this model.
Xiang Tian 0001, Khaled Benkrid
FPT2
2007 Design and Implementation of a Highly Parameterised FPGA-Based Skeleton for Pairwise Biological Sequence Alignment
abstract
This paper presents the design and implementation of a generic and highly parameterised FPGA-based skeleton for pairwise biological sequence alignment. The skeleton is parameterised in terms of the sequence symbol type i.e. DNA, RNA, or protein sequences, the sequence lengths, the match score i.e. the score attributed to a symbol match or the penalty attributed to a mismatch or gap, and the matching task. Instances of the skeleton implement the Smith-Waterman and the Needleman-Wunsch algorithms. The skeleton has been captured in the Handel-C language which makes it FPGA-platform-independent. It implements the sequence alignment algorithm in hand using a pipeline of basic processing elements, which are tailored to the supplied parameters. Actual hardware implementations of the Smith-Waterman algorithm for protein sequence alignment achieve speed-ups in excess of 100:1 compared to equivalent standard desktop software implementations.
Khaled Benkrid, Ying Liu 0003, Abdsamad Benkrid
FCCM1
2007 High Performance Biosequence Database Scanning using FPGAs
abstract
This paper presents the design and implementation of a generic and highly parameterised FPGA-based core for pairwise biological sequence alignment. The core is captured in the Handel-C language, which allows for high level software-like descriptions of hardware architectures. It implements the sequence alignment algorithm in hand using a pipeline of basic processing elements. This results in high performance FPGA implementations tailored to the algorithm in hand. For instance, actual hardware implementations of the Smith-Waterman algorithm for protein sequence alignment achieve speed-ups in excess of 100:1 compared to equivalent standard PC-based software implementations.
Khaled Benkrid, Ying Liu 0003, Abdsamad Benkrid
ICASSP (1)1
2007 Efficient FPGA hardware development: A multi-language approach
Khaled Benkrid, Abdsamad Benkrid, Samir Belkacemi
J. Syst. Archit.1
2006 Handling finite length signals borders in two-channel multirate filter banks for perfect reconstruction
Abdsamad Benkrid, Khaled Benkrid
Signal Process.2
2005 An integrated framework for the high level design of high performance signal processing circuits on FPGAs (abstract only)
abstract
This paper proposes an integrated framework for the high level design of high performance signal processing algorithms' implementations on FPGAs. The framework emerged from a constant need to rapidly implement increasingly complicated algorithms on FPGAs while maintaining the high performance needed in many real time digital signal processing applications. This is particularly important for application developers who often rely on iterative and interactive development methodologies.The central idea behind the proposed framework is to dynamically integrate high performance structural hardware description languages with higher level hardware languages in other to help satisfy the dual requirement of high level design and high performance implementation. The paper illustrates this by integrating two environments: Celoxica's Handel-C language, and HIDE, a structural hardware environment developed at the Queen's University of Belfast.
Khaled Benkrid, Samir Belkacemi
FPGA1
2004 From application descriptions to hardware in seconds: a logic-based approach to bridging the gap
abstract
This paper presents a high-level hardware description environment developed at Queen's University, Belfast, U.K., which aims to bridge the gap between application design and hardware description. The environment, called application-to-hardware (A2H), allows for efficient compilation of high-level application descriptions to field programmable gate array (FPGA) hardware in the form of EDIF netlist in seconds. A key concept in bridging the gap while retaining the hardware efficiency, is that of hardware skeletons. A hardware skeleton is a parameterized description of a task-specific architecture, to which the user can supply not only value parameters but also functions or even other skeletons. A skeleton contains built-in rules, which capture optimizations specific to the target hardware at the implementation phase. The rule-based logic programming language Prolog has been chosen as the base notation for the A2H environment. This paper includes descriptions of hardware skeletons abstractions in the particular context of image processing applications. The current implementation of our system targets Xilinx XC4000 and Virtex series FPGAs.
Khaled Benkrid, Danny Crookes
IEEE Trans. Very Large Scale Integr. Syst.1
2003 A Logic Based Hardware Development Environment
abstract
This paper presents a logic-based approach to hardware abstraction and composition based on the logic programming language Prolog. This is an attempt to satisfy the dual requirement of abstract hardware design and hardware efficiency. Central to this approach is a hardware description environment called HIDE, which provides more abstract and elegant hardware descriptions and compositions than are possible in traditional hardware description languages such as VHDL or Verilog. The environment enables highly scaleable and parameterized composition of blocks using a small set of constructors e.g. 'horizontal' and 'vertical' for 2D circuit abstractions and the novel 'above' constructor for 3D circuit compositions. It also generates preplaced configurations in EDIF (and VHDL) format for Xilinx FPGAs (field programmable gate arrays).
Samir Belkacemi, Khaled Benkrid, Danny Crookes
FCCM2
2003 Design and Implementation of a Generic 2-D Orthogonal Discrete Wavelet Transform on FPGA
abstract
This paper gives a design framework for the implementation of the 2D (two-dimensional) orthogonal discrete wavelet transform (DWT) on FPGA (field programmable gate array). The architecture is based on the pyramid algorithm analysis. It maps spatially the multistage filter banks of the DWT on Xilinx Virtex-e FPGA family using on chip buffering. The architecture takes advantage from the low rate of the high transform stages to reuse the logic. In this paper, we propose an FIR structure to handle the computation along the borders using symmetry extension, a new BlockRam configuration for multi ports shift register, and a mathematical approach to predict and reduce the error dynamic range due to wordlength rounding. For an MxM image size input, our architecture has a period of M/sup 2/ clock cycles, and requires the minimum storage size. The architecture is highly scalable for different filter lengths and number of octaves. The implementation results for a specific 2D Doubechies-4 wavelet transform are included.
Abdsamad Benkrid, Khaled Benkrid, Danny Crookes
FCCM2
2003 A Novel FIR Filter Architecture for Efficient Signal Boundary Handling on Xilinx VIRTEX FPGAs
abstract
FIR (Finite Impulse Response) filters are often used in digital signal processing. This paper presents architecture for FIR filters on Xilinx Virtex FPGAs (field programmable gate arrays). The architecture is particularly useful for handling the problem of signal boundaries filtering, which occurs in finite length signal processing (e.g. image processing). Based on a bit parallel arithmetic, our architecture is fully scalable and parameterized. It cleverly exploits the Shift Register Logic (SRL16) component of the Virtex family. The implementation leads to considerable area savings compared to the conventional implementation (based on a hard router) with no speed penalty. A case study based on the implementation of the standard low filter of the Daubechies-8 wavelet on Xilinx Virtex-E FPGAs is presented.
Abdsamad Benkrid, Khaled Benkrid, Danny Crookes
FCCM2
2003 A logic based approach to hardware abstraction
abstract
This paper presents a novel approach to hardware abstraction based on the logic programming language Prolog. This is an attempt to satisfy the dual requirement of abstract hardware design and hardware efficiency. Central to this approach is a hardware description environment called HIDE, which provides more abstract hardware descriptions and compositions than are possible in traditional hardware description languages such as VHDL or Verilog. HIDE enables highly scaleable and parameterised composition of blocks using a small set of abstract constructors such as the horizontal and vertical constructors for 2D circuit abstractions, and the novel above constructor for 3D circuit abstractions. It also generates pre-placed configurations in EDIF (and VHDL) format for Xilinx FPGAs. The paper presents the syntax and semantics of HIDE and illustrates our logic-based approach in the construction of a high performance Matrix Multiplier core for Xilinx Virtex FPGAs.
Khaled Benkrid, Samir Belkacemi, Danny Crookes
FPGA1
2003 Design framework for the implementation of the 2-D orthogonal discrete wavelet transform on FPGA
abstract
This paper gives a design framework for the implementation of the 2-D Orthogonal Discrete Wavelet Transform (DWT) on FPGA. The architecture is based on the Pyramid Algorithm Analysis. Our architecture spatially maps the multistage filter banks of the DWT onto the Xilinx Virtex-E FPGA family. In this paper we propose a novel FIR structure to handle the computation along the borders using symmetric extension. The paper includes a new detailed mathematical approach to determine the architecture's dynamic range as well as predicting and reducing the error dynamic range due to wordlength rounding. For an NxN image size input, our architecture has a period of N2 clock cycles, and requires only the minimum storage size. The architecture is highly scalable for different filter lengths and number of octaves. The implementation results for a specific 2-D Daubechies-4 Wavelet Transform are included.
Abdsamad Benkrid, Danny Crookes, Khaled Benkrid
FPGA3
2003 A single-FPGA implementation of image connected component labelling
abstract
This paper describes an architecture based on a serial iterative algorithm for Image Connected Component Labelling with a hardware complexity O(N) for an NxN image. The algorithm iteratively scans the input image, performing a recursive non-zero maximum neighbourhood operation. A complete forward pass is followed by an inverse pass in which the image is scanned in reverse order. The process is repeated until no change in the image occurs. The algorithm has been coded in Handel C language and targeted to a Celoxica RC1000-PP PCI board, which is based on a Virtex XCV2000E-6 FPGA. The whole design was fully implemented and tested on real hardware in less than 24 man-hours of work. The implementation speed (pixel throughput) is virtually independent of the image size and is equal to ~34MHz. For 1024x1024 input images, the whole circuit consumes 566 Slices and 5 BlockRAMs and can run at 34 MHz, leading to a 32 pass/sec performance.
Khaled Benkrid, S. Sukhsawas, Danny Crookes, Samir Belkacemi
FPGA1
2003 Design and Implementation of a Novel FIR Filter Architecture with Boundary Handling on Xilinx VIRTEX FPGAs
Abdsamad Benkrid, Khaled Benkrid, Danny Crookes
FPL2
2003 An FPGA-Based Image Connected Component Labeller
Khaled Benkrid, S. Sukhsawas, Danny Crookes, Abdsamad Benkrid
FPL1
2003 A novel approach for diminishing and predicting the error dynamic range in finite wordlength FIR based architectures
abstract
This paper analyses the effects of fixed-point arithmetic in FIR filter based architectures, based on the roundoff statistical noise model. A novel approach, which allows diminishing the error dynamic range and predicting its value according to the wordlength precision, is suggested. This permits the user to preset the fraction precision according to the sought architecture's precision. The efficiency of this approach is demonstrated through the 2D DWT biorthogonal 9&7 transform.
Abdsamad Benkrid, Khaled Benkrid, Danny Crookes
ICASSP (2)2
2002 A Prolog-Based Hardware Development Environment
Khaled Benkrid, Danny Crookes, Abdsamad Benkrid, Samir Belkacemi
FPL1
2002 HIDE: a logic based hardware intelligent description environment
abstract
This paper presents a high-level hardware description environment based on the logic programming language Prolog, called HIDE. The latter has been designed in an attempt to address the problem of abstract hardware design and hardware efficiency. HIDE provides more abstract hardware descriptions and compositions than are possible in traditional hardware description languages such as VHDL or Verilog. It enables highly scaleable and parameterised composition of blocks using a small set of constructors (e.g. horizontal, vertical composition), and generates pre-placed configurations in EDIF format for Xilinx Virtex FPGAs. The paper presents the syntax and semantics of HIDE and illustrates its use in the construction of a bit parallel multiplier core for Xilinx Virtex FPGAs.
Samir Belkacemi, Khaled Benkrid, Danny Crookes
FPT2
2002 A multiplier-less FPGA core for image algebra neighbourhood operations
abstract
This paper presents the design and implementation of a high-level generator of optimised FPGA configurations for Image Algebra (IA) neighbourhood operations. These configurations are parameterised and scaleable in terms of the IA operation itself the window size, the window coefficients, the input pixel word length and the image size. The window coefficients of the neighbourhood operations are represented as sum/subtract of power of twos in Canonical Signed Digit (CSD) representation, which means that the usually costly multiplication operation can be easily implemented using a small number of simple shift-and-add operations, leading to considerable hardware savings. EDIF netlists are generated automatically from high-level descriptions of the IA operations in /spl sim/1 sec. These are specifically optimised for Xilinx XC4000 chips, although implementations for other targets can also be easily realised.
Khaled Benkrid
FPT1
2002 Design and implementation of a novel architecture for symmetric FIR filters with boundary handling on Xilinx Virtex FPGAs
abstract
Symmetric FIR filters, which provide linear phases, are frequently used in digital signal processing. This paper presents the design and implementation of a novel architecture for symmetric FIR filters on Xilinx Virtex FPGAs. The architecture is particularly useful for handling the problem of processing signal boundaries, which occurs in finite length signal processing (e.g. image processing). Based on bit parallel arithmetic, our architecture is fully scalable and parameterised. It takes into account the details of the symmetry and exploits the features of Xilinx Virtex FPGAs. The implementation leads to considerable area savings compared to conventional implementations (based on a hard router), at the expense of using a clock doubler, which reduces the overall processing speed. The latter is however still high enough to achieve real time performance. Moreover, our architecture can match the speed of a conventional implementation if the filter output is going to be decimated, as it is the case in multirate applications (e.g. wavelets).
Abdsamad Benkrid, Khaled Benkrid, Danny Crookes
FPT2
2002 Towards a general framework for FPGA based image processing using hardware skeletons
Khaled Benkrid, Danny Crookes, Abdsamad Benkrid
Parallel Comput.1
2001 Design and Implementation of a Generic 2-D Biorthogonal Discrete Wavelet Transform on an FPGA
Abdsamad Benkrid, Danny Crookes, Khaled Benkrid
FCCM3
2001 High Level Programming for FPGA Based Image and Video Processing Using Hardware Skeletons
Khaled Benkrid, Danny Crookes, Abdsamad Benkrid
FCCM1
2000 High level programming for real time FPGA based video processing
abstract
The inherent reprogrammability of field programmable gate arrays (FPGAs) gives them some of the flexibility of software while keeping the performance advantages of an application specific hardware solution. However, the main disadvantage of FPGAs is the low level of their programming model. Although software tools have been drastically improved since the early days of this new technology, they still require the user to think at the hardware level rather than at the algorithmic level. To bridge the gap between the application and implementation levels, we present a high level software environment for FPGA based real time video processing, which aims to hide hardware details completely from the user. Our approach is to provide a flexible FPGA-based image processing coprocessor with a very high level programming interface based on the core operators of image algebra. Our system has been successfully implemented on VISICOM's VigraVision/sup TM/ PCI board giving real time processing of video data.
Khaled Benkrid, Danny Crookes, Abdsamad Benkrid
ICASSP1
1999 A high level FPGA-based abstract machine for image processing
Ahmed Bouridane, Danny Crookes, Paul Donachy, Khalid Alotaibi, Khaled Benkrid
J. Syst. Archit.5