Habib Mehrez

dblp:56/2953 · DBLP profile ↗
← Back
29ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0002-0692-1754ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 1 since 2021Software engineering, systems software and programming languages · 3Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Reconfigurable computing and FPGAs · 72% Electronic design automation · 19% Hardware accelerators and domain-specific architectures · 4%
Network and information security
1 paper
Cryptographic primitives and cryptanalysis · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA architecture
0.232010
Heterogeneous-ASIF: an application specific inflexible FPGA using heterogeneous logic blocks (abstract only) · FPGA 2010
Configuration tools for a new multilevel hierarchical FPGA · FPGA 2006
A multilevel hierarchical interconnection structure for FPGA · FPGA 2006
Cryptographic primitives and cryptanalysis › authenticated encryption
AES-GCM
0.212014
Towards high performance GHASH for pipelined AES-GCM using FPGAs (abstract only) · FPGA 2014
Cryptographic primitives and cryptanalysis
authenticated encryption
0.212014
Towards high performance GHASH for pipelined AES-GCM using FPGAs (abstract only) · FPGA 2014
Cryptographic primitives and cryptanalysis › authenticated encryption
GHASH
0.212014
Towards high performance GHASH for pipelined AES-GCM using FPGAs (abstract only) · FPGA 2014
Reconfigurable computing and FPGAs › FPGA-based hardware security
FPGA-based cryptographic implementation
0.212014
Towards high performance GHASH for pipelined AES-GCM using FPGAs (abstract only) · FPGA 2014
Reconfigurable computing and FPGAs › multi-FPGA system
inter-FPGA communication
0.212014
Future inter-FPGA communication architecture for multi-FPGA based prototyping (abstract only) · FPGA 2014
Reconfigurable computing and FPGAs › FPGA prototyping
multi-FPGA prototyping
0.212014
Future inter-FPGA communication architecture for multi-FPGA based prototyping (abstract only) · FPGA 2014
Reconfigurable computing and FPGAs › FPGA architecture
heterogeneous logic blocks
0.112010
Heterogeneous-ASIF: an application specific inflexible FPGA using heterogeneous logic blocks (abstract only) · FPGA 2010
Electronic design automation › physical design › placement and routing
FPGA placement and routing
0.122006
Configuration tools for a new multilevel hierarchical FPGA · FPGA 2006
A multilevel hierarchical interconnection structure for FPGA · FPGA 2006
Electronic design automation
physical design
0.122006
Configuration tools for a new multilevel hierarchical FPGA · FPGA 2006
A multilevel hierarchical interconnection structure for FPGA · FPGA 2006
Electronic design automation › physical design › routing
FPGA routing
0.112006
A multilevel hierarchical interconnection structure for FPGA · FPGA 2006
Reconfigurable computing and FPGAs › FPGA architecture
hierarchical FPGA
0.112006
A multilevel hierarchical interconnection structure for FPGA · FPGA 2006
Electronic design automation › physical design › routing › FPGA routing
switch box design
0.112006
A multilevel hierarchical interconnection structure for FPGA · FPGA 2006
Hardware accelerators and domain-specific architectures
cryptographic accelerator
0.112014
Towards high performance GHASH for pipelined AES-GCM using FPGAs (abstract only) · FPGA 2014
Cloud and datacenter computing › resource management › resource multiplexing
time-division multiplexing
0.112014
Future inter-FPGA communication architecture for multi-FPGA based prototyping (abstract only) · FPGA 2014
Interconnection networks and networks-on-chip › interconnect architecture
reconfigurable interconnect
0.012006
Configuration tools for a new multilevel hierarchical FPGA · FPGA 2006

Methods — techniques the papers use, named apart from their topics

pipelining · 0.4karatsuba-ofman algorithm · 0.4time-division multiplexing · 0.2multi-gigabit transceiver · 0.2rent's rule · 0.1pathfinder routing · 0.1partitioning · 0.1clustering · 0.1
YearPublicationVenuePosition
2021 Pre-Silicon Verification Using Multi-FPGA Platforms: A Review
Umer Farooq 0001, Habib Mehrez
J. Electron. Test.2
2016 Using Timing-Driven Inter-FPGA Routing for Multi-FPGA Prototyping Exploration
abstract
Multi-FPGA prototyping, because of its low cost, high speed, and real world testing is quite popular today for pre-silicon verification of increasingly complex designs. In this work, we present a novel exploration flow that is used to analyze and optimize the multi-FPGA based prototyping of complex digital designs. In this flow, an end-to-end experience starting from benchmark generation to optimized inter-FPGA routing is given. For inter-FPGA routing, timing-driven approach is used instead of previously used routability-driven approach. Ten large designs are generated using generic tools of the flow and then effect of number of FPGAs on board, number of inter-FPGA tracks is observed on the performance of generated designs. Extensive experimentation reveals that FPGA board with six FPGAs gives best system frequency results. Furthermore, execution time comparison between routability and timing-driven approach reveals that compared to routability-driven approach, timing-driven approach consumes, on average, 46% less time while giving same or better frequency results.
Umer Farooq 0001, Roselyne Chotin-Avot, Muhammad Moazam Azeem, Zouha Cherif, Maminionja Ravoson, Saqib Khan, Habib Mehrez
DSD7
2016 Exploration of Mesh-Based FPGA Architecture: Comparison of 2D and 3D Technologies in Terms of Power, Area and Performance
abstract
In this paper, we propose a 2D and 3D interconnect network based on a Mesh-of-Clusters (MoC) topology for the implementation of an efficient Field Programmable Gate Arrays (FPGA) architecture. Proposed MoC-based FPGA architecture presents a new hierarchical Switch Box (SBs) and depopulated intra-cluster interconnect based on the Butterfly-Fat-Tree (BFT) topology. Long routing wires which span multiple SBs in every row and column were used in order to improve performance. By adjusting the percentage of long wire and span, we can design and build 3D high density MoC-based FPGA. To design 3D MoC-based FPGAs, we cut the 2D FPGA into two equal FPGA dies and we adjust the long wire span factor to connect the two dies. Then, these long wire segments are converted as 3D through silicon via (TSV). We present also a design methodology and CAD tools to explore the performance of proposed 2D and 3D MoC-based FPGA architectures in term of power, energy, area and delay. Experimental results with large benchmarks show that with 3D MoC-based FPGA the average gains in terms frequency, energy and area are 23%, 37% and 47% respectively, compared to 2D MoC-based FPGA.
Sonda Chtourou, Zied Marrakchi, Emna Amouri, Vinod Pangracious, Habib Mehrez, Mohamed Abid
PDP5
2016 Inter-FPGA routing environment for performance exploration of multi-FPGA systems
abstract
Multi-FPGA platforms are a popular choice today for complex system prototyping because they offer high execution speed, low cost, and real world testing experience. However, performance of multi-FPGA based systems is severely affected by widening logic to I/O gap in FPGAs. In order to address the performance issue, in this work, we propose an exploration and optimization flow for multi-FPGA based prototyping that gives an end-to-end experience starting from benchmark generation to optimized inter-FPGA routing. Using generic tools of the flow, ten large benchmarks are generated. Then, through a generic novel inter-FPGA routing environment, effect of variation of number of FPGAs as well as number of inter-FPGA tracks on the performance of a target design is explored. For performance exploration and optimization, five different FPGA boards are utilized where number of FPGAs on board are varied from two to six. Moreover, for each board four different inter-FPGA track combinations are used. Experimental results reveal that multi-FPGA boards with inter-FPGA tracks corresponding optimally to the cut net requirements of benchmarks under consideration give best frequency results. Furthermore, frequency comparison between different boards shows that FPGA board with six FPGAs gives, on average, best frequency results. Finally, we also perform frequency-price analysis which shows that board with four FPGAs gives better frequency-price tradeoff as compared to other FPGA boards under consideration.
Umer Farooq 0001, Roselyne Chotin-Avot, Muhammad Moazam Azeem, Maminionja Ravoson, Mariem Turki, Habib Mehrez
RSP6
2014 Performance Comparison between Multi-FPGA Prototyping Platforms: Hardwired Off-the-Shelf, Cabling, and Custom
abstract
We can classify multi-FPGA prototyping platforms in three categories: hardwired off-the-shelf, cabling and custom. Three points are developed in this paper. Firstly, an automatic design flow is proposed to generate a cabling platform and a custom platform for a given design. Then, the optimal width of cables for a cabling multi-FPGA platform is explored. Finally, the performances of these three multi-FPGA platforms are compared. The results show that the cabling platform achieves up to 82% gain in performance, and the custom platform achieves up to 100%, compared to the hardwired off-the-shelf platform. The custom platform achieves up to 20% gain in performance over the cabling platform. Therefore the results show that, apart from some stringent constraints (such as deployment cost or specific frequency needed), the relatively new cabling paradigm with the proposed automatic, inter-FPGA tracks distribution tool, offers an attractive alternative compared to the two other platforms.
Qingshan Tang, Matthieu Tuna, Habib Mehrez
FCCM3
2014 Towards high performance GHASH for pipelined AES-GCM using FPGAs (abstract only)
abstract
AES-GCM has been utilized in various security applications. It consists of two components: an Advanced Encryption Standard (AES) engine and a Galois Hash (GHASH) core. The performance of the system is determined by the GHASH architecture because of the inherent computation feedback. This paper introduces a modification for the pipelined Karatsuba Ofman Algorithm (KOA)-based GHASH. In particular, the computation feedback is removed by analyzing the complexity of the computation process. The proposed GHASH core is evaluated with three different implementations of AES ( BRAMs-based SubBytes, composite field-based SubBytes, and LUT-based SubBytes). The presented AES-GCM architectures are implemented using Xilinx Virtex5 FPGAs. Our comparison to previous work reveals that our architectures are more performance-efficient (Thr. /Slices).
Karim M. Abdellatif, Roselyne Chotin-Avot, Zied Marrakchi, Habib Mehrez, Qingshan Tang
FPGA4
2014 Future inter-FPGA communication architecture for multi-FPGA based prototyping (abstract only)
abstract
Multi-FPGA boards are widely used for rapid system prototyping. Even though the prototyping is trying to reach the maximum performance, the performance is limited by the inter-FPGA communication. As the capacity per I/O for each FPGA generation is increasing, FPGA I/Os are becoming a scarce resource. The design is divided into several parts, each part's capacity fits in a single FPGA. Signals crossing design's parts located in different FPGAs are called cut nets. In order to resolve pin limitation problem, cut nets are sent between FPGAs in pipelined way using the Time-Division-Multiplexing technique. The maximum number of cut nets passing through one FPGA I/O is called the TDM ratio. There are two multiplexing architectures used for multi-FPGA based prototyping: Logic Multiplexing and ISERDES/OSERDES. In this paper, a new multiplexing architecture Multi-Gigabit Transceiver (MGT) is proposed. Experiments are done in a multi-FPGA board with the testbench LFSR to validate the achieved performance. Assume that all the FPGA I/Os used for inter-FPGA communication are MGT capable in the future. Analyses show that the proposed multiplexing architecture can achieve higher performance when the TDM ratio exceeds 67. The gain in performance of the proposed architecture over the existing architecture augments as the TDM ratio increases.
Qingshan Tang, Matthieu Tuna, Habib Mehrez
FPGA3
2014 Balancing WDDL dual-rail logic in a tree-based FPGA to enhance physical security
abstract
The Tree-based FPGA offers better density and timing determinism than traditional mesh-based FPGA. Moreover, thanks to its multilevel structure, it offers greater easiness to balance dual signals in terms of routing resources number. In this paper, we study the use of the Wave Dynamic Differential Logic (WDDL) on a custom tree-based FPGA of 2048 cells. The WDDL technique offers an effective way to withstand Differential Power Attacks (DPA). However, the effectiveness of this countermeasure is guaranteed provided a symmetry is maintained between the routing of both the direct and complementary paths, which is very hard to achieve in FPGA. Thus, balancing aware Computer-Aided Design (CAD) tools must be developed. In this work, we propose first adjacent placement and balancing-aware routing techniques for tree-based FPGA to counter the routing unbalance. Then side channel analyses are performed on FPGA circuit implementing PRESENT crypto-processor. Experimental results show that the balancing methods enhance the design security against side channel attacks.
Emna Amouri, Shivam Bhasin, Yves Mathieu, Tarik Graba, Jean-Luc Danger, Habib Mehrez
FPL6
2014 Improve defect tolerance in a cluster of a SRAM-based Mesh of Cluster FPGA using hardware redundancy
abstract
The technological evolution involves a higher number of physical defects in circuits after manufacturing. One of the future challenge is to find a way to use a maximum of defected manufactured circuits. In this paper, multiple techniques are proposed to avoid defects in the cluster local interconnect of a SRAM-based Mesh of Clusters FPGA. Using defect tolerance, area and timing metrics, two previous hardware redundancy strategies are evaluated on the Mesh of Clusters architecture : Fine Grain Redundancy (FGR) and Improved Fine Grain Redundancy (IFGR). We show that using these techniques on a cluster of a Mesh of Clusters architecture permits to tolerate 8 times more defects than on an industrial Mesh FPGA with a low area overhead (-6% for FGR and 22% for IFGR) and a low increase of Critical Path Delay (CPD)(6% for FGR and 2% for IFGR). We also proposed three new redundancy strategies using spare resources : Distributed Feedbacks (DF) for crossbar down, Adapted Fine Grain Redundancy (AFGR) to avoid defective multiplexers and Upward Redundant Multiplexer (URM) for the crossbar up. Compared to the Mesh of Clusters architecture without defect tolerance techniques, the best trade off between defect tolerance (36.4%), area overhead (11.56%) and CPD (+7.46%) is obtained using AFGR. Using the other methods permits to considerably limit the area overhead (10.4% with URM) with a lesser number of defective elements bypassed (18% max).
Adrien Blanchardon, Roselyne Chotin-Avot, Habib Mehrez, Emna Amouri
FPL3
2014 Low-power comb decimation filter for RF Sigma-Delta ADCs
abstract
An efficient multi-rate multi-stage architecture for the Comb decimation filter of Sigma-Delta ADCs is presented. Polyphase decomposition in all stages is used to reduce the operating frequency of the Comb filter. A systematic design procedure is developed in order to generate all possible combinations for the decimation factor of each stage. A third order Comb decimation filter with a total decimation factor of 16 is taken as a design example. The eight possible architectures are generated in two different CMOS processes. The performance of the generated architectures are compared in terms of power consumption, area and maximum operating frequency.
Alp Kiliç, Delaram Haghighitalab, Habib Mehrez, Hassan Aboushady
ISCAS3
2013 A defect-tolerant cluster in a mesh SRAM-based FPGA
abstract
In this paper, we propose the implementation of multiple defect-tolerant techniques on an SRAM-based FPGA. These techniques include redundancy at both the logic block and intra-cluster interconnect. In the logic block, redundancy is implemented at the multiplexer level. Its efficiency is analyzed by injecting a single defect at the output of a multiplexer, considering all possible locations and input combinations. While at the interconnect level, fine grain redundancy is introduced which not only bypasses defects but also increases routability. Taking advantage of the sparse intra-cluster interconnect structures, routability is further improved by efficient distribution of feedback paths allowing more flexibility in the connections among logic blocks. Emulation results show a significant improvement of about 15% and 34% in the robustness of logic block and intra-cluster interconnect respectively. Furthermore, the impact of these hardening schemes on the testability of the FPGA cluster for manufacturing defects is also investigated in terms of maximum achievable fault coverage and the respective cost.
Arwa Ben Dhia, Saif-Ur Rehman, Adrien Blanchardon, Lirida A. B. Naviner, Mounir Benabdenbi, Roselyne Chotin-Avot, Emna Amouri, Habib Mehrez, Zied Marrakchi
FPT8
2013 Design and optimization of heterogeneous tree-based FPGA using 3D technology
abstract
The CMOS technology scaling has greatly improved the overall performance and density of Field Programmable Gate Arrays (FPGAs). However, when looking at the performance metrics such as speed, area and power consumption, the gap is generally very wide for FPGAs compared to application specific integrated circuits (ASICs) mainly due to the programmable interconnect overhead. We propose a 3-dimensional (3D) design methodology using horizontal design partitioning to vertically stack heterogeneous FPGA designs based on a Tree-based multilevel FPGA architecture. We describe the 3D design and optimization methodology to improve speed, interconnect area and power consumption using Tezzaron's 3D stacking technology.
Vinod Pangracious, Zied Marrakchi, Habib Mehrez
FPT3
2013 Physical design exploration of 3D tree-based FPGA architecture
abstract
An innovative 3D physical design exploration methodology for Tree-based FPGA architecture is presented in this paper. In a Tree-based FPGA architecture, the interconnects are arranged in a multidimensional network with the logic unites and switch blocks placed at different levels, using a Butterfly-Fat Tree network topology. A 3D physical design exploration methodology leverage on Through Silicon Via (TSVs) using a horizontal break-point to re-distribute the Tree interconnects into multiple stacked active silicon layers proposed in this paper.
Vinod Pangracious, Emna Amouri, Habib Mehrez, Zied Marrakchi
ACM Great Lakes Symposium on VLSI3
2013 Routing algorithm for multi-FPGA based systems using multi-point physical tracks
abstract
Multi-FPGA boards suffer from large timing delays in inter-FPGA physical tracks compared to intra-FPGA track delays, as well as a limited bandwidth between FPGAs due to the limited number of I/Os per FPGA. In order to tackle this problem, an algorithm which routes multi-terminal nets in multi-point tracks is proposed in this paper to spare FPGA I/Os. Experiments are conducted using Gaisler Research Benchmarks. Firstly, each testbench will be implemented in an off-the-shelf board. The results show that the system frequency can be increased in the off-the-shelf board by the proposed routing algorithm. Secondly, an automatic design flow which generates a custom multi-FPGA board is enhanced by generating multi-point tracks in the board, and each testbench will be implemented with the proposed routing algorithm in custom boards. The results show that the system frequency is improved in the custom board with both 2- and multi-point tracks.
Qingshan Tang, Matthieu Tuna, Habib Mehrez
RSP3
2013 Exploring redundant arithmetics in computer-aided design of arithmetic datapaths
Sophie Dupuis, Roselyne Chotin-Avot, Habib Mehrez
Integr.3
2012 Design for prototyping of a parameterizable cluster-based Multi-Core System-on-Chip on a multi-FPGA board
abstract
Reaching a physical limitation in terms of power consumption, the new chosen axis to increase the performance is to scale the number of processing elements (PEs) instead of increasing the speed of processors. These complex systems often take the form of a Multi-Core System-on-Chip (MCSoC) in which individual nodes are connected using a Network-on-Chip (NoC). FPGA-based prototyping is no longer optional to test these very large designs. A multi-FPGA prototyping board, which is a collection of FPGAs, is used when the logic capacity of a single FPGA is insufficient. Nevertheless, multi-FPGA boards come with some constraints, which lower the system frequency. However, cluster-based MCSoCs have a great value: their flexibility in terms of form factor. In this paper, a parameterizable cluster-based MCSoC with a mesh topology has been designed. An indicator based on the average Rent's Rule of cluster-based MCSoCs has been proposed to predict the best form factor in order to obtain the maximum prototyping system frequency. The targeted prototyping platform is a six Virtex-5 multi-FPGA board. The results show that for a given total number of PEs in cluster-based MCSoCs, a lower indicator can result in a higher system frequency.
Qingshan Tang, Habib Mehrez, Matthieu Tuna
RSP2
2011 Application-Specific FPGA using heterogeneous logic blocks
abstract
This work presents a new automatic mechanism to explore the solution space between Field Programmable Gate Arrays (FPGAs) and Application-Specific Integrated Circuits (ASICs). This new solution is termed as an Application-Specific Inflexible FPGA (ASIF) [Parvez et al. 2009]. An ASIF can be considered as an FPGA with reduced flexibility, or as a reconfigurable ASIC that can implement a set of application circuits which will operate at mutually exclusive times. Execution of different application circuits can be switched by loading their respective bitstream on an ASIF. An ASIF that is reduced from a heterogeneous FPGA is termed as a heterogeneous ASIF. It is shown that a standard-cell-based heterogeneous ASIF for a set of 10 opencore application circuits is 9.6 times smaller than a single-driver mesh-based heterogeneous FPGA. The area gap between ASIC and ASIF is not too significant; however, it can be reduced by designing repeatedly used components of ASIF in full-custom. Unlike an ASIC, an ASIF is a reprogrammable device that can be used to reprogram new or modified circuits at a limited scale.
Husain Parvez, Zied Marrakchi, Alp Kiliç, Habib Mehrez
ACM Trans. Reconfigurable Technol. Syst.4
2010 Heterogeneous-ASIF: an application specific inflexible FPGA using heterogeneous logic blocks (abstract only)
abstract
International audience
Husain Parvez, Zied Marrakchi, Habib Mehrez
FPGA3
2009 ASIF: Application Specific Inflexible FPGA
abstract
An application specific inflexible FPGA (ASIF) is an FPGA with reduced flexibility that can implement a set of application circuits which will operate at mutually exclusive times. These circuits are efficiently placed and routed on an FPGA to minimize the total routing switches required by the architecture. Later all the unused routing switches are removed from the FPGA to generate an ASIF. An ASIF for a set of 17 MCNC benchmark circuits is found to be 5.43 times (81.5%) smaller than a mesh-based unidirectional FPGA required to map any of these circuits.
Husain Parvez, Zied Marrakchi, Habib Mehrez
FPT3
2008 A new coarse-grained FPGA architecture exploration environment
abstract
This paper presents an exploration environment for the design of 2D island-style coarse grained FPGA architectures. An architecture description file defines various architectural parameters including the definition of new coarse grained blocks, the positioning of blocks in the architecture and the selection of routing network. Once the initial architecture is defined, a software flow places and routes a target netlist on the generated architecture. The placement cost of a netlist is optimized either by changing the position of netlist instances on its respective blocks or by changing the position of blocks on the architecture. A single FPGA architecture can also be obtained for mapping a set of netlists at mutually exclusive times. It has been found that the sum of the placement costs of all the netlists is found to be minimum if all the netlists are used to get a single architecture. A set of DSP test-benches is used to show the effectiveness of the various techniques used in this work.
Husain Parvez, Zied Marrakchi, Umer Farooq 0001, Habib Mehrez
FPT4
2008 Efficient tree topology for FPGA interconnect network
abstract
This paper presents an improved Tree-based architecture that unifies two unidirectional programmable networks: A predictible downward network based on the Butter y-Fat-Tree topology, and an upward network using hierarchy. Studies based on Rent's Rule show that switch requirements in this architecture grow slower than in traditional Mesh topologies. New tools are developed to place and route several benchmark circuits on this architecture. Experimental results show that the Tree-based architecture can implement MCNC benchmark circuits with an average gain of 54% in total area compared with Mesh architecture.
Zied Marrakchi, Hayder Mrabet, Emna Amouri, Habib Mehrez
ACM Great Lakes Symposium on VLSI4
2007 Efficient Mesh of Tree Interconnect for FPGA Architecture
abstract
In this paper we present a new mesh of tree FPGA architecture, where clusters are surrounded by a mesh style interconnect and each cluster local interconnect is equivalent to a depopulated tree-based topology. The particularity of the architecture allows to retain the distinction between mesh and tree levels in the mapping phase. This has an important impact on run time saving and tool simplification. Nevertheless an efficient interconnect distribution must be found between both levels, to reach a tradeoff between interconnect reduction and routability. With the proposed Mesh of Tree architecture, we divided the required run time by 3 and reduced the routing interconnect by 24%, compared to the clustered VPR-style mesh architecture.
Zied Marrakchi, Hayder Mrabet, Christian Masson, Habib Mehrez
FPT4
2007 Mesh of Tree: Unifying Mesh and MFPGA for Better Device Performances
abstract
In this paper we present a new clustered mesh FPGA architecture where each cluster local interconnect is implemented as an MFPGA tree network. Unlike previous clustered mesh architectures, the mesh of tree allows us to consider large clusters sizes (thanks to MFPGA depopulated local interconnect). Experimentation shows that we obtain a reduction of 14% in switches number and 2 times in the placement and routing run time. Furthermore, compared to MFPGA, the mesh of tree achieves full mutability of all MCNC benchmarks since we can easily control both clusters LUTs occupation and mesh channel width
Zied Marrakchi, Hayder Mrabet, Christian Masson, Habib Mehrez
NOCS4
2006 A multilevel hierarchical interconnection structure for FPGA
abstract
Creation of large FPGAs needs radical efficient changes in architecture to improve speed, density and software mapping time. Based on industry experience with standard ASICs, we believe that partitioning and hierarchy become an obligation for FPGA hardware and software developments. As an alternative we propose a new Multilevel hierarchical FPGA (MFPGA) architecture where logic blocks and routing resources are sparsely partitioned into multilevel clustered structure. Since the routing resources consume most of the FPGA area, we focus on interconnect check. We try to achieve the best area efficiency by balancing interconnect and logic block utilization. The proposed MFPGA interconnect unifies two unidirectional programmable networks: A downward network based on the Butterfly-Fat-Tree topology, and an upward network that uses hierarchy. The Downward network uses linear populated and unidirectional switch boxes and gives one path from each wire-source in the top to each leaf (logic block) in the lowest level. The upward network connects the logic blocks outputs and the input pads to the different levels of the downward network. Studies based on the Rent's Rule show that wiring and switch requirements in the MFPGA grow slower than in traditional topologies. We used MCNC benchmark circuits to compare the switch and area requirements between our MFPGA architecture and the traditional mesh topology. New software tools for placement and routing were developed to conduct this study on the MFPGA architecture. Expermimental results show that MFPGA can implement circuits with fewer switches and a smaller total area than mesh architecture.
Hayder Mrabet, Zied Marrakchi, Pierre Souillot, Habib Mehrez
FPGA4
2006 Configuration tools for a new multilevel hierarchical FPGA
abstract
In this paper we evaluate a new multi-level hierarchical MFPGA. The specific architecture includes two unidirectional programmable networks: A downward network based on the Butterfly-Fat-Tree topology and a special hierarchical upward network. The Downward network uses linear populated and unidirectional Switch Boxes (SBs) and gives one path from each wire-source in the top to reach a leaf (Logic Block: LB) in the lowest level. The upward network connects the LBs output and the input Pads to the SBs situated in different levels of the downward network. New tools are developed to program the new architecture. The global placement approach uses a combination of clustering and partitioning with adaptations to deal with the multi-level interconnect topology. First we run a multi-level bottom-up clustering to reduce external connections. Second we run a multi-level top-down refinement to reduce signals bandwidth of clusters in each level. A detailed placer defines the position of each LB inside a cluster and considers more complex routing constraints. The router is an adaptation of Pathfinder. The global routing consists on selecting the level to use. Signals routing is immediate since path to reach a destination is predictable and unique. Results are based on the MCNC benchmarks and they quantify the LB occupancy and routability. Comparison with the traditional symmetric Manhattan mesh architecture shows that MFPGA can implement circuits with fewer switches and a smaller total area.
Zied Marrakchi, Hayder Mrabet, Habib Mehrez
FPGA3
2006 Performances improvement of FPGA using novel multilevel hierarchical interconnection structure
abstract
This paper presents a new Multilevel hierarchical FPGA (MFPGA) architecture that unifies two unidirectional programmable networks: A predictible downward network based on the Butterfly-Fat-Tree topology, and an upward network using hierarchy. Studies based on the Rent's Rule show that wiring and switch requirements in the MFPGA grow slower than in traditional topologies. New tools are developed to place and route several benchmark circuits on this architecture. Experimental results based on the MCNC benchmarks show that MFPGA can implement circuits with an average gain of 40% in total area compared with mesh architecture.
Hayder Mrabet, Zied Marrakchi, Pierre Souillot, Habib Mehrez
ICCAD4
2000 A family of redundant multipliers dedicated to fast computation for signal processing
abstract
In view of the performance achieved through the use of the redundant addition, it would appear interesting to generalize the redundant notations (Carry Save, Borrow Save). To achieve this we require, as well the addition, a multiplication satisfying these notations. This paper presents the design of a set of multipliers spanning all possible I/O combinations, in redundant and conventional notations. We also describe the associated architectures and details of our method, which is based on parameterizable IP cores. The functions developed offer superior performance over conventional multipliers.
Yannick Dumonteix, Habib Mehrez
ISCAS2
1998 On portable macrocell FPU generators for division and square root operators complying to the full IEEE-754 standard
abstract
In this paper, we investigate the design of macrocell generators of division and square root floating-point operators. The number representation used in our operators is the IEEE-754-1985 standard for binary floating-point numbers. The design and implementation of the generators rely on a powerful multi-view macroblock generator tool called GenOptim. This computer-aided design (CAD) tool is able to output a set of different descriptions for several VLSI technologies as well as field programmable gate arrays (FPGAs). The division and square root operators described in this paper use the signed-binary-digit representation. We start first by describing the operators for the significand, then we investigate the IEEE floating-point operators. Throughout this paper, and wherever appropriate, we present the implementation results using the GenOptim environment.
Mourad Aberbour, A. Houelle, Habib Mehrez, Nicolas Vaucher, Alain Guyot
IEEE Trans. Very Large Scale Integr. Syst.3
1995 Application of fast layout synthesis environment to dividers evaluation
abstract
Experience has shown that generator programs are quite often written by VLSI designers, as they hold the empirical knowledge better than anyone. However, their ability does not necessarily include programming and debugging skills: these designers have to focus on the problem at hand not on the tools or the language they use to solve it. GenOptim has been created to quickly design efficient IEEE 754 floating-point macro-cell generators that do not rely on particular target technologies. Whereas the design of fast and efficient adders, multipliers and shifters is well understood division and square root remain a serious design challenge. GenOptim was used to quickly evaluate new divider architectures.>
A. Houelle, Habib Mehrez, Nicolas Vaucher, Luis A. Montalvo, Alain Guyot
IEEE Symposium on Computer Arithmetic2