Fayez Gebali

dblp:69/4864 · DBLP profile ↗
← Back
30ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0001-5189-3409ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 2 first-authorComputer networks · 12 · 1 first-authorArtificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 33% Hardware accelerators and domain-specific architectures · 25% Electronic design automation · 18%
Computer networks
2 papers
Network performance modeling · 61% Optical networks · 30% Internet architecture and protocols · 9%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
array processor
0.532017
Design Space Exploration of 2-D Processor Array Architectures for Similarity Distance Computation · IEEE Trans. Parallel Distributed Syst. 2017
Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm · IEEE Trans. Parallel Distributed Syst. 2011
Processor Array Architectures for Deep Packet Classification · IEEE Trans. Parallel Distributed Syst. 2006
Electronic design automation
design space exploration
0.312017
Design Space Exploration of 2-D Processor Array Architectures for Similarity Distance Computation · IEEE Trans. Parallel Distributed Syst. 2017
Hardware accelerators and domain-specific architectures
systolic array
0.312017
Design Space Exploration of 2-D Processor Array Architectures for Similarity Distance Computation · IEEE Trans. Parallel Distributed Syst. 2017
Hardware accelerators and domain-specific architectures
cryptographic accelerator
0.112011
Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm · IEEE Trans. Parallel Distributed Syst. 2011
Integrated circuit design › digital circuit design › arithmetic circuit design
modular multiplication
0.112011
Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm · IEEE Trans. Parallel Distributed Syst. 2011
Integrated circuit design › digital circuit design › arithmetic circuit design › modular multiplication
montgomery multiplication
0.112011
Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm · IEEE Trans. Parallel Distributed Syst. 2011
Network performance modeling › loss systems
blocking probability
0.112006
A New Analytical Model for Computing Blocking Probability in Optical Burst Switching Networks · IEEE J. Sel. Areas Commun. 2006
Network performance modeling
markov chain model
0.112006
A New Analytical Model for Computing Blocking Probability in Optical Burst Switching Networks · IEEE J. Sel. Areas Commun. 2006
Optical networks › optical switching
optical burst switching
0.112006
A New Analytical Model for Computing Blocking Probability in Optical Burst Switching Networks · IEEE J. Sel. Areas Commun. 2006
Parallel and multicore computing › parallel algorithms
string matching
0.112006
Processor Array Architectures for Deep Packet Classification · IEEE Trans. Parallel Distributed Syst. 2006
Processor architecture and microarchitecture › parallel computer organization
systolic array design
0.112006
Processor Array Architectures for Deep Packet Classification · IEEE Trans. Parallel Distributed Syst. 2006
Energy-efficient computing › dynamic power reduction
glitch reduction
0.012011
Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm · IEEE Trans. Parallel Distributed Syst. 2011
Energy-efficient computing
low-power design
0.012011
Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm · IEEE Trans. Parallel Distributed Syst. 2011
Internet architecture and protocols › packet processing
packet classification
0.012006
Processor Array Architectures for Deep Packet Classification · IEEE Trans. Parallel Distributed Syst. 2006

Methods — techniques the papers use, named apart from their topics

linear scheduling · 0.63-d computation domain · 0.6affine scheduling · 0.2data dependence graph · 0.1regular iterative expression · 0.1markov chain · 0.1discrete-event simulation · 0.1
YearPublicationVenuePosition
2020 Animal Species Recognition Using Deep Learning
Mai Ibraheam, Fayez Gebali, Kin Fun Li, Leonard Sielecki
AINA2
2020 A situation refinement model for complex event processing
Alaa Alakari, Kin Fun Li, Fayez Gebali
Knowl. Based Syst.3
2020 Performance Analysis of Multiuser FSO/RF Network Under Non-Equal Priority With $P$ -Persistence Protocol
abstract
This paper presents and analyzes a novel multiuser network based on hybrid free-space optical (FSO)/radiofrequency (RF) transmission system, where every user is serviced by a primary FSO link. When more than one FSO link fail, the central node services these corresponding users of non-equal priority by using a common backup RF link according to a p persistence servicing protocol. A novel discrete-time Markov chain model is developed for the proposed network, where different transmission rates over RF and FSO links are assumed. We investigate the throughput from central node to the user, the average size of the transmit buffer allocated for every user, the frame queuing delay in the transmit buffer, the efficiency of the queuing system, the frame loss probability, and the RF link utilization. Numerical examples show that transmitting a data frame with probability p when using the common backup RF link achieves considerable network performance improvement while ensuring high-priority users enjoy better performance. Meanwhile, the performance of low-priority users approaches the performance when using equal priority protocol to serve all the remote users.
Tamer Rakia, Fayez Gebali, Hong-Chuan Yang, Mohamed-Slim Alouini
IEEE Trans. Wirel. Commun.2
2019 Thermal-aware network-on-chips: Single- and cross-layered approaches
Mostafa Said, Ahmed Shalaby 0001, Fayez Gebali
Future Gener. Comput. Syst.3
2019 Hardware Trojan Detection Using Reconfigurable Assertion Checkers
abstract
In this paper, reconfigurable assertion checkers (RACs) are proposed for hardware Trojan detection on a system-on-chip (SoC). A modified circuit design flow to incorporate RACs into an SoC is proposed. A flexible RAC architecture is developed to implement different types of assertions. Case studies of mapping a simple assertion checker into RAC are given. The utilization of RAC in detecting hardware Trojans is demonstrated in two advanced microcontroller bus architecture (AMBA) case studies: a denial-of-service (DoS) hardware Trojan and a leak of information hardware Trojan. In both cases, our proposed RAC was able to detect the hardware Trojan whenever the Trojan was triggered.
Uthman Alsaiari, Fayez Gebali
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Throughput analysis of point-to-multi-point hybric FSO/RF network
abstract
This paper presents and analyzes a point-to-multi-point (P2MP) network that uses a number of free-space optical (FSO) links for data transmission from the central node to the different remote nodes. A common backup radio-frequency (RF) link is used by the central node for data transmission to any remote node in case of the failure of any one of FSO links. We develop a cross-layer Markov chain model to study the throughput from central node to a tagged remote node. Numerical examples are presented to compare the performance of the proposed P2MP hybrid FSO/RF network with that of a P2MP FSO-only network and show that the P2MP Hybrid FSO/RF network achieves considerable performance improvement over the P2MP FSO-only network.
Tamer Rakia, Fayez Gebali, Hong-Chuan Yang, Mohamed-Slim Alouini
ICC2
2017 Design Space Exploration of 2-D Processor Array Architectures for Similarity Distance Computation
abstract
We present a systematic methodology for exploring the design space of similarity distance computation in machine learning algorithms. Previous architectures proposed in the literature have been obtained using ad hoc techniques that do not allow for design space exploration. The size and dimensionality of the input datasets have not been taken into consideration in previous works. This may result in impractical designs that are not amenable for hardware implementation. The methodology presented in this work is used to obtain the 3-D computation domain of the similarity distance computation algorithm. A scheduling function determines whether an algorithm variable is pipelined or broadcast. Four linear scheduling functions are presented, and six possible 2-D processor array architectures are obtained and classified based on the size and dimensionality of the input datasets. The obtained designs are analyzed in terms of speed and area, and compared with previously obtained designs. The proposed designs achieve better time and area complexities.
Awos Kanan, Fayez Gebali, Atef Ibrahim
IEEE Trans. Parallel Distributed Syst.2
2016 Hardware Covert Attacks and Countermeasures
abstract
Computing platforms deployed in many critical infrastructures, such as smart grid, financial systems, governmental organizations etc., are subjected to security attacks and potentially devastating consequences. Computing platforms often get attacked 'physically' by an intruder accessing stored information, studying the internal structure of the hardware or injecting a fault. Even if the attackers fail to gain sensitive information stored in hardware, they may be able to disrupt the hardware or deny service leading to other kinds of security failures in the system. Hardware attacks could be covert or overt, based on the awareness of the intended system. This work classifies existing hardware attacks. Focusing mainly on covert attacks, they are quantified using a proposed schema. Different countermeasure techniques are proposed to prevent such attacks.
Jahnabi Phukan, Kin Fun Li, Fayez Gebali
AINA3
2015 Efficient Scalable Serial Multiplier Over GF(2m) Based on Trinomial
abstract
This brief presents a novel low-complexity scalable serial architecture for finite field multiplication over GF(2m) based on irreducible trinomial. This architecture was explored by applying nonlinear technique that allows the designer, using progressive product reduction technique, to control the workload per processor and also allows the communication overhead between processors to be reduced. By comparing the ASIC implementation of the proposed structure to some of the previously published structures, the proposed structure have at least 71.7% lower area and at least 89.9% lower power compared with most of them. This makes the proposed design more suitable for constrained implementations of cryptographic primitives in resource constrained applications, such as smart cards, handheld devices, and implantable medical devices.
Fayez Gebali, Atef Ibrahim
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Systolic Array Architectures for Sunar-Koç Optimal Normal Basis Type II Multiplier
abstract
We present linear and nonlinear techniques for design exploration of an iterative algorithm. The nonlinear techniques allow control of processor workload and control of communication between processors. The algorithm considered is the Sunar-Koç optimal normal basis type II multiplication algorithm. Six systolic arrays are obtained. General formulas are provided for each design so that the operation of the system can be determined for a given GF(2m). The proposed architectures have been implemented using 45-nm CMOS technology and compared with published architectures. The results show that the proposed designs have at least 44.4% lower total computation time compared with the designs of all bit serial multipliers, while having slightly larger area delay product (ADP), up to 19.1%, compared with some of the bit serial multipliers and having smaller ADP values compared with most of the digit serial ones. Moreover, they have at least 46% lower power delay product compared with all bit serial and digit serial multipliers.
Atef Ibrahim, Fayez Gebali, Turki F. Al-Somani
IEEE Trans. Very Large Scale Integr. Syst.2
2014 Reliability analysis and fault tolerance for hypercube multi-computer networks
Mostafa I. H. Abd-El-Barr, Fayez Gebali
Inf. Sci.2
2011 Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm
abstract
This paper presents a systematic methodology for exploring possible processor arrays of scalable radix 4 modular Montgomery multiplication algorithm. In this methodology, the algorithm is first expressed as a regular iterative expression, then the algorithm data dependence graph and a suitable affine scheduling function are obtained. Four possible processor arrays are obtained and analyzed in terms of speed, area, and power consumption. To reduce power consumption, we applied low power techniques for reducing the glitches and the Expected Switching Activity (ESA) of high fan-out signals in our processor array architectures. The resulting processor arrays are compared to other efficient ones in terms of area, speed, and power consumption.
Atef Ibrahim, Fayez Gebali, Hamed Elsimary, Amin M. Nassar
IEEE Trans. Parallel Distributed Syst.2
2010 Networks-on-chip topology optimization subject to power, delay, and reliability constraints
abstract
In this paper, we present a novel approach in Networks-on-Chip topology optimization, by considering the network power consumption, packet transmission delay, and system reliability, simultaneously. We use the Particle Swarm Optimization technique to acquire the most suitable topology architecture, which achieves maximum reliability as well as minimum delay and power consumption. The optimization problem, which considers six design variables: network topology architecture, traffic distribution, processing elements' mapping, noise power, voltage swing, and probability of edge failure, is validated through a case study of an H.263-encoder MP3-decoder.
Haytham Elmiligi, Ahmed A. Morgan, M. Watheq El-Kharashi, Fayez Gebali
ISCAS4
2010 Multi-objective optimization for Networks-on-Chip architectures using Genetic Algorithms
abstract
Networks-on-Chip (NoC) architecture design faces a trade-off between different conflicting metrics. In this paper, we target one aspect of this trade-off: area versus average delay. The NoC architecture generation is formulated as a two-objective optimization problem and a Genetic Algorithm (GA)-based technique is used to solve it. According to the application requirements and the design constraints, the optimization process could be controlled by the designer by specifying weight factors for area and delay. As a proof of concept, our technique is applied to three real applications with different number of cores. Results show that the proposed solution is a promising way to achieve the best architecture with respect to both area and delay.
Ahmed A. Morgan, Haytham Elmiligi, M. Watheq El-Kharashi, Fayez Gebali
ISCAS4
2010 Enhanced Busy-Tone-Assisted MAC Protocol for Wireless Ad Hoc Networks
abstract
In wireless multihop ad hoc networks, the hidden terminal problem severely degrades the overall performance. On the other hand, most existing solutions cause larger blocking areas and hence a more severe exposed-terminal problem. In this paper, we present a new Enhanced Busy-tone Multiple Access (EBTMA) medium access control (MAC) protocol. The proposed protocol minimizes the negative impact of both the hidden-terminal and the exposed-terminal problems with the assistance of an out-of-band busy tone signal. The new protocol can also enhance the reliability of packet broadcast and multicast which are very important for many network control functions such as routing. Unlike the previous busy-tone schemes, such as the original Busy-tone Multiple Access (BTMA) protocol, our proposed protocol uses a non-interfering busy- tone signal in a short period of time, in order to notify all hidden terminals without blocking a large number of nodes for a long time. Our analysis, verified by simulation results, demonstrates that the proposed MAC protocol outperforms the existing ones and it can greatly improve the performance of wireless ad hoc networks. In addition, the proposed EBTMA protocol can co-exist with the existing 802.11 MAC protocol, so it can be incrementally deployed.
Ahmad Ali Abdullah, Lin Cai 0001, Fayez Gebali
VTC Fall3
2009 Modeling the Throughput and Delay in Wireless Multihop Ad Hoc Networks
abstract
In wireless multihop ad hoc networks, the use of the RTS/CTS mechanism does not completely eliminate the hidden-terminal problem. Considering the hidden-terminal problem adds complexity to the existing analysis for single-hop networks. In this paper, we provide precise and accurate analytical models for quantifying the throughput and delay in wireless multihop ad hoc networks. The proposed analysis is applicable to many wireless MAC protocols and applications. The accuracy of our analytical models are verified by extensive NS-2 simulations. Our analysis reveals how the throughput and delay in wireless multihop ad hoc networks are affected by the hidden-terminals and by the transmission and interference ranges of wireless devices. These results are important for network planning and protocol optimization in wireless multihop ad hoc networks.
Ahmad Ali Abdullah, Fayez Gebali, Lin Cai 0001
GLOBECOM2
2009 Cross-Layer Modeling of Wireless Ad Hoc Networks in the Presence of Channel Noise
abstract
We present a new analytical model of MAC layer for wireless ad hoc networks that takes into account channel bit errors and frame retry limits for a two-way handshaking mechanism. This model offers flexibility to address key design issues such as the effects of traffic parameters and possible improvements for wireless ad hoc networks. We illustrate that an important parameter affecting the performance of binary exponential back-off is the initial back-off window size. We show that for a low bit error rate the throughput increases and, at the same time, the delay and the average energy to transmit the frame is reduced. Results show also that the negative acknowledgment (NAK)-based model proves more useful for a high bit error rate.
K. M. Jamil Khayyat, Fayez Gebali
GLOBECOM2
2009 Targeting spam control on middleboxes: Spam detection based on layer-3 e-mail content classification
Muhammad N. Marsono, M. Watheq El-Kharashi, Fayez Gebali
Comput. Networks3
2009 A spam rejection scheme during SMTP sessions based on layer-3 e-mail classification
Muhammad N. Marsono, M. Watheq El-Kharashi, Fayez Gebali
J. Netw. Comput. Appl.3
2008 Power-aware topology optimization for networks-on-chips
abstract
The choice of a network topology for a Networks-on-Chip based application significantly impacts its power consumption. In this paper, we propose a new methodology to reduce the total power consumption of the global router-to-router links by selecting the optimal network topology. The proposed methodology merges three mapping approaches: network partitioning, standard topology mapping, and long-range insertion. Analytical power models for global links are studied at different levels of abstraction. The proposed methodology is validated through a case study. Experimental results show the power consumption improvement compared to related work.
Haytham Elmiligi, Ahmed A. Morgan, M. Watheq El-Kharashi, Fayez Gebali
ISCAS4
2008 Quality of service support in wireless local area network with error control protocol
abstract
In this paper, we applied an error control protocol for wireless local area network in medium access control. Hiperlan\2 random access phase is taken as an example. We applied quality of service support in the random access phase. Analytical model is developed for the backoff strategy with error control protocol. The performance metrics are shown.
Abdelsalam B. Amer, Fayez Gebali, Yousry Abdel-Hamid
LCN2
2008 Backoff Strategies in Hiperlan2 with Error Control Protocol
abstract
In this paper, we study the impact of random access phase in Hiperlan\2 medium access control frame. We propose and study several backoff strategies for the unsuccessful users in the random access phase. Markov chain modeling is used to model the random access channels. We also applied the error control to the transmission state. Instead of assuming the channel is error free, we studied the impact when the channel is in error. We developed analytical model and we applied this model for simple backoff strategy model. A comparison is made between these models in terms of their throughput, delay, energy, etc. The performance of the error control model is shown in terms of average number of retransmissions and the efficiency.
Abdelsalam B. Amer, Fayez Gebali, Yousry Abdel-Hamid
VTC Fall2
2006 Binary LNS-based naive Bayes hardware classifier for spam control
abstract
We propose a hardware architecture for a naive Bayes classifier in the context of e-mail classification for spam control. Our proposal presents a word-serial naive Bayes classifier architecture that utilizes the logarithmic number system (LNS) to reduce the computational complexity. We present the hardware architecture for non-iterative binary LNS recoding using a look-up table approach. Our design was synthesized targeting an Altera Stratix CPLD device. The synthesized classifier was functionally verified with a MATLAB implementation. Our binary LNS naive Bayes classifier exhibits high e-mail classification throughput of more than 30 thousands e-mails per second.
Muhammad N. Marsono, M. Watheq El-Kharashi, Fayez Gebali
ISCAS3
2006 An FPGA implementation of the flexible triangle search algorithm for block based motion estimation
abstract
In this paper a hardware architecture for the implementation of the flexible triangle search algorithm (FTS) using FPGAs is proposed. The FTS is a fast block-matching algorithm for motion estimation proposed in previous work, which can be used for video compression. The FTS finds the best matching blocks between two frames using a search triangle which changes its direction and size through a set of operations. These operations provide the triangle with the necessary flexibility to locate the best matching block. Simulation results indicate that the FTS reduces the number of block matching operations compared with other fast block matching algorithms without affecting quality or compression ratio of the compressed bitstream. In this paper, a hardware architecture for a FPGA implementation of the FTS algorithm is proposed. This architecture is simulated and tested using VHDL and synthesized using Xilinx ISE for the Xilinx Spartan3 device. The results obtained were compared to an FPGA implementation of the full search (FS) algorithm. Results indicates that the FTS FPGA implementation requires less number of gates than FS and the required number of cycles needed to complete motion search for one block is much lower. This indicates that the proposed implementation is fast and requires less hardware and power than existing ones.
Mohamed M. Rehan, M. Watheq El-Kharashi, Panajotis Agathoklis, Fayez Gebali
ISCAS4
2006 A New Analytical Model for Computing Blocking Probability in Optical Burst Switching Networks
abstract
This paper presents a new analytical model for calculating the blocking probability in Just-Enough-Time (JET)-based optical burst switching networks. Relationship to the problem of calculating the reservation probability in advance reservation systems is also discussed. The proposed analytical model takes into consideration the effects of the burst offset time and the burst length on the blocking probability. We use a (M+1)-state non-homogenous Markov chain to describe the state of an output link carrying M wavelength channels. In addition, we model each wavelength channel by a 2-state Markov chain. The offset time is drawn from a specified distribution so that wavelength reservation requests, made before a given time, build up to be a workload whose mean value declines with the reservation starting time. Furthermore, we express the blocking probability in terms of first passage time distributions to account for the burst length. To verify its accuracy, the model results are compared with the results of a sophisticated discrete-event simulation model. The model results were found to be in satisfactory agreement with simulation results.
Ayman Kaheel, Hussein M. Alnuweiri, Fayez Gebali
IEEE J. Sel. Areas Commun.3
2006 Processor Array Architectures for Deep Packet Classification
abstract
This paper presents a systematic technique for expressing a string search algorithm as a regular iterative expression to explore all possible processor arrays for deep packet classification. The computation domain of the algorithm is obtained and three affine scheduling functions are presented. The technique allows some of the algorithm variables to be pipelined while others are broadcast over system-wide buses. Nine possible processor array structures are obtained and analyzed in terms of speed, area, power, and I/O timing requirements. Time complexities are derived analytically and through extensive numerical simulations. The proposed designs exhibit optimum speed and area complexities. The processor arrays are compared with previously derived processor arrays for the string matching problem.
Fayez Gebali, A. N. M. Ehtesham Rafiq
IEEE Trans. Parallel Distributed Syst.1
2005 Markov chain analysis of collaborative codes in random multi access communication systems
Fayez Gebali, A-Imam Al-Sammak
Comput. Commun.1
2004 Analytical evaluation of blocking probability in optical burst switching networks
abstract
In this paper we present a new analytical model for evaluating the blocking probability in Just-Enough-Time-based optical burst switching networks. The proposed analytical model takes into consideration the effects of the burst offset time and the burst length on the blocking probability. We use a (M+1)-state nonhomogenous Markov chain to describe the state of an output link carrying M wavelength channels. In addition, we model each wavelength channel by a 2-state Markov chain. The offset time is drawn from a specified distribution so that wavelength reservation requests, made before a given time, build up to be a workload whose mean value declines with the reservation starting time. Furthermore, we express the blocking probability in terms of first passage time distributions to account for the burst length. To verify its accuracy, the model results are compared with the results of a sophisticated discrete-event simulation model. The model results were found to be in satisfactory agreement with simulation results.
Ayman Kaheel, Hussein M. Alnuweiri, Fayez Gebali
ICC3
2004 A new analytical model for computing blocking probability in optical burst switching networks
abstract
This work presents a new analytical model for calculating the blocking probability in just-enough-time (JET)-based optical burst switching networks. Relationship to the problem of calculating the reservation probability in advance-reservation systems is also discussed. The proposed analytical model takes into consideration the effects of the offset time and the burst length on the blocking probability. We model the wavelength channel by a 2-state nonhomogenous Markov chain. The offset time is drawn from a specified distribution so that wavelength reservation requests, made before a given time, build up to be a workload whose mean value declines with the reservation starting time. In addition, we express the blocking probability in terms of first passage time distributions to account for the burst length. To verify its accuracy, the model results are compared with the results of a discrete-event simulation model. The model results were found to be in satisfactory agreement with simulation results.
Ayman Kaheel, Hussein M. Alnuweiri, Fayez Gebali
ISCC3
2004 A fast string search algorithm for deep packet classification
A. N. M. Ehtesham Rafiq, M. Watheq El-Kharashi, Fayez Gebali
Comput. Commun.3