Atef Ibrahim

dblp:71/9913 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
1since 2021 · last 2022
0000-0002-1115-4051ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-authorComputer networks · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 37% Integrated circuit design · 30% Parallel and multicore computing · 18%
Network and information security
1 paper
Cryptographic primitives and cryptanalysis · 100%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
systolic array
0.722020
Unified and Scalable Digit-Serial Systolic Array for Multiplication and Division Over GF (2m) · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Design Space Exploration of 2-D Processor Array Architectures for Similarity Distance Computation · IEEE Trans. Parallel Distributed Syst. 2017
Cryptographic primitives and cryptanalysis
finite field arithmetic
0.412020
Unified and Scalable Digit-Serial Systolic Array for Multiplication and Division Over GF (2m) · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Integrated circuit design
digital circuit design
0.412020
Unified and Scalable Digit-Serial Systolic Array for Multiplication and Division Over GF (2m) · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Parallel and multicore computing
array processor
0.422017
Design Space Exploration of 2-D Processor Array Architectures for Similarity Distance Computation · IEEE Trans. Parallel Distributed Syst. 2017
Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm · IEEE Trans. Parallel Distributed Syst. 2011
Electronic design automation
design space exploration
0.312017
Design Space Exploration of 2-D Processor Array Architectures for Similarity Distance Computation · IEEE Trans. Parallel Distributed Syst. 2017
Hardware accelerators and domain-specific architectures
cryptographic accelerator
0.112011
Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm · IEEE Trans. Parallel Distributed Syst. 2011
Integrated circuit design › digital circuit design › arithmetic circuit design
modular multiplication
0.112011
Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm · IEEE Trans. Parallel Distributed Syst. 2011
Integrated circuit design › digital circuit design › arithmetic circuit design › modular multiplication
montgomery multiplication
0.112011
Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm · IEEE Trans. Parallel Distributed Syst. 2011
Energy-efficient computing › dynamic power reduction
glitch reduction
0.012011
Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm · IEEE Trans. Parallel Distributed Syst. 2011
Energy-efficient computing
low-power design
0.012011
Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm · IEEE Trans. Parallel Distributed Syst. 2011

Methods — techniques the papers use, named apart from their topics

stein's algorithm · 0.9digit-serial systolic array · 0.9linear scheduling · 0.63-d computation domain · 0.6data dependence graph · 0.1affine scheduling · 0.1
YearPublicationVenuePosition
2022 Systolic Processor Core for Finite-Field Multiplication and Squaring in Cryptographic Processors of IoT Edge Devices
abstract
Internet of Things (IoT) edge devices’ security is one of the main barriers to use IoT applications on a large scale. Securing these devices is mainly based on using primitive cryptographic algorithms. The hardware implementation of the cryptographic algorithms should be managed to be suitable for these resource-constrained devices. Finite-field arithmetic operations are at the heart of the cryptographic algorithms, and their efficient implementation directly affects the whole performance of the cryptographic algorithm. Filed multiplication operation is the core of the most finite-field arithmetic operations, such as squaring, inversion, and division. Therefore, this article mainly concentrates on efficiently implementing a resource-constrained unified processor core that simultaneously performs multiplication and squaring operations to reduce hardware resources. The offered processor core has a digit-serial systolic structure providing the designer with flexibility to manage the area, delay, and consumed energy to be suitable for IoT edge devices. ASIC results of the developed design and the reported efficient ones indicate that the proposed structure has a meaningful saving in the area and consumed energy for all embedded word-sizes$v$. The area achieves a reduction varying from 44.7% to 97.46% at$v=8$, 35.7% to 95.4% at$v=16$, and 57.1% to 95.6% at$v=32$. Also, the energy realizes a reduction ranging from 24.0% to 97.7% at$v=8$, 7.6% to 97.2% at$v=16$, and 25.4% to 98.1% at$v=32$. That makes it more suitable for embedded and IoT applications that impose more restrictions on the area and consumed energy.
Atef Ibrahim
IEEE Internet Things J.1
2020 Unified and Scalable Digit-Serial Systolic Array for Multiplication and Division Over GF (2m)
abstract
This brief offers a new unified and scalable digit-serial systolic array structure to implement the unified Stein's multiplication and division algorithm. The proposed structure is flexible enough to help the designer select the required number of processing elements and manage the latency of the multiplication/division operations. Thus, the proposed design can realize the required time performance with minimum space complexity. The implementation results of the proposed design and the previously reported digit-serial competitor designs display that the proposed scalable architecture has better performance for 32-bit embedded cryptographic processors that need reasonable performance with a small footprint.
Atef Ibrahim
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Blockchain in internet-of-things: a necessity framework for security, reliability, transparency, immutability and liability
abstract
Blockchain is a distributed operation and information supervision technology programmed initially for Bitcoin cryptocurrency. The awareness in Blockchain technology is rapidly growing since the notion was invented in the year 2008. The motivation for the concentration in Blockchain is its significant characteristics that deliver security, privacy, and information reliability devoid of any additional system regulating the communications, and consequently it generates fascinating research domains, specifically from the viewpoint of methodological difficulties and restrictions. This study discovers the wide‐ranging Blockchain technology and studies it's perspective with respect to ‘ internet‐of‐things ’ controlled nodes. A resilient prototype method has been programmed that reveals a basic system exhausting Blockchain. The outcome illustrates that the established method is functional in test‐bed environment.
Usman Tariq, Atef Ibrahim, Tariq Ahamed Ahanger, Yassine Bouteraa, Ahmed M. Elmogy
IET Commun.2
2017 Design Space Exploration of 2-D Processor Array Architectures for Similarity Distance Computation
abstract
We present a systematic methodology for exploring the design space of similarity distance computation in machine learning algorithms. Previous architectures proposed in the literature have been obtained using ad hoc techniques that do not allow for design space exploration. The size and dimensionality of the input datasets have not been taken into consideration in previous works. This may result in impractical designs that are not amenable for hardware implementation. The methodology presented in this work is used to obtain the 3-D computation domain of the similarity distance computation algorithm. A scheduling function determines whether an algorithm variable is pipelined or broadcast. Four linear scheduling functions are presented, and six possible 2-D processor array architectures are obtained and classified based on the size and dimensionality of the input datasets. The obtained designs are analyzed in terms of speed and area, and compared with previously obtained designs. The proposed designs achieve better time and area complexities.
Awos Kanan, Fayez Gebali, Atef Ibrahim
IEEE Trans. Parallel Distributed Syst.3
2015 Efficient Scalable Serial Multiplier Over GF(2m) Based on Trinomial
abstract
This brief presents a novel low-complexity scalable serial architecture for finite field multiplication over GF(2m) based on irreducible trinomial. This architecture was explored by applying nonlinear technique that allows the designer, using progressive product reduction technique, to control the workload per processor and also allows the communication overhead between processors to be reduced. By comparing the ASIC implementation of the proposed structure to some of the previously published structures, the proposed structure have at least 71.7% lower area and at least 89.9% lower power compared with most of them. This makes the proposed design more suitable for constrained implementations of cryptographic primitives in resource constrained applications, such as smart cards, handheld devices, and implantable medical devices.
Fayez Gebali, Atef Ibrahim
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Systolic Array Architectures for Sunar-Koç Optimal Normal Basis Type II Multiplier
abstract
We present linear and nonlinear techniques for design exploration of an iterative algorithm. The nonlinear techniques allow control of processor workload and control of communication between processors. The algorithm considered is the Sunar-Koç optimal normal basis type II multiplication algorithm. Six systolic arrays are obtained. General formulas are provided for each design so that the operation of the system can be determined for a given GF(2m). The proposed architectures have been implemented using 45-nm CMOS technology and compared with published architectures. The results show that the proposed designs have at least 44.4% lower total computation time compared with the designs of all bit serial multipliers, while having slightly larger area delay product (ADP), up to 19.1%, compared with some of the bit serial multipliers and having smaller ADP values compared with most of the digit serial ones. Moreover, they have at least 46% lower power delay product compared with all bit serial and digit serial multipliers.
Atef Ibrahim, Fayez Gebali, Turki F. Al-Somani
IEEE Trans. Very Large Scale Integr. Syst.1
2011 Processor Array Architectures for Scalable Radix 4 Montgomery Modular Multiplication Algorithm
abstract
This paper presents a systematic methodology for exploring possible processor arrays of scalable radix 4 modular Montgomery multiplication algorithm. In this methodology, the algorithm is first expressed as a regular iterative expression, then the algorithm data dependence graph and a suitable affine scheduling function are obtained. Four possible processor arrays are obtained and analyzed in terms of speed, area, and power consumption. To reduce power consumption, we applied low power techniques for reducing the glitches and the Expected Switching Activity (ESA) of high fan-out signals in our processor array architectures. The resulting processor arrays are compared to other efficient ones in terms of area, speed, and power consumption.
Atef Ibrahim, Fayez Gebali, Hamed Elsimary, Amin M. Nassar
IEEE Trans. Parallel Distributed Syst.1