Ahmed A. Wahba

dblp:193/3104 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
0since 2021 · last 2017
0000-0003-2943-821XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Integrated circuit design · 87% Processor architecture and microarchitecture · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Integrated circuit design
digital circuit design
0.312017
Area Efficient and Fast Combined Binary/Decimal Floating Point Fused Multiply Add Unit · IEEE Trans. Computers 2017
Integrated circuit design › digital circuit design › arithmetic circuit design
floating-point unit
0.312017
Area Efficient and Fast Combined Binary/Decimal Floating Point Fused Multiply Add Unit · IEEE Trans. Computers 2017
Processor architecture and microarchitecture › computer arithmetic › floating-point arithmetic
fused multiply-add
0.112017
Area Efficient and Fast Combined Binary/Decimal Floating Point Fused Multiply Add Unit · IEEE Trans. Computers 2017

Methods — techniques the papers use, named apart from their topics

rounding-while-redundant · 0.3redundant arithmetic · 0.3leading zeros detection · 0.3column reduction · 0.3
YearPublicationVenuePosition
2017 Area Efficient and Fast Combined Binary/Decimal Floating Point Fused Multiply Add Unit
abstract
In this work we present a new 64-bit floating point Fused Multiply Add (FMA) unit that can perform both binary and decimal addition, multiplication, and fused-multiply-add operations. The presented FMA has 6 percent less delay than the fastest stand-alone decimal unit and 23 percent less area than both binary and decimal units together. These results were achieved by the use of: 1) column by column reduction to reduce the partial products in the multiplier tree, 2) a new leading zeros detector that produces its output in base-3 to simplify the normalization shifting in the binary datapath, 3) the use of a redundant adder to perform the final addition, 4) using a new rounding-while-redundant technique to hide the rounding delay and remove it from the critical path, and 5) using a new simple conversion technique from redundant to binary/decimal.
Ahmed A. Wahba, Hossam A. H. Fahmy
IEEE Trans. Computers1