Paul Metzgen

dblp:58/2023 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
0since 2021 · last 2005
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Electronic design automation · 60% Processor architecture and microarchitecture · 17% Reconfigurable computing and FPGAs · 17%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation › design optimization
area optimization
0.112005
Multiplexer restructuring for FPGA implementation cost reduction · DAC 2005
Electronic design automation › logic synthesis
FPGA synthesis
0.112005
Multiplexer restructuring for FPGA implementation cost reduction · DAC 2005
Electronic design automation
logic synthesis
0.112005
Multiplexer restructuring for FPGA implementation cost reduction · DAC 2005
Processor architecture and microarchitecture › arithmetic unit
arithmetic logic unit
0.012004
A high performance 32-bit ALU for programmable logic · FPGA 2004
Reconfigurable computing and FPGAs › FPGA-based processor implementation
soft-core processor
0.012004
A high performance 32-bit ALU for programmable logic · FPGA 2004
Integrated circuit design
low-power circuit design
0.012004
A high performance 32-bit ALU for programmable logic · FPGA 2004

Methods — techniques the papers use, named apart from their topics

multiplexer restructuring · 0.1LUT optimization · 0.1logic element mapping · 0.0
YearPublicationVenuePosition
2005 Multiplexer restructuring for FPGA implementation cost reduction
abstract
This paper presents a novel synthesis algorithm that reduces the area needed for implementing multiplexers on an FPGA by an average of 18%. This is achieved by reducing the number of Lookup Tables (LUTs) needed to implement multiplexers. The algorithm relies on reimplementing 2:1 multiplexer trees using efficient 4:1 multiplexers. The key to the algorithm's performance lies in exploiting the observation that most multiplexers occur in busses. New optimizations are employed which pay a small cost in logic that is shared across the bus to achieve a reduction in the logic required for every bit of the bus.
Paul Metzgen, Dominic Nancekievill
DAC1
2004 A high performance 32-bit ALU for programmable logic
abstract
The Arithmetic-Logic-Unit (ALU) is at the heart of a modern microprocessor, and its size and speed are often significant contributors to the overall processor's cost and performance. This paper presents the design of the ALU used in Altera's NIOS 2.0 soft processor implemented on Altera's Apex 20KE FPGA architecture. This ALU enabled the 32-bit NIOS 2.0 to consume only 1200 LEs and run at 85MHz. This is a 50% size reduction and 70% speed improvement over its predecessor, NIOS 1.1.The Logic-element (LE) is the basic building block within the Apex architecture. Making full use of the advanced features of the LE has resulted in this novel ALU design. A functional representation of the logic is used to describe how the ALU performs the core set of NIOS instructions, and an LE representation shows the amount of logic-resources needed for the implementation. The cost of additional features such as a barrel-shifter and custom instructions is also described.Likely worst-case delays for different routing and logic elements are used to estimate the ALU's speed. Further speed and size optimizations are also presented from which it is possible to create ALU ranging in speed from 87 MHz to over 100 MHz.
Paul Metzgen
FPGA1
2002 Strassen's matrix multiplication for customisable processors
abstract
Strassen's algorithm is an efficient method for multiplying large matrices. We explore various ways of mapping Strassen's algorithm into reconfigurable hardware that contains one or more customisable instruction processors. Our approach has been implemented using Nios processors with custom instructions and with custom-designed coprocessors, taking advantage of the additional logic and memory blocks available on a reconfigurable platform.
Henry M. D. Ip, James D. Low, Peter Y. K. Cheung, George A. Constantinides, Wayne Luk, Shay Ping Seng, Paul Metzgen
FPT7