David R. Galloway

dblp:78/3398 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Reconfigurable computing and FPGAs · 64% Electronic design automation · 34% Processor architecture and microarchitecture · 1%
Software engineering, system software, and programming languages
1 paper
Runtime systems and virtual machines · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA architecture
0.212016
The Stratix™ 10 Highly Pipelined FPGA Architecture · FPGA 2016
Electronic design automation › physical design › placement and routing
FPGA place-and-route
0.212016
The Stratix™ 10 Highly Pipelined FPGA Architecture · FPGA 2016
Reconfigurable computing and FPGAs › FPGA architecture
adaptive logic module
0.112005
The Stratix II logic and routing architecture · FPGA 2005
Reconfigurable computing and FPGAs
FPGA routing architecture
0.112005
The Stratix II logic and routing architecture · FPGA 2005
Reconfigurable computing and FPGAs
multi-FPGA system
0.011997
The Transmogrifier-2: A 1 Million Gate Rapid Prototyping System · FPGA 1997
Reconfigurable computing and FPGAs
rapid prototyping
0.011997
The Transmogrifier-2: A 1 Million Gate Rapid Prototyping System · FPGA 1997
Interconnection networks and networks-on-chip
network topology
0.011997
The Transmogrifier-2: A 1 Million Gate Rapid Prototyping System · FPGA 1997
Processor architecture and microarchitecture
instruction set architecture
0.011986
Swamp: A Fast Processor for Smalltalk-80 · OOPSLA 1986
Processor architecture and microarchitecture
microprogramming
0.011986
Swamp: A Fast Processor for Smalltalk-80 · OOPSLA 1986

Methods — techniques the papers use, named apart from their topics

pipelining · 0.2circuit retiming · 0.2arithmetic structure design · 0.1LUT partitioning · 0.1programmable clock · 0.0partial crossbar · 0.0micro-address prediction · 0.0bit-slice ALU · 0.0
YearPublicationVenuePosition
2016 The Stratix™ 10 Highly Pipelined FPGA Architecture
abstract
This paper describes architectural enhancements in the Altera Stratix? 10 HyperFlex? FPGA architecture, fabricated in the Intel 14nm FinFET process. Stratix 10 includes ubiquitous flip-flops in the routing to enable a high degree of pipelining. In contrast to the earlier architectural exploration of pipelining in pass-transistor based architectures, the direct drive routing fabric in Stratix-style FPGAs enables an extremely low-cost pipeline register. The presence of ubiquitous flip-flops simplifies circuit retiming and improves performance. The availability of predictable retiming affects all stages of the cluster, place and route flow. Ubiquitous flip-flops require a low-cost clock network with sufficient flexibility to enable pipelining of dozens of clock domains. Different cost/performance tradeoffs in a pipelined fabric and use of a 14nm process, lead to other modifications to the routing fabric and the logic element. User modification of the design enables even higher performance, averaging 2.3X faster in a small set of designs.
David M. Lewis, Gordon R. Chiu, Jeffrey Chromczak, David R. Galloway, Ben Gamsa, Valavan Manohararajah, Ian Milton, Tim Vanderhoek, John Van Dyken
FPGA4
2005 The Stratix II logic and routing architecture
abstract
This paper describes the Altera Stratix II™ logic and routing architecture. This architecture features a novel adaptive logic module (ALM) that is based on a 6-LUT, but can be partitioned into two smaller LUTs to efficiently implement circuits containing a range of LUT sizes that arises in conventional synthesis flows. This provides a performance increase of 15% in the Stratix II architecture while reducing area by 2%. The ALM also includes a more powerful arithmetic structure that can perform two bits of arithmetic per ALM, and perform a sum of up to three inputs. The routing fabric adds a new set of fast inputs to the routing multiplexers for another 3% improvement in performance, while other improvements in routing efficiency cause another 6% reduction in area. These changes in combination with other circuit and architecture changes in Stratix II contribute 27% of an overall 51% performance improvement (including architecture and process improvement). The architecture changes reduce area by 10% in the same process, and by 50% after including process migration.
David M. Lewis, Elias Ahmed, Gregg Baeckler, Vaughn Betz, Mark Bourgeault, David Cashman, David R. Galloway, Mike Hutton, Christopher Lane, Andy Lee, Paul Leventis, Sandy Marquardt, Cameron McClintock, Ketan Padalia, Bruce Pedersen, Giles Powell, Boris Ratchev, Srinivas Reddy, Jay Schleicher, Kevin Stevens, Richard Yuan, Richard Cliff, Jonathan Rose
FPGA7
2005 The Transmogrifier-4: An FPGA-Based Hardware Development System with Multi-Gigabyte Memory Capacity and High Host and Memory Bandwidth
Joshua Fender, Jonathan Rose, David R. Galloway
FPT3
1998 The Transmogrifier-2: a 1 million gate rapid-prototyping system
abstract
This paper describes the Transmogrifier-2 (TM-2), a second-generation multifield programmable gate array (FPGA) rapid-prototyping system. The largest version of the system will comprise 16 boards that each contain two Altera 10K50 FPGA's, four I-Cube interconnect chips, and up to 8 Mbytes of memory. The inter-FPGA routing architecture of the TM-2 uses a novel interconnect structure, a nonuniform partial crossbar, that provides a constant delay between any two FPGA's in the system. The TM-2 architecture is modular and scalable, meaning that systems of various sizes can be constructed from copies of the same board, while maintaining routability and the constant delay feature. Other features include a system-level programmable clock that allows single-cycle access to off-chip memory, and programmable clock waveforms with edge resolution of 10 ns. The first Transmogrifier-2 boards have been manufactured and are functional. They have recently been used successfully in some simple graphics acceleration applications.
David M. Lewis, David R. Galloway, Marcus van Ierssel, Jonathan Rose, Paul Chow
IEEE Trans. Very Large Scale Integr. Syst.2
1997 The Transmogrifier-2: A 1 Million Gate Rapid Prototyping System
abstract
This paper describes the Transmogrifier-2, a second generation multi-FPGA system. The largest version of the system will comprise 16 boards that each contain two Altera 10K50 FPGAs, four I-cube interconnect chips, and up to 8 Mbytes of memory. The inter-FPGA routing architecture of the TM-2 uses a novel interconnect structure, a non-uniform partial crossbar, that provides a constant delay between any two FPGAs in the system. The TM-2 architecture is modular and scalable, meaning that various sized systems can be constructed from the same board, while maintaining routability and the constant delay feature. Other features include a system-level programmable clock that allows single-cycle access to off-chip memory, and programmable clock waveforms with resolution to 10ns. The first Transmogrifier-2 boards have been manufactured and are functional. They have recently been used successfully in some simple graphics acceleration applications.
David M. Lewis, David R. Galloway, Marcus van Ierssel, Jonathan Rose, Paul Chow
FPGA2
1995 The Transmogrifier C hardware description language and compiler for FPGAs
abstract
The Transmogrifier C hardware description language is almost identical to the C programming language, making it attractive to the large community of C-language programmers. This paper describes the semantics of the language and presents a Transmogrifier C compiler that targets the Xilinx 4000 FPGA. The compiler is operational and has produced several working circuits, including a graphics display driver.
David R. Galloway
FCCM1
1986 Swamp: A Fast Processor for Smalltalk-80
abstract
A processor for the Smalltalk-80↑ programming language is described. This machine is implemented using a standard bit slice ALU and sequencer, TTL MSI, and NMOS LSI RAMS. It executes an instruction set similar to the Smalltalk-80 virtual machine instruction set. The data paths of the machine are optimized for rapid Smalltalk-80 execution by the inclusion of a context cache, tag checking, and a hardware method cache. Each context is only partly initialized when created, and has no memory allocated for it until a possibly non-LIFO reference to it is created. The machine is microprogrammed, and uses a simple next micro-address prediction strategy to obtain most of the performance of pipelining without the attendant complexity. The machine can execute simple instructions at over 7M bytecodes per second and has a predicted average throughput of 1.9M bytecodes per second.
David M. Lewis, David R. Galloway, Robert J. Francis, Brian W. Thomson
OOPSLA2