Greg Faanes

dblp:17/5131 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Interconnection networks and networks-on-chip · 41% Processor architecture and microarchitecture · 23% High-performance computing · 23%
Computer networks
1 paper
Routing and switching · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
distributed memory systems
0.112012
Cray cascade: a scalable HPC system based on a Dragonfly network · SC 2012
Interconnection networks and networks-on-chip › network topology › low-diameter topology
dragonfly network
0.112012
Cray cascade: a scalable HPC system based on a Dragonfly network · SC 2012
Interconnection networks and networks-on-chip
network topology
0.112012
Cray cascade: a scalable HPC system based on a Dragonfly network · SC 2012
Memory systems › shared memory
distributed shared memory
0.112007
The Cray BlackWidow: a highly scalable vector multiprocessor · SC 2007
Processor architecture and microarchitecture › vector processor
vector multiprocessor
0.112007
The Cray BlackWidow: a highly scalable vector multiprocessor · SC 2007
Routing and switching
adaptive routing
0.012012
Cray cascade: a scalable HPC system based on a Dragonfly network · SC 2012
Processor architecture and microarchitecture
instruction set architecture
0.012000
Vector instruction set support for conditional operations · ISCA 2000
Processor architecture and microarchitecture › SIMD
vector instruction set
0.012000
Vector instruction set support for conditional operations · ISCA 2000
High-performance computing
supercomputing
0.012007
The Cray BlackWidow: a highly scalable vector multiprocessor · SC 2007
Memory systems
cache
0.011994
Cache performance in vector supercomputers · SC 1994
Memory systems › cache
cache behavior
0.011994
Cache performance in vector supercomputers · SC 1994
Processor architecture and microarchitecture › instruction set architecture › instruction set extension
multimedia instruction set
0.012000
Vector instruction set support for conditional operations · ISCA 2000
Processor architecture and microarchitecture
SIMD
0.012000
Vector instruction set support for conditional operations · ISCA 2000
Memory systems
memory hierarchy
0.011994
Cache performance in vector supercomputers · SC 1994
High-performance computing › supercomputer architecture
vector supercomputer
0.011994
Cache performance in vector supercomputers · SC 1994

Methods — techniques the papers use, named apart from their topics

simulation · 0.3prototype performance measurement · 0.3system architecture design · 0.1performance evaluation · 0.1trace-driven analysis · 0.0
YearPublicationVenuePosition
2012 Cray cascade: a scalable HPC system based on a Dragonfly network
abstract
Higher global bandwidth requirement for many applications and lower network cost have motivated the use of the Dragonfly network topology for high performance computing systems. In this paper we present the architecture of the Cray Cascade system, a distributed memory system based on the Dragonfly [1] network topology. We describe the structure of the system, its Dragonfly network and the routing algorithms. We describe a set of advanced features supporting both mainstream high performance computing applications and emerging global address space programing models. We present a combination of performance results from prototype systems and simulation data for large systems. We demonstrate the value of the Dragonfly topology and the benefits obtained through extensive use of adaptive routing.
Greg Faanes, Abdulla Bataineh, Duncan Roweth, Tom Court, Edwin Froese, Robert Alverson, Tim Johnson, Joe Kopnick, Mike Higgins, James Reinhard
SC1
2007 The Cray BlackWidow: a highly scalable vector multiprocessor
abstract
This paper describes the system architecture of the Cray BlackWidow scalable vector multiprocessor. The BlackWidow system is a distributed shared memory (DSM) architecture that is scalable to 32K processors, each with a 4-way dispatch scalar execution unit and an 8-pipe vector unit capable of 20.8 Gflops for 64-bit operations and 41.6 Gflops for 32-bit operations at the prototype operating frequency of 1.3 GHz. Global memory is directly accessible with processor loads and stores and is globally coherent. The system supports thousands of outstanding references to hide remote memory latencies, and provides a rich suite of built-in synchronization primitives. Each BlackWidow node is implemented as a 4-way SMP with up to 128 Gbytes of DDR2 main memory capacity. The system supports common programming models such as MPI and OpenMP, as well as global address space languages such as UPC and CAF. We describe the system architecture and microarchitecture of the processor, memory controller, and router chips. We give preliminary performance results and discuss design tradeoffs.
Dennis Abts, Abdulla Bataineh, Steve Scott, Greg Faanes, James L. Schwarzmeier, Eric Lundberg, Tim Johnson, Mike Bye, Gerald Schwoerer
SC4
2000 Vector instruction set support for conditional operations
abstract
Vector instruction sets are receiving renewed interest because of their applicability to multimedia. Current multimedia instruction sets use short vectors with SIMD implementations, but long vector, pipelined implementations have a number of advantages and are a logical next step in multimedia ISA development. Support for conditional operations (as occur in loops containing IF statements) is an important aspect of a vector ISA. Seven ISA alternatives for implementing conditional operations are systematically explored. Performance considerations are discussed through evaluation of a typical IF loop over a range of vector lengths and true conditional values. An approach using masked operations is shown to be one of the better methods, especially if its implementation is able to skip over blocks of false mask bits. Additional analyses of complex IF loops and parallel pipeline implementations support the masked operation approach. The paper concludes with a practical implementation of masked operations that skips over power-of-2-length blocks of false values. This implementation is simpler than skipping arbitrary-length blocks and provides similar performance. 1.
James E. Smith 0001, Greg Faanes, Rabin A. Sugumar
ISCA2
1994 Cache performance in vector supercomputers
abstract
Traditional supercomputers use a flat multi-bank SRAM memory organization to supply high bandwidth at low latency. Most other computers use a hierarchical organization with a small SRAM cache and a slower, cheaper DRAM for the main memory. Such systems rely heavily on data locality for achieving optimum performance. This paper evaluates cache-based memory systems for vector supercomputers. We develop a simulation model for a cache-based version of the Cray Research C90 and use the NAS parallel benchmarks to provide a large-scale workload. We show that while caches reduce memory traffic and improve the performance of plain DRAM memory, they still lag behind cacheless SRAM. We identify the performance bottlenecks in DRAM-based memory systems and quantify their contribution to program performance degradation. We find the data fetch strategy to be a significant parameter affecting performance, we evaluate the performance of several fetch policies, and we show that small fetch sizes improve performance by maximizing the use of available memory bandwidth.>
Leonidas I. Kontothanassis, Rabin A. Sugumar, Greg Faanes, James E. Smith 0001, Michael L. Scott
SC3