Timothy J. Knight

dblp:69/567 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
0since 2021 · last 2008
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 31% Hardware reliability and fault tolerance · 23% Processor architecture and microarchitecture · 20%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
parallel programming runtimes
0.112008
A portable runtime interface for multi-level memory hierarchies · PPoPP 2008
Compilers and program optimization › memory optimization
memory hierarchy optimization
0.112006
Sequoia: programming the memory hierarchy · SC 2006
Parallel and multicore computing
parallel programming models
0.112006
Sequoia: programming the memory hierarchy · SC 2006
Processor architecture and microarchitecture › dataflow architecture
stream architecture
0.122005
Merrimac: Supercomputing with Streams · SC 2003
Fault Tolerance Techniques for the Merrimac Streaming Supercomputer · SC 2005
Hardware reliability and fault tolerance
soft errors
0.112005
Fault Tolerance Techniques for the Merrimac Streaming Supercomputer · SC 2005
Hardware reliability and fault tolerance › soft errors
soft error resilience
0.112005
Fault Tolerance Techniques for the Merrimac Streaming Supercomputer · SC 2005
Memory systems
data movement
0.012008
A portable runtime interface for multi-level memory hierarchies · PPoPP 2008
Memory systems
memory hierarchy
0.012008
A portable runtime interface for multi-level memory hierarchies · PPoPP 2008
High-performance computing
scientific computing systems
0.012005
Fault Tolerance Techniques for the Merrimac Streaming Supercomputer · SC 2005
Processor architecture and microarchitecture › data-parallel architecture
stream processor
0.012005
Fault Tolerance Techniques for the Merrimac Streaming Supercomputer · SC 2005
High-performance computing › supercomputing
supercomputing systems
0.012003
Merrimac: Supercomputing with Streams · SC 2003

Methods — techniques the papers use, named apart from their topics

compiler scheduling · 0.1bulk operation manipulation · 0.1runtime composition · 0.1compiler target interface · 0.1software fault tolerance · 0.1reconfigurability · 0.1hardware redundancy · 0.1
YearPublicationVenuePosition
2008 A portable runtime interface for multi-level memory hierarchies
abstract
We present a platform independent runtime interface for moving data and computation through parallel machines with multi-level memory hierarchies. We show that this interface can be used as a compiler target and can be implemented easily and efficiently on a variety of platforms. The interface design allows us to compose multiple runtimes, achieving portability across machines with multiple memory levels. We demonstrate portability of programs across machines with two memory levels with runtime implementations for multi-core/SMP machines, the STI Cell Broadband Engine, a distributed memory cluster, and disk systems. We also demonstrate portability across machines with multiple memory levels by composing runtimes and running on a cluster of SMP nodes, out-of-core algorithms on a Sony Playstation 3 pulling data from disk, and a cluster of Sony Playstation 3's. With this uniform interface, we achieve good performance for our applications and maximize bandwidth and computational resources on these system configurations.
Mike Houston, Ji Young Park, Manman Ren, Timothy J. Knight, Kayvon Fatahalian, Alex Aiken, William J. Dally, Pat Hanrahan
PPoPP4
2007 Compilation for explicitly managed memory hierarchies
abstract
We present a compiler for machines with an explicitly managed memory hierarchy and suggest that a primary role of any compiler for such architectures is to manipulate and schedule a hierarchy of bulk operations at varying scales of the application and of the machine. We evaluate the performance of our compiler using several benchmarks running on a Cell processor.
Timothy J. Knight, Ji Young Park, Manman Ren, Mike Houston, Mattan Erez, Kayvon Fatahalian, Alex Aiken, William J. Dally, Pat Hanrahan
PPoPP1
2006 Sequoia: programming the memory hierarchy
abstract
We present Sequoia, a programming language designed to facilitate the development of memory hierarchy aware parallel programs that remain portable across modern machines featuring different memory hierarchy configurations. Sequoia abstractly exposes hierarchical memory in the programming model and provides language mechanisms to describe communication vertically through the machine and to localize computation to particular memory locations within it. We have implemented a complete programming system, including a compiler and runtime systems for Cell processor-based blade systems and distributed memory clusters, and demonstrate efficient performance running Sequoia programs on both of these platforms.
Kayvon Fatahalian, Daniel Reiter Horn, Timothy J. Knight, Larkhoon Leem, Mike Houston, Ji Young Park, Mattan Erez, Manman Ren, Alex Aiken, William J. Dally, Pat Hanrahan
SC3
2005 Fault Tolerance Techniques for the Merrimac Streaming Supercomputer
abstract
As device scales shrink, higher transistor counts are available while soft-errors, even in logic, become a major concern. A new class of architectures, such as Merrimac and the IBM Cell, take advantage of the higher transistor count by exposing control, communication, and a large number of functional-units at the architectural level, thus achieving high performance and efficiency. This paper explores soft-error fault tolerance in the context of these computeintensive architectures, which differ significantly from their control-intensive CPU counterparts. The main goal of the proposed schemes for Merrimac is to conserve the critical and costly off-chip bandwidth and on-chip storage resources, while maintaining high peak and sustained performance. We achieve this by allowing for reconfigurability and relying on programmer input. The processor is either run at full peak performance employing software fault-tolerance methods, or reduced performance with hardware redundancy. We present several methods, their analysis, and detailed case studies.
Mattan Erez, Nuwan Jayasena, Timothy J. Knight, William J. Dally
SC3
2003 Merrimac: Supercomputing with Streams
abstract
Merrimac uses stream architecture and advanced interconnection networks to give an order of magnitude more performance per unit cost than cluster-based scientific computers built from the same technology. Organizing the computation into streams and exploiting the resulting locality using a register hierarchy enables a stream architecture to reduce the memory bandwidth required by representative applications by an order of magnitude or more. Hence a processing node with a fixed bandwidth (expensive) can support an order of magnitude more arithmetic units (inexpensive). This in turn allows a given level of performance to be achieved with fewer nodes (a 1-PFLOPS machine, for example, with just 8,192 nodes) resulting in greater reliability, and simpler system management. We sketch the design of Merrimac, a streaming scientific computer that can be scaled from a $20K 2 TFLOPS workstation to a $20M 2 PFLOPS supercomputer and present the results of some initial application experiments on this architecture.
William J. Dally, Francois Labonte, Pat Hanrahan, Jung Ho Ahn, Jayanth Gummaraju, Mattan Erez, Nuwan Jayasena, Ian Buck, Timothy J. Knight, Ujval J. Kapasi
SC10