David A. Kranz

dblp:53/6629 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
0since 2021 · last 1999
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Parallel and multicore computing · 51% Memory systems · 27% Processor architecture and microarchitecture · 19%
Software engineering, system software, and programming languages
4 papers
Compilers and program optimization · 46% Programming languages and type systems · 39% Runtime systems and virtual machines · 15%

Topics — the 18 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.131999
The MIT Alewife Machine · Proc. IEEE 1999
Automatic Partitioning of Parallel Loops and Data Arrays for Distributed Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 1995
The MIT Alewife Machine: Architecture and Performance · ISCA 1995
Parallel and multicore computing
parallel programming models
0.031999
The MIT Alewife Machine · Proc. IEEE 1999
Integrating Message-Passing and Shared-Memory: Early Experience · PPoPP 1993
Mul-T: A High-Performance Parallel Lisp · PLDI 1989
Processor architecture and microarchitecture
multiprocessor architecture
0.031995
The MIT Alewife Machine: Architecture and Performance · ISCA 1995
Integrating Message-Passing and Shared-Memory: Early Experience · PPoPP 1993
APRIL: A Processor Architecture for Multiprocessing · ISCA 1990
Memory systems › shared memory
distributed shared memory
0.011999
The MIT Alewife Machine · Proc. IEEE 1999
Parallel and multicore computing › parallel programming models › hybrid programming models
shared memory and message passing
0.011999
The MIT Alewife Machine · Proc. IEEE 1999
Parallel and multicore computing › parallel programming models
message passing
0.021995
Integrating Message-Passing and Shared-Memory: Early Experience · PPoPP 1993
The MIT Alewife Machine: Architecture and Performance · ISCA 1995
Processor architecture and microarchitecture
memory latency tolerance
0.021999
The MIT Alewife Machine · Proc. IEEE 1999
APRIL: A Processor Architecture for Multiprocessing · ISCA 1990
Parallel and multicore computing › multiprocessor system
scalable multiprocessor
0.011995
The MIT Alewife Machine: Architecture and Performance · ISCA 1995
Parallel and multicore computing › parallel programming models
shared-memory parallelization
0.011993
Integrating Message-Passing and Shared-Memory: Early Experience · PPoPP 1993
Parallel and multicore computing
parallel scheduling
0.011991
Lazy Task Creation: A Technique for Increasing the Granularity of Parallel Programs · IEEE Trans. Parallel Distributed Syst. 1991
Parallel and multicore computing › parallel programming models › task parallelism
task granularity control
0.011991
Lazy Task Creation: A Technique for Increasing the Granularity of Parallel Programs · IEEE Trans. Parallel Distributed Syst. 1991
Interconnection networks and networks-on-chip › network topology
mesh network
0.011999
The MIT Alewife Machine · Proc. IEEE 1999
Processor architecture and microarchitecture
multithreading
0.011990
APRIL: A Processor Architecture for Multiprocessing · ISCA 1990
Parallel and multicore computing › parallel computing › parallel programming languages
parallel lisp
0.011989
Mul-T: A High-Performance Parallel Lisp · PLDI 1989
Compilers and program optimization › parallelization
automatic parallelization
0.011995
Automatic Partitioning of Parallel Loops and Data Arrays for Distributed Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 1995
Programming languages and type systems › lambda calculus
combinators
0.011984
A Combinator-Based Compiler for a Functional Language · POPL 1984
Programming languages and type systems › language implementation
graph reduction
0.011984
A Combinator-Based Compiler for a Functional Language · POPL 1984
Parallel and multicore computing › synchronization
synchronization support
0.011990
APRIL: A Processor Architecture for Multiprocessing · ISCA 1990

Methods — techniques the papers use, named apart from their topics

prefetching · 0.0block multithreading · 0.0linear algebra · 0.0lattice theory · 0.0data footprint analysis · 0.0software-extended coherent shared memory · 0.0load-based inlining · 0.0lazy task creation · 0.0coherent caching · 0.0performance modeling · 0.0
YearPublicationVenuePosition
1999 The MIT Alewife Machine
abstract
A variety of models for parallel architectures, such as shared memory, message passing, and data flow, have converged in the recent past to a hybrid architecture form called distributed shared memory (DSM). Alewife, an early prototype of such DSM architectures, uses hybrid software and hardware mechanisms to support coherent shared memory, efficient user level messaging, fine grain synchronization, and latency tolerance. Alewife supports up to 512 processing nodes connected over a scalable and cost effective mesh network at a constant cost per node. Four mechanisms combine to achieve Alewife's goals of scalability and programmability: software extended coherent shared memory provides a global, linear address space; integrated message passing allows compiler and operating system designers to provide efficient communication and synchronization; support for fine grain computation allows many processors to cooperate on small problem sizes; and latency tolerance mechanisms-including block multithreading and prefetching-mask unavoidable delays due to communication. Extensive results from microbenchmarks, together with over a dozen complete applications running on a 32-node prototype, demonstrate that integrating message passing with shared memory enables a cost efficient solution to the cache coherence problem and provides a rich set of programming primitives. Our results further show that messaging and shared memory operations are both important because each helps the programmer to achieve the best performance for various machine configurations.
Anant Agarwal, Ricardo Bianchini, David Chaiken, Fred Chong, Kirk L. Johnson, David A. Kranz, John Kubiatowicz, Beng-Hong Lim, Kenneth Mackenzie, Donald Yeung
Proc. IEEE6
1995 The MIT Alewife Machine: Architecture and Performance
abstract
Alewife is a multiprocessor architecture that supports up to 512 processing nodes connected over a scalable and cost-effective mesh network at a constant cost per node. The MIT Alewife machine, a prototype implementation of the architecture, demonstrates that a parallel system can be both scalable and programmable. Four mechanisms combine to achieve these goals: software-extended coherent shared memory provides a global, linear address space; integrated message passing allows compiler and operating system designers to provide efficient communication and synchronization; support for fine-grain computation allows many processors to cooperate on small problem sizes; and latency tolerance mechanisms --- including block multithreading and prefetching --- mask unavoidable delays due to communication.Microbenchmarks, together with over a dozen complete applications running on the 32-node prototype, help to analyze the behavior of the system. Analysis shows that integrating message passing with shared memory enables a cost-efficient solution to the cache coherence problem and provides a rich set of programming primitives. Block multithreading and prefetching improve performance by up to 25% individually, and 35% together. Finally, language constructs that allow programmers to express fine-grain synchronization can improve performance by over a factor of two.
Anant Agarwal, Ricardo Bianchini, David Chaiken, Kirk L. Johnson, David A. Kranz, John Kubiatowicz, Beng-Hong Lim, Kenneth Mackenzie, Donald Yeung
ISCA5
1995 Automatic Partitioning of Parallel Loops and Data Arrays for Distributed Shared-Memory Multiprocessors
abstract
Presents a theoretical framework for automatically partitioning parallel loops to minimize cache coherency traffic on shared-memory multiprocessors. While several previous papers have looked at hyperplane partitioning of iteration spaces to reduce communication traffic, the problem of deriving the optimal tiling parameters for minimal communication in loops with general affine index expressions has remained open. Our paper solves this open problem by presenting a method for deriving an optimal hyperparallelepiped tiling of iteration spaces for minimal communication in multiprocessors with caches. We show that the same theoretical framework can also be used to determine optimal tiling parameters for both data and loop partitioning in distributed memory multicomputers. Our framework uses matrices to represent iteration and data space mappings and the notion of uniformly intersecting references to capture temporal locality in array references. We introduce the notion of data footprints to estimate the communication traffic between processors and use linear algebraic methods and lattice theory to compute precisely the size of data footprints. We have implemented this framework in a compiler for Alewife, a distributed shared-memory multiprocessor.>
Anant Agarwal, David A. Kranz, Venkat Natarajan
IEEE Trans. Parallel Distributed Syst.2
1993 Automatic Partitioning of Parallel Loops for Cache-Coherent Multiprocessors
abstract
This paper presents a theoretical framework for automatically partitioning parallel loops to minimize cache coherency traffic on shared-memory multiprocessors. While several previous papers have looked at hyperplane partitioning of iteration spaces to reduce communication traffic, the problem of deriving the optimal tiling parameters for minimal communication in loops with general affine index expressions and multiple arrays has remained open. Our paper solves this open problem by presenting a method for deriving an optimal hyperparallelepiped tiling of iteration spaces for minimal communication in multiprocessors with caches. We show that the same theoretical framework can also be used to determine optimal tiling parameters for data and loop partitioning in distributed memory multiprocessors without caches. Like previous papers, our framework uses matrices to represent iteration and data space mappings and the notion of uniformly intersecting references to capture temporal locality in array references. We introduce the notion of data footprints to estimate the communication traffic between processors and use lattice theory to compute precisely the size of data footprints. We have implemented a subset of this framework in a compiler for the Alewife machine.
Anant Agarwal, David A. Kranz, Venkat Natarajan
ICPP (1)2
1993 Integrating Message-Passing and Shared-Memory: Early Experience
abstract
This paper discusses some of the issues involved in implementing a shared-address space programming model on large-scale, distributed-memory multiprocessors. While such a programming model can be implemented on both shared-memory and message-passing architectures, we argue that the transparent, coherent caching of global data provided by many shared-memory architectures is of crucial importance. Because message-passing mechanisms ar much more efficient than shared-memory loads and stores for certain types of interprocessor communication and synchronization operations, hwoever, we argue for building multiprocessors that efficiently support both shared-memory and message-passing mechnisms. We describe an architecture, Alewife, that integrates support for shared-memory and message-passing through a simple interface; we expect the compiler and runtime system to cooperate in using appropriate hardware mechanisms that are most efficient for specific operations. We report on both integrated and exclusively shared-memory implementations of our runtime system and two applications. The integrated runtime system drastically cuts down the cost of communication incurred by the scheduling, load balancing, and certain synchronization operations. We also present preliminary performance results comparing the two systems.
David A. Kranz, Kirk L. Johnson, Anant Agarwal, John Kubiatowicz, Beng-Hong Lim
PPoPP1
1991 Lazy Task Creation: A Technique for Increasing the Granularity of Parallel Programs
abstract
When a parallel algorithm is written naturally, the resulting program often produces tasks of a finer grain than an implementation can exploit efficiently. Two solutions to the granularity problem that combine parallel tasks dynamically at runtime are discussed. The simpler load-based inlining method, in which tasks are combined based on dynamic bad level, is rejected in favor of the safer and more robust lazy task creation method, in which tasks are created only retroactively as processing results become available. The strategies grew out of work on Mul-T, an efficient parallel implementation of Scheme, but could be used with other languages as well. Mul-T implementations of lazy task creation are described for two contrasting machines, and performance statistics that show the method's effectiveness are presented. Lazy task creation is shown to allow efficient execution of naturally expressed algorithms of a substantially finer grain than possible with previous parallel Lisp systems.>
Eric Mohr, David A. Kranz, Robert H. Halstead Jr.
IEEE Trans. Parallel Distributed Syst.2
1990 APRIL: A Processor Architecture for Multiprocessing
abstract
Processors in large-scale multiprocessors must be able to tolerate large communication latencies and synchronization delays. This paper describes the architecture of a rapid-context-switching processor called APRIL with support for fine-grain threads and synchronization. APRIL achieves high single-thread performance and supports virtual dynamic threads. A commercial RISC-based implementation of APRIL and a run-time software system that can switch contexts in about 10 cycles is described. Measurements taken for several parallel applications on an APRIL simulator show that the overhead for supporting parallel tasks based on futures is reduced by a factor of two over a corresponding implementation on the Encore Multimax. The scalability of a multiprocessor based on APRIL is explored using a performance model. We show that the SPARC-based implementation of APRIL can achieve close to 80% processor utilization with as few as three resident threads per processor in a large-scale cache-based machine with an average base network latency of 55 cycles.
Anant Agarwal, Beng-Hong Lim, David A. Kranz, John Kubiatowicz
ISCA3
1989 Mul-T: A High-Performance Parallel Lisp
abstract
Mul-T is a parallel Lisp system, based on Multilisp's future construct, that has been developed to run on an Encore Multimax multiprocessor. Mul-T is an extended version of the Yale T system and uses the T system's ORBIT compiler to achieve “production quality” performance on stock hardware — about 100 times faster than Multilisp. Mul-T shows that futures can be implemented cheaply enough to be useful in a production-quality system. Mul-T is fully operational, including a user interface that supports managing groups of parallel tasks.
David A. Kranz, Robert H. Halstead Jr., Eric Mohr
PLDI1
1984 A Combinator-Based Compiler for a Functional Language
abstract
Article Free Access Share on A combinator-based compiler for a functional language Authors: Paul Hudak Yale University, Department of Computer Science Yale University, Department of Computer ScienceView Profile , David Kranz Yale University, Department of Computer Science Yale University, Department of Computer ScienceView Profile Authors Info & Claims POPL '84: Proceedings of the 11th ACM SIGACT-SIGPLAN symposium on Principles of programming languagesJanuary 1984 Pages 122–132https://doi.org/10.1145/800017.800523Online:15 January 1984Publication History 24citation0DownloadsMetricsTotal Citations24Total Downloads0Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Paul Hudak, David A. Kranz
POPL2