EDBT 2026 Demo / reviewers in the wild / expert
David A. Kranz
dblp:53/6629
· DBLP profile ↗
9ranked-venue papers
2as first author
0since 2021 · last 1999
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Parallel and multicore computing · 51% Memory systems · 27% Processor architecture and microarchitecture · 19% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 46% Programming languages and type systems · 39% Runtime systems and virtual machines · 15% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache coherence |
0.1 | 3 | 1999 | The MIT Alewife Machine · Proc. IEEE 1999 Automatic Partitioning of Parallel Loops and Data Arrays for Distributed Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 1995 The MIT Alewife Machine: Architecture and Performance · ISCA 1995 |
Parallel and multicore computing
parallel programming models |
0.0 | 3 | 1999 | The MIT Alewife Machine · Proc. IEEE 1999 Integrating Message-Passing and Shared-Memory: Early Experience · PPoPP 1993 Mul-T: A High-Performance Parallel Lisp · PLDI 1989 |
Processor architecture and microarchitecture
multiprocessor architecture |
0.0 | 3 | 1995 | The MIT Alewife Machine: Architecture and Performance · ISCA 1995 Integrating Message-Passing and Shared-Memory: Early Experience · PPoPP 1993 APRIL: A Processor Architecture for Multiprocessing · ISCA 1990 |
Memory systems › shared memory
distributed shared memory |
0.0 | 1 | 1999 | The MIT Alewife Machine · Proc. IEEE 1999 |
Parallel and multicore computing › parallel programming models › hybrid programming models
shared memory and message passing |
0.0 | 1 | 1999 | The MIT Alewife Machine · Proc. IEEE 1999 |
Parallel and multicore computing › parallel programming models
message passing |
0.0 | 2 | 1995 | Integrating Message-Passing and Shared-Memory: Early Experience · PPoPP 1993 The MIT Alewife Machine: Architecture and Performance · ISCA 1995 |
Processor architecture and microarchitecture
memory latency tolerance |
0.0 | 2 | 1999 | The MIT Alewife Machine · Proc. IEEE 1999 APRIL: A Processor Architecture for Multiprocessing · ISCA 1990 |
Parallel and multicore computing › multiprocessor system
scalable multiprocessor |
0.0 | 1 | 1995 | The MIT Alewife Machine: Architecture and Performance · ISCA 1995 |
Parallel and multicore computing › parallel programming models
shared-memory parallelization |
0.0 | 1 | 1993 | Integrating Message-Passing and Shared-Memory: Early Experience · PPoPP 1993 |
Parallel and multicore computing
parallel scheduling |
0.0 | 1 | 1991 | Lazy Task Creation: A Technique for Increasing the Granularity of Parallel Programs · IEEE Trans. Parallel Distributed Syst. 1991 |
Parallel and multicore computing › parallel programming models › task parallelism
task granularity control |
0.0 | 1 | 1991 | Lazy Task Creation: A Technique for Increasing the Granularity of Parallel Programs · IEEE Trans. Parallel Distributed Syst. 1991 |
Interconnection networks and networks-on-chip › network topology
mesh network |
0.0 | 1 | 1999 | The MIT Alewife Machine · Proc. IEEE 1999 |
Processor architecture and microarchitecture
multithreading |
0.0 | 1 | 1990 | APRIL: A Processor Architecture for Multiprocessing · ISCA 1990 |
Parallel and multicore computing › parallel computing › parallel programming languages
parallel lisp |
0.0 | 1 | 1989 | Mul-T: A High-Performance Parallel Lisp · PLDI 1989 |
Compilers and program optimization › parallelization
automatic parallelization |
0.0 | 1 | 1995 | Automatic Partitioning of Parallel Loops and Data Arrays for Distributed Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 1995 |
Programming languages and type systems › lambda calculus
combinators |
0.0 | 1 | 1984 | A Combinator-Based Compiler for a Functional Language · POPL 1984 |
Programming languages and type systems › language implementation
graph reduction |
0.0 | 1 | 1984 | A Combinator-Based Compiler for a Functional Language · POPL 1984 |
Parallel and multicore computing › synchronization
synchronization support |
0.0 | 1 | 1990 | APRIL: A Processor Architecture for Multiprocessing · ISCA 1990 |
Methods — techniques the papers use, named apart from their topics
prefetching · 0.0block multithreading · 0.0linear algebra · 0.0lattice theory · 0.0data footprint analysis · 0.0software-extended coherent shared memory · 0.0load-based inlining · 0.0lazy task creation · 0.0coherent caching · 0.0performance modeling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 1999 | The MIT Alewife MachineabstractA variety of models for parallel architectures, such as shared memory, message passing, and data flow, have converged in the recent past to a hybrid architecture form called distributed shared memory (DSM). Alewife, an early prototype of such DSM architectures, uses hybrid software and hardware mechanisms to support coherent shared memory, efficient user level messaging, fine grain synchronization, and latency tolerance. Alewife supports up to 512 processing nodes connected over a scalable and cost effective mesh network at a constant cost per node. Four mechanisms combine to achieve Alewife's goals of scalability and programmability: software extended coherent shared memory provides a global, linear address space; integrated message passing allows compiler and operating system designers to provide efficient communication and synchronization; support for fine grain computation allows many processors to cooperate on small problem sizes; and latency tolerance mechanisms-including block multithreading and prefetching-mask unavoidable delays due to communication. Extensive results from microbenchmarks, together with over a dozen complete applications running on a 32-node prototype, demonstrate that integrating message passing with shared memory enables a cost efficient solution to the cache coherence problem and provides a rich set of programming primitives. Our results further show that messaging and shared memory operations are both important because each helps the programmer to achieve the best performance for various machine configurations. Anant Agarwal, Ricardo Bianchini, David Chaiken, Fred Chong, Kirk L. Johnson, David A. Kranz, John Kubiatowicz, Beng-Hong Lim, Kenneth Mackenzie, Donald Yeung |
Proc. IEEE | 6 |
| 1995 | The MIT Alewife Machine: Architecture and PerformanceabstractAlewife is a multiprocessor architecture that supports up to 512 processing nodes connected over a scalable and cost-effective mesh network at a constant cost per node. The MIT Alewife machine, a prototype implementation of the architecture, demonstrates that a parallel system can be both scalable and programmable. Four mechanisms combine to achieve these goals: software-extended coherent shared memory provides a global, linear address space; integrated message passing allows compiler and operating system designers to provide efficient communication and synchronization; support for fine-grain computation allows many processors to cooperate on small problem sizes; and latency tolerance mechanisms --- including block multithreading and prefetching --- mask unavoidable delays due to communication.Microbenchmarks, together with over a dozen complete applications running on the 32-node prototype, help to analyze the behavior of the system. Analysis shows that integrating message passing with shared memory enables a cost-efficient solution to the cache coherence problem and provides a rich set of programming primitives. Block multithreading and prefetching improve performance by up to 25% individually, and 35% together. Finally, language constructs that allow programmers to express fine-grain synchronization can improve performance by over a factor of two. Anant Agarwal, Ricardo Bianchini, David Chaiken, Kirk L. Johnson, David A. Kranz, John Kubiatowicz, Beng-Hong Lim, Kenneth Mackenzie, Donald Yeung |
ISCA | 5 |
| 1995 | Automatic Partitioning of Parallel Loops and Data Arrays for Distributed Shared-Memory MultiprocessorsabstractPresents a theoretical framework for automatically partitioning parallel loops to minimize cache coherency traffic on shared-memory multiprocessors. While several previous papers have looked at hyperplane partitioning of iteration spaces to reduce communication traffic, the problem of deriving the optimal tiling parameters for minimal communication in loops with general affine index expressions has remained open. Our paper solves this open problem by presenting a method for deriving an optimal hyperparallelepiped tiling of iteration spaces for minimal communication in multiprocessors with caches. We show that the same theoretical framework can also be used to determine optimal tiling parameters for both data and loop partitioning in distributed memory multicomputers. Our framework uses matrices to represent iteration and data space mappings and the notion of uniformly intersecting references to capture temporal locality in array references. We introduce the notion of data footprints to estimate the communication traffic between processors and use linear algebraic methods and lattice theory to compute precisely the size of data footprints. We have implemented this framework in a compiler for Alewife, a distributed shared-memory multiprocessor.> Anant Agarwal, David A. Kranz, Venkat Natarajan |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1993 | Automatic Partitioning of Parallel Loops for Cache-Coherent MultiprocessorsabstractThis paper presents a theoretical framework for automatically partitioning parallel loops to minimize cache coherency traffic on shared-memory multiprocessors. While several previous papers have looked at hyperplane partitioning of iteration spaces to reduce communication traffic, the problem of deriving the optimal tiling parameters for minimal communication in loops with general affine index expressions and multiple arrays has remained open. Our paper solves this open problem by presenting a method for deriving an optimal hyperparallelepiped tiling of iteration spaces for minimal communication in multiprocessors with caches. We show that the same theoretical framework can also be used to determine optimal tiling parameters for data and loop partitioning in distributed memory multiprocessors without caches. Like previous papers, our framework uses matrices to represent iteration and data space mappings and the notion of uniformly intersecting references to capture temporal locality in array references. We introduce the notion of data footprints to estimate the communication traffic between processors and use lattice theory to compute precisely the size of data footprints. We have implemented a subset of this framework in a compiler for the Alewife machine. Anant Agarwal, David A. Kranz, Venkat Natarajan |
ICPP (1) | 2 |
| 1993 | Integrating Message-Passing and Shared-Memory: Early ExperienceabstractThis paper discusses some of the issues involved in implementing a shared-address space programming model on large-scale, distributed-memory multiprocessors. While such a programming model can be implemented on both shared-memory and message-passing architectures, we argue that the transparent, coherent caching of global data provided by many shared-memory architectures is of crucial importance. Because message-passing mechanisms ar much more efficient than shared-memory loads and stores for certain types of interprocessor communication and synchronization operations, hwoever, we argue for building multiprocessors that efficiently support both shared-memory and message-passing mechnisms. We describe an architecture, Alewife, that integrates support for shared-memory and message-passing through a simple interface; we expect the compiler and runtime system to cooperate in using appropriate hardware mechanisms that are most efficient for specific operations. We report on both integrated and exclusively shared-memory implementations of our runtime system and two applications. The integrated runtime system drastically cuts down the cost of communication incurred by the scheduling, load balancing, and certain synchronization operations. We also present preliminary performance results comparing the two systems. David A. Kranz, Kirk L. Johnson, Anant Agarwal, John Kubiatowicz, Beng-Hong Lim |
PPoPP | 1 |
| 1991 | Lazy Task Creation: A Technique for Increasing the Granularity of Parallel ProgramsabstractWhen a parallel algorithm is written naturally, the resulting program often produces tasks of a finer grain than an implementation can exploit efficiently. Two solutions to the granularity problem that combine parallel tasks dynamically at runtime are discussed. The simpler load-based inlining method, in which tasks are combined based on dynamic bad level, is rejected in favor of the safer and more robust lazy task creation method, in which tasks are created only retroactively as processing results become available. The strategies grew out of work on Mul-T, an efficient parallel implementation of Scheme, but could be used with other languages as well. Mul-T implementations of lazy task creation are described for two contrasting machines, and performance statistics that show the method's effectiveness are presented. Lazy task creation is shown to allow efficient execution of naturally expressed algorithms of a substantially finer grain than possible with previous parallel Lisp systems.> Eric Mohr, David A. Kranz, Robert H. Halstead Jr. |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1990 | APRIL: A Processor Architecture for MultiprocessingabstractProcessors in large-scale multiprocessors must be able to tolerate large communication latencies and synchronization delays. This paper describes the architecture of a rapid-context-switching processor called APRIL with support for fine-grain threads and synchronization. APRIL achieves high single-thread performance and supports virtual dynamic threads. A commercial RISC-based implementation of APRIL and a run-time software system that can switch contexts in about 10 cycles is described. Measurements taken for several parallel applications on an APRIL simulator show that the overhead for supporting parallel tasks based on futures is reduced by a factor of two over a corresponding implementation on the Encore Multimax. The scalability of a multiprocessor based on APRIL is explored using a performance model. We show that the SPARC-based implementation of APRIL can achieve close to 80% processor utilization with as few as three resident threads per processor in a large-scale cache-based machine with an average base network latency of 55 cycles. Anant Agarwal, Beng-Hong Lim, David A. Kranz, John Kubiatowicz |
ISCA | 3 |
| 1989 | Mul-T: A High-Performance Parallel LispabstractMul-T is a parallel Lisp system, based on Multilisp's future construct, that has been developed to run on an Encore Multimax multiprocessor. Mul-T is an extended version of the Yale T system and uses the T system's ORBIT compiler to achieve “production quality” performance on stock hardware — about 100 times faster than Multilisp. Mul-T shows that futures can be implemented cheaply enough to be useful in a production-quality system. Mul-T is fully operational, including a user interface that supports managing groups of parallel tasks. David A. Kranz, Robert H. Halstead Jr., Eric Mohr |
PLDI | 1 |
| 1984 | A Combinator-Based Compiler for a Functional LanguageabstractArticle Free Access Share on A combinator-based compiler for a functional language Authors: Paul Hudak Yale University, Department of Computer Science Yale University, Department of Computer ScienceView Profile , David Kranz Yale University, Department of Computer Science Yale University, Department of Computer ScienceView Profile Authors Info & Claims POPL '84: Proceedings of the 11th ACM SIGACT-SIGPLAN symposium on Principles of programming languagesJanuary 1984 Pages 122–132https://doi.org/10.1145/800017.800523Online:15 January 1984Publication History 24citation0DownloadsMetricsTotal Citations24Total Downloads0Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Paul Hudak, David A. Kranz |
POPL | 2 |