Robert H. Kuhn

dblp:61/2881 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 54% High-performance computing · 20% Interconnection networks and networks-on-chip · 20%
Software engineering, system software, and programming languages
1 paper
Program analysis · 50% Compilers and program optimization · 50%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
cluster computing
0.011992
Low Copy Message Passing on the Alliant CAMPUS/800 · SC 1992
Parallel and multicore computing › parallel programming models
message passing
0.011992
Low Copy Message Passing on the Alliant CAMPUS/800 · SC 1992
Interconnection networks and networks-on-chip › interprocessor communication
message passing protocol
0.011992
Low Copy Message Passing on the Alliant CAMPUS/800 · SC 1992
Parallel and multicore computing
parallel programming models
0.011992
Low Copy Message Passing on the Alliant CAMPUS/800 · SC 1992
Parallel and multicore computing › parallel architecture
massively parallel processor
0.011992
Low Copy Message Passing on the Alliant CAMPUS/800 · SC 1992
Processor architecture and microarchitecture › instruction set architecture › RISC
RISC processor
0.011992
Low Copy Message Passing on the Alliant CAMPUS/800 · SC 1992
Compilers and program optimization › program transformation
compiler transformations
0.011981
Dependence Graphs and Compiler Optimizations · POPL 1981
Program analysis › program representation
dependence graphs
0.011981
Dependence Graphs and Compiler Optimizations · POPL 1981
Parallel and multicore computing › parallelization strategies
algorithm transformation
0.011980
Efficient Mapping of Algorthims To Single-Stage Interconnections · ISCA 1980
Parallel and multicore computing › parallel algorithms › parallel algorithm design
parallel algorithm mapping
0.011980
Efficient Mapping of Algorthims To Single-Stage Interconnections · ISCA 1980
Processor architecture and microarchitecture
instruction-level parallelism
0.011981
Dependence Graphs and Compiler Optimizations · POPL 1981

Methods — techniques the papers use, named apart from their topics

protocol optimization · 0.0low copy message passing · 0.0dependence analysis · 0.0algorithm transformation · 0.0
YearPublicationVenuePosition
2012 Hierarchical overlapped tiling
abstract
This paper introduces hierarchical overlapped tiling, a transformation that applies loop tiling and fusion to conventional loops. Overlapped tiling is a useful transformation to reduce communication overhead, but it may also generate a significant amount of redundant computation. Hierarchical overlapped tiling performs overlapped tiling hierarchically to balance communication overhead and redundant computation, and thus has the potential to provide better performance.
Xing Zhou 0002, Jean-Pierre Giacalone, María Jesús Garzarán, Robert H. Kuhn, David A. Padua
CGO4
2006 Speeding up NGB with distributed file streaming framework
abstract
Grid computing provides a very rich environment for scientific calculations. In addition to the challenges it provides, it also offers new opportunities for optimization. In this paper we have utilized DFS (distributed file streaming) framework to speed up NAS grid benchmark workflows. By studying I/O patterns of NGB codes we have identified program locations where it is possible to overlap computation and data workflow phases. By integrating DFS into NGB, we demonstrate a useful method of improving overall workflow efficiency by streaming the output of the current process to make an input of the following stage, reducing a workflow to a series of distributed producer consumer stages. DFS framework eliminates file transfers and in the process makes process scheduling more efficient, leading to overall performance improvements in the turnaround time for HC (helical chain) data flow graph under Globus grid environment with the embedded DFS over the original version of the benchmark
Zhiteng Huang, Hrabri L. Rajic, Robert H. Kuhn
IPDPS5
1992 Low Copy Message Passing on the Alliant CAMPUS/800
abstract
With modern massively parallel processors consisting of hundreds of RISC (reduced instruction set computer) processors, a significant portion of program run time is being spent in passing messages between processors. The authors address the problem of reducing unnecessary copies in message passing protocols. Message passing that avoids using copies in the application program and the transport mechanism is called low copy message passing. The authors describe the low copy message passing library on the Alliant CAMPUS/800 system. First, a way to reduce the number of copies in a message passing protocol is proposed. In gross machine independent terms, the message passing protocol can be reduced from seven major steps to one. The authors then describe how application programmers can use the low copy message passing library. Finally, a brief survey of applications is presented which shows how applications use the low copy library.>
Charles M. Burns, Robert H. Kuhn, Eric J. Werme
SC2
1981 Dependence Graphs and Compiler Optimizations
abstract
Dependence graphs can be used as a vehicle for formulating and implementing compiler optimizations. This paper defines such graphs and discusses two kinds of transformations. The first are simple rewriting transformations that remove dependence arcs. The second are abstraction transformations that deal more globally with a dependence graph. These transformations have been implemented and applied to several different types of high-speed architectures.
David J. Kuck, Robert H. Kuhn, David A. Padua, Bruce Leasure, Michael Wolfe
POPL2
1980 Efficient Mapping of Algorthims To Single-Stage Interconnections
abstract
In this paper, we consider the problem of restructuring or transforming algorithms to efficiently use a single-stage interconnection network. All algorithms contain some freedom in the way they are mapped to a machine. We use this freedom to show that superior interconnection efficiency can be obtained by implementing the interconnections required by the algorithm within the context of the algorithm rather than attempting to implement each request individually. The interconnection considered is the bidirectional shuffle-shift. It is shown that two algorithm transformations are useful for implementing several lower triangular and tridiagonal system algorithms on the shuffle-shift network. Of the 14 algorithms considered, 85% could be implemented on this network. The transformations developed to produce these results are described. They are general-purpose in nature and can be applied to a much larger class of algorithms.
Robert H. Kuhn
ISCA1