Edgar T. Kalns

dblp:66/1669 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
0since 2021 · last 1995
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 63% High-performance computing · 20% Interconnection networks and networks-on-chip · 15%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › data distribution
data decomposition
0.021995
Processor Mapping Techniques Toward Efficient Data Redistribution · IEEE Trans. Parallel Distributed Syst. 1995
Issues in scalable library design for massively parallel computers · SC 1993
Parallel and multicore computing › data distribution
data redistribution
0.011995
Processor Mapping Techniques Toward Efficient Data Redistribution · IEEE Trans. Parallel Distributed Syst. 1995
High-performance computing
distributed memory systems
0.011995
Processor Mapping Techniques Toward Efficient Data Redistribution · IEEE Trans. Parallel Distributed Syst. 1995
Parallel and multicore computing › processor allocation
processor mapping
0.011995
Processor Mapping Techniques Toward Efficient Data Redistribution · IEEE Trans. Parallel Distributed Syst. 1995
Parallel and multicore computing › parallel libraries
communication library
0.011992
ComPaSS: Efficient Communication Services for Scalable Architectures · SC 1992
Parallel and multicore computing › parallel computing › parallel communication
global communication
0.011992
ComPaSS: Efficient Communication Services for Scalable Architectures · SC 1992
Interconnection networks and networks-on-chip
multicast
0.011992
ComPaSS: Efficient Communication Services for Scalable Architectures · SC 1992
Interconnection networks and networks-on-chip › routing algorithms
wormhole routing
0.011992
ComPaSS: Efficient Communication Services for Scalable Architectures · SC 1992
Parallel and multicore computing
runtime optimization
0.011995
Processor Mapping Techniques Toward Efficient Data Redistribution · IEEE Trans. Parallel Distributed Syst. 1995
Parallel and multicore computing › parallel architecture
massively parallel processor
0.011992
ComPaSS: Efficient Communication Services for Scalable Architectures · SC 1992

Methods — techniques the papers use, named apart from their topics

performance measurement · 0.0layered library design · 0.0
YearPublicationVenuePosition
1995 Processor Mapping Techniques Toward Efficient Data Redistribution
abstract
Run-time data redistribution can enhance algorithm performance in distributed-memory machines. Explicit redistribution of data can be performed between algorithm phases when a different data decomposition is expected to deliver increased performance for a subsequent phase of computation. Redistribution, however, represents increased program overhead as algorithm computation is discontinued while data are exchanged among processor memories. In this paper, we present a technique that minimizes the amount of data exchange for BLOCK to CYCLIC(c) (or vice-versa) redistributions of arbitrary number of dimensions. Preserving the semantics of the target (destination) distribution pattern, the technique manipulates the data to logical processor mapping of the target pattern. When implemented on an IBM SP, the mapping technique demonstrates redistribution performance improvements of approximately 40% over traditional data to processor mapping. Relative to the traditional mapping technique, the proposed method affords greater flexibility in specifying precisely which data elements are redistributed and which elements remain on-processor.
Edgar T. Kalns, Lionel M. Ni
IEEE Trans. Parallel Distributed Syst.1
1994 ComPaSS: A Communication Package for Scalable Software Design
Hong Xu 0005, Edgar T. Kalns, Philip K. McKinley, Lionel M. Ni
J. Parallel Distributed Comput.2
1993 Evaluation of Data Distirbution Patterns in Distributed-Memory Machines
abstract
Determining an appropriate data distribution among different memories is critical to the performance of data-parallel programs on distributedmemory machines. By analyzing the computational load of data arrays and the communication cornplexity of various data movement operations in a program, this paper suggests a first-order cost model for determining a small set of appropriate data distribution patterns among many possible choices. A new data distribution specification, name! y CYBLOCK, is proposed to enhance the expressiveness of data distribution specifications being proposed in High Performance Fortran. Cost analysis of two case studies: a linear system solver and a Purdue-set benchmark loop, are used to illustrate the proposed evaluation method. The model correctly predicts the relative performance of the case studies when implemented with various regular data distributions on an nCUBE- 2 multicomputer.
Edgar T. Kalns, Hong Xu 0005, Lionel M. Ni
ICPP (2)1
1993 Issues in scalable library design for massively parallel computers
abstract
This paper examines some crtttcal tssues raised in the design of libraries for MPCS, such as scalability, portability, recompilation, and flexibility.we adUocate a layered structure of it brary design, comprising a high-level language layer, a machine-independent node layer, a machine-dependent node layer, and an object code layer for diflerent demands and requirements.We discuss the impact of various data decomposition strategies on program performance and the computation and communication analysts techniques employed at different layers.We also propose the concept of the range of scalability as a metric for selectzng the most appropriate implementation.A linear system solver based on the Gaussian elimination method is used as an example to illustrate various design alternates.nessee and Oak Ridge National Laboratory attempts to provide a scalable LAPACK package for dense and banded matrix computations [3].The ComPaSS library being developed at Michigan State University attempts to provide a set of scalable communication
Lionel M. Ni, Hong Xu 0005, Edgar T. Kalns
SC3
1992 ComPaSS: Efficient Communication Services for Scalable Architectures
abstract
The authors describe the initial implementation of the ComPaSS communication library to support scalable software development in massively parallel processors. ComPaSS provides high-level global communication operations for both data manipulation and process control, many of which are based on a small set of low-level communication primitives. The ComPaSS library is unique in that these low-level operations are provably optimal for a class of architectures representative of many commercial scalable systems-in particular, those using wormhole routing and n-dimensional mesh network topologies. The authors concentrate on the multicast component of the ComPaSS library, which is useful in several data parallel operations. The design of the multicast primitive is described, and an example of its use in a data parallel application is given. Improvements in performance resulting from use of the library on a 64-node nCUBE-2 are presented.>
Philip K. McKinley, Hong Xu 0005, Edgar T. Kalns, Lionel M. Ni
SC3