Duncan H. Lawrie

dblp:32/1299 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
0since 2021 · last 1994
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 5 first-authorSoftware engineering, systems software and programming languages · 2Theory of computation · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
13 papers
Performance modeling and evaluation · 35% Interconnection networks and networks-on-chip · 21% Memory systems · 20%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 79% Operating systems · 21%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
0.011993
The Cedar System and an Initial Performance Study · ISCA 1993
Performance modeling and evaluation › parallel system performance
multiprocessor performance evaluation
0.011993
The Cedar System and an Initial Performance Study · ISCA 1993
Performance modeling and evaluation
parallel performance evaluation
0.011993
The Cedar System and an Initial Performance Study · ISCA 1993
Interconnection networks and networks-on-chip › switching network
multistage interconnection network
0.051987
Performance Analysis of Redundant-Path Networks for Multiprocessor Systems · ACM Trans. Comput. Syst. 1985
A Class of Redundant Path Multistage Interconnection Networks · IEEE Trans. Computers 1983
An Easily Controlled Network for Frequently Used Permutation · IEEE Trans. Computers 1981
Parallel and multicore computing › parallel computing
parallel programming languages
0.011988
Cedar Fortran and other Vector and parallel Fortran dialects · SC 1988
Processor architecture and microarchitecture
vector processor
0.011988
Cedar Fortran and other Vector and parallel Fortran dialects · SC 1988
Memory systems
cache design
0.011987
Multiprocessor Cache Design Considerations · ISCA 1987
Interconnection networks and networks-on-chip › network contention
hot-spot contention
0.011987
Distributing Hot-Spot Addressing in Large-Scale Multiprocessors · IEEE Trans. Computers 1987
Memory systems › memory interference
memory contention
0.011987
Distributing Hot-Spot Addressing in Large-Scale Multiprocessors · IEEE Trans. Computers 1987
Memory systems › cache
multiprocessor cache
0.011987
Multiprocessor Cache Design Considerations · ISCA 1987
Interconnection networks and networks-on-chip
network contention
0.011987
Distributing Hot-Spot Addressing in Large-Scale Multiprocessors · IEEE Trans. Computers 1987
Performance modeling and evaluation
markov models
0.011985
Performance Analysis of Redundant-Path Networks for Multiprocessor Systems · ACM Trans. Comput. Syst. 1985
Performance modeling and evaluation › benchmarking
parallel benchmark
0.011993
The Cedar System and an Initial Performance Study · ISCA 1993
Memory systems › memory access
parallel memory access
0.021982
The Prime Memory System for Array Access · IEEE Trans. Computers 1982
Access and Alignment of Data in an Array Processor · IEEE Trans. Computers 1975
Performance modeling and evaluation
workload characterization
0.011993
The Cedar System and an Initial Performance Study · ISCA 1993
Interconnection networks and networks-on-chip › switching network › multistage interconnection network
shuffle-exchange network
0.021981
An Easily Controlled Network for Frequently Used Permutation · IEEE Trans. Computers 1981
Access and Alignment of Data in an Array Processor · IEEE Trans. Computers 1975
Parallel and multicore computing › multiprocessor system
large-scale multiprocessor
0.021987
Distributing Hot-Spot Addressing in Large-Scale Multiprocessors · IEEE Trans. Computers 1987
Multiprocessor Cache Design Considerations · ISCA 1987
Hardware reliability and fault tolerance
network fault tolerance
0.011983
A Class of Redundant Path Multistage Interconnection Networks · IEEE Trans. Computers 1983
Parallel and multicore computing
multiprocessor system
0.021987
Multiprocessor Cache Design Considerations · ISCA 1987
Performance Analysis of Redundant-Path Networks for Multiprocessor Systems · ACM Trans. Comput. Syst. 1985
Parallel and multicore computing
parallel algorithms
0.011982
A Practical Algorithm for the Solution of Triangular Systems on a Parallel Processing System · IEEE Trans. Computers 1982
Memory systems › memory architecture
prime memory system
0.011982
The Prime Memory System for Array Access · IEEE Trans. Computers 1982
Compilers and program optimization › memory optimization
data locality optimization
0.011981
On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981
Compilers and program optimization
program transformation
0.011981
On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981
Storage systems
paging performance
0.011981
On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981
Interconnection networks and networks-on-chip
permutation network
0.011981
An Easily Controlled Network for Frequently Used Permutation · IEEE Trans. Computers 1981
Memory systems › memory management
virtual memory
0.011981
On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981
Processor architecture and microarchitecture › multiprocessor architecture
multiprocessor design
0.011980
High-Speed Multiprocessors and Compilation Techniques · IEEE Trans. Computers 1980
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.011988
Cedar Fortran and other Vector and parallel Fortran dialects · SC 1988
Memory systems › memory architecture
interleaved memory
0.011977
On the Effective Bandwidth of Parallel Memories · IEEE Trans. Computers 1977
Performance modeling and evaluation › performance model construction
memory system performance modeling
0.011977
On the Effective Bandwidth of Parallel Memories · IEEE Trans. Computers 1977

Methods — techniques the papers use, named apart from their topics

performance measurement · 0.0benchmarking methodology · 0.0source-to-source transformation · 0.0request combining · 0.0program analysis · 0.0LRU replacement · 0.0markov modeling · 0.0indexing hardware · 0.0data alignment switches · 0.0LU decomposition · 0.0
YearPublicationVenuePosition
1994 Editor's Notice
Duncan H. Lawrie
IEEE Trans. Parallel Distributed Syst.1
1994 Editor's Notice
Duncan H. Lawrie
IEEE Trans. Parallel Distributed Syst.1
1994 Introduction of New Associate Editor
Duncan H. Lawrie
IEEE Trans. Parallel Distributed Syst.1
1993 The Cedar System and an Initial Performance Study
abstract
In this paper, we give an overview of the Cedar multiprocessor and present recent performance results. These include the performance of some computational kernels and the Perfect Benchmarks. We also present a methodology for judging parallel system performance and apply this methodology to Cedar, Cray YMP-8, and Thinking Machines CM-5.
David J. Kuck, Edward S. Davidson, Duncan H. Lawrie, Ahmed H. Sameh, Chuanqi Zhu, Alexander V. Veidenbaum, Jeff Konicek, Pen-Chung Yew, Kyle A. Gallivan, William Jalby, Harry A. G. Wijshoff, Randall Bramley, Ulrike Meier Yang, Perry A. Emrath, David A. Padua, Rudolf Eigenmann, Jay P. Hoeflinger, Greg P. Jaxon, Zhiyuan Li 0001, T. Murphy, John T. Andrews, Stephen W. Turner
ISCA3
1990 Cedar Fortran and other vector and parallel Fortran dialects
Mark D. Guzzi, David A. Padua, Jay P. Hoeflinger, Duncan H. Lawrie
J. Supercomput.4
1988 Cedar Fortran and other Vector and parallel Fortran dialects
abstract
The development of vector and multiprocessor language constructs in Fortran is outlined. The significant architectures, their languages, and optimizers are described. A description is given of Cedar Fortran, the language for the Cedar multiprocessor, a hierarchical, shared-memory, vector multiprocessor currently under development.>
Mark D. Guzzi, Jay P. Hoeflinger, David A. Padua, Duncan H. Lawrie
SC4
1987 : Data Prefetching In Shared Memory Multiprocessors
Roland L. Lee, Pen-Chung Yew, Duncan H. Lawrie
ICPP3
1987 Multiprocessor Cache Design Considerations
abstract
In this paper, cache design is explored for large high-performance multiprocessors with hundreds or thousands of processors and memory modules interconnected by a pipe-lined multi-stage network. The majority of the multiprocessor cache studies in the literature exclusively focus on the issue of cache coherence enforcement. However, there are other characteristics unique to such multiprocessors which create an environment for cache performance that is very different from that of many uniprocessors.
Roland L. Lee, Pen-Chung Yew, Duncan H. Lawrie
ISCA3
1987 Distributing Hot-Spot Addressing in Large-Scale Multiprocessors
abstract
When a large number of processors try to access a common variable, referred to as hot-spot accesses in [6], not only can the resulting memory contention seriously degrade performance, but it can also cause tree saturation in the interconnection network which blocks both hot and regular requests alike. It is shown in [6] that even if only a small percentage of all requests are to a hot-spot, these requests can cause very serious performances problems, and networks that do the necessary combining of requests are suggested to keep the interconnection network and memory contention from becoming a bottleneck.
Pen-Chung Yew, Nian-Feng Tzeng, Duncan H. Lawrie
IEEE Trans. Computers3
1986 Distributing Hot-Spot Addressing in Large Scale Multiprocessor
Pen-Chung Yew, Nian-Feng Tzeng, Duncan H. Lawrie
ICPP3
1985 Performance Analysis of Redundant-Path Networks for Multiprocessor Systems
abstract
Performance of a class of multistage interconnection networks employing redundant paths is investigated. Redundant path networks provide significant tolerance to faults at minimal costs; in this paper improvements in performance and very graceful degradation are also shown to result from the availability of redundant paths. A Markov model is introduced for the operation of these networks in the circuit-switched mode and is solved numerically to obtain the performance measures of interest. The structure of the networks that provide maximal performance is also characterized.
Krishnan Padmanabhan, Duncan H. Lawrie
ACM Trans. Comput. Syst.2
1985 Corrections to "The Computation and Communication Complexity of a Parallel Banded System Solver"
abstract
No abstract available.
Duncan H. Lawrie, Ahmed H. Sameh
ACM Trans. Math. Softw.1
1984 The computation and communication complexity of a parallel banded system solver
abstract
We present an algorithm for solving banded positive defimte linear systems on a multiprocessor computer whose number of processors p is much less than the order of the system n.Assuming that the banded matrix, of bandwidth 2m + 1, is stored in the global memory by diagonals as several onedlmensmnal arrays, we consider the time required by several alignment networks for allocating the appropriate data to the local memory of each processor.We demonstrate that the time required in this preprocessmg stage does not exceed that required by the algorithm provided we use a shuffle exchange, a plpelmed shuffle exchange, or a crossbar switch.Once the data are allocated in the local memories, the algorithm requires only a "nearest neighbor" alignment network to achieve a total time of O(m'~n/p).The total cost of the algorithm is minimized when p ~ ~.
Duncan H. Lawrie, Ahmed H. Sameh
ACM Trans. Math. Softw.1
1983 Cedar : A Large Scale Multiprocessor
Daniel Gajski, David J. Kuck, Duncan H. Lawrie, Ahmed H. Sameh
ICPP3
1983 Fault Tolerance Schemes in Shuffle-Exchange Type Interconnection Networks
Krishnan Padmanabhan, Duncan H. Lawrie
ICPP2
1983 A Class of Redundant Path Multistage Interconnection Networks
abstract
A general class of fault-tolerant multistage interconnection networks is presented, wherein fault-tolerance is achieved by providing multiple disjoint paths between every input and output. These networks are derived from the Omega networks and as such retain all the connection properties of the parent networks in the absence of faults. An R-path network in this class can tolerate (R-1) arbitrary faults in the intermediate stages of the network at a cost that is far less than providing R copies of the original network. Different techniques for constructing such networks are presented and relevant properties and control algorithms are investigated.
Krishnan Padmanabhan, Duncan H. Lawrie
IEEE Trans. Computers2
1982 Performance of packet switching in buffered single-stage shuffle-exchange networks
Pin-Yee Chen, Pen-Chung Yew, Duncan H. Lawrie
ICDCS3
1982 A fault tolerant interconnection network using error correcting codes
J. Edward Lilienkamp, Duncan H. Lawrie, Pen-Chung Yew
ICPP2
1982 The Prime Memory System for Array Access
abstract
In this paper we describe a memory system designed for parallel array access. The system is based on the use of a prime nwnber of memories and a powerful combination of indexing hardware and data alignment switches. Particular emphasis is placed on the indexing equations and their implementation.
Duncan H. Lawrie, Chandra R. Vora
IEEE Trans. Computers1
1982 A Practical Algorithm for the Solution of Triangular Systems on a Parallel Processing System
abstract
An algorithm is presented for a more efficient and implementable solution of triangular systems on a parallel (SIMD) computer which requires 0(log (N)) fewer processing cycles than the best previous results, where N is the system size. We will also show that the data can be accessed and aligned in the same order of time using as many memory units as processors and Ω networks for data alignment. (Previous results dealing with this type of algorithm have not dealt in any detail with the problem of data access and alignment.)
Robert K. Montoye, Duncan H. Lawrie
IEEE Trans. Computers2
1981 On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations
abstract
It is possible to improve the paging performance of a program by applying transformations to the source program that improve data access locality. We discuss this subject in general terms, including automation of these transformations, and present a number of such transforms. This is followed by experimental results which indicate that these transformations are indeed effective. Use of practical, simple memory management policies like the fixed allocation local lru replacement algorithm leads to average improvements over untransformed programs of a factor of 10 in space-time cost, and a factor of 5 in memory size. Multiprogramming questions are also discussed.
Walid A. Abu-Sufah, David J. Kuck, Duncan H. Lawrie
IEEE Trans. Computers3
1981 An Easily Controlled Network for Frequently Used Permutation
abstract
A π network, which is a concatenation of 2 Ω networks [2], along with a simple control algorithm is proposed. This network is capable of performing all Ω network realizable permutations and the bit-permute-complement (BPC) class of permutations[5] in 0(log N) time. The control algorithm is actually a multiple-pass control algorithm on the Ω network, which is more general than Pease's LU decomposition method [6] and Lenfant's decomposition method[4].
Pen-Chung Yew, Duncan H. Lawrie
IEEE Trans. Computers2
1980 High-Speed Multiprocessors and Compilation Techniques
abstract
The purpose of this paper is to present some ideas on multiprocessor design and on automatic translation of sequential programs into parallel programs for multiprocessors. With respect to machine design, two subjects are discussed. First, a multiprocessor allowing parallelism at a very low level is sketched and then, a brief discussion on the interconnection network is presented.
David A. Padua, David J. Kuck, Duncan H. Lawrie
IEEE Trans. Computers3
1977 On the Effective Bandwidth of Parallel Memories
abstract
The object of this paper is to bring together several models of interleaved or parallel memory systems and to expose some of the underlying assumptions about the address streams in each model. We derive the performance for each model, either analytically or by simulation, and discuss why it yields better or worse performance than other models (e.g., because of dependencies in the address stream or hardware queues, etc.). We also show that the performance of a properly designed system can be a linear rather than a square root function of the number of memories and processors.
Donald Y. Chang, David J. Kuck, Duncan H. Lawrie
IEEE Trans. Computers3
1975 Access and Alignment of Data in an Array Processor
abstract
This paper discusses the design of a primary memory system for an array processor which allows parallel, conflict-free access to various slices of data (e.g., rows, columns, diagonals, etc.), and subsequent alignment of these data for processing. Memory access requirements for an array processor are discussed in general terms and a set of common requirements are defined. The ability to meet these requirements is shown to depend on the number of independent memory units and on the mapping of the data in these memories. Next, the need to align these data for processing is demonstrated and various alignment requirements are defined. Hardware which can perform this alignment function is discussed, e.g., permutation, indexing, switching or sorting networks, and a network (the omega network) based on Stone's shuffle-exchange operation [1] is presented. Construction of this network is described and many of its useful properties are proven. Finally, as an example of these ideas, an array processor is shown which allows conflict-free access and alignment of rows, columns, diagonals, backward diagonals, and square blocks in row or column major order, as well as certain other special operations.
Duncan H. Lawrie
IEEE Trans. Computers1