Michael E. Wolf

dblp:02/1821 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
0since 2021 · last 1996
0000-0001-6385-9128ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
4 papers
Compilers and program optimization · 86% Program analysis · 8% Programming languages and type systems · 4%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 56% High-performance computing · 38% Parallel and multicore computing · 6%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
loop transformation
0.031996
Combining Loop Transformations Considering Caches and Scheduling · MICRO 1996
A Loop Transformation Theory and an Algorithm to Maximize Parallelism · IEEE Trans. Parallel Distributed Syst. 1991
A Data Locality Optimizing Algorithm · PLDI 1991
Compilers and program optimization
instruction scheduling
0.011996
Combining Loop Transformations Considering Caches and Scheduling · MICRO 1996
Compilers and program optimization › instruction scheduling
software pipelining
0.011996
Combining Loop Transformations Considering Caches and Scheduling · MICRO 1996
Program analysis
data dependence analysis
0.011991
A Loop Transformation Theory and an Algorithm to Maximize Parallelism · IEEE Trans. Parallel Distributed Syst. 1991
Compilers and program optimization › memory optimization
data locality optimization
0.011991
A Data Locality Optimizing Algorithm · PLDI 1991
Compilers and program optimization › parallelization › automatic parallelization
loop parallelization
0.011991
A Loop Transformation Theory and an Algorithm to Maximize Parallelism · IEEE Trans. Parallel Distributed Syst. 1991
Compilers and program optimization › loop transformation
unimodular transformation
0.011991
A Data Locality Optimizing Algorithm · PLDI 1991
High-performance computing › numerical linear algebra
blocked algorithms
0.011991
The Cache Performance and Optimizations of Blocked Algorithms · ASPLOS 1991
Memory systems › memory access optimization
cache blocking
0.011991
The Cache Performance and Optimizations of Blocked Algorithms · ASPLOS 1991
Memory systems › cache
cache performance
0.011991
The Cache Performance and Optimizations of Blocked Algorithms · ASPLOS 1991
High-performance computing
performance optimization
0.011991
The Cache Performance and Optimizations of Blocked Algorithms · ASPLOS 1991
Memory systems
cache
0.021996
Combining Loop Transformations Considering Caches and Scheduling · MICRO 1996
A Data Locality Optimizing Algorithm · PLDI 1991
Programming languages and type systems
language design
0.011987
Extensions for Multi-Module Records in Conventional Programming Languages · POPL 1987
Parallel and multicore computing
parallel programming models
0.011991
A Loop Transformation Theory and an Algorithm to Maximize Parallelism · IEEE Trans. Parallel Distributed Syst. 1991
Requirements engineering and software design
modularity
0.011987
Extensions for Multi-Module Records in Conventional Programming Languages · POPL 1987

Methods — techniques the papers use, named apart from their topics

search-based transformation selection · 0.0machine cycle time model · 0.0wavefronting · 0.0tiling · 0.0skewing · 0.0reversal · 0.0loop interchange · 0.0dependence analysis · 0.0cache simulation · 0.0
YearPublicationVenuePosition
1996 Combining Loop Transformations Considering Caches and Scheduling
abstract
The performance of modern microprocessors is greatly affected by cache behavior, instruction scheduling, register allocation and loop overhead. High level loop transformations such as fission, fusion, tiling, interchanging and outer loop unrolling (e.g., unroll and jam) are well known to be capable of improving all these aspects of performance. Difficulties arise because these machine characteristics and these optimizations are highly interdependent. Interchanging two loops might, for example, improve cache behavior but make it impossible to allocate registers in the inner loop. Similarly, unrolling or interchanging a loop might individually hurt performance but doing both simultaneously might help performance. Little work has been published on how to combine these transformations into an efficient and effective compiler algorithm. In this paper we present a model that estimates total machine cycle time taking into account cache misses, software pipelining, register pressure and loop overhead. We then develop an algorithm to intelligently search through the various possible transformations, using our machine model to select the set of transformations leading to the best overall performance. We have implemented this algorithm as part of the MIPSPro commercial compiler system. We give experimental results showing that our approach is both effective and efficient in optimizing numerical programs.
Michael E. Wolf, Dror E. Maydan, Ding-Kai Chen
MICRO1
1991 The Cache Performance and Optimizations of Blocked Algorithms
abstract
article Free Access Share on The cache performance and optimizations of blocked algorithms Authors: Monica D. Lam View Profile , Edward E. Rothberg View Profile , Michael E. Wolf View Profile Authors Info & Claims ACM SIGOPS Operating Systems ReviewVolume 25Issue Special IssueApr. 1991 pp 63–74https://doi.org/10.1145/106974.106981Published:01 April 1991Publication History 704citation3,371DownloadsMetricsTotal Citations704Total Downloads3,371Last 12 Months344Last 6 weeks57 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Monica S. Lam, Edward E. Rothberg, Michael E. Wolf
ASPLOS3
1991 A Data Locality Optimizing Algorithm
abstract
This paper proposes an algorithm that improves the locality of a loop nest by transforming the code via interchange, reversal, skewing and tiling.The loop transformation rrlgorithm is based on two concepts: a mathematical formulation of reuse and locality, and a loop transformation theory that unifies the various transforms as unimodular matrix tmnsfonnations.The algorithm haa been implemented in the SUIF (Stanford University Intermediate Format) compiler, and is successful in optimizing codes such as matrix multiplication, successive over-relaxation (SOR), LU decomposition without pivoting, and Givens QR factorization.Performance evaluation indicates that locatity optimization is especially crucial for scaling up the performance of parallel code.
Michael E. Wolf, Monica S. Lam
PLDI1
1991 A Loop Transformation Theory and an Algorithm to Maximize Parallelism
abstract
An approach to transformations for general loops in which dependence vectors represent precedence constraints on the iterations of a loop is presented. Therefore, dependences extracted from a loop nest must be lexicographically positive. This leads to a simple test for legality of compound transformations: any code transformation that leaves the dependences lexicographically positive is legal. The loop transformation theory is applied to the problem of maximizing the degree of coarse- or fine-grain parallelism in a loop nest. It is shown that the maximum degree of parallelism can be achieved by transforming the loops into a nest of coarsest fully permutable loop nests and wavefronting the fully permutable nests. The canonical form of coarsest fully permutable nests can be transformed mechanically to yield maximum degrees of coarse- and/or fine-grain parallelism. The efficient heuristics can find the maximum degrees of parallelism for loops whose nesting level is less than five.>
Michael E. Wolf, Monica S. Lam
IEEE Trans. Parallel Distributed Syst.1
1987 Extensions for Multi-Module Records in Conventional Programming Languages
abstract
An extended record facility is described that supports multi-module records by providing:
David R. Cheriton, Michael E. Wolf
POPL2