Elana D. Granston

dblp:28/4888 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
0since 2021 · last 2002
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 81% Program analysis · 19%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › memory optimization
data layout optimization
0.011993
To copy or not to copy: a compile-time technique for assessing when data copying should be used to eliminate cache conflicts · SC 1993
Memory systems
cache
0.011993
To copy or not to copy: a compile-time technique for assessing when data copying should be used to eliminate cache conflicts · SC 1993
Memory systems › cache › cache miss
cache conflict misses
0.011993
To copy or not to copy: a compile-time technique for assessing when data copying should be used to eliminate cache conflicts · SC 1993
Compilers and program optimization › dependence analysis
array access analysis
0.011991
Detecting redundant accesses to array data · SC 1991
Compilers and program optimization
compiler analysis
0.011991
Detecting redundant accesses to array data · SC 1991
Compilers and program optimization › memory optimization
memory access optimization
0.011991
Detecting redundant accesses to array data · SC 1991
Program analysis › code quality analysis
redundancy detection
0.011991
Detecting redundant accesses to array data · SC 1991

Methods — techniques the papers use, named apart from their topics

simulation · 0.0compile-time analysis · 0.0flow analysis · 0.0dependence analysis · 0.0
YearPublicationVenuePosition
2002 Affinity-based cluster assignment for unrolled loops
abstract
To compete performance-wise, modern VLIW processors must have fast clock rates and high instruction-level parallelism (ILP). Partitioning resources (functional units and registers) into clusters allows the processor to be clocked faster, but operand transfers across clusters can easily become a bottleneck. Increasing the number of functional units increases the potential ILP, but only helps if the functional units can be kept busy.To support these features, optimizations such as loop unrolling must be applied to expose ILP, and instructions must be explicitly assigned to clusters to minimize cross-cluster transfers. In an architecture with homogeneous clusters, the number of functional units of a given type is typically a multiple of the number of clusters. Thus, it is common to unroll a loop so that the number of copies of the loop body is a multiple of the number of clusters. The result is that there is a natural mapping of instructions to clusters, which is often the best mapping. While this mapping can be obvious by inspection, we have found that existing cluster assignment algorithms often miss this natural split. The consequence is an excessive number of inter-cluster transfers, which slows down the loop.Because we were unable to find an existing cluster-assignment algorithm that performed well for unrolled loops, we developed our own. Our Affinity-Based Clustering (ABC) algorithm has been implemented in a production compiler for the Texas Instruments TMS320C6000, a two-cluster VLIW architecture. It is tailored for exploiting the patterns that result from either manual or compiler-based unrolling. As demonstrated experimentally, it performs well, even when post-unrolling optimizations partially obscure the natural split.
Gayathri Krishnamurthy, Elana D. Granston, Eric Stotzer
ICS2
1993 Managing Pages in Shared Virtual Memory Systems: Getting the Compiler into the Game
abstract
In large-scale multiprocessors, whether loosely or tightly coupled, some memory is cheaper to access than other memory. Because direct management of memory on these machines is quite burdensome to the programmer, much research effort has been directed toward providing a shared virtual memory (SVM) interface. Clearly, the success of this endeavor depends heavily on the efficiency of page management strategies. To date, this has been primarily the responsibility of the operating system, and secondarily that of the hardware. Unfortunately, delaying page management decisions entirely until run time can lead to an unacceptable loss of efficiency, due to poor data layout and memory reference patterns that are fixed by the end of compile time. For this reason, programmer assistance has been occasionally solicited. However, this disrupts the SVM abstraction. Moreover, many of these problems may be addressable at the compiler level instead. This is especially promising for array-based languages where compiler-based, analytical technology is most mature. Surprisingly, this possibility is largely unexplored. In this paper, we discuss the issue of compiler involvement in areas ranging from loop transformations and scheduling issues, to data layout strategies, page placement decisions, access pattern analysis, and use of run time system directives.
Elana D. Granston, Harry A. G. Wijshoff
International Conference on Supercomputing1
1993 To copy or not to copy: a compile-time technique for assessing when data copying should be used to eliminate cache conflicts
abstract
To reduce conflict misses, this technique, the data layout in a cache is adjusted by copying array files into temporary arrays that exhibit better cache behavior. This approach incurs a cost proportional to the amount of data being copied. To date, there has been no discussion regarding either this tradeoff or the problem of determining what and when to copy. The authors present a compile-time technique for making this determination and present a selective copying strategy based on this methodology. Preliminary experimental results demonstrate that, because of the sensitivity of cache conflicts to small changes in problem size and base addresses, selective copying can lead to better overall performance than either no copying, complete copying, or copying based on manually applied heuristics.
Olivier Temam, Elana D. Granston, William Jalby
SC2
1991 An Integrated Hardware/Software Solution for Effective Management of Local Storage in High-Performance Systems
Elana D. Granston, Alexander V. Veidenbaum
ICPP (2)1
1991 Detecting redundant accesses to array data
abstract
Alleviating memory access delays is crucial to harnessing the potential of hierarchical-memory, highperformance systems, especially vector and paral!el systems.In typica!numerics!applications, a significant portion of global memory data -lrafic arises from accesses to blocks of array elements or regions.Memory access delays due to such trafic can be reduced by using compile-time information to detect when iocal data can be reused, thereby eliminating redundant global memory accesses.In this paper, we present a compile-time algorithm that applies combined flow and dependence analysis to programs with vector and paraL lel constructs to detect such redundancies across loops nests, and in the presence of conditionals.We also show how this information can be used to eliminate redundancies.
Elana D. Granston, Alexander V. Veidenbaum
SC1
1990 Compiler-directed data prefetching in multiprocessors with memory hierarchies
abstract
Memory hierarchies are used by multiprocessor systems to reduce large memory access times. It is necessary to automatically manage such a hierarchy, to obtain effective memory utilization. In this paper, we discuss the various issues involved in obtaining an optimal memory management strategy for a memory hierarchy. We present an algorithm for finding the earliest point in a program that a block of data can be prefetched. This determination is based on the control and data dependencies in the program. Such a method is an integral part of more general memory management algorithms. We demonstrate our method's potential by using static analysis to estimate the performance improvement afforded by our prefetching strategy and to analyze the reference patterns in a set of Fortran benchmarks. We also study the effectiveness of prefetching in a realistic shared-memory system using an RTL-level simulator and real codes. This differs from previous studies by considering prefetching benefits in the presence of network contention.
Edward H. Gornish, Elana D. Granston, Alexander V. Veidenbaum
ICS2