Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jennifer-Ann M. Anderson

dblp:a/JenniferAnnMAnderson · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
0since 2021 · last 2000
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 5 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Memory systems · 41% Performance modeling and evaluation · 36% Energy-efficient computing · 12%
Software engineering, system software, and programming languages
7 papers
Compilers and program optimization · 69% Operating systems · 23% Runtime systems and virtual machines · 8%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation › profiling
continuous profiling
0.021997
Continuous Profiling: Where Have All the Cycles Gone? · ACM Trans. Comput. Syst. 1997
Continuous Profiling: Where Have All the Cycles Gone? · SOSP 1997
Performance modeling and evaluation
profiling
0.021997
Continuous Profiling: Where Have All the Cycles Gone? · ACM Trans. Comput. Syst. 1997
Continuous Profiling: Where Have All the Cycles Gone? · SOSP 1997
Compilers and program optimization › memory optimization
data locality optimization
0.021997
Data Distribution Support on Distributed Shared Memory Multiprocessors · PLDI 1997
Global Optimizations for Parallelism and Locality on Scalable Parallel Machines · PLDI 1993
Energy-efficient computing
energy measurement
0.012000
Quantifying the energy consumption of a pocket computer and a Java virtual machine · SIGMETRICS 2000
Parallel and multicore computing
data distribution
0.011997
Data Distribution Support on Distributed Shared Memory Multiprocessors · PLDI 1997
Memory systems › shared memory
distributed shared memory
0.011997
Data Distribution Support on Distributed Shared Memory Multiprocessors · PLDI 1997
Memory systems
cache
0.011996
Compiler-Directed Page Coloring for Multiprocessors · ASPLOS 1996
Memory systems › cache › cache miss
cache conflict misses
0.011996
Compiler-Directed Page Coloring for Multiprocessors · ASPLOS 1996
Memory systems › virtual memory management
page coloring
0.011996
Compiler-Directed Page Coloring for Multiprocessors · ASPLOS 1996
Memory systems › memory management
virtual memory
0.011996
Compiler-Directed Page Coloring for Multiprocessors · ASPLOS 1996
Compilers and program optimization › parallelization
automatic parallelization
0.011995
Data and Computation Transformations for Multiprocessors · PPoPP 1995
Compilers and program optimization › memory optimization
data layout transformation
0.011995
Data and Computation Transformations for Multiprocessors · PPoPP 1995
Memory systems › cache
cache performance
0.011995
Data and Computation Transformations for Multiprocessors · PPoPP 1995
Compilers and program optimization
parallelizing compiler
0.011993
Global Optimizations for Parallelism and Locality on Scalable Parallel Machines · PLDI 1993
Runtime systems and virtual machines › virtual machine implementation
java virtual machine
0.012000
Quantifying the energy consumption of a pocket computer and a Java virtual machine · SIGMETRICS 2000
Performance modeling and evaluation
workload characterization
0.011997
Continuous Profiling: Where Have All the Cycles Gone? · SOSP 1997
Compilers and program optimization › compiler optimization
compiler-directed optimization
0.011996
Compiler-Directed Page Coloring for Multiprocessors · ASPLOS 1996
Performance modeling and evaluation › parallel system performance
multiprocessor performance
0.011996
Compiler-Directed Page Coloring for Multiprocessors · ASPLOS 1996
Parallel and multicore computing › multiprocessor system
scalable parallel computers
0.011993
Global Optimizations for Parallelism and Locality on Scalable Parallel Machines · PLDI 1993

Methods — techniques the papers use, named apart from their topics

energy measurement · 0.1data acquisition · 0.1sampling-based profiling · 0.0error detection · 0.0compiler directives · 0.0cache coherence · 0.0simulation · 0.0compiler-directed page coloring · 0.0data transformation framework · 0.0compiler-based parallelization · 0.0
YearPublicationVenuePosition
2000 Quantifying the energy consumption of a pocket computer and a Java virtual machine
abstract
In this paper, we examine the energy consumption of a state-of-the-art pocket computer. Using a data acquisition system, we measure the energy consumption of the Itsy Pocket Computer, developed by Compaq Computer Corporation's Palo Alto Research Labs. We begin by showing that the energy usage characteristics of the Itsy differ markedly from that of a notebook computer. Then, since we expect that flexible software environments will become increasingly prevalent on pocket computers, we consider applications running in a Java environment. In particular, we explain some of the Java design tradeoffs applicable to pocket computers, and quantify their energy costs. For the design options we considered and the three workloads we studied, we find a maximum change in energy use of 25%.
Keith I. Farkas, Jason Flinn, Godmar Back, Dirk Grunwald, Jennifer-Ann M. Anderson
SIGMETRICS5
1997 Data Distribution Support on Distributed Shared Memory Multiprocessors
abstract
Cache-coherent multiprocessors with distributed shared memory are becoming increasingly popular for parallel computing. However, obtaining high performance on these machines mquires that an application execute with good data locality. In addition to making efiective use of caches, it is often necessary to distribute data structures across the local memories of the processing nodes, thereby reducing the latency of cache misses.We have designed a set of abstractions for performing data distribution in the context of explicitly parallel programs and implemented them within the SGI MIPSpro compiler system. Our system incorporates many unique features to enhance both programmability and performance. We address the former by providing a very simple programmming model with extensive support for error detection. Regarding performance, we carefully design the user abstractions with the underlying compiler optimizations in mind, we incorporate several optimization techniques to generate efficient code for accessing distributed data, and we provide a tight integration of these techniques with other optimizations within the compiler Our initial experience suggests that the directives are easy to use and can yield substantial performance gains, in some cases by as much as a factor of 3 over the same codes without distribution.
Rohit Chandra, Ding-Kai Chen, Robert Cox, Dror E. Maydan, Nenad Nedeljkovic, Jennifer-Ann M. Anderson
PLDI6
1997 Continuous Profiling: Where Have All the Cycles Gone?
abstract
Article Continuous profiling: where have all the cycles gone? Share on Authors: Jennifer M. Anderson View Profile , Lance M. Berc View Profile , Jeffrey Dean View Profile , Sanjay Ghemawat View Profile , Monika R. Henzinger View Profile , Shun-Tak A. Leung View Profile , Richard L. Sites View Profile , Mark T. Vandevoorde View Profile , Carl A. Waldspurger View Profile , William E. Weihl View Profile Authors Info & Claims SOSP '97: Proceedings of the sixteenth ACM symposium on Operating systems principlesOctober 1997 Pages 1–14https://doi.org/10.1145/268998.266637Published:01 October 1997 175citation1,209DownloadsMetricsTotal Citations175Total Downloads1,209Last 12 Months17Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jennifer-Ann M. Anderson, Lance M. Berc, Jeffrey Dean, Sanjay Ghemawat, Monika Henzinger, Shun-Tak Leung, Richard L. Sites, Mark T. Vandevoorde, Carl A. Waldspurger, William E. Weihl
SOSP1
1997 Continuous Profiling: Where Have All the Cycles Gone?
abstract
This article describes the Digital Continuous Profiling Infrastructure, a sampling-based profiling system designed to run continuously on production systems. The system supports multiprocessors, works on unmodified executables, and collects profiles for entire systems, including user programs, shared libraries, and the operating system kernel. Samples are collected at a high rate (over 5200 samples/sec. per 333MHz processor), yet with low overhead (1–3% slowdown for most workloads). Analysis tools supplied with the profiling system use the sample data to produce a precise and accurate accounting, down to the level of pipeline stalls incurred by individual instructions, of where time is bring spent. When instructions incur stalls, the tools identify possible reasons, such as cache misses, branch mispredictions, and functional unit contention. The fine-grained instruction-level analysis guides users and automated optimizers to the causes of performance problems and provides important insights for fixing them.
Jennifer-Ann M. Anderson, Lance M. Berc, Jeffrey Dean, Sanjay Ghemawat, Monika Henzinger, Shun-Tak Leung, Richard L. Sites, Mark T. Vandevoorde, Carl A. Waldspurger, William E. Weihl
ACM Trans. Comput. Syst.1
1996 Compiler-Directed Page Coloring for Multiprocessors
abstract
This paper presents a new technique, compiler-directed page coloring, that eliminates conflict misses in multiprocessor applications. It enables applications to make better use of the increased aggregate cache size available in a multiprocessor. This technique uses the compiler's knowledge of the access patterns of the parallelized applications to direct the operating system's virtual memory page mapping strategy. We demonstrate that this technique can lead to significant performance improvements over two commonly used page mapping strategies for machines with either direct-mapped or two-way set-associative caches. We also show that it is complementary to latency-hiding techniques such as prefetching.We implemented compiler-directed page coloring in the SUIF parallelizing compiler and on two commercial operating systems. We applied the technique to the SPEC95fp benchmark suite, a representative set of numeric programs. We used the SimOS machine simulator to analyze the applications and isolate their performance bottlenecks. We also validated these results on a real machine, an eight-processor 350MHz Digital AlphaServer. Compiler-directed page coloring leads to significant performance improvements for several applications. Overall, our technique improves the SPEC95fp rating for eight processors by 8% over Digital UNIX's page mapping policy and by 20% over a page coloring, a standard page mapping policy. The SUIF compiler achieves a SPEC95fp ratio of 57.4, the highest ratio to date.
Edouard Bugnion, Jennifer-Ann M. Anderson, Todd C. Mowry, Mendel Rosenblum, Monica S. Lam
ASPLOS2
1995 Unified Compilation Techniques for Shared and Distributed Address Space Machines
abstract
Parallel machines with shared address spaces are easy to program because they provide hardware support that allows each processor to transparently access non-local data. However, obtaining scalable performance can be difficult due to memory access and synchronization overhead. In this paper, we use profiling and simulation studies to identify the sources of parallel overhead. We demonstrate that compilation techniques for distributed address space machines can be very effective when used in compilers for shared address space machines. Automatic data decomposition can co-locate data and computation to improve locality. Data reorganization transformations can reduce harmful cache effects. Communication analysis can eliminate barrier synchronization. We present a set of unified compilation techniques that exemplify this convergence in compilers for shared and distributed address space machines, and illustrate their effectiveness using two example applications. 1 Introduction Until recent...
Chau-Wen Tseng, Jennifer-Ann M. Anderson, Saman P. Amarasinghe, Monica S. Lam
International Conference on Supercomputing2
1995 Data and Computation Transformations for Multiprocessors
abstract
Effective memory hierarchy utilization is critical to the performance of modern multiprocessor architectures. We havedeveloped the first compiler system that fully automatically parallelizes sequential programs and changes the original array layouts to improve memory system performance. Our optimization algorithm consists of two steps. The first step chooses the parallelization and computation assignment such that synchronization and data sharing are minimized. The second step then restructures the layout of the data in the shared address space with an algorithm that is based on a new data transformation framework. We ran our compiler on a set of application programs and measured their performance on the Stanford DASH multiprocessor. Our results show that the compiler can effectively optimize parallelism in conjunction with memory subsystem performance. 1 Introduction In the last decade, microprocessor speeds have been steadily improving at a rate of 50% to 100% every year[16]. Meanwh...
Jennifer-Ann M. Anderson, Saman P. Amarasinghe, Monica S. Lam
PPoPP1
1993 Global Optimizations for Parallelism and Locality on Scalable Parallel Machines
abstract
Data locality is critical to achieving high performance on large-scale parallel machines. Non-local data accesses result in communication that can greatly impact performance. Thus the mapping, or decomposition, of the computation and data onto the processors of a scalable parallel machine is a key issue in compiling programs for these architectures. This paper describes a compiler algorithm that automatically finds computation and data decompositions that optimize both parallelism and locality. This algorithm is designed for use with both distributed and shared address space machines. The scope of our algorithm is dense matrix computations where the array accesses are affine functions of the loop indices. Our algorithm can handle programs with general nestings of parallel and sequential loops. We present a mathematical framework that enables us to systematically derive the decompositions. Our algorithm can exploit parallelism in both fully parallelizable loops as well as loops that require explicit synchronization. The algorithm will trade off extra degrees of parallelism to eliminate communication. If communication is needed, the algorithm will try to introduce the least expensive forms of communication into those parts of the program that are least frequently executed. 1
Jennifer-Ann M. Anderson, Monica S. Lam
PLDI1