Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sally A. McKee

dblp:m/SallyAMcKee · DBLP profile ↗
← Back
73ranked-venue papers
6as first author
1since 2021 · last 2022
0000-0003-0514-3767ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 60 · 5 first-authorSoftware engineering, systems software and programming languages · 11 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 2Computer networks · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
17 papers
Memory systems · 20% Distributed systems · 17% Performance modeling and evaluation · 16%
Network and information security
1 paper
Systems and software security · 61% Hardware security and side channels · 30% Cryptographic primitives and cryptanalysis · 9%
Software engineering, system software, and programming languages
3 papers
Program analysis · 88% Compilers and program optimization · 11% Operating systems · 1%

Topics — the 30 heaviest of 47, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › datacenter architecture
datacenter server architecture
0.622018
Venice: An Effective Resource Sharing Architecture for Data Center Servers · ACM Trans. Comput. Syst. 2018
Venice: Exploring server architectures for effective resource sharing · HPCA 2016
Distributed systems
resource sharing
0.622018
Venice: An Effective Resource Sharing Architecture for Data Center Servers · ACM Trans. Comput. Syst. 2018
Venice: Exploring server architectures for effective resource sharing · HPCA 2016
Systems and software security › memory safety
control-flow integrity
0.412019
RAGuard: An Efficient and User-Transparent Hardware Mechanism against ROP Attacks · ACM Trans. Archit. Code Optim. 2019
Systems and software security
return-oriented programming defense
0.412019
RAGuard: An Efficient and User-Transparent Hardware Mechanism against ROP Attacks · ACM Trans. Archit. Code Optim. 2019
Program analysis › static analysis › abstract interpretation
abstract domain design
0.312018
Verifying Reliability Properties Using the Hyperball Abstract Domain · ACM Trans. Program. Lang. Syst. 2018
Program analysis › static analysis
abstract interpretation
0.312018
Verifying Reliability Properties Using the Hyperball Abstract Domain · ACM Trans. Program. Lang. Syst. 2018
Interconnection networks and networks-on-chip › interconnect architecture
communication fabric
0.312018
Venice: An Effective Resource Sharing Architecture for Data Center Servers · ACM Trans. Comput. Syst. 2018
Hardware reliability and fault tolerance
soft errors
0.312018
Verifying Reliability Properties Using the Hyperball Abstract Domain · ACM Trans. Program. Lang. Syst. 2018
Electronic design automation
design space exploration
0.122008
Efficient architectural design space exploration via predictive modeling · ACM Trans. Archit. Code Optim. 2008
Efficiently exploring architectural design spaces via predictive modeling · ASPLOS 2006
Performance modeling and evaluation
workload characterization
0.122017
Main Memory in HPC: Do We Need More or Could We Live with Less? · ACM Trans. Archit. Code Optim. 2017
Identifying and Exploiting Spatial Regularity in Data Memory References · SC 2003
Memory systems
DRAM
0.142001
Dynamic Access Ordering for Streamed Computations · IEEE Trans. Computers 2000
Design of a Parallel Vector Access Unit for SDRAM Memory Systems · HPCA 2000
Access Order and Effective Bandwidth for Streams on a Direct Rambus Memory · HPCA 1999
Memory systems
3d-stacked memory
0.112017
Main Memory in HPC: Do We Need More or Could We Live with Less? · ACM Trans. Archit. Code Optim. 2017
Distributed systems
fault tolerance
0.112008
Compiler-enhanced incremental checkpointing for OpenMP applications · PPoPP 2008
Performance modeling and evaluation
performance prediction
0.112008
Efficient architectural design space exploration via predictive modeling · ACM Trans. Archit. Code Optim. 2008
Memory systems
cache
0.112007
METRIC: Memory tracing via dynamic binary rewriting to identify cache inefficiencies · ACM Trans. Program. Lang. Syst. 2007
Performance modeling and evaluation › tracing
memory reference tracing
0.112007
METRIC: Memory tracing via dynamic binary rewriting to identify cache inefficiencies · ACM Trans. Program. Lang. Syst. 2007
Performance modeling and evaluation › parallel system performance
parallel performance modeling
0.112007
Methods of inference and learning for performance modeling of parallel applications · PPoPP 2007
Performance modeling and evaluation › statistical analysis
statistical inference
0.112007
Methods of inference and learning for performance modeling of parallel applications · PPoPP 2007
Memory systems
memory controller
0.122001
The Impulse Memory Controller · IEEE Trans. Computers 2001
Design of a Parallel Vector Access Unit for SDRAM Memory Systems · HPCA 2000
Memory systems › memory management › virtual memory
superpage promotion
0.122001
Reevaluating Online Superpage Promotion with Hardware Support · HPCA 2001
Online superpage promotion revisited (poster) · SIGMETRICS 2000
Memory systems › memory management
virtual memory
0.122001
Reevaluating Online Superpage Promotion with Hardware Support · HPCA 2001
Online superpage promotion revisited (poster) · SIGMETRICS 2000
Memory systems
memory access optimization
0.012003
Identifying and Exploiting Spatial Regularity in Data Memory References · SC 2003
Performance modeling and evaluation › profiling
memory access profiling
0.012003
Identifying and Exploiting Spatial Regularity in Data Memory References · SC 2003
Memory systems
memory bandwidth
0.021999
Access Order and Effective Bandwidth for Streams on a Direct Rambus Memory · HPCA 1999
Access Ordering and Memory-Conscious Cache Utilization · HPCA 1995
Memory systems › memory access
indirect addressing
0.012001
The Impulse Memory Controller · IEEE Trans. Computers 2001
Memory systems › memory management › virtual memory › address translation
TLB
0.012001
Reevaluating Online Superpage Promotion with Hardware Support · HPCA 2001
Memory systems › memory management › virtual memory
TLB reach
0.012001
Reevaluating Online Superpage Promotion with Hardware Support · HPCA 2001
Processor architecture and microarchitecture › memory system microarchitecture
memory ordering
0.012000
Dynamic Access Ordering for Streamed Computations · IEEE Trans. Computers 2000
Memory systems › DRAM › DRAM architecture
SDRAM
0.012000
Design of a Parallel Vector Access Unit for SDRAM Memory Systems · HPCA 2000
Memory systems › memory access patterns
vector access
0.012000
Design of a Parallel Vector Access Unit for SDRAM Memory Systems · HPCA 2000

Methods — techniques the papers use, named apart from their topics

submodular optimization · 0.7true random number generator · 0.4physical unclonable function · 0.4MAC binding · 0.4resource-joining mechanisms · 0.3inter-channel collaboration · 0.3hardware prototype · 0.3memory footprint profiling · 0.3hardware prototyping · 0.2incremental checkpointing · 0.2compiler analysis · 0.2predictive modeling · 0.1simulation sampling · 0.1online algorithm · 0.0
YearPublicationVenuePosition
2022 Posits and the state of numerical representations in the age of exascale and edge computing
abstract
Abstract Growing constraints on memory utilization, power consumption, and I/O throughput have increasingly become limiting factors to the advancement of high performance computing (HPC) and edge computing applications. IEEE‐754 floating‐point types have been the de facto standard for floating‐point number systems for decades, but the drawbacks of this numerical representation leave much to be desired. Alternative representations are gaining traction, both in HPC and machine learning environments. Posits have recently been proposed as a drop‐in replacement for the IEEE‐754 floating‐point representation. We survey the state‐of‐the‐art and state‐of‐the‐practice in the development and use of posits in edge computing and HPC. The current literature supports posits as a promising alternative to traditional floating‐point systems, both as a stand‐alone replacement and in a mixed‐precision environment. Development and standardization of the posit type is ongoing, and much research remains to explore the application of posits in different domains, how to best implement them in hardware, and where they fit with other numerical representations.
Alexandra Poulos, Sally A. McKee, Jon Calhoun 0001
Softw. Pract. Exp.2
2019 RAGuard: An Efficient and User-Transparent Hardware Mechanism against ROP Attacks
abstract
Control-flow integrity (CFI) is a general method for preventing code-reuse attacks, which utilize benign code sequences to achieve arbitrary code execution. CFI ensures that the execution of a program follows the edges of its predefined static Control-Flow Graph: any deviation that constitutes a CFI violation terminates the application. Despite decades of research effort, there are still several implementation challenges in efficiently protecting the control flow of function returns (Return-Oriented Programming attacks). The set of valid return addresses of frequently called functions can be large and thus an attacker could bend the backward-edge CFI by modifying an indirect branch target to another within the valid return set. This article proposes RAGuard, an efficient and user-transparent hardware-based approach to prevent Return-Oreiented Programming attacks. RAGuard binds a message authentication code (MAC) to each return address to protect its integrity. To guarantee the security of the MAC and reduce runtime overhead: RAGuard (1) computes the MAC by encrypting the signature of a return address with AES-128, (2) develops a key management module based on a Physical Unclonable Function (PUF) and a True Random Number Generator (TRNG), and (3) uses a dedicated register to reduce MACs’ load and store operations of leaf functions. We have evaluated our mechanism based on the open-source LEON3 processor and the results show that RAGuard incurs acceptable performance overhead and occupies reasonable area.
Rui Hou 0001, Wei Song 0002, Sally A. McKee, Zhen Jia 0001, Chen Zheng 0001, Mingyu Chen 0001, Lixin Zhang 0002, Dan Meng 0002
ACM Trans. Archit. Code Optim.4
2019 Understanding Processors Design Decisions for Data Analytics in Homogeneous Data Centers
abstract
Our global economy increasingly depends on our ability to gather, analyze, link, and compare very large data sets. Keeping up with such big data poses challenges in terms of both computational performance and energy efficiency, and motivates different approaches to explore data center systems and architectures. To better understand the processor design decisions in context of data analytics in data centers, we conduct comprehensive evaluations using representative data analaytics workloads on representative conventional multi-core and many-core processors. After a comprehensive analysis of performance, power, energy efficiency and performance-cost efficiency, we have the following observations: contrasted with the conventional wisdom that uses wimpy many-core processors to improve energy-efficiency, the brawny multi-core processors with SMT (simultaneous multithreading) and dynamic overclocking technologies outperform the counterparts in terms of not only execution time, but also energy-efficiency for most of data analytics workloads in our experiments.
Zhen Jia 0001, Wanling Gao, Yingjie Shi, Sally A. McKee, Zhenyan Ji, Jianfeng Zhan, Lei Wang 0004, Lixin Zhang 0002
IEEE Trans. Big Data4
2018 XOS: An Application-Defined Operating System for Datacenter Computing
abstract
Rapid growth of datacenter (DC) scale, urgency of cost control, increasing workload diversity, and huge software investment protection place unprecedented demands on the operating system (OS) efficiency, scalability, performance isolation, and backward-compatibility. The traditional OSes are not built to work with deep-hierarchy software stacks, large numbers of cores, tail latency guarantee, and increasingly rich variety of applications seen in modern DCs, and thus they struggle to meet the demands of such workloads. This paper presents XOS, an application-defined OS for modern DC servers. Our design moves resource management out of the OS kernel, supports customizable kernel subsystems in user space, and enables elastic partitioning of hardware resources. Specifically, XOS leverages modern hardware support for virtualization to move resource management functionality out of the conventional kernel and into user space, which lets applications achieve near bare-metal performance. We implement XOS on top of Linux to provide backward compatibility. XOS speeds up a set of DC workloads by up to 1.6× over our baseline Linux on a 24-core server, and outperforms the state-of-the-art Dune by up to 3.3× in terms of virtual memory management. In addition, XOS demonstrates good scalability and strong performance isolation.
Chen Zheng 0001, Lei Wang 0004, Sally A. McKee, Lixin Zhang 0002, Hainan Ye, Jianfeng Zhan
IEEE BigData3
2018 Venice: An Effective Resource Sharing Architecture for Data Center Servers
abstract
Consolidated server racks are quickly becoming the standard infrastructure for engineering, business, medicine, and science. Such servers are still designed much in the way when they were organized as individual, distributed systems. Given that many fields rely on big-data analytics substantially, its cost-effectiveness and performance should be improved, which can be achieved by flexibly allowing resources to be shared across nodes. Here we describe Venice, a family of data-center server architectures that includes a strong communication substrate as a first-class resource. Venice supports a diverse set of resource-joining mechanisms that enables applications to leverage non-local resources efficiently. We have constructed a hardware prototype to better understand the implications of design decisions about system support for resource sharing. We use it to measure the performance of at-scale applications and to explore performance, power, and resource-sharing transparency tradeoffs (i.e., how many programming changes are needed). We analyze these tradeoffs for sharing memory, accelerators, and NICs. We find that reducing/hiding latency is particularly important, the chosen communication channels should match the sharing access patterns of the applications, and of which we can improve performance by exploiting inter-channel collaboration.
Boyan Zhao, Rui Hou 0001, Jianbo Dong, Michael C. Huang 0001, Sally A. McKee, Qianlong Zhang, Yueji Liu, Lixin Zhang 0002, Dan Meng 0002
ACM Trans. Comput. Syst.5
2018 Verifying Reliability Properties Using the Hyperball Abstract Domain
abstract
Modern systems are increasingly susceptible to soft errors that manifest themselves as bit flips and possibly alter the semantics of an application. We would like to measure the quality degradation on semantics due to such bit flips, and thus we introduce a Hyperball abstract domain that allows us to determine the worst-case distance between expected and actual results. Similar to intervals, hyperballs describe a connected and dense space. The semantics of low-level code in the presence of bit flips is hard to accurately describe in such a space. We therefore combine the Hyperball domain with an existing affine system abstract domain that we extend to handle bit flips, which are introduce as disjunctions. Bit-flips can reduce the precision of our analysis, and we therefor introduce the Scale domain as a disjunctive refinement to minimize precision loss. This domain bounds the number of disjunctive elements by quantifying the over-approximation of different partitions and uses submodular optimization to find a good partitioning (within a bound of optimal). We evaluate these domains to show benefits and potential problems. For the application we examine here, adding the Scale domain to the Hyperball abstraction improves accuracy by up to two orders of magnitude. Our initial results demonstrate the feasibility of this approach, although we would like to further improve execution efficiency.
Jacob Lidman, Sally A. McKee
ACM Trans. Program. Lang. Syst.2
2017 Main Memory in HPC: Do We Need More or Could We Live with Less?
abstract
An important aspect of High-Performance Computing (HPC) system design is the choice of main memory capacity. This choice becomes increasingly important now that 3D-stacked memories are entering the market. Compared with conventional Dual In-line Memory Modules (DIMMs), 3D memory chiplets provide better performance and energy efficiency but lower memory capacities. Therefore, the adoption of 3D-stacked memories in the HPC domain depends on whether we can find use cases that require much less memory than is available now. This study analyzes the memory capacity requirements of important HPC benchmarks and applications. We find that the High-Performance Conjugate Gradients (HPCG) benchmark could be an important success story for 3D-stacked memories in HPC, but High-Performance Linpack (HPL) is likely to be constrained by 3D memory capacity. The study also emphasizes that the analysis of memory footprints of production HPC applications is complex and that it requires an understanding of application scalability and target category, i.e., whether the users target capability or capacity computing. The results show that most of the HPC applications under study have per-core memory footprints in the range of hundreds of megabytes, but we also detect applications and use cases that require gigabytes per core. Overall, the study identifies the HPC applications and use cases with memory footprints that could be provided by 3D-stacked memory chiplets, making a first step toward adoption of this novel technology in the HPC domain.
Darko Zivanovic, Milan Pavlovic, Milan Radulovic, Hyunsung Shin, Jong-Pil Son, Sally A. McKee, Paul M. Carpenter, Petar Radojkovic, Eduard Ayguadé
ACM Trans. Archit. Code Optim.6
2016 Redesigning a tagless access buffer to require minimal ISA changes
abstract
Energy efficiency is a first-order design goal for nearly all classes of processors, but it is particularly important in mobile and embedded systems. Data caches in such systems account for a large portion of the processor's energy usage, and thus techniques to improve the energy efficiency of the cache hierarchy are likely to have high impact. Our prior work reduced data cache energy via a tagless access buffer (TAB) that sits at the top of the cache hierarchy. Strided memory references are redirected from the level-one data cache (L1D) to the smaller, more energy-efficient TAB. These references need not access the data translation lookaside buffer (DTLB), and they can avoid unnecessary transfers from lower levels of the memory hierarchy. The original TAB implementation requires changing the immediate field of load and store instructions, necessitating substantial ISA modifications. Here we present a new TAB design that requires minimal instruction set changes, gives software more explicit control over TAB resource management, and remains compatible with legacy (non-TAB) code. With a line size of 32 bytes, a four-line TAB can eliminate 31% of L1D accesses, on average. Together, the new TAB, L1D, and DTLB use 22% less energy than a TAB-less hierarchy, and the TAB system decreases execution time by 1.7%.
Carlos Sanchez, Peter Gavin, Daniel Moreau, Magnus Själander, David B. Whalley, Per Larsson-Edefors, Sally A. McKee
CASES7
2016 Venice: Exploring server architectures for effective resource sharing
abstract
Consolidated server racks are quickly becoming the backbone of IT infrastructure for science, engineering, and business, alike. These servers are still largely built and organized as when they were distributed, individual entities. Given that many fields increasingly rely on analytics of huge datasets, it makes sense to support flexible resource utilization across servers to improve cost-effectiveness and performance. We introduce Venice, a family of data-center server architectures that builds a strong communication substrate as a first-class resource for server chips. Venice provides a diverse set of resource-joining mechanisms that enables user programs to efficiently leverage non-local resources. To better understand the implications of design decisions about system support for resource sharing we have constructed a hardware prototype that allows us to more accurately measure end-to-end performance of at-scale applications and to explore tradeoffs among performance, power, and resource-sharing transparency. We present results from our initial studies analyzing these tradeoffs when sharing memory, accelerators, or NICs. We find that it is particularly important to reduce or hide latency, that data-sharing access patterns should match the features of the communication channels employed, and that inter-channel collaboration can be exploited for better performance.
Jianbo Dong, Rui Hou 0001, Michael C. Huang 0001, Tao Jiang 0010, Boyan Zhao, Sally A. McKee, Xiaosong Cui, Lixin Zhang 0002
HPCA6
2016 Extending On-chip Interconnects for rack-level remote resource access
abstract
The need to perform data analytics on exploding data volumes coupled with the rapidly changing workloads in cloud computing places great pressure on data-center servers. To improve hardware resource utilization across servers within a rack, we propose Direct Extension of On-chip Interconnects (DEOI), a high-performance and efficient architecture for remote resource access among server nodes. DEOI extends an SoC server node's on-chip interconnect to access resources in adjacent nodes with no protocol changes, allowing remote memory and network resources to be used as if they were local. Our results on a four-node FPGA prototype show that the latency of user-level, cross-node, random reads to DEOI-connected remote memory is as low as 1.16µs, which beats current commercial technologies. We exploit DEOI remote access to improve performance of the Redis in-memory key-value framework by 47%. When using DEOI to access remote network resources, we observe an 8.4% average performance degradation and only a 2.52µs ping-pong latency disparity compared to using local assets. These results suggest that DEOI can be a promising mechanism for increasing both performance and efficiency in next-generation data-center servers.
Yisong Chang, Ke Zhang 0017, Sally A. McKee, Lixin Zhang 0002, Mingyu Chen 0001, Liqiang Ren, Zhiwei Xu 0002
ICCD3
2016 A Methodology for Modeling Dynamic and Static Power Consumption for Multicore Processors
abstract
System designers and application programmersmust consider trade-offs between performance and energy. Making energy-aware decisions when designing an application or runtime system requires quantitative information about power consumed by different processor components. We present a methodology to model static and dynamic power consumption of individual cores and the uncore components, and we validate our power model for both sequential and parallel benchmarks at different voltage-frequency pairs on an Intel Haswell platform. Our power models yield the following insights about energy-efficient scaling. (1) We show that uncore energy accounts for up to 74% of total energy. In particular, uncore static energy can be as high as 61% of total energy, potentially making it a major source of energy inefficiency. (2) We find that the frequency at which an application expends the lowest energy depends on how memory-bound it is. (3) We demonstrate that even though using more cores may improve performance, the energy consumed by stalled cores during serial portions of theprogram can make using fewer cores more energy-efficient.
Bhavishya Goel, Sally A. McKee
IPDPS2
2016 Agave: A benchmark suite for exploring the complexities of the Android software stack
abstract
Traditional suites used for benchmarking high-performance computing platforms or for architectural design space exploration use much simpler virtual memory layouts and multitasking/ multithreading schemes, which means that they cannot be used to study the complex interactions among the layers of the Android software stack. To demonstrate this, we present memory reference and concurrency data showing how Android applications differ from traditional C benchmarks. We propose the Agave suite of open-source applications as the basis for a standard, multipurpose Android benchmark suite. We make all sources and tools available in hopes that the community will adopt and build on this initial version of Agave.
Martin K. Brown, Zachary Yannes, Michael Lustig, Mazdak Sanati, Sally A. McKee, Gary S. Tyson, Steven K. Reinhardt
ISPASS5
2015 Exploiting Program Semantics to Place Data in Hybrid Memory
abstract
Large-memory applications like data analytics and graph processing benefit from extended memory hierarchies, and hybrid DRAM/NVM (non-volatile memory) systems represent an attractive means by which to increase capacity at reasonable performance/energy tradeoffs. Compared to DRAM, NVMs generally have longer latencies and higher energies for writes, which makes careful data placement essential for efficient system operation. Data placement strategies that resort to monitoring all data accesses and migrating objects to dynamically adjust data locations incur high monitoring overhead and unnecessary memory copies due to mispredicted migrations. We find that program semantics (specifically, global access characteristics) can effectively guide initial data placement with respect to memory types, which, in turn, makes run-time migration more efficient. We study a combined offline/online placement scheme that uses access profiling information to place objects statically and then selectively monitors run-time behaviors to optimize placements dynamically. We present a software/hardware cooperative framework, 2PP, and evaluate it with respect to state-of-the-art migratory placement, finding that it improves performance by an average of 12.1%. Furthermore, 2PP improves energy efficiency by up to 51.8%, and by an average of 18.4%. It does so by reducing run-time monitoring and migration overheads.
Dejun Jiang 0001, Sally A. McKee, Jin Xiong, Mingyu Chen 0001
PACT3
2015 Adapting Memory Hierarchies for Emerging Datacenter Interconnects
Tao Jiang 0010, Rui Hou 0001, Jianbo Dong, Lin Chai, Sally A. McKee, Lixin Zhang 0002, Ninghui Sun
J. Comput. Sci. Technol.5
2014 Digging deeper into cluster system logs for failure prediction and root cause diagnosis
abstract
As the sizes of supercomputers and data centers grow towards exascale, failures become normal. System logs play a critical role in the increasingly complex tasks of automatic failure prediction and diagnosis. Many methods for failure prediction are based on analyzing event logs for large scale systems, but there is still neither a widely used one to predict failures based on both non-fatal and fatal events, nor a precise one that uses fine-grained information (such as failure type, node location, related application, and time of occurrence). A deeper and more precise log analysis technique is needed. We propose a three-step approach to draw out event dependencies and to identify failure-event generating processes. First, we cluster frequent event sequences into event groups based on common events. Then we infer causal dependencies between events in each event group. Finally, we extract failure rules based on the observation that events of the same event types, on the same nodes or from the same applications have similar operational behaviors. We use this rich information to improve failure prediction. Our approach semi-automates diagnosing the root causes of failure events, making it a valuable tool for system administrators.
Xiaoyu Fu, Sally A. McKee, Jianfeng Zhan, Ninghui Sun
CLUSTER3
2014 DTail: a flexible approach to DRAM refresh management
abstract
DRAM cells must be refreshed (or rewritten) periodically to maintain data integrity, and as DRAM density grows, so does the refresh time and energy. Not all data need to be refreshed with the same frequency, though, and thus some refresh operations can safely be delayed. Tracking such information allows the memory controller to reduce refresh costs by judiciously choosing when to refresh different rows
Zehan Cui, Sally A. McKee, Zhongbin Zha, Yungang Bao, Mingyu Chen 0001
ICS2
2014 Performance and Energy Analysis of the Restricted Transactional Memory Implementation on Haswell
abstract
Hardware transactional memory implementations are becoming increasingly available. For instance, the Intel Core i7 4770 implements Restricted Transactional Memory (RTM) support for Intel Transactional Synchronization Extensions (TSX). In this paper, we present a detailed evaluation of RTM performance and energy expenditure. We compare RTM behavior to that of the TinySTM software transactional memory system, first by running micro benchmarks, and then by running the STAMP benchmark suite. We find that which system performs better depends heavily on the workload characteristics. We then conduct a case study of two STAMP applications to assess the impact of programming style on RTM performance and to investigate what kinds of software optimizations can help overcome RTM's hardware limitations.
Bhavishya Goel, J. Rubén Titos Gil, Anurag Negi, Sally A. McKee, Per Stenström
IPDPS4
2014 QBLESS: A case for QoS-aware bufferless NoCs
abstract
Datacenters consolidate diverse applications to improve utilization. However when multiple applications are co-located on such platforms, contention for shared resources like Networks-on-Chip (NoCs) can degrade the performance of latency-critical online services (high-priority applications). Recently proposed bufferless NoCs have the advantages of requiring less area and power, but they pose challenges in quality-of-service (QoS) support, which usually relies on buffer-based virtual channels (VCs). We propose QBLESS, a QoS-aware bufferless NoC scheme for datacenters. QBLESS consists of two components: a routing mechanism (QBLESS-R) that can substantially reduce flit deflection for high-priority applications, and a congestion-control mechanism (QBLESS-CC) that guarantees performance for high-priority applications and improves overall system throughput. We use trace-driven simulation to model a 64-core system, finding that when compared to BLESS, a previous state-of-the-art bufferless NoC design, QBLESS improves performance of high-priority applications by an average of 33.2%.
Zhicheng Yao, Xiufeng Sui, Tianni Xu, Jiuyue Ma, Sally A. McKee, Binzhang Fu, Yungang Bao
IWQoS6
2013 Improving data access efficiency by using a tagless access buffer (TAB)
abstract
The need for energy efficiency continues to grow for many classes of processors, including those for which performance remains vital. Data cache is crucial for good performance, but it also represents a significant portion of the processor's energy expenditure. We describe the implementation and use of a tagless access buffer (TAB) that greatly improves data access energy efficiency while slightly improving performance. The compiler recognizes memory reference patterns within loops and allocates these references to a TAB. This combined hardware/software approach reduces energy usage by (1) replacing many level-one data cache (L1D) accesses with accesses to the smaller, more power-efficient TAB; (2) removing the need to perform tag checks or data translation lookaside buffer (DTLB) lookups for TAB accesses; and (3) reducing DTLB lookups when transferring data between the L1D and the TAB. Accesses to the TAB occur earlier in the pipeline, and data lines are prefetched from lower memory levels, which result in a small performance improvement. In addition, we can avoid many unnecessary block transfers between other memory hierarchy levels by characterizing how data in the TAB are used. With a combined size equal to that of a conventional 32-entry register file, a four-entry TAB eliminates 40% of L1D accesses and 42% of DTLB accesses, on average. This configuration reduces data-access related energy by 35% while simultane-ously decreasing execution time by 3%.
Alen Bardizbanyan, Peter Gavin, David B. Whalley, Magnus Själander, Per Larsson-Edefors, Sally A. McKee, Per Stenström
CGO6
2012 Topic 2: Performance Prediction and Evaluation
Allen D. Malony, Helen D. Karatza, William J. Knottenbelt, Sally A. McKee
Euro-Par4
2012 Parallelizing more Loops with Compiler Guided Refactoring
abstract
The performance of many parallel applications relies not on instruction-level parallelism but on loop-level parallelism. Unfortunately, automatic parallelization of loops is a fragile process, many different obstacles affect or prevent it in practice. To address this predicament we developed an interactive compilation feedback system that guides programmers in iteratively modifying their application source code. This helps leverage the compiler's ability to generate loop-parallel code. We employ our system to modify two sequential benchmarks dealing with image processing and edge detection, resulting in scalable parallelized code that runs up to 8.3 times faster on an eight-core Intel Xeon 5570 system and up to 12.5 times faster on a quad-core IBM POWER6 system. Benchmark performance varies significantly between the systems. This suggests that semi-automatic parallelization should be combined with target-specific optimizations. Furthermore, comparing the first benchmark to manually-parallelized, hand-optimized pthreads and OpenMP versions, we find that code generated using our approach typically outperforms the pthreads code (within 93-339%). It also performs competitively against the OpenMP code (within 75-111%). The second benchmark outperforms manually-parallelized and optimized OpenMP code (within 109-242%).
Per Larsen, Razya Ladelsky, Jacob Lidman, Sally A. McKee, Sven Karlsson, Ayal Zaks
ICPP4
2012 An LTE Uplink Receiver PHY benchmark and subframe-based power management
abstract
With the proliferation of mobile phones and other mobile internet appliances, the application area of baseband processing continues to grow in importance. Much academic research addresses the underlying mathematics, but little has been published on the design of systems to execute baseband workloads. Most systems research is conducted within companies who go to great lengths to protect their intellectual property. We present an open-source LTE Uplink Receiver PHY benchmark with a realistic representation of the baseband processing of an LTE base station, and we demonstrate its usefulness in investigating resource management strategies to conserve power on a TILEPro64. By estimating the workload of each subframe and using these estimates to control power-gating, we reduce power consumption by more than 24% (11% on average) compared to executing the benchmark with no estimation-guided resource management. By making available a benchmark containing no proprietary algorithms, we enable a broader community to conduct research both in baseband processing and on the systems that are used to execute such workloads.
Magnus Själander, Sally A. McKee, Peter Brauer, David Engdal, András Vajda
ISPASS2
2012 Active memory controller
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Sally A. McKee, Ali Ibrahim, Michael A. Parker, Xiaowei Jiang
J. Supercomput.4
2011 SoftBeam: Precise tracking of transient faults and vulnerability analysis at processor design time
abstract
To study system reliability of a next-generation system, we undertake a soft error vulnerability study for a next-generation microprocessor design. Starting from design data for the entire processor, we extend the microprocessor verification methodology to study soft error propagation through microprocessor logic into the architected processor state. We use soft error injection into randomly selected latch bits to (1) identify areas for improvement, (2) derate technology susceptibility by architectural, microarchitectural, and logic masking resulting in increased soft error resilience; and (3) identify areas where microarchitectural data corruption can be tolerated as performance degradation without impact on correctness, yielding even greater soft error resilience. Based on these results, we reduce design vulnerability to soft errors by factors ranging from 2 for an execution unit to more than 32 for a memory management unit.
Michael Gschwind, Valentina Salapura, Catherine Trammell, Sally A. McKee
ICCD4
2011 Power-Aware Resource Scheduling in Base Stations
abstract
Base band stations for Long Term Evolution (LTE) communication processing tend to rely on over-provisioned resources to ensure that peak demands can be met. These systems must meet user Quality of Service expectations, but during non-peak workloads, for instance, many of the cores could be placed in low-power modes. One key property of such application-specific systems is that they execute frequent, short-lived tasks. Sophisticated resource management and task scheduling approaches suffer intolerable overhead costs in terms of time and expense, and thus lighter-weight and more efficient strategies are essential to both saving power and meeting performance expectations. To this end, we develop a flexible, non-propietary LTE workload model to drive our resource management studies. Here we describe our experimental infrastructure and present early results that underscore the promise of our approach along with its implications on future hardware/software codesign.
Magnus Själander, Sally A. McKee, Bhavishya Goel, Peter Brauer, David Engdal, András Vajda
MASCOTS2
2010 Designing OS for HPC Applications: Scheduling
abstract
Operating systems have historically been implemented as independent layers between hardware and applications. User programs communicate with the OS through a set of well defined system calls, and do not have direct access to the hardware. The OS, in turn, communicates with the underlying architecture via control registers. Except for these interfaces, the three layers are practically oblivious to each other. While this structure improves portability and transparency, it may not deliver optimal performance. This is especially true for High Performance Computing (HPC) systems, where modern parallel applications and multi-core architectures pose new challenges in terms of performance, power consumption, and system utilization. The hardware, the OS, and the applications can no longer remain isolated, and instead should cooperate to deliver high performance with minimal power consumption. In this paper we present our experience with the design and implementation of High Performance Linux (HPL), an operating system designed to optimize the performance of HPC applications running on a state-of-the-art compute cluster. We show how characterizing parallel applications through hardware and software performance counters drives the design of the OS and how including knowledge about the architecture improves performance and efficiency. We perform experiments on a dual-socket IBM POWER6 machine, showing performance improvements and stability (performance variation of 2.11% on average) for NAS, a widely used parallel benchmark suite.
Roberto Gioiosa, Sally A. McKee, Mateo Valero
CLUSTER2
2010 Comparing Scalability Prediction Strategies on an SMP of CMPs
Matthew Curtis-Maury, Sally A. McKee, Filip Blagojevic, Dimitrios S. Nikolopoulos, Bronis R. de Supinski, Martin Schulz 0001
Euro-Par (1)3
2010 An approach to resource-aware co-scheduling for CMPs
abstract
We develop real-time scheduling techniques for improving performance and energy for multiprogrammed workloads that scale non-uniformly with increasing thread counts. Multithreaded programs generally deliver higher throughput than single-threaded programs on chip multiprocessors, but performance gains from increasing threads decrease when there is contention for shared resources. We use analytic metrics to derive local search heuristics for creating efficient multiprogrammed, multithreaded workload schedules. Programs are allocated fewer cores than requested, and scheduled to space-share the CMP to improve global throughput. Our holistic approach attempts to co-schedule programs that complement each other with respect to shared resource consumption. We find application co-scheduling for performance and energy in a resource-aware manner achieves better results than solely targeting total throughput or concurrently co-scheduling all programs. Our schedulers improve overall energy delay (E*D) by a factor of 1.5 over time-multiplexed gang scheduling.
Major Bhadauria, Sally A. McKee
ICS2
2009 Accomodating Diversity in CMPs with Heterogeneous Frequencies
Major Bhadauria, Vincent M. Weaver, Sally A. McKee
HiPEAC3
2009 Revisiting Cache Block Superloading
Matthew A. Watkins, Sally A. McKee, Lambert Schaelicke
HiPEAC2
2009 Code density concerns for new architectures
abstract
Reducing a program's instruction count can improve cache behavior and bandwidth utilization, lower power consumption, and increase overall performance. Nonetheless, code density is an often overlooked feature in studying processor architectures. We hand-optimize an assembly language embedded benchmark for size on 21 different instruction set architectures, finding up to a factor of three difference in code sizes from ISA alone. We find that the architectural features that contribute most heavily to code density are instruction length, number of registers, availability of a zero register, bit-width, hardware divide units, number of instruction operands, and the availability of unaligned loads and stores. We extend our results to investigate operating system, compiler, and system library effects on code density. We find that the executable starting address, executable format, and system call interface all affect program size. While ISA effects are important, the efficiency of the entire system stack must be taken into account when developing a new dense instruction set architecture.
Vincent M. Weaver, Sally A. McKee
ICCD2
2009 PARSEC: hardware profiling of emerging workloads for CMP design
abstract
No abstract available.
Major Bhadauria, Vincent M. Weaver, Sally A. McKee
ICS3
2009 Cancellation of loads that return zero using zero-value caches
abstract
The speed gap between processor and memory continues to limit performance. To address this problem, we explore the potential of eliminating Zero Loads -- loads accessing memory locations that contain the value "zero" -- to improve performance and energy dissipation. Our study shows that such loads comprise as many as 18% of the total number of dynamic loads. We show that a significant fraction of zero loads ends up on the critical memory-access path in out-of-order cores. We propose a non-speculative microarchitectural technique -- Zero-Value Cache (ZVC) -- to capitalize on zero loads and explore critical design options of such caches. We show that with modest investment (typically a 512-byte structure), we can obtain speedups up to 32%. Most importantly, zero-value caches never cause performance loss.
Md. Mafijul Islam, Sally A. McKee, Per Stenström
ICS2
2009 Prediction-based power estimation and scheduling for CMPs
abstract
No abstract available.
Major Bhadauria, Sally A. McKee
ICS3
2009 Compiler-enhanced incremental checkpointing for OpenMP applications
abstract
As modern supercomputing systems reach the peta-flop performance range, they grow in both size and complexity. This makes them increasingly vulnerable to failures from a variety of causes. Checkpointing is a popular technique for tolerating such failures, enabling applications to periodically save their state and restart computation after a failure. Although a many automated system-level checkpointing solutions are currently available to HPC users, manual application-level checkpointing remains more popular due to its superior performance. This paper improves performance of automated checkpointing via a compiler analysis for incremental checkpointing. This analysis, which works with both sequential and OpenMP applications, reduces checkpoint sizes by as much as 80% and enables asynchronous checkpointing.
Greg Bronevetsky, Daniel Marques, Keshav Pingali, Sally A. McKee, Radu Rugina
IPDPS4
2009 Machine learning based online performance prediction for runtime parallelization and task scheduling
abstract
With the emerging many-core paradigm, parallel programming must extend beyond its traditional realm of scientific applications. Converting existing sequential applications as well as developing next-generation software requires assistance from hardware, compilers and runtime systems to exploit parallelism transparently within applications. These systems must decompose applications into tasks that can be executed in parallel and then schedule those tasks to minimize load imbalance. However, many systems lack a priori knowledge about the execution time of all tasks to perform effective load balancing with low scheduling overhead. In this paper, we approach this fundamental problem using machine learning techniques first to generate performance models for all tasks and then applying those models to perform automatic performance prediction across program executions. We also extend an existing scheduling algorithm to use generated task cost estimates for online task partitioning and scheduling. We implement the above techniques in the pR framework, which transparently parallelizes scripts in the popular R language, and evaluate their performance and overhead with both a real-world application and a large number of synthetic representative test scripts. Our experimental results show that our proposed approach significantly improves task partitioning and scheduling, with maximum improvements of 21.8%, 40.3% and 22.1% and average improvements of 15.9%, 16.9% and 4.2% for LMM (a real R application) and synthetic test cases with independent and dependent tasks, respectively.
Jiangtian Li, Xiaosong Ma, Martin Schulz 0001, Bronis R. de Supinski, Sally A. McKee
ISPASS6
2008 Evolutionary system for prediction and optimization of hardware architecture performance
abstract
The design of computer architectures is a very complex problem. The multiple parameters make the number of possible combinations extremely high.Many researchers have used simulation, although it is a slow solution since evaluating a single point of the search space can take hours. In this work we propose using evolutionary multilayer perceptron (MLP) to compute the performance of an architecture parameter settings. Instead of exploring the search space, simulating many configurations, our method randomly selects some architecture configurations; those are simulated to obtain their performance, and then an artificial neural network is trained to predict the remaining configurations performance. Results obtained show a high accuracy of the estimations using a simple method to select the configurations we have to simulate to optimize the MLP. In order to explore the search space, we have designed a genetic algorithm that uses the MLP as fitness function to find the niche where the best architecture configurations (those with higher performance) are located. Our models need only a small fraction of the design space, obtaining small errors and reducing required simulation by two orders of magnitude.
Pedro A. Castillo, Juan Julián Merelo Guervós, Miquel Moretó, Francisco J. Cazorla, Mateo Valero, Antonio Mora García, Juan Luis Jiménez Laredo, Sally A. McKee
IEEE Congress on Evolutionary Computation8
2008 Archer: A Community Distributed Computing Infrastructure for Computer Architecture Research and Education
Renato J. O. Figueiredo, P. Oscar Boykin, José A. B. Fortes, Tao Li 0006, Jie-Kwon Peir, David Wolinsky, Lizy Kurian John, David R. Kaeli, David J. Lilja, Sally A. McKee, Gokhan Memik, Alain J. Roy, Gary S. Tyson
CollaborateCom10
2008 Using Dynamic Binary Instrumentation to Generate Multi-platform SimPoints: Methodology and Accuracy
Vincent M. Weaver, Sally A. McKee
HiPEAC2
2008 A projection-based optimization framework for abstractions with application to the unstructured mesh domain
abstract
Computational scientists often must choose between the greater programming productivity of high-level abstractions, such as matrices and mesh entities, and the greater execution efficiency of low-level constructs. Performance is degraded when abstraction indirection introduces overhead and hinders compiler analysis. This can be overcome by targeting the semantics, rather than the implementation, of abstractions. Raising operators specified by a domain expert project an application from an implementation space to an abstraction space, where optimizations leverage domain semantics to complement conservative analyses. Raising operators define a domain-specific intermediate representation, which optimizations target for improved portability. Following optimization, transformed code is reified as a concrete implementation via lowering operators. We have developed a framework to implement this optimization strategy, which we use to introduce two domain-specific unstructured mesh optimizations. The first uses an inspector/executor approach to avoid costly traversals over a static mesh by memoizing the relatively few references required for mathematical computations. The executor phase accesses stored entities without incurring the indirections. The second optimization lowers object-based mesh access and iteration to a low-level implementation, which uses integer-based access and iteration.
Brian S. White, Sally A. McKee, Daniel J. Quinlan
ICS2
2008 Compiler-enhanced incremental checkpointing for OpenMP applications
abstract
As modern supercomputing systems reach peta-flop performance they grow in both size and complexity, becoming increasingly vulnerable to failures. Checkpointing is a popular technique for tolerating such failures. Although a variety of automated system-level checkpointing solutions are currently available to HPC users, manual application-level checkpointing remains more popular due to its superior performance. This paper improves performance of automated checkpointing by presenting a compiler analysis for incremental checkpointing. This analysis, which works with both sequential and OpenMP applications, significantly reduces checkpoint sizes and enables asynchronous checkpointing.
Greg Bronevetsky, Daniel Marques, Keshav Pingali, Radu Rugina, Sally A. McKee
PPoPP5
2008 Efficient architectural design space exploration via predictive modeling
abstract
Efficiently exploring exponential-size architectural design spaces with many interacting parameters remains an open problem: the sheer number of experiments required renders detailed simulation intractable. We attack this via an automated approach that builds accurate predictive models. We simulate sampled points, using results to teach our models the function describing relationships among design parameters. The models can be queried and are very fast, enabling efficient design tradeoff discovery. We validate our approach via two uniprocessor sensitivity studies, predicting IPC with only 1--2% error. In an experimental study using the approach, training on 1% of a 250-K-point CMP design space allows our models to predict performance with only 4--5% error. Our predictive modeling combines well with techniques that reduce the time taken by each simulation experiment, achieving net time savings of three-four orders of magnitude.
Engin Ipek, Sally A. McKee, Rich Caruana, Bronis R. de Supinski, Martin Schulz 0001
ACM Trans. Archit. Code Optim.2
2007 A Phase-Adaptive Approach to Increasing Cache Performance
Matthew A. Watkins, Sally A. McKee, Lambert Schaelicke
PACT2
2007 Identifying energy-efficient concurrency levels using machine learning
abstract
Multicore microprocessors have been largely motivated by the diminishing returns in performance and the increased power consumption of single-threaded ILP microprocessors. With the industry already shifting from multicore to many-core microprocessors, software developers must extract more thread-level parallelism from applications. Unfortunately, low power-efficiency and diminishing returns in performance remain major obstacles with many cores. Poor interaction between software and hardware, and bottlenecks in shared hardware structures often prevent scaling to many cores, even in applications where a high degree of parallelism is potentially available. In some cases, throwing additional cores at a problem may actually harm performance and increase power consumption. Better use of otherwise limitedly beneficial cores by software components such as hypervisors and operating systems can improve system-wide performance and reliability, even in cases where power consumption is not a main concern. In response to these observations, we evaluate an approach to throttle concurrency in parallel programs dynamically. We throttle concurrency to levels with higher predicted efficiency from both performance and energy standpoints, and we do so via machine learning, specifically artificial neural networks (ANNs). One advantage of using ANNs over similar techniques previously explored is that the training phase is greatly simplified, thereby reducing the burden on the end user. Using machine learning in the context of concurrency throttling is novel. We show that ANNs are effective for identifying energy-efficient concurrency levels in multithreaded scientific applications, and we do so using physical experimentation on a state-of-the-art quad-core Xeon platform.
Matthew Curtis-Maury, Sally A. McKee, Filip Blagojevic, Dimitrios S. Nikolopoulos, Bronis R. de Supinski, Martin Schulz 0001
CLUSTER3
2007 Leveraging High Performance Data Cache Techniques to Save Power in Embedded Systems
Major Bhadauria, Sally A. McKee, Gary S. Tyson
HiPEAC2
2007 Methods of inference and learning for performance modeling of parallel applications
abstract
Increasing system and algorithmic complexity combined with a growing number of tunable application parameters pose significant challenges for analytical performance modeling. We propose a series of robust techniques to address these challenges. In particular, we apply statistical techniques such as clustering, association, and correlation analysis, to understand the application parameter space better. We construct and compare two classes of effective predictive models: piecewise polynomial regression and artifical neural networks. We compare these techniques with theoretical analyses and experimental results. Overall, both regression and neural networks are accurate with median error rates ranging from 2.2 to 10.5 percent. The comparable accuracy of these models suggest differentiating features will arise from ease of use, transparency, and computational efficiency.
Benjamin C. Lee, David Brooks 0001, Bronis R. de Supinski, Martin Schulz 0001, Sally A. McKee
PPoPP6
2007 Predicting parallel application performance via machine learning approaches
abstract
Abstract Consistently growing architectural complexity and machine scales make the creation of accurate performance models for large‐scale applications increasingly challenging. Traditional analytic models are difficult and time consuming to construct, and are often unable to capture full system and application complexity. To address these challenges, we automatically build models based on execution samples. We use multilayer neural networks, because they can represent arbitrary functions and handle noisy inputs robustly. In this paper we focus on two well‐known parallel applications whose variations in execution times are not well understood: SMG 2000, a semicoarsening multigrid solver, and HPL, an open‐source implementation of LINPACK. We sparsely sample performance data on two radically different platforms across large, multidimensional parameter spaces and show that our models based on these data can predict performance within 2% to 7% of actual application runtimes. Copyright © 2007 John Wiley & Sons, Ltd.
Engin Ipek, Sally A. McKee, Bronis R. de Supinski, Martin Schulz 0001, Rich Caruana
Concurr. Comput. Pract. Exp.3
2007 Editorial to special issue on reliable computing
abstract
No abstract available.
Sally A. McKee
ACM J. Emerg. Technol. Comput. Syst.1
2007 METRIC: Memory tracing via dynamic binary rewriting to identify cache inefficiencies
abstract
With the diverging improvements in CPU speeds and memory access latencies, detecting and removing memory access bottlenecks becomes increasingly important. In this work we present METRIC, a software framework for isolating and understanding such bottlenecks using partial access traces. METRIC extracts access traces from executing programs without special compiler or linker support. We make four primary contributions. First, we present a framework for extracting partial access traces based on dynamic binary rewriting of the executing application. Second, we introduce a novel algorithm for compressing these traces. The algorithm generates constant space representations for regular accesses occurring in nested loop structures. Third, we use these traces for offline incremental memory hierarchy simulation. We extract symbolic information from the application executable and use this to generate detailed source-code correlated statistics including per-reference metrics, cache evictor information, and stream metrics. Finally, we demonstrate how this information can be used to isolate and understand memory access inefficiencies. This illustrates a potential advantage of METRIC over compile-time analysis for sample codes, particularly when interprocedural analysis is required.
Jaydeep Marathe, Frank Mueller 0001, Tushar Mohan, Sally A. McKee, Bronis R. de Supinski, Andy B. Yoo
ACM Trans. Program. Lang. Syst.4
2006 Efficiently exploring architectural design spaces via predictive modeling
abstract
Architects use cycle-by-cycle simulation to evaluate design choices and understand tradeoffs and interactions among design parameters. Efficiently exploring exponential-size design spaces with many interacting parameters remains an open problem: the sheer number of experiments renders detailed simulation intractable. We attack this problem via an automated approach that builds accurate, confident predictive design-space models. We simulate sampled points, using the results to teach our models the function describing relationships among design parameters. The models produce highly accurate performance estimates for other points in the space, can be queried to predict performance impacts of architectural changes, and are very fast compared to simulation, enabling efficient discovery of tradeoffs among parameters in different regions. We validate our approach via sensitivity studies on memory hierarchy and CPU design spaces: our models generally predict IPC with only 1-2% error and reduce required simulation by two orders of magnitude. We also show the efficacy of our technique for exploring chip multiprocessor (CMP) design spaces: when trained on a 1% sample drawn from a CMP design space with 250K points and up to 55x performance swings among different system configurations, our models predict performance with only 4-5% error on average. Our approach combines with techniques to reduce time per simulation, achieving net time savings of three-four orders of magnitude.
Engin Ipek, Sally A. McKee, Rich Caruana, Bronis R. de Supinski, Martin Schulz 0001
ASPLOS2
2006 Dynamic program phase detection in distributed shared-memory multiprocessors
abstract
We present a novel hardware mechanism for dynamic program phase detection in distributed shared-memory (DSM) multiprocessors. We show that successful hardware mechanisms for phase detection in uniprocessors do not necessarily work well in DSM systems, since they lack the ability to incorporate the parallel application's global execution information and memory access behavior based on data distribution. We then propose a hardware extension to a well-known uniprocessor mechanism that significantly improves phase detection in the context of DSM multiprocessors. The resulting mechanism is modest in size and complexity, and is transparent to the parallel application.
Engin Ipek, José F. Martínez, Bronis R. de Supinski, Sally A. McKee, Martin Schulz 0001
IPDPS4
2005 An Approach to Performance Prediction for Parallel Applications
Engin Ipek, Bronis R. de Supinski, Martin Schulz 0001, Sally A. McKee
Euro-Par4
2005 Beyond Basic Region Caching: Specializing Cache Structures for High Performance and Energy Conservation
Michael J. Geiger, Sally A. McKee, Gary S. Tyson
HiPEAC2
2005 Improving the computational intensity of unstructured mesh applications
abstract
Although unstructured mesh algorithms are a popular means of solving problems across a broad range of disciplines---from texture mapping to computational fluid dynamics---they are often dominated not by computation, but by mesh overhead. Our study of an object-oriented mesh-based benchmark reveals that 72% of its execution time is spent on mesh-related operations, such as iterating over faces or chasing pointers. We report a series of optimizations---some traditional, some novel---that dramatically improve the benchmark's computational intensity---the ratio of floating point operations to memory accesses. This improvement is attributable to an eight-fold reduction in memory operations and results in a 4.7x speedup in execution time.Our work demonstrates that common subexpression elimination and code motion are important optimizations for mesh-based codes. However, conservative analysis prevents their application. We discuss these barriers to analysis and argue that an understanding of mesh semantics complements more traditional analyses, such as pointer alias analysis, and certifies the correctness of these optimizations. Our identification of overheads in mesh-based codes, optimizations that address them, and limitations of current compiler analyses are required for our eventual goal of automating these optimizations in a semantics-aware compiler.
Brian S. White, Sally A. McKee, Bronis R. de Supinski, Brian Miller 0001, Daniel J. Quinlan, Martin Schulz 0001
ICS2
2004 Formal hardware specification languages for protocol compliance verification
abstract
The advent of the system-on-chip and intellectual property hardware design paradigms makes protocol compliance verification increasingly important to the success of a project. One of the central tools in any verification project is the modeling language, and we survey the field of candidate languages for protocol compliance verification, limiting our discussion to languages originally intended for hardware and software design and verification activities. We frame our comparison by first constructing a taxonomy of these languages, and then by discussing the applicability of each approach to the compliance verification problem. Each discussion includes a summary of the development of the language, an evaluation of the language's utility for our problem domain, and, where feasible, an example of how the language might be used to specify hardware protocols. Finally, we make some general observations regarding the languages considered.
Annette Bunker, Ganesh Gopalakrishnan, Sally A. McKee
ACM Trans. Design Autom. Electr. Syst.3
2003 METRIC: Tracking Down Inefficiencies in the Memory Hierarchy via Binary Rewriting
abstract
We present METRIC, an environment for determining memory inefficiencies by examining data traces. METRIC is designed to alter the performance behavior of applications that are mostly constrained by their latency to resolve memory references. We make four primary contributions. First, we present methods to extract partial data traces from running applications by observing their memory behavior via dynamic binary rewriting. Second, we present a methodology to represent partial data traces in constant space for regular references through a novel technique for online compression of reference streams. Third, we employ offline cache simulation to derive indications about memory performance bottlenecks from partial data traces. By exploiting summarized memory metrics, by-reference metrics as well as cache evictor information, we can pin-point the sources of performance problems. Fourth, we demonstrate the ability to derive opportunities for optimizations and assess their benefits in several experiments resulting in up to 40% lower miss ratios.
Jaydeep Marathe, Frank Mueller 0001, Tushar Mohan, Bronis R. de Supinski, Sally A. McKee, Andy B. Yoo
CGO5
2003 An MPEG-4 performance study for non-SIMD, general purpose architectures
abstract
MPEG-4 is an important international standard with wide applicability. This paper focuses on MPEG-4's main profile, video, whose approach allows more efficiency in coding and more flexibility in managing heterogeneous media objects than previous MPEG standards. This study presents evidence to support the assertion that for non-SIMD architectures and computational models, most memory-system optimizations will have little effect on MPEG-4 performance. This paper makes two contributions. First, it serves as an independent confirmation that for current, general-purpose architectures, MPEG-4 video is computation bound (just like most other media processing applications). Second, our findings should prove useful to other researchers and practitioners considering how to (or how not to) optimize MPEG-4 performance.
Sally A. McKee, Zhen Fang 0002, Mateo Valero
ISPASS1
2003 Identifying and Exploiting Spatial Regularity in Data Memory References
abstract
The growing processor/memory performance gap causes the performance of many codes to be limited by memory accesses. If known to exist in an application, strided memory accesses forming streams can be targeted by optimizations such as prefetching, relocation, remapping, and vector loads. Undetected, they can be a significant source of memory stalls in loops. Existing stream-detection mechanisms either require special hardware, which may not gather statistics for subsequent analysis, or are limited to compile-time detection of array accesses in loops. Formally, little treatment has been accorded to the subject; the concept of locality fails to capture the existence of streams in a program's memory accesses. The contributions of this paper are as follows. First, we define spatial regularity as a means to discuss the presence and effects of streams. Second, we develop measures to quantify spatial regularity, and we design and implement an on-line, parallel algorithm to detect streams - and hence regularity - in running applications. Third, we use examples from real codes and common benchmarks to illustrate how derived stream statistics can be used to guide the application of profile-driven optimizations. Overall, we demonstrate the benefits of our novel regularity metric as an instrument to detect potential for code optimizations affecting memory performance.
Tushar Mohan, Bronis R. de Supinski, Sally A. McKee, Frank Mueller 0001, Andy B. Yoo, Martin Schulz 0001
SC3
2002 Computation regrouping: restructuring programs for temporal data cache locality
abstract
Data access costs contribute significantly to the execution time of applications with complex data structures. As the latency of memory accesses becomes high relative to processor cycle times, application performance is increasingly limited by memory performance. In some situations it may be reasonable to trade increased computation costs for reduced memory costs. The contributions of this paper are three-fold: we provide a detailed analysis of the memory performance of a set of seven, memory-intensive benchmarks; we describe Computation Regrouping, a general, source-level approach to improving the overall performance of these applications by improving temporal locality to reduce cache and TLB miss ratios (and thus memory stall times); and we demonstrate significant performance improvements from applying Computation Regrouping to our suite of seven benchmarks. With Computation Regrouping, we observe an average speedup of 1.97, with individual speedups ranging from 1.26 to 3.03. Most of this improvement comes from eliminating memory stall time.
Venkata K. Pingali, Sally A. McKee, Wilson C. Hsieh, John B. Carter
ICS2
2001 Reevaluating Online Superpage Promotion with Hardware Support
abstract
Typical translation lookaside buffers (TLBs) can map a far smaller region of memory than application footprints demand, and the cost of handling TLB misses therefore limits the performance of an increasing number of applications. This bottleneck can be mitigated by the use of superpages, multiple adjacent virtual memory pages that can be mapped with a single TLB entry that extend TLB reach without significantly increasing size or cost. We analyze hardware/software tradeoff for dynamically creating superpages. This study extends previous work by using execution-driven simulation to compare creating superpages via copying with remapping pages within the memory controller and by examining how the tradeoffs change when moving front a single-issue to a superscalar processor model. We find that remapping-based promotion outperforms copying-based promotion, often significantly. Copying-based promotion is slightly more effective on superscalar processors than on single-issue processors, and the relative performance of remapping-based promotion on the two platform is application-dependent.
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Wilson C. Hsieh, Sally A. McKee
HPCA5
2001 The Impulse Memory Controller
abstract
Impulse is a memory system architecture that adds an optional level of address indirection at the memory controller. Applications can use this level of indirection to remap their data structures in memory. As a result, they can control how their data is accessed and cached, which can improve cache and bus utilization. The Impulse design does not require any modification to processor, cache, or bus designs since all the functionality resides at the memory controller. As a result, Impulse can be adopted in conventional systems without major system changes. We describe the design of the Impulse architecture and how an Impulse memory system can be used in a variety of ways to improve the performance of memory-bound applications. Impulse can be used to dynamically create superpages cheaply, to dynamically recolor physical pages, to perform strided fetches, and to perform gathers and scatters through indirection vectors. Our performance results demonstrate the effectiveness of these optimizations in a variety of scenarios. Using Impulse can speed up a range of applications from 20 percent to over a factor of 5. Alternatively, Impulse can be used by the OS for dynamic superpage creation; the best policy for creating superpages using Impulse outperforms previously known superpage creation policies.
Lixin Zhang 0002, Zhen Fang 0002, Michael A. Parker, Binu K. Mathew, Lambert Schaelicke, John B. Carter, Wilson C. Hsieh, Sally A. McKee
IEEE Trans. Computers8
2000 Design of a Parallel Vector Access Unit for SDRAM Memory Systems
abstract
We are attacking the memory bottleneck by building a "smart" memory controller that improves effective memory bandwidth, bus utilization, and cache efficiency by letting applications dictate how their data is accessed and cached. This paper describes a parallel vector access unit (PVA), the vector memory subsystem that efficiently "gathers" sparse, strided data structures in parallel on a multi-bank SDRAM memory. We have validated our PVA design via gate-level simulation, and have evaluated its performance via functional simulation and formal analysis. On unit-stride vectors, PVA performance equals or exceeds that of an SDRAM system optimized for cache line fills. On vectors with larger strides, the PVA is up to 32.8 times faster. Our design is up to 3.3 times faster than a pipelined, serial SDRAM memory system that gathers sparse vector data, and the gathering mechanism is two to five times faster than in other PVAs with similar goals. Our PVA only slightly increases hardware complexity with respect to these other systems, and the scalable design is appropriate for a range of computing platforms, from vector supercomputers to commodity PCs.
Binu K. Mathew, Sally A. McKee, John B. Carter, Al Davis
HPCA2
2000 Hardware-only stream prefetching and dynamic access ordering
abstract
Memory system bottlenecks limit performance for many applications, and computations with strided access patterns are among the hardest hit. The streams used in such applications have extremely poor cache behavior. These access patterns have the advantage of being predictable, though, and this can be exploited to improve the efficiency of the memory subsystem in two ways: memory latencies can be masked by prefetching stream data, and the latencies can be reduced by reordering stream accesses to exploit parallelism and locality within the DRAMs. Many researchers have studied hardware prefetching in its various forms. Others have examined dynamic memory scheduling to help bridge the performance gap between processors and DRAM memory systems. This study builds on these results, combining a stride-based reference prediction table, a mechanism that prefetches L2 cache lines, and a memory controller that dynamically schedules accesses to a Direct Rambus memory subsystem. We find that such a s...
Chengqiang Zhang, Sally A. McKee
ICS2
2000 Profiling I/O Interrupts in Modern Architectures
abstract
As applications grow increasingly communication-oriented, interrupt performance quickly becomes a crucial component of high performance I/O system design. At the same time, accurately measuring interrupt handler performance is difficult with the traditional simulation, instrumentation, or statistical sampling approaches. One of the most important components of interrupt performance is cache behavior. This paper presents a portable method for measuring the cache effects of I/O interrupt handling using hardware performance counters. The method is demonstrated on two commercial platforms with different architectures, the SGI Origin 200 and the Sun Ultra-1. This case study uses the methodology to measure the overhead of the two most common forms of interrupts: disk and network interrupts. It demonstrates that the method works well and is reasonably robust. In addition, the results show that network interrupts have larger cache footprints than disk interrupts, and behave very differently on both platforms, due to significant differences in OS organization.
Lambert Schaelicke, Al Davis, Sally A. McKee
MASCOTS3
2000 Online superpage promotion revisited (poster)
abstract
No abstract available.
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Sally A. McKee, Wilson C. Hsieh
SIGMETRICS4
2000 Algorithmic foundations for a parallel vector access memory system
abstract
This paper presents mathematical foundations for the design of a memory controller subcomponent that helps to bridge the processor/memory performance gap for applications with strided access patterns. The Parallel Vector Access (PVA) unit exploits the regularity of vectors or streams to access them efficiently in parallel on a multi-bank SDRAM memory system. The PVA unit performs scatter/gather operations so that only the elements accessed by the application are transmitted across the system bus. Vector operations are broadcast in parallel to all memory banks, each of which implements an efficient algorithm to determine which vector elements it holds. Earlier performance evaluations have demonstrated that our PVA implementation loads elements up to 32.8 times faster than a conventional memory system and 3.3 times faster than a pipelined vector unit, without hurting the performance of normal cache-line fills. Here we present the underlying PVA algorithms for both word interleaved and cache-line inter-leaved memory systems.
Binu K. Mathew, Sally A. McKee, John B. Carter, Al Davis
SPAA2
2000 Dynamic Access Ordering for Streamed Computations
abstract
Memory bandwidth is rapidly becoming the limiting performance factor for many applications, particularly for streaming computations such as scientific vector processing or multimedia (de)compression. Although these computations lack the temporal locality of reference that makes traditional caching schemes effective, they have predictable access patterns. Since most modern DRAM components support modes that make it possible to perform some access sequences faster than others, the predictability of the stream accesses makes it possible to reorder them to get better memory performance. We describe a Stream Memory Controller (SMC) system that combines compile-time detection of streams with execution-time selection of the access order and issue. The SMC effectively prefetches read-streams, buffers write-streams, and reorders the accesses to exploit the existing memory bandwidth as much as possible. Unlike most other hardware prefetching or stream buffer designs, this system does not increase bandwidth requirements. The SMC is practical to implement, using existing compiler technology and requiring only a modest amount of special purpose hardware. We present simulation results for fast-page mode and Rambus DRAM memory systems and we describe a prototype system with which we have observed performance improvements for inner loops by factors of 13 over traditional access methods.
Sally A. McKee, William A. Wulf, James H. Aylor, Robert H. Klenke, Maximo H. Salinas, Sung I. Hong, Dee A. B. Weikle
IEEE Trans. Computers1
1999 Access Order and Effective Bandwidth for Streams on a Direct Rambus Memory
abstract
Processor speeds are increasing rapidly and memory speeds are not keeping up. Streaming computations (such as multimedia or scientific applications) are among those whose performance is most limited by the memory bottleneck. Rambus hopes to bridge the processor/memory performance gap with a recently introduced DRAM that can deliver up to 1.6 Gbytes/sec. We analyze the performance of these interesting new memory devices on the inner loops of streaming computations, both for traditional memory controllers that treat all DRAM transactions as random cacheline accesses, and for controllers augmented with streaming hardware. For our benchmarks, we find that accessing unit-stride streams in cacheline bursts in the natural order of the computation exploits from 44-76% of the peak bandwidth of a memory system composed of a single Direct RDRAM device, and that accessing streams via a streaming mechanism with a simple access ordering scheme can improve performance by factors of 1.18 to 2.25.
Sung I. Hong, Sally A. McKee, Maximo H. Salinas, Robert H. Klenke, James H. Aylor, William A. Wulf
HPCA2
1998 Caches as Filters: A New Approach to Cache Analysis
abstract
As the processor-memory performance gap continues to grow, so does the need for effective tools and metrics to guide the design of efficient memory hierarchies to bridge that gap. Aggregate statistics of cache performance can be useful for comparison, but they give us little insight into how to improve the design of a particular component. We propose a different approach to cache analysis-viewing caches as filters-and present two new metrics for analyzing cache behavior: instantaneous hit rate and instantaneous locality. We demonstrate how these measures can give us insight into the reference pattern of an executing program, and show an application of these measures in analyzing the effectiveness of the second level cache of a particular memory hierarchy.
Dee A. B. Weikle, Sally A. McKee, William A. Wulf
MASCOTS2
1996 Design and Evaluation of Dynamic Access Ordering Hardware
abstract
Memory bandwidth is rapidly becoming the limiting performance factor for many applications, particularly for streaming computations such as scientific vector processing or multimedia (de)compression. Although these computations lack the temporal locality of reference that makes caches effective, they have predictable access patterns. Since most modern DRAM components support modes that make it possible to perform some access sequences faster than others, the predictability of the stream accesses makes it possible to reorder them to get better memory performance. We describe and evaluate a Stream Memory Controller system that combines compile-time detection of streams with execution-time selection of the access order and issue. The technique is practical to implement, using existing compiler technology and requiring only a modest amount of special-purpose hardware. With our prototype system, we have observed performance improvements by factors of 13 over normal caching. 1. INTRODUCTION...
Sally A. McKee, Assaji Aluwihare, Benjamin H. Clark, Robert H. Klenke, Trevor C. Landon, Christopher W. Oliver, Maximo H. Salinas, Adam E. Szymkowiak, Kenneth L. Wright, William A. Wulf, James H. Aylor
International Conference on Supercomputing1
1995 Bounds on Memory Bandwidth in Streamed Computations
Sally A. McKee, William A. Wulf, Trevor C. Landon
Euro-Par1
1995 Access Ordering and Memory-Conscious Cache Utilization
abstract
As processor speeds increase relative to memory speeds, memory bandwidth is rapidly becoming the limiting performance, factor for many applications. Several approaches to bridging this performance gap have been suggested. This paper examines one approach, access ordering, and pushes its limits to determine bounds on memory performance. We present several access-ordering schemes, and compare their performance, developing analytic models and partially validating these with benchmark timings on the Intel i860XR.>
Sally A. McKee, William A. Wulf
HPCA1
1993 Toward a Steiner engine: enhanced serial and parallel implementations of the iterated 1-Steiner MRST algorithm
abstract
The minimum rectilinear Steiner tree (MRST) problem is known to be NP-hard, and the best performing MRST heuristic to date is the Iterated 1-Steiner (I1S) method recently proposed by A.B. Kahng and G. Robins (1992). The authors develop a straightforward, efficient implementation of I1S, achieving speedup factors of over 200 compared to previous implementations. They also propose a parallel implementation of I1S that achieves high parallel speedup on K processors. Extensive empirical testing confirms the viability of the approach, which allows the benchmarking of I1S on nets containing several hundred pins.>
Tim Barrera, Jeff Griffith, Sally A. McKee, Gabriel Robins, Tongtong Zhang
Great Lakes Symposium on VLSI3