John B. Carter

dblp:c/JohnBCarter · DBLP profile ↗
← Back
46ranked-venue papers
9as first author
0since 2021 · last 2016
0000-0002-1395-0254ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 38 · 7 first-authorSoftware engineering, systems software and programming languages · 7 · 3 first-authorComputer networks · 5Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
21 papers
Memory systems · 59% Energy-efficient computing · 12% Cloud and datacenter computing · 10%
Computer networks
6 papers
Datacenter networks · 34% Routing and switching · 28% Transport protocols and congestion control · 16%
Software engineering, system software, and programming languages
4 papers
Operating systems · 100%

Topics — the 30 heaviest of 74, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.372008
Extending CC-NUMA systems to support write update optimizations · SC 2008
An Adaptive Cache Coherence Protocol Optimized for Producer-Consumer Sharing · HPCA 2007
Interconnect-Aware Coherence Protocols for Chip Multiprocessors · ISCA 2006
Cloud and datacenter computing › multi-tenancy
multi-tenant datacenter
0.212016
AC/DC TCP: Virtual Congestion Control Enforcement for Datacenter Networks · SIGCOMM 2016
Energy-efficient computing
power management
0.222011
Active management of timing guardband to save energy in POWER7 · MICRO 2011
Architecting for power management: The IBM POWER7TM approach · HPCA 2010
Datacenter networks
load balancing
0.212015
Presto: Edge-based Load Balancing for Fast Datacenter Networks · SIGCOMM 2015
Memory systems › cache coherence
cache coherence protocol
0.232008
Extending CC-NUMA systems to support write update optimizations · SC 2008
An Adaptive Cache Coherence Protocol Optimized for Producer-Consumer Sharing · HPCA 2007
Interconnect-Aware Coherence Protocols for Chip Multiprocessors · ISCA 2006
Routing and switching › traffic engineering
congestion-aware rerouting
0.212014
Planck: millisecond-scale monitoring and control for commodity networks · SIGCOMM 2014
Datacenter networks
lossless ethernet
0.212014
Practical DCB for improved data center networks · INFOCOM 2014
Network management and operations
network monitoring
0.212014
Planck: millisecond-scale monitoring and control for commodity networks · SIGCOMM 2014
Transport protocols and congestion control
TCP variants
0.212014
Practical DCB for improved data center networks · INFOCOM 2014
Routing and switching
traffic engineering
0.212014
Planck: millisecond-scale monitoring and control for commodity networks · SIGCOMM 2014
Memory systems
DRAM
0.232012
Tiered Memory: An Iso-Power Memory Architecture to Address the Memory Power Wall · IEEE Trans. Computers 2012
Design of a Parallel Vector Access Unit for SDRAM Memory Systems · HPCA 2000
The Impulse Memory Controller · IEEE Trans. Computers 2001
Datacenter networks
data center network topology
0.112012
PAST: scalable ethernet for data centers · CoNEXT 2012
Datacenter networks › data center network topology
fat-tree
0.112012
PAST: scalable ethernet for data centers · CoNEXT 2012
Network performance modeling › protocol performance analysis › routing performance
routing scalability
0.112012
PAST: scalable ethernet for data centers · CoNEXT 2012
Routing and switching › routing protocol
spanning tree routing
0.112012
PAST: scalable ethernet for data centers · CoNEXT 2012
Memory systems
tiered memory
0.112012
Tiered Memory: An Iso-Power Memory Architecture to Address the Memory Power Wall · IEEE Trans. Computers 2012
Hardware reliability and fault tolerance
timing guardband
0.112011
Active management of timing guardband to save energy in POWER7 · MICRO 2011
Operating systems › resource management
memory management
0.132009
Dynamic hardware-assisted software-controlled page placement to manage capacity allocation and sharing within large caches · HPCA 2009
Online superpage promotion revisited (poster) · SIGMETRICS 2000
Increasing TLB Reach Using Superpages Backed by Shadow Memory · ISCA 1998
Energy-efficient computing › power management
dynamic power management
0.112010
Architecting for power management: The IBM POWER7TM approach · HPCA 2010
Memory systems › shared memory
distributed shared memory
0.162006
Efficient address remapping in distributed shared-memory systems · ACM Trans. Archit. Code Optim. 2006
Techniques for Reducing Consistency-Related Communication in Distributed Shared-Memory Systems · ACM Trans. Comput. Syst. 1995
Implementation and Performance of Munin · SOSP 1991
Operating systems › resource management › memory management
page coloring
0.112009
Dynamic hardware-assisted software-controlled page placement to manage capacity allocation and sharing within large caches · HPCA 2009
Memory systems
cache management
0.112009
Dynamic hardware-assisted software-controlled page placement to manage capacity allocation and sharing within large caches · HPCA 2009
Memory systems › cache design
non-uniform cache architecture
0.112009
Dynamic hardware-assisted software-controlled page placement to manage capacity allocation and sharing within large caches · HPCA 2009
Memory systems
page placement
0.112009
Dynamic hardware-assisted software-controlled page placement to manage capacity allocation and sharing within large caches · HPCA 2009
Memory systems › memory management
address remapping
0.122006
Efficient address remapping in distributed shared-memory systems · ACM Trans. Archit. Code Optim. 2006
Impulse: Building a Smarter Memory Controller · HPCA 1999
Processor architecture and microarchitecture
multicore design
0.112008
Extending CC-NUMA systems to support write update optimizations · SC 2008
Memory systems › cache coherence
write-update protocol
0.112008
Extending CC-NUMA systems to support write update optimizations · SC 2008
Memory systems
memory controller
0.132001
The Impulse Memory Controller · IEEE Trans. Computers 2001
Design of a Parallel Vector Access Unit for SDRAM Memory Systems · HPCA 2000
Impulse: Building a Smarter Memory Controller · HPCA 1999
Memory systems › memory management
virtual memory
0.132001
Reevaluating Online Superpage Promotion with Hardware Support · HPCA 2001
Online superpage promotion revisited (poster) · SIGMETRICS 2000
Increasing TLB Reach Using Superpages Backed by Shadow Memory · ISCA 1998
Transport protocols and congestion control › explicit congestion control
DCTCP
0.112016
AC/DC TCP: Virtual Congestion Control Enforcement for Datacenter Networks · SIGCOMM 2016

Methods — techniques the papers use, named apart from their topics

virtual switch implementation · 0.7shadow address space · 0.2hardware-assisted page placement · 0.2port mirroring · 0.2deadlock-free routing · 0.2TCP congestion control · 0.2simulation · 0.2cycle-accurate simulation · 0.2per-address spanning tree routing · 0.1data migration · 0.1DRAM self-refresh · 0.1energyscale methodology · 0.1speculative update mechanism · 0.1model checking · 0.1online algorithm · 0.0protocol implementation · 0.0
YearPublicationVenuePosition
2016 AC/DC TCP: Virtual Congestion Control Enforcement for Datacenter Networks
abstract
Multi-tenant datacenters are successful because tenants can seamlessly port their applications and services to the cloud. Virtual Machine (VM) technology plays an integral role in this success by enabling a diverse set of software to be run on a unified underlying framework. This flexibility, however, comes at the cost of dealing with out-dated, inefficient, or misconfigured TCP stacks implemented in the VMs. This paper investigates if administrators can take control of a VM's TCP congestion control algorithm without making changes to the VM or network hardware. We propose AC/DC TCP, a scheme that exerts fine-grained control over arbitrary tenant TCP stacks by enforcing per-flow congestion control in the virtual switch (vSwitch). Our scheme is light-weight, flexible, scalable and can police non-conforming flows. In our evaluation the computational overhead of AC/DC TCP is less than one percentage point and we show implementing an administrator-defined congestion control algorithm in the vSwitch (i.e., DCTCP) closely tracks its native performance, regardless of the VM's TCP stack.
Keqiang He, Eric Rozner, Kanak Agarwal 0001, Yu Gu 0001, Wes Felter, John B. Carter, Aditya Akella
SIGCOMM6
2015 Presto: Edge-based Load Balancing for Fast Datacenter Networks
abstract
Datacenter networks deal with a variety of workloads, ranging from latency-sensitive small flows to bandwidth-hungry large flows. Load balancing schemes based on flow hashing, e.g., ECMP, cause congestion when hash collisions occur and can perform poorly in asymmetric topologies. Recent proposals to load balance the network require centralized traffic engineering, multipath-aware transport, or expensive specialized hardware. We propose a mechanism that avoids these limitations by (i) pushing load-balancing functionality into the soft network edge (e.g., virtual switches) such that no changes are required in the transport layer, customer VMs, or networking hardware, and (ii) load balancing on fine-grained, near-uniform units of data (flowcells) that fit within end-host segment offload optimizations used to support fast networking speeds. We design and implement such a soft-edge load balancing scheme, called Presto, and evaluate it on a 10 Gbps physical testbed. We demonstrate the computational impact of packet reordering on receivers and propose a mechanism to handle reordering in the TCP receive offload functionality. Presto's performance closely tracks that of a single, non-blocking switch over many workloads and is adaptive to failures and topology asymmetry.
Keqiang He, Eric Rozner, Kanak Agarwal 0001, Wes Felter, John B. Carter, Aditya Akella
SIGCOMM5
2014 OpenSample: A Low-Latency, Sampling-Based Measurement Platform for Commodity SDN
Junho Suh, Ted Taekyoung Kwon, Colin Dixon, Wes Felter, John B. Carter
ICDCS5
2014 Practical DCB for improved data center networks
abstract
Storage area networking is driving commodity data center switches to support lossless Ethernet (DCB). Unfortunately, to enable DCB for all traffic on arbitrary network topologies, we must address several problems that can arise in lossless networks, e.g., large buffering delays, unfairness, head of line blocking, and deadlock. We propose TCP-Bolt, a TCP variant that not only addresses the first three problems but reduces flow completion times by as much as 70%. We also introduce a simple, practical deadlock-free routing scheme that eliminates deadlock while achieving aggregate network throughput within 15% of ECMP routing. This small compromise in potential routing capacity is well worth the gains in flow completion time. We note that our results on deadlock-free routing are also of independent interest to the storage area networking community. Further, as our hardware testbed illustrates, these gains are achievable today, without hardware changes to switches or NICs.
Brent E. Stephens, Alan L. Cox, Ankit Singla, John B. Carter, Colin Dixon, Wes Felter
INFOCOM4
2014 Planck: millisecond-scale monitoring and control for commodity networks
abstract
Software-defined networking introduces the possibility of building self-tuning networks that constantly monitor network conditions and react rapidly to important events such as congestion. Unfortunately, state-of-the-art monitoring mechanisms for conventional networks require hundreds of milliseconds to seconds to extract global network state, like link utilization or the identity of "elephant" flows. Such latencies are adequate for responding to persistent issues, e.g., link failures or long-lasting congestion, but are inadequate for responding to transient problems, e.g., congestion induced by bursty workloads sharing a link. In this paper, we present Planck, a novel network measurement architecture that employs oversubscribed port mirroring to extract network information at 280 µs--7 ms timescales on a 1 Gbps commodity switch and 275 µs--4 ms timescales on a 10 Gbps commodity switch,over 11x and 18x faster than recent approaches, respectively (and up to 291x if switch firmware allowed buffering to be disabled on some ports). To demonstrate the value of Planck's speed and accuracy, we use it to drive a traffic engineering application that can reroute congested flows in milliseconds. On a 10 Gbps commodity switch, Planck-driven traffic engineering achieves aggregate throughput within 1--4% of optimal for most workloads we evaluated, even with flows as small as 50 MiB, an improvement of up to 53% over previous schemes.
Jeff Rasley, Brent E. Stephens, Colin Dixon, Eric Rozner, Wes Felter, Kanak Agarwal 0001, John B. Carter, Rodrigo Fonseca
SIGCOMM7
2012 PAST: scalable ethernet for data centers
abstract
We present PAST, a novel network architecture for data center Ethernet networks that implements a Per-Address Spanning Tree routing algorithm. PAST preserves Ethernet's self-configuration and mobility support while increasing its scalability and usable bandwidth. PAST is explicitly designed to accommodate unmodified commodity hosts and Ethernet switch chips. Surprisingly, we find that PAST can achieve performance comparable to or greater than Equal-Cost Multipath (ECMP) forwarding, which is currently limited to layer-3 IP networks, without any multipath hardware support. In other words, the hardware and firmware changes proposed by emerging standards like TRILL are not required for high-performance, scalable Ethernet networks. We evaluate PAST on Fat Tree, HyperX, and Jellyfish topologies, and show that it is able to capitalize on the advantages each offers. We also describe an OpenFlow-based implementation of PAST in detail.
Brent E. Stephens, Alan L. Cox, Wes Felter, Colin Dixon, John B. Carter
CoNEXT5
2012 Tiered Memory: An Iso-Power Memory Architecture to Address the Memory Power Wall
abstract
Moore's Law improvement in transistor density is driving a rapid increase in the number of cores per processor. DRAM device capacity and energy efficiency are increasing at a slower pace, so the importance of DRAM power is increasing. This problem presents system designers with two nominal options when designing future systems: 1) decrease off-chip memory capacity and bandwidth per core or 2) increase the fraction of system power allocated to main memory. Reducing capacity and bandwidth leads to imbalanced systems with poor processor utilization for noncache-resident applications, so designers have chosen to increase DRAM power budget. This choice has been viable to date, but is fast running into a memory power wall. To address the looming memory power wall problem, we propose a novel iso-power tiered memory architecture that supports 2-3X more memory capacity for the same power budget as traditional designs by aggressively exploiting low-power DRAM modes. We employ two "tiers” of DRAM, a "hot” tier with active DRAM and a "cold” tier in which DRAM is placed in self-refresh mode. The DRAM capacity of each tier is adjusted dynamically based on aggregate workload requirements and the most frequently accessed data are migrated to the "hot” tier. This design allows larger memory capacities at a fixed power budget while mitigating the performance impact of using low-power DRAM modes. We target our solution at server consolidation scenarios where physical memory capacity is typically the primary factor limiting the number of virtual machines a server can support. Using iso-power tiered memory, we can run 3× as many virtual machines, achieving a 250 percent improvement in average aggregate performance, compared to a conventional memory design with the same power budget.
Kshitij Sudan, Karthick Rajamani, Wei Huang 0004, John B. Carter
IEEE Trans. Computers4
2012 Active memory controller
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Sally A. McKee, Ali Ibrahim, Michael A. Parker, Xiaowei Jiang
J. Supercomput.3
2011 Active management of timing guardband to save energy in POWER7
abstract
Microprocessor voltage levels include substantial margin to deal with process variation, system power supply variation, workload induced thermal and voltage variation, aging, random uncertainty, and test inaccuracy. This margin allows the microprocessor to operate correctly during worst-case conditions, but during typical conditions it is larger than necessary and wastes energy. We present a mechanism that reduces excess voltage margin by (1) introducing a critical path monitor (CPM) circuit that measures available timing margin in real-time, (2) coupling the CPM output to the clock generation circuit to adjust clock frequency within cycles in response to excess or inadequate timing margin, and (3) adjusting the processor voltage level periodically in firmware to achieve a specified average clock frequency target. We implemented this mechanism in a prototype IBM POWER7 server. During better-than-worst case conditions our guardband management mechanism reduces the average voltage setting 137-152 mV below nominal, resulting in average processor power reduction of 24% with no performance loss while running industry-standard benchmarks.
Charles Lefurgy, Alan J. Drake, Michael S. Floyd, Malcolm Allen-Ware, Bishop Brock, José A. Tierno, John B. Carter
MICRO7
2011 Reliability-aware energy management for hybrid storage systems
abstract
Modern disk-based storage systems are not energy proportional, because disks consume almost as much power when idle (but spinning) as they do when actively accessing data. We combine a power-aware, solid-state (flash) cache and a reliability-aware disk spindown mechanism to significantly improve storage energy proportionality without hurting disk reliability, data integrity, or performance. We evaluated the resulting power- and reliability-aware hybrid flash-disk RAID storage array and found that it reduces energy consumption by 85% compared to a similar-cost, similar-performance typical configuration of all SAS drives that are never spun down. Our design also achieves almost 50% energy savings compared to hybrid flash-disk systems tuned for performance or that do not take full advantage of opportunities for safe spindown. Further, unlike most previous work that exploits spindown to save energy, we limit the rate at which disks are spun down to avoid premature mechanical failures, whereas reliability-unaware spindown algorithms can exceed manufacturer waranteed lifetime spindown limits in as little as one year.
Wes Felter, Anthony Hylick, John B. Carter
MSST3
2010 Architecting for power management: The IBM POWER7TM approach
abstract
The POWER7 processor is the newest member of the IBM POWER®family of server processors. With greater than 4X the peak performance and the same power budget as the previous generation POWER6®, POWER7 will deliver impressive energy-efficiency boosts. The improved peak energy-efficiency is accompanied by a wide array of new features in the processor and system designs that advance IBM's EnergyScaleTMdynamic power management methodology. This paper provides an overview of these new features, which include better sensing, more advanced power controls, improved scalability for power management, and features to address the diverse needs of the full range of POWER servers from blades to supercomputers. We also highlight three challenges that need attention from a range of systems design and research teams: (i) power management in highly virtualized environments, (ii) power (in)efficiency of systems software and applications, and (iii) memory power costs, especially for servers with large memory footprints.
Malcolm Allen-Ware, Karthick Rajamani, Michael S. Floyd, Bishop Brock, Juan C. Rubio, Freeman L. Rawson III, John B. Carter
HPCA7
2010 Power-performance management on an IBM POWER7 server
abstract
The processor and cooling subsystems of high-performance servers consume a significant portion of total system power. In this paper, we use the server energy-efficiency benchmark SPECpower_ssj2008 to assess dynamic power management strategies for these sub-systems on an IBM POWER 750 platform.
Karthick Rajamani, Freeman L. Rawson III, Malcolm Allen-Ware, Heather Hanson, John B. Carter, Todd Rosedahl, Andrew J. Geissler, Guillermo J. Silva, Hong Hua
ISLPED5
2009 Dynamic hardware-assisted software-controlled page placement to manage capacity allocation and sharing within large caches
abstract
In future multi-cores, large amounts of delay and power will be spent accessing data in large L2/L3 caches. It has been recently shown that OS-based page coloring allows a non-uniform cache architecture (NUCA) to provide low latencies and not be hindered by complex data search mechanisms. In this work, we extend that concept with mechanisms that dynamically move data within caches. The key innovation is the use of a shadow address space to allow hardware control of data placement in the L2 cache while being largely transparent to the user application and off-chip world. These mechanisms allow the hardware and OS to dynamically manage cache capacity per thread as well as optimize placement of data shared by multiple threads. We show an average IPC improvement of 10-20% for multi-programmed workloads with capacity allocation policies and an average IPC improvement of 8% for multi-threaded workloads with policies for shared page placement.
Manu Awasthi, Kshitij Sudan, Rajeev Balasubramonian, John B. Carter
HPCA4
2009 A look inside IBM's green data center research
abstract
With data centers using 10--30 times more energy per square foot than office space, data center energy use doubling every 5 years, and delayed capital investments in new power plants—energy efficiency is becoming a key design constraint for modern servers and data centers. This talk will discuss a wide variety of energy management research being performed within IBM Research, including topics such as low-power process technology, novel processor energy management features, memory power management, system-level power shifting, power-aware task placement, power-aware virtualization/consolidation, partition-level power management, energy-efficient high-performance storage, energy-efficient cooling and electrical distribution/backup technologies, and datacenter modeling and optimization. The talk will conclude with a discussion of the key challenges that remain and some important open research topics.
John B. Carter
ISLPED1
2008 Extending CC-NUMA systems to support write update optimizations
abstract
Processor stalls and protocol messages caused by coherence misses limit the performance of shared memory applications. Modern multiprocessors employ write-invalidate coherence protocols, which induce read misses to ensure consistency. Previous research has shown that an invalidate protocol is not optimal for all memory access patterns - an update protocol can significantly outperform an invalidate protocol when data is heavily shared or accessed in predictable patterns. However, update protocols can generate excessive network traffic and are difficult to build on a scalable (non-bus) interconnect. To obtain the benefits of both invalidate and update protocols, we built a speculative sequentially consistent write- update mechanism on top of a write-invalidate protocol. To ensure coherence, a processor wishing to write to a block of data uses a traditional write-invalidate protocol to obtain exclusive access to the block before modifying it. To improve performance, the writing processor can later self- downgrade the modified block to the shared state and flush it back to its home node, which forwards the new data to processors that it predicts are likely to consume the data. We present a practical and cost-effective design for extending CC-NUMA systems to support this speculative update mechanism that requires no changes to the processor core, bus interface, or memory consistency model. We also present two hardware-efficient mechanisms for detecting access patterns that benefit from the speculative update mechanism, stable reader set and stream. We evaluate our update mechanisms on a wide range of scientific benchmarks and commercial applications. Using a cycle-accurate execution-driven simulator of a future 16-node SGI multiprocessor, we find that the mechanisms proposed in this paper reduce the average remote miss rate by 30%, reduce network traffic by 15%, and improve performance by 10%, and in no case hurt performance.
Liqun Cheng, John B. Carter
SC2
2007 An Adaptive Cache Coherence Protocol Optimized for Producer-Consumer Sharing
abstract
Shared memory multiprocessors play an increasingly important role in enterprise and scientific computing facilities. Remote misses limit the performance of shared memory applications, and their significance is growing as network latency increases relative to processor speeds. This paper proposes two mechanisms that improve shared memory performance by eliminating remote misses and/or reducing the amount of communication required to maintain coherence. We focus on improving the performance of applications that exhibit producer-consumer sharing. We first present a simple hardware mechanism for detecting producer-consumer sharing. We then describe a directory delegation mechanism whereby the "home node" of a cache line can be delegated to a producer node, thereby converting 3-hop coherence operations into 2-hop operations. We then extend the delegation mechanism to support speculative updates for data accessed in a producer-consumer pattern, which can convert 2-hop misses into local misses, thereby eliminating the remote memory latency. Both mechanisms can be implemented without changes to the processor. We evaluate our directory delegation and speculative update mechanisms on seven benchmark programs that exhibit producer-consumer sharing using a cycle-accurate execution-driven simulator of a future 16-node SGI multiprocessor. We find that the mechanisms proposed in this paper reduce the average remote miss rate by 40%, reduce network traffic by 15%, and improve performance by 21%. Finally, we use Murphi to verify that each mechanism is error-free and does not violate sequential consistency
Liqun Cheng, John B. Carter, Donglai Dai
HPCA2
2007 Active memory operations
abstract
The performance of modern microprocessors is increasingly limited by their inability to hide main memory latency. The problem is worse in large-scale shared memory systems, where remote memory latencies are hundreds, and soon thousands, of processor cycles. To mitigate this problem, we propose the use of Active Memory Operations (AMOs), in which select operations can be sent to and executed on the home memory controller of data. AMOs can eliminate significant number of coherence messages, minimize intranode and internode memory traffic, and create opportunities for parallelism. Our implementation of AMOs is cache-coherent and requires no changes to the processor core or DRAM chips.
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Ali Ibrahim, Michael A. Parker
ICS3
2006 Program phase detection and exploitation
abstract
Studies of application behavior reveal the nested repetition of large and small program phases, with significant variation among phases in such characteristics as memory reference patterns, memory and energy usage, I/O activity, and occupancy of micro-architectural resources. In this project, we study theories and techniques for reliably predicting and exploiting phased behavior, so an advanced execution environment may allocate resources in a way that better matches program needs, or to transform programs so that their needs better match the available resources. In this paper, we present the basic components of the study and report the progress in the past half year
Chen Ding 0001, Sandhya Dwarkadas, Michael C. Huang 0001, John B. Carter
IPDPS5
2006 Interconnect-Aware Coherence Protocols for Chip Multiprocessors
abstract
Improvements in semiconductor technology have made it possible to include multiple processor cores on a single die. Chip Multi-Processors (CMP) are an attractive choice for future billion transistor architectures due to their low design complexity, high clock frequency, and high throughput. In a typical CMP architecture, the L2 cache is shared by multiple cores and data coherence is maintained among private L1s. Coherence operations entail frequent communication over global on-chip wires. In future technologies, communication between different L1s will have a significant impact on overall processor performance and power consumption. On-chip wires can be designed to have different latency, bandwidth, and energy properties. Likewise, coherence protocol messages have different latency and bandwidth needs. We propose an interconnect composed of wires with varying latency, bandwidth, and energy characteristics, and advocate intelligently mapping coherence operations to the appropriate wires. In this paper, we present a comprehensive list of techniques that allow coherence protocols to exploit a heterogeneous interconnect and evaluate a subset of these techniques to show their performance and power-efficiency potential. Most of the proposed techniques can be implemented with a minimum complexity overhead.
Liqun Cheng, Naveen Muralimanohar, Karthik Ramani, Rajeev Balasubramonian, John B. Carter
ISCA5
2006 Efficient address remapping in distributed shared-memory systems
abstract
As processor performance continues to improve at a rate much higher than DRAM and network performance, we are approaching a time when large-scale distributed shared memory systems will have remote memory latencies measured in tens of thousands of processor cycles. The Impulse memory system architecture adds an optional level of address indirection at the memory controller. Applications can use this level of indirection to control how data is accessed and cached and thereby improve cache and bus utilization and reduce the number of memory accesses required. Previous Impulse work focuses on uniprocessor systems and relies on software to flush processor caches when necessary to ensure data coherence. In this paper, we investigate an extension of Impulse to multiprocessor systems that extends the coherence protocol to maintain data coherence without requiring software-directed cache flushing. Specifically, the multiprocessor Impulse controller can gather/scatter data across the network while its coherence protocol guarantees that each gather request gets coherent data and each scatter request updates every coherent replica in the system. Our simulation results demonstrate that the proposed system can significantly outperform conventional systems, achieving an average speedup of 9X on four memory-bound benchmarks on a 32-processor system.
Lixin Zhang 0002, Michael A. Parker, John B. Carter
ACM Trans. Archit. Code Optim.3
2005 Flexible Consistency for Wide Area Peer Replication
abstract
The lack of a flexible consistency management solution hinders P2P implementation of applications involving updates, such as read-write file sharing, directory services, online auctions and wide area collaboration. Managing mutable shared data in a P2P setting requires a consistency solution that can operate efficiently over variable-quality failure-prone networks, support pervasive replication for scaling, and give peers autonomy to tune consistency to their sharing needs and resource constraints. Existing solutions lack one or more of these features. In this paper, we describe a new consistency model for P2P sharing of mutable data called composable consistency, and outline its implementation in a wide area middleware file service called Swarm¹. Composable consistency lets applications compose consistency semantics appropriate for their sharing needs by combining a small set of primitive options. Swarm implements these options efficiently to support scalable, pervasive, failure-resilient, wide-area replication behind a simple yet flexible interface. We present two applications to demonstrate the expressive power and effectiveness of composable consistency: a wide area file system that outperforms Coda in providing close-to-open consistency overWANs, and a replicated BerkeleyDB database that reaps order-of-magnitude performance gains by relaxing consistency for queries and updates.
Sai Susarla, John B. Carter
ICDCS2
2005 Fast Barriers for Scalable ccNUMA Systems
abstract
The contributions of this paper are threefold. First, we identify and quantify the performance deficiencies of conventional barrier implementations when they are executed on real (non-idealized) hardware. Second, we propose a queue-based barrier algorithm that has effectively O(1) time complexity as measured in round trip message latencies. Third, we demonstrate how matching the barrier implementation to the way that modern shared memory systems operate can improve performance dramatically by exploiting a hardware write-update (PUT) mechanism for signaling. The resulting barrier algorithm only costs one serialized round trip message latency to perform a barrier operation across N processors. Using a cycle-accurate execution-driven simulator of a future-generation SGI multiprocessor, we show that with no special hardware support our queue-based barrier outperforms OpenMP's LL/SC-based barrier implementation by a factor of 7.9 on 256 processors. With hardware that supports a coherent PUT operation, our queue-based barrier outperforms OpenMP barriers by a factor of 94 and outperforms barriers based on SGI's memory controller-based atomic operations by a factor of 6.5 on 256 processors.
Liqun Cheng, John B. Carter
ICPP2
2005 Fast synchronization on shared-memory multiprocessors: An architectural approach
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Liqun Cheng, Michael A. Parker
J. Parallel Distributed Comput.3
2004 Highly Efficient Synchronization Based on Active Memory Operations
abstract
Summary form only given. Synchronization is a crucial operation in many parallel applications. As network latency approaches thousands of processor cycles for large scale multiprocessors, conventional synchronization techniques are failing to keep up with the increasing demand for scalable and efficient synchronization operations. We present a mechanism that allows atomic synchronization operations to be executed on the home memory controller of the synchronization variable. By performing atomic operations near where the data resides, our proposed mechanism can significantly reduce the number of network messages required by synchronization operations. Our proposed design also enhances performance by using fine-grained updates to selectively "push " the results of offloaded synchronization operations back to processors when they complete (e.g., when a barrier count reaches the desired value). We use the proposed mechanism to optimize two of the most widely used synchronization operations, barriers and spin locks. Our simulation results show that the proposed mechanism outperforms conventional implementations based on load-linked/store-conditional, processor-centric atomic instructions, conventional memory-side atomic instructions, or active messages. It speeds up conventional barriers by up to 2.1 (4 processors) to 61.9 (256 processors) and spin locks by a factor of up to 2.0 (4 processors) to 10.4 (256 processors).
Lixin Zhang 0002, Zhen Fang 0002, John B. Carter
IPDPS3
2002 Computation regrouping: restructuring programs for temporal data cache locality
abstract
Data access costs contribute significantly to the execution time of applications with complex data structures. As the latency of memory accesses becomes high relative to processor cycle times, application performance is increasingly limited by memory performance. In some situations it may be reasonable to trade increased computation costs for reduced memory costs. The contributions of this paper are three-fold: we provide a detailed analysis of the memory performance of a set of seven, memory-intensive benchmarks; we describe Computation Regrouping, a general, source-level approach to improving the overall performance of these applications by improving temporal locality to reduce cache and TLB miss ratios (and thus memory stall times); and we demonstrate significant performance improvements from applying Computation Regrouping to our suite of seven benchmarks. With Computation Regrouping, we observe an average speedup of 1.97, with individual speedups ranging from 1.26 to 3.03. Most of this improvement comes from eliminating memory stall time.
Venkata K. Pingali, Sally A. McKee, Wilson C. Hsieh, John B. Carter
ICS4
2001 Reevaluating Online Superpage Promotion with Hardware Support
abstract
Typical translation lookaside buffers (TLBs) can map a far smaller region of memory than application footprints demand, and the cost of handling TLB misses therefore limits the performance of an increasing number of applications. This bottleneck can be mitigated by the use of superpages, multiple adjacent virtual memory pages that can be mapped with a single TLB entry that extend TLB reach without significantly increasing size or cost. We analyze hardware/software tradeoff for dynamically creating superpages. This study extends previous work by using execution-driven simulation to compare creating superpages via copying with remapping pages within the memory controller and by examining how the tradeoffs change when moving front a single-issue to a superscalar processor model. We find that remapping-based promotion outperforms copying-based promotion, often significantly. Copying-based promotion is slightly more effective on superscalar processors than on single-issue processors, and the relative performance of remapping-based promotion on the two platform is application-dependent.
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Wilson C. Hsieh, Sally A. McKee
HPCA3
2001 The Impulse Memory Controller
abstract
Impulse is a memory system architecture that adds an optional level of address indirection at the memory controller. Applications can use this level of indirection to remap their data structures in memory. As a result, they can control how their data is accessed and cached, which can improve cache and bus utilization. The Impulse design does not require any modification to processor, cache, or bus designs since all the functionality resides at the memory controller. As a result, Impulse can be adopted in conventional systems without major system changes. We describe the design of the Impulse architecture and how an Impulse memory system can be used in a variety of ways to improve the performance of memory-bound applications. Impulse can be used to dynamically create superpages cheaply, to dynamically recolor physical pages, to perform strided fetches, and to perform gathers and scatters through indirection vectors. Our performance results demonstrate the effectiveness of these optimizations in a variety of scenarios. Using Impulse can speed up a range of applications from 20 percent to over a factor of 5. Alternatively, Impulse can be used by the OS for dynamic superpage creation; the best policy for creating superpages using Impulse outperforms previously known superpage creation policies.
Lixin Zhang 0002, Zhen Fang 0002, Michael A. Parker, Binu K. Mathew, Lambert Schaelicke, John B. Carter, Wilson C. Hsieh, Sally A. McKee
IEEE Trans. Computers6
2000 Design of a Parallel Vector Access Unit for SDRAM Memory Systems
abstract
We are attacking the memory bottleneck by building a "smart" memory controller that improves effective memory bandwidth, bus utilization, and cache efficiency by letting applications dictate how their data is accessed and cached. This paper describes a parallel vector access unit (PVA), the vector memory subsystem that efficiently "gathers" sparse, strided data structures in parallel on a multi-bank SDRAM memory. We have validated our PVA design via gate-level simulation, and have evaluated its performance via functional simulation and formal analysis. On unit-stride vectors, PVA performance equals or exceeds that of an SDRAM system optimized for cache line fills. On vectors with larger strides, the PVA is up to 32.8 times faster. Our design is up to 3.3 times faster than a pipelined, serial SDRAM memory system that gathers sparse vector data, and the gathering mechanism is two to five times faster than in other PVAs with similar goals. Our PVA only slightly increases hardware complexity with respect to these other systems, and the scalable design is appropriate for a range of computing platforms, from vector supercomputers to commodity PCs.
Binu K. Mathew, Sally A. McKee, John B. Carter, Al Davis
HPCA3
2000 Online superpage promotion revisited (poster)
abstract
No abstract available.
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Sally A. McKee, Wilson C. Hsieh
SIGMETRICS3
2000 Algorithmic foundations for a parallel vector access memory system
abstract
This paper presents mathematical foundations for the design of a memory controller subcomponent that helps to bridge the processor/memory performance gap for applications with strided access patterns. The Parallel Vector Access (PVA) unit exploits the regularity of vectors or streams to access them efficiently in parallel on a multi-bank SDRAM memory system. The PVA unit performs scatter/gather operations so that only the elements accessed by the application are transmitted across the system bus. Vector operations are broadcast in parallel to all memory banks, each of which implements an efficient algorithm to determine which vector elements it holds. Earlier performance evaluations have demonstrated that our PVA implementation loads elements up to 32.8 times faster than a conventional memory system and 3.3 times faster than a pipelined vector unit, without hurting the performance of normal cache-line fills. Here we present the underlying PVA algorithms for both word interleaved and cache-line inter-leaved memory systems.
Binu K. Mathew, Sally A. McKee, John B. Carter, Al Davis
SPAA3
1999 Impulse: Building a Smarter Memory Controller
abstract
Impulse is a new memory system architecture that adds two important features to a traditional memory controller. First, Impulse supports application-specific optimizations through configurable physical address remapping. By remapping physical addresses, applications control how their data is accessed and cached, improving their cache and bus utilization. Second, Impulse supports prefetching at the memory controller, which can hide much of the latency of DRAM accesses. In this paper we describe the design of the Impulse architecture, and show how an Impulse memory system can be used to improve the performance of memory-bound programs. For the NAS conjugate gradient benchmark, Impulse improves performance by 67%. Because it requires no modification to processor, cache, or bus designs, Impulse can be adopted in conventional systems. In addition to scientific applications, we expect that Impulse will benefit regularly strided memory-bound applications of commercial importance, such as database and multimedia programs.
John B. Carter, Wilson C. Hsieh, Leigh Stoller, Mark R. Swanson, Lixin Zhang 0002, Erik Brunvand, Al Davis, Chen-Chi Kuo, Ravindra Kuramkote, Michael A. Parker, Lambert Schaelicke, Terry Tateyama
HPCA1
1999 MP-LOCKs: Replacing H/W Synchronization Primitives with Message Passing
abstract
Shared memory programs guarantee the correctness of concurrent accesses to shared data using interprocessor synchronization operations. The most common synchronization operators are locks, which are traditionally implemented via a mix of shared memory accesses and hardware synchronization primitives like test-and-set. In this paper, we argue that synchronization operations implemented using fast message passing and kernel-embedded lock managers are an attractive alternative to dedicated synchronization hardware. We propose three message passing lock (MP-LOCK) algorithms (centralized, distributed, and reactive) and provide implementation guidelines. MP-LOCKs reduce the design complexity and runtime occupancy of DSM controllers and can exploit software's inherent flexibility to adapt to differing applications lock access patterns. We compared the performance of MP-LOCKs with two common shared memory lock algorithms: test-and-test-and-set and MCS locks and found that MP-LOCKs scale better. For machines with 16 to 32 nodes, applications using MP-LOCKs ran up to 186% faster than the same applications with shared memory locks. For small systems (up to 8 nodes), three applications with MP-LOCKs slow down by no more than 18%, while the other two slowed by no more than 180% due to higher software overhead. We conclude that locks based on message passing should be considered as a replacement for hardware locks in future scalable multiprocessors that support efficient message passing mechanisms.
Chen-Chi Kuo, John B. Carter, Ravindra Kuramkote
HPCA2
1998 Design alternatives for shared memory multiprocessors
abstract
We consider the design alternatives available for building the next generation DSM machine (e.g., the choice of memory architecture, network technology, and amount and location of per-node remote data cache). To investigate this design space, we have simulated five applications on a wide variety of possible DSM architectures that employ significantly different caching techniques. We also examine the impact of using a special purpose system interconnect designed specifically to support low latency DSM operation versus using a powerful off the shelf system interconnect. We found that two architectures have the best combination of good average performance and reasonable worst case performance: CC-NUMA employing a moderate sized DRAM remote access cache (RAC) and a hybrid CC-NUMA/S-COMA architecture called AS-COMA or adaptive S-COMA. Both pure CC-NUMA and pure S-COMA have serious performance problems for some applications, while CC-NUMA employing an SRAM RAC does not perform as well as the two architectures that employ larger DRAM caches. The paper concludes with several recommendations to designers of next generation DSM machines, complete with a discussion of the issues that led to each recommendation so that designers can decide which ones are relevant to them given changes in technology and corporate priorities.
John B. Carter, Chen-Chi Kuo, Ravindra Kuramkote, Mark R. Swanson
HiPC1
1998 Making Distributed Shared Memory Simple, Yet Efficient
abstract
Recent research on distributed shared memory (DSM) has focussed on improving performance by reducing the communication overhead of DSM. Features added include lazy release consistency based coherence protocols and new interfaces that give programmers the ability to hand tune communication. These features have increased DSM performance at the expense of requiring increasingly complex DSM systems or increasingly cumbersome programming. They have also increased the computation overhead of DSM, which has partially offset the communication related performance gains. We chose to implement a simple DSM system, Quarks, with an eye towards hiding most computation overhead while using a very low latency transport layer to reduce the effect of communication overhead. The resulting performance is comparable to that of far more complex DSM systems, such as Treadmarks and Cashmere.
Mark R. Swanson, Leigh Stoller, John B. Carter
HIPS3
1998 Khazana: An Infrastructure for Building Distributed Services
abstract
Essentially all distributed systems, applications, and services at some level boil down to the problem of managing distributed shared state. Unfortunately, while the problem of managing distributed shared state is shared by many applications, there is no common means of managing the data-every application devises its own solution. We have developed Khazana, a distributed service exporting the abstraction of a distributed persistent globally shared store that applications can use to store their shared state. Khazana is responsible for performing many of the common operations needed by distributed applications, including replication, consistency management, fault recovery, access control and location management. Using Khazana as a form of middleware, distributed applications can be quickly developed from corresponding uniprocessor applications through the insertion of Khazana data access and synchronization operations.
John B. Carter, Anand Ranganathan, Sai Susarla
ICDCS1
1998 ASCOMA: An Adaptive Hybrid Shared Memory Architecture
abstract
Scalable shared memory multiprocessors traditionally use either a cache coherent non-uniform memory access (CC-NUMA) or simple cache-only memory architecture (S-COMA) memory architecture. Recently, hybrid architectures that combine aspects of both CC-NUMA and S-COMA have emerged. We present two improvements over other hybrid architectures. The first improvement is a page allocation algorithm that prefers S-COMA pages at low memory pressures. Once the local free page pool is drained, additional pages are mapped in CC-NUMA mode until they suffer sufficient remote misses to warrant upgrading to S-COMA mode. The second improvement is a page replacement algorithm that dynamically backs off the rate of page remappings from CC-NUMA to S-COMA mode at high memory pressure. This design dramatically reduces the amount of kernel overhead and the number of induced cold misses caused by needless thrashing of the page cache. The resulting hybrid architecture is called adaptive S-COMA (AS-COMA). AS-COMA exploits the best of S-COMA and CC-NUMA, performing like an S-COMA machine at low memory pressure and like a CC-NUMA machine at high memory pressure. AS-COMA outperforms CC-NUMA under almost all conditions, and outperforms other hybrid architectures by up to 17% at low memory pressure and up to 90% at high memory pressure.
Chen-Chi Kuo, John B. Carter, Ravindra Kuramkote, Mark R. Swanson
ICPP2
1998 Increasing TLB Reach Using Superpages Backed by Shadow Memory
abstract
The amount of memory that can be accessed without causing a translation lookaside buffer (TLB) fault, the reach of a TLB, is failing to keep pace with the increasingly large working sets of applications. We propose to extend TLB reach via a novel Memory Controller TLB (MTLB) that lets us aggressively create superpages from non-contiguous, unaligned regions of physical memory. This flexibility increases the OS's ability to use superpages on arbitrary application data. The MTLB supports shadow pages, regions of physical address space for which the MTLB remaps accesses to "real" physical pages. The MTLB preserves per-base-page referenced and dirty bits, which enables the OS to swap shadow-backed superpages a page at a time, unlike conventional superpages. Simulation of five applications, including two SPECint95 benchmarks, demonstrated that a modest-sized MTLB improves performance of applications with moderate-to-high TLB miss rates by 5-20%. Simulation also showed that this mechanism can more than double the effective reach of a processor TLB with no modification to the processor MMU.
Mark R. Swanson, Leigh Stoller, John B. Carter
ISCA3
1995 Distributed shared memory: where we are and where we should be headed
abstract
It has been almost ten years since the birth of the first distributed shared memory (DSM) system, Ivy. While significant progress has been made in the area of improving the performance of DSM, and DSM has been the focus of several dozen PhD theses, its overall impact on "real" users and applications has been small. The goal of this paper is to present our position on what remains to be done before DSM will have a significant impact on real applications. More specifically, we reflect on what we believe have been the major advances in the area, what the important outstanding problems are, and what work needs to be done. Finally, we describe a modest step towards solving these problems, the Quarks DSM system.
John B. Carter, Dilip Khandekar, Linus Kamb
HotOS1
1995 An Argument for Simple COMA
abstract
We present design details and some initial performance results of a novel scalable shared memory multiprocessor architecture. This architecture features the automatic data migration and replication capabilities of cache-only memory architecture (COMA) machines, without the accompanying hardware complexity. A software layer manages cache space allocation at a page-granularity-similarly to distributed virtual shared memory (DVSM) systems, leaving simpler hardware to maintain shared memory coherence at a cache line granularity. By reducing the hardware complexity, the machine cost and development time are reduced. We call the resulting hybrid hardware and software multiprocessor architecture Simple COMA. Preliminary results indicate that the performance of Simple COMA is comparable to that of more complex contemporary all hardware designs.>
Ashley Saulsbury, Tim Wilkinson, John B. Carter, Anders Landin
HPCA3
1995 An argument for simple COMA
Ashley Saulsbury, Tim Wilkinson, John B. Carter, Anders Landin
Future Gener. Comput. Syst.3
1995 Design of the Munin Distributed Shared Memory System
John B. Carter
J. Parallel Distributed Comput.1
1995 Techniques for Reducing Consistency-Related Communication in Distributed Shared-Memory Systems
abstract
Distributed shared memory (DSM) is an abstraction of shared memory on a distributed-memory machine. Hardware DSM systems support this abstraction at the architecture level; software DSM systems support the abstraction within the runtime system. One of the key problems in building an efficient software DSM system is to reduce the amount of communication needed to keep the distributed memories consistent. In this article we present four techniques for doing so: software release consistency; multiple consistency protocols; write-shared protocols; and an update-with-timeout mechanism. These techniques have been implemented in the Munin DSM system. We compare the performance of seven Munin application programs: first to their performance when implemented using message passing, and then to their performance when running on a conventional software DSM system that does not embody the preceding techniques. On a 16-processor cluster of workstations, Munin's performance is within 5% of message passing for four out of the seven applications. For the other three, performance is within 29 to 33%. Detailed analysis of two of these three applications indicates that the addition of a function-shipping capability would bring their performance to within 7% of the message-passing performance. Compared to a conventional DSM system, Munin achieves performance improvements ranging from a few to several hundred percent, depending on the application.
John B. Carter, John K. Bennett, Willy Zwaenepoel
ACM Trans. Comput. Syst.1
1991 Implementation and Performance of Munin
abstract
Munin is a distributed shared memory (DSM) system that allows shared memory parallel programs to be executed efficiently on distributed memory multiprocessors. Munin is unique among existing DSM systems in its use of multiple consistency protocols and in its use of release consistency. In Munin, shared program variables are annotated with their expected access pattern, and these annotations are then used by the runtime system to choose a consistency protocol best suited to that access pattern. Release consistency allows Munin to mask network latency and reduce the number of messages required to keep memory consistent. Munin's multiprotocol release consistency is implemented in software using a delayed update queue that buffers and merges pending outgoing writes. A sixteen-processor prototype of Munin is currently operational. We evaluate its implementation and describe the execution of two Munin programs that achieve performance within ten percent of message passing implementations of the same programs. Munin achieves this level of performance with only minor annotations to the shared memory programs.
John B. Carter, John K. Bennett, Willy Zwaenepoel
SOSP1
1990 Adaptive Software Cache Management for Distributed Shared Memory Architectures
abstract
An adaptive cache coherence mechanism exploits semantic information about the expected or observed access behavior of particular data objects. We contend that, in distributed shared memory systems, adaptive cache coherence mechanisms will outperform static cache coherence mechanisms. We have examined the sharing and synchronization behavior of a variety of shared memory parallel programs. We have found that the access patterns of a large percentage of shared data objects fall in a small number of categories for which efficient software coherence mechanisms exist. In addition, we have performed a simulation study that provides two examples of how an adaptive caching mechanism can take advantage of semantic information.
John K. Bennett, John B. Carter, Willy Zwaenepoel
ISCA2
1990 Munin: Distributed Shared Memory Based on Type-Specific Memory Coherence
abstract
We are developing Munin, a system that allows programs written for shared memory multiprocessors to be executed efficiently on distributed memory machines. Munin attempts to overcome the architectural limitations of shared memory machines, while maintaining their advantages in terms of ease of programming. Our system is unique in its use of loosely coherent memory, based on the partial order specified by a shared memory parallel program, and in its use of type-specific memory coherence. Instead of a single memory coherence mechanism for all shared data objects, Munin employs several different mechanisms, each appropriate for a different class of shared data object. These type-specific mechanisms are part of a runtime system that accepts hints from the user or the compiler to determine the coherence mechanism to be used for each object. This paper focuses on the design and use of Munin's memory coherence mechanisms, and compares our approach to previous work in this area.
John K. Bennett, John B. Carter, Willy Zwaenepoel
PPoPP2
1989 Optimistic Implementation of Bulk Data Transfer Protocols
abstract
During a bulk data transfer over a high speed network, there is a high probability that the next packet received from the network by the destination host is the next packet in the transfer. An optimistic implementation of a bulk data transfer protocol takes advantage of this observation by instructing the network interface on the destination host to deposit the data of the next packet immediately into its anticipated final location. No copying of the data is required in the common case, and overhead is greatly reduced.
John B. Carter, Willy Zwaenepoel
SIGMETRICS1