Karthick Rajamani

dblp:71/1431 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorComputer networks · 2Software engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Memory systems · 37% Energy-efficient computing · 29% Cloud and datacenter computing · 22%
Software engineering, system software, and programming languages
2 papers
Operating systems · 70% Runtime systems and virtual machines · 30%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
power management
0.632019
A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019
Accurate Fine-Grained Processor Power Proxies · MICRO 2012
Architecting for power management: The IBM POWER7TM approach · HPCA 2010
Energy-efficient computing
datacenter power management
0.522019
A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019
Smarter data center power monitoring and management · SenSys 2011
Memory systems
non-volatile memory
0.412020
Temperature Aware Adaptations for Improved Read Reliability in STT-MRAM Memory Subsystem · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Hardware reliability and fault tolerance › memory reliability
read reliability
0.412020
Temperature Aware Adaptations for Improved Read Reliability in STT-MRAM Memory Subsystem · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM
0.412020
Temperature Aware Adaptations for Improved Read Reliability in STT-MRAM Memory Subsystem · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Energy-efficient computing › datacenter power management
server power capping
0.412019
A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019
Memory systems › cache management
cache isolation
0.312018
dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018
Memory systems
cache management
0.312018
dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018
Memory systems › cache management
cache partitioning
0.312018
dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018
Cloud and datacenter computing › virtualization
container
0.312018
Iron: Isolating Network-based CPU in Container Environments · NSDI 2018
Cloud and datacenter computing › workload isolation
container isolation
0.312018
Iron: Isolating Network-based CPU in Container Environments · NSDI 2018
Cloud and datacenter computing
performance isolation
0.312018
dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018
Cloud and datacenter computing
virtualization
0.312018
dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018
Memory systems
DRAM
0.322020
Tiered Memory: An Iso-Power Memory Architecture to Address the Memory Power Wall · IEEE Trans. Computers 2012
Temperature Aware Adaptations for Improved Read Reliability in STT-MRAM Memory Subsystem · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Energy-efficient computing
power-performance tradeoff
0.112012
Does lean imply green?: a study of the power performance implications of Java runtime bloat · SIGMETRICS 2012
Memory systems
tiered memory
0.112012
Tiered Memory: An Iso-Power Memory Architecture to Address the Memory Power Wall · IEEE Trans. Computers 2012
Performance modeling and evaluation
workload characterization
0.112012
Does lean imply green?: a study of the power performance implications of Java runtime bloat · SIGMETRICS 2012
Energy-efficient computing › energy measurement
power monitoring
0.112011
Smarter data center power monitoring and management · SenSys 2011
Distributed systems › fault tolerance
high availability
0.112019
A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019
Energy-efficient computing › power management
dynamic power management
0.112010
Architecting for power management: The IBM POWER7TM approach · HPCA 2010
Operating systems › system security › operating system security › protection mechanism › isolation
resource isolation
0.112018
Iron: Isolating Network-based CPU in Container Environments · NSDI 2018
Processor architecture and microarchitecture
multicore design
0.112018
dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018
Memory systems › cache
shared last-level cache
0.112018
dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018
Runtime systems and virtual machines
garbage collection
0.012012
Does lean imply green?: a study of the power performance implications of Java runtime bloat · SIGMETRICS 2012
Cloud and datacenter computing › resource management
cloud resource management
0.012012
Accurate Fine-Grained Processor Power Proxies · MICRO 2012
Energy-efficient computing › power management
memory power management
0.012012
Tiered Memory: An Iso-Power Memory Architecture to Address the Memory Power Wall · IEEE Trans. Computers 2012
Cloud and datacenter computing › virtualization › virtual machine management
server consolidation
0.012012
Tiered Memory: An Iso-Power Memory Architecture to Address the Memory Power Wall · IEEE Trans. Computers 2012
Memory systems › shared memory
distributed shared memory
0.021999
Adaptive protocols for software distributed shared memory · Proc. IEEE 1999
Trade-offs Between False Sharing and Aggregation in Software Distributed Shared Memory · PPoPP 1997
Cloud and datacenter computing › virtualization
virtual machine
0.012010
Architecting for power management: The IBM POWER7TM approach · HPCA 2010
Memory systems › cache coherence
false sharing
0.011997
Trade-offs Between False Sharing and Aggregation in Software Distributed Shared Memory · PPoPP 1997

Methods — techniques the papers use, named apart from their topics

read disturb analysis · 0.4bit error rate estimation · 0.4priority-aware scheduling · 0.4power capping · 0.4dynamic cache management · 0.3Intel CAT · 0.3equi-performance power reduction · 0.3controlled experiment · 0.3power model training · 0.1on-chip power sensing · 0.1
YearPublicationVenuePosition
2023 REFORM: Increase alerts value using data driven approach
abstract
We introduce REFORM as a data-driven approach to assess alert quality and increase the value of alerts, reduce alert noise and to identify gaps in alert coverage. In the context of monitoring cloud infrastructure and services, surfacing the right set of alerts to human operators is critical to ensure timely intervention for issues that may otherwise result in significant impact to customers as well as for avoiding operator burnout.To assess alert quality we focus on the notion of actionable alerts – a specific alert is categorized as actionable or not based on historic evidence of operator actions having been taken in response to such alerts. Using alert data collected over a six month period from a combination of cloud environments that include pre-production and production systems, alerts are categorized as true positive, false positive or false negative with respect to their actionability. Based on this, we then introduce and quantify precision and miss rates per alert category (trigger condition) from the historic data. Noisy, non-actionable alerts are identified using precision. Gaps in coverage are identified using miss rate, or failure of monitoring to detect issues.A prescriptive approach is then provided to increase the alerts’ value that includes refinement of alert definitions, replacement of alerts that are no longer actionable and identification of alerting coverage gaps. We monitored our REFORM system over three months and found that the volume of noisy, non-actionable alerts were significantly reduced and furthermore, this resulted in significantly higher intervention rates from the operators.
Anupama Jagannathan, Christopher M. Dye, Karthick Rajamani, Chris Galtenberg, Benjamin Luong, Egan Ford
IC2E3
2022 Architecture slack exploitation for phase classification and performance estimation in server-class processors
Diyanesh Chinnakkonda, Karthick Rajamani, M. B. Srinivas
J. Parallel Distributed Comput.2
2020 Temperature Aware Adaptations for Improved Read Reliability in STT-MRAM Memory Subsystem
abstract
Spin-transfer torque magneto-resistive random-access memory (STT-MRAM) is an exciting new emerging technology, being considered as a strong candidate to fill the gaps in the existing memory hierarchy between DRAM and the secondary memory. STT-MRAM has adequate endurance. However, unresolved write switching and read reliability issues still exist at the functional operating temperature corners. One biggest challenge is that the read bit error rate (RBER) is not at an acceptable level for system reliability across the wide operating temperature range. We present an STT-MRAM memory subsystem that is fully compatible with existing DDR-based DIMM designs and evaluate read disturb and read sense bit-error rate (BER) under various operating temperature conditions. We propose temperature aware adaptive techniques for reliable reads at the rank level. The proposed temperature adaptation technique improves overall reliability of the DDR4 STT-MRAM-based memory subsystem with an optimal read current considering an acceptable 64-byte cacheline BER. Our full system simulations show 1000× order of improvements toward a cell raw read disturb BER along with 5% reduction in memory power and less than 1% impact on overall system performance.
Saravanan Sethuraman, T. Venkata Kalyan, Karthick Rajamani, Chitra K. Subramanian, Kyu-Hyoun Kim, Hillery C. Hunter, M. B. Srinivas
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 A Scalable Priority-Aware Approach to Managing Data Center Server Power
abstract
Power management is a key component of modern data center design. Power managers must (1) ensure the costand energy-efficient utilization of the data center infrastructure, (2) maintain availability of the services provided by the center, and (3) address environmental concerns associated with the center's power consumption. While several power management techniques have been proposed and deployed in production data centers, there are still many challenges to comprehensive data center power management. This is particularly true in public cloud environments, where different jobs have different priority levels, and where high availability is critical. One example of the challenges facing public cloud data centers involves power capping. As power delivery must be highly reliable and tolerate wide variation in the load drawn by the data center components, the power infrastructure (e.g., power supplies, circuit breakers, UPS) has high redundancy and overprovisioning. During normal operation (i.e., typical server power demands, and no failures in the center), the power infrastructure is significantly underutilized. Power capping is a common solution to reduce this underutilization, by allowing more servers to be added safely (i.e., without power shortfalls) to the existing power infrastructure, and throttling power consumption in the infrequent cases where the demanded power exceeds the provisioned power capacity to avoid shortfalls. However, state-of-the-art power capping solutions are (1) not directly applicable to the redundant power infrastructure used in highly-available data centers; and (2) oblivious to differing workload priorities across the entire center when power consumption needs to be throttled, which can unnecessarily slow down high-priority work. To address this need, we develop CapMaestro, a new power management architecture with three key features for public cloud data centers. First, CapMaestro is designed to work with multiple power feeds (i.e., sources), and exploits server-level power capping to independently cap the load on each feed of a server. Second, CapMaestro uses a scalable, global priority-aware power capping approach, which accounts for power capacity at each level of the power distribution hierarchy. It exploits the underutilization of commonly-employed redundant power infrastructure at each level of the hierarchy to safely accommodate a much greater number of servers. Third, CapMaestro exploits stranded power (i.e., power budgets that are not utilized) in redundant power infrastructure to boost the performance of workloads in the data center. We add CapMaestro to a real cloud data center control plane, and demonstrate the effectiveness of all three key features. Using a large-scale data center simulation, we demonstrate that CapMaestro significantly and safely increases the number of servers for existing infrastructure. We also call out other key technical challenges the industry faces in data center power management.
Yang Li 0183, Charles Lefurgy, Karthick Rajamani, Malcolm Allen-Ware, Guillermo J. Silva, Daniel D. Heimsoth, Saugata Ghose, Onur Mutlu
HPCA3
2018 dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service
abstract
In the modern multi-tenant cloud, resource sharing increases utilization but causes performance interference between tenants. More generally, performance isolation is also relevant in any multi-workload scenario involving shared resources. Last level cache (LLC) on processors is shared by all CPU cores in x86, thus the cloud tenants inevitably suffer from the cache flush by their noisy neighbors running on the same socket. Intel Cache Allocation Technology (CAT) provides a mechanism to assign cache ways to cores to enable cache isolation, but its static configuration can result in underutilized cache when a workload cannot benefit from its allocated cache capacity, and/or lead to sub-optimal performance for workloads that do not have enough assigned capacity to fit their working set.
Karthick Rajamani, Wes Felter, Juan C. Rubio, Yang Li 0183
EuroSys2
2018 Iron: Isolating Network-based CPU in Container Environments
Junaid Khalid, Eric Rozner, Wes Felter, Karthick Rajamani, Aditya Akella
NSDI5
2012 Accurate Fine-Grained Processor Power Proxies
abstract
There are not yet practical and accurate ways to directly measure core power in a microprocessor. This limits the granularity of measurement and control for computer power management. We overcome this limitation by presenting an accurate runtime per-core power proxy which closely estimates true core power. This enables new fine-grained microprocessor power management techniques at the core level. For example, cloud environments could manage and bill virtual machines for energy consumption associated with the core. The power model underlying our power proxy also enables energy-efficiency controllers to perform what-if analysis, instead of merely reacting to current conditions. We develop and validate a methodology for accurate power proxy training at both chip and core levels. Our implementation of power proxies uses on-chip logic in a high-performance multi-core processor and associated platform firmware. The power proxies account for full voltage and frequency ranges, as well as chip-to-chip process variations. For fixed clock frequency operation, a mean unsigned error of 1.8% for fine-grained 32ms samples across all workloads was achieved. For an interval of an entire workload, we achieve an average error of-0.2%. Similar results were achieved for voltage-scaling scenarios, too. We also present two sample applications of the power proxy: (1) per-core power billing for cloud computing services, and (2) simultaneous runtime energy saving comparisons among different power management policies without running each policy separately.
Wei Huang 0004, Charles Lefurgy, William Kuk, Alper Buyuktosunoglu, Michael S. Floyd, Karthick Rajamani, Malcolm Allen-Ware, Bishop Brock
MICRO6
2012 Does lean imply green?: a study of the power performance implications of Java runtime bloat
abstract
The presence of software bloat in large flexible software systems can hurt energy efficiency. However, identifying and mitigating bloat is fairly effort intensive. To enable such efforts to be directed where there is a substantial potential for energy savings, we investigate the impact of bloat on power consumption under different situations. We conduct the first systematic experimental study of the joint power-performance implications of bloat across a range of hardware and software configurations on modern server platforms. The study employs controlled experiments to expose different effects of a common type of Java runtime bloat, excess temporary objects, in the context of the SPECPower_ssj2008 workload. We introduce the notion of equi-performance power reduction to characterize the impact, in addition to peak power comparisons. The results show a wide variation in energy savings from bloat reduction across these configurations. Energy efficiency benefits at peak performance tend to be most pronounced when bloat affects a performance bottleneck and non-bloated resources have low energy-proportionality. Equi-performance power savings are highest when bloated resources have a high degree of energy proportionality.
Suparna Bhattacharya, Karthick Rajamani, K. Gopinath
SIGMETRICS2
2012 Tiered Memory: An Iso-Power Memory Architecture to Address the Memory Power Wall
abstract
Moore's Law improvement in transistor density is driving a rapid increase in the number of cores per processor. DRAM device capacity and energy efficiency are increasing at a slower pace, so the importance of DRAM power is increasing. This problem presents system designers with two nominal options when designing future systems: 1) decrease off-chip memory capacity and bandwidth per core or 2) increase the fraction of system power allocated to main memory. Reducing capacity and bandwidth leads to imbalanced systems with poor processor utilization for noncache-resident applications, so designers have chosen to increase DRAM power budget. This choice has been viable to date, but is fast running into a memory power wall. To address the looming memory power wall problem, we propose a novel iso-power tiered memory architecture that supports 2-3X more memory capacity for the same power budget as traditional designs by aggressively exploiting low-power DRAM modes. We employ two "tiers” of DRAM, a "hot” tier with active DRAM and a "cold” tier in which DRAM is placed in self-refresh mode. The DRAM capacity of each tier is adjusted dynamically based on aggregate workload requirements and the most frequently accessed data are migrated to the "hot” tier. This design allows larger memory capacities at a fixed power budget while mitigating the performance impact of using low-power DRAM modes. We target our solution at server consolidation scenarios where physical memory capacity is typically the primary factor limiting the number of virtual machines a server can support. Using iso-power tiered memory, we can run 3× as many virtual machines, achieving a 250 percent improvement in average aggregate performance, compared to a conventional memory design with the same power budget.
Kshitij Sudan, Karthick Rajamani, Wei Huang 0004, John B. Carter
IEEE Trans. Computers2
2011 Smarter data center power monitoring and management
abstract
This demonstration presents a power panel level power monitoring and management (PMM) system developed at IBM Research. The ultimate goal of this project is to develop a low-cost, high accuracy, non-intrusive and retrofittable data center power management system.
Wael El-Essawy, Malcolm Allen-Ware, Karthick Rajamani, Juan C. Rubio, Michael A. Schappert, Tom W. Keller, Hendrik F. Hamann
SenSys4
2010 Adaptive energy management features of the POWER7TM processor
Michael S. Floyd, Bishop Brock, Malcolm Allen-Ware, Karthick Rajamani, Alan J. Drake, Charles Lefurgy, Lorena Pesantez
Hot Chips Symposium4
2010 Architecting for power management: The IBM POWER7TM approach
abstract
The POWER7 processor is the newest member of the IBM POWER®family of server processors. With greater than 4X the peak performance and the same power budget as the previous generation POWER6®, POWER7 will deliver impressive energy-efficiency boosts. The improved peak energy-efficiency is accompanied by a wide array of new features in the processor and system designs that advance IBM's EnergyScaleTMdynamic power management methodology. This paper provides an overview of these new features, which include better sensing, more advanced power controls, improved scalability for power management, and features to address the diverse needs of the full range of POWER servers from blades to supercomputers. We also highlight three challenges that need attention from a range of systems design and research teams: (i) power management in highly virtualized environments, (ii) power (in)efficiency of systems software and applications, and (iii) memory power costs, especially for servers with large memory footprints.
Malcolm Allen-Ware, Karthick Rajamani, Michael S. Floyd, Bishop Brock, Juan C. Rubio, Freeman L. Rawson III, John B. Carter
HPCA2
2010 Power-performance management on an IBM POWER7 server
abstract
The processor and cooling subsystems of high-performance servers consume a significant portion of total system power. In this paper, we use the server energy-efficiency benchmark SPECpower_ssj2008 to assess dynamic power management strategies for these sub-systems on an IBM POWER 750 platform.
Karthick Rajamani, Freeman L. Rawson III, Malcolm Allen-Ware, Heather Hanson, John B. Carter, Todd Rosedahl, Andrew J. Geissler, Guillermo J. Silva, Hong Hua
ISLPED1
2008 Power management solutions for computer systems and datacenters
abstract
The growing power and cooling requirements of high-density computing systems pose significant challenges for the design and operation of computers and their facilities. The rising operating expenses for datacenters demand the implementation of energy-efficient technologies and the best power management solutions. This tutorial addresses power management and cooling solutions from the individual computer system level to the datacenter. The audience will learn about the fundamental nature of the problems, approaches to developing solutions, available commercial solutions, and current research directions.
Karthick Rajamani, Charles Lefurgy, Soraya Ghiasi, Juan C. Rubio, Heather Hanson, Tom W. Keller
ISLPED1
2007 Power, Performance, and Thermal Management for High-Performance Systems
abstract
In future high-performance systems it will be essential to balance often-conflicting objectives of performance, power, energy, and temperature under variable workload and environmental conditions. In this work, we describe a goal-driven approach that conveys multiple expectations to managers that dynamically tune operating states to best meet those demands. We show the benefit of a concise goal specification for complex objectives and the feasibility of managing multiple constraints while maintaining high performance and safe operation. We evaluate key features of our approach with a prototype implementation on a Pentium M platform with Red Hat Enterprise 4 that controls voltage and frequency scaling to achieve the desired performance, power and temperature goals.
Heather Hanson, Stephen W. Keckler, Karthick Rajamani, Soraya Ghiasi, Freeman L. Rawson III, Juan C. Rubio
IPDPS3
2007 Thermal response to DVFS: analysis with an Intel Pentium M
abstract
Increasing power density in computing systems from laptops to servers has spurred interest in dynamic thermal management. Based on the success of dynamic voltage and frequency scaling (DVFS) in managing power and energy, DVFS may be a viable option for thermal management, as well. However, publicly available data on the thermal effects of DVFS are very limited. In this work, we characterize the thermal response of Intel Pentium M system to DVFS, identifying the response timescale and influence of factors beyond voltage and frequency on processor temperature.
Heather Hanson, Stephen W. Keckler, Soraya Ghiasi, Karthick Rajamani, Freeman L. Rawson III, Juan C. Rubio
ISLPED4
2005 A performance-conserving approach for reducing peak power consumption in server systems
abstract
The combination of increasing component power consumption, a desire for denser systems, and the required performance growth in the face of technology-scaling issues are posing enormous challenges for powering and cooling of server systems. The challenges are directly linked to the peak power consumption of servers.Our solution, Power Shifting, reduces the peak power consumption of servers minimizing the impact on performance. We reduce peak power consumption by using workload-guided dynamic allocation of power among components incorporating real-time performance feedback, activity-related power estimation techniques, and performance-sensitive activity-regulation mechanisms to enforce power budgets.We apply our techniques to a computer system with a single processor and memory. Power shifting adds a system power manager with a dynamic, global view of the system's power consumption to continuously re-budget the available power amongst the two components. Our contributions include:• Demonstration of the greater effectiveness of dynamic power allocation over static budgeting,• Evaluation of different power shifting policies,• Analysis of system and workload factors critical to successful power shifting, and• Proposal of performance-sensitive power budget enforcement mechanisms that ensure system reliability.
Wes Felter, Karthick Rajamani, Tom W. Keller, Cosmin Rusu
ICS2
2003 On evaluating request-distribution schemes for saving energy in server clusters
abstract
Power-performance optimization is a relatively new problem area particularly in the context of server clusters. Power-aware request distribution is a method of scheduling service requests among servers in a cluster so that energy consumption is minimized, while maintaining a particular level of performance. Energy efficiency is obtained by powering-down some servers when the desired quality of service can be met with fewer servers. We have found that it is critical to take into account the system and workload factors during both the design and the evaluation of such request distribution schemes. We identify the key system and workload factors that impact such policies and their effectiveness in saving energy. We measure a web cluster running an industry-standard commercial web workload to demonstrate that understanding this system-workload context is critical to performing valid evaluations and even for improving the energy-saving schemes.
Karthick Rajamani, Charles Lefurgy
ISPASS1
1999 Efficient Mining for Association Rules with Relational Database Systems
abstract
With the tremendous growth of large scale data repositories, a need for integrating the exploratory techniques of data mining with the capabilities of relational systems to efficiently handle large volumes of data has now risen. We look at the performance of the most prevalent association rule mining algorithm-Apriori with IBM's DB2 Universal Database system. We show that a multi-column (MC) data model is preferable over the commonly used single column (SC) data model for association rule mining. We obtain factors of 4.8 to 6 improvement in performance for the MC data model over commercial implementations for the SC data model. We provide a new relational operator called Combinations, for efficient SQL implementation of Apriori in the database engine-this results in trivial parallelizability, reliability, and portability for the mining application.
Karthick Rajamani, Alan L. Cox, Balakrishna R. Iyer, Atul Chadha
IDEAS1
1999 Extending the Applicability of Association Rules
Karthick Rajamani, Sam Yuan Sung, Alan L. Cox
PAKDD1
1999 Adaptive protocols for software distributed shared memory
abstract
We demonstrate the benefits of software shared memory protocols that adapt at run time to the memory access patterns observed in the applications. This adaptation is automatic-no user annotations are required-and does not rely on compiler support or special hardware. We investigate adaptation between singleand multiple-writer protocols, dynamic aggregation of pages into a larger transfer unit, and adaptation between invalidate and update. Our results indicate that adaptation between single- and multiple-writer and dynamic page aggregation are clearly beneficial. The results for the adaptation between invalidate and update are less compelling, showing at best gains similar to the dynamic aggregation adaptation and at worst serious performance deterioration.
Cristiana Amza, Alan L. Cox, Sandhya Dwarkadas, Li-Jie Jin, Karthick Rajamani, Willy Zwaenepoel
Proc. IEEE5
1997 Trade-offs Between False Sharing and Aggregation in Software Distributed Shared Memory
abstract
Software Distributed Shared Memory (DSM) systems based on virtual memory techniques traditionally use the hardware page as the consistency unit. The large size of the hardware page is considered to be a performance bottleneck because of the implied false sharing overheads. Instead, we show that in the presence of a relaxed consistency model and a multiple writer protocol, a large consistency unit is generally not detrimental to performance. We study the tradeoffs between false sharing and aggregation effects when using large consistency units. In this context, this paper makes three separate contributions:1. We document the cost of false sharing in terms of extra messages and extra data being communicated. We find that, for the applications considered, when the virtual memory page is used as the consistency unit, the number of extra messages is small, while the amount of extra data can be substantial.2. We evaluate the performance when the consistency unit is increased to a multiple of the virtual memory page size. For most applications and data sets, the performance improves, except when the false sharing effects include extra messages or a large amount of extra data.3. We present a new algorithm for dynamically aggregating pages. In our algorithm, the aggregated pages do not necessarily need to be contiguous. In all cases, the performance of our dynamic aggregation algorithm is similar to that achieved with the best static page size.These results were obtained by measuring the performance of eight applications on the TreadMarks distributed shared memory system. The hardware platform used is a network of 166Mhz Pentiums connected by a switched 100Mbps Ethernet network.
Cristiana Amza, Alan L. Cox, Karthick Rajamani, Willy Zwaenepoel
PPoPP3