Shekhar Srikantaiah

dblp:18/1418 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 6 first-authorSoftware engineering, systems software and programming languages · 5 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
14 papers
Memory systems · 31% Storage systems · 31% Processor architecture and microarchitecture · 10%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › flash and SSD › flash memory management
garbage collection
0.622020
Design of a Host Interface Logic for GC-Free SSDs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
HIOS: A host interface I/O scheduler for Solid State Disks · ISCA 2014
Processor architecture and microarchitecture
chip multiprocessor
0.672011
Coordinated power management of voltage islands in CMPs · SIGMETRICS 2010
CPM in CMPs: Coordinated Power Management in Chip-Multiprocessors · SC 2010
A case for integrated processor-cache partitioning in chip multiprocessors · SC 2009
Memory systems › cache management
shared cache management
0.552012
Courteous cache sharing: being nice to others in capacity management · DAC 2012
A case for integrated processor-cache partitioning in chip multiprocessors · SC 2009
Dynamic storage cache allocation in multi-server architectures · SC 2009
Memory systems
cache management
0.552012
Courteous cache sharing: being nice to others in capacity management · DAC 2012
MorphCache: A Reconfigurable Adaptive Multi-level Cache hierarchy · HPCA 2011
SHARP control: controlled shared cache management in chip multiprocessors · MICRO 2009
Storage systems
i/o scheduling
0.412020
Design of a Host Interface Logic for GC-Free SSDs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Embedded and real-time systems › real-time scheduling
quality-of-service-aware scheduling
0.412020
Design of a Host Interface Logic for GC-Free SSDs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Storage systems › flash and SSD
solid-state drive
0.412020
Design of a Host Interface Logic for GC-Free SSDs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Memory systems › cache management
cache partitioning
0.222011
MorphCache: A Reconfigurable Adaptive Multi-level Cache hierarchy · HPCA 2011
A case for integrated processor-cache partitioning in chip multiprocessors · SC 2009
Cloud and datacenter computing › resource management
shared resource management
0.222011
METE: meeting end-to-end QoS in multicores through system-wide resource management · SIGMETRICS 2011
SHARP control: controlled shared cache management in chip multiprocessors · MICRO 2009
Memory systems › cache management › storage caching
storage cache management
0.222011
QoS aware storage cache management in multi-server environments · PPoPP 2011
Dynamic storage cache allocation in multi-server architectures · SC 2009
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.222010
Coordinated power management of voltage islands in CMPs · SIGMETRICS 2010
CPM in CMPs: Coordinated Power Management in Chip-Multiprocessors · SC 2010
Energy-efficient computing
power management
0.222010
Coordinated power management of voltage islands in CMPs · SIGMETRICS 2010
CPM in CMPs: Coordinated Power Management in Chip-Multiprocessors · SC 2010
Storage systems
flash and SSD
0.212014
HIOS: A host interface I/O scheduler for Solid State Disks · ISCA 2014
Cloud and datacenter computing
quality of service
0.212014
HIOS: A host interface I/O scheduler for Solid State Disks · ISCA 2014
Storage systems › flash and SSD › SSD performance
SSD I/O scheduling
0.212014
HIOS: A host interface I/O scheduler for Solid State Disks · ISCA 2014
Storage systems
storage reliability
0.212014
HIOS: A host interface I/O scheduler for Solid State Disks · ISCA 2014
Processor architecture and microarchitecture
multicore design
0.222011
METE: meeting end-to-end QoS in multicores through system-wide resource management · SIGMETRICS 2011
Cache topology aware computation mapping for multicores · PLDI 2010
Memory systems › cache management
cache capacity management
0.112012
Courteous cache sharing: being nice to others in capacity management · DAC 2012
Storage systems › i/o architecture › i/o subsystem
host interface
0.112020
Design of a Host Interface Logic for GC-Free SSDs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Memory systems › memory hierarchy
cache hierarchy
0.112011
MorphCache: A Reconfigurable Adaptive Multi-level Cache hierarchy · HPCA 2011
Compilers and program optimization › loop transformation
loop distribution
0.112010
Cache topology aware computation mapping for multicores · PLDI 2010
Memory systems › memory management › virtual memory
address translation
0.112010
Synergistic TLBs for High Performance Address Translation in Chip Multiprocessors · MICRO 2010
Memory systems
cache
0.112010
Cache topology aware computation mapping for multicores · PLDI 2010
Integrated circuit design › clocking
multiple clock domain
0.112010
CPM in CMPs: Coordinated Power Management in Chip-Multiprocessors · SC 2010
Memory systems › memory management › virtual memory › address translation › TLB
TLB design
0.112010
Synergistic TLBs for High Performance Address Translation in Chip Multiprocessors · MICRO 2010
Memory systems › memory management
virtual memory
0.112010
Synergistic TLBs for High Performance Address Translation in Chip Multiprocessors · MICRO 2010
Distributed systems › distributed system architecture
multi-server architecture
0.112009
Dynamic storage cache allocation in multi-server architectures · SC 2009
Memory systems › cache › cache miss
cache miss classification
0.112008
Adaptive set pinning: managing shared caches in chip multiprocessors · ASPLOS 2008
Parallel and multicore computing › thread-level parallelism
multithreaded workloads
0.122012
Courteous cache sharing: being nice to others in capacity management · DAC 2012
MorphCache: A Reconfigurable Adaptive Multi-level Cache hierarchy · HPCA 2011
Performance modeling and evaluation › workload characterization
multiprogrammed workloads
0.012012
Courteous cache sharing: being nice to others in capacity management · DAC 2012

Methods — techniques the papers use, named apart from their topics

simulation · 0.5feedback control theory · 0.2full-system simulation · 0.2priority-based thread scheduling · 0.1resource management · 0.1qos scheduling · 0.1max-flow algorithm · 0.1workload characterization · 0.1two-tier control · 0.1control theory · 0.1
YearPublicationVenuePosition
2020 Design of a Host Interface Logic for GC-Free SSDs
abstract
Garbage collection (GC) and resource contention on I/O buses (channels) are among the critical bottlenecks in solid-state drives (SSDs) that cannot be easily hidden. Most existing I/O scheduling algorithms in the host interface logic (HIL) of state-of-the-art SSDs are oblivious to such low-level performance bottlenecks in SSDs. As a result, SSDs may violate quality of service (QoS) requirements by not being able to meet the deadlines of I/O requests. In this paper, we propose a novel host interface I/O scheduler that is both GC aware and QoS aware. The proposed scheduler redistributes the GC overheads across noncritical I/O requests and reduces channel resource contention. Our experiments with workloads from various application domains revealed that the proposed client-level SSD scheduler reduces the standard deviation for latency by 52.5% and the worst-case latency by 86.6%, compared to the state-of-the-art I/O schedulers used for the HIL. In addition, for I/O requests smaller than a superpage, the proposed scheduler avoids channel resource conflicts and reduces latency by 29.2% in comparison to the state-of-the-art I/O schedulers. Furthermore, we present an extension of the proposed I/O scheduler for enterprise SSDs based on the NVMe protocol.
Myoungsoo Jung, Wonil Choi, Miryeong Kwon, Shekhar Srikantaiah, Joonhyuk Yoo, Mahmut T. Kandemir
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2014 HIOS: A host interface I/O scheduler for Solid State Disks
abstract
Garbage collection (GC) and resource contention on I/O buses (channels) are among the critical bottlenecks in Solid State Disks (SSDs) that cannot be easily hidden. Most existing I/O scheduling algorithms in the host interface logic (HIL) of state-of-the-art SSDs are oblivious to such low-level performance bottlenecks in SSDs. As a result, SSDs may violate quality of service (QoS) requirements by not being able to meet the deadlines of I/O requests. In this paper, we propose a novel host interface I/O scheduler that is both GC-aware and QoS-aware. The proposed scheduler redistributes the GC overheads across non-critical I/O requests and reduces channel resource contention. Our experiments with workloads from various application domains reveal that the proposed scheduler reduces the standard deviation for latency over state-of-the-art I/O schedulers used in the HIL by 52.5%, and the worst-case latency by 86.6%. In addition, for I/O requests with sizes smaller than a superpage, our proposed scheduler avoids channel resource conflicts and reduces latency by 29.2% compared to the state-of-the-art.
Myoungsoo Jung, Wonil Choi, Shekhar Srikantaiah, Joonhyuk Yoo, Mahmut T. Kandemir
ISCA3
2012 PEPON: performance-aware hierarchical power budgeting for NoC based multicores
abstract
Targeting NoC based multicores, we propose a two-level power budget distribution mechanism, called PEPON, where the first level distributes the overall power budget of the multicore system among various types of on-chip resources like the cores, caches, and NoC, and the second level determines the allocation of power to individual instances of each type of resource. Both these distributions are oriented towards maximizing workload performance without exceeding the specified power budget. Extensive experimental evaluations of the proposed power distribution scheme using a full system simulation and detailed power models emphasize the importance of power budget partitioning at both levels. Specifically, our results show that the proposed scheme can provide up to 29% performance improvement as compared to no power budgeting, and performs 13% better than a competing scheme, under the same chip-wide power cap.
Akbar Sharifi, Asit K. Mishra, Shekhar Srikantaiah, Mahmut T. Kandemir, Chita R. Das
PACT3
2012 Courteous cache sharing: being nice to others in capacity management
abstract
This paper proposes a cache management scheme for multiprogrammed, multithreaded applications, with the objective of obtaining maximum performance for both individual applications and the multithreaded workload mix. In this scheme, each individual application's performance is improved by increasing the priority of its slowest thread, while the overall system performance is improved by ensuring that each individual application's performance benefit does not come at the cost of a significant degradation to other application's threads that are sharing the same cache. Averaged over six workloads, our shared cache management scheme improves the performance of the combination of applications by 18%. These improvements across applications in each mix are also fair, as indicated by average fair speedup improvements of 10% across the threads of each application (averaged over all the workloads).
Akbar Sharifi, Shekhar Srikantaiah, Mahmut T. Kandemir, Mary Jane Irwin
DAC2
2011 Adaptive QoS Decomposition and Control for Storage Cache Management in Multi-server Environments
abstract
Poor I/O performance can prevent an application from scaling to a large number of nodes even if the computation is parallelized appropriately. Therefore, improving I/O performance of large-scale parallel applications is very important. Caching recently and frequently accessed I/O blocks in memory is a widely used technique for improving I/O performance of these applications on high-end machines. However, simultaneous storage cache accesses of multiple applications may lead to unacceptable degradations in application performance due to interferences at the storage cache layer. As a result, efficient management of storage cache space across multiple I/O servers among competing applications is critical in order to ensure performance quality of service (QoS) to individual applications. In this paper, we propose a novel two-step approach to the management of the storage caches to provide predictable performance in multi-server storage architectures: (1)An adaptive QoS decomposition and optimization step uses max-flow algorithm to determine the best decomposition of application-level QoS to sub-QoSs such that the application performance is optimized, and (2) A storage cache allocation step uses feedback control theory to allocates hared storage cache space such that the specified QoSs are satisfied throughout the execution. Our experimental evaluation indicates that, on an average, our approach improves the I/O throughput of applications by 48.6%, 29.2%, and 20.7%, respectively, over the uncontrolled partitioning, fair share and uniform decomposition schemes. We also observed 31.4%, 20.2%, and 44.7% improvements by our approach, in our global metric, called the fair speedup metric, against the fair share, uncontrolled partitioning and uniform decomposition schemes, respectively.
Ramya Prabhakar, Shekhar Srikantaiah, Rajat Garg, Mahmut T. Kandemir
CCGRID2
2011 MorphCache: A Reconfigurable Adaptive Multi-level Cache hierarchy
abstract
Given the diverse range of application characteristics that chip multiprocessors (CMPs) need to cater to, a “one-cache-topology-fits-all” design philosophy will clearly be inadequate. In this paper, we propose MorphCache, a Reconfigurable Adaptive Multi-level Cache hierarchy. Mor-phCache dynamically tunes a multi-level cache topology in a CMP to allow significantly different cache topologies to exist on the same architecture. Starting from per-core L2 and L3 cache slices as the basic design point, MorphCache alters the cache topology dynamically by merging or splitting cache slices and modifying the accessibility of different cache slice groups to different cores in a CMP. We evaluated MorphCache on a 16 core CMP on a full system simulator and found that it significantly improves both average throughput and harmonic mean of speedups of diverse multithreaded and multiprogrammed workloads. Specifically, our results show that MorphCache improves throughput of the multiprogrammed mixes by 29.9% over a topology with all-shared L2 and L3 caches and 27.9% over a topology with per core private L2 cache and shared L3 cache. In addition, we also compared MorphCache to partitioning a single shared cache at each level using promotion/insertion pseudo-partitioning (PIPP) [28] and managing per-core private cache at each level using dynamic spill receive caches (DSR) [18]. We found that MorphCache improves average throughput by 6.6% over PIPP and by 5.7% over DSR when applied to both L2 and L3 caches.
Shekhar Srikantaiah, Emre Kultursay, Tao Zhang 0032, Mahmut T. Kandemir, Mary Jane Irwin, Yuan Xie 0001
HPCA1
2011 Improving shared cache behavior of multithreaded object-oriented applications in multicores
abstract
Understanding shared cache performance when executing multithreaded object-oriented applications and optimizing these applications for multicores have not received much attention. In this paper, we first quantify the intra-thread and inter-thread cache line (block) reuse characteristics of a set of multithreaded C++ programs when executed in shared cache based multicores. Our results show that, as far as shared on-chip caches are concerned, inter-thread cache line (block) reuse distances are much higher than intra-thread cache line reuse distances. We study the impact of these characteristics on the hit/miss behavior of the shared last-level cache on a commercial multicore machine. We then show that, by rearranging accesses to the objects shared across different threads and to the objects stored in nearby memory locations, inter-thread (temporal and spatial) object reuse distances can be reduced, which in turn helps to reduce inter-thread cache line reuse distances. The results we collected using eight multithreaded applications show that our proposed shared cache-aware code restructuring strategy can reduce misses in the last-level on-chip cache of a commercial multicore machine by 25.4%, on average. These savings in cache misses translate in turn to average execution time improvement of 11.9%.
Mahmut T. Kandemir, Shekhar Srikantaiah, Seung Woo Son 0001
ICCAD2
2011 Feedback control based cache reliability enhancement for emerging multicores
abstract
Focusing on data reliability, we propose a control theory centric approach designed to improve transient error resilience in shared caches of emerging multicores while satisfying performance goals. The proposed scheme takes, as input, two quality of service (QoS) specifications: performance QoS and reliability QoS. The first of these indicates the minimum workload-wide cache (L2) hit rate value acceptable, whereas the second one captures the reliability bound on an application basis, with the help of a metric called the Reads-with-Replica (RwR). We present an extensive experimental evaluation of the proposed scheme on various workloads formed using the applications from the SPEC2006 benchmark suite. The proposed scheme is able to satisfy, in most of the tested cases, both performance and reliability QoS targets, by successfully modulating the total size of the data replication area and partitioning of this area among the co-runner applications. The collected results also show that our scheme achieves consistent improvements under different values of the major simulation parameters.
Hui Zhao 0013, Akbar Sharifi, Shekhar Srikantaiah, Mahmut T. Kandemir
ICCAD3
2011 QoS aware storage cache management in multi-server environments
abstract
In this paper, we propose a novel two-step approach to the management of the storage caches to provide predictable performance in multi-server storage architectures: (1) An adaptive QoS decomposition and optimization step uses max-flow algorithm to determine the best decomposition of application-level QoS to sub-QoSs such that the application performance is optimized, and (2) A storage cache allocation step uses feedback control theory to allocate shared storage cache space such that the specified QoSs are satisfied throughout the execution.
Ramya Prabhakar, Shekhar Srikantaiah, Rajat Garg, Mahmut T. Kandemir
PPoPP2
2011 METE: meeting end-to-end QoS in multicores through system-wide resource management
abstract
Management of shared resources in emerging multicores for achieving predictable performance has received considerable attention in recent times. In general, almost all these approaches attempt to guarantee a certain level of performance QoS (weighted IPC, harmonic speedup, etc) by managing a single shared resource or at most a couple of interacting resources. A fundamental shortcoming of these approaches is the lack of coordination between these shared resources to satisfy a system level QoS. This is undesirable because providing end-to-end QoS in future multicores is essential for supporting wide-spread adoption of these architectures in virtualized servers and cloud computing systems. An initial step towards such an end-to-end QoS support in multicores is to ensure that at least the major computational and memory resources on-chip are managed efficiently in a coordinated fashion.
Akbar Sharifi, Shekhar Srikantaiah, Asit K. Mishra, Mahmut T. Kandemir, Chita R. Das
SIGMETRICS2
2010 SRP: Symbiotic Resource Partitioning of the Memory Hierarchy in CMPs
Shekhar Srikantaiah, Mahmut T. Kandemir
HiPEAC1
2010 Adaptive multi-level cache allocation in distributed storage architectures
abstract
Increasing complexity of large-scale applications and continuous increases in data set sizes of such applications combined with slow improvements in disk access latencies has resulted in I/O becoming a performance bottleneck. While there are several ways of improving I/O access latencies of dataintensive applications, one of the promising approaches has been using different layers of the I/O subsystem to cache recently and/or frequently used data so that the number of I/O requests accessing the disk is reduced. These different layers of caches across the storage hierarchy introduce the need for efficient cache management schemes to derive maximum performance benefits. Several state-of-the-art multi-level storage cache management schemes focus on optimizing aggregate hit rate or overall I/O latency, while being agnostic to Service Level Objectives (SLOs). Also, most of the existing works focus on different cache replacement algorithms for managing storage caches and discuss different exclusive caching techniques in the context of multilevel cache hierarchy. However, the orthogonal problem of storage cache space allocation to multiple, simultaneously-running applications in a multi-level hierarchy of storage caches with multiple storage servers has remained an open research problem. In this work, using a combination of per-application latency model and a linear programming model, we proportion storage caches dynamically among multiple concurrently-executing applications across the different levels of the storage hierarchy and across multiple servers to provide isolation to applications while satisfying the application level SLOs. Further, our algorithm improves the overall system performance significantly.
Ramya Prabhakar, Shekhar Srikantaiah, Mahmut T. Kandemir, Christina M. Patrick
ICS2
2010 Synergistic TLBs for High Performance Address Translation in Chip Multiprocessors
abstract
Translation Look-aside Buffers (TLBs) are vital hardware support for virtual memory management in high performance computer systems and have a momentous influence on overall system performance. Numerous techniques to reduce TLB miss latencies including the impact of TLB size, associativity, multilevel hierarchies, super pages, and prefetching have been well studied in the context of uniprocessors. However, with Chip Multiprocessors (CMPs) becoming the standard design point of processor architectures, it is imperative that we review the design and organization of TLBs in the context of CMPs. In this paper, we propose to improve system performance by means of a novel way of organizing TLBs called Synergistic TLBs. Synergistic TLB is different from per-core private TLB organization in three ways: (i) it provides capacity sharing of TLBs by facilitating storing of victim translations from one TLB in another to emulate a distributed shared TLB (DST), (ii) it supports translation migration for maximizing the utilization of TLB capacity, and (iii) it supports translation replication to avoid excess latency for remote TLB accesses. We explore all the design points in this design space and find that an optimal point exists for high performance address translation. Our evaluation with both multiprogrammed (SPEC 2006 applications) and multithreaded workloads (PARSEC applications) shows that Synergistic TLBs can eliminate, respectively, 44.3% and 31.2% of the TLB misses, on average. It also improves the weighted speedup of multiprogrammed application mixes by 25.1% and performance of multithreaded applications by 27.3%, on average.
Shekhar Srikantaiah, Mahmut T. Kandemir
MICRO1
2010 Cache topology aware computation mapping for multicores
abstract
The main contribution of this paper is a compiler based, cache topology aware code optimization scheme for emerging multicore systems. This scheme distributes the iterations of a loop to be executed in parallel across the cores of a target multicore machine and schedules the iterations assigned to each core. Our goal is to improve the utilization of the on-chip multi-layer cache hierarchy and to maximize overall application performance. We evaluate our cache topology aware approach using a set of twelve applications and three different commercial multicore machines. In addition, to study some of our experimental parameters in detail and to explore future multicore machines (with higher core counts and deeper on-chip cache hierarchies), we also conduct a simulation based study. The results collected from our experiments with three Intel multicore machines show that the proposed compiler-based approach is very effective in enhancing performance. In addition, our simulation results indicate that optimizing for the on-chip cache hierarchy will be even more important in future multicores with increasing numbers of cores and cache levels.
Mahmut T. Kandemir, Taylan Yemliha, Sai Prashanth Muralidhara, Shekhar Srikantaiah, Mary Jane Irwin
PLDI4
2010 CPM in CMPs: Coordinated Power Management in Chip-Multiprocessors
abstract
Multiple clock domain architectures have recently been proposed to alleviate the power problem in CMPs by having different frequency/voltage values assigned to each domain based on workload requirements. However, accurate allocation of power to these voltage/frequency islands based on time varying workload characteristics as well as controlling the power consumption at the provisioned power level is quite non-trivial. Toward this end, we propose a two-tier feedback-based control theoretic solution. Our first-tier consists of a global power manager that allocates power targets to individual islands based on the workload dynamics. The power consumptions of these islands are in turn controlled by a second-tier, consisting of local controllers that regulate island power using dynamic voltage and frequency scaling in response to workload requirements.
Asit K. Mishra, Shekhar Srikantaiah, Mahmut T. Kandemir, Chita R. Das
SC2
2010 Coordinated power management of voltage islands in CMPs
abstract
Multiple clock domain architectures have recently been proposed to alleviate the power problem in CMPs by having different frequency/voltage values assigned to each domain based on workload requirements. However, accurate allocation of power to these voltage/frequency islands based on time varying workload characteristics as well as controlling the power consumption at the provisioned power level is non-trivial. Toward this end, we propose a two-tier feedback-based control theoretic solution. Our first-tier consists of a global power manager that allocates power targets to individual islands based on the workload dynamics. The power consumptions of these islands are in turn controlled by a second-tier, consisting of local controllers that regulate island power using dynamic voltage and frequency scaling in response to workload requirements.
Asit K. Mishra, Shekhar Srikantaiah, Mahmut T. Kandemir, Chita R. Das
SIGMETRICS2
2009 SHARP control: controlled shared cache management in chip multiprocessors
abstract
Shared resources in a chip multiprocessors (CMPs) pose unique challenges to the seamless adoption of CMPs in virtualization environments and high performance computing systems. While sharing resources like on-chip last level cache is generally beneficial due to increased resource utilization, lack of control over management of these resources can lead to loss of determinism, faded performance isolation, and an overall lack of the notion of Quality of Service (QoS) provided to individual applications. This has direct ramifications on adhering to service level agreements in environments involving consolidation of multiple heterogeneous workloads. Although providing QoS in presence of shared resources has been addressed in the literature, it has been commonly observed that reservation of resources for QoS leads to under-utilization of resources.
Shekhar Srikantaiah, Mahmut T. Kandemir, Qian Wang 0029
MICRO1
2009 Dynamic storage cache allocation in multi-server architectures
abstract
We introduce a dynamic and efficient shared cache management scheme, called Maxperf, that manages the aggregate cache space in multi-server storage architectures such that the service level objectives (SLOs) of concurrently executing applications are satisfied and any spare cache capacity is proportionately allocated according to the marginal gains of the applications to maximize performance. We use a combination of Neville's algorithm and linear-programming-model to discover the required storage cache partition size, on each server, for every application accessing that server. Experimental results show that our algorithm enforces partitions to provide stronger isolation to applications, meets application level SLOs even in the presence of dynamically changing storage cache requirements, and improves I/O latency of individual applications as well as the overall I/O latency significantly compared to two alternate storage cache management schemes, and a state-of-the-art single server storage cache management scheme extended to multi-server architecture.
Ramya Prabhakar, Shekhar Srikantaiah, Christina M. Patrick, Mahmut T. Kandemir
SC2
2009 A case for integrated processor-cache partitioning in chip multiprocessors
abstract
Existing cache partitioning schemes are designed in a manner oblivious to the implicit processor partitioning enforced by the operating system. This paper examines an operating system directed integrated processor-cache partitioning scheme that partitions both the available processors and the shared cache in a chip multiprocessor among different multi-threaded applications. Extensive simulations using a set of multiprogrammed workloads show that our integrated processor-cache partitioning scheme facilitates achieving better performance isolation as compared to state of the art hardware/software based solutions. Specifically, our integrated processor-cache partitioning approach performs, on an average, 20.83% and 14.14% better than equal partitioning and the implicit partitioning enforced by the underlying operating system, respectively, on the fair speedup metric on an 8 core system. We also compare our approach to processor partitioning alone and a state-of-the-art cache partitioning scheme and our scheme fares 8.21% and 9.19% better than these schemes on a 16 core system.
Shekhar Srikantaiah, Reetuparna Das, Asit K. Mishra, Chita R. Das, Mahmut T. Kandemir
SC1
2008 Adaptive set pinning: managing shared caches in chip multiprocessors
abstract
As part of the trend towards Chip Multiprocessors (CMPs) for the next leap in computing performance, many architectures have explored sharing the last level of cache among different processors for better performance-cost ratio and improved resource allocation. Shared cache management is a crucial CMP design aspect for the performance of the system. This paper first presents a new classification of cache misses - CII: Compulsory, Inter-processor and Intra-processor misses - for CMPs with shared caches to provide a better understanding of the interactions between memory transactions of different processors at the level of shared cache in a CMP. We then propose a novel approach, called set pinning, for eliminating inter-processor misses and reducing intra-processor misses in a shared cache. Furthermore, we show that an adaptive set pinning scheme improves over the benefits obtained by the set pinning scheme by significantly reducing the number of off-chip accesses. Extensive analysis of these approaches with SPEComp 2001 benchmarks is performed using a full system simulator. Our experiments indicate that the set pinning scheme achieves an average improvement of 22.18% in the L2 miss rate while the adaptive set pinning scheme reduces the miss rates by an average of 47.94% as compared to the traditional shared cache scheme. They also improve the performance by 7.24% and 17.88% respectively.
Shekhar Srikantaiah, Mahmut T. Kandemir, Mary Jane Irwin
ASPLOS1
2008 Integrated code and data placement in two-dimensional mesh based chip multiprocessors
abstract
As transistor sizes continue to shrink and the number of transistors per chip keeps increasing, chip multiprocessors (CMPs) are becoming a promising alternative to remain on the current performance trajectory for both high-end systems and embedded systems. Since future technologies offer the promise of being able to integrate billions of transistors on a chip, the prospects of having hundreds to thousands of processors on a single chip along with an underlying memory hierarchy and an interconnection system is entirely feasible. This paper proposes a compiler directed integrated code and data placement scheme for two-dimensional mesh based CMP architectures. The proposed approach uses a Code-Data Affinity Graph (CDAG) to represent the relationship between loop iterations and array data and then assigns the sets of loop iterations to processing cores and sets of data blocks to on-chip memories. During the mapping process, the on-chip memory capacity and load imbalance across different cores and the topology of the NoC are taken into account. In this paper, we present two variants of our approach: depth-first placement (DFP) and breadth-first placement (BFP), and compare them to three alternate code/data mapping schemes. The experimental evaluation shows that our CDAG based placement schemes are very successful in practice, achieving average performance improvements of 19.9% (DFP) and 16.8% (BFP), and average energy improvements of 29.7% (DFP) and 27.8% (BFP).
Taylan Yemliha, Shekhar Srikantaiah, Mahmut T. Kandemir, Mustafa Karaköy, Mary Jane Irwin
ICCAD2
2008 SPM management using Markov chain based data access prediction
abstract
Leveraging the power of scratchpad memories (SPMs) available in most embedded systems today is crucial to extract maximum performance from application programs. While regular accesses like scalar values and array expressions with affine subscript functions have been tractable for compiler analysis (to be prefetched into SPM), irregular accesses like pointer accesses and indexed array accesses have not been easily amenable for compiler analysis. This paper presents an SPM management technique using Markov chain based data access prediction for such irregular accesses. Our approach takes advantage of inherent, but hidden reuse in data accesses made by irregular references. We have implemented our proposed approach using an optimizing compiler. In this paper, we also present a thorough comparison of our different dynamic prediction schemes with other SPM management schemes. SPM management using our approaches produces 12.7% to 28.5% improvements in performance across a range of applications with both regular and irregular access patterns, with an average improvement of 20.8%.
Taylan Yemliha, Shekhar Srikantaiah, Mahmut T. Kandemir, Ozcan Ozturk 0001
ICCAD2