Ramya Prabhakar

dblp:49/7475 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 7 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Storage systems · 44% Memory systems · 41% Distributed systems · 9%

Topics — the 6 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › cache management › storage caching
storage cache management
0.332011
Virtual I/O caching: dynamic storage cache management for concurrent workloads · SC 2011
QoS aware storage cache management in multi-server environments · PPoPP 2011
Dynamic storage cache allocation in multi-server architectures · SC 2009
Storage systems
backup storage
0.312018
Can't We All Get Along? Redesigning Protection Storage for Modern Workloads · USENIX ATC 2018
Storage systems
storage reliability
0.112018
Can't We All Get Along? Redesigning Protection Storage for Modern Workloads · USENIX ATC 2018
Distributed systems › distributed system architecture
multi-server architecture
0.112009
Dynamic storage cache allocation in multi-server architectures · SC 2009
Memory systems › cache management
shared cache management
0.112009
Dynamic storage cache allocation in multi-server architectures · SC 2009
Storage systems
distributed storage
0.012011
Virtual I/O caching: dynamic storage cache management for concurrent workloads · SC 2011

Methods — techniques the papers use, named apart from their topics

max-flow algorithm · 0.1feedback control theory · 0.1cache management · 0.1neville's algorithm · 0.1linear programming · 0.1
YearPublicationVenuePosition
2018 Can't We All Get Along? Redesigning Protection Storage for Modern Workloads
Yamini Allu, Fred Douglis, Mahesh Kamat, Ramya Prabhakar, Philip Shilane, Rahul Ugale
USENIX ATC4
2013 Disk-Cache and Parallelism Aware I/O Scheduling to Improve Storage System Performance
abstract
Modern large computing systems employ sophisticated disk I/O systems that are configured to deliver high-throughput, low-latency disk I/O to multiple clients accessing them. However, due to potential interferences among concurrent I/O accesses issued by multiple clients, a disk-cache and disk-level parallelism unaware I/O scheduling algorithm employed by the operating system/storage controller may have a significant impact on both system throughput and I/O latency. In this paper, we propose two fundamentally new disk I/O scheduling techniques. The first technique, called DCAP, performs I/O scheduling in a disk cache aware and parallelism aware manner. The key idea in DCAP is to process simultaneous requests to different disks from the same application/priority class together and reorder them so that they have the highest number of hits in the disk cache. We then propose an enhanced version of DCAP called DCAP-G, that aggregates requests into service groups to alleviate the problem of request starvation that may occur in DCAP in certain cases. We evaluate both DCAP and DCAP-G using a set of I/O workloads from production-based enterprise systems as well as high-performance computing domain. In addition, we also compare the performance of our algorithms to previously proposed I/O scheduling algorithms. Our evaluation shows that, averaged across all our workloads, DCAP improves the average I/O response time, taking maximum advantage of disk access locality and exploiting parallelism among concurrent accesses to multiple disks, by 14.9% over an I/O scheduler that schedules requests on a first-come-first-served (FCFS) basis and also improves by 6.5% over a previously proposed locality-optimal I/O scheduler (SPCTF). In addition to these improvements, DCAP-G improves the average I/O response time by 6.6% over DCAP, leading to an overall 20.7% and 12.0% improvement over FCFS, and SPCTF, respectively.
Ramya Prabhakar, Mahmut T. Kandemir, Myoungsoo Jung
IPDPS1
2012 MROrchestrator: A Fine-Grained Resource Orchestration Framework for MapReduce Clusters
abstract
Efficient resource management in data centers and clouds running large distributed data processing frameworks like MapReduce is crucial for enhancing the performance of hosted applications and increasing resource utilization. However, existing resource scheduling schemes in Hadoop MapReduce allocate resources at the granularity of fixed-size, static portions of nodes, called slots. In this work, we show that MapReduce jobs have widely varying demands for multiple resources, making the static and fixed-size slot-level resource allocation a poor choice both from the performance and resource utilization standpoints. Furthermore, lack of coordination in the management of multiple resources across nodes prevents dynamic slot reconfiguration, and leads to resource contention. Motivated by this, we propose MROrchestrator, a MapReduce resource Orchestrator framework, which can dynamically identify resource bottlenecks, and resolve them through fine-grained, coordinated, and on-demand resource allocations. We have implemented MROrchestrator on two 24-node native and virtualized Hadoop clusters. Experimental results with a suite of representative MapReduce benchmarks demonstrate up to 38% reduction in job completion times, and up to 25% increase in resource utilization. We further demonstrate the performance boost in existing resource managers like NGM and Mesos, when augmented with MROrchestrator.
Bikash Sharma, Ramya Prabhakar, Seung-Hwan Lim, Mahmut T. Kandemir, Chita R. Das
IEEE CLOUD2
2012 On Urgency of I/O Operations
abstract
Many high-performance parallel file systems and storage hierarchies employ multilayer storage caches in an attempt to reduce data access latencies. In current storage cache hierarchies, all data requests are treated uniformly and hit/miss characteristics are dictated only by the degree of reuse exhibited by data blocks. In reality however, different I/O operations may have different urgencies (criticalities), and in particular, some I/O operations can be delayed without having a major impact on overall application performance. Motivated by this observation, we define the concept of I/O operation urgency (criticality) and study the critical latencies of I/O operations for a set of seven high-performance applications that manipulate disk-resident data sets. We propose and experimentally evaluate three profile-based strategies for exploiting urgent I/O operations in managing storage caches. The results collected with these schemes on both two-tier and three-tier systems indicate that significant performance improvements are possible if one could exploit urgencies of different I/O operations in managing storage caches.
Mahmut T. Kandemir, Taylan Yemliha, Ramya Prabhakar, Myoungsoo Jung
CCGRID3
2012 Taking Garbage Collection Overheads Off the Critical Path in SSDs
Myoungsoo Jung, Ramya Prabhakar, Mahmut T. Kandemir
Middleware2
2011 Adaptive QoS Decomposition and Control for Storage Cache Management in Multi-server Environments
abstract
Poor I/O performance can prevent an application from scaling to a large number of nodes even if the computation is parallelized appropriately. Therefore, improving I/O performance of large-scale parallel applications is very important. Caching recently and frequently accessed I/O blocks in memory is a widely used technique for improving I/O performance of these applications on high-end machines. However, simultaneous storage cache accesses of multiple applications may lead to unacceptable degradations in application performance due to interferences at the storage cache layer. As a result, efficient management of storage cache space across multiple I/O servers among competing applications is critical in order to ensure performance quality of service (QoS) to individual applications. In this paper, we propose a novel two-step approach to the management of the storage caches to provide predictable performance in multi-server storage architectures: (1)An adaptive QoS decomposition and optimization step uses max-flow algorithm to determine the best decomposition of application-level QoS to sub-QoSs such that the application performance is optimized, and (2) A storage cache allocation step uses feedback control theory to allocates hared storage cache space such that the specified QoSs are satisfied throughout the execution. Our experimental evaluation indicates that, on an average, our approach improves the I/O throughput of applications by 48.6%, 29.2%, and 20.7%, respectively, over the uncontrolled partitioning, fair share and uniform decomposition schemes. We also observed 31.4%, 20.2%, and 44.7% improvements by our approach, in our global metric, called the fair speedup metric, against the fair share, uncontrolled partitioning and uniform decomposition schemes, respectively.
Ramya Prabhakar, Shekhar Srikantaiah, Rajat Garg, Mahmut T. Kandemir
CCGRID1
2011 Multilayer Cache Partitioning for Multiprogram Workloads
Mahmut T. Kandemir, Ramya Prabhakar, Mustafa Karaköy
Euro-Par (1)2
2011 Provisioning a Multi-tiered Data Staging Area for Extreme-Scale Machines
abstract
Massively parallel scientific applications, running on extreme-scale supercomputers, produce hundreds of terabytes of data per run, driving the need for storage solutions to improve their I/O performance. Traditional parallel file systems (PFS) in high performance computing (HPC) systems are unable to keep up with such high data rates, creating a storage wall. In this work, we present a novel multi-tiered storage architecture comprising hybrid node-local resources to construct a dynamic data staging area for extreme-scale machines. Such a staging ground serves as an impedance matching device between applications and the PFS. Our solution combines diverse resources (e.g., DRAM, SSD) in such a way as to approach the performance of the fastest component technology and the cost of the least expensive one. We have developed an automated provisioning algorithm that aids in meeting the check pointing performance requirement of HPC applications, by using a least-cost storage configuration. We evaluate our approach using both an implementation on a large scale cluster and a simulation driven by six-years worth of Jaguar supercomputer job-logs, and show that our approach, by choosing an appropriate storage configuration, achieves 41.5% cost savings with only negligible impact on performance.
Ramya Prabhakar, Sudharshan S. Vazhkudai, Youngjae Kim 0001, Ali Raza Butt, Mahmut T. Kandemir
ICDCS1
2011 SRC: virtual i/o caching: dynamic storage cache management for concurrent workloads
abstract
A leading cause of unpredictable application performance in distributed systems is contention at the storage layer, where resources are multiplexed among concurrent data intensive workloads. We target the shared storage cache, used to alleviate disk I/O bottlenecks, and propose a new caching paradigm to improve performance.
Michael R. Frasca, Ramya Prabhakar
ICS2
2011 QoS aware storage cache management in multi-server environments
abstract
In this paper, we propose a novel two-step approach to the management of the storage caches to provide predictable performance in multi-server storage architectures: (1) An adaptive QoS decomposition and optimization step uses max-flow algorithm to determine the best decomposition of application-level QoS to sub-QoSs such that the application performance is optimized, and (2) A storage cache allocation step uses feedback control theory to allocate shared storage cache space such that the specified QoSs are satisfied throughout the execution.
Ramya Prabhakar, Shekhar Srikantaiah, Rajat Garg, Mahmut T. Kandemir
PPoPP1
2011 Virtual I/O caching: dynamic storage cache management for concurrent workloads
abstract
A leading cause of reduced or unpredictable application performance in distributed systems is contention at the storage layer, where resources are multiplexed among many concurrent data intensive workloads. We target the shared storage cache, used to alleviate disk I/O bottlenecks, and propose a new caching paradigm to both improve performance and reduce memory requirements for HPC storage systems.
Michael R. Frasca, Ramya Prabhakar, Padma Raghavan, Mahmut T. Kandemir
SC2
2010 Adaptive multi-level cache allocation in distributed storage architectures
abstract
Increasing complexity of large-scale applications and continuous increases in data set sizes of such applications combined with slow improvements in disk access latencies has resulted in I/O becoming a performance bottleneck. While there are several ways of improving I/O access latencies of dataintensive applications, one of the promising approaches has been using different layers of the I/O subsystem to cache recently and/or frequently used data so that the number of I/O requests accessing the disk is reduced. These different layers of caches across the storage hierarchy introduce the need for efficient cache management schemes to derive maximum performance benefits. Several state-of-the-art multi-level storage cache management schemes focus on optimizing aggregate hit rate or overall I/O latency, while being agnostic to Service Level Objectives (SLOs). Also, most of the existing works focus on different cache replacement algorithms for managing storage caches and discuss different exclusive caching techniques in the context of multilevel cache hierarchy. However, the orthogonal problem of storage cache space allocation to multiple, simultaneously-running applications in a multi-level hierarchy of storage caches with multiple storage servers has remained an open research problem. In this work, using a combination of per-application latency model and a linear programming model, we proportion storage caches dynamically among multiple concurrently-executing applications across the different levels of the storage hierarchy and across multiple servers to provide isolation to applications while satisfying the application level SLOs. Further, our algorithm improves the overall system performance significantly.
Ramya Prabhakar, Shekhar Srikantaiah, Mahmut T. Kandemir, Christina M. Patrick
ICS1
2010 Automated Tracing of I/O Stack
Seong Jo Kim, Seung Woo Son 0001, Ramya Prabhakar, Mahmut T. Kandemir, Christina M. Patrick, Wei-keng Liao, Alok N. Choudhary
EuroMPI4
2009 Markov Model Based Disk Power Management for Data Intensive Workloads
abstract
In order to meet the increasing demands of present and upcoming data-intensive computer applications, there has been a major shift in the disk subsystem, which now consists of more disks with higher storage capacities and higher rotational speeds. These have made the disk subsystem a major consumer of power, making disk power management an important issue. People have considered the option of spinning down the disk during periods of idleness or serving the requests at lower rotational speeds when performance is not an issue. Accurately predicting future disk idle periods is crucial to such schemes. This paper presents a novel disk-idleness prediction mechanism based on Markov models and explains how this mechanism can be used in conjunction with a three-speed disk. Our experimental evaluation using a diverse set of workloads indicates that (i) prediction accuracies achieved by the proposed scheme are very good (87.5% on average); (ii) it generates significant energy savings over the traditional power-saving method of spinning down the disk when idle (35.5% on average); (iii) it performs better than a previously proposed multi-speed disk management scheme (19% on average); and (iv) the performance penalty is negligible (less than 1% on average). Overall, our implementation and experimental evaluation using both synthetic disk traces and traces extracted from real applications demonstrate the feasibility of a Markov-model-based approach to saving disk power.
Rajat Garg, Seung Woo Son 0001, Mahmut T. Kandemir, Padma Raghavan, Ramya Prabhakar
CCGRID5
2009 MPISec I/O: Providing Data Confidentiality in MPI-I/O
abstract
Applications performing scientific computations or processing streaming media benefit from parallel I/O significantly, as they operate on large data sets that require large I/O. MPI-I/O is a commonly used library interface in parallel applications to perform I/O efficiently. Optimizations like collective-I/O embedded in MPI-I/O allow multiple processes executing in parallel to perform I/O by merging requests of other processes and sharing them later. In such a scenario, preserving confidentiality of disk-resident data from unauthorized accesses by processes without significantly impacting performance of the application is a challenging task. In this paper, we evaluate the impact of ensuring data-confidentiality in MPI-I/O on the performance of parallel applications and provide an enhanced interface, called MPISec I/O, which brings an average overhead of only 5.77% over MPI-I/O in the best case, and about 7.82% in the average case.
Ramya Prabhakar, Christina M. Patrick, Mahmut T. Kandemir
CCGRID1
2009 Dynamic storage cache allocation in multi-server architectures
abstract
We introduce a dynamic and efficient shared cache management scheme, called Maxperf, that manages the aggregate cache space in multi-server storage architectures such that the service level objectives (SLOs) of concurrently executing applications are satisfied and any spare cache capacity is proportionately allocated according to the marginal gains of the applications to maximize performance. We use a combination of Neville's algorithm and linear-programming-model to discover the required storage cache partition size, on each server, for every application accessing that server. Experimental results show that our algorithm enforces partitions to provide stronger isolation to applications, meets application level SLOs even in the presence of dynamically changing storage cache requirements, and improves I/O latency of individual applications as well as the overall I/O latency significantly compared to two alternate storage cache management schemes, and a state-of-the-art single server storage cache management scheme extended to multi-server architecture.
Ramya Prabhakar, Shekhar Srikantaiah, Christina M. Patrick, Mahmut T. Kandemir
SC1