Shriram Rajagopalan

dblp:14/10315 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
0since 2021 · last 2020
0000-0002-0368-5207ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-authorComputer networks · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 4Software engineering, systems software and programming languages · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Cloud and datacenter computing · 48% Distributed systems · 37% Memory systems · 13%
Computer networks
4 papers
Software-defined and programmable networks · 100%
Databases, data mining, and information retrieval
2 papers
Distributed and cloud data management · 57% Database system architecture and tuning · 43%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software-defined and programmable networks
network function virtualization
1.142020
NFVnice: Dynamic Backpressure and Scheduling for NFV Service Chains · IEEE/ACM Trans. Netw. 2020
NFVnice: Dynamic Backpressure and Scheduling for NFV Service Chains · SIGCOMM 2017
Flurries: Countless Fine-Grained NFs for Flexible Per-Flow Customization · CoNEXT 2016
Software-defined and programmable networks › network function virtualization › service function chaining
service chain scheduling
0.722020
NFVnice: Dynamic Backpressure and Scheduling for NFV Service Chains · IEEE/ACM Trans. Netw. 2020
NFVnice: Dynamic Backpressure and Scheduling for NFV Service Chains · SIGCOMM 2017
Cloud and datacenter computing
resource management
0.722020
NFVnice: Dynamic Backpressure and Scheduling for NFV Service Chains · IEEE/ACM Trans. Netw. 2020
NFVnice: Dynamic Backpressure and Scheduling for NFV Service Chains · SIGCOMM 2017
Cloud and datacenter computing › job scheduling
CPU scheduling
0.412020
NFVnice: Dynamic Backpressure and Scheduling for NFV Service Chains · IEEE/ACM Trans. Netw. 2020
Distributed systems › replication
database replication
0.322013
RemusDB: transparent high availability for database systems · VLDB J. 2013
RemusDB: Transparent High Availability for Database Systems · Proc. VLDB Endow. 2011
Distributed systems
replication
0.322013
RemusDB: transparent high availability for database systems · VLDB J. 2013
RemusDB: Transparent High Availability for Database Systems · Proc. VLDB Endow. 2011
Software-defined and programmable networks › network function virtualization
service function chaining
0.212016
Flurries: Countless Fine-Grained NFs for Flexible Per-Flow Customization · CoNEXT 2016
Distributed and cloud data management
high availability
0.212013
RemusDB: transparent high availability for database systems · VLDB J. 2013
Operating systems › resource management › memory management
cache management
0.212013
Whose cache line is it anyway?: operating system support for live detection and repair of false sharing · EuroSys 2013
Memory systems
cache coherence
0.212013
Whose cache line is it anyway?: operating system support for live detection and repair of false sharing · EuroSys 2013
Memory systems › cache coherence
false sharing
0.212013
Whose cache line is it anyway?: operating system support for live detection and repair of false sharing · EuroSys 2013
Distributed systems
fault tolerance
0.212013
RemusDB: transparent high availability for database systems · VLDB J. 2013
Distributed systems › fault tolerance
high availability
0.112011
RemusDB: Transparent High Availability for Database Systems · Proc. VLDB Endow. 2011
Software-defined and programmable networks
programmable data plane
0.112016
Flurries: Countless Fine-Grained NFs for Flexible Per-Flow Customization · CoNEXT 2016
Distributed systems
consensus
0.012013
RemusDB: transparent high availability for database systems · VLDB J. 2013
Parallel and multicore computing › parallel programming models
shared-memory parallelization
0.012013
Whose cache line is it anyway?: operating system support for live detection and repair of false sharing · EuroSys 2013
Cloud and datacenter computing
virtualization
0.012013
Split/Merge: System Support for Elastic Execution in Virtual Middleboxes · NSDI 2013

Methods — techniques the papers use, named apart from their topics

rate proportional scheduling · 1.4backpressure · 1.4cgroups · 0.7cgroup · 0.7runtime repair · 0.3live detection · 0.3
YearPublicationVenuePosition
2020 NFVnice: Dynamic Backpressure and Scheduling for NFV Service Chains
abstract
Managing Network Function (NF) service chains requires careful system resource management. We propose NFVnice, a user space NF scheduling and service chain management framework to provide fair, efficient and dynamic resource scheduling capabilities on Network Function Virtualization (NFV) platforms. The NFVnice framework monitors load on a service chain at high frequency (1000Hz) and employs backpressure to shed load early in the service chain, thereby preventing wasted work. Borrowing concepts such as rate proportional scheduling from hardware packet schedulers, CPU shares are computed by accounting for heterogeneous packet processing costs of NFs, I/O, and traffic arrival characteristics. By leveraging cgroups, a user space process scheduling abstraction exposed by the operating system, NFVnice is capable of controlling when network functions should be scheduled. NFVnice improves NF performance by complementing the capabilities of the OS scheduler but without requiring changes to the OS's scheduling mechanisms. Our controlled experiments show that NFVnice provides the appropriate rate-cost proportional fair share of CPU to NFs and significantly improves NF performance (throughput and latency) by reducing wasted work across an NF chain, compared to using the default OS scheduler. NFVnice achieves this even for heterogeneous NFs with vastly different computational costs and for heterogeneous workloads.
Sameer G. Kulkarni, Wei Zhang 0052, Jinho Hwang, Shriram Rajagopalan, K. K. Ramakrishnan, Timothy Wood 0001, Mayutan Arumaithurai, Xiaoming Fu 0001
IEEE/ACM Trans. Netw.4
2017 NFVnice: Dynamic Backpressure and Scheduling for NFV Service Chains
abstract
Managing Network Function (NF) service chains requires careful system resource management. We propose NFVnice, a user space NF scheduling and service chain management framework to provide fair, efficient and dynamic resource scheduling capabilities on Network Function Virtualization (NFV) platforms. The NFVnice framework monitors load on a service chain at high frequency (1000Hz) and employs backpressure to shed load early in the service chain, thereby preventing wasted work. Borrowing concepts such as rate proportional scheduling from hardware packet schedulers, CPU shares are computed by accounting for heterogeneous packet processing costs of NFs, I/O, and traffic arrival characteristics. By leveraging cgroups, a user space process scheduling abstraction exposed by the operating system, NFVnice is capable of controlling when network functions should be scheduled. NFVnice improves NF performance by complementing the capabilities of the OS scheduler but without requiring changes to the OS's scheduling mechanisms. Our controlled experiments show that NFVnice provides the appropriate rate-cost proportional fair share of CPU to NFs and significantly improves NF performance (throughput and loss) by reducing wasted work across an NF chain, compared to using the default OS scheduler. NFVnice achieves this even for heterogeneous NFs with vastly different computational costs and for heterogeneous workloads.
Sameer G. Kulkarni, Wei Zhang 0052, Jinho Hwang, Shriram Rajagopalan, K. K. Ramakrishnan, Timothy Wood 0001, Mayutan Arumaithurai, Xiaoming Fu 0001
SIGCOMM4
2016 Flurries: Countless Fine-Grained NFs for Flexible Per-Flow Customization
abstract
The combination of Network Function Virtualization (NFV) and Software Defined Networking (SDN) allows flows to be flexibly steered through efficient processing pipelines. As deployment of NFV becomes more prevalent, the need to provide fine-grained customization of service chains and flow-level performance guarantees will increase, even as the diversity of Network Functions (NFs) rises. Existing NFV approaches typically route wide classes of traffic through pre-configured service chains. While this aggregation improves efficiency, it prevents flexibly steering and managing performance of flows at a fine granularity.
Wei Zhang 0052, Jinho Hwang, Shriram Rajagopalan, K. K. Ramakrishnan, Timothy Wood 0001
CoNEXT3
2016 Gremlin: Systematic Resilience Testing of Microservices
abstract
Modern Internet applications are being disaggregated into a microservice-based architecture, with services being updated and deployed hundreds of times a day. The accelerated software life cycle and heterogeneity of language runtimes in a single application necessitates a new approach for testing the resiliency of these applications in production infrastructures. We present Gremlin, a framework for systematically testing the failure-handling capabilities of microservices. Gremlin is based on the observation that microservices are loosely coupled and thus rely on standard message-exchange patterns over the network. Gremlin allows the operator to easily design tests and executes them by manipulating inter-service messages at the network layer. We show how to use Gremlin to express common failure scenarios and how developers of an enterprise application were able to discover previously unknown bugs in their failure-handling code without modifying the application.
Victor Heorhiadi, Shriram Rajagopalan, Hani Jamjoom, Michael K. Reiter, Vyas Sekar
ICDCS2
2016 A look at the dynamics of the JavaScript package ecosystem
abstract
The node package manager (npm) serves as the frontend to a large repository of JavaScript-based software packages, which foster the development of currently huge amounts of server-side Node. js and client-side JavaScript applications. In a span of 6 years since its inception, npm has grown to become one of the largest software ecosystems, hosting more than 230, 000 packages, with hundreds of millions of package installations every week. In this paper, we examine the npm ecosystem from two complementary perspectives: 1) we look at package descriptions, the dependencies among them, and download metrics, and 2) we look at the use of npm packages in publicly available applications hosted on GitHub. In both perspectives, we consider historical data, providing us with a unique view on the evolution of the ecosystem. We present analyses that provide insights into the ecosystem's growth and activity, into conflicting measures of package popularity, and into the adoption of package versions over time. These insights help understand the evolution of npm, design better package recommendation engines, and can help developers understand how their packages are being used.
Erik Wittern, Philippe Suter, Shriram Rajagopalan
MSR3
2013 Pico replication: a high availability framework for middleboxes
abstract
Middleboxes are being rearchitected to be service oriented, composable, extensible, and elastic. Yet system-level support for high availability (HA) continues to introduce significant performance overhead. In this paper, we propose Pico Replication (PR), a system-level framework for middleboxes that exploits their flow-centric structure to achieve low overhead, fully customizable HA. Unlike generic (virtual machine level) techniques, PR operates at the flow level. Individual flows can be checkpointed at very high frequencies while the middlebox continues to process other flows. Furthermore, each flow can have its own checkpoint frequency, output buffer and target for backup, enabling rich and diverse policies that balance---per-flow---performance and utilization. PR leverages OpenFlow to provide near instant flow-level failure recovery, by dynamically rerouting a flow's packets to its replication target. We have implemented PR and a flow-based HA policy. In controlled experiments, PR sustains checkpoint frequencies of 1000Hz, an order of magnitude improvement over current VM replication solutions. As a result, PR drastically reduces the overhead on end-to-end latency from 280% to 15.5% and throughput overhead from 99.5% to 3.2%.
Shriram Rajagopalan, Dan Williams 0001, Hani Jamjoom
SoCC1
2013 Whose cache line is it anyway?: operating system support for live detection and repair of false sharing
abstract
As hardware parallelism continues to increase, CPU caches can no longer be considered as a transparent, hardware-level performance optimization. Cache impact on performance, in particular in the face of false sharing, is completely dependent on the software that is executing. To effectively support parallel workloads on cache coherent hardware, the operating system must begin to treat the CPU cache like other shared hardware resources, and manage it appropriately.
Mihir Nanavati, Mark Spear, Nathan Taylor, Shriram Rajagopalan, Dutch T. Meyer, William Aiello, Andy Warfield
EuroSys4
2013 Escape Capsule: Explicit State Is Robust and Scalable
Shriram Rajagopalan, Dan Williams 0001, Hani Jamjoom, Andy Warfield
HotOS1
2013 Split/Merge: System Support for Elastic Execution in Virtual Middleboxes
Shriram Rajagopalan, Dan Williams 0001, Hani Jamjoom, Andy Warfield
NSDI1
2013 RemusDB: transparent high availability for database systems
Umar Farooq Minhas, Shriram Rajagopalan, Brendan Cully, Ashraf Aboulnaga, Kenneth Salem, Andy Warfield
VLDB J.2
2012 SecondSite: disaster tolerance as a service
abstract
This paper describes the design and implementation of SecondSite, a cloud-based service for disaster tolerance. SecondSite extends the Remus virtualization-based high availability system by allowing groups of virtual machines to be replicated across data centers over wide-area Internet links. The goal of the system is to commodify the property of availability, exposing it as a simple tick box when configuring a new virtual machine. To achieve this in the wide area, we have had to tackle the related issues of replication traffic bandwidth, reliable failure detection across geographic regions and traffic redirection over a wide-area network without compromising on transparency and consistency.
Shriram Rajagopalan, Brendan Cully, Ryan O'Connor, Andy Warfield
VEE1
2011 Schema Mediation in Peer Data Management Systems
abstract
Peer Data Management Systems (PDMSs) allow the efficient sharing of data between peers with overlapping sources of information. These sources share data through mappings between peers. In current systems, queries are asked over each peer's local schema and then translated using the mappings between peers. While this allows the data to be accessed uniformly, users lack access to information that is not in their own schemas. In this paper, we propose a light-weight, automatic method to create a mediated schema in a PDMS. Our work benefits PDMSs by allowing access to more data and without unduly stressing the peer's resources or requiring additional resources such as ontologies. We present our system — MePSys, which creates a mediated schema in PDMSs automatically using the existing mappings provided to translate queries. We further discuss how to update the mediated schema in a stable state, i.e. after the system setup period.
Rachel Pottinger, Cody Brown, Shriram Rajagopalan
Int. J. Cooperative Inf. Syst.4
2011 RemusDB: Transparent High Availability for Database Systems
Umar Farooq Minhas, Shriram Rajagopalan, Brendan Cully, Ashraf Aboulnaga, Kenneth Salem, Andy Warfield
Proc. VLDB Endow.2