Raja R. Sambasivan

dblp:99/5200 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
4since 2021 · last 2026
0000-0001-7940-403XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Computer networks · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
2 papers
Routing and switching · 84% Network management and operations · 16%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Performance modeling and evaluation · 48% Storage systems · 48% Distributed systems · 4%
Software engineering, system software, and programming languages
1 paper
Services computing and microservices · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Services computing and microservices
microservice architecture
0.712023
Lifting the veil on Meta's microservice architecture: Analyses of topology and request workflows · USENIX ATC 2023
Routing and switching › inter-domain routing
BGP
0.312017
Bootstrapping evolvability for inter-domain routing with D-BGP · SIGCOMM 2017
Routing and switching
inter-domain routing
0.312017
Bootstrapping evolvability for inter-domain routing with D-BGP · SIGCOMM 2017
Storage systems
distributed storage
0.222010
A Transparently-Scalable Metadata Service for the Ursa Minor Storage System · USENIX ATC 2010
Ursa Minor: Versatile Cluster-based Storage · FAST 2005
Network management and operations › performance management
performance diagnosis
0.112011
Diagnosing Performance Changes by Comparing Request Flows · NSDI 2011
Storage systems › file systems › distributed file system
metadata service
0.112010
A Transparently-Scalable Metadata Service for the Ursa Minor Storage System · USENIX ATC 2010
Performance modeling and evaluation
workload characterization
0.122007
Modeling the relative fitness of storage · SIGMETRICS 2007
//TRACE: Parallel Trace Replay with Approximate Causal Events · FAST 2007
Performance modeling and evaluation
storage performance modeling
0.112007
Modeling the relative fitness of storage · SIGMETRICS 2007
Performance modeling and evaluation › simulation
trace replay
0.112007
//TRACE: Parallel Trace Replay with Approximate Causal Events · FAST 2007
Storage systems › distributed storage
storage cluster
0.112005
Ursa Minor: Versatile Cluster-based Storage · FAST 2005
Distributed systems
distributed coordination
0.012010
A Transparently-Scalable Metadata Service for the Ursa Minor Storage System · USENIX ATC 2010
Storage systems
metadata management
0.012010
A Transparently-Scalable Metadata Service for the Ursa Minor Storage System · USENIX ATC 2010

Methods — techniques the papers use, named apart from their topics

empirical study · 0.7user study · 0.2causal events · 0.1black-box modeling · 0.1versatility · 0.1
YearPublicationVenuePosition
2026 Dynamic read \u0026 write optimization with TurtleKV
Tony Astolfi, Vidya Silai, Darby Huye, Raja R. Sambasivan, Johes Bater
Proc. VLDB Endow.5
2024 Systemizing and Mitigating Topological Inconsistencies in Alibaba's Microservice Call-graph Datasets
abstract
Alibaba's 2021 and 2022 microservice datasets are the only publicly available sources of request-workflow traces from a large-scale microservice deployment. They have the potential to strongly influence future research as they provide much-needed visibility into industrial microservices' characteristics. We conduct the first systematic analyses of both datasets to help facilitate their use by the community. We find that the 2021 dataset contains numerous inconsistencies preventing accurate reconstruction of full trace topologies. The 2022 dataset also suffers from inconsistencies, but at a much lower rate. Tools that strictly follow Alibaba's specs for constructing traces from these datasets will silently ignore these inconsistencies, misinforming researchers by creating traces of the wrong sizes and shapes. Tools that discard traces with inconsistencies will discard many traces. We present Casper, a construction method that uses redundancies in the datasets to sidestep the inconsistencies. Compared to an approach that discards traces with inconsistencies, Casper accurately reconstructs an additional 25.5% of traces in the 2021 dataset (going from 58.32% to 83.82%) and an additional 12.18% in the 2022 dataset (going from 86.42% to 98.6%).
Darby Huye, Raja R. Sambasivan
ICPE3
2023 Lifting the veil on Meta's microservice architecture: Analyses of topology and request workflows
Darby Huye, Yuri Shkuro, Raja R. Sambasivan
USENIX ATC3
2021 Automating instrumentation choices for performance problems in distributed applications with VAIF
abstract
Developers use logs to diagnose performance problems in distributed applications. However, it is difficult to know a priori where logs are needed and what information in them is needed to help diagnose problems that may occur in the future. We present the Variance-driven Automated Instrumentation Framework (VAIF), which runs alongside distributed applications. In response to newly-observed performance problems, VAIF automatically searches the space of possible instrumentation choices to enable the logs needed to help diagnose them. To work, VAIF combines distributed tracing (an enhanced form of logging) with insights about how response-time variance can be decomposed on the critical-path portions of requests' traces. We evaluate VAIF by using it to localize performance problems in OpenStack and HDFS. We show that VAIF can localize problems related to slow code paths, resource contention, and problematic third-party code while enabling only 3-34% of the total tracing instrumentation.
Mert Toslali, Emre Ates, Alex Ellis, Darby Huye, Samantha Puterman, Ayse K. Coskun, Raja R. Sambasivan
SoCC9
2019 D3N: A multi-layer cache for the rest of us
abstract
Current caching methods for improving the performance of big-data jobs assume high (e.g., full bi-section) bandwidth; however many enterprise data centers and co-location facilities have large network imbalances due to over-subscription and incremental networking upgrades. We describe D3N, a multi-layer cooperative caching architecture that mitigates network imbalances by caching data on the access side of each layer of a hierarchical network topology, adaptively adjusting cache sizes of each layer based on observed workload patterns and network congestion. We have added (and submitted upstream) a 2-layer D3N cache to the Ceph RADOS Gateway; read bandwidth achieves the 5GB/s speed of our SSDs, and we show that it substantially improves big-data job performance while reducing network traffic.
Emine Ugur Kaynar, Mania Abdi, Mohammad Hossein Hajkazemi, Ata Turk, Raja R. Sambasivan, Larry Rudolph, Peter Desnoyers, Orran Krieger
IEEE BigData5
2019 An automated, cross-layer instrumentation framework for diagnosing performance problems in distributed applications
abstract
Diagnosing performance problems in distributed applications is extremely challenging. A significant reason is that it is hard to know where to place instrumentation a priori to help diagnose problems that may occur in the future. We present the vision of an automated instrumentation framework, Pythia, that runs alongside deployed distributed applications. In response to a newly-observed performance problem, Pythia searches the space of possible instrumentation choices to enable the instrumentation needed to help diagnose it. Our vision for Pythia builds on workflow-centric tracing, which records the order and timing of how requests are processed within and among a distributed application's nodes (i.e., records their workflows). It uses the key insight that localizing the sources high performance variation within the workflows of requests that are expected to perform similarly gives insight into where additional instrumentation is needed.
Emre Ates, Lily Sturmann, Mert Toslali, Orran Krieger, Richard Megginson, Ayse K. Coskun, Raja R. Sambasivan
SoCC7
2017 Bootstrapping evolvability for inter-domain routing with D-BGP
abstract
The Internet's inter-domain routing infrastructure, provided today by BGP, is extremely rigid and does not facilitate the introduction of new inter-domain routing protocols. This rigidity has made it incredibly difficult to widely deploy critical fixes to BGP. It has also depressed ASes' ability to sell value-added services or replace BGP entirely with a more sophisticated protocol. Even if operators undertook the significant effort needed to fix or replace BGP, it is likely the next protocol will be just as difficult to change or evolve. To help, this paper identifies two features needed in the routing infrastructure (i.e., within any inter-domain routing protocol) to facilitate evolution to new protocols. To understand their utility, it presents D-BGP, a version of BGP that incorporates them.
Raja R. Sambasivan, David Tran-Lam, Aditya Akella, Peter Steenkiste
SIGCOMM1
2016 Principled workflow-centric tracing of distributed systems
abstract
Workflow-centric tracing captures the workflow of causally-related events (e.g., work done to process a request) within and among the components of a distributed system. As distributed systems grow in scale and complexity, such tracing is becoming a critical tool for understanding distributed system behavior. Yet, there is a fundamental lack of clarity about how such infrastructures should be designed to provide maximum benefit for important management tasks, such as resource accounting and diagnosis. Without research into this important issue, there is a danger that workflow-centric tracing will not reach its full potential. To help, this paper distills the design space of workflow-centric tracing and describes key design choices that can help or hinder a tracing infrastructures utility for important tasks. Our design space and the design choices we suggest are based on our experiences developing several previous workflow-centric tracing infrastructures.
Raja R. Sambasivan, Ilari Shafer, Jonathan Mace, Benjamin H. Sigelman, Rodrigo Fonseca, Gregory R. Ganger
SoCC1
2015 Bootstrapping Evolvability for Inter-Domain Routing
abstract
It is extremely difficult to deploy newinter-domain routing protocols in today's Internet. As a result, the Internet's baseline protocol for connectivity, BGP, has remained largely unchanged, despite known significant flaws. The difficulty of deploying new protocols has also depressed opportunities for (currently commoditized) transit providers to provide value-added routing services. To help, we identify the key deployment models under which new protocols are introduced and the requirements each poses for enabling their usage goals. Based on these requirements, we argue for two modifications to BGP that will greatly improve support for new routing protocols.
Raja R. Sambasivan, David Tran-Lam, Aditya Akella, Peter Steenkiste
HotNets1
2013 Specialized Storage for Big Numeric Time Series
Ilari Shafer, Raja R. Sambasivan, Anthony Rowe 0001, Gregory R. Ganger
HotStorage2
2013 Visualizing Request-Flow Comparison to Aid Performance Diagnosis in Distributed Systems
abstract
Distributed systems are complex to develop and administer, and performance problem diagnosis is particularly challenging. When performance degrades, the problem might be in any of the system's many components or could be a result of poor interactions among them. Recent research efforts have created tools that automatically localize the problem to a small number of potential culprits, but research is needed to understand what visualization techniques work best for helping distributed systems developers understand and explore their results. This paper compares the relative merits of three well-known visualization approaches (side-by-side, diff, and animation) in the context of presenting the results of one proven automated localization technique called request-flow comparison. Via a 26-person user study, which included real distributed systems developers, we identify the unique benefits that each approach provides for different problem types and usage modes.
Raja R. Sambasivan, Ilari Shafer, Michelle L. Mazurek, Gregory R. Ganger
IEEE Trans. Vis. Comput. Graph.1
2011 Diagnosing Performance Changes by Comparing Request Flows
Raja R. Sambasivan, Alice X. Zheng, Michael De Rosa, Elie Krevat, Spencer Whitman, Michael Stroucken, Lianghong Xu, Gregory R. Ganger
NSDI1
2010 A Transparently-Scalable Metadata Service for the Ursa Minor Storage System
Shafeeq Sinnamohideen, Raja R. Sambasivan, James Hendricks, Likun Liu, Gregory R. Ganger
USENIX ATC2
2007 //TRACE: Parallel Trace Replay with Approximate Causal Events
Michael P. Mesnier, Matthew Wachs, Raja R. Sambasivan, Julio López 0002, James Hendricks, Gregory R. Ganger, David R. O'Hallaron
FAST3
2007 Modeling the relative fitness of storage
abstract
Relative fitness is a new black-box approach to modeling the performance of storage devices. In contrast with an absolute model that predicts the performance of a workload on a given storage device, a relative fitness model predicts performance differences between a pair of devices. There are two primary advantages to this approach. First, because are lative fitness model is constructed for a device pair, the application-device feedback of a closed workload can be captured (e.g., how the I/O arrival rate changes as the workload moves from device A to device B). Second, a relative fitness model allows performance and resource utilization to be used in place of workload characteristics. This is beneficial when workload characteristics are difficult to obtain or concisely express (e.g., rather than describe the spatio-temporal characteristics of a workload, one could use the observed cache behavior of device A to help predict the performance of B.
Michael P. Mesnier, Matthew Wachs, Raja R. Sambasivan, Alice X. Zheng, Gregory R. Ganger
SIGMETRICS3
2005 Ursa Minor: Versatile Cluster-based Storage
Michael Abd-El-Malek, William V. Courtright II, Chuck Cranor, Gregory R. Ganger, James Hendricks, Andrew J. Klosterman, Michael P. Mesnier, Manish Prasad, Brandon Salmon, Raja R. Sambasivan, Shafeeq Sinnamohideen, John D. Strunk, Eno Thereska, Matthew Wachs, Jay J. Wylie
FAST10
2005 Replication policies for layered clustering of NFS servers
abstract
Layered clustering offers cluster-like load balancing for unmodified NFS or CIFS servers. Read requests sent to a busy server can be offloaded to other servers holding replicas of the accessed files. This paper explores a key design question for this approach; which files should be replicated? We find that the popular policy of replicating read-only files offers little benefit. A policy that replicates read-only portions of read-mostly files, however, implicitly coordinates with client cache invalidations and thereby allows almost all read operations to be offloaded. In a read-heavy trace, 75% of all operations and 52% of all data transfers can be offloaded.
Raja R. Sambasivan, Andrew J. Klosterman, Gregory R. Ganger
MASCOTS1