VLDB 2026 Research / reviewers in the wild / expert
Sanjay G. Rao
dblp:r/SanjayGRao
· DBLP profile ↗
68ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0003-4825-4352ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 51 · 12 since 2021Systems, architecture and hardware · 10 · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Security and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beagle: Auto-tuning Performance Diagnosis for Live Video Streaming
Chandan Bothra, Leonardo Teixeira, Zahaib Akhtar, Sanjay G. Rao, Bruno Ribeiro 0001 |
IWQoS | 4 |
| 2026 | Elastispec: Formalizing Enterprise Firewall Management with Informal and Elastic Specifications
Chenan Wen, Yizhan Qing, Curt P. Jansen, Xiaokang Qiu, Sanjay G. Rao |
SIGCOMM | 5 |
| 2025 | Hattrick: Solving Multi-Class TE using Neural ModelsabstractWhile recent work shows ML-based approaches are a promising alternative to conventional optimization methods for Traffic Engineering (TE), existing research is limited to a single traffic class. In this paper, we present Hattrick, the first ML-based approach for handling multiple traffic classes, a key requirement of cloud and ISP WANs. As part of Hattrick we have developed (i) a novel neural architecture aligned with the sequence of optimization problems in multiclass TE; and (ii) a variant of classical multitask learning methods to deal with the unique challenge of optimizing multiple metrics that have a precedence relationship. Evaluations on a large private WAN and other public datasets show Hattrick outperforms state-of-the-art optimization-based multiclass TE methods by better coping with prediction error - e.g., for GEANT, Hattrick outperforms SWAN by 5.48% to 19.3% across classes when considering the traffic that can be supported 99% of the time. Abd AlRhman AlQiam, Zhuocong Li, Satyajeet Ahuja, Zhaodong Wang, Ying Zhang 0022, Sanjay G. Rao, Bruno Ribeiro 0001, Mohit Tawarmalani |
SIGCOMM | 6 |
| 2025 | Establishing Trust for Using Natural Language for Intent-Based NetworkingabstractTodays enterprise networks wrestle with accommodating an ever-growing number of devices of different types, supporting increasingly demanding applications and ever more complex services, and protecting their users from sophisticated and disrupting cyber threats. In response, a proposed architectural approach for improving network management, referred to as Intent-Based Networking (IBN), has attracted significant attention. It is built on the premise that network operators specify network policies in natural language and the network correctly translates these spoken intents (e.g., policies) into proper device-specific configurations that are then deployed across the network to reliably act on the operators expressed intents. Unfortunately, IBN has not yet fully delivered on its promise of automated, fast, and reliable policy deployment, mainly due to the significant challenges that the reliance on methods from Natural Language Processing (NLP) or more recent techniques from Machine Learning (ML) and Artificial Intelligence (AI) poses for unambiguously and accurately translating the myriad of intents that operators can express in natural language into “trustworthy” device configurations. This paper uses LUMI, a recently designed end-to-end prototype of a system that allows operators “to manage their network by talking to the network”, as an illustrative case study. In particular, we use it to elaborate on the different functionalities such systems should have to realize IBNs vision of automating the fast deployment of policies. At the same time, we leverage LUMI to highlight the extra efforts that are required to ensure that the deployed policies can be entrusted to accurately express and execute the operators original intents. Arthur Selle Jacobs, Ricardo J. Pfitscher, Rafael Hengen Ribeiro, Lisandro Z. Granville, Ronaldo A. Ferreira, Walter Willinger, Sanjay G. Rao |
IEEE Trans. Netw. Serv. Manag. | 7 |
| 2024 | Leo: Online ML-based Traffic Classification at Multi-Terabit Line Rate
Syed Usman Jafri, Sanjay G. Rao, Vishal Shrivastav, Mohit Tawarmalani |
NSDI | 2 |
| 2024 | Transferable Neural WAN TE for Changing TopologiesabstractRecently, researchers have proposed ML-driven traffic engineering (TE) schemes where a neural network model is used to produce TE decisions in lieu of conventional optimization solvers. Unfortunately existing ML-based TE schemes are not explicitly designed to be robust to topology changes that may occur due to WAN evolution, failures or planned maintenance. In this paper, we present HARP, a neural model for TE explicitly capable of handling variations in topology including those not observed in training. HARP is designed with two principles in mind: (i) ensure invariances to natural input transformations (e.g., permutations of node ids, tunnel reordering); and (ii) align neural architecture to the optimization model. Evaluations on a multi-week dataset of a large private WAN show HARP achieves an MLU at most 11% higher than optimal over 98% of the time despite encountering significantly different topologies in testing relative to training data. Further, comparisons with state-of-the-art ML-based TE schemes indicate the importance of the mechanisms introduced by HARP to handle topology variability. Finally, when predicted traffic matrices are provided, HARP outperforms classic optimization solvers achieving a median reduction in MLU of 5 to 10% on the true traffic matrix. Abd AlRhman AlQiam, Yuanjun Yao, Zhaodong Wang, Satyajeet Ahuja, Ying Zhang 0022, Sanjay G. Rao, Bruno Ribeiro 0001, Mohit Tawarmalani |
SIGCOMM | 6 |
| 2023 | Veritas: Answering Causal Queries from Video Streaming TracesabstractIn this paper, we consider the task of answering what-if questions in the context of adaptive bit rate (ABR) video streaming without access to randomized control trials (RCTs) (e.g., no A/B testing) - i.e., given recorded data of an existing deployed system, what would be the performance impact if we changed its design. Our work makes three contributions. First, we show the problem is challenging since data may only be available for a single ABR algorithm without RCTs, and since it is necessary to deal with the cascading effects that past ABR decisions have on future decisions. Next we present Veritas, the first framework that tackles causal reasoning for video streaming without requiring data collected through RCTs. Integral to Veritas is an easy-to-interpret domain-specific ML model that relates the latent stochastic process (intrinsic bandwidth that the video session can achieve) to actual observations (download times), while exploiting counterfactual queries via abduction using the observed TCP states (e.g., congestion window) for blocking the cascading dependencies. Third, we evaluate Veritas's ability to accurately answer a wide range of what-if questions using emulation experiments, and data of real video sessions from Puffer. The results show that (i) Veritas accurately tackles a wider range of what-if questions (e.g., change of buffer size or video quality) that existing approaches cannot; (ii) Veritas without RCT training data achieves performance comparable or better than a recent parallel approach that requires RCT data; and (iii) in many scenarios Veritas achieves accuracy close to an ideal oracle. Chandan Bothra, Jianfei Gao 0001, Sanjay G. Rao, Bruno Ribeiro 0001 |
SIGCOMM | 3 |
| 2023 | Dragonfly: Higher Perceptual Quality For Continuous 360° Video PlaybackabstractWhen streaming 360° video, it is possible to reduce bandwidth by 5× with approaches that spatially segment video into tiles and only stream the user's viewport. Unfortunately, it is difficult to accurately predict a user's viewport even 2--3 seconds before playback. This results in rebuffering events owing to misprediction of a user's viewport or network bandwidth dips, which hurts interactive experience. However, avoiding rebuffering by naively skipping tiles that do not arrive by the playback deadline may lead to incomplete viewports and degraded experience. Ehab Ghabashneh, Chandan Bothra, Ramesh Govindan, Antonio Ortega, Sanjay G. Rao |
SIGCOMM | 5 |
| 2023 | Comparative Synthesis: Learning Near-Optimal Network Designs by QueryabstractWhen managing wide-area networks, network architects must decide how to balance multiple conflicting metrics, and ensure fair allocations to competing traffic while prioritizing critical traffic. The state of practice poses challenges since architects must precisely encode their intent into formal optimization models using abstract notions such as utility functions, and ad-hoc manually tuned knobs. In this paper, we present the first effort to synthesize optimal network designs with indeterminate objectives using an interactive program-synthesis-based approach. We make three contributions. First, we present comparative synthesis, an interactive synthesis framework which produces near-optimal programs (network designs) through two kinds of queries (Validate and Compare), without an objective explicitly given. Second, we develop the first learning algorithm for comparative synthesis in which a voting-guided learner picks the most informative query in each iteration. We present theoretical analysis of the convergence rate of the algorithm. Third, we implemented Net10Q, a system based on our approach, and demonstrate its effectiveness on four real-world network case studies using black-box oracles and simulation experiments, as well as a pilot user study comprising network researchers and practitioners. Both theoretical and experimental results show the promise of our approach. Yanjun Wang 0010, Xiaokang Qiu, Sanjay G. Rao |
Proc. ACM Program. Lang. | 5 |
| 2022 | Flexile: meeting bandwidth objectives almost alwaysabstractWide-area cloud provider networks must support the bandwidth requirements of network traffic despite failures. Existing traffic engineering (TE) schemes perform no better than an approach that optimally routes traffic for each failure scenario. We show that this results in sub-optimal routing decisions that hurt performance, and are potentially unfair to some traffic across scenarios. To tackle this, we develop Flexile, which exploits and discovers opportunities to improve network performance by prioritizing certain traffic in each failure state so that it can meet its bandwidth requirements. Flexile seeks to minimize a desired percentile of loss across all traffic flows, while modeling diverse needs of different traffic classes. To achieve this, Flexile consists of (i) an offline phase that identifies which failure states are critical for each flow; and (ii) an online phase, which on failure allocates bandwidth prioritizing critical flows for that failure state, while also judiciously allocating bandwidth to non-critical flows. For tractability, Flexile's offline phase uses a decomposition algorithm aided with problem-specific accelerations. Evaluations using real topologies, and validated with emulation testbed experiments, show that Flexile outperforms state-of-the-art TE schemes including SWAN, SMORE, and Teavar in reducing flow loss at desired percentiles by 46% or more in the median case. Sanjay G. Rao, Mohit Tawarmalani |
CoNEXT | 3 |
| 2022 | A microscopic view of bursts, buffer contention, and loss in data centersabstractManaging data center networks with low loss requires understanding traffic dynamics at short (millisecond) time-scales, especially the burstiness of traffic, and to what extent bursts contend for switch buffer resources. Yet, monitoring traffic over such intervals is a challenge at scale. Ehab Ghabashneh, Yimeng Zhao, Cristian Lumezanu, Neil Spring, Srikanth Sundaresan, Sanjay G. Rao |
IMC | 6 |
| 2021 | Chimera: exploiting UAS flight path information to optimize heterogeneous data transmissionabstractUnmanned Aerial Systems (UAS) collect and transmit data such as live video and radar images, which have different latency and reliability requirements, over wireless links that exhibit much performance variability. In this paper, we make three contributions. First, we show through a characterization of two real-world UAS flight datasets that there is significant opportunity to optimize data transmission in UAS settings by exploiting knowledge of UAS flight paths. Second, we developed Chimera, a system that taps into this opportunity while transmitting heterogeneous data streams over UAS networks. Chimera learns a model online that relates UAS network throughput to the flight path, and combines the model with a control framework that optimizes transmissions based on long-range throughput prediction. Third, with a combination of emulation and simulation experiments using real-world flight traces, we show Chimera’s effectiveness. Specifically, Chimera reduces penalties related to dropped radar images by 72.4%−100% compared to an algorithm agnostic to flight path information, and achieves an average bitrate of 90.5% compared to an optimal scheme that knows the exact future throughput, with only a minimal increase in radar images dropped. Russell Shirey, Sanjay G. Rao, Shreyas Sundaram |
ICNP | 2 |
| 2021 | Optimizing Quality of Experience for Long-Range UAS Video StreamingabstractThere is much emerging interest in operating Unmanned Aerial Systems (UAS) at long-range distances. Unfortunately, it is unclear whether network connectivity at these distances is sufficient to enable applications with stringent performance needs. In this paper, we consider this question in the context of video streaming, an important UAS use-case. We make three contributions. First, we characterize network data collected from real-world UAS flight tests. Our results show that while dropouts (i.e., extended periods of poor performance) present challenges, there is potential to enable video streaming with modest delays, and correlation of throughput with flight path (both distance and orientation) provides new opportunities. Second, we present Proteus, the first system for video streaming targeted at long-range UAS settings. Proteus is distinguished from Adaptive Bit Rate (ABR) algorithms developed for Internet settings by explicitly accounting for dropouts, and leveraging flight path information. Third, through experiments with real-world flight traces on an emulation test-bed, we show that at distances of around 4 miles, Proteus reduces the fraction of a viewing session encountering rebuffering from 14.33% to 1.57%, while also significantly improving well-accepted composite video delivery metrics. Overall, our results show promise for enabling video streaming with dynamic UAS networks at long-range distances. Russell Shirey, Sanjay G. Rao, Shreyas Sundaram |
IWQoS | 2 |
| 2021 | Hey, Lumi! Using Natural Language for Intent-Based Network Management
Arthur Selle Jacobs, Ricardo J. Pfitscher, Rafael Hengen Ribeiro, Ronaldo A. Ferreira, Lisandro Z. Granville, Walter Willinger, Sanjay G. Rao |
USENIX ATC | 7 |
| 2021 | Karma: Cost-Effective Geo-Replicated Cloud Storage with Dynamic Enforcement of Causal ConsistencyabstractCausal consistency has emerged as an attractive middle-ground to architecting cloud storage systems, as it allows for high availability and low latency, while supporting semantics stronger than eventual consistency. However, causally-consistent cloud storage systems have seen limited deployment in practice. A key factor is these systems employ full replication of all the data in all the data centers (DCs), incurring high cost. A simple extension of current causal systems to support partial replication by clustering DCs into rings incurs availability and latency problems. We propose Karma, the first system to enable causal consistency for partitioned data stores while achieving the cost advantages of partial replication without the availability and latency problems of the simple extension. Our evaluation with 64 servers emulating 8 geo-distributed DCs shows that Karma (i) incurs much lower cost than a fully-replicated causal store (obviously due to the lower replication factor); and (ii) offers higher availability and better performance than the above partial-replication extension at similar costs. Tariq Mahmood 0005, Shankaranarayanan Puzhavakath Narayanan, Sanjay G. Rao, T. N. Vijaykumar, Mithuna Thottethodi |
IEEE Trans. Cloud Comput. | 3 |
| 2020 | Exploring the interplay between CDN caching and video streaming performanceabstractContent Delivery Networks (CDNs) are critical for optimizing Internet video delivery. In this paper, we characterize how CDNs serve video content, and the implications for video performance, especially emerging 4K video streaming. Our work is based on measurements of multiple well known video publishers served by top CDNs. Our results show that (i) video chunks in a session often see heterogeneous behavior in terms of whether they hit in the CDN, and which layer they are served from; and (ii) application throughput can vary significantly even across chunks in a session based on where they are served from. The differences while sensitive to client location and CDN can sometimes be significant enough to impact the viability of 4K streaming. We consider the implications for Adaptive bit rate (ABR) algorithms given they rely on throughput prediction, and are agnostic to whether objects hit in the CDN and where. We evaluate the potential benefits of exposing where a video chunk is served from to the client ABR algorithm in the context of the widely studied model predictive control (MPC) algorithm. Emulation experiments show the approach has the potential to reduce prediction inaccuracies, and enhance video streaming performance. Ehab Ghabashneh, Sanjay G. Rao |
INFOCOM | 2 |
| 2020 | PCF: Provably Resilient Flexible RoutingabstractRecently, traffic engineering mechanisms have been developed that guarantee that a network (cloud provider WAN, or ISP) does not experience congestion under failures. In this paper, we show that existing congestion-free mechanisms, notably FFC, achieve performance far short of the network's intrinsic capability. We propose PCF, a set of novel congestion-free mechanisms to bridge this gap. PCF achieves these goals by better modeling network structure, and by carefully enhancing the flexibility of network response while ensuring that the performance under failures can be tractably modeled. All of PCF's schemes involve relatively light-weight operations on failures, and many of them can be realized using a local proportional routing scheme similar to FFC. We show PCF's effectiveness through formal theoretical results, and empirical experiments over 21 Internet topologies. PCF's schemes provably out-perform FFC, and in practice, can sustain higher throughput than FFC by a factor of 1.11X to 1.5X on average across the topologies, while providing a benefit of 2.6X in some cases. Sanjay G. Rao, Mohit Tawarmalani |
SIGCOMM | 2 |
| 2019 | Learning Network Design Objectives Using A Program Synthesis ApproachabstractWhile the networking community has extensively tackled network design problems using optimization or other techniques (e.g., in areas such as traffic-engineering, and resource allocation), much of this work focuses on efficiently generating designs assuming well-defined objectives. In this paper, we argue that in practice, the objectives of a network design task may not be easy to specify for an architect. We argue for, and present a structured approach where the objectives of a network design task are learnt through iterative interactions with the architect. Our approach is inspired by a programming-by-examples approach that has seen success in the programming languages community. However, conventional program synthesis techniques do not apply because in our context a user can only provide a relative comparison between multiple choices on which one is more desirable, rather than provide an exact output for a given input. We propose a novel comparative synthesis approach to tackle these challenges. We sketch the approach, present promising preliminary results, and discuss future research questions. Yanjun Wang 0010, Xiaokang Qiu, Sanjay G. Rao |
HotNets | 4 |
| 2019 | Cost-aware Multi Data-Center Bulk Transfers in the Cloud from a Customer-Side PerspectiveabstractMany cloud applications (e.g., data backup and replication, video distribution) require dissemination of large volumes of data from a source data-center to multiple geographically distributed data-centers. Given the high costs of wide-area bandwidth, the overall cost of inter-data-center communication is a major concern in such scenarios. While previous works have focused on optimizing the costs of bulk transfer, most of them use the charging models of Internet service providers, typically based on the 95th percentile of bandwidth consumption. However, public Cloud Service Providers (CSP) follow very different models to charge their customers. First, the cost for transmission is flat and depends on the location of the source and receiver data-centers. Second, CSPs offer discounts once customer transfers exceed certain volume thresholds per data-center. We present a systematic framework, CloudMPcast, that exploits these two aspects of cloud pricing schemes. CloudMPcast constructs overlay distribution trees for bulk-data transfer that both optimizes dollar costs of distribution, and ensures end-to-end data transfer times are not affected. CloudMPCast monitors TCP throughputs between data-centers and only proposes alternative trees that respect original transfer times. After an extensive measurement study, the cost savings range from 10 to 60 percent for both Azure and EC2 infrastructures, which potentially translates to millions of dollars a year assuming realistic demands. José Luis García-Dorado, Sanjay G. Rao |
IEEE Trans. Cloud Comput. | 2 |
| 2018 | ACCORD: Automated Change Coordination across Independently Administered Cloud ServicesabstractIt is very hard to coordinate changes across independently administered cloud services in a dependable manner due to several features of its service-oriented architecture: (i) services are often unaware of how a change will affect other services; (ii) impacted services may respond to changes in diverse ways; and (iii) the asynchronous nature of cross-service communication can introduce subtle errors. To tackle these challenges, our major contribution in this paper is a platform for Automated Change COoRDination (ACCORD) across independently administered cloud services. ACCORD (i) provides the abstractions and protocols for services to explicitly register direct dependencies on shared resources and automatically tracks cross-service transitive dependencies; (ii) allows each service to specify custom change coordination policies; and (iii) enables dependable change coordination in several real-world use-cases with minimal overhead to cloud administrators. Tariq Mahmood 0005, Bharath Balasubramanian, Mithuna Thottethodi, Sanjay G. Rao, Kaustubh R. Joshi |
IEEE CLOUD | 4 |
| 2018 | Understanding Video Management Planes
Zahaib Akhtar, Yun Seong Nam, Jessica Chen, Ramesh Govindan, Ethan Katz-Bassett, Sanjay G. Rao, Jibin Zhan, Hui Zhang 0001 |
Internet Measurement Conference | 6 |
| 2018 | Oboe: auto-tuning video ABR algorithms to network conditionsabstractMost content providers are interested in providing good video delivery QoE for all users, not just on average. State-of-the-art ABR algorithms like BOLA and MPC rely on parameters that are sensitive to network conditions, so may perform poorly for some users and/or videos. In this paper, we propose a technique called Oboe to auto-tune these parameters to different network conditions. Oboe pre-computes, for a given ABR algorithm, the best possible parameters for different network conditions, then dynamically adapts the parameters at run-time for the current network conditions. Using testbed experiments, we show that Oboe significantly improves BOLA, MPC, and a commercially deployed ABR. Oboe also betters a recently proposed reinforcement learning based ABR, Pensieve, by 24% on average on a composite QoE metric, in part because it is able to better specialize ABR behavior across different network states. Zahaib Akhtar, Yun Seong Nam, Ramesh Govindan, Sanjay G. Rao, Jessica Chen, Ethan Katz-Bassett, Bruno Ribeiro 0001, Jibin Zhan, Hui Zhang 0001 |
SIGCOMM | 4 |
| 2017 | NutShell: Scalable Whittled Proxy Execution for Low-Latency Web over Cellular NetworksabstractDespite much recent progress, Web page latencies over cellular networks remain much higher than those over wired networks. Proxies that execute Web page JavaScript (JS) and push objects needed by the client can reduce latency. However, a key concern is the scalability of the proxy which must execute JS for many concurrent users. In this paper, we propose to scale the proxies, focusing on a design where the proxy's execution is solely to push the needed objects and the client completely executes the page as normal. Such redundant execution is a simple, yet effective approach to cutting network latencies, which dominate page load delays in cellular settings. We develop whittling, a technique to identify and execute in the proxy only the JS code necessary to identify and push the objects required for the client page load, while skipping other code. Whittling is closely related to program slicing, but with the important distinction that it is acceptable to approximate the program slice in the proxy given the client's complete execution. Experiments with top Alexa Web pages show NutShell can sustain, on average, 27\% more user requests per second than a proxy performing fully redundant execution, while preserving, and sometimes enhancing, the latency benefits. Ashiwan Sivakumar, Yun Seong Nam, Shankaranarayanan Puzhavakath Narayanan, Vijay Gopalakrishnan, Sanjay G. Rao, Subhabrata Sen, Mithuna Thottethodi, T. N. Vijaykumar |
MobiCom | 6 |
| 2017 | Robust Validation of Network Designs under Uncertain Demands and Failures
Yiyang Chang, Sanjay G. Rao, Mohit Tawarmalani |
NSDI | 2 |
| 2017 | Alpaca: Compact Network Policies With Attribute-Encoded AddressesabstractIn enterprise networks, policies (e.g., QoS or security) are often defined based on the categorization of hosts along dimensions, such as the organizational role of the host (faculty versus student) and department (engineering versus sales). While current best practices (virtual local area networks) help when hosts are categorized along a single dimension, policy may often need to be expressed along multiple orthogonal dimensions. In this paper, we make three contributions. First, we argue for attribute-encoded IPs (ACIPs), where the IP address allocation process in enterprises considers attributes of a host along all policy dimensions. ACIPs enable flexible policy specification in a manner that may not otherwise be feasible owing to the limited size of switch rule-tables. Second, we present Alpaca, algorithms for realizing ACIPs under practical constraints of limited-length IP addresses. Our algorithms can be applied to different switch architectures, and we provide bounds on their performance. Third, we demonstrate the importance and viability of ACIPs on data collected from real campus networks. Nanxi Kang, Ori Rottenstreich, Sanjay G. Rao, Jennifer Rexford |
IEEE/ACM Trans. Netw. | 3 |
| 2016 | Reducing Latency Through Page-aware Management of Web Objects by Content Delivery NetworksabstractAs popular web sites turn to content delivery networks (CDNs) for full-site delivery, there is an opportunity to improve the end-user experience by optimizing the delivery of entire web pages, rather than just individual objects. In particular, this paper explores page-structure-aware strategies for placing objects in CDN cache hierarchies. The key idea is that the objects in a web page that have the largest impact on page latency should be served out of the closest or fastest caches in the hierarchy. We present schemes for identifying these objects and develop mechanisms to ensure that they are served with higher priority by the CDN, while balancing traditional CDN concerns such as optimizing the delivery of popular objects and minimizing bandwidth costs. To establish a baseline for evaluating improvements in page latencies, we collect and analyze publicly visible HTTP headers that reveal the distribution of objects among the various levels of a major CDN's cache hierarchy. Through extensive experiments on 83 real-world web pages, we show that latency reductions of over 100 ms can be obtained for 30% of the popular pages, with even larger reductions for the less popular pages. Using anonymized server logs provided by the CDN, we show the feasibility of reducing capacity and staleness misses of critical objects by 60% with minimal increase in overall miss rates, and bandwidth overheads of under 0.02%. Shankaranarayanan Puzhavakath Narayanan, Yun Seong Nam, Ashiwan Sivakumar, Balakrishnan Chandrasekaran 0002, Bruce M. Maggs, Sanjay G. Rao |
SIGMETRICS | 6 |
| 2015 | Alpaca: compact network policies with attribute-carrying addressesabstractIn enterprise networks, policies (e.g., QoS or security) are often defined based on the categorization of hosts along dimensions such as the organizational role of the host (faculty vs. student), and department (engineering vs. sales). While current best practices (VLANs) help when hosts are categorized along a single dimension, policy may often need to be expressed along multiple orthogonal dimensions. In this paper, we make three contributions. First, we argue for Attribute-Carrying IPs (ACIPs), where the IP address allocation process in enterprises considers attributes of a host along all policy dimensions. ACIPs enable flexible policy specification in a manner that may not otherwise be feasible owing to the limited size of switch rule-tables. Second, we present Alpaca, algorithms for realizing ACIPs under practical constraints of limited-length IP addresses. Our algorithms can be applied to different switch architectures, and we provide bounds on their performance. Third, we demonstrate the importance and viability of ACIPs on data collected from real campus networks. Nanxi Kang, Ori Rottenstreich, Sanjay G. Rao, Jennifer Rexford |
CoNEXT | 3 |
| 2015 | Application-specific configuration selection in the cloud: Impact of provider policy and potential of systematic testingabstractProvider policy (e.g., bandwidth rate limits, virtualization, CPU scheduling) can significantly impact application performance in cloud environments. This paper takes a first step towards understanding the impact of provider policy and tackling the complexity of selecting configurations that can best meet the cost and performance requirements of applications. We make three contributions. First, we conduct a measurement study spanning a 19 months period of a wide variety of applications on Amazon EC2 to understand issues involved in configuration selection. Our results show that provider policy can impact communication and computation performance in unpredictable ways. Moreover, seemingly sensible rules of thumb are inappropriate - e.g., VMs with latest hardware or larger VM sizes do not always provide the best performance. Second, we systematically characterize the overheads and resulting benefits of a range of testing strategies for configuration selection. A key focus of our characterization is understanding the overheads of a testing approach in the face of variability in performance across deployments and measurements. Finally, we present configuration pruning and short-listing techniques for minimizing testing overheads. Evaluations on a variety of compute, bandwidth and data intensive applications validate the effectiveness of these techniques in selecting good configurations with low overheads. Mohammad Y. Hajjat, Yiyang Chang, T. S. Eugene Ng, Sanjay G. Rao |
INFOCOM | 5 |
| 2015 | Measuring and characterizing the performance of interactive multi-tier cloud applicationsabstractIn this paper, we conduct a detailed study characterizing the performance of multi-tier web applications on commercial cloud platforms and evaluate the potential of techniques to improve the resilience of such applications to performance fluctuations in the cloud. In contrast to prior works that have studied the performance of individual cloud services or that of compute-intensive scientific applications (e.g., map-reduce based), our study focuses on multi-tier web applications. Our work is conducted in the context of four real-world web applications which we instrumented to collect the overall response time and the time spent in each application tier, for each transaction. Our results indicate that cloud applications undergo frequent periods of poor performance that typically (i) are short-lived lasting a few minutes; and (ii) may be attributed to a small subset of application components, though different subsets may be involved at different times. While geo-distributing applications can help mitigate performance variability, coarse-grained approaches that merely choose the best performing data-center (DC) provide only modest benefits. More significant benefits could accrue, however, if combination of cloud services located across multiple datacenters (DCs) are chosen to serve each request. Mohammad Y. Hajjat, Shankaranarayanan Puzhavakath Narayanan, Ashiwan Sivakumar, Sanjay G. Rao |
LANMAN | 4 |
| 2015 | VIDalizer: An energy efficient video streamerabstractRecent years have witnessed a significant rise in the number, duration and variety of video contents, which contribute to the bulk of internet traffic. With increase in smartphone and tablet users, watching videos on mobile devices has become one of its most popular use cases. These devices live on limited battery energy which is still a major bottleneck and a source of user dissatisfaction during video playback. In this paper we introduce an intermediate framework called VIDalizer for power efficient video delivery to smartphones and tablets. This almost transparent to the user, battery aware framework takes away some of the video processing overhead from the device and intelligently tunes its parameters customized for the mobile device while delivering the video using a novel transport protocol. Our preliminary results show that this framework can significantly reduce energy consumption up to 45%–55% of a mobile device without compromising user experience. Arnab Raha, Subrata Mitra, Vijay Raghunathan, Sanjay G. Rao |
WCNC | 4 |
| 2015 | A flow measurement architecture to preserve application structure
Myungjin Lee, Mohammad Y. Hajjat, Ramana Rao Kompella, Sanjay G. Rao |
Comput. Networks | 4 |
| 2014 | PARCEL: Proxy Assisted BRowsing in Cellular networks for Energy and Latency reductionabstractToday's web page download process is ill suited to cellular networks resulting in high page load times and radio energy usage. While there have been notable prior attempts at tackling the challenge with assistance from proxies (cloud), achieving a responsive and energy efficient browsing experience remains an elusive goal. In this paper, we make a fresh attempt at addressing the challenge by proposing PARCEL. PARCEL splits functionality between the mobile device and the proxy based on their strengths, and in a manner distinct from both traditional browsers and existing cloud-heavy approaches. We conduct extensive evaluations over an operational LTE network using a prototype implementation of PARCEL. Our results show that PARCEL reduces page load times by 49.6%, and radio energy consumption by 65% compared to traditional mobile web browsers. Further, our results show PARCEL continues to perform well under client interactions, owing to its judicious functionality split. Ashiwan Sivakumar, Shankaranarayanan Puzhavakath Narayanan, Vijay Gopalakrishnan, Seungjoon Lee, Sanjay G. Rao, Subhabrata Sen |
CoNEXT | 5 |
| 2014 | Performance Sensitive Replication in Geo-distributed Cloud DatastoresabstractModern web applications face stringent requirements along many dimensions including latency, scalability, and availability. In response, several geo-distributed cloud data stores have emerged in recent years. Customizing data stores to meet application SLAs is challenging given the scale of applications, and their diverse and dynamic workloads. In this paper, we tackle these challenges in the context of quorum-based systems (e.g. Amazon Dynamo, Cassandra), an important class of cloud storage systems. We present models that optimize percentiles of response time under normal operation and under a data-center (DC) failure. Our models consider factors like the geographic spread of users, DC locations, consistency requirements and inter-DC communication costs. We evaluate our models using real-world traces of three applications: Twitter, Wikipedia and Go Walla on a Cassandra cluster deployed in Amazon EC2. Our results confirm the importance and effectiveness of our models, and highlight the benefits of customizing replication in cloud datastores. Shankaranarayanan Puzhavakath Narayanan, Ashiwan Sivakumar, Sanjay G. Rao, Mohit Tawarmalani |
DSN | 3 |
| 2013 | D-tunes: self tuning datastores for geo-distributed interactive applicationsabstractModern internet applications have resulted in users sharing data with each other in an interactive fashion. These applications have very stringent service level agreements (SLAs) which place tight constraints on the performance of the underlying geo-distributed datastores. Deploying these systems in the cloud to meet such constraints is a challenging task, as application architects have to strike an optimal balance among different contrasting objectives such as maintaining consistency between multiple replicas, minimizing access latency and ensuring high availability. Achieving these objectives requires carefully configuring a number of low-level parameters of the datastores, such as the number of replicas, which DCs contain which data, and the underlying consistency protocol parameters. In this work, we adopt a systematic approach where we develop analytical models that capture the performance of a datastore based on application workload and build a system that can automatically configure the datastore for optimal performance. Shankaranarayanan Puzhavakath Narayanan, Ashiwan Sivakumar, Sanjay G. Rao, Mohit Tawarmalani |
SIGCOMM | 3 |
| 2013 | Dynamic Request Splitting for Interactive Cloud ApplicationsabstractDeploying interactive applications in the cloud is a challenge due to the high variability in performance of cloud services. In this paper, we present Dealer — a system that helps geo-distributed, interactive and multi-tier applications meet their stringent requirements on response time despite such variability. Our approach is motivated by the fact that, at any time, only a small number of application components of large multi-tier applications experience poor performance. Dealer continually monitors the performance of individual components and communication latencies between them to build a global view of the application. In serving any given request, Dealer seeks to minimize user response times by picking the best combination of replicas (potentially located across different data centers). While Dealer requires modifications to application code, we show the changes required are modest. Our evaluations on two multi-tier applications using real cloud deployments indicate the 90%ile of response times could be reduced by more than a factor of 6 under natural cloud dynamics. Our results indicate the cost of inter-data-center traffic with Dealer is minor, and that Dealer can in fact be used to reduce the overall operational costs of applications by up to 15% by leveraging the difference in billing plans of cloud instances. Mohammad Y. Hajjat, Shankaranarayanan Puzhavakath Narayanan, David A. Maltz, Sanjay G. Rao, Kunwadee Sripanidkulchai |
IEEE J. Sel. Areas Commun. | 4 |
| 2013 | Characterization of community based-P2P systems and implications for traffic localization
Ruben Torres, Marco Mellia, Maurizio M. Munafò, Sanjay G. Rao |
Peer-to-Peer Netw. Appl. | 4 |
| 2012 | Dealer: application-aware request splitting for interactive cloud applicationsabstractDeploying interactive applications in the cloud is a challenge due to the high variability in performance of cloud services. In this paper, we present Dealer-- a system that helps geo-distributed, interactive and multi-tier applications meet their stringent requirements on response time despite such variability. Our approach is motivated by the fact that, at any time, only a small number of application components of large multi-tier applications experience poor performance. Dealer abstracts application structure as a component graph, with nodes being application components and edges capturing inter-component communication patterns. Dealer continually monitors the performance of individual component replicas and communication latencies between replica pairs. In serving any given user request, Dealer seeks to minimize user response times by picking the best combination of replicas (potentially located across different data-centers). While Dealer does require modifications to application code, we show through integration with two multi-tier applications that the changes required are modest. Our evaluations on two multi-tier applications using real cloud deployments indicate the 90%ile of application response times could be reduced by a factor of 3 under natural cloud dynamics compared to conventional data-center redirection techniques which are agnostic of application structure. Mohammad Y. Hajjat, Shankaranarayanan Puzhavakath Narayanan, David A. Maltz, Sanjay G. Rao, Kunwadee Sripanidkulchai |
CoNEXT | 4 |
| 2012 | Modeling complexity of enterprise routing designabstractEnterprise networks often have complex routing designs given the need to meet a wide set of resiliency, security and routing policies. In this paper, we take the position that minimizing design complexity must be an explicit objective of routing design. We take a first step to this end by presenting a systematic approach for modeling and reasoning about complexity in enterprise routing design. We make three contributions. First, we present a framework for precisely defining objectives of routing design, and for reasoning about how a combination of routing design primitives (e.g. routing instances, static routes, and route filters etc.) will meet the objectives. Second, we show that it is feasible to quantitatively measure the complexity of a routing design by modeling individual routing design primitives, and leveraging configuration complexity metrics [5]. Our approach helps understand how individual design choices made by operators impact configuration complexity, and can enable quantifying design complexity in the absence of configuration files. Third, we validate our model and demonstrate its utility through a longitudinal analysis of the evolution of the routing design of a large campus network over the last three years. We show how our models can enable comparison of the complexity of multiple routing designs that meet the same objective, guide operators in making design choices that can lower complexity, and enable what-if analysis to assess the potential impact of a configuration change on routing design complexity. Xin Sun 0002, Sanjay G. Rao, Geoffrey G. Xie |
CoNEXT | 2 |
| 2012 | The internet-wide impact of P2P traffic localization on ISP profitabilityabstractWe conduct a detailed simulation study to examine how localizing P2P traffic within network boundaries impacts the profitability of an ISP. A distinguishing aspect of our work is the focus on Internet-wide implications, i.e., how adoption of localization within an ISP affects both itself and other ISPs. Our simulations are based on detailed models that estimate inter-autonomous-system (AS) P2P traffic and inter-AS routing, localization models that predict the extent to which P2P traffic is reduced, and pricing models that predict the impact of changes in traffic on the profit of an ISP. We evaluate our models by using a large-scale crawl of BitTorrent containing over 138 million users sharing 2.75 million files. Our results show that the benefits of localization must not be taken for granted. Some of our key findings include: 1) residential ISPs can actually lose money when localization is employed, and some of them will not see increased profitability until other ISPs employ localization; 2) the reduction in costs due to localization will be limited for small ISPs and tends to grow only logarithmically with client population; and 3) some ISPs can better increase profitability through alternate strategies to localization by taking advantage of the business relationships they have with other ISPs. Jeff Seibert, Ruben Torres, Marco Mellia, Maurizio M. Munafò, Cristina Nita-Rotaru, Sanjay G. Rao |
IEEE/ACM Trans. Netw. | 6 |
| 2011 | Dissecting Video Server Selection Strategies in the YouTube CDNabstractIn this paper, we conduct a detailed study of the YouTube CDN with a view to understanding the mechanisms and policies used to determine which data centers users download video from. Our analysis is conducted using week-long datasets simultaneously collected from the edge of five networks - two university campuses and three ISP networks - located in three different countries. We employ state-of-the-art delay-based geolocation techniques to find the geographical location of YouTube servers. A unique aspect of our work is that we perform our analysis on groups of related YouTube flows. This enables us to infer key aspects of the system design that would be difficult to glean by considering individual flows in isolation. Our results reveal that while the RTT between users and data centers plays a role in the video server selection process, a variety of other factors may influence this selection including load-balancing, diurnal effects, variations across DNS servers within a network, limited availability of rarely accessed video, and the need to alleviate hot-spots that may arise due to popular video content. Ruben Torres, Alessandro Finamore, Jin Ryong Kim, Marco Mellia, Maurizio M. Munafò, Sanjay G. Rao |
ICDCS | 6 |
| 2011 | YouTube everywhere: impact of device and infrastructure synergies on user experienceabstractIn this paper we present a complete measurement study that compares YouTube traffic generated by mobile devices (smart-phones,tablets) with traffic generated by common PCs (desktops, notebooks, netbooks). We investigate the users' behavior and correlate it with the system performance. Our measurements are performed using unique data sets which are collected from vantage points in nation-wide ISPs and University campuses from two countries in Europe and the U.S. Alessandro Finamore, Marco Mellia, Maurizio M. Munafò, Ruben Torres, Sanjay G. Rao |
Internet Measurement Conference | 5 |
| 2011 | RelSamp: Preserving application structure in sampled flow measurementsabstractThe Internet has significantly evolved in the number and variety of applications. Network operators need mechanisms to constantly monitor and study these applications. Given modern applications routinely consist of several flows, potentially to many different destinations, existing measurement approaches such as Sampled NetFlow sample only a few flows per application session. To address this issue, in this paper, we introduce RelSamp architecture that implements the notion of related sampling where flows that are part of the same application session are given higher probability. In our evaluation using real traces, we show that RelSamp achieves 5-10x more flows per application session compared to Sampled NetFlow for the same effective number of sampled packets. We also show that behavioral and statistical classification approaches such as BLINC, SVM and C4.5 achieve up to 50% better classification accuracy compared to Sampled NetFlow, while not breaking existing management tasks such as volume estimation. Myungjin Lee, Mohammad Y. Hajjat, Ramana Rao Kompella, Sanjay G. Rao |
INFOCOM | 4 |
| 2011 | A cost-benefit framework for judicious enterprise network redesignabstractRecent works, have shown the benefits of a systematic approach to designing enterprise networks. However, these works are limited to the design of greenfield (newly deployed) networks, or to incremental evolution of existing networks without altering prior design decisions. In this paper, we focus on redesigning existing networks, allowing for changes to existing decisions. Such redesign (migration) may be desirable from the perspective of improved network performance or lower complexity. However, the key challenge is that the costs of redesign may be high due to the presence of complex dependencies between network configurations. We consider these issues in the context of virtual local area networks (VLANs), an important area of enterprise network design. We make three contributions. First, we present a model to capture VLAN redesign costs. Such costs may arise from the need to reconfigure policies (e.g., security policies) to reflect the changes in VLAN design and ensure the continued correctness of the network. Second, we present a framework that enables operators to systematically determine the best strategies to redesign VLANs so the desired performance goals may be achieved while the costs of redesign are minimized. Finally, we demonstrate the effectiveness of our approach using data obtained from a large-scale campus network. Xin Sun 0002, Sanjay G. Rao |
INFOCOM | 2 |
| 2011 | A design for securing data delivery in mesh-based peer-to-peer streaming
Jeff Seibert, Xin Sun 0002, Cristina Nita-Rotaru, Sanjay G. Rao |
Comput. Networks | 4 |
| 2011 | Towards systematic design of enterprise networksabstractEnterprise networks are important, with size and complexity even surpassing carrier networks. Yet, the design of enterprise networks remains ad hoc and poorly understood. In this paper, we show how a systematic design approach can handle two key areas of enterprise design: virtual local area networks (VLANs) and reachability control. We focus on these tasks given their complexity, prevalence, and time-consuming nature. Our contributions are threefold. First, we show how these design tasks may be formulated in terms of network-wide performance, security, and resilience requirements. Our formulations capture the correctness and feasibility constraints on the design, and they model each task as one of optimizing desired criteria subject to the constraints. The optimization criteria may further be customized to meet operator-preferred design strategies. Second, we develop a set of algorithms to solve the problems that we formulate. Third, we demonstrate the feasibility and value of our systematic design approach through validation on a large-scale campus network with hundreds of routers and VLANs. Yu-Wei Eric Sung, Xin Sun 0002, Sanjay G. Rao, Geoffrey G. Xie, David A. Maltz |
IEEE/ACM Trans. Netw. | 3 |
| 2010 | A Systematic Approach for Evolving VLAN DesignsabstractEnterprise networks are large and complex, and their designs must be frequently altered to adapt to changing organizational needs. The process of redesigning and reconfiguring enterprise networks is ad-hoc and error-prone, and configuration errors could cause serious issues such as network outages. In this paper, we take a step towards systematic evolution of network designs in the context of virtual local area networks (VLANs). We focus on VLANs given their importance and prevalence, the frequent need to change VLAN designs, and the time-consuming and error-prone process of making changes. We present algorithms for common design tasks encountered in evolving VLANs such as deciding which VLAN a new host must be assigned to. Our algorithms trade off multiple criteria such as broadcast traffic costs, and costs associated with maintaining spanning trees for each VLAN in the network, while honoring correctness and feasibility constraints on the design. Our algorithms also enable automatic detection of network-wide dependencies which must be factored when reconfiguring VLANs. We evaluate our algorithms on longitudinal snapshots of configuration files of a large-scale operational campus network obtained over a two year period. Our results show that our algorithms can produce significantly better designs than current practice, while avoiding errors and minimizing human work. Our unique data-sets also enable us to characterize VLAN related change activity in real networks, an important contribution in its own right. Xin Sun 0002, Yu-Wei Eric Sung, Sunil Krothapalli, Sanjay G. Rao |
INFOCOM | 4 |
| 2010 | Cloudward bound: planning for beneficial migration of enterprise applications to the cloudabstractIn this paper, we tackle challenges in migrating enterprise services into hybrid cloud-based deployments, where enterprise operations are partly hosted on-premise and partly in the cloud. Such hybrid architectures enable enterprises to benefit from cloud-based architectures, while honoring application performance requirements, and privacy restrictions on what services may be migrated to the cloud. We make several contributions. First, we highlight the complexity inherent in enterprise applications today in terms of their multi-tiered nature, large number of application components, and interdependencies. Second, we have developed a model to explore the benefits of a hybrid migration approach. Our model takes into account enterprise-specific constraints, cost savings, and increased transaction delays and wide-area communication costs that may result from the migration. Evaluations based on real enterprise applications and Azure-based cloud deployments show the benefits of a hybrid migration approach, and the importance of planning which components to migrate. Third, we shed insight on security policies associated with enterprise applications in data centers. We articulate the importance of ensuring assurable reconfiguration of security policies as enterprise applications are migrated to the cloud. We present algorithms to achieve this goal, and demonstrate their efficacy on realistic migration scenarios. Mohammad Y. Hajjat, Xin Sun 0002, Yu-Wei Eric Sung, David A. Maltz, Sanjay G. Rao, Kunwadee Sripanidkulchai, Mohit Tawarmalani |
SIGCOMM | 5 |
| 2010 | Preventing DDoS attacks on internet servers exploiting P2P systems
Xin Sun 0002, Ruben Torres, Sanjay G. Rao |
Comput. Networks | 3 |
| 2010 | On-demand waypoints for live P2P video broadcasting
Aditya Ganjam, Sanjay G. Rao, Kunwadee Sripanidkulchai, Jibin Zhan, Hui Zhang 0001 |
Peer-to-Peer Netw. Appl. | 2 |
| 2010 | On the feasibility of exploiting P2P systems to launch DDoS attacks
Xin Sun 0002, Ruben Torres, Sanjay G. Rao |
Peer-to-Peer Netw. Appl. | 3 |
| 2009 | Extracting Network-Wide Correlated Changes from Longitudinal Configuration Data
Yu-Wei Eric Sung, Sanjay G. Rao, Subhabrata Sen, Stephen Leggett |
PAM | 2 |
| 2009 | Modeling and understanding end-to-end class of service policies in operational networksabstractBusiness and economic considerations are driving the extensive use of service differentiation in Virtual Private Networks (VPNs) operated for business enterprises today. The resulting Class of Service (CoS) designs embed complex policy decisions based on the described priorities of various applications, extent of bandwidth availability, and cost considerations. These inherently complex high-level policies are realized through low-level router configurations. The configuration process is tedious and error-prone given the highly intertwined nature of CoS configuration, the multiple router configurations over which the policies are instantiated, and the complex access control lists (ACLs) involved. Our contributions include (i) a formal approach to modeling CoS policies from router configuration files in a precise manner; (ii) a practical and computationally efficient tool that can determine the CoS treatment received by an arbitrary set of flows across multiple routers; and (iii) a validation of our approach in enabling applications such as troubleshooting, auditing, and visualization of network-wide CoS design, using router configuration data from a cross-section of 150 diverse enterprise VPNs. To our knowledge, this is the first effort aimed at modeling and analyzing CoS configurations. Yu-Wei Eric Sung, Carsten Lund, Mark Lyn, Sanjay G. Rao, Subhabrata Sen |
SIGCOMM | 4 |
| 2009 | Configuration management at massive scale: system design and experienceabstractThe development and maintenance of network device configurations is one of the central challenges faced by large network providers. Current network management systems fail to meet this challenge primarily because of their inability to adapt to rapidly evolving customer and provider-network needs, and because of mismatches between the conceptual models of the tools and the services they must support. In this paper, we present the Presto configuration management system that attempts to address these failings in a comprehensive and flexible way. Developed for and used during the last 5 years within a large ISP network, Presto constructs device-native configurations based on the composition of configlets representing different services or service options. Configlets are compiled by extracting and manipulating data from external systems as directed by the Presto configuration scripting and template language. We outline the configuration management needs of large-scale network providers, introduce the PRESTO system and configuration language, and reflect upon our experiences developing PRESTO configured VPN and VoIP services. In doing so, we describe how PRESTO promotes healthy configuration management practices. William Enck, Thomas Moyer, Patrick D. McDaniel, Subhabrata Sen, Panagiotis Sebos, Sylke Spoerel, Albert G. Greenberg, Yu-Wei Eric Sung, Sanjay G. Rao, William Aiello |
IEEE J. Sel. Areas Commun. | 9 |
| 2008 | Towards systematic design of enterprise networksabstractEnterprise networks are important, with size and complexity even surpassing carrier networks. Yet, the design of enterprise networks is ad-hoc and poorly understood. In this paper, we show how a systematic design approach can handle two key areas of enterprise design: virtual local area networks (VLANs) and reachability control. We focus on these tasks given their complexity, prevalence, and time-consuming nature. Our contributions are three-fold. First, we show how these design tasks may be formulated in terms of network-wide performance, security, and resilience requirements. Our formulations capture the correctness and feasibility constraints on the design, and they model each task as one of optimizing desired criteria subject to the constraints. The optimization criteria may further be customized to meet operator-preferred design strategies. Second, we develop a set of algorithms to solve the problems that we formulate. Third, we demonstrate the feasibility and value of our systematic design approach through validation on a large-scale campus network with hundreds of routers and VLANs. Yu-Wei Eric Sung, Sanjay G. Rao, Geoffrey G. Xie, David A. Maltz |
CoNEXT | 2 |
| 2008 | Opportunities and Challenges of Peer-to-Peer Internet Video BroadcastabstractThere have been tremendous efforts and many technical innovations in supporting real-time video streaming in the past two decades, but cost-effective large-scale video broadcast has remained an elusive goal. Internet protocol (IP) multicast represented an earlier attempt to tackle this problem but failed largely due to concerns regarding scalability, deployment, and support for higher level functionality. Recently, peer-to-peer based broadcast has emerged as a promising technique, which has been shown to be cost effective and easy to deploy. This new paradigm brings a number of unique advantages such as scalability, resilience, and effectiveness in coping with dynamics and heterogeneity. While peer-to-peer applications such as file download and voice-over-IP have gained tremendous popularity, video broadcast is still in its early stages, and its full potential remains to be seen. This paper reviews the state-of-the-art of peer-to-peer Internet video broadcast technologies. We describe the basic taxonomy of peer-to-peer broadcast and summarize the major issues associated with the design of broadcast overlays. We closely examine two approaches - tree-based and data-driven - and discuss their fundamental tradeoff and potential for large-scale deployment. Lastly, we outline the key challenges and open problems and highlight possible avenues for future directions. Jiangchuan Liu, Sanjay G. Rao, Bo Li 0001, Hui Zhang 0001 |
Proc. IEEE | 2 |
| 2007 | Enabling Confidentiality of Data Delivery in an Overlay Broadcasting SystemabstractMost prior work on the use of key management algorithms to enable confidentiality of video delivery has been conducted in the context of IP Multicast. In this paper, we consider the unique challenges and opportunities of integrating key management algorithms in an overlay multicast system. We conduct a systematic and extensive performance evaluation of strategies for key dissemination in the context of an operational overlay broadcasting system on the Planetlab testbed using real traces of join/leave dynamics. We show that leveraging TCP in each hop of the overlay dissemination structure can significantly simplify reliable key dissemination. The performance can be further enhanced if convergence properties of overlays are considered. We show that using two specialized dissemination structures, one for data and one for keys, potentially achieves low overhead for key dissemination without sacrificing application performance. To our knowledge, this is the first paper to study key management schemes in an overlay context using real implementation and Internet experiments and the first to consider issues in resilient key dissemination with overlays. Ruben Torres, Xin Sun 0002, Aaron Walters, Cristina Nita-Rotaru, Sanjay G. Rao |
INFOCOM | 5 |
| 2007 | Configuration Management at Massive Scale: System Design and Experience
William Enck, Patrick D. McDaniel, Subhabrata Sen, Panagiotis Sebos, Sylke Spoerel, Albert G. Greenberg, Sanjay G. Rao, William Aiello |
USENIX ATC | 7 |
| 2007 | Enabling Confidentiality of Data Delivery in an Overlay Broadcasting SystemabstractIn this paper, we present an extensive study of key dissemination schemes in an overlay multicast context, and the first to involve actual implementation, real traces, and performance in Internet environments. Given that rekey traffic has stronger resilience requirements and is burstier than data traffic, we consider whether data and keys must be distributed using the same overlay or using two separate dissemination structures. Our key findings are: (i) a coupled architecture is effective in achieving resilient key dissemination. Using TCP in each hop of the dissemination structure (an opportunity unique to overlays) is effective in achieving resiliency in end-to-end key delivery. The performance can be further enhanced if convergence properties of overlays are considered; and (ii) a coupled architecture optimized for data delivery has high overheads, while a coupled architecture optimized for key delivery may not honor access bandwidth constraints of nodes. Distributing data and keys using separate overlays achieves low overhead for key dissemination while honoring access bandwidth constraints of nodes. Ruben Torres, Xin Sun 0002, Aaron Walters, Cristina Nita-Rotaru, Sanjay G. Rao |
IEEE J. Sel. Areas Commun. | 5 |
| 2007 | Enabling Contribution Awareness in an Overlay Broadcasting SystemabstractWe consider the design of bandwidth-demanding broadcasting applications using overlays in environments characterized by hosts with limited and asymmetric bandwidth, and significant heterogeneity in upload bandwidth. Such environments are critical to consider to extend the applicability of overlay multicast to mainstream Internet environments where insufficient bandwidth exists to support all hosts, but have not received adequate attention from the research community. We leverage the multitree framework and design heuristics to enable it to consider host contribution and operate in bandwidth-scarce environments. Our extensions seek to simultaneously achieve good utilization of system resources, performance to hosts commensurate to their contributions, and consistent performance. We have implemented the system and conducted an Internet evaluation on PlanetLab using real traces from previous operational deployments of an overlay broadcasting system. Our results indicate for these traces, our heuristics can improve the performance of high contributors by 10-240% and facilitate equitable bandwidth distribution among hosts with similar contributions. Yu-Wei Eric Sung, Michael A. Bishop 0001, Sanjay G. Rao |
IEEE Trans. Multim. | 3 |
| 2006 | Considering Priority in Overlay Multicast Protocols Under Heterogeneous EnvironmentsabstractAbstract — Hosts participating in overlay multicast applications have a wide range of heterogeneity in bandwidth and participation characteristics. In this paper, we highlight and show the need to systematically consider prioritization as a key criterion in the design of protocols for overlay multicast. We identify trade-offs in the design of prioritization heuristics in two important contexts. The first part of the paper considers prioritization strategies in the context of heterogeneity in node outgoing bandwidth and node stay time durations, and a lack of correlation between the two dimensions. The second part of the paper considers bandwidth allocation and prioritization policies with multi-tree data delivery in environments with heterogeneity in outgoing bandwidth and a certain degree of altruistic behavior. We conduct a systematic study of the trade-offs using both real trace data, and sensitivity studies using synthetic workloads. To the best of our knowledge, this is the first work to identify and study these trade-offs, and to demonstrate the potential benefits of the resulting prioritization heuristics. I. Michael A. Bishop 0001, Sanjay G. Rao, Kunwadee Sripanidkulchai |
INFOCOM | 2 |
| 2006 | Enabling contribution awareness in an overlay broadcasting systemabstractWe consider the design of bandwidth-demanding broadcasting applications using overlays in environments characterized by hosts with limited and asymmetric bandwidth, and significant heterogeneity in outgoing bandwidth. Such environments are critical to consider to extend the applicability of overlay multicast to mainstream Internet environments where insufficient bandwidth exists to support all hosts, but have not received adequate attention from the research community. We leverage the multi-tree framework and design heuristics to enable it to consider host contribution and operate in bandwidth-scarce environments. Our extensions seek to simultaneously achieve good utilization of system resources, performance to hosts commensurate to their contributions, and consistent performance. We have implemented the system and conducted an Internet evaluation on Planet-Lab using real traces from previous operational deployments of an overlay broadcasting system. Our results indicate for these traces, our heuristics can improve the performance of high contributors by 10-240% and facilitate equitable bandwidth distribution among hosts with similar contributions. Yu-Wei Eric Sung, Michael A. Bishop 0001, Sanjay G. Rao |
SIGCOMM | 3 |
| 2004 | Early Experience with an Internet Broadcast System Based on Overlay Multicast
Yang-Hua Chu, Aditya Ganjam, T. S. Eugene Ng, Sanjay G. Rao, Kunwadee Sripanidkulchai, Jibin Zhan, Hui Zhang 0001 |
USENIX ATC, General Track | 4 |
| 2003 | Measurement-Based Optimization Techniques for Bandwidth-Demanding Peer-to-Peer SystemsabstractMeasurement-based optimization is one important strategy to improve the performance of bandwidth-demanding peer-to-peer systems. However, to date, we have little quantitative knowledge of how well basic lightweight measurement-based techniques such as RTT probing, 10KB TCP probing, and bottleneck bandwidth probing may work in practice in the peer-to-peer environment. By conducting trace-based analyses, we find that the basic techniques can help achieve 40 to 50% optimal performance. To deepen our understanding, we analyze some of the intrinsic properties of these techniques. Our analyses reveal the inherent difficulty of the peer selection problem due to the extreme heterogeneity in the peer-to-peer environment, and that the basic techniques are limited because their primary strength lies in eliminating the low-performance peers rather than reliably identifying the best-performing one. However, our analyses also reveal two key insights that can potentially be exploited by applications. First, for adaptive applications that can continuously change communication peers, the basic techniques are highly effective in guiding the adaption process. In our experiments, typically an 80% optimal peer can be found by trying less than 5 candidates. Secondly, we find that the basic techniques are highly complementary and can potentially be combined to better identify a high-performance peer, thus even applications that cannot adapt may benefit. Using media file sharing and overlay multicast streaming as case studies, we have systematically experimented with several simple combined peer selection techniques. Our results show that for the nonadaptive media file sharing application, a simple combined technique can boost performance to 60% optimal. In contrast, for the continuously adaptive overlay multicast application, we find that a basic technique with even low-fidelity network information is sufficient to ensure good performance. We believe our findings will help guide the future designs of high-performance peer-to-peer systems. T. S. Eugene Ng, Yang-Hua Chu, Sanjay G. Rao, Kunwadee Sripanidkulchai, Hui Zhang 0001 |
INFOCOM | 3 |
| 2002 | A case for end system multicastabstractThe conventional wisdom has been that Internet protocol (IP) is the natural protocol layer for implementing multicast related functionality. However, more than a decade after its initial proposal, IP multicast is still plagued with concerns pertaining to scalability, network management, deployment, and support for higher layer functionality such as error, flow, and congestion control. We explore an alternative architecture that we term end system multicast, where end systems implement all multicast related functionality including membership management and packet replication. This shifting of multicast support from routers to end systems has the potential to address most problems associated with IP multicast. However, the key concern is the performance penalty associated with such a model. In particular, end system multicast introduces duplicate packets on physical links and incurs larger end-to-end delays than IP multicast. We study these performance concerns in the context of the Narada protocol. In Narada, end systems self-organize into an overlay structure using a fully distributed protocol. Further, end systems attempt to optimize the efficiency of the overlay by adapting to network dynamics and by considering application level performance. We present details of Narada and evaluate it using both simulation and Internet experiments. Our results indicate that the performance penalties are low both from the application and the network perspectives. We believe the potential benefits of transferring multicast functionality from end systems to routers significantly outweigh the performance penalty incurred. Yang-Hua Chu, Sanjay G. Rao, Srinivasan Seshan, Hui Zhang 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2001 | Enabling conferencing applications on the internet using an overlay muilticast architectureabstractIn response to the serious scalability and deployment concerns with IP Multicast, we and other researchers have advocated an alternate architecture for supporting group communication applications over the Internet where all multicast functionality is pushed to the edge. We refer to such an architecture as End System Multicast. While End System Multicast has several potential advantages, a key concern is the performance penalty associated with such a design. While preliminary simulation results conducted in static environments are promising, they have yet to consider the challenging performance requirements of real world applications in a dynamic and heterogeneous Internet environment.In this paper, we explore how Internet environments and application requirements can influence End System Multicast design. We explore these issues in the context of audio and video conferencing: an important class of applications with stringent performance requirements. We conduct an extensive evaluation study of schemes for constructing overlay networks on a wide-area test-bed of about twenty hosts distributed around the Internet. Our results demonstrate that it is important to adapt to both latency and bandwidth while constructing overlays optimized for conferencing applications. Further, when relatively simple techniques are incorporated into current self-organizing protocols to enable dynamic adaptation to latency and bandwidth, the performance benefits are significant. Our results indicate that End System Multicast is a promising architecture for enabling performance-demanding conferencing applications in a dynamic and heterogeneous Internet environment. Sanjay G. Rao, Srinivasan Seshan, Hui Zhang 0001 |
SIGCOMM | 2 |
| 2000 | A case for end system multicastabstractThe conventional wisdom has been that IP is the natural protocol layer for implementing multicast related functionality. However, ten years after its initial proposal, IP Multicast is still plagued with concerns pertaining to scalability, network management, deployment and support for higher layer functionality such as error, flow and congestion control. In this paper, we explore an alternative architecture for small and sparse groups, where end systems implement all multicast related functionality including membership management and packet replication. We call such a scheme End System Multicast. This shifting of multicast support from routers to end systems has the potential to address most problems associated with IP Multicast. However, the key concern is the performance penalty associated with such a model. In particular, End System Multicast introduces duplicate packets on physical links and incurs larger end-to-end delay than IP Multicast. In this paper, we study this question in the context of the Narada protocol. In Narada, end systems self-organize into an overlay structure using a fully distributed protocol. In addition, Narada attempts to optimize the efficiency of the overlay based on end-to-end measurements. We present details of Narada and evaluate it using both simulation and Internet experiments. Preliminary results are encouraging. In most simulations and Internet experiments, the delay and bandwidth penalty are low. We believe the potential benefits of repartitioning multicast functionality between end systems and routers significantly outweigh the performance penalty incurred. Yang-Hua Chu, Sanjay G. Rao, Hui Zhang 0001 |
SIGMETRICS | 2 |
| 1999 | Fast Techniques for the Optimal Smoothing of Stored Video
Sanjay G. Rao |
Multim. Syst. | 1 |
| 1993 | Run-Time Support and Storage Management for Memory-Mapped Persistent ObjectsabstractThe authors present the design and implementation of a persistent store called SPOMS. SPOMS is a runtime system that provides a store for persistent objects and is language independent. The objects are created via calls to SPOMS, and, when used, SPOMS directly maps them into the spaces of all requesting processes. The objects are stored in native format and are concurrently sharable. The store can handle distributed applications. The system uses the concept of a compiled class to manage persistent objects. The compiled class is a template that is used to create and store objects in a language independent manner and so that object reuse can occur without recompilation or relinking of an application that uses it. A prototype of SPOMS has been built on top of the Mach operating system. The motivations, the design, and implementation details are presented. Related and future work are discussed.> Bruce Millard, Partha Dasgupta, Sanjay G. Rao, Ravindra Kuramkote |
ICDCS | 3 |