Aditya Akella

dblp:a/AdityaAkella · DBLP profile ↗
← Back
157ranked-venue papers
15as first author
48since 2021 · last 2026
0000-0002-5920-170XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 101 · 11 first-author · 22 since 2021Systems, architecture and hardware · 28 · 4 first-author · 12 since 2021Software engineering, systems software and programming languages · 17 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Security and privacy · 3Databases, data management, data science and information retrieval · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Canopy: Property-Driven Learning for Congestion Control
abstract
Learning-based congestion controllers offer better adaptability compared to traditional heuristics. However, the unreliability of learning techniques can cause learning-based controllers to behave poorly, creating a need for formal guarantees. While methods for formally verifying learned congestion controllers exist, these methods offer binary feedback that cannot optimize the controller toward better behavior. We improve this state-of-the-art via Canopy, a new property-driven framework that integrates learning with formal reasoning in the learning loop. Canopy uses novel quantitative certification with an abstract interpreter to guide the training process, rewarding models, and evaluating robust and safe model performance on worst-case inputs. Our evaluation demonstrates that unlike state-of-the-art learned controllers, Canopy-trained controllers provide both adaptability and worst-case reliability across a range of network conditions.
Divyanshu Saxena, Rohit Dwivedula, Kshiteej Mahajan, Swarat Chaudhuri, Aditya Akella
EuroSys6
2026 Reforge: Low-Latency Distributed GNN Serving with Selective Embedding Recomputation
Geon-Woo Kim, Donghyun Kim 0002, Jeongyoon Moon, Henry Liu, Tarannum Khan, Anand Iyer, Daehyeok Kim, Aditya Akella
IPDPS8
2026 SYMPHONY: Enabling Compute-Memory Disaggregation in LLM Serving Systems
Bodun Hu, Anyong Mao, Aditya Akella, Shivaram Venkataraman
NSDI4
2026 UNUM: A New Framework for Network Control
Nihal Sharma, Debajit Chakraborty, Jeffrey Zhou, Aditya Akella, Sanjay Shakkottai
NSDI6
2026 Express Lane to Efficiency and Reliability: Multi-Dimensional Control in Meta's Express Backbone Network
Faisal Iqbal, Vitaly Neganov, Brian Chang, Rong Rong, Yuanjun Yao, Marek Denis, Alexandru Manea, Anton Marchenko, Ulas Kozat, Aditya Akella, Ying Zhang 0022
NSDI11
2026 Towards Performance Robustness for Microservices
Divyanshu Saxena, Gaurav Vipat, Jingbo Wang 0006, Isil Dillig, Sanjay Shakkottai, Aditya Akella
NSDI7
2026 MatchBox: A Semantic Foundation for Data Plane Portability
abstract
Match-action tables are the core abstraction underlying network packet-processing systems, from fixed-function switches to eBPF-based software dataplanes. However, their concrete syntax and semantics vary widely across programming environments, reflecting differences in hardware generations, engineering practices, and vendor design choices. This syntactic and semantic variation renders portability of match-action tables across environments a persistent challenge. This paper presents MatchBox, a system for translating match-action tables across heterogeneous environments. At its core is the Match Algebra , a compositional formalism for concisely and declaratively expressing transformations on match-action tables. To ensure unambiguous semantics, MatchBox introduces a static type system based on guarded functional dependencies (GFDs) that guarantees that every well-typed Match Algebra expression denotes a well-defined function. From such specifications, the MatchBox compiler efficiently computes compact target tables that are semantically faithful. Across case studies in programmable switches, multi-cloud firewalls, and eBPF systems, MatchBox enables concise, declarative portability specifications and realizes them as compact target tables.
Eric Hayden Campbell, Robert Zhang 0003, Divyanshu Saxena, Aditya Akella, Isil Dillig
Proc. ACM Program. Lang.4
2026 Virtual Slicing: Achieving Control Plane Availability and Traffic Engineering Efficiency in Data Centers
abstract
Many proposals have demonstrated the efficiency advantages of software-defined networking (SDN) in managing data center networks. Common practices employ centralized traffic engineering (TE) in the SDN control plane to optimize load balancing and throughput. Meanwhile, for high availability purposes, the control plane is partitioned to ensure the impact of a single faulty controller is contained. However, the interaction between these two aspects is often overlooked. In particular, we show that the current control plane partitioning approach leads to imbalanced link loads and degraded application performance. To address this issue, we proposevirtual slicing, a new control plane partitioning scheme. Virtual slicing achieves desirable traffic engineering performance while retaining the availability guarantees from the current approach. Virtual slicing is implemented and evaluated with real-world and synthetic traffic traces on production spine-free data center networks. Results show that virtual slicing reduces tail link utilizations by up to 28.4%, and improves flow completion times by up to 36%.
Brian Chang, Keqiang He, Shawn Shuoshuo Chen, Mingyang Zhang 0005, Wenfei Wu, Fan Wu 0006, Chen Tian 0001, Aditya Akella
IEEE Trans. Netw.9
2025 StitchLLM: Serving LLMs, One Block at a Time
abstract
Bodun Hu, Shuozhe Li, Saurabh Agarwal, Myungjin Lee, Akshay Jajoo, Jiamin Li, Le Xu, Geon-Woo Kim, Donghyun Kim, Hong Xu, Amy Zhang, Aditya Akella. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Bodun Hu, Shuozhe Li, Myungjin Lee, Akshay Jajoo, Jiamin Li 0002, Geon-Woo Kim, Donghyun Kim 0002, Hong Xu 0001, Amy Zhang 0001, Aditya Akella
ACL (1)12
2025 Copper and Wire: Bridging Expressiveness and Performance for Service Mesh Policies
Divyanshu Saxena, William Zhang 0002, Shankara Pailoor, Isil Dillig, Aditya Akella
ASPLOS (1)5
2025 Large Language Models as Realistic Microservice Trace Generators
abstract
Workload traces are essential to understand complex computer systems' behavior and manage processing and memory resources.Since real-world traces are hard to obtain, synthetic trace generation is a promising alternative.This paper proposes a first-of-a-kind approach that relies on training a large language model (LLM) to generate synthetic workload traces, specifically microservice call graphs.To capture complex and arbitrary hierarchical structures and implicit constraints in such traces, we propose to train LLMs to generate recursively, making call graph generation a sequence of more manageable steps.To further enforce learning constraints on the traces and generate uncommon situations, we apply additional instruction tuning steps to align our model with the desired trace features.With this method, we train TraceLLM, an LLM for microservice trace generation, and demonstrate that it produces diverse, realistic traces under varied conditions, outperforming existing approaches in both accuracy and validity.The synthetically generated traces can effectively replace real data to optimize important microservice management tasks.Additionally, TraceLLM adapts to downstream trace-related tasks, such as predicting key trace features and infilling missing data.
Donghyun Kim 0002, Sriram Ravula, Taemin Ha, Alexandros G. Dimakis, Daehyeok Kim, Aditya Akella
EMNLP6
2025 Man-Made Heuristics Are Dead. Long Live Code Generators!
abstract
Policy design for various systems controllers has conventionally been a manual process, with domain experts carefully tailoring heuristics for the specific instance in which the policy will be deployed. In this paper, we re-imagine policy design via a novel automated search technique fueled by recent advances in generative models, specifically Large Language Model (LLM)-driven code generation. We outline the design and implementation of PolicySmith, a framework that applies LLMs to synthesize instance-optimal heuristics. We apply PolicySmith to two long-standing systems policies - web caching and congestion control, highlighting the opportunities unraveled by this LLM-driven heuristic search. For caching, PolicySmith discovers heuristics that outperform established baselines on standard open-source traces. For congestion control, we show that PolicySmith can generate safe policies that integrate directly into the Linux kernel.
Rohit Dwivedula, Divyanshu Saxena, Aditya Akella, Swarat Chaudhuri, Daehyeok Kim
HotNets3
2025 How I learned to stop worrying and love learned OS policies
abstract
While machine learning has been adopted across various fields, its ability to outperform traditional heuristics in operating systems is often met with justified skepticism. Concerns about unsafe decisions, opaque debugging processes, and the challenges of integrating ML into the kernel---given its stringent latency constraints and inherent complexity --- make practitioners understandably cautious. This paper introduces Guardrails for the OS, a framework that allows kernel developers to declaratively specify system-level properties and define corrective actions to address property violations. The framework facilitates the compilation of these guardrails into monitors capable of running within the kernel. In this work, we establish the foundation for Guardrails, detailing its core abstractions, examining the problem space, and exploring potential solutions.
Divyanshu Saxena, Sujay Yadalam, Yeonju Ro, Rohit Dwivedula, Eric Hayden Campbell, Aditya Akella, Christopher J. Rossbach, Michael Swift
HotOS7
2025 CONGO: Compressive Online Gradient Optimization
abstract
We address the challenge of zeroth-order online convex optimization where the objective function's gradient exhibits sparsity, indicating that only a small number of dimensions possess non-zero gradients. Our aim is to leverage this sparsity to obtain useful estimates of the objective function's gradient even when the only information available is a limited number of function samples. Our motivation stems from the optimization of large-scale queueing networks that process time-sensitive jobs. Here, a job must be processed by potentially many queues in sequence to produce an output, and the service time at any queue is a function of the resources allocated to that queue. Since resources are costly, the end-to-end latency for jobs must be balanced with the overall cost of the resources used. While the number of queues is substantial, the latency function primarily reacts to resource changes in only a few, rendering the gradient sparse. We tackle this problem by introducing the Compressive Online Gradient Optimization framework which allows compressive sensing methods previously applied to stochastic optimization to achieve regret bounds with an optimal dependence on the time horizon without the full problem dimension appearing in the bound. For specific algorithms, we reduce the samples required per gradient estimate to scale with the gradient's sparsity factor rather than its full dimensionality. Numerical simulations and real-world microservices benchmarks demonstrate CONGO's superiority over gradient descent approaches that do not account for sparsity.
Jeremy Carleton, Prathik Vijaykumar, Divyanshu Saxena, Dheeraj Narasimha, Srinivas Shakkottai, Aditya Akella
ICLR6
2025 HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model Training
abstract
Training large language models (LLMs) increasingly relies on geographically distributed accelerators, causing prohibitive communication costs across regions and uneven utilization of heterogeneous hardware. We propose HALoS, a hierarchical asynchronous optimization framework that tackles these issues by introducing local parameter servers (LPSs) within each region and a global parameter server (GPS) that merges updates across regions. This hierarchical design minimizes expensive inter-region communication, reduces straggler effects, and leverages fast intra-region links. We provide a rigorous convergence analysis for HALoS under non-convex objectives, including theoretical guarantees on the role of hierarchical momentum in asynchronous training. Empirically, HALoS attains up to 7.5× faster convergence than synchronous baselines in geo-distributed LLM training and improves upon existing asynchronous methods by up to 2.1×. Crucially, HALoS preserves the model quality of fully synchronous SGD—matching or exceeding accuracy on standard language modeling and downstream benchmarks—while substantially lowering total training time. These results demonstrate that hierarchical, server-side update accumulation and global model merging are powerful tools for scalable, efficient training of new-era LLMs in heterogeneous, geo-distributed environments.
Geon-Woo Kim, Shashidhar Gandham, Omar Baldonado, Adithya Gangidi, Pavan Balaji, Zhangyang Wang, Aditya Akella
ICML8
2025 On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention for Long-Context LLM Serving
abstract
Large language models (LLMs) excel at capturing global token dependencies via self-attention but face prohibitive compute and memory costs on lengthy inputs. While sub-quadratic methods (e.g., linear attention) can reduce these costs, they often degrade accuracy due to overemphasizing recent tokens. In this work, we first propose dual-state linear attention (DSLA), a novel design that maintains two specialized hidden states—one for preserving historical context and one for tracking recency—thereby mitigating the short-range bias typical of linear-attention architectures. To further balance efficiency and accuracy under dynamic workload conditions, we introduce DSLA-Serve, an online adaptive distillation framework that progressively replaces Transformer layers with DSLA layers at inference time, guided by a sensitivity-based layer ordering. DSLA-Serve uses a chained fine-tuning strategy to ensure that each newly converted DSLA layer remains consistent with previously replaced layers, preserving the overall quality. Extensive evaluations on commonsense reasoning, long-context QA, and text summarization demonstrate that DSLA-Serve yields 2.3$\times$ faster inference than Llama2-7B and 3.0$\times$ faster than the hybrid Zamba-7B, while retaining comparable performance across downstream tasks. Our ablation studies show that DSLA’s dual states capture both global and local dependencies, addressing the historical-token underrepresentation seen in prior linear attentions.
Yeonju Ro, Zhenyu Zhang 0015, Souvik Kundu 0009, Zhangyang Wang, Aditya Akella
ICML5
2025 ConfigBot: Adaptive Resource Allocation for Robot Applications in Dynamic Environments
abstract
The growing use of service robots in dynamic environments requires flexible management of on-board compute resources to optimize the performance of diverse tasks such as navigation, localization, and perception. Current robot deployments often rely on static OS configurations and system over-provisioning. However, they are suboptimal because they ignore variations in resource usage, leading to system-wide issues like robot instability or inefficient resource utilization. This paper presents ConfigBot, a novel system designed to adaptively reconfigure robot applications to meet a predefined performance specification by leveraging runtime profiling and automated configuration tuning. Through experiments on multiple real robots, each running a different stack with diverse performance requirements, which could be context-dependent, we illustrate ConfigBot's efficacy in maintaining system stability and optimizing resource allocation. Our findings highlight the promise of automatic system configuration tuning for robot deployments, including adaptation to dynamic changes. Code available at: https://github.com/ldos-project/configbot
Rohit Dwivedula, Sadanand Modak, Aditya Akella, Joydeep Biswas, Daehyeok Kim, Christopher J. Rossbach
IROS3
2025 MTP: Transport for In-Network Computing
Rohan Vardekar, Balajee Vamanan, Brent E. Stephens, Aditya Akella
NSDI5
2025 Enabling Portable and High-Performance SmartNIC Programs with Alkali
Mihir Shah, Yiying Zhang 0005, Daehyeok Kim, Aditya Akella
NSDI7
2024 MOSEL: Inference Serving Using Dynamic Modality Selection
abstract
Rapid advancements over the years have helped machine learning models reach previously hardto-achieve goals, sometimes even exceeding human capabilities.However, achieving desired accuracy comes at the cost of larger model sizes and increased computational demands.Thus, serving predictions from these models to meet any latency and cost requirements of applications remains a key challenge, despite recent work in building inference serving systems as well as algorithmic approaches that dynamically adapt models based on inputs.Our paper introduces a new form of dynamism, modality selection, where we adaptively choose modalities from inference inputs while maintaining the model quality.We introduce MOSEL, an automated inference serving system for multi-modal ML models that carefully picks input modalities per request based on resource availability, as we as user-defined service level objectives (SLOs).MOSEL extensively leverages modality configurations, improving system throughput by 3.6× with an accuracy guarantee.It also reduces job completion times by 11× compared to modalityagnostic approaches.
Bodun Hu, Jeongyoon Moon, Neeraja J. Yadwadkar, Aditya Akella
EMNLP5
2024 FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
abstract
Autoregressive Large Language Models (e.g., LLaMa, GPTs) are omnipresent achieving remarkable success in language understanding and generation.However, such impressive capability typically comes with a substantial model size, which presents significant challenges for autoregressive token-by-token generation.To mitigate computation overload incurred during generation, several early-exit and layer-dropping strategies have been proposed.Despite some promising success due to the redundancy across LLMs layers on metrics like Rough-L/BLUE, our careful knowledgeintensive evaluation unveils issues such as generation collapse, hallucination, and noticeable performance drop even at the trivial exit ratio of ∼ 10-15% of layers.We attribute these errors primarily to ineffective handling of the KV cache through state copying during early exit.In this work, we observe the saturation of computationally expensive feed-forward blocks of LLM layers and propose FFN-SkipLLM, which is a novel fine-grained skip strategy for autoregressive LLMs.FFN-SkipLLM leverages an input-adaptive feed-forward skipping approach that can skip ∼ 25-30% of FFN blocks of LLMs with marginal change in performance on knowledge-intensive generation tasks without any requirement to handle the KV cache.Our extensive experiments and ablation studies across benchmarks like MT-Bench, Factoid-QA, and variable-length text summarization illustrate how our simple and easy-touse method can facilitate faster autoregressive decoding.
Ajay Jaiswal, Bodun Hu, Lu Yin 0006, Yeonju Ro, Tianlong Chen 0001, Shiwei Liu 0003, Aditya Akella
EMNLP7
2024 Balancing Sdn Control Plane Availability and Traffic Engineering Efficiency in Data Centers
abstract
Many proposals have demonstrated the efficiency advantages of software-defined networking (SDN) in managing data center networks. Common practices employ centralized traffic engineering (TE) in the SDN control plane to optimize load balancing and throughput. Meanwhile, for high availability purposes, the control plane is partitioned to ensure the impact of a single faulty controller is contained. However, the interaction between these two aspects is often overlooked. In particular, we show that the current control plane partitioning approach leads to imbalanced link loads and degraded application performance. To address this issue, we propose virtual slicing, a new control plane partitioning scheme. Virtual slicing achieves desirable traffic engineering performance while retaining the availability guarantees from the current approach. Virtual slicing is implemented and evaluated with real-world and synthetic traffic traces on production spine-free data center networks. Results show that virtual slicing reduces tail link utilizations by up to 28.4 %, and improves flow completion times by up to 36 %.
Brian Chang, Keqiang He, Shawn Shuoshuo Chen, Mingyang Zhang 0005, Wenfei Wu, Aditya Akella
ICNP7
2024 Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
Ruisi Cai, Yeonju Ro, Geon-Woo Kim, Peihao Wang, Babak Ehteshami Bejnordi, Aditya Akella, Zhangyang Wang
NeurIPS6
2024 CASSINI: Network-Aware Job Scheduling in Machine Learning Clusters
Sudarsanan Rajasekaran, Manya Ghobadi, Aditya Akella
NSDI3
2024 Learned load balancing
Brian Chang, Kausik Subramanian, Loris D'Antoni, Aditya Akella
Theor. Comput. Sci.4
2023 Towards a Machine Learning-Assisted Kernel with LAKE
abstract
The complexity of modern operating systems (OSes), rapid diversification of hardware, and steady evolution of machine learning (ML) motivate us to explore the potential of ML to improve decision-making in OS kernels. We conjecture that ML can better manage tradeoff spaces for subsystems such as memory management and process and I/O scheduling that currently rely on hand-tuned heuristics to provide reasonable average-case performance. We explore the replacement of heuristics with ML-driven decision-making in five kernel subsystems, consider the implications for kernel design, shared OS-level components, and access to hardware acceleration. We identify obstacles, address challenges and characterize tradeoffs for the benefits ML can provide that arise in kernel-space. We find that use of specialized hardware such as GPUs is critical to absorbing the additional computational load required by ML decisioning, but that poor accessibility of accelerators in kernel space is a barrier to adoption. We also find that the benefits of ML and acceleration for OSes is subsystem-, workload- and hardware-dependent, suggesting that using ML in kernels will require frameworks to help kernel developers navigate new tradeoff spaces. We address these challenge by building a system called LAKE for supporting ML and exposing accelerators in kernel space. LAKE includes APIs for feature collection and management across abstraction layers and module boundaries. LAKE provides mechanisms for managing the variable profitability of acceleration, and interfaces for mitigating contention for resources between user and kernel space. We show that an ML-backed I/O latency predictor can have its inference time reduced by up to 96% with acceleration.
Henrique Fingler, Isha Tarte, Hangchen Yu, Ariel Szekely, Bodun Hu, Aditya Akella, Christopher J. Rossbach
ASPLOS (2)6
2023 Yama: Providing Performance Isolation for Black-Box Offloads
abstract
The sharing of clusters with various on-NIC offloads by high-level entities (users, containers, etc.) has become increasingly common. Performance isolation across these entities is desired because the offloads can become bottlenecks due to the limited capacity of hardware. However, the existing works that provide scheduling and resource management to NIC offloads all require customization of the NIC or offloads, while commodity off-the-shelf NICs and offloads with proprietary implementation have been widely deployed in datacenters. This paper presents Yama, the first solution to enable per-entity isolation in the sharing of such black-box NIC offloads. Yama provides a generic framework that captures a common abstraction to the operation of most offloads, which allows operators to incorporate existing offloads. The framework proactively probes for the performance of the offloads with auxiliary workload and enforces isolation at the initiator side. Yama also accommodates chained offloads. Our evaluation shows that 1) Yama achieves per-entity max-min fairness for various types of offloads and in complicated offload chaining scenarios; 2) Yama quickly converges to changes in equilibrium and 3) Yama adds negligible overhead to application workload.
Divyanshu Saxena, Brent E. Stephens, Aditya Akella
SoCC4
2023 Auxo: Efficient Federated Learning via Scalable Client Clustering
abstract
Federated learning (FL) is an emerging machine learning (ML) paradigm that enables heterogeneous edge devices to collaboratively train ML models without revealing their raw data to a logically centralized server. However, beyond the heterogeneous device capacity, FL participants often exhibit differences in their data distributions, which are not independent and identically distributed (Non-IID). Many existing works present point solutions to address issues like slow convergence, low final accuracy, and bias in FL, all stemming from client heterogeneity.
Fan Lai 0001, Yinwei Dai, Aditya Akella, Harsha V. Madhyastha, Mosharaf Chowdhury
SoCC4
2023 Lowering the Pre-training Tax for Gradient-based Subset Training: A Lightweight Distributed Pre-Training Toolkit
abstract
Training data and model sizes are increasing exponentially. One way to reduce training time and resources is to train with a carefully selected subset of the full dataset. Prior work uses the gradient signals obtained during a warm-up or “pre-training" phase over the full dataset, for determining the core subset; if the pre-training phase is too small, the gradients obtained are chaotic and unreliable. As a result, the pre-training phase itself incurs significant time/resource overhead, and prior work has not gone beyond hyperparameter search to reduce pre-training time. Our work explicitly aims to reduce this $\textbf{pre-training tax}$ in gradient-based subset training. We develop a principled, scalable approach for pre-training in a distributed setup. Our approach is $\textit{lightweight}$ and $\textit{minimizes communication}$ between distributed worker nodes. It is the first to utilize the concept of model-soup based distributed training $\textit{at initialization}$. The key idea is to minimally train an ensemble of models on small, disjointed subsets of the data; we further employ data-driven sparsity and data augmentation for local worker training to boost ensemble diversity. The centralized model, obtained at the end of pre-training by merging the per-worker models, is found to offer stabilized gradient signals to select subsets, on which the main model is further trained. We have validated the effectiveness of our method through extensive experiments on CIFAR-10/100, and ImageNet, using ResNet and WideResNet models. For example, our approach is shown to achieve $\textbf{15.4$\times$}$ pre-training speedup and $\textbf{2.8$\times$}$ end-to-end speedup on CIFAR10 and ResNet18 without loss of accuracy. The code is at https://github.com/moonbucks/LiPT.git.
Yeonju Ro, Zhangyang Wang, Vijay Chidambaram, Aditya Akella
ICML4
2023 LogNIC: A High-Level Performance Model for SmartNICs
abstract
SmartNICs have become an indispensable communication fabric and computing substrate in today’s data centers and enterprise clusters, providing in-network computing capabilities for traversed packets and benefiting a range of applications across the system stack. Building an efficient SmartNIC-assisted solution is generally non-trivial and tedious as it requires programmers to understand the SmartNIC architecture, refactor application logic to match the device’s capabilities and limitations, and correlate an application execution with traffic characteristics. A high-level SmartNIC performance model can decouple the underlying SmartNIC hardware device from its offloaded software implementations and execution contexts, thereby drastically simplifying and facilitating the development process. However, prior architectural models can hardly be applied due to their limited capabilities in dissecting the SmartNIC-offloaded program’s complexity, capturing the nondeterministic overlapping between computation and I/O, and perceiving diverse traffic profiles.
Zerui Guo, Yuebin Bai, Daehyeok Kim, Michael M. Swift, Aditya Akella, Ming Liu 0027
MICRO6
2023 RingLeader: Efficiently Offloading Intra-Server Orchestration to NICs
Adney Cardoza, Tarannum Khan, Yeonju Ro, Brent E. Stephens, Hassan M. G. Wassel, Aditya Akella
NSDI7
2023 Better Together: Jointly Optimizing ML Collective Scheduling and Execution Planning using SYNDICATE
Kshiteej Mahajan, Ching-Hsiang Chu, Srinivas Sridharan 0002, Aditya Akella
NSDI4
2023 Shockwave: Fair and Efficient Cluster Scheduling for Dynamic Adaptation in Machine Learning
Rui Pan 0003, Tarannum Khan, Shivaram Venkataraman, Aditya Akella
NSDI5
2023 Darwin: Flexible Learning-based CDN Caching
abstract
Cache management is critical for Content Delivery Networks (CDNs), impacting their performance and operational costs. Most production CDNs apply static, hand-tuned caching policy parameters at cache servers, such as admission frequency or size thresholds for the Hot Object Caches (HOC) of their system. However, these static policies fall short when a server is faced with unpredictable traffic pattern changes, even when policies employ multiple control parameters/knobs. Recent approaches have proposed learning-based solutions to dynamically adjust policy parameters, but they are limited in action space, caching objectives, or impose high overhead. We propose Darwin, a CDN cache management system that is robust to traffic pattern changes and can flexibly optimize different caching objectives with unrestricted action spaces. Darwin employs a three-stage pipeline involving traffic pattern feature collection, unsupervised clustering for classification, and neural bandit expert selection to choose the optimal caching policy. Through extensive simulations, experiments using an Apache Traffic Server (ATS)-based prototype, and theoretical analysis, we show that Darwin achieves significant performance gain w.r.t. different objectives such as maximizing object hit rates and minimizing disk writes, while simultaneously adapting to traffic pattern shifts. Darwin imposes negligible overhead and achieves high throughput compared to the state-of-the-art.
Nihal Sharma, Tarannum Khan, Brian Chang, Aditya Akella, Sanjay Shakkottai, Ramesh K. Sitaraman
SIGCOMM6
2023 Resilient Baseband Processing in Virtualized RANs with Slingshot
abstract
In cellular networks, there is a growing adoption of virtualized radio access networks (vRANs), where operators are replacing the traditional specialized hardware for RAN processing with software running on commodity servers. Today's vRAN deployments lack resilience, since there is no support for vRAN failover or upgrades without long service interruptions. Enabling these features for vRANs is challenging because of their strict real-time latency requirements and black-box nature. Slingshot is a new system that transparently provides resilience for the vRAN's most performance-critical layer: the physical layer (PHY). We design new techniques for realtime workload migration with fast RAN protocol middle-boxes, and realtime RAN failure detection. A key insight in our design is to view the transient disruptions from resilience events to RAN computation state and I/O similarly to regular wireless signal impairments, and leverage the inherent resilience of cellular networks to these events. Experiments with a state-of-the-art 5G vRAN testbed show that Slingshot handles PHY failover with no disruption to video conferencing, and under 110 ms disruption to a TCP connection, and it also enables zero-downtime upgrades.
Nikita Lazarev, Anuj Kalia, Daehyeok Kim, Ilias Marinos, Francis Y. Yan, Christina Delimitrou, Zhiru Zhang, Aditya Akella
SIGCOMM9
2023 ChainedFilter: Combining Membership Filters by Chain Rule
abstract
Membership (membership query/membership testing) is a fundamental problem across databases, networks and security. However, previous research has primarily focused on either approximate solutions, such as Bloom Filters, or exact methods, like perfect hashing and dictionaries, without attempting to develop an integral theory. In this paper, we propose a unified and complete theory, namely chain rule, for general membership problems, which encompasses both approximate and exact membership as extreme cases. Building upon the chain rule, we introduce a straightforward yet versatile algorithm framework, namely ChainedFilter, to combine different elementary filters without losing information. Our evaluation results demonstrate that ChainedFilter improves performance of many applications including static dictionary, lossless data compression, Cuckoo Hashing, LSM-Tree and Learned Filters.
Liuhui Wang, Jianan Ji, Yuhan Wu 0001, Yikai Zhao 0001, Tong Yang 0003, Aditya Akella
Proc. ACM Manag. Data8
2023 Blink-hash: An Adaptive Hybrid Index for In-Memory Time-Series Databases
abstract
High-speed data ingestion is critical in time-series workloads that are driven by the growth of Internet of Things (IoT) applications. We observe that traditional tree-based indexes encounter severe scalability bottlenecks for time-series workloads that insert monotonically increasing timestamp keys into an index; all insertions go to a small memory region that sees extremely high contention. In this work, we present a new index design, B link -hash, that enhances a tree-based index with hash leaf nodes to mitigate the contention of monotonic insertions --- insertions go to random locations within a hash node (which is much larger than a B+-tree node) to reduce conflicts. We develop further optimizations (median approximation and lazy split) to accelerate hash node splits. We also develop structure adaptation optimizations to dynamically convert a hash node to B+-tree nodes for good scan performance. Our evaluation shows that B link -hash achieves up to 91.3× higher throughput than conventional indexes in a time-series workload that monotonically inserts timestamps into an index, while showing comparable scan performance to a well-optimized B+-tree.
Hokeun Cha, Xiangpeng Hao, Tianzheng Wang 0001, Huanchen Zhang, Aditya Akella, Xiangyao Yu
Proc. VLDB Endow.5
2022 Jiffy: elastic far-memory for stateful serverless analytics
abstract
Stateful serverless analytics can be enabled using a remote memory system for inter-task communication, and for storing and exchanging intermediate data. However, existing systems allocate memory resources at job granularity---jobs specify their memory demands at the time of the submission; and, the system allocates memory equal to the job's demand for the entirety of its lifetime. This leads to resource underutilization and/or performance degradation when intermediate data sizes vary during job execution.
Anurag Khandelwal, Yupeng Tang, Rachit Agarwal 0001, Aditya Akella, Ion Stoica
EuroSys4
2022 Memory deduplication for serverless computing with Medes
abstract
Serverless platforms today impose rigid trade-offs between resource use and user-perceived performance. Limited controls, provided via toggling sandboxes between warm and cold states and keep-alives, force operators to sacrifice significant resources to achieve good performance. We present a serverless framework, Medes, that breaks the rigid trade-off and allows operators to navigate the trade-off space smoothly. Medes leverages the fact that the warm sandboxes running on serverless platforms have a high fraction of duplication in their memory footprints. We exploit these redundant chunks to develop a new sandbox state, called a dedup state, that is more memory-efficient than the warm state and faster to restore from than the cold state. We develop novel mechanisms to identify memory redundancy at minimal overhead while ensuring that the dedup containers' memory footprint is small. Finally, we develop a simple sandbox management policy that exposes a narrow, intuitive interface for operators to trade-off performance for memory by jointly controlling warm and dedup sandboxes. Detailed experiments with a prototype using real-world serverless workloads demonstrate that Medes can provide up to 1×-2.75× improvements in the end-to-end latencies. The benefits of Medes are enhanced in memory pressure situations, where Medes can provide up to 3.8× improvements in end-to-end latencies. Medes achieves this by reducing the number of cold starts incurred by 10--50% against the state-of-the-art baselines.
Divyanshu Saxena, Arjun Singhvi, Junaid Khalid, Aditya Akella
EuroSys5
2022 Impact of RoCE Congestion Control Policies on Distributed Training of DNNs
abstract
Ahstract-RDMA over Converged Ethernet (RoCE) has gained significant attraction for datacenter networks due to its compatibility with conventional Ethernet-based fabric. However, the RDMA protocol is efficient only on (nearly) lossless networks, emphasizing the vital role of congestion control on RoCE networks. Unfortunately, the native RoCE congestion control scheme, based on Priority Flow Control (PFC), suffers from many drawbacks such as unfairness, head-of-line-blocking, and deadlock. Therefore, in recent years many schemes have been proposed to provide additional congestion control for RoCE networks to minimize PFC drawbacks. However, these schemes are proposed for general datacenter environments. In contrast to the general datacenters that are built using commodity hardware and run general-purpose workloads, high-performance distributed training platforms deploy high-end accelerators and network components and exclusively run training workloads using collectives (All-Reduce, All-To-All) communication libraries for communication. Furthermore, these platforms usually have a private network, separating their communication traffic from the rest of the datacenter traffic. Scalable topology-aware collective algorithms are inherently designed to avoid incast patterns and balance traffic optimally. These distinct features necessitate revisiting previously proposed congestion control schemes for general-purpose datacenter environments. In this paper, we thoroughly analyze some of the state-of-the-art RoCE congestion control schemes (DCQCN, DCTCP, TIMELY, and HPCC) vs. PFC when running on distributed training platforms. Our results indicate that pre-viously proposed RoCE congestion control schemes have little impact on the end-to-end performance of training workloads, motivating the necessity of designing an optimized, yet low-overhead, congestion control scheme based on the characteristics of distributed training platforms and workloads.
Tarannum Khan, Saeed Rashidi, Srinivas Sridharan 0002, Pallavi Shurpali, Aditya Akella, Tushar Krishna
HOTI5
2022 Congestion control in machine learning clusters
abstract
This paper argues that fair-sharing, the holy grail of congestion control algorithms for decades, is not necessarily a desirable property in Machine Learning (ML) training clusters. We demonstrate that for a specific combination of jobs, introducing unfairness improves the training time for all competing jobs. We call this specific combination of jobs compatible and define the compatibility criterion using a novel geometric abstraction. Our abstraction rolls time around a circle and rotates the communication phases of jobs to identify fully compatible jobs. Using this abstraction, we demonstrate up to 1.3× improvement in the average training iteration time of popular ML models. We advocate that resource management algorithms should take job compatibility on network links into account. We then propose three directions to ameliorate the impact of network congestion in ML training clusters: (i) an adaptively unfair congestion control scheme, (ii) priority queues on switches, and (iii) precise flow scheduling.
Sudarsanan Rajasekaran, Manya Ghobadi, Gautam Kumar 0001, Aditya Akella
HotNets4
2021 Atoll: A Scalable Low-Latency Serverless Platform
abstract
With user-facing apps adopting serverless computing, good latency performance of serverless platforms has become a strong fundamental requirement. However, it is difficult to achieve this on platforms today due to the design of their underlying control and data planes that are particularly ill-suited to short-lived functions with unpredictable arrival patterns. We present Atoll, a serverless platform, that overcomes the challenges via a ground-up redesign of the control and data planes. In Atoll, each app is associated with a latency deadline. Atoll achieves its per-app request latency goals by: (a) partitioning the cluster into (semi-global scheduler, worker pool) pairs, (b) performing deadline-aware scheduling and proactive sandbox allocation, and (c) using a load balancing layer to do sandbox-aware routing, and automatically scale the semi-global schedulers per app. Our results show that Atoll reduces missed deadlines by ~66x and tail latencies by ~3x compared to state-of-the-art alternatives.
Arjun Singhvi, Arjun Balasubramanian, Kevin Houck, Mohammed Danish Shaikh, Shivaram Venkataraman, Aditya Akella
SoCC6
2021 TCP is Harmful to In-Network Computing: Designing a Message Transport Protocol (MTP)
abstract
This paper presents the motivation and design of MTP, a new offload-friendly message transport protocol. Existing transport protocols like TCP, MPTCP, and UDP/Quic all have key limitations when used in a network that may potentially offload computation from end-servers into NICs, switches, and other network devices. To enable important new in-network computing use cases and correct congestion control in the face of ever changing network paths and application replicas, MTP introduces a new message transport protocol design and pathlet congestion control, a new approach where end-hosts explicitly communicate messaging information to network devices and network devices explicitly communicate network path and congestion information back to end-hosts.
Brent E. Stephens, Darius Grassi, Hamidreza Almasi, Balajee Vamanan, Aditya Akella
HotNets6
2021 A Vision for Runtime Programmable Networks
abstract
Our community has made significant progress in developing programmable network infrastructure, starting from the control plane and expanding to the data plane. As a latest trend, network devices are becoming runtime programmable while serving live traffic. This allows for reprogramming of individual device programs at fine-grained timescales to add or remove network functions. Many applications and services, however, need control over a combination of devices, including end host stacks, NICs, and switches, to accomplish their goals. We lay out our vision for runtime programmable networks, building upon device-level features to provide live, network-wide, runtime reprogramming. A whole-stack approach is needed with new programming models, compiler support, and network management abstractions. We outline a research agenda as a call to arms to the community.
Jiarong Xing, Yiming Qiu 0001, Kuo-Feng Hsu, Matty Kadosh, Alan Lo, Aditya Akella, Thomas E. Anderson, Arvind Krishnamurthy, T. S. Eugene Ng, Ang Chen 0001
HotNets7
2021 Running BGP in Data Centers at Scale
Anubhavnidhi Abhashkumar, Kausik Subramanian, Alexey Andreyev, Hyojeong Kim, Nanda Kishore Salem, Petr Lapukhov, Aditya Akella, Hongyi Zeng
NSDI8
2021 Whiz: Data-Driven Analytics Execution
Robert Grandl, Arjun Singhvi, Raajay Viswanathan, Aditya Akella
NSDI4
2021 ATP: In-network Aggregation for Multi-tenant Learning
ChonLam Lao, Yanfang Le, Kshiteej Mahajan, Wenfei Wu, Aditya Akella, Michael M. Swift
NSDI6
2021 CliqueMap: productionizing an RMA-based distributed caching system
abstract
Distributed in-memory caching is a key component of modern Internet services. Such caches are often accessed via remote procedure call (RPC), as RPC frameworks provide rich support for productionization, including protocol versioning, memory efficiency, auto-scaling, and hitless upgrades. However, full-featured RPC limits performance and scalability as it incurs high latencies and CPU overheads. Remote Memory Access (RMA) offers a promising alternative, but meeting productionization requirements can be a significant challenge with RMA-based systems due to limited programmability and narrow RMA primitives.
Arjun Singhvi, Aditya Akella, Maggie Anderson, Rob Cauble, Harshad Deshmukh, Dan Gibson, Milo M. K. Martin, Amanda Strominger, Thomas F. Wenisch, Amin Vahdat
SIGCOMM2
2020 SNF: serverless network functions
abstract
Our work addresses how a cloud provider can offer Network Functions (NF) as a Service, or NFaaS, using the emerging serverless computing paradigm. Serverless computing has the right NFaaS building blocks - usage-based billing, event-driven programming model and elastic scaling. But we identify two core limitations of existing serverless platforms that undermine support for NFaaS - coupling of the billing and work assignment granularities, and state sharing via an external store. Our framework, SNF, overcomes these limitations via two ideas. SNF allocates work at the granularity of flowlets observed in network traffic, whereas billing and programming occur at a finer level. SNF embellishes serverless platforms with ephemeral local state that lasts for the flowlet duration and supports high performance state operations. We demonstrate that our SNF prototype matches utilization closely with demand and reduces tail packet processing latency substantially compared to alternatives.
Arjun Singhvi, Junaid Khalid, Aditya Akella, Sujata Banerjee
SoCC3
2020 Network-accelerated distributed machine learning for multi-tenant settings
abstract
Many distributed machine learning (DML) workloads are increasingly being run in shared clusters. Training in such clusters can be impeded by unexpected compute and network contention, resulting in stragglers. We present MLfabric, a contention-aware DML system that manages the performance of a DML job running in a shared cluster. The DML application hands all network communication (gradient and model transfers) to the MLfabric communication library. MLfabric then carefully orders transfers to improve convergence, opportunistically aggregates them at idle DML workers to improve resource efficiency, and replicates them to support new notions of fault tolerance, while systematically accounting for compute stragglers and network contention. We find that MLfabric achieves up to 3x speed-up in training large deep learning models in realistic dynamic cluster settings.
Raajay Viswanathan, Arjun Balasubramanian, Aditya Akella
SoCC3
2020 AED: incrementally synthesizing policy-compliant and manageable configurations
abstract
When updating router configurations, network operators often attempt to meet a variety of management objectives (e.g., maintaining structural similarity across devices), while also ensuring all forwarding policies are correctly satisfied. Our tool, AED, automates this process. AED models configuration updates as a collection of syntax tree additions and removals, and formulates an innovative system of SMT (Satisfiability Modulo Theory) constraints that encode configurations' structure and interaction with routing algorithms. Operators express management objectives in a high-level language, and AED translates these to "soft" constraints that are maximally satisfied. Evaluations on real and synthetic network configurations show that AED can update networks with tens of routers and hundreds of policies in under a minute, and AED outperforms both hand-crafted updates and state-of-the-art tools in meeting management objectives.
Anubhavnidhi Abhashkumar, Aaron Gember, Aditya Akella
CoNEXT3
2020 Tiramisu: Fast Multilayer Network Verification
Anubhavnidhi Abhashkumar, Aaron Gember, Aditya Akella
NSDI3
2020 Themis: Fair and Efficient GPU Cluster Scheduling
Kshiteej Mahajan, Arjun Balasubramanian, Arjun Singhvi, Shivaram Venkataraman, Aditya Akella, Amar Phanishayee, Shuchi Chawla 0001
NSDI5
2020 Liveness Verification of Stateful Network Functions
Farnaz Yousefi, Anubhavnidhi Abhashkumar, Kausik Subramanian, Kartik Hans, Soudeh Ghorbani, Aditya Akella
NSDI6
2020 Automated Verification of Customizable Middlebox Properties with Gravel
Kaiyuan Zhang 0001, Danyang Zhuo, Aditya Akella, Arvind Krishnamurthy, Xi Wang 0005
NSDI3
2020 PANIC: A High-Performance Programmable NIC for Multi-tenant Networks
Kiran Patel, Brent E. Stephens, Anirudh Sivaraman, Aditya Akella
OSDI5
2020 Detecting network load violations for distributed control planes
abstract
One of the major challenges faced by network operators pertains to whether their network can meet input traffic demand, avoid overload, and satisfy service-level agreements. Automatically verifying if no network links are overloaded is complicated---requires modeling frequent network failures, complex routing and load-balancing technologies, and evolving traffic requirements. We present QARC, a distributed control plane abstraction that can automatically verify whether a control plane may cause link-load violations under failures. QARC is fully automatic and can help operators program networks that are more resilient to failures and upgrade the network to avoid violations. We apply QARC to real datacenter and ISP networks and find interesting cases of load violations. QARC can detect violations in under an hour.
Kausik Subramanian, Anubhavnidhi Abhashkumar, Loris D'Antoni, Aditya Akella
PLDI4
2020 1RMA: Re-envisioning Remote Memory Access for Multi-tenant Datacenters
abstract
Remote Direct Memory Access (RDMA) plays a key role in supporting performance-hungry datacenter applications. However, existing RDMA technologies are ill-suited to multi-tenant datacenters, where applications run at massive scales, tenants require isolation and security, and the workload mix changes over time. Our experiences seeking to operationalize RDMA at scale indicate that these ills are rooted in standard RDMA's basic design attributes: connectionorientedness and complex policies baked into hardware.
Arjun Singhvi, Aditya Akella, Dan Gibson, Thomas F. Wenisch, Monica Wong-Chan, Sean Clark 0003, Milo M. K. Martin, Moray McLaren, Prashant Chandra, Rob Cauble, Hassan M. G. Wassel, Behnam Montazeri, Simon L. Sabato, Joel Scherpelz, Amin Vahdat
SIGCOMM2
2019 On the Impact of Cluster Configuration on RoCE Application Design
abstract
RDMA over Converged Ethernet (RoCE) allows RDMA-enabled NICs to operate in datacenter networks. This study focuses on identifying how different aspects of datacenter cluster configuration impact the latency, and throughput, and CPU utilization of different ways of transferring data in RoCE (RDMA verbs). We look into the impact of colocated applications competing for both the CPU and access to the NIC as well as the impact of the network MTU. We find that RDMA applications do not fairly share the NIC, large frames should not be used, and that correct verb choice is dependent on many variables, including application access patterns, object size, and the load of both the local and remote CPU.
Yanfang Le, Mojtaba MalekpourShahraki, Brent E. Stephens, Aditya Akella, Michael M. Swift
APNet4
2019 Correctness and Performance for Stateful Chained Network Functions
Junaid Khalid, Aditya Akella
NSDI2
2019 Loom: Flexible and Efficient NIC Packet Scheduling
Brent E. Stephens, Aditya Akella, Michael M. Swift
NSDI2
2019 The Design and Operation of CloudLab
Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson 0004, Kirk Webb, Aditya Akella, Kuang-Ching Wang, Glenn Ricart, Lawrence H. Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, Prabodh Mishra
USENIX ATC11
2018 RoGUE: RDMA over Generic Unconverged Ethernet
abstract
RDMA over Converged Ethernet (RoCE) promises low latency and low CPU utilization over commodity networks, and is attractive for cloud infrastructure services. Current implementations require Priority Flow Control (PFC) that uses backpressure-based congestion control to provide lossless networking to RDMA. Unfortunately, PFC compromises network stability. As a result, RoCE's adoption has been slow and requires complex network management. Recent efforts, such as DCQCN, reduce the risk to the network, but do not completely solve the problem.
Yanfang Le, Brent E. Stephens, Arjun Singhvi, Aditya Akella, Michael M. Swift
SoCC4
2018 Your Programmable NIC Should be a Programmable Switch
abstract
Today's NICs are becoming programmable ("smart"). To support new network protocols, services, and offloads, there are NICs today that have on-board FPGAs, embedded processors, programmable forwarding pipelines, and specialized engines to support features like RDMA. Unfortunately, existing programmable NICs have a number of key limitations. It is difficult to chain offloads, schedule competing accesses to shared resources, and support functions that require variable processing time and thus may not run at line-rate.
Brent E. Stephens, Aditya Akella, Michael M. Swift
HotNets2
2018 Iron: Isolating Network-based CPU in Container Environments
Junaid Khalid, Eric Rozner, Wes Felter, Karthick Rajamani, Aditya Akella
NSDI7
2018 Dynamic Query Re-Planning using QOOP
Kshiteej Mahajan, Mosharaf Chowdhury, Aditya Akella, Shuchi Chawla 0001
OSDI3
2018 Smurf: Self-Service String Matching Using Random Forests
abstract
We argue that more attention should be devoted to developing self-service string matching (SM) solutions, which lay users can easily use. We show that Falcon, a self-service entity matching (EM) solution, can be applied to SM and is more accurate than current self-service SM solutions. However, Falcon often asks lay users to label many string pairs (e.g., 770-1050 in our experiments). This is expensive, can significantly compound labeling mistakes, and takes a long time. We developed Smurf, a self-service SM solution that reduces the labeling effort by 43-76%, yet achieves comparable F 1 accuracy. The key to make Smurf possible is a novel solution to efficiently execute a random forest (that Smurf learns via active learning with the lay user) over two sets of strings. This solution uses RDBMS-style plan optimization to reuse computations across the trees in the forest. As such, Smurf significantly advances self-service SM and raises interesting future directions for self-service EM and scalable random forest execution over structured data.
Paul Suganthan G. C., Adel Ardalan, AnHai Doan, Aditya Akella
Proc. VLDB Endow.4
2017 Low Latency Software Rate Limiters for Cloud Networks
abstract
A lot of recent work has focused on reducing in network queueing latency in datacenter networks. In this paper, we focus on a less explored topic --- latency increases caused by queueing in rate limiters on the end-host. First, we show that latency can be increased by an order of magnitude by rate limiters in cloud networks. To solve this problem, we extend ECN marking into rate limiters and use a datacenter congestion control algorithm --- DCTCP. Unfortunately, while this reduces latency, it also leads to throughput oscillation. Thus, this solution is not sufficient. In this paper, we also analyze the specific reasons that ECN marking in software rate limiters leads to the throughput oscillation problem. Finally, we propose two potential solutions to design software rate limiters that can achieve stable high throughput and low latency.
Keqiang He, Weite Qin, Wenfei Wu, Tian Pan 0001, Chengchen Hu, Jiao Zhang 0002, Brent E. Stephens, Aditya Akella, Ying Zhang 0022
APNet10
2017 UNO: uniflying host and smart NIC offload for flexible packet processing
abstract
Increasingly, smart Network Interface Cards (sNICs) are being used in data centers to offload networking functions (NFs) from host processors thereby making these processors available for tenant applications. Modern sNICs have fully programmable, energy-efficient multi-core processors on which many packet processing functions, including a full-blown programmable switch, can run. However, having multiple switch instances deployed across the host hypervisor and the attached sNICs makes controlling them difficult and data plane operations more complex.
Yanfang Le, Hyunseok Chang, Sarit Mukherjee, Limin Wang 0010, Aditya Akella, Michael M. Swift, T. V. Lakshman
SoCC5
2017 Supporting Diverse Dynamic Intent-based Policies using Janus
abstract
Existing network policy abstractions handle basic group based reachability and access control list based security policies. However, QoS policies as well as dynamic policies are also important and not representing them in the high level policy abstraction poses serious limitations. At the same time, efficiently configuring and composing group based QoS and dynamic policies present significant technical challenges, such as (a) maintaining group granularity during configuration, (b) dealing with network-bandwidth contention among policies from distinct writers and (c) dealing with multiple path changes corresponding to dynamically changing policies, group membership and end-point mobility. In this paper we propose Janus, a system which makes two major contributions. First, we extend the prior policy graph abstraction model to represent complex QoS and dynamic tateful/temporal policies. Second, we convert the policy configuration problem into an optimization problem with the goal of maximizing the number of satisfied and configured policies, and minimizing the number of path changes under dynamic environments. To solve this, Janus presents several novel heuristic algorithms. We evaluate our system using a diverse set of bandwidth policies and network topologies. Our experiments demonstrate that Janus can achieve near-optimal solutions in a reasonable amount of time.
Anubhavnidhi Abhashkumar, Joon-Myung Kang, Sujata Banerjee, Aditya Akella, Ying Zhang 0022, Wenfei Wu
CoNEXT4
2017 Granular Computing and Network Intensive Applications: Friends or Foes?
abstract
Computing/infrastructure as a service continues to evolve with bare metal, virtual machines, containers and now serverless granular computing service offerings. Granular computing enables developers to decompose their applications into smaller logical units or functions, and run them on small, low cost and short lived computation containers without having to worry about setting up servers - hence the term serverless computing. While serverless environments can be used very cost effectively for large scale parallel processing data analytics applications, it is less clear if network intensive packet processing applications can also benefit from these new computing services as they do not share the same characteristics. This paper examines the architectural constraints as well as current serverless implementations to develop a position on this topic and influence the next generation of computing services. We support our position through measurement and experimentation on Amazon's AWS Lambda service with a few popular network functions.
Arjun Singhvi, Sujata Banerjee, Yotam Harchol, Aditya Akella, Mark Peek, Pontus Rydin
HotNets4
2017 Genesis: synthesizing forwarding tables in multi-tenant networks
abstract
Operators in multi-tenant cloud datacenters require support for diverse and complex end-to-end policies, such as, reachability, middlebox traversals, isolation, traffic engineering, and network resource management. We present Genesis, a datacenter network management system which allows policies to be specified in a declarative manner without explicitly programming the network data plane. Genesis tackles the problem of enforcing policies by synthesizing switch forwarding tables. It uses the formal foundations of constraint solving in combination with fast off-the-shelf SMT solvers. To improve synthesis performance, Genesis incorporates a novel search strategy that uses regular expressions to specify properties that leverage the structure of datacenter networks, and a divide-and-conquer synthesis procedure which exploits the structure of policy relationships. We have prototyped Genesis, and conducted experiments with a variety of workloads on real-world topologies to demonstrate its performance.
Kausik Subramanian, Loris D'Antoni, Aditya Akella
POPL3
2017 Bootstrapping evolvability for inter-domain routing with D-BGP
abstract
The Internet's inter-domain routing infrastructure, provided today by BGP, is extremely rigid and does not facilitate the introduction of new inter-domain routing protocols. This rigidity has made it incredibly difficult to widely deploy critical fixes to BGP. It has also depressed ASes' ability to sell value-added services or replace BGP entirely with a more sophisticated protocol. Even if operators undertook the significant effort needed to fix or replace BGP, it is likely the next protocol will be just as difficult to change or evolve. To help, this paper identifies two features needed in the routing infrastructure (i.e., within any inter-domain routing protocol) to facilitate evolution to new protocols. To understand their utility, it presents D-BGP, a version of BGP that incorporates them.
Raja R. Sambasivan, David Tran-Lam, Aditya Akella, Peter Steenkiste
SIGCOMM3
2017 Automatically Repairing Network Control Planes Using an Abstract Representation
abstract
The forwarding behavior of computer networks is governed by the configuration of distributed routing protocols and access filters---collectively known as the network control plane. Unfortunately, control plane configurations are often buggy, causing networks to violate important policies: e.g., specific traffic classes (defined in terms of source and destination endpoints) should always be able to reach their destination, or always traverse a waypoint. Manually repairing these configurations is daunting because of their inter-twined nature across routers, traffic classes, and policies.
Aaron Gember, Aditya Akella, Ratul Mahajan, Hongqiang Harry Liu
SOSP2
2017 Titan: Fair Packet Scheduling for Commodity Multiqueue NICs
Brent E. Stephens, Arjun Singhvi, Aditya Akella, Michael M. Swift
USENIX ATC3
2016 Paving the Way for NFV: Simplifying Middlebox Modifications Using StateAlyzr
Junaid Khalid, Aaron Gember, Roney Michael, Anubhavnidhi Abhashkumar, Aditya Akella
NSDI5
2016 Efficiently Delivering Online Services over Integrated Infrastructure
Hongqiang Harry Liu, Raajay Viswanathan, Matt Calder, Aditya Akella, Ratul Mahajan, Jitendra Padhye, Ming Zhang 0005
NSDI4
2016 Altruistic Scheduling in Multi-Resource Clusters
Robert Grandl, Mosharaf Chowdhury, Aditya Akella, Ganesh Ananthanarayanan
OSDI3
2016 GRAPHENE: Packing and Dependency-Aware Scheduling for Data-Parallel Clusters
Robert Grandl, Srikanth Kandula, Sriram Rao, Aditya Akella, Janardhan Kulkarni
OSDI4
2016 CLARINET: WAN-Aware Optimization for Analytics Queries
Raajay Viswanathan, Ganesh Ananthanarayanan, Aditya Akella
OSDI3
2016 Fast Control Plane Analysis Using an Abstract Representation
abstract
Networks employ complex, and hence error-prone, routing control plane configurations. In many cases, the impact of errors manifests only under failures and leads to devastating effects. Thus, it is important to proactively verify control plane behavior under arbitrary link failures. State-of-the-art verifiers are either too slow or impractical to use for such verification tasks. In this paper we propose a new high level abstraction for control planes, ARC, that supports fast control plane analyses under arbitrary failures. ARC can check key invariants without generating the data plane--which is the main reason for current tools' ineffectiveness. This is possible because of the nature of verification tasks and the constrained nature of control plane designs in networks today. We develop algorithms to derive a network's ARC from its configuration files. Our evaluation over 314 networks shows that ARC computation is quick, and that ARC can verify key invariants in under 1s in most cases, which is orders-of-magnitude faster than the state-of-the-art.
Aaron Gember, Raajay Viswanathan, Aditya Akella, Ratul Mahajan
SIGCOMM3
2016 AC/DC TCP: Virtual Congestion Control Enforcement for Datacenter Networks
abstract
Multi-tenant datacenters are successful because tenants can seamlessly port their applications and services to the cloud. Virtual Machine (VM) technology plays an integral role in this success by enabling a diverse set of software to be run on a unified underlying framework. This flexibility, however, comes at the cost of dealing with out-dated, inefficient, or misconfigured TCP stacks implemented in the VMs. This paper investigates if administrators can take control of a VM's TCP congestion control algorithm without making changes to the VM or network hardware. We propose AC/DC TCP, a scheme that exerts fine-grained control over arbitrary tenant TCP stacks by enforcing per-flow congestion control in the virtual switch (vSwitch). Our scheme is light-weight, flexible, scalable and can police non-conforming flows. In our evaluation the computational overhead of AC/DC TCP is less than one percentage point and we show implementing an administrator-defined congestion control algorithm in the vSwitch (i.e., DCTCP) closely tracks its native performance, regardless of the VM's TCP stack.
Keqiang He, Eric Rozner, Kanak Agarwal 0001, Yu Gu 0001, Wes Felter, John B. Carter, Aditya Akella
SIGCOMM7
2016 Enhancing Video Accessibility and Availability Using Information-Bound References
abstract
Users are often frustrated when they cannot view video links shared via blogs, social networks, and shared bookmark sites on their devices or suffer performance and usability problems when doing so. While other versions of the same content better suited to their device and network constraints may be available on other third-party hosting sites, these remain unusable because users cannot efficiently discover these and verify that these variants match the content publisher's original intent. Our vision is to enable consumers to leverage verifiable alternatives from different hosting sites that are best suited to their constraints to deliver a high quality of experience and enable content publishers to reach a wide audience with diverse operating conditions with minimal upfront costs. To this end, we make a case for information-bound references or IBRs that bind references to video content to the underlying information that a publisher wants to convey, decoupled from details such as protocols, hosts, file names, or the underlying bits. This paper addresses key challenges in the design and implementation of IBR generation and resolution mechanisms, and presents an evaluation of the benefits IBRs offer.
Ashok Anand, Athula Balachandran, Aditya Akella, Vyas Sekar, Srinivasan Seshan
IEEE/ACM Trans. Netw.3
2015 Seeing through Network-Protocol Obfuscation
abstract
Censorship-circumvention systems are designed to help users bypass Internet censorship. As more sophisticated deep-packet-inspection (DPI) mechanisms have been deployed by censors to detect circumvention tools, activists and researchers have responded by developing network protocol obfuscation tools. These have proved to be effective in practice against existing DPI and are now distributed with systems such as Tor. In this work, we provide the first in-depth investigation of the detectability of in-use protocol obfuscators by DPI. We build a framework for evaluation that uses real network traffic captures to evaluate detectability, based on metrics such as the false-positive rate against background (i.e., non obfuscated) traffic. We first exercise our framework to show that some previously proposed attacks from the literature are not as effective as a censor might like. We go on to develop new attacks against five obfuscation tools as they are configured in Tor, including: two variants of obfsproxy, FTE, and two variants of meek. We conclude by using our framework to show that all of these obfuscation mechanisms could be reliably detected by a determined censor with sufficiently low false-positive rates for use in many censorship settings.
Liang Wang 0023, Kevin P. Dyer, Aditya Akella, Thomas Ristenpart, Thomas Shrimpton
CCS3
2015 Bootstrapping Evolvability for Inter-Domain Routing
abstract
It is extremely difficult to deploy newinter-domain routing protocols in today's Internet. As a result, the Internet's baseline protocol for connectivity, BGP, has remained largely unchanged, despite known significant flaws. The difficulty of deploying new protocols has also depressed opportunities for (currently commoditized) transit providers to provide value-added routing services. To help, we identify the key deployment models under which new protocols are introduced and the requirements each poses for enabling their usage goals. Based on these requirements, we argue for two modifications to BGP that will greatly improve support for new routing protocols.
Raja R. Sambasivan, David Tran-Lam, Aditya Akella, Peter Steenkiste
HotNets3
2015 A case for application-managed cache for browser
abstract
Mobile web usage has significantly increased in last few years. There has been a lot of emphasis on providing good web page performance for mobile devices. Client-side caching can play a significant role in providing good web page performance, but unfortunately, traditional browser caches lack in various aspects leading to sub-optimal performance. More specifically, web applications do not have control on caching, e.g., which resources to cache, how to cache, etc., leading to ineffective cache utilization. Recently, HTML5 has introduced number of persistent storage APIs, that can provide required control for web applications. We evaluate these HTML5 storage options on various devices, and find that they can also meet the performance criteria of caching; in fact, some of the HTML5 storage APIs, e.g., localStorage, can provide even better performance than browser cache. Based on these insights, we make a case for application-managed hierarchical client-side cache, called HCache, that leverages these storage options as backends. We propose a novel API that allows web application developers to intelligently control the caching behavior and the usage of these storage options transparently. Our experiments with a prototype show that HCache can improve web page performance by up to 60%.
Ashok Anand, Mehrdad Reshadi, Bowei Du, Hariharan Kolam, Sharad Jaiswal, Aditya Akella
ICME6
2015 Management Plane Analytics
abstract
While it is generally held that network management is tedious and error-prone, it is not well understood which specific management practices increase the risk of failures. Indeed, our survey of 51 network operators reveals a significant diversity of opinions, and our characterization of the management practices in the 850+ networks of a large online service provider shows significant diversity in prevalent practices. Motivated by these observations, we develop a management plane analytics (MPA) framework that an organization can use to: (i) infer which management practices impact network health, and (ii) develop a predictive model of health, based on observed practices, to improve network management. We overcome the challenges of sparse and skewed data by aggregating data from many networks, reducing data dimensionality, and oversampling minority cases. Our learned models predict network health with an accuracy of 76-89%, and our causal analysis uncovers some high impact practices that operators thought had a low impact on network health. Our tool is publicly available, so organizations can analyze their own management practices.
Aaron Gember, Wenfei Wu, Xiujun Li, Aditya Akella, Ratul Mahajan
Internet Measurement Conference4
2015 PerfSight: Performance Diagnosis for Software Dataplanes
abstract
The advent of network functions virtualization (NFV) means that data planes are no longer simply composed of routers and switches. Instead they are very complex and involve a variety of sophisticated packet processing elements that reside on the OSes and software running on compute servers where network functions (NFs) are hosted. In this paper, we argue that these new "software data planes" are susceptible to at least three new classes of performance problems. To diagnose such problems, we design, implement and evaluate, PerfSight, a ground-up system that works by extracting comprehensive low-level information regarding packet processing and I/O performance of the various elements in the software data plane. Name then analyzes the information gathered in various dimensions (e.g., across all VMs on a machine, or all VMs deployed by a tenant). By looking across aggregates, we show that it becomes possible to detect and diagnose key performance problems. Experimental results show that our framework can result in accurate detection of the root causes of key performance problems in software data planes, and it imposes very little overhead.
Wenfei Wu, Keqiang He, Aditya Akella
Internet Measurement Conference3
2015 Presto: Edge-based Load Balancing for Fast Datacenter Networks
abstract
Datacenter networks deal with a variety of workloads, ranging from latency-sensitive small flows to bandwidth-hungry large flows. Load balancing schemes based on flow hashing, e.g., ECMP, cause congestion when hash collisions occur and can perform poorly in asymmetric topologies. Recent proposals to load balance the network require centralized traffic engineering, multipath-aware transport, or expensive specialized hardware. We propose a mechanism that avoids these limitations by (i) pushing load-balancing functionality into the soft network edge (e.g., virtual switches) such that no changes are required in the transport layer, customer VMs, or networking hardware, and (ii) load balancing on fine-grained, near-uniform units of data (flowcells) that fit within end-host segment offload optimizations used to support fast networking speeds. We design and implement such a soft-edge load balancing scheme, called Presto, and evaluate it on a 10 Gbps physical testbed. We demonstrate the computational impact of packet reordering on receivers and propose a mechanism to handle reordering in the TCP receive offload functionality. Presto's performance closely tracks that of a single, non-blocking switch over many workloads and is adaptive to failures and topology asymmetry.
Keqiang He, Eric Rozner, Kanak Agarwal 0001, Wes Felter, John B. Carter, Aditya Akella
SIGCOMM6
2015 Network Policy Whiteboarding and Composition
abstract
We present Policy Graph Abstraction (PGA) that graphically expresses network policies and service chain requirements, just as simple as drawing whiteboard diagrams. Different users independently draw policy graphs that can constrain each other. PGA graph clearly captures user intents and invariants and thus facilitates automatic composition of overlapping policies into a coherent policy.
Jeongkeun Lee, Joon-Myung Kang, Chaithan Prakash, Yoshio Turner, Aditya Akella, Charles Clark, Yadi Ma, Puneet Sharma 0001, Ying Zhang 0022
SIGCOMM5
2015 PGA: Using Graphs to Express and Automatically Reconcile Network Policies
abstract
Software Defined Networking (SDN) and cloud automation enable a large number of diverse parties (network operators, application admins, tenants/end-users) and control programs (SDN Apps, network services) to generate network policies independently and dynamically. Yet existing policy abstractions and frameworks do not support natural expression and automatic composition of high-level policies from diverse sources. We tackle the open problem of automatic, correct and fast composition of multiple independently specified network policies. We first develop a high-level Policy Graph Abstraction (PGA) that allows network policies to be expressed simply and independently, and leverage the graph structure to detect and resolve policy conflicts efficiently. Besides supporting ACL policies, PGA also models and composes service chaining policies, i.e., the sequence of middleboxes to be traversed, by merging multiple service chain requirements into conflict-free composed chains. Our system validation using a large enterprise network policy dataset demonstrates practical composition times even for very large inputs, with only sub-millisecond runtime latencies.
Chaithan Prakash, Jeongkeun Lee, Yoshio Turner, Joon-Myung Kang, Aditya Akella, Sujata Banerjee, Charles Clark, Yadi Ma, Puneet Sharma 0001, Ying Zhang 0022
SIGCOMM5
2015 Low Latency Geo-distributed Data Analytics
abstract
Low latency analytics on geographically distributed datasets (across datacenters, edge clusters) is an upcoming and increasingly important challenge. The dominant approach of aggregating all the data to a single datacenter significantly inflates the timeliness of analytics. At the same time, running queries over geo-distributed inputs using the current intra-DC analytics frameworks also leads to high query response times because these frameworks cannot cope with the relatively low and variable capacity of WAN links. We present Iridium, a system for low latency geo-distributed analytics. Iridium achieves low query response times by optimizing placement of both data and tasks of the queries. The joint data and task placement optimization, however, is intractable. Therefore, Iridium uses an online heuristic to redistribute datasets among the sites prior to queries' arrivals, and places the tasks to reduce network bottlenecks during the query's execution. Finally, it also contains a knob to budget WAN usage. Evaluation across eight worldwide EC2 regions using production queries show that Iridium speeds up queries by 3× -- 19× and lowers WAN usage by 15% -- 64% compared to existing baselines.
Qifan Pu, Ganesh Ananthanarayanan, Peter Bodík, Srikanth Kandula, Aditya Akella, Paramvir Bahl, Ion Stoica
SIGCOMM5
2015 Latency in Software Defined Networks: Measurements and Mitigation Techniques
abstract
We conduct a comprehensive measurement study of switch control plane latencies using four types of production SDN switches. Our measurements show that control actions, such as rule installation, have surprisingly high latency, due to both software implementation inefficiencies and fundamental traits of switch hardware. We also propose three measurement-driven latency mitigation techniques---optimizing route selection, spreading rules across switches, and reordering rule installations---to effectively tame the flow setup latencies in SDN.
Keqiang He, Junaid Khalid, Aaron Gember, Chaithan Prakash, Aditya Akella, Li Erran Li, Marina Thottan
SIGMETRICS6
2014 Design patterns for tunable and efficient SSD-based indexes
abstract
A number of data-intensive systems require using random hash-based indexes of various forms, e.g., hash tables, Bloom filters, and locality sensitive hash tables. In this paper, we present general SSD optimization techniques that can be used to design a variety of such indexes while ensuring higher performance and easier tunability than specialized state-of-the-art approaches. We leverage two key SSD innovations: a) rearranging the data layout on the SSD to combine multiple read requests into one page read, and b) intelligently reordering requests to exploit inherent parallelism in the architecture of SSDs. We build three different indexes using these techniques, and we conduct extensive studies showing their superior performance, lower CPU/memory footprint, and tunability compared to state-of-the-art systems.
Ashok Anand, Aaron Gember, Collin Engstrom, Aditya Akella
ANCS4
2014 A Highly Available Software Defined Fabric
abstract
Existing SDNs rely on a collection of intricate, mutually-dependent mechanisms to implement a logically centralized control plane. These cyclical dependencies and lack of clean separation of concerns can impact the availability of SDNs, such that a handful of link failures could render entire portions of an SDN non-functional. This paper shows why and when this could happen, and makes the case for taking a fresh look at architecting SDNs for robustness to faults from the ground up. Our approach carefully synthesizes various key distributed systems ideas -- in particular, reliable flooding, global snapshots, and replicated controllers. We argue informally that it can offer high availability in the face of a variety of network failures, but much work needs to be done to make our approach scalable and general. Thus, our paper represents a starting point for a broader discussion on approaches for building highly available SDNs.
Aditya Akella, Arvind Krishnamurthy
HotNets1
2014 A Call to Arms for Management Plane Analytics
abstract
Over the last few decades, the networking community has developed numerous techniques for understanding how real networks behave through analyzing their data and control planes. In this paper, we call upon the community to similarly develop techniques to analyze the network management plane, that is, activities that underlie network design and operation. Such analytics can shed light on why a network behaves as observed and the relative merits of different management practices. While the management plane is often not directly observable, we argue that many relevant aspects can be inferred through data that most networks already gather (e.g., snapshots of configurations). Using preliminary analysis of such data from many large networks, we demonstrate the feasibility and the value of management plane analytics.
Aditya Akella, Ratul Mahajan
HotNets1
2014 WhoWas: A Platform for Measuring Web Deployments on IaaS Clouds
abstract
Public infrastructure-as-a-service (IaaS) clouds such as Amazon EC2 and Microsoft Azure host an increasing number of web services. The dynamic, pay-as-you-go nature of modern IaaS systems enable web services to scale up or down with demand, and only pay for the resources they need. We are unaware, however, of any studies reporting on measurements of the patterns of usage over time in IaaS clouds as seen in practice. We fill this gap, offering a measurement platform that we call WhoWas. Using active, but lightweight, probing, it enables associating web content to public IP addresses on a day-by-day basis. We exercise WhoWas to provide the first measurement study of churn rates in EC2 and Azure, the efficacy of IP blacklists for malicious activity in clouds, the rate of adoption of new web software by public cloud customers, and more.
Liang Wang 0023, Antonio Nappa, Juan Caballero, Thomas Ristenpart, Aditya Akella
Internet Measurement Conference5
2014 OpenNF: enabling innovation in network function control
abstract
Network functions virtualization (NFV) together with software-defined networking (SDN) has the potential to help operators satisfy tight service level agreements, accurately monitor and manipulate network traffic, and minimize operating expenses. However, in scenarios that require packet processing to be redistributed across a collection of network function (NF) instances, simultaneously achieving all three goals requires a framework that provides efficient, coordinated control of both internal NF state and network forwarding state. To this end, we design a control plane called OpenNF. We use carefully designed APIs and a clever combination of events and forwarding updates to address race conditions, bound overhead, and accommodate a variety of NFs. Our evaluation shows that OpenNF offers efficient state control without compromising flexibility, and requires modest additions to NFs.
Aaron Gember, Raajay Viswanathan, Chaithan Prakash, Robert Grandl, Junaid Khalid, Aditya Akella
SIGCOMM7
2014 Multi-resource packing for cluster schedulers
abstract
Tasks in modern data parallel clusters have highly diverse resource requirements, along CPU, memory, disk and network. Any of these resources may become bottlenecks and hence, the likelihood of wasting resources due to fragmentation is now larger. Today's schedulers do not explicitly reduce fragmentation. Worse, since they only allocate cores and memory, the resources that they ignore (disk and network) can be over-allocated leading to interference, failures and hogging of cores or memory that could have been used by other tasks. We present Tetris, a cluster scheduler that packs, i.e., matches multi-resource task requirements with resource availabilities of machines so as to increase cluster efficiency (makespan). Further, Tetris uses an analog of shortest-running-time-first to trade-off cluster efficiency for speeding up individual jobs. Tetris' packing heuristics seamlessly work alongside a large class of fairness policies. Trace-driven simulations and deployment of our prototype on a 250 node cluster shows median gains of 30% in job completion time while achieving nearly perfect fairness.
Robert Grandl, Ganesh Ananthanarayanan, Srikanth Kandula, Sriram Rao, Aditya Akella
SIGCOMM5
2013 Harmony: coordinating network, compute, and storage in software-defined clouds
abstract
The progress of a big data job is often a function of storage, networking and processing. Hence, for efficient job execution, it is important to collectively optimize all three components. Prior proposals [1], in contrast, have focused on mainly on one or two of the three components. This narrow focus constraints the extent to which these proposals can support efficient operation of big data applications.
Robert Grandl, Yizheng Chen 0005, Junaid Khalid, Suli Yang, Ashok Anand, Theophilus Benson, Aditya Akella
SoCC7
2013 Virtual network diagnosis as a service
abstract
Today's cloud network platforms allow tenants to construct sophisticated virtual network topologies among their VMs on a shared physical network infrastructure. However, these platforms provide little support for tenants to diagnose problems in their virtual networks. Network virtualization hides the underlying infrastructure from tenants as well as prevents deploying existing network diagnosis tools. This paper makes a case for providing virtual network diagnosis as a service in the cloud. We identify a set of technical challenges in providing such a service and propose a Virtual Network Diagnosis (VND) framework. VND exposes abstract configuration and query interfaces for cloud tenants to troubleshoot their virtual networks. It controls software switches to collect flow traces, distributes traces storage, and executes distributed queries for different tenants for network diagnosis. It reduces the data collection and processing overhead by performing local flow capture and on-demand query execution. Our experiments validate VND's functionality and shows its feasibility in terms of quick service response and acceptable overhead; our simulation proves the VND architecture scales to the size of a real data center network.
Wenfei Wu, Aditya Akella, Anees Shaikh
SoCC3
2013 Enhancing video accessibility and availability using information-bound references
abstract
Users are often frustrated when they cannot view video links shared via blogs, social networks, and shared bookmark sites on their devices or suffer performance and usability problems when doing so. While other versions of the same content better suited to their device and network constraints may be available on other third-party hosting sites, these remain unusable because users cannot efficiently discover these and verify that these variants match the content publisher's original intent. Our vision is to enable consumers to leverage verifiable alternatives from different hosting sites that are best suited to their constraints to deliver a high quality of experience and enable content publishers to reach a wide audience with diverse operating conditions with minimal upfront costs. To this end, we make a case for information-bound references or IBRs that bind references to video content to the underlying information that a publisher wants to convey, decoupled from details such as protocols, hosts, file names, or the underlying bits. This paper addresses key challenges in the design and implementation of IBR generation and resolution mechanisms, and presents an evaluation of the benefits IBRs offer.
Ashok Anand, Athula Balachandran, Aditya Akella, Vyas Sekar, Srinivasan Seshan
CoNEXT3
2013 Analyzing the potential benefits of CDN augmentation strategies for internet video workloads
abstract
Video viewership over the Internet is rising rapidly, and market predictions suggest that video will comprise over 90\% of Internet traffic in the next few years. At the same time, there have been signs that the Content Delivery Network (CDN) infrastructure is being stressed by ever-increasing amounts of video traffic. To meet these growing demands, the CDN infrastructure must be designed, provisioned and managed appropriately. Federated telco-CDNs and hybrid P2P-CDNs are two content delivery infrastructure designs that have gained significant industry attention recently. We observed several user access patterns that have important implications to these two designs in our unique dataset consisting of 30 million video sessions spanning around two months of video viewership from two large Internet video providers. These include partial interest in content, regional interests, temporal shift in peak load and patterns in evolution of interest. We analyze the impact of our findings on these two designs by performing a large scale measurement study. Surprisingly, we find significant amount of synchronous viewing behavior for Video On Demand (VOD) content, which makes hybrid P2P-CDN approach feasible for VOD and suggest new strategies for CDNs to reduce their infrastructure costs. We also find that federation can significantly reduce telco-CDN provisioning costs by as much as 95%.
Athula Balachandran, Vyas Sekar, Aditya Akella, Srinivasan Seshan
Internet Measurement Conference3
2013 Next stop, the cloud: understanding modern web service deployment in EC2 and azure
abstract
An increasingly large fraction of Internet services are hosted on a cloud computing system such as Amazon EC2 or Windows Azure. But to date, no in-depth studies about cloud usage by Internet services has been performed. We provide a detailed measurement study to shed light on how modern web service deployments use the cloud and to identify ways in which cloud-using services might improve these deployments. Our results show that: 4% of the Alexa top million use EC2/Azure; there exist several common deployment patterns for cloud-using web service front ends; and services can significantly improve their wide-area performance and failure tolerance by making better use of existing regional diversity in EC2. Driving these analyses are several new datasets, including one with over 34 million DNS records for Alexa websites and a packet capture from a large university network.
Keqiang He, Alexis Fisher, Liang Wang 0023, Aaron Gember, Aditya Akella, Thomas Ristenpart
Internet Measurement Conference5
2013 Adaptive data transmission in the cloud
abstract
Data centers provide resources for a broad range of services, such as web search, email, web sites, etc., each with different delay requirements. For example, web search should cater to users' requests quickly, while data backup has no special requirement on completion time. Different applications also introduce flows with very different properties (e.g., size and duration). The default method of transport in data centers, namely TCP, treats flows equally, forcing equal share of the bottleneck network bandwidth. This fairness property leads to poor outcomes for time-sensitive applications. A better solution is to allocate more bandwidth to time-sensitive applications. However, the state-of-the-art approaches that do this all require forklift changes to data center networking gear. In some cases, substantial changes need to be made to end-system stacks and applications as well. In this paper, we argue that a simple modification to TCP can help better meet the requirements of latency-sensitive applications in the data center. No modification to end-systems, applications or networking gear is necessary. We motivate our Adaptive TCP (ATCP) design using measurements of real data center traffic. We analytically derive the parameters to use in our proposed modification to TCP. Finally, we use extensive simulations in NS2 to show the benefits of ATCP.
Wenfei Wu, Yizheng Chen 0005, Ramakrishnan Durairajan, Ashok Anand, Aditya Akella
IWQoS6
2013 An information-aware QoE-centric mobile video cache
abstract
Recent years have seen a tremendous growth in the volume of video traffic in mobile settings. In this paper, we present the design of a mobile video-centric proxy cache, named iProxy, that offers improved performance in terms of both hit rates and streaming quality. Our thesis in designing iProxy is that we need to elevate the traditional view of caching from "data" to "information" in order to optimally meet the stringent requirements of video streaming in mobile settings. iProxy relies on recent advances on information-bound references (IBRs) to collapse multiple related cache entries into a single one, improving hitrate while lowering storage costs. iProxy incorporates a novel dynamic linear rate adaptation scheme to ensure high stream quality in face of channel diversity and device heterogeneity. Our evaluation of iProxy using realistic traffic traces shows that it can improve hitrate, but we need to use novel information-aware replacement policies for optimal performance. We show that our linear encoder can adapt well to changes in bandwidth, and yield better bit rates, lower buffering and lower start up delays than state-of-the-art schemes.
Shan-Hsiang Shen, Aditya Akella
MobiCom2
2013 Developing a predictive model of quality of experience for internet video
abstract
Improving users' quality of experience (QoE) is crucial for sustaining the advertisement and subscription based revenue models that enable the growth of Internet video. Despite the rich literature on video and QoE measurement, our understanding of Internet video QoE is limited because of the shift from traditional methods of measuring video quality (e.g., Peak Signal-to-Noise Ratio) and user experience (e.g., opinion scores). These have been replaced by new quality metrics (e.g., rate of buffering, bitrate) and new engagement centric measures of user experience (e.g., viewing time and number of visits). The goal of this paper is to develop a predictive model of Internet video QoE. To this end, we identify two key requirements for the QoE model: (1) it has to be tied in to observable user engagement and (2) it should be actionable to guide practical system design decisions. Achieving this goal is challenging because the quality metrics are interdependent, they have complex and counter-intuitive relationships to engagement measures, and there are many external factors that confound the relationship between quality and engagement (e.g., type of video, user connectivity). To address these challenges, we present a data-driven approach to model the metric interdependencies and their complex relationships to engagement, and propose a systematic framework to identify and account for the confounding factors. We show that a delivery infrastructure that uses our proposed model to choose CDN and bitrates can achieve more than 20\% improvement in overall user engagement compared to strawman approaches.
Athula Balachandran, Vyas Sekar, Aditya Akella, Srinivasan Seshan, Ion Stoica, Hui Zhang 0001
SIGCOMM3
2013 Design and implementation of a framework for software-defined middlebox networking
abstract
No abstract available.
Aaron Gember, Robert Grandl, Junaid Khalid, Aditya Akella
SIGCOMM4
2013 FCP: a flexible transport framework for accommodating diversity
abstract
Transport protocols must accommodate diverse application and network requirements. As a result, TCP has evolved over time with new congestion control algorithms such as support for generalized AIMD, background flows, and multipath. On the other hand, explicit congestion control algorithms have been shown to be more efficient. However, they are inherently more rigid because they rely on in-network components. Therefore, it is not clear whether they can be made flexible enough to support diverse application requirements. This paper presents a flexible framework for network resource allocation, called FCP, that accommodates diversity by exposing a simple abstraction for resource allocation. FCP incorporates novel primitives for end-point flexibility (aggregation and preloading) into a single framework and makes economics-based congestion control practical by explicitly handling load variations and by decoupling it from actual billing. We show that FCP allows evolution by accommodating diversity and ensuring coexistence, while being as efficient as existing explicit congestion control algorithms.
Dongsu Han, Robert Grandl, Aditya Akella, Srinivasan Seshan
SIGCOMM3
2013 Understanding internet video viewing behavior in the wild
abstract
Over the past few years video viewership over the Internet has risen dramatically and market predictions suggest that video will account for more than 50% of the traffic over the Internet in the next few years. Unfortunately, there has been signs that the Content Delivery Network (CDN) infrastructure is being stressed with the increasing video viewership load. Our goal in this paper is to provide a first step towards a principled understanding of how the content delivery infrastructure must be designed and provisioned to handle the increasing workload by analyzing video viewing behaviors and patterns in the wild. We analyze various viewing behaviors using a dataset consisting of over 30 million video sessions spanning two months of viewership from two large Internet video providers. In these preliminary results, we observe viewing patterns that have significant impact on the design of the video delivery infrastructure.
Athula Balachandran, Vyas Sekar, Aditya Akella, Srinivasan Seshan
SIGMETRICS3
2012 ECOS: leveraging software-defined networks to support mobile application offloading
abstract
Offloading has emerged as a promising idea to allow resource-constrained mobile devices to access intensive applications, without performance or energy costs, by leveraging external computing resources. This could be particularly useful in enterprise contexts where running line-of-business applications on mobile devices can enhance enterprise operations. However, we must address three practical roadblocks to make offloading amenable to adoption by enterprises: (i) ensuring privacy and trustworthiness of offload, (ii) decoupling offloading systems from their reliance on the availability of dedicated resources and (iii) accommodating offload at scale. We present the design and implementation of ECOS, an enterprise-centric offloading framework that leverages Software-Defined Networking to augment prior offloading proposals and address these limitations. ECOS functions as an application running at an enterprise-wide controller to allocate resources to mobile applications based on privacy and performance requirements, to ensure fairness, and to enforce security constraints. Experiments using a prototype based on Android and OpenFlow establish the effectiveness of our approach.
Aaron Gember, Chris Dragga, Aditya Akella
ANCS3
2012 A quest for an Internet video quality-of-experience metric
abstract
An imminent challenge that content providers, CDNs, third-party analytics and optimization services, and video player designers in the Internet video ecosystem face is the lack of a single "gold standard" to evaluate different competing solutions. Existing techniques that describe the quality of the encoded signal or controlled studies to measure opinion scores do not translate directly into user experience at scale. Recent work shows that measurable performance metrics such as buffering, startup time, bitrate, and number of bitrate switches impact user experience. However, converting these observations into a quantitative quality-of-experience metric turns out to be challenging since these metrics are interrelated in complex and sometimes counter-intuitive ways, and their relationship to user experience can be unpredictable. To further complicate things, many confounding factors are introduced by the nature of the content itself (e.g., user interest, genre). We believe that the issue of interdependency can be addressed by casting this as a machine learning problem to build a suitable predictive model from empirical observations. We also show that setting up the problem based on domain-specific and measurement-driven insights can minimize the impact of the various confounding factors to improve the prediction performance.
Athula Balachandran, Vyas Sekar, Aditya Akella, Srinivasan Seshan, Ion Stoica, Hui Zhang 0001
HotNets3
2012 Toward software-defined middlebox networking
abstract
Current middlebox (MB) management mechanisms are clumsy and unsuitable for taking full advantage of new MB deployment models and diverse MB functionality. Instead, we advocate for mechanisms that help exercise unified control over the key factors influencing MB operations. Our goal is to realize a software-defined MB networking framework to simplify management of complex, diverse functionalities and engender rich deployments. We discuss the major challenges that arise---representing, manipulating, and knowledgeably controlling MB state---and we present initial thoughts on the appropriate abstractions and interfaces to address them.
Aaron Gember, Prathmesh Prabhu, Zainab Ghadiyali, Aditya Akella
HotNets4
2012 Obtaining in-context measurements of cellular network performance
abstract
Network service providers, and other parties, require an accurate understanding of the performance cellular networks deliver to users. In particular, they often seek a measure of the network performance users experience solely when they are interacting with their device---a measure we call in-context. Acquiring such measures is challenging due to the many factors, including time and physical context, that influence cellular network performance. This paper makes two contributions. First, we conduct a large scale measurement study, based on data collected from a large cellular provider and from hundreds of controlled experiments, to shed light on the issues underlying in-context measurements. Our novel observations show that measurements must be conducted on devices which (i) recently used the network as a result of user interaction with the device, (ii) remain in the same macro-environment (e.g., indoors and stationary), and in some cases the same micro-environment (e.g., in the user's hand), during the period between normal usage and a subsequent measurement, and (iii) are currently sending/ receiving little or no user-generated traffic. Second, we design and deploy a prototype active measurement service for Android phones based on these key insights. Our analysis of 1650 measurements gathered from 12 volunteer devices shows that the system is able to obtain average throughput measurements that accurately quantify the performance experienced during times of active device and network usage.
Aaron Gember, Aditya Akella, Jeffrey Pang, Alexander Varshavsky, Ramón Cáceres
Internet Measurement Conference2
2012 DECOR: A distributed coordinated resource monitoring system
abstract
Network resources are often limited, so how to use them efficiently is an issue that arises in many important scenarios. Many recent proposals rely on a central controller to carefully orchestrate resources across multiple network locations. The central controller gathers network information and relative levels of usage of different resources and calculates optimized task allocation arrangements to maximize some global benefit. Examples of architectures that use this framework include coordinated sampling (cSamp [1]) and redundancy elimination (SmartRE [2]). However, a centralized solution creates practical problems as it is susceptible to overload, and the controller is a single point of failure. In this paper, we present a distributed solution called decor that achieves global optimization based on local information that closes to centralized approaches in terms of performance. In decor, the responsibility of resource monitoring and information gathering is spread among multiple nodes; thus, no single point is overloaded. Allocation of tasks is also done in a similar distributed fashion. decor can easily scale up to large networks, and the partial network failures do not affect DECOR's functioning in other parts of the network. decor can be applied to most of path-based applications. We describe in detail how to apply it to distributed SmartRE and implement it in the Click software router.
Shan-Hsiang Shen, Aditya Akella
IWQoS2
2012 RPT: Re-architecting Loss Protection for Content-Aware Networks
Dongsu Han, Ashok Anand, Aditya Akella, Srinivasan Seshan
NSDI3
2012 XIA: Efficient Support for Evolvable Internetworking
Dongsu Han, Ashok Anand, Fahad R. Dogar, Hyeontaek Lim, Michel Machado, Arvind Mukundan, Wenfei Wu, Aditya Akella, David G. Andersen, John W. Byers, Srinivasan Seshan, Peter Steenkiste
NSDI9
2011 CloudNaaS: a cloud networking platform for enterprise applications
abstract
Enterprises today face several challenges when hosting line-of-business applications in the cloud. Central to many of these challenges is the limited support for control over cloud network functions, such as, the ability to ensure security, performance guarantees or isolation, and to flexibly interpose middleboxes in application deployments. In this paper, we present the design and implementation of a novel cloud networking system called CloudNaaS. Customers can leverage CloudNaaS to deploy applications augmented with a rich and extensible set of network functions such as virtual network isolation, custom addressing, service differentiation, and flexible interposition of various middleboxes. CloudNaaS primitives are directly implemented within the cloud infrastructure itself using high-speed programmable network elements, making CloudNaaS highly efficient. We evaluate an OpenFlow-based prototype of CloudNaaS and find that it can be used to instantiate a variety of network functions in the cloud, and that its performance is robust even in the face of large numbers of provisioned services and link/device failures.
Theophilus Benson, Aditya Akella, Anees Shaikh, Sambit Sahu
SoCC2
2011 MicroTE: fine grained traffic engineering for data centers
abstract
The effects of data center traffic characteristics on data center traffic engineering is not well understood. In particular, it is unclear how existing traffic engineering techniques perform under various traffic patterns, namely how do the computed routes differ from the optimal routes. Our study reveals that existing traffic engineering techniques perform 15% to 20% worse than the optimal solution. We find that these techniques suffer mainly due to their inability to utilize global knowledge about flow characteristics and make coordinated decision for scheduling flows.
Theophilus Benson, Ashok Anand, Aditya Akella, Ming Zhang 0005
CoNEXT3
2011 XIA: an architecture for an evolvable and trustworthy internet
abstract
Motivated by limitations in today's host-based IP network architecture, recent studies have proposed clean-slate network architectures centered around alternative first-class principals, such as content, services, or users. However, much like the host-centric IP design, elevating one principal type above others hinders communication between other principals and inhibits the network's capability to evolve. Our work presents the eXpressive Internet Architecture (XIA), an architecture with native support for multiple principals and the ability to evolve its functionality to accommodate new, as yet unforeseen, principals over time. XIA also provides intrinsic security: communicating entities validate that their underlying intent was satisfied correctly without relying on external databases or configuration.
Ashok Anand, Fahad R. Dogar, Dongsu Han, Hyeontaek Lim, Michel Machado, Wenfei Wu, Aditya Akella, David G. Andersen, John W. Byers, Srinivasan Seshan, Peter Steenkiste
HotNets8
2011 The evolution of network configuration: a tale of two campuses
abstract
Studying network configuration evolution can improve our understanding of the evolving complexity of networks and can be helpful in making network configuration less error-prone. Unfortunately, the nature of changes that operators make to network configuration is poorly understood. Towards improving our understanding, we examine and analyze five years of router, switch, and firewall configurations from two large campus networks using the logs from version control systems used to store the configurations. We study how network configuration is distributed across different network operations tasks and how the configuration for each task evolves over time, for different types of devices and for different locations in the network. To understand the trends of how configuration evolves over time, we study the extent to which configuration for various tasks are added, modified, or deleted. We also study whether certain devices experience configuration changes more frequently than others, as well as whether configuration changes tend to focus on specific portions of the configuration (or on specific tasks). We also investigate when network operators make configuration changes of various types. Our results concerning configuration changes can help the designers of configuration languages understand which aspects of configuration might be more automated or tested more rigorously and may ultimately help improve configuration languages.
Hyojoon Kim, Theophilus Benson, Aditya Akella, Nick Feamster
Internet Measurement Conference3
2011 REfactor-ing content overhearing to improve wireless performance
abstract
Many systems have leveraged the broadcast nature of wireless radios to improve wireless capacity and performance. While conventional approaches have focused on overhearing entire packets, recent designs have argued that focusing on overheard content may be more effective. Unfortunately, key design choices in these approaches limit them from fully leveraging the benefits of overhearing content. We propose a cleaner refactoring of functionality where-in overhearing is realized at the sub-packet payload level through the use of IP-layer redundancy elimination. We show that this dramatically improves the effectiveness of prior overhearing based approaches and enables new designs, e.g., enhanced network coding, where content overhearing can be more effectively integrated to improve performance. Realizing the benefits of IP-layer content overhearing requires us to overcome challenges arising from the probabilistic nature of wireless reception (which could lead to inconsistent state) and the limited resources on wireless devices. We overcome these challenges through careful data structure and wireless redundancy elimination designs. We evaluate the effectiveness of our system using experimentation on real traces. We find that our design is highly effective: e.g., it can improve goodput by nearly 25% and air time utilization by nearly 20%.
Shan-Hsiang Shen, Aaron Gember, Ashok Anand, Aditya Akella
MobiCom4
2011 A Comparative Study of Handheld and Non-handheld Traffic in Campus Wi-Fi Networks
Aaron Gember, Ashok Anand, Aditya Akella
PAM3
2011 Demystifying configuration challenges and trade-offs in network-based ISP services
abstract
ISPs are increasingly offering a variety of network-based services such as VPN, VPLS, VoIP, Virtual-Wire and DDoS protection. Although both enterprise and residential networks are rapidly adopting these services, there is little systematic work on the design challenges and trade-offs ISPs face in providing them. The goal of our paper is to understand the complexity underlying the layer-3 design of services and to highlight potential factors that hinder their introduction, evolution and management. Using daily snapshots of configuration and device metadata collected from a tier-1 ISP, we examine the logical dependencies and special cases in device configurations for five different network-based services. We find: (1) the design of the core data-plane is usually service-agnostic and simple, but the control-planes for different services become more complex as services evolve; (2) more crucially, the configuration at the service edge inevitably becomes more complex over time, potentially hindering key management issues such as service upgrades and troubleshooting; and (3) there are key service-specific issues that also contribute significantly to the overall design complexity. Thus, the high prevalent complexity could impede the adoption and growth of network-based services. We show initial evidence that some of the complexity can be mitigated systematically.
Theophilus Benson, Aditya Akella, Aman Shaikh
SIGCOMM2
2011 De-ossifying internet routing through intrinsic support for end-network and ISP selfishness
abstract
We present the S4R supplemental routing system to address the constraints BGP places on ISPs and stub network alike. Technical soundness and economic viability are equal first class design requirements for S4R. In S4R, ISPs announce links connecting different parts of the Internet. ISPs can selfishly price their links to attract maximal amount of traffic. Stub networks can selfishly select paths that best meet their requirements at the lowest cost. We design a variety of practical algorithms for ISP and stub network response that strike a balance between accommodating selfishness of all participants and ensuring efficient and stable operation overall. We employ large scale simulations over realistic scenarios to show that S4R operates at a close-to-optimal state and that it encourages broad participation from stubs and ISPs.
Aditya Akella, Shuchi Chawla 0001, Holly Esquivel, Chitra Muthukrishnan
SIGMETRICS1
2010 A case for information-bound referencing
abstract
Links and content references form the foundation of the way that users interact today. Unfortunately, the links used today (URLs) are fragile since they tightly specify a protocol, host, and filename. Some past efforts have decoupled this binding to a certain degree; e.g., creating links that bind to byte-level data. We argue that these systems do not go far enough. Our key observation is that users really care about the intent of the referenced link and are relatively agnostic to the byte-level representation. Based on this observation, we argue that references should be bound to the underlying information associated with the referenced content. We call such references Information-Bound References (IBR). In this paper, we focus on the challenges of creating IBRs for multimedia data, since these form a dominant fraction of Internet traffic today. We explore the trade-offs of various alternatives for generating and using IBRs. We identify that it is possible to adapt multimedia fingerprinting algorithms in the literature to generate IBRs.
Ashok Anand, Aditya Akella, Vyas Sekar, Srinivasan Seshan
HotNets2
2010 Using strongly typed networking to architect for tussle
abstract
Today's networks discriminate towards or against traffic for a wide range of reasons, and in response end users and their applications increasingly attempt to evade monitoring and control, resulting in an ongoing tussle whose roots run deep. In this work we explore an architectural paradigm that can accommodate such tussles in a systematic and transparent fashion. The key idea at the core of our design is strongly typed networking: the notion that application messages contain type information that fully describes the content being transferred. Our framework allows for transparency between parties which then leads to dialog and choice for both users and service providers. While in the early stages, we provide a possible framework for directly addressing the tussle between end users and "the network" without resorting to an ever-increasing degree of obfuscation and inference.
Chitra Muthukrishnan, Vern Paxson, Mark Allman, Aditya Akella
HotNets4
2010 Network traffic characteristics of data centers in the wild
abstract
Although there is tremendous interest in designing improved networks for data centers, very little is known about the network-level traffic characteristics of data centers today. In this paper, we conduct an empirical study of the network traffic in 10 data centers belonging to three different categories, including university, enterprise campus, and cloud data centers. Our definition of cloud data centers includes not only data centers employed by large online service providers offering Internet-facing applications but also data centers used to host data-intensive (MapReduce style) applications). We collect and analyze SNMP statistics, topology and packet-level traces. We examine the range of applications deployed in these data centers and their placement, the flow-level and packet-level transmission properties of these applications, and their impact on network and link utilizations, congestion and packet drops. We describe the implications of the observed traffic patterns for data center internal traffic engineering as well as for recently proposed architectures for data center networks.
Theophilus Benson, Aditya Akella, David A. Maltz
Internet Measurement Conference2
2010 EndRE: An End-System Redundancy Elimination Service for Enterprises
Bhavish Agarwal, Aditya Akella, Ashok Anand, Athula Balachandran, Pushkar V. Chitnis, Chitra Muthukrishnan, Ramachandran Ramjee, George Varghese
NSDI2
2010 Cheap and Large CAMs for High Performance Data-Intensive Networked Systems
Ashok Anand, Chitra Muthukrishnan, Steven Kappes, Aditya Akella, Suman Nath
NSDI4
2010 Flexible multimedia content retrieval using InfoNames
abstract
Multimedia content is a dominant fraction of Internet usage today. At the same time, there is significant heterogeneity in video presentation modes and operating conditions of Internet-enabled devices that access such content. Users are often interested in the content, rather than the specific sources or the formats. The host-centric format of the current Internet does not support these requirements naturally. Neither do the recent data-centric naming proposals, since they rely on naming content based on raw byte-level hashing schemes. We argue that to meet these requirements, enabling content retrieval mechanisms to name and query directly for the underlying information is a good way forward. In addition to decoupling content from available sources and transfer protocols, these "information-aware names" or InfoNames explicitly decouple the information from content presentation factors as well. We envision an InfoName Resolution System (IRS) to resolve location based on InfoNames, while taking into account the operating conditions of devices. In this demo, we present an application to show how InfoNames can serve as presentation-invariant and portable names to fetch video content independent of device capabilities and resource constraints.
Ashok Anand, Aditya Akella, Athula Balachandran, Vyas Sekar, Srinivasan Seshan
SIGCOMM3
2010 Cooperative interdomain traffic engineering using Nash bargaining and decomposition
Gireesh Shrimali, Aditya Akella, Almir Mutapcic
IEEE/ACM Trans. Netw.2
2009 Mining policies from enterprise network configuration
abstract
Few studies so far have examined the nature of reachability policies in enterprise networks. A better understanding of reachability policies could both inform future approaches to network design as well as current network configuration mechanisms. In this paper, we introduce the notion of a policy unit, which is an abstract representation of how the policies implemented in a network apply to different network hosts. We develop an approach for reverse-engineering a network's policy units from its router configuration. We apply this approach to the configurations of five productions networks, including three university and two private enterprises. Through our empirical study, we validate that policy units capture useful characteristics of a network's policy. We also obtain insights into the nature of the policies implemented in modern enterprises. For example, we find most hosts in these networks are subject to nearly identical reachability policies at Layer 3.
Theophilus Benson, Aditya Akella, David A. Maltz
Internet Measurement Conference2
2009 Unraveling the Complexity of Network Management
Theophilus Benson, Aditya Akella, David A. Maltz
NSDI2
2009 SmartRE: an architecture for coordinated network-wide redundancy elimination
abstract
Application-independent Redundancy Elimination (RE), or identifying and removing repeated content from network transfers, has been used with great success for improving network performance on enterprise access links. Recently, there is growing interest for supporting RE as a network-wide service. Such a network-wide RE service benefits ISPs by reducing link loads and increasing the effective network capacity to better accommodate the increasing number of bandwidth-intensive applications. Further, a networkwide RE service democratizes the benefits of RE to all end-to-end traffic and improves application performance by increasing throughput and reducing latencies.
Ashok Anand, Vyas Sekar, Aditya Akella
SIGCOMM3
2008 Avoiding File System Micromanagement with Range Writes
Ashok Anand, Sayandeep Sen, Andrew Krioukov, Florentina I. Popovici, Aditya Akella, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, Suman Banerjee 0001
OSDI5
2008 Packet caches on routers: the implications of universal redundant traffic elimination
abstract
Many past systems have explored how to eliminate redundant transfers from network links and improve network efficiency. Several of these systems operate at the application layer, while the more recent systems operate on individual packets. A common aspect of these systems is that they apply to localized settings, e.g. at stub network access links. In this paper, we explore the benefits of deploying packet-level redundant content elimination as a universal primitive on all Internet routers. Such a universal deployment would immediately reduce link loads everywhere. However, we argue that far more significant network-wide benefits can be derived by redesigning network routing protocols to leverage the universal deployment. We develop "redundancy-aware" intra- and inter-domain routing algorithms and show that they enable better traffic engineering, reduce link usage costs, and enhance ISPs' responsiveness to traffic variations. In particular, employing redundancy elimination approaches across redundancy-aware routes can lower intra and inter-domain link loads by 10-50%. We also address key challenges that may hinder implementation of redundancy elimination on fast routers. Our current software router implementation can run at OC48 speeds.
Ashok Anand, Archit Gupta, Aditya Akella, Srinivasan Seshan, Scott Shenker
SIGCOMM3
2008 Remote Profiling of Resource Constraints of Web Servers Using Mini-Flash Crowds
Pratap Ramamurthy, Vyas Sekar, Aditya Akella, Balachander Krishnamurthy, Anees Shaikh
USENIX ATC3
2008 On the performance benefits of multihoming route control
Aditya Akella, Bruce M. Maggs, Srinivasan Seshan, Anees Shaikh
IEEE/ACM Trans. Netw.1
2008 Corrections to "on the performance benefits of multihoming route control"
Aditya Akella, Bruce M. Maggs, Srinivasan Seshan, Anees Shaikh, Ramesh K. Sitaraman
IEEE/ACM Trans. Netw.1
2007 Traffic-Aware Channel Assignment in Enterprise Wireless LANs
abstract
Campus and enterprise wireless networks are increasingly characterized by ubiquitous coverage and rising traffic demands. Efficiently assigning channels to access points (APs) in these networks can significantly affect the performance and capacity of the WLANs. The state-of-the-art approaches assign channels statically, without considering prevailing traffic demands. In this paper, we show that the quality of a channel assignment can be improved significantly by incorporating observed traffic demands at APs and clients into the assignment process. We refer to this astraffic-aware channel assignment. We conduct extensive trace-driven and synthetic simulations and identify deployment scenarios where traffic-awareness is likely to be of great help, and scenarios where the benefit is minimal. We address key practical issues in using traffic-awareness, including measuring an interference graph, handling non-binary interference, collecting traffic demands, and predicting future demands based on historical information. We present an implementation of our assignment scheme for a 25-node WLAN testbed. Our testbed experiments show that traffic-aware assignment offers superior network performance under a wide range of real network configurations. On the whole, our approach is simple yet effective. It can be incorporated into existing WLANs with little modification to existing wireless nodes and infrastructure.
Eric Rozner, Yogita Mehta, Aditya Akella, Lili Qiu
ICNP3
2007 Cooperative Inter-Domain Traffic Engineering Using Nash Bargaining and Decomposition
abstract
We present a new inter-domain traffic engineering protocol based on the concepts of Nash bargaining and dual decomposition. Under this scheme, ISPs use an iterative procedure to jointly optimize a social cost function, referred to as the Nash product. We show that the global optimization problem can be separated into sub-problems by introducing appropriate shadow prices on the inter-domain flows. These sub-problems can then be solved independently and in a decentralized manner by the individual ISPs. Our approach does not require the ISPs to share any sensitive internal information (such as network topology or link weights). More importantly, our approach is provably Pareto-efficient and fair. Therefore, we believe that our approach is highly amenable to adoption by ISPs when compared to past naive approaches. We conduct simulation studies of our approach over several real ISP topologies. Our evaluation shows that the approach converges quickly, offers equitable performance improvements to ISPs, is significantly better than unilateral approaches (e.g. hot potato routing) and offers the same performance as a centralized solution with full knowledge.
Gireesh Shrimali, Aditya Akella, Almir Mutapcic
INFOCOM2
2007 Self-management in chaotic wireless deployments
Aditya Akella, Glenn Judd, Srinivasan Seshan, Peter Steenkiste
Wirel. Networks1
2006 Achieving Good End-to-End Service Using Bill-Pay
Cristian Estan, Aditya Akella, Suman Banerjee 0001
HotNets2
2006 Flow-Cookies: Using Bandwidth Amplification to Defend Against DDoS Flooding Attacks
abstract
This paper describes flow-cookies which defend against DDoS flooding attacks using bandwidth amplification. "Flow-cookies" is a mechanism in which a Website can reliably send filtering requests to a cooperating node in the network, leveraging its protection bandwidth. In this approach, a third party provider installs a flow-cookies enabled middlebox called the cookie box, in the network at a high bandwidth link. All traffic to or from the protected Web server must traverse the cookie box. The cookie box guarantees that all packets that pass between it and the server belong to a legitimate TCP flow with a valid sender. This implementation is able to operate at gigabit speeds including per-packet IP filtering of millions of addresses. This approach is also very effective against high volume SYN flooding attacks
Martín Casado, Aditya Akella, Niels Provos
IWQoS3
2006 SANE: A Protection Architecture for Enterprise Networks
Martín Casado, Tal Garfinkel, Aditya Akella, Michael J. Freedman, Dan Boneh, Nick McKeown
USENIX Security Symposium3
2005 Self-management in chaotic wireless deployments
abstract
Over the past few years, wireless networking technologies have made vast forays into our daily lives. Today, one can find 802.11 hardware and other personal wireless technology employed at homes, shopping malls, coffee shops and airports. Present-day wireless network deployments bear two important properties: they are unplanned, with most access points (APs) deployed by users in a spontaneous manner, resulting in highly variable AP densities; and they are unmanaged, since manually configuring and managing a wireless network is very complicated. We refer to such wireless deployments as being chaotic.In this paper, we present a study of the impact of interference in chaotic 802.11 deployments on end-client performance. First, using large-scale measurement data from several cities, we show that it is not uncommon to have tens of APs deployed in close proximity of each other. Moreover, most APs are not configured to minimize interference with their neighbors. We then perform trace-driven simulations to show that the performance of end-clients could suffer significantly in chaotic deployments. We argue that end-client experience could be significantly improved by making chaotic wireless networks self-managing. We design and evaluate automated power control and rate adaptation algorithms to minimize interference among neighboring APs, while ensuring robust end-client performance.
Aditya Akella, Glenn Judd, Srinivasan Seshan, Peter Steenkiste
MobiCom1
2004 On the responsiveness of DNS-based network control
abstract
For the last few years, large Web content providers interested in improving their scalability and availability have increasingly turned to three techniques: mirroring, content distribution, and ISP multihoming. The Domain Name System (DNS) has gained a prominent role in the way each of these techniques directs client requests to achieve the goals of scalability and availability. The DNS is thought to offer the transparent and agile control necessary to react quickly to ISP link failures or phenomenon such as flash crowds.
Jeffrey Pang, Aditya Akella, Anees Shaikh, Balachander Krishnamurthy, Srinivasan Seshan
Internet Measurement Conference2
2004 Availability, usage, and deployment characteristics of the domain name system
abstract
The Domain Name System (DNS) is a critical part of the Internet's infrastructure, and is one of the few examples of a robust, highlyscalable, and operational distributed system. Although a few studies have been devoted to characterizing its properties, such as its workload and the stability of the top-level servers, many key components of DNS have not yet been examined. Based on large-scale measurements taken from servers in a large content distribution network, we present a detailed study of key characteristics of the DNS infrastructure, such as load distribution, availability, and deployment patterns of DNS servers. Our analysis includes both local DNS servers and servers in the authoritative hierarchy. We find that (1) the vast majority of users use a small fraction of deployed name servers, (2) the availability of most name servers is high, and (3) there exists a larger degree of diversity in local DNS server deployment and usage than for authoritative servers. Furthermore, we use our DNS measurements to draw conclusions about federated infrastructures in general. We evaluate and discuss the impact of federated deployment models on future systems, such as Distributed Hash Tables.
Jeffrey Pang, James Hendricks, Aditya Akella, Roberto De Prisco, Bruce M. Maggs, Srinivasan Seshan
Internet Measurement Conference3
2004 A comparison of overlay routing and multihoming route control
abstract
The limitations of BGP routing in the Internet are often blamed for poor end-to-end performance and prolonged connectivity interruptions. Recent work advocates using overlays to effectively bypass BGP's path selection in order to improve performance and fault tolerance. In this paper, we explore the possibility that intelligent control of BGP routes, coupled with ISP multihoming, can provide competitive end-to-end performance and reliability. Using extensive measurements of paths between nodes in a large content distribution network, we compare the relative benefits of overlay routing and multihoming route control in terms of round-trip latency, TCP connection throughput, and path availability. We observe that the performance achieved by route control together with multihoming to three ISPs (3-multihoming), is within 5-15% of overlay routing employed in conjunction 3-multihoming, in terms of both end-to-end RTT and throughput. We also show that while multihoming cannot offer the nearly perfect resilience of overlays, it can eliminate almost all failures experienced by a singly-homed end-network. Our results demonstrate that, by leveraging the capability of multihoming route control, it is not necessary to circumvent BGP routing to extract good wide-area performance and availability from the existing routing system.
Aditya Akella, Jeffrey Pang, Bruce M. Maggs, Srinivasan Seshan, Anees Shaikh
SIGCOMM1
2004 Multihoming Performance Benefits: An Experimental Evaluation of Practical Enterprise Strategies
Aditya Akella, Srinivasan Seshan, Anees Shaikh
USENIX ATC, General Track1
2003 The Impact of False Sharing on Shared Congestion Management
abstract
Several recent proposals for sharing congestion information across concurrent flows between end-systems overlook an important problem: two or more flows sharing congestion state may in fact not share the same bottleneck. In this paper, we categorize the origins of this false sharing into two distinct cases: (i) networks with QoS enhancements such as differentiated services, where a flow classifier segregates flows into different queues, and (ii) networks with path diversity where different flows to the same destination address are routed differently. We evaluate the impact of false sharing on flow performance and investigate how false sharing can be detected by a sender. We discuss how a sender must respond upon detecting false sharing. Our results show that persistent overload can be avoided with window-based congestion control even for extreme false sharing, but higher bandwidth flows run at a slower rate. We find that delay and reordering statistics can be used to develop robust detectors of false sharing and are superior to those based on loss patterns. We also find that it is markedly easier to detect and react to false sharing than it is to start by isolating flows and merge their congestion state afterward.
Aditya Akella, Srinivasan Seshan, Hari Balakrishnan
ICNP1
2003 An empirical evaluation of wide-area internet bottlenecks
abstract
Conventional wisdom has been that the performance limitations in the current Internet lie at the edges of the network -- i.e last mile connectivity to users, or access links of stub ASes. As these links are upgraded, however, it is important to consider where new bottlenecks and hot-spots are likely to arise. In this paper, we address this question through an investigation of non-access bottlenecks. These are links within carrier ISPs or between neighboring carriers that could potentially constrain the bandwidth available to long-lived TCP flows. Through an extensive measurement study, we discover, classify, and characterize bottleneck links (primarily in the U.S.) in terms of their location, latency, and available capacity.We find that about 50% of the Internet paths explored have a non-access bottleneck with available capacity less than 50 Mbps, many of which limit the performance of well-connected nodes on the Internet today. Surprisingly, the bottlenecks identified are roughly equally split between intra-ISP links and peering links between ISPs. Also, we find that low-latency links, both intra-ISP and peering, have a significant likelihood of constraining available bandwidth. Finally, we discuss the implications of our findings on related issues such as choosing an access provider and optimizing routes through the network. We believe that these results could be valuable in guiding the design of future network services, such as overlay routing, in terms of which links or paths to avoid (and how to avoid them) in order to improve performance.
Aditya Akella, Srinivasan Seshan, Anees Shaikh
Internet Measurement Conference1
2003 Scaling properties of the Internet graph
abstract
As the Internet grows in size, it becomes crucial to understand how the speeds of links in the network must improve in order to sustain the pressure of new end-nodes being added each day. Although the speeds of links in the core and at the edges roughly improve according to Moore's law, this improvement alone might not be enough. Indeed, the structure of the Internet graph and routing in the network might necessitate much faster improvements in the speeds of key links in the network.In this paper, using a combination of analysis and extensive simulations, we show that the worst congestion in the Internet in fact scales poorly with the network size (n1+Ω(1), where n is the number of nodes), when shortest-path routing is used. We also show, somewhat surprisingly, that policy-based routing does not exacerbate the maximum congestion when compared to shortest-path routing.Our results show that it is crucial to identify ways to alleviate this congestion to avoid some links from being perpetually congested. To this end, we show that the congestion scaling properties of the Internet graph can be improved dramatically by introducing moderate amounts of redundancy in the graph in terms of parallel edges between pairs of adjacent nodes.
Aditya Akella, Shuchi Chawla 0001, Arvind Kannan, Srinivasan Seshan
PODC1
2003 A measurement-based analysis of multihoming
abstract
Multihoming has traditionally been employed by stub networks to enhance the reliability of their network connectivity. With the advent of commercial "intelligent route control" products, stubs now leverage multihoming to improve performance. Although multihoming is widely used for reliability and, increasingly for performance, not much is known about the tangible benefits that multihoming can offer, or how these benefits can be fully exploited. In this paper, we aim to quantify the extent to which multihomed networks can leverage performance and reliability benefits from connections to multiple providers. We use data collected from servers belonging to the Akamai content distribution network to evaluate performance benefits from two distinct perspectives of multihoming: high-volume content-providers which transmit large volumes of data to many distributed clients, and enterprises which primarily receive data from the network. In both cases, we find that multihoming can improve performance significantly and that not choosing the right set of providers could result in a performance penalty as high as 40%. We also find evidence of diminishing returns in performance when more than four providers are considered for multihoming. In addition, using a large collection of measurements, we provide an analysis of the reliability benefits of multihoming. Finally, we provide guidelines on how multihomed networks can choose ISPs, and discuss practical strategies of using multiple upstream connections to achieve optimal performance benefits.
Aditya Akella, Bruce M. Maggs, Srinivasan Seshan, Anees Shaikh, Ramesh K. Sitaraman
SIGCOMM1
2003 An empirical evaluation of wide-area internet bottlenecks
abstract
Performance limitations in the current Internet are thought to lie at the edges of the network -- i.e last mile connectivity to users, or access links of stub ASes. As these links are upgraded, however, it is important to consider where new bottlenecks and hot-spots are likely to arise. Through an extensive measurement study, we discover, classify and characterize non-access bottleneck links in terms of their location, latency and available capacity. We find that nearly half of the paths explored have a non-access bottleneck with available capacity less than 50 Mbps. The bottlenecks identified are roughly equally split between intra-ISP links and links between ISPs. These results have implications on issues such as the choice of access providers and route optimization.
Aditya Akella, Srinivasan Seshan, Anees Shaikh
SIGMETRICS1
2002 Selfish behavior and stability of the internet: a game-theoretic analysis of TCP
abstract
For years, the conventional wisdom [7, 22] has been that the continued stability of the Internet depends on the widespread deployment of "socially responsible" congestion control. In this paper, we seek to answer the following fundamental question: If network end-points behaved in a selfish manner, would the stability of the Internet be endangered?.We evaluate the impact of greedy end-point behavior through a game-theoretic analysis of TCP. In this "TCP Game" each flowattempts to maximize the throughput it achieves by modifying its congestion control behavior. We use a combination of analysis and simulation to determine the Nash Equilibrium of this game. Our question then reduces to whether the network operates efficiently at these Nash equilibria.Our findings are twofold. First, in more traditional environments -- where end-points use TCP Reno-style loss recovery and routers use drop-tail queues -- the Nash Equilibria are reasonably efficient. However, when endpoints use more recent variations of TCP (e.g., SACK) and routers employ either RED or drop-tail queues, the Nash equilibria are very inefficient. This suggests that the Internet of the past could remain stable in the face of greedy end-user behavior, but the Internet of today is vulnerable to such behavior. Second, we find that restoring the efficiency of the Nash equilibria in these settings does not require heavy-weight packet scheduling techniques (e.g., Fair Queuing) but instead can be done with a very simple stateless mechanism based on CHOKe [21].
Aditya Akella, Srinivasan Seshan, Richard M. Karp, Scott Shenker, Christos H. Papadimitriou
SIGCOMM1