VLDB 2026 Research / reviewers in the wild / expert
Sameer G. Kulkarni
dblp:185/5705
· DBLP profile ↗
23ranked-venue papers
6as first author
13since 2021 · last 2024
0000-0003-4727-6875ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 4 first-author · 6 since 2021Systems, architecture and hardware · 4 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PePC: Popularity Based Early Predictive Caching in Named Data NetworksabstractCaching technique used in Information Centric/Named Data Networks (ICN/NDN) governs the response time. Cache capacity constraints at routers have led to investigations on different caching mechanisms to improve effective caching and performance in terms of improved cache hits and response time for requested contents. However, most caching methods remain oblivious to the dynamics of cache occupancy. In this paper, we describe a new caching technique which predicts whether a new content has to be cached or not considering the current occupancy level of the cache. Our prediction based approach is inspired by the Random Early Detection (RED) method used for queue management. Similar to RED, our predictive caching algorithm bases its decision to cache a content using the average cache occupancy and also takes into account the content popularity. When the cache occupancy is low, we cache every possible content, and with the increasing cache occupancy, the decision to cache the content is decided based on the content popularity and the occupancy threshold parameters. We perform simulation based studies using discrete event simulator to assess its performance. We also compare the performance of our predictive caching method with five different popular caching methods used in Named Data Networks to show its superiority over others. Neminath Hubballi, Pankaj Chaudhary, Sameer G. Kulkarni |
CCNC | 3 |
| 2024 | Privacy Performance Trade-off in Web ServicesabstractSecurity and Privacy have become fundamental requirements of modern Internet services. Over the years, both Hypertext Transfer Protocol (HTTP) and Transport Layer Security (TLS) have evolved significantly to meet the performance, privacy and security demands of the web services. However, the usage of Service Name Identity (SNI) in TLS carry service-related information in plain-text, which potentially reveal the user’s activity and compromise the privacy. In this work, we analyse the performance, security and privacy trade-offs offered by the recent developments in HTTP and TLS protocols namely HTTP/3 and TLS1.3. Our results indicate the end-to-end performance of HTTP/3 and HTTP/2 to be very similar, but HTTP/3 offers better security and privacy. Further, we quantify the overheads associated with HTTP/3 and find that the computational complexity with HTTP/3 for SNI obfuscation and extraction from ‘ClientHello’ packets is nearly 10 times more than HTTP/2. Further, we find that the user-space implementations of QUIC in HTTP/3 are more compute-intensive and prone to be unstable. We conclude that a leaner alternative would be the adoption of "Encrypted ClientHello" (ECH), that proposes to overcome this privacy issue by extending TLS 1.3, where all the information that could potentially reveal the service type is encrypted using a public key. The widespread adoption of TLS 1.3 with ECH is imperative to enable complete privacy in web services. S. Hari Hara Sudhan, Manjesh Kumar Hanawal, Sameer G. Kulkarni |
LCN | 3 |
| 2024 | Demo: Security Vulnerabilities and Network Service Disruptions with HTTP/3abstractIn this work, we meticulously examine and demonstrate the security vulnerabilities associated with HTTP/3 and the adversities it brings to the operations of the network services (middleboxes). HTTP/3 is built using the new QUIC transport protocol to introduce enhancements to web communication by leveraging the QUIC protocols secure and privacy focused features such as connection migration, passive latency monitoring, congestion control, flow control, and support for multiple streams.In the course of our investigation, we unveil unintended vulnerabilities inherent in the QUIC protocol. Specifically, we demonstrate that the passive latency monitoring feature in the QUIC protocol exposes a covert channel that can be exploited for reliable covert communication. Furthermore, we reveal that the QUIC connection migration feature disrupts the functionality of critical network functions, such as NAT/NAPT, leading to a denial-of-service vulnerability. We provide a practical demonstration of this denial-of-service vulnerability in a NAT network. Our findings highlight the need for comprehensive and robust security solutions to address the outlined vulnerabilities in HTTP/3. S. Hari Hara Sudhan, Sameer G. Kulkarni |
LCN | 2 |
| 2024 | D-STACK: High Throughput DNN Inference by Effective Multiplexing and Spatio-Temporal Scheduling of GPUsabstractHardware accelerators such as GPUs are required for real-time, low latency inference with Deep Neural Networks (DNN). Providing inference services in the cloud can be resource intensive, and effectively utilizing accelerators in the cloud is important. Spatial multiplexing of the GPU, while limiting the GPU resources (GPU%) to each DNN to the right amount, leads to higher GPU utilization and higher inference throughput. Right-sizing the GPU for each DNN the optimal batching of requests to balance throughput and service level objectives (SLOs), and maximizing throughput by appropriately scheduling DNNs are still significant challenges.This article introduces a dynamic and fair spatio-temporal scheduler (D-STACK) for multiple DNNs to run in the GPU concurrently. We develop and validate a model that estimates the parallelism each DNN can utilize and a lightweight optimization formulation to find an efficient batch size for each DNN. Our holistic inference framework provides high throughput while meeting application SLOs. We compare D-STACK with other GPU multiplexing and scheduling methods (e.g., NVIDIA Triton, Clipper, Nexus), using popular DNN models. Our controlled experiments with multiplexing several popular DNN models achieve up to$1.6\times$improvement in GPU utilization and up to$4\times$improvement in inference throughput. Aditya Dhakal, Sameer G. Kulkarni, K. K. Ramakrishnan |
IEEE Trans. Cloud Comput. | 2 |
| 2023 | COUNSEL: Cloud Resource Configuration Management using Deep Reinforcement LearningabstractInternet Clouds are essentially service factories that offer various networked services through different service models, viz., Infrastructure, Platform, Software, and Functions as a Service. Meeting the desired service level objectives (SLOs) while ensuring efficient resource utilization requires significant efforts to provision the associated cloud resources correctly and on time. Therefore, one of the critical issues for any cloud service provider is resource configuration management. On one end, i.e., from the cloud operator's perspective, resource management affects overall resource utilization and efficiency. In contrast, from the cloud user/customer perspective, resource configuration affects the performance, cost, and offered SLOs. However, the state-of-the-art solutions for finding the configurations are limited to a single component or handle static workloads. Further, these solutions are computationally expensive and introduce profiling overhead, limiting scalability. Therefore, we propose COUNSEL, a deep reinforcement learning-based framework to handle the dynamic workloads and efficiently manage the configurations of an arbitrary multi-component service. We evaluate COUNSEL with three initial policies: over-provisioning, under-provisioning, and expert provisioning. In all the cases, COUNSEL eliminates the profiling overhead and achieves the average reward between 20 - 60% without violating the SLOs and budget constraints. Moreover, the inference time of COUNSEL has a constant time complexity. Adithya Hegde, Sameer G. Kulkarni, Abhinandan S. Prasad |
CCGrid | 2 |
| 2023 | eNCache: Improving content delivery with cooperative caching in Named Data Networking
Pankaj Chaudhary, Neminath Hubballi, Sameer G. Kulkarni |
Comput. Networks | 3 |
| 2022 | Spectrum-Efficiency Analysis for Trench-Assisted and Heterogeneous-Index Multicore Fiber NetworksabstractAdvanced multicore fiber (MCF) structures such as trench-assisted (TA)-MCF and heterogeneous index (HI)-MCF have been designed to reduce inter-core crosstalk (XT) in MCFs. However, optical networking studies have been primarily focused on the most common MCF structure, namely, homogeneous (HOM)MCF, where the high XT-levels limit the transmission reach (TR) of lightpaths. In this work, we investigate XT, XT-margin, and modulation selection for the TA-MCF and HI-MCF networks. Calculations for estimation of XT for different MCF structures, core-types, and modulation schemes are presented. We perform routing, spectrum, modulation, and core assignment (RSMCA) under incremental traffic using proactive XT-management schemes, and analyze the spectrum-efficiency of TA-MCF and HI-MCF networks. Simulations indicate that TA-MCF and HI-MCF structures can significantly improve the spectrum-efficiency in backbone optical networks, thus achieving higher lightpath establishment. We also highlight some of the open aspects of TA-MCF and HI-MCF networks. Anuj Agrawal, Sameer G. Kulkarni |
ICC | 2 |
| 2022 | Slice-Tune: a system for high performance DNN autotuningabstractAutotuning DNN models prior to their deployment is an essential but time-consuming task. Using expensive (and power-hungry) GPU and TPU accelerators efficiently is also key. Since DNNs do not always use a GPU fully, spatial multiplexing of multiple models can provide just the right amount of GPU resources for each DNN. We find that a DNN model tuned with the maximum GPU resources has higher inference latency if less GPU resources are available at inference time. We present methods to tune a DNN model, so that we provide the right amount of accelerator resources during tuning. Thus, even when a wide range of GPU resources are available at inference time, the tuned model achieves low inference latency. Further, existing autotuning frameworks take a long time to tune a model due to inefficient utilization of the client and server-side CPU and GPU. Our system, Slice-Tune, improves several autotuning frameworks to efficiently use system resources by re-thinking the partitioning of tasks between the client and server (where models are profiled on the server GPU), in a Kubernetes environment. We increase parallelism during tuning by sharding the tuning model across multiple tuning application instances, providing concurrent tuning of different operators of a model. We also scale server instances to achieve better GPU multiplexing. Slice-Tune reduces DNN autotuning time in a single GPU and in GPU clusters. Slice-Tune decreases DNN autotuning time by up to 75%, and increase autotuning throughput by a factor of 5, across 3 different autotuning frameworks (TVM, Ansor, and Chameleon). Aditya Dhakal, K. K. Ramakrishnan, Sameer G. Kulkarni, Puneet Sharma 0001, Junguk Cho |
Middleware | 3 |
| 2021 | Primitives Enhancing GPU Runtime Support for Improved DNN PerformanceabstractDeep neural networks (DNNs) are increasingly used for real-time inference, requiring low latency, but require significant computational power as they continue to increase in complexity. Edge clouds promise to offer lower latency due to their proximity to end users and having powerful accelerators like GPUs to provide the computation power needed for DNNs. But it is also important to ensure that the edge-cloud resources are utilized well. For this, multiplexing several DNN models through spatial sharing of the GPU can substantially improve edge-cloud resource usage. Typical GPU runtime environments have significant interactions with the CPU, to transfer data to the GPU, for CPU-GPU synchronization on inference task completions, etc. These result in overheads. We present a DNN inference framework with a set of software primitives that reduce the overhead for DNN inference, increase GPU utilization and improve performance, with lower latency and higher throughput. Our first primitive uses the GPU DMA effectively, reducing the CPU cycles spent to transfer the data to the GPU. A second primitive uses asynchronous ‘events' for faster task completion notification. GPU runtimes typically preclude fine-grained user control on GPU resources, causing long GPU downtimes when adjusting resources. Our third primitive supports overlapping of model-loading and execution, thus allowing GPU resource re-allocation with very little GPU idle time. Our other primitives increase inference throughput by improving scheduling and processing more requests. Overall, our primitives decrease inference latency by more than 35% and increase DNN throughput by 2-3x. Aditya Dhakal, Sameer G. Kulkarni, K. K. Ramakrishnan |
CLOUD | 2 |
| 2021 | Mu: An Efficient, Fair and Responsive Serverless Framework for Resource-Constrained Edge CloudsabstractServerless computing platforms simplify development, deployment, and automated management of modular software functions. However, existing serverless platforms typically assume an over-provisioned cloud, making them a poor fit for Edge Computing environments where resources are scarce. In this paper we propose a redesigned serverless platform that comprehensively tackles the key challenges for serverless functions in a resource constrained Edge Cloud. Viyom Mittal, Shixiong Qi, Ratnadeep Bhattacharya, Xiaosu Lyu, Sameer G. Kulkarni, Dan Li 0001, Jinho Hwang, K. K. Ramakrishnan, Timothy Wood 0001 |
SoCC | 6 |
| 2021 | SmartWatch: accurate traffic analysis and flow-state tracking for intrusion prevention using SmartNICsabstractDespite advances in network security, attacks targeting mission critical systems and applications remain a significant problem for network and datacenter providers. Existing telemetry platforms detect volumetric attacks at terabit scales using approximation techniques and coarse grain analysis. However, the prevalence of low and slow attacks that require very little bandwidth, makes flow-state tracking critical to overall attack mitigation. Traffic queries deployed on network switches are often limited by hardware constraints, preventing them from carrying out flow tracking features required to detect stealthy attacks. Such attacks can go undetected in the midst of high traffic volumes. Sourav Panda, Yixiao Feng, Sameer G. Kulkarni, K. K. Ramakrishnan, Nick G. Duffield, Laxmi N. Bhuyan |
CoNEXT | 3 |
| 2021 | Analyzing Open-Source Serverless Platforms: Characteristics and Performance (S)abstractServerless computing is increasingly popular because of its lower cost and easier deployment.Several cloud service providers (CSPs) offer serverless computing on their public clouds, but it may bring the vendor lock-in risk.To avoid this limitation, many open-source serverless platforms come out to allow developers to freely deploy and manage functions on self-hosted clouds.However, building effective functions requires much expertise and thorough comprehension of platform frameworks and features that affect performance.It is a challenge for a service developer to differentiate and select the appropriate serverless platform for different demands and scenarios.Thus, we elaborate the frameworks and event processing models of four popular open-source serverless platforms and identify their salient idiosyncrasies.We analyze the root causes of performance differences between different service exporting and auto-scaling modes on those platforms.Further, we provide several insights for future work, such as auto-scaling and metric collection.Index Terms-cloud computing, Sameer G. Kulkarni, K. K. Ramakrishnan, Dan Li 0001 |
SEKE | 2 |
| 2021 | Assessing Container Network Interface Plugins: Functionality, Performance, and ScalabilityabstractKubernetes, an open-source container orchestration platform, has been widely adopted by cloud service providers (CSPs) for its advantages in simplifying container deployment, scalability, and scheduling. Networking is one of the central components of Kubernetes, providing connectivity between different Pods (a group of containers) both within the same host and across hosts. To bootstrap Kubernetes networking, the Container Network Interface (CNI) provides a unified interface for the interaction between container runtimes. There are several CNI implementations, available as open-source `CNI plugins'. While they differ in functionality and performance, it is a challenge for a cloud provider to differentiate and choose the appropriate plugin for their environment. In this article, we compare the various open-source CNI plugins available from the community, qualitatively, and through detailed quantitative measurements. With our experimental evaluation, we analyze the overheads and bottlenecks for each CNI plugin, especially because of the interaction with the datapath/iptables as well as the host network stack. Overlay tunnel offload support in the network interface card plays a significant role in achieving the good performance of CNIs that use overlay tunnels for inter-host Pod-to-Pod communication. We also study scalability with an increasing number of Pods, as well as with HTTP workloads, and briefly evaluate Pod startup latency. Our measurement results inform the outline of an ideal CNI environment for Kubernetes. Shixiong Qi, Sameer G. Kulkarni, K. K. Ramakrishnan |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2020 | GSLICE: controlled spatial sharing of GPUs for a scalable inference platformabstractThe increasing demand for cloud-based inference services requires the use of Graphics Processing Unit (GPU). It is highly desirable to utilize GPU efficiently by multiplexing different inference tasks on the GPU. Batched processing, CUDA streams and Multi-process-service (MPS) help. However, we find that these are not adequate for achieving scalability by efficiently utilizing GPUs, and do not guarantee predictable performance. Aditya Dhakal, Sameer G. Kulkarni, K. K. Ramakrishnan |
SoCC | 2 |
| 2020 | Machine Learning at the Edge: Efficient Utilization of Limited CPU/GPU Resources by MultiplexingabstractEdge clouds can provide very responsive services for end-user devices that require more significant compute capabilities than they have. But edge cloud resources such as CPUs and accelerators such as GPUs are limited and must be shared across multiple concurrently running clients. However, multiplexing GPUs across applications is challenging. Further, edge servers are likely to require considerable amounts of streaming data to be processed. Getting that data from the network stream to the GPU can be a bottleneck, limiting the amount of work GPUs do. Finally, the lack of prompt notification of job completion from GPU also results in ineffective GPU utilization. We propose a framework that addresses these challenges in the following manner. We utilize spatial sharing of GPUs to multiplex the GPU more efficiently. While spatial sharing of GPU can increase GPU utilization, the uncontrolled spatial sharing currently available with state-of-the-art systems such as CUDA-MPS can cause interference between applications, resulting in unpredictable latency. Our framework utilizes controlled spatial sharing of GPU, which limits the interference across applications. Our framework uses the GPU DMA engine to offload data transfer to GPU, therefore preventing CPU from being bottleneck while transferring data from the network to GPU. Our framework uses the CUDA event library to have timely, low overhead GPU notifications. Preliminary experiments show that we can achieve low DNN inference latency and improve DNN inference throughput by a factor of ~ 1.4. Aditya Dhakal, Sameer G. Kulkarni, K. K. Ramakrishnan |
ICNP | 2 |
| 2020 | A SmartNIC-Accelerated Monitoring Platform for In-band Network TelemetryabstractRecent developments in In-band Network Telemetry (INT) provide granular monitoring of performance and load on network elements by collecting information in the data plane. INT enables traffic sources to embed telemetry instructions in data packets, avoiding separate probing or infrequent management-based monitoring. INT sink nodes track and collect metrics by retrieving INT metadata instructions appended by different sources of INT information. However, tracking the INT state in packets arriving at the sink is both compute intensive (requiring complex operations on each packet), and challenging for the standard P4 match-action packet processing pipeline to maintain line-rate. We propose a network telemetry platform in which the INT sink is implemented using distinct (C-based) algorithms on a SmartNIC in the monitoring host, complementing the P4 packet processing pipeline. This design accelerates packet processing and handles complex INT-related operations more efficiently than P4 match-action processing alone. While the P4 pipeline parses INT headers, a general-purpose Micro-C algorithms performs complex INT tasks (e.g. aggregation, event-detection, notification, etc.). We demonstrate that partitioning of INT processing significantly reduces processing overhead vs. a P4-on1y implementation, providing accurate, timely and almost loss-free event notification. Yixiao Feng, Sourav Panda, Sameer G. Kulkarni, K. K. Ramakrishnan, Nick G. Duffield |
LANMAN | 3 |
| 2020 | Managing State for Failure Resiliency in Network Function VirtualizationabstractEnsuring high scalability (elastic scale-out and consolidation), as well as high availability (failure resiliency) are critical in encouraging adoption of software-based network functions (NFs). In recent years, two paradigms have evolved in terms of the way the NFs manage their state - namely the Stateful (state is coupled with the NF instance) and a Stateless (state is externalized to a datastore) manner. These two paradigms present unique challenges and opportunities for ensuring high scalability and high availability of NFs and NF chains. In this work, we assess the impact on ensuring the correctness of NF state including the implications of non-determinism in packet processing, and carefully analyze and present the benefits and disadvantages of the two state management paradigms. We leverage OpenNetVM and Redis in-memory datastore to implement both state management paradigms and empirically compare the two. Although the stateless paradigm is desirable for elastic scaling, our experimental results show that, even at line-rate packet processing (10 Gbps), stateful NFs can achieve chain-level failover across servers in a LAN incurring less than 10% performance. The state-of-the-art stateless counterparts incur severe throughput penalties. We observe 30-85% overhead on normal processing, depending on the mode of state updated to the externalized datastore. Sameer G. Kulkarni, K. K. Ramakrishnan, Timothy Wood 0001 |
LANMAN | 1 |
| 2020 | Understanding Container Network Interface Plugins: Design Considerations and PerformanceabstractKubernetes, an open-source container orchestration platform, has been widely adopted by cloud service providers (CSPs) for its advantages in simplifying container deployment, scalability and scheduling. Networking is one of the central components of Kubernetes, providing connectivity between different pods (group of containers) both within the same host and across hosts. To bootstrap Kubernetes networking, the Container Network Interface (CNI) provides a unified interface for the interaction between container runtimes. There are several CNI implementations, available as open-source ‘CNI plugins’. While they differ in functionality and performance, it is a challenge for a cloud provider to differentiate and choose the appropriate plugin for their environment. In this paper, we compare the various open source CNI plugins available from the community, qualitatively and through detailed quantitative measurements. With our experimental evaluation, we analyze the overheads and bottlenecks for each CNI plugin, as a result of the network model it implements, interaction with the host network protocol stack and the network policies implemented in iptables rules. The choice of the CNI plugin may also be based on whether intra-host or inter-host communication dominates. Shixiong Qi, Sameer G. Kulkarni, K. K. Ramakrishnan |
LANMAN | 2 |
| 2020 | REINFORCE: Achieving Efficient Failure Resiliency for Network Function Virtualization-Based ServicesabstractEnsuring high availability (HA) for software-based networks is a critical design feature that will help the adoption of software-based network functions (NFs) in production networks. It is important for NFs to avoid outages and maintain mission-critical operations. However, HA support for NFs on the critical data path can result in unacceptable performance degradation. We present REINFORCE, an integrated framework to support efficient resiliency for NF service chains. REINFORCE includes timely failure detection and consistent failover mechanisms. REINFORCE replicates state to standby NFs (local and remote) while enforcing correctness. It minimizes the number of state transfers by exploiting the concept of external synchrony, and leverages opportunistic batching and multi-buffering to optimize performance. Experimental results show that, even at line-rate packet processing (10 Gbps), REINFORCE achieves chain-level failover across servers in a LAN within 10ms, incurring less than 10% performance overhead, and adds average latency only ~400 μs, with a worst-case latency of less than 1ms. REINFORCE also recovers from software failures within the same node in less than 100 μs, incurring less than 1% performance overhead and adds less than 5 μs latency during normal operation. Sameer G. Kulkarni, Guyue Liu, K. K. Ramakrishnan, Mayutan Arumaithurai, Timothy Wood 0001, Xiaoming Fu 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2020 | NFVnice: Dynamic Backpressure and Scheduling for NFV Service ChainsabstractManaging Network Function (NF) service chains requires careful system resource management. We propose NFVnice, a user space NF scheduling and service chain management framework to provide fair, efficient and dynamic resource scheduling capabilities on Network Function Virtualization (NFV) platforms. The NFVnice framework monitors load on a service chain at high frequency (1000Hz) and employs backpressure to shed load early in the service chain, thereby preventing wasted work. Borrowing concepts such as rate proportional scheduling from hardware packet schedulers, CPU shares are computed by accounting for heterogeneous packet processing costs of NFs, I/O, and traffic arrival characteristics. By leveraging cgroups, a user space process scheduling abstraction exposed by the operating system, NFVnice is capable of controlling when network functions should be scheduled. NFVnice improves NF performance by complementing the capabilities of the OS scheduler but without requiring changes to the OS's scheduling mechanisms. Our controlled experiments show that NFVnice provides the appropriate rate-cost proportional fair share of CPU to NFs and significantly improves NF performance (throughput and latency) by reducing wasted work across an NF chain, compared to using the default OS scheduler. NFVnice achieves this even for heterogeneous NFs with vastly different computational costs and for heterogeneous workloads. Sameer G. Kulkarni, Wei Zhang 0052, Jinho Hwang, Shriram Rajagopalan, K. K. Ramakrishnan, Timothy Wood 0001, Mayutan Arumaithurai, Xiaoming Fu 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2019 | Living on the Edge: Serverless Computing and the Cost of Failure ResiliencyabstractServerless computing platforms have gained popularity because they allow easy deployment of services in a highly scalable and cost-effective manner. By enabling just-in-time startup of container-based services, these platforms can achieve good multiplexing and automatically respond to traffic growth, making them particularly desirable for edge cloud data centers where resources are scarce. Edge cloud data centers are also gaining attention because of their promise to provide responsive, low-latency shared computing and storage resources. Bringing serverless capabilities to edge cloud data centers must continue to achieve the goals of low latency and reliability. The reliability guarantees provided by serverless computing however are weak, with node failures causing requests to be dropped or executed multiple times. Thus serverless computing only provides a best effort infrastructure, leaving application developers responsible for implementing stronger reliability guarantees at a higher level. Current approaches for providing stronger semantics such as “exactly once” guarantees could be integrated into serverless platforms, but they come at high cost in terms of both latency and resource consumption. As edge cloud services move towards applications such as autonomous vehicle control that require strong guarantees for both reliability and performance, these approaches may no longer be sufficient. In this paper we evaluate the latency, throughput, and resource costs of providing different reliability guarantees, with a focus on these emerging edge cloud platforms and applications. Sameer G. Kulkarni, Guyue Liu, K. K. Ramakrishnan, Timothy Wood 0001 |
LANMAN | 1 |
| 2018 | REINFORCE: achieving efficient failure resiliency for network function virtualization based servicesabstractEnsuring high availability (HA) for software-based networks is a critical design feature that will help the adoption of software-based network functions (NFs) in production networks. It is important for NFs to avoid outages and maintain mission-critical operations. However, HA support for NFs on the critical data path can result in unacceptable performance degradation. We present REINFORCE, an integrated framework to support efficient resiliency for NFs and NF service chains. REINFORCE includes timely failure detection and consistent failover mechanisms. REINFORCE replicates state to standby NFs (local and remote) while enforcing correctness. It minimizes the number of state transfers by exploiting the concept of external synchrony, and leverages opportunistic batching and multi-buffering to optimize performance. Experimental results show that, even at line-rate packet processing (10 Gbps), REINFORCE achieves chain-level failover across servers in a LAN (or within the same node) within 10ms (100/μs), incurring less than 10% (1%) performance overhead, and adds average latency of only ~400/μs (5/μs), with a worst-case latency of less than 1ms (10/μs). Sameer G. Kulkarni, Guyue Liu, K. K. Ramakrishnan, Mayutan Arumaithurai, Timothy Wood 0001, Xiaoming Fu 0001 |
CoNEXT | 1 |
| 2017 | NFVnice: Dynamic Backpressure and Scheduling for NFV Service ChainsabstractManaging Network Function (NF) service chains requires careful system resource management. We propose NFVnice, a user space NF scheduling and service chain management framework to provide fair, efficient and dynamic resource scheduling capabilities on Network Function Virtualization (NFV) platforms. The NFVnice framework monitors load on a service chain at high frequency (1000Hz) and employs backpressure to shed load early in the service chain, thereby preventing wasted work. Borrowing concepts such as rate proportional scheduling from hardware packet schedulers, CPU shares are computed by accounting for heterogeneous packet processing costs of NFs, I/O, and traffic arrival characteristics. By leveraging cgroups, a user space process scheduling abstraction exposed by the operating system, NFVnice is capable of controlling when network functions should be scheduled. NFVnice improves NF performance by complementing the capabilities of the OS scheduler but without requiring changes to the OS's scheduling mechanisms. Our controlled experiments show that NFVnice provides the appropriate rate-cost proportional fair share of CPU to NFs and significantly improves NF performance (throughput and loss) by reducing wasted work across an NF chain, compared to using the default OS scheduler. NFVnice achieves this even for heterogeneous NFs with vastly different computational costs and for heterogeneous workloads. Sameer G. Kulkarni, Wei Zhang 0052, Jinho Hwang, Shriram Rajagopalan, K. K. Ramakrishnan, Timothy Wood 0001, Mayutan Arumaithurai, Xiaoming Fu 0001 |
SIGCOMM | 1 |