Sourav Panda

dblp:261/2022 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
4since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2022 Synergy: A SmartNIC Accelerated 5G Dataplane and Monitor for Mobility Prediction
abstract
The 5G user plane function (UPF) is a critical inter-connection point between the data network and cellular network infrastructure. It governs the packet processing performance of the 5G core network. UPFs also need to be flexible to support several key control plane operations. Existing UPFs typically run on general-purpose CPUs, but have limited performance because of the overheads of host-based forwarding. We design Synergy, a novel 5G UPF running on SmartNICs that provides high throughput and low latency. It also supports monitoring functionality to gather critical data on user sessions for the prediction and optimization of handovers during user mobility. The SmartNIC UPF efficiently buffers data packets during handover and paging events by using a two-level flow-state access mechanism. This enables maintaining flow-state for a very large number of flows, thus providing very low latency for control and data planes and high throughput packet forwarding. Mobility prediction can reduce the handover delay by pre-populating state in the UPF and other core NFs. Synergy performs handover predictions based on an existing recurrent neural network model. Synergy's mobility predictor helps us achieve 2.32× lower average handover latency. Buffering in the SmartNIC, rather than the host, during paging and handover events reduces packet loss rate by at least 2.04×. Compared to previous approaches to building programmable switch-based UPFs, Synergy speeds up control plane operations such as handovers because of the low P4-programming latency leveraging tight coupling between SmartNIC and host.
Sourav Panda, K. K. Ramakrishnan, Laxmi N. Bhuyan
ICNP1
2022 CloudCluster: Unearthing the Functional Structure of a Cloud Service
Weiwu Pang, Sourav Panda, Muhammad J. Amjad, Christophe Diot, Ramesh Govindan
NSDI2
2021 SmartWatch: accurate traffic analysis and flow-state tracking for intrusion prevention using SmartNICs
abstract
Despite advances in network security, attacks targeting mission critical systems and applications remain a significant problem for network and datacenter providers. Existing telemetry platforms detect volumetric attacks at terabit scales using approximation techniques and coarse grain analysis. However, the prevalence of low and slow attacks that require very little bandwidth, makes flow-state tracking critical to overall attack mitigation. Traffic queries deployed on network switches are often limited by hardware constraints, preventing them from carrying out flow tracking features required to detect stealthy attacks. Such attacks can go undetected in the midst of high traffic volumes.
Sourav Panda, Yixiao Feng, Sameer G. Kulkarni, K. K. Ramakrishnan, Nick G. Duffield, Laxmi N. Bhuyan
CoNEXT1
2021 pMACH: Power and Migration Aware Container scHeduling
abstract
Data center workload fluctuations need periodic, but careful scheduling to minimize power consumption while meeting the task completion time requirements. Existing data center scheduling systems tightly pack containers to save power. However, with the growth of multi-tiered applications, there is a significant need to account for the affinity between application components, to minimize communication overheads and latency. Centralized container scheduling systems using graph partitioning algorithms cause a significant number of task migrations, with associated downtime.We design pMACH, a novel distributed container scheduling scheme for optimizing both power and task completion time in data centers. It minimizes task migrations and packs frequently communicating containers together without overloading servers. pMACH operates at peak energy efficiency, thus reducing energy consumption while also providing greater headroom for unpredictable workload spikes. We also propose in-network monitoring using smartNICs (sNIC) to measure the communications and then perform scheduling in a hierarchical, parallelized framework to achieve high performance and scalability. pMACH is based on incremental partitioning and it leverages the previous scheduling decision to significantly reduce the number of containers moved between servers, avoiding application downtime.Both testbed measurements and large-scale trace-driven simulations show that pMACH saves at least 13.44% more power compared to previous scheduling systems. It speeds task completion, reducing the 95th percentile by a factor of 1.76-2.11 compared to existing container scheduling schemes. Compared to other static graph-based approaches, our incremental partitioning technique reduces migrations per epoch by 82%.
Sourav Panda, K. K. Ramakrishnan, Laxmi N. Bhuyan
ICNP1
2020 A SmartNIC-Accelerated Monitoring Platform for In-band Network Telemetry
abstract
Recent developments in In-band Network Telemetry (INT) provide granular monitoring of performance and load on network elements by collecting information in the data plane. INT enables traffic sources to embed telemetry instructions in data packets, avoiding separate probing or infrequent management-based monitoring. INT sink nodes track and collect metrics by retrieving INT metadata instructions appended by different sources of INT information. However, tracking the INT state in packets arriving at the sink is both compute intensive (requiring complex operations on each packet), and challenging for the standard P4 match-action packet processing pipeline to maintain line-rate. We propose a network telemetry platform in which the INT sink is implemented using distinct (C-based) algorithms on a SmartNIC in the monitoring host, complementing the P4 packet processing pipeline. This design accelerates packet processing and handles complex INT-related operations more efficiently than P4 match-action processing alone. While the P4 pipeline parses INT headers, a general-purpose Micro-C algorithms performs complex INT tasks (e.g. aggregation, event-detection, notification, etc.). We demonstrate that partitioning of INT processing significantly reduces processing overhead vs. a P4-on1y implementation, providing accurate, timely and almost loss-free event notification.
Yixiao Feng, Sourav Panda, Sameer G. Kulkarni, K. K. Ramakrishnan, Nick G. Duffield
LANMAN2
2020 Error Vulnerabilities and Fault Recovery in Deep-Learning Frameworks for Hardware Accelerators
abstract
Hardware accelerators such as GP-GPUs, Tensor Cores, and Deep-Learning Accelerators (DLA) are increasingly being used in real-time settings such as autonomous vehicles (AVs). In such deployments, any software errors and process failures in hardware systems can lead to critical faults in AVs. Therefore, assessing and mitigating hardware accelerator faults are critical requirements for safety-critical systems. Past work on this subject focused on simulated and injected software and hardware faults to understand and analyze the behavior of the software stack and the entire system. However, programming errors and process failures caused when using software frameworks must also be considered. In this paper, we present experiments which show that widely used deep-learning frameworks are vulnerable to programming mistakes and errors. We first focus on memory-related programming errors caused by applications using deep-learning frameworks that facilitate high-performance inferencing. We next find that a reset to recover from any fault imposes significant time penalties in reloading a pre-trained deep neural network model. To reduce these fault recovery times, we propose fault recovery mechanisms that checkpoint and resume the network based on the inference stage when an error is detected. Finally, we substantiate the practical feasibility of our approach and evaluate the improvement in recovery times11A demo video clip demonstrating our recovery algorithm has been uploaded to Youtube: https://www.youtube.com/watch?v=xwUYdJdA5oM.. We use a case-study with real-world applications on an Nvidia GeForce GTX 1070 GPU and an Nvidia Xavier embedded platform, which is commonly used by multiple automotive OEMs.
Iljoo Baek, Sourav Panda, Nandha Kishore Srinivasan, Soheil Samii, Ragunathan Rajkumar
RTCSA3