VLDB 2026 Research / reviewers in the wild / expert
Rinku Shah
dblp:146/0131
· DBLP profile ↗
10ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0001-9823-4515ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling Hardware-Software Co-Design for Large-Scale AI Systems via Fine-Grained GPU Memory TracingabstractMemory traces capturing GPU operations over modern scale-up fabrics (e.g., load/store accesses, ordering events, synchronization, and communication behavior) enable a broad set of research directions for next-generation AI infrastructure. Palak Mishra, Rajat Bhardwaj, Yash Verma, Ramanjeet Singh, Rinku Shah |
SIGCOMM | 5 |
| 2025 | Securing In-Network Traffic Control Systems with P4AuthabstractIn-network traffic control systems built on programmable data planes enhance network performance. However, these systems also increase the attack surface and are vulnerable to attacks not seen before. We focus on a problem that stems from the fact that a programmable switch data plane trusts and processes the messages from upper layers in the switch software (OS, SDK, drivers) and from neighbor nodes in the network. Since these messages can update the state maintained in the data plane, which can influence traffic control decisions, it is important to protect such messages from adversaries aiming to degrade performance, compromise privacy, bypass security, or, worst case, network outage.In this paper, we present P4Auth, a key-based protection mechanism that ensures the authenticity and integrity of such messages in in-network systems making fast traffic control decisions. Our key idea is to move key-based security primitives to the switch data plane so that it reduces the trusted computing base and exposure to switch software vulnerabilities while enabling faster checks in the data plane. To realize this idea, we design and develop an authentication protocol, secure key exchange mechanism, and associated data plane primitives. We prototype P4Auth for Intel Tofino and understand the overheads of P4Auth. We also demonstrate how P4Auth protects two in-network systems from man-in-the-middle (MitM) adversaries. Ranjitha K., Medha Rachel Panna, Stavan Nilesh Christian, Karuturi Havya Sree, Sri Hari Malla, Dheekshitha Bheemanath, Rinku Shah, Praveen Tammana |
DSN | 7 |
| 2025 | Detecting Manipulation to Table Rules in the Programmable Data Planes
Ranjitha K., Karuturi Havya Sree, Devansh Garg, Stavan Nilesh Christian, Dheekshitha Bheemanath, Rinku Shah, Praveen Tammana |
Networking | 6 |
| 2024 | DL3: Adaptive Load Balancing for Latency-critical Edge Cloud ApplicationsabstractOn-premise edge cloud provides opportunities to enable ML-based latency-critical services to resource-constrained end devices. The edge services are deployed as loosely coupled microservices using cloud orchestrators like Kubernetes, and a load balancer distributes requests from an upstream microservice instance (client) across many downstream microservice instances (servers). However, in a shared environment, transient and sporadic delay events are common due to contention for host and network resources (e.g., high load on servers, high network queuing delays). To meet low latency requirements of edge services, the load balancer should quickly adapt to such delay events and adjust routing decisions (e.g., pick the best downstream instance among all). In this paper, we propose DL3, a distributed load balancer (LB) that quickly adapts to server load and network queuing delays by adjusting routing decisions so that the requests are forwarded to the best possible servers. The key idea is to enable LB with visibility into both servers’ load and transient delays on network paths toward the servers. We prototype DL3on a Kubernetes-managed edge cloud cluster and evaluated its performance for a latency-sensitive ML-based object detection service. Our preliminary results show that DL3improves tail response time by 33% compared to the state-of-the-art load balance mechanism. Prashanth P. S, Ranjitha K., Arjun Temura, Rinku Shah, Praveen Tammana |
CNSM | 5 |
| 2023 | In-Network Probabilistic Monitoring Primitives under the Influence of Adversarial Network InputsabstractNetwork management tasks heavily rely on network telemetry data. Programmable data planes provide novel ways to collect this telemetry data efficiently using probabilistic data structures like bloom filters and their variants. Despite the benefits of the data structures (and associated data plane primitives), their exposure increases the attack surface. That is, they are at risk of adversarial network inputs. Harish S. A, K. Shiv Kumar, Anibrata Majee, Amogh Bedarakota, Praveen Tammana, Pravein G. Kannan, Rinku Shah |
APNet | 7 |
| 2022 | Packet Processing Algorithm Identification using Program EmbeddingsabstractTo keep up with the network speeds, many recent works propose to offload network functions to SmartNICs. The process involves identifying packet-processing algorithms in a network function program then offloading them to appropriate accelerators available on SmartNICs. This process is often done manually for each architecture and is error-prone and laborious. In this work, we propose an automated solution to identify algorithms in network function programs. We model our approach as a classification problem of Machine Learning (ML) and propose using sophisticated program embeddings for representing the network function programs. We also identify the limited availability of datasets and propose a way of extrapolating them by systematically generating equivalent programs using (existing) compiler transformations in popular compiler infrastructures. Our approach relies on modeling programs as embeddings, uses ML models trained on such extrapolated datasets, and shows superior results over the recent works. S. VenkataKeerthy, Yashas Andaluri, Sayan Dey, Rinku Shah, Praveen Tammana, Ramakrishna Upadrasta |
APNet | 4 |
| 2021 | Leveraging Programmable Dataplanes for a High Performance 5G User Plane FunctionabstractEmerging 5G applications require a dataplane that has a high forwarding throughput and low processing latency, in addition to low cost and power consumption. To meet these requirements, the state-of-the-art 5G User Plane Functions (UPFs) are built over high performance packet I/O mechanisms like the Data Plane Development Kit (DPDK), and further offload some functionality to programmable dataplane hardware. In this paper, we design and implement several standards-compliant UPF prototypes, beginning with a software-only DPDK-based UPF, progressing to designs which offload different functions to programmable hardware. We evaluate and compare the performance of these designs, to highlight the costs and benefits of these offloads. Our results show that offload techniques employed in prior work help improve performance in certain scenarios, but also have their limitations. Overcoming these limitations and fully realizing the power of programmable hardware requires offloading more complex functionality than is done today. Our work presents a preliminary implementation towards a comprehensive programmable dataplane-accelerated 5G UPF. Abhik Bose, Diptyaroop Maji, Prateek Agarwal, Nilesh Unhale, Rinku Shah, Mythili Vutukuru |
APNet | 5 |
| 2018 | pcube: Primitives for Network Data Plane ProgrammingabstractP4 is a domain specific language to configure packet processing pipelines in programmable dataplane switches, and is a powerful idea towards realizing the goal of flexible software-defined networks. This paper presents pcube, a framework that provides a set of primitives to simplify the development of P4-based dataplane applications. pcube provides primitives for loops, summations, and other common operations on indexed state variables, which can be embedded within P4 code and unrolled by the pcube preprocessor. pcube also provides primitives to synchronize state variables across switches in distributed dataplane applications, which are automatically translated into P4 code to send and receive synchronization messages across multiple switches by pcube. We build example dataplane applications such as a distributed load balancer in our framework, and show that using pcube reduces the programming effort (in term of lines of code) significantly-by a factor of up to 5.4x. Rinku Shah, Aniket Shirke, Akash Trehan, Mythili Vutukuru, Purushottam Kulkarni |
ICNP | 1 |
| 2018 | Cuttlefish: Hierarchical SDN Controllers with Adaptive OffloadabstractOffloading computation to local controllers (closer to switches) has been a popular approach to designing scalable SDN controllers. We observe that, in addition to the offload of local switch-specific state, a subset of global state can also be offloaded to, and accessed at local controllers with suitable synchronization. We present the design and implementation of Cuttlefish, an SDN controller framework that adaptively offloads a portion of the application state (and computation) to local controllers. Cuttlefish uses developer-specified input to identify control messages that can be correctly processed at local controllers, and makes offloading decisions based on the cost of synchronizing the offloaded state across controllers. SDN applications use the Cuttlefish API to access the offloaded state, and Cuttlefish transparently manages the state synchronization, and redirection of control messages to the appropriate (central or local) controller. We have implemented Cuttlefish using the Floodlight SDN controller. Our evaluation shows that Cuttlefish applications achieve ~2X higher control plane throughput and ~50% lower control plane latency as compared to the traditional SDN design. Rinku Shah, Mythili Vutukuru, Purushottam Kulkarni |
ICNP | 1 |
| 2017 | Devolve-Redeem: Hierarchical SDN Controllers with Adaptive OffloadingabstractTowards improving SDN control plane scalability, past work has proposed SDN controller frameworks that offload computation which depends on local state to controllers residing on the switches. Our work identifies another type of computation that can be offloaded to local controllers: that which depends on state that is generated globally but can be used within local controllers with loose synchronization. Because using such state locally incurs a synchronization cost, such offload makes sense only when the benefits of the offload out-weigh the synchronization cost. We present the design and implementation of Devolve-Redeem, an SDN controller framework that can offload computation to local controllers depending on the mix of various control messages in the incoming traffic. The offload decision in our framework is made by computing a cost metric that captures the relative costs of processing every control message at the central and local controllers, taking into account synchronization costs. The SDN application developer using our framework writes a single application that runs at both the central and local controllers, using our state management API to access offloadable state. Our framework migrates between various offload modes using the computed cost metric, by manipulating the rules in the SDN switches that forward control messages to the controllers. Our framework also transparently handles state synchronization between central and local controllers in a manner that is consistent with the offload mode. We have implemented the SDN-based LTE EPC application in our framework, and experiments with our prototype demonstrate the effectiveness of our adaptive offload framework. Rinku Shah, Mythili Vutukuru, Purushottam Kulkarni |
APNet | 1 |