Sri Pramodh Rachuri

dblp:266/1005 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0001-9330-0552ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Constellate: Establishing the opportunity for Distributed Unit pooling in real-world 5G Radio Access Networks
abstract
As the adoption of Virtualized Radio Access Networks (vRAN) is gaining momentum in 5 G networks, Mobility Network Operators are considering a Centralized RAN (CRAN) architecture that moves the baseband functions to a far-edge cloud in order to gain dimensioning flexibility, resiliency and improved RAN performance. However, there have been limited studies on the benefits of centralization in improving RAN compute utilization, especially in the context of pooling the compute-intensive Distributed Unit (DU) resources. In this paper, we present the first study on the benefits of pooling in improving DU server utilization. Using longitudinal traces from a real-world 5G network, we show that significant Capex and Opex gains of $\mathbf{8 4 \%}$ and $\mathbf{9 4 \%}$, respectively, can be obtained through fine-grained pooling at a granularity of 1 second. We also present an affinitybased and dynamic pooling algorithm that can reduce the pooling overheads while still achieving significant pooling gains.
Sri Pramodh Rachuri, Anshul Gandhi, Gueyoung Jung, Shankaranarayanan Puzhavakath Narayanan, Alex Zelezniak
MASCOTS1
2024 EcoEdgeInfer: Dynamically Optimizing Latency and Sustainability for Inference on Edge Devices
abstract
The use of Deep Neural Networks (DNNs) has skyrocketed in recent years. While its applications have brought many benefits and use cases, they also have a significant environmental impact due to the high energy consumption of DNN execution. It has already been acknowledged in the literature that training DNNs is computationally expensive and requires large amounts of energy. However, the energy consumption of DNN inference is still an area that has not received much attention, yet. With the increasing adoption of online tools, the usage of inference has significantly grown and will likely continue to grow. Unlike training, inference is user-facing, requires low latency, and is used more frequently. As such, edge devices are being considered for DNN inference due to their low latency and privacy benefits. In this context, inference on edge is a timely area that requires closer attention to regulate its energy consumption. We present EcoEdgeInfer, a system that balances performance and sustainability for DNN inference on edge devices. Our core component of EcoEdgeInfer is an adaptive optimization algorithm, EcoGD, that strategically and quickly sweeps through the hardware and software configuration space to find the jointly optimal configuration that can minimize energy consumption and latency. EcoGD is agile by design, and adapts the configuration parameters in response to time-varying and unpredictable inference workload. We evaluate EcoEdgeInfer on different DNN models using real-world traces and show that EcoGD consistently outperforms existing baselines, lowering energy consumption by 31% and reducing tail latency by 14%, on average.
Sri Pramodh Rachuri, Nazeer Shaik, Mehul Choksi, Anshul Gandhi
SEC1
2024 OVIDA: Orchestrator for Video Analytics on Disaggregated Architecture
abstract
Millions of video cameras are deployed globally across major cities for learning-based video analytic (VA) applications, such as object detection. Video streams from the cameras are either sent over the wide-area network to be processed by the cloud or are (at least partially) processed in a local edge workstation, incurring significant latency and elevated financial costs. In this paper, to minimize reliance on the cloud and overcome the unavailability of high-compute workstations on edge, we investigate the use of heterogeneous and distributed embedded devices as edge nodes shared by multiple cameras to fully serve the video processing needs of a VA application (without requiring cloud support). We present OVIDA, an edge-only orchestrator to deploy VA application(s) on a distributed edge environment to maximize accuracy. Given the resource-constrained nature of edge nodes, OVIDA disaggregates the VA application pipeline into multiple modules. OVIDA's core functionality and contributions are: (i) optimizing the placement and replication of the VA application modules across the edge nodes to maximize the throughput, and in turn, accuracy; and (ii) an adaptive model selection algorithm for VA modules based on accuracy-throughput tradeoff to maximize accuracy in response to varying load conditions. To further improve performance, OVIDA employs a central-queue-based design (instead of the usual push-based design), which also obviates the need for complex load balancing algorithms. We implement OVIDA on top of Kubernetes and evaluate its performance for three VA applications, supported over a heterogeneous edge cluster under varying network conditions. When compared against several baselines in our evaluation, we achieve throughput and accuracy gains of at least 51% and 28%.
Manavjeet Singh, Sri Pramodh Rachuri, Bryan Bo Cao, Venkata Bhumireddy, Francesco Bronzino, Samir Ranjan Das, Anshul Gandhi, Shubham Jain 0003
SEC2
2022 Optimizing Near-Data Processing for Spark
abstract
Resource disaggregation (RD) is an emerging paradigm for data center computing whereby resource-optimized servers are employed to minimize resource fragmentation and improve resource utilization. Apache Spark deployed under the RD paradigm employs a cluster of compute-optimized servers to run executors and a cluster of storage-optimized servers to host the data on HDFS. However, the network transfer from storage to compute cluster becomes a severe bottleneck for big data processing. Near-data processing (NDP) is a concept that aims to alleviate network load in such cases by offloading (or "pushing down") some of the compute tasks to the storage cluster. Employing NDP for Spark under the RD paradigm is challenging because storage-optimized servers have limited computational resources and cannot host the entire Spark processing stack. Further, even if such a lightweight stack could be developed and deployed on the storage cluster, it is not entirely obvious which Spark queries would benefit from pushdown, and which tasks of a given query should be pushed down to storage.This paper presents the design and implementation of a near-data processing system for Spark, SparkNDP, that aims to address the aforementioned challenges. SparkNDP works by implementing novel NDP Spark capabilities on the storage cluster using a lightweight library of SQL operators and then developing an analytical model to help determine which Spark tasks should be pushed down to storage based on the current network and system state. Simulation and prototype implementation results show that SparkNDP can help reduce Spark query execution times when compared to both the default approach of not pushing down any tasks to storage and the outright NDP approach of pushing all tasks to storage.
Sri Pramodh Rachuri, Arun Gantasala, Prajeeth Emanuel, Anshul Gandhi, Robert Foley, Peter Puhov, Theodoros Gkountouvas, Hui Lei 0001
ICDCS1
2021 Multiarmed-Bandit-Based Decentralized Computation Offloading in Fog-Enabled IoT
abstract
The Internet-of-Things (IoT) environments have hard real-time tasks that need execution within fixed deadlines. As IoT devices consist of a myriad of sensors, each task is composed of multiple interdependent subtasks. Toward this, the cloud and fog computing platforms have the potential of facilitating these IoT sensor nodes (SNs) in accommodating complex operations with minimum delay. To further reduce operational latencies, we breakdown the high-level tasks into smaller subtasks and form a directed acyclic task graph (DATG). Initially, the SNs offload their tasks to a nearby fog node (FN) based on a greedy choice. The greedy formulation helps in selecting the FN in linear time while avoiding combinatorial optimizations at the SN, which saves time as well as energy. IoT environments are highly dynamic, which mandates the need for adaptive solutions. At the chosen FN, depending on the dependencies on the DATGs, its corresponding deadlines, and the varying conditions of the other FNs, we propose an ϵ-greedy nonstationary multiarmed bandit-based scheme (D2CIT) for online task allocation among them. The online learning D2CIT scheme allows the FN to autonomously select a set of FNs for distributing the subtasks among themselves and executes the subtasks in parallel with minimum latency, energy, and resource usage. Simulation results show that D2CIT offers a reduction in latency by 17% compared to traditional fog computing schemes. Additionally, upon comparison with existing online learning-based task offloading solutions in fog environments, D2CIT offers an improved speedup of 59% due to the induced parallelism.
Sudip Misra, Sri Pramodh Rachuri, Pallav Kumar Deb, Anandarup Mukherjee
IEEE Internet Things J.2