VLDB 2026 Research / reviewers in the wild / expert
Anshul Gandhi
dblp:62/2403
· DBLP profile ↗
47ranked-venue papers
11as first author
20since 2021 · last 2026
0000-0002-2351-6523ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 7 first-author · 14 since 2021Software engineering, systems software and programming languages · 7 · 3 first-author · 1 since 2021Computer networks · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tensor Parallelism for State-Space Models
Anurag Dutt, Nimit Shah, Hazem Masarani, Anshul Gandhi |
ICDCS | 4 |
| 2025 | Constellate: Establishing the opportunity for Distributed Unit pooling in real-world 5G Radio Access NetworksabstractAs the adoption of Virtualized Radio Access Networks (vRAN) is gaining momentum in 5 G networks, Mobility Network Operators are considering a Centralized RAN (CRAN) architecture that moves the baseband functions to a far-edge cloud in order to gain dimensioning flexibility, resiliency and improved RAN performance. However, there have been limited studies on the benefits of centralization in improving RAN compute utilization, especially in the context of pooling the compute-intensive Distributed Unit (DU) resources. In this paper, we present the first study on the benefits of pooling in improving DU server utilization. Using longitudinal traces from a real-world 5G network, we show that significant Capex and Opex gains of $\mathbf{8 4 \%}$ and $\mathbf{9 4 \%}$, respectively, can be obtained through fine-grained pooling at a granularity of 1 second. We also present an affinitybased and dynamic pooling algorithm that can reduce the pooling overheads while still achieving significant pooling gains. Sri Pramodh Rachuri, Anshul Gandhi, Gueyoung Jung, Shankaranarayanan Puzhavakath Narayanan, Alex Zelezniak |
MASCOTS | 2 |
| 2025 | Foreword - Special Issue - MASCOTS 2023
Mariacarla Calzarossa, Anshul Gandhi |
Perform. Evaluation | 2 |
| 2025 | Editorial: Special issue on Performance Analysis and Evaluation of Systems for Artificial Intelligence
Anshul Gandhi, Shaolei Ren |
Perform. Evaluation | 1 |
| 2024 | KACE: Kernel-Aware Colocation for Efficient GPU Spatial SharingabstractGPU spatial sharing among jobs is an effective approach to increase resource utilization and reduce the monetary and environmental costs of running deep learning workloads. While hardware support for GPU spatial sharing already exists, accurately predicting GPU interference between colocated workloads remains a concern. This makes it challenging to improve GPU utilization by sharing the GPU between workloads without severely impacting their performance. Existing approaches to identify and mitigate GPU interference often require extensive profiling and/or hardware modifications, making them difficult to deploy in practice. Bing-Shiun Han, Tathagata Paul, Zhenhua Liu 0002, Anshul Gandhi |
SoCC | 4 |
| 2024 | EcoEdgeInfer: Dynamically Optimizing Latency and Sustainability for Inference on Edge DevicesabstractThe use of Deep Neural Networks (DNNs) has skyrocketed in recent years. While its applications have brought many benefits and use cases, they also have a significant environmental impact due to the high energy consumption of DNN execution. It has already been acknowledged in the literature that training DNNs is computationally expensive and requires large amounts of energy. However, the energy consumption of DNN inference is still an area that has not received much attention, yet. With the increasing adoption of online tools, the usage of inference has significantly grown and will likely continue to grow. Unlike training, inference is user-facing, requires low latency, and is used more frequently. As such, edge devices are being considered for DNN inference due to their low latency and privacy benefits. In this context, inference on edge is a timely area that requires closer attention to regulate its energy consumption. We present EcoEdgeInfer, a system that balances performance and sustainability for DNN inference on edge devices. Our core component of EcoEdgeInfer is an adaptive optimization algorithm, EcoGD, that strategically and quickly sweeps through the hardware and software configuration space to find the jointly optimal configuration that can minimize energy consumption and latency. EcoGD is agile by design, and adapts the configuration parameters in response to time-varying and unpredictable inference workload. We evaluate EcoEdgeInfer on different DNN models using real-world traces and show that EcoGD consistently outperforms existing baselines, lowering energy consumption by 31% and reducing tail latency by 14%, on average. Sri Pramodh Rachuri, Nazeer Shaik, Mehul Choksi, Anshul Gandhi |
SEC | 4 |
| 2024 | OVIDA: Orchestrator for Video Analytics on Disaggregated ArchitectureabstractMillions of video cameras are deployed globally across major cities for learning-based video analytic (VA) applications, such as object detection. Video streams from the cameras are either sent over the wide-area network to be processed by the cloud or are (at least partially) processed in a local edge workstation, incurring significant latency and elevated financial costs. In this paper, to minimize reliance on the cloud and overcome the unavailability of high-compute workstations on edge, we investigate the use of heterogeneous and distributed embedded devices as edge nodes shared by multiple cameras to fully serve the video processing needs of a VA application (without requiring cloud support). We present OVIDA, an edge-only orchestrator to deploy VA application(s) on a distributed edge environment to maximize accuracy. Given the resource-constrained nature of edge nodes, OVIDA disaggregates the VA application pipeline into multiple modules. OVIDA's core functionality and contributions are: (i) optimizing the placement and replication of the VA application modules across the edge nodes to maximize the throughput, and in turn, accuracy; and (ii) an adaptive model selection algorithm for VA modules based on accuracy-throughput tradeoff to maximize accuracy in response to varying load conditions. To further improve performance, OVIDA employs a central-queue-based design (instead of the usual push-based design), which also obviates the need for complex load balancing algorithms. We implement OVIDA on top of Kubernetes and evaluate its performance for three VA applications, supported over a heterogeneous edge cluster under varying network conditions. When compared against several baselines in our evaluation, we achieve throughput and accuracy gains of at least 51% and 28%. Manavjeet Singh, Sri Pramodh Rachuri, Bryan Bo Cao, Venkata Bhumireddy, Francesco Bronzino, Samir Ranjan Das, Anshul Gandhi, Shubham Jain 0003 |
SEC | 8 |
| 2024 | Representation Similarity: A Better Guidance of DNN Layer Sharing for Edge Computing without TrainingabstractEdge computing has emerged as an alternative to reduce transmission and processing delay and preserve privacy of the video streams. However, the ever-increasing complexity of Deep Neural Networks (DNNs) used in video-based applications (e.g. object detection) exerts pressure on memory-constrained edge devices. Model merging is proposed to reduce the DNNs' memory footprint by keeping only one copy of merged layers' weights in memory. In existing model merging techniques, (i) only architecturally identical layers can be shared; (ii) requires computationally expensive retraining in the cloud; (iii) assumes the availability of ground truth for retraining. The re-evaluation of a merged model's performance, however, requires a validation dataset with ground truth, typically runs at the cloud. Common metrics to guide the selection of shared layers include the size or computational cost of shared layers or representation size. We propose a new model merging scheme by sharing representations (i.e., outputs of layers) at the edge, guided by representation similarity S. We show that S is extremely highly correlated with merged model's accuracy with Pearson Correlation Coefficient |r| > 0.94 than other metrics, demonstrating that representation similarity can serve as a strong validation accuracy indicator without ground truth. We present our preliminary results of the newly proposed model merging scheme with identified challenges, demonstrating a promising research future direction. Bryan Bo Cao, Manavjeet Singh, Anshul Gandhi, Samir Ranjan Das, Shubham Jain 0003 |
MobiCom | 4 |
| 2024 | OPPerTune: Post-Deployment Configuration Tuning of Services Made Easy
Gagan Somashekar, Karan Tandon, Anush Kini, Chieh-Chun Chang, Petr Husak, Ranjita Bhagwan, Mayukh Das, Anshul Gandhi, Nagarajan Natarajan |
NSDI | 8 |
| 2024 | GAMMA: Graph Neural Network-Based Multi-Bottleneck Localization for Microservices ApplicationsabstractMicroservices architecture is quickly replacing monolithic and multi-tier architectures as the implementation choice for large-scale web applications as it allows independent development, scalability, and maintenance. However, even with careful node scheduling and scaling, the microservices applications are still vulnerable to performance degradation due to unexpected (dependent or independent) events like anomalous node behavior, workload interference, or sudden spikes in requests or retries. These events can adversely affect the performance of one or more microservices (bottlenecks), degrading the overall application performance. To ensure a good customer experience and avoid revenue loss, it is crucial to detect and mitigate all bottlenecks swiftly. Gagan Somashekar, Anurag Dutt, Mainak Adak, Tania Lorido-Botran, Anshul Gandhi |
WWW | 5 |
| 2024 | Accelerating multi-tier storage cache simulations using knee detection
Tyler Estro, Mário Antunes 0001, Pranav Bhandari, Anshul Gandhi, Geoffrey H. Kuenning, Carl A. Waldspurger, Avani Wildani, Erez Zadok |
Perform. Evaluation | 4 |
| 2023 | Guiding Simulations of Multi-Tier Storage Caches Using Knee DetectionabstractSimulating storage cache hierarchies enables efficient exploration of their configuration space, including diverse topologies, parameters and policies, and devices with varied performance characteristics, while avoiding expensive physical experiments. Miss Ratio Curves (MRCs) efficiently characterize the performance of a cache over a range of cache sizes. These useful tools reveal “key points” for cache simulation, such as knees in the curve that immediately follow sharp cliffs. Unfortunately, there are no automated techniques for efficiently finding key points in MRCs, and the cross-application of existing knee-detection algorithms yields inaccurate results. We present a multi-stage framework that identifies key points in any MRC, for both stack-based (e.g., LRU) and more sophis-ticated eviction algorithms (e.g., ARC). Our approach quickly locates candidates using efficient hash-based sampling, curve simplification, knee detection, and novel post-processing filters. We introduce Z-Method, a new multi-knee detection algorithm that employs statistical outlier detection to choose promising points robustly and efficiently. We evaluate our framework against seven other knee-detection algorithms, using both ARC and LRU MRCs from 106 diverse real-world workloads, and apply it to identify key points in multi-tier MRCs. Compared to naive approaches, our framework reduces the total number of points needed to accurately identify the best two-tier cache hierarchies by an average factor of approximately$5.5\times$for ARC and$7.7\times$for LRU. Tyler Estro, Mário Antunes 0001, Pranav Bhandari, Anshul Gandhi, Geoffrey H. Kuenning, Carl A. Waldspurger, Avani Wildani, Erez Zadok |
MASCOTS | 4 |
| 2023 | Efficient and accurate Lyapunov function-based truncation technique for multi-dimensional Markov chains with applications to discriminatory processor sharing and priority queues
Gagan Somashekar, Mohammad Delasay, Anshul Gandhi |
Perform. Evaluation | 3 |
| 2022 | Optimizing Near-Data Processing for SparkabstractResource disaggregation (RD) is an emerging paradigm for data center computing whereby resource-optimized servers are employed to minimize resource fragmentation and improve resource utilization. Apache Spark deployed under the RD paradigm employs a cluster of compute-optimized servers to run executors and a cluster of storage-optimized servers to host the data on HDFS. However, the network transfer from storage to compute cluster becomes a severe bottleneck for big data processing. Near-data processing (NDP) is a concept that aims to alleviate network load in such cases by offloading (or "pushing down") some of the compute tasks to the storage cluster. Employing NDP for Spark under the RD paradigm is challenging because storage-optimized servers have limited computational resources and cannot host the entire Spark processing stack. Further, even if such a lightweight stack could be developed and deployed on the storage cluster, it is not entirely obvious which Spark queries would benefit from pushdown, and which tasks of a given query should be pushed down to storage.This paper presents the design and implementation of a near-data processing system for Spark, SparkNDP, that aims to address the aforementioned challenges. SparkNDP works by implementing novel NDP Spark capabilities on the storage cluster using a lightweight library of SQL operators and then developing an analytical model to help determine which Spark tasks should be pushed down to storage based on the current network and system state. Simulation and prototype implementation results show that SparkNDP can help reduce Spark query execution times when compared to both the default approach of not pushing down any tasks to storage and the outright NDP approach of pushing all tasks to storage. Sri Pramodh Rachuri, Arun Gantasala, Prajeeth Emanuel, Anshul Gandhi, Robert Foley, Peter Puhov, Theodoros Gkountouvas, Hui Lei 0001 |
ICDCS | 4 |
| 2022 | Are mobiles ready for BBR?abstractBBR is a new congestion control algorithm that has seen widespread Internet adoption in recent years with an estimated 40% of Internet traffic volume as BBR traffic. While many studies examine the performance and fairness of BBR on desktops and servers, there is still a question of how BBR would behave on mobile devices. This is especially important because mobiles represent a large segment of Internet devices. In this work, we study the potential performance bottlenecks of BBR if it were to be deployed on Android devices. We compare the performance of BBR and the default congestion control algorithm Cubic for different devices and device configurations. We find that BBR performs poorly compared to Cubic, especially under low-end device configurations. Further investigation reveals that this poor performance is because of packet pacing which is enabled in BBR by default. Pacing increases the computational overhead, which can affect performance for low-end devices. To address this problem, we propose a first cut solution that modifies BBR's pacing behavior to improve performance while still retaining the benefits of packet pacing. Santiago Vargas, Gautham Gunapati, Anshul Gandhi, Aruna Balasubramanian |
IMC | 3 |
| 2022 | Truncating Multi-Dimensional Markov Chains With Accuracy GuaranteeabstractThe ability to obtain the steady-state probability distribution of a Markov chain is invaluable for modern service providers who aim to satisfy arbitrary tail performance requirements. However, it is often challenging and even intractable to obtain the steady-state distribution for several classes of Markov chains, such as multi-dimensional and infinite state-space Markov chains with state-dependent transitions. Two examples include the M/M/1 with Discriminatory Processor Sharing (DPS) and the preemptive M/M/c with multiple priority classes and customer abandonment. This paper proposes a Lyapunov function-based state-space truncation technique for such Markov chains. Our technique leverages the available moments, or bounds on moments, of the state variables of the Markov chain to obtain tight truncation bounds while satisfying arbitrary probability mass guarantees for the truncated chain. We demonstrate the efficacy of our technique for the multi-dimensional DPS and M/M/c priority queue with abandonment and highlight the significant reduction in state space (as much as 72%) afforded by our approach compared to the state-of-the-art. Gagan Somashekar, Mohammad Delasay, Anshul Gandhi |
MASCOTS | 3 |
| 2022 | User-Centric Interference-Aware Load Balancing for Cloud-Deployed ApplicationsabstractVMs deployed in cloud environments are prone to performance interference due to dynamic and unpredictable contention for shared physical resources among colocated tenants. Current provider-centric solutions, such as careful co-scheduling of VMs and/or VM migration, require a priori profiling of customer VMs, which is infeasible in public clouds. Further, such solutions are not always aware of the user's SLO requirements or application bottlenecks. This paper presents DIAL, an interference-aware load balancing framework that can directly be employed by cloud users without requiring any assistance from the provider. The key idea behind DIAL is to infer the demand for contended resources on the physical hosts, which is otherwise hidden from users. Estimates of the colocated load are then used to dynamically shift load away from compromised VMs without violating the application's tail latency SLOs. We implement DIAL for web and online analytical processing applications, and show, via experimental results on OpenStack and AWS clouds, that DIAL can reduce tail latencies by as much as 70 percent compared to existing solutions. Seyyed Ahmad Javadi, Anshul Gandhi |
IEEE Trans. Cloud Comput. | 2 |
| 2021 | ServerMore: Opportunistic Execution of Serverless Functions in the CloudabstractServerless computing allows customers to submit their jobs to the cloud for execution, with the resource provisioning being taken care of by the cloud provider. Serverless functions are often short-lived and have modest resource requirements, thereby presenting an opportunity to improve server utilization by colocating with latency-sensitive customer workloads. This paper presents ServerMore, a server-level resource manager that opportunistically colocates customer serverless jobs with serverful customer VMs. ServerMore dynamically regulates the CPU, memory bandwidth, and LLC resources on the server to ensure that the colocation between serverful and serverless workloads does not impact application tail latencies. By selectively admitting serverless functions and inferring the performance of black-box serverful workloads, ServerMore improves resource utilization on average by 35.9% to 245% compared to prior works; while having a minimal impact on the latency of both serverful applications and serverless functions. Amoghavarsha Suresh, Anshul Gandhi |
SoCC | 2 |
| 2021 | Towards optimal placement and scheduling of DNN operations with PestoabstractThe increasing size of Deep Neural Networks (DNNs) has necessitated the use of multiple GPUs to host a single DNN model, a practice commonly referred to as model parallelism. The key challenge for model parallelism is to efficiently and effectively partition the DNN model across GPUs to avoid communication overheads while maximizing the GPU utilization, with the end-goal of minimizing the training time of DNN models. Existing approaches either take a long time(hours or even days) to find an effective partition or settle for sub-optimal partitioning, invariably increasing the end-to-end training effort. In this paper, we design and implement Pesto, a fast and near-optimal model placement technique for automatically partitioning arbitrary DNNs across multiple GPUs. The key idea in Pesto is to jointly optimize the model placement and scheduling at the fine-grained operation level to minimize inter-GPU communication while maximizing the opportunity to parallelize the model across GPUs. By carefully formulating the problem as an integer program, Pesto can provide the optimal placement and scheduling. We implement Pesto in TensorFlow and show that Pesto can reduce model training time by up to 31% compared to state-of-the-art approaches, across several large DNN models. Ubaid Ullah Hafeez, Xiao Sun 0004, Anshul Gandhi, Zhenhua Liu 0002 |
Middleware | 3 |
| 2021 | BBR Bufferbloat in DASH VideoabstractBBR is a new congestion control algorithm and is seeing increased adoption especially for video traffic. BBR solves the bufferbloat problem in legacy loss-based congestion control algorithms where application performance drops considerably when router buffers are deep. BBR regulates traffic such that router queues don’t build up to avoid the bufferbloat problem while still maintaining high throughput. However, our analysis shows that video applications experience significantly poor performance when using BBR under deep buffers. In fact, we find that video traffic sees inflated latencies because of long queues at the router, ultimately degrading video performance. To understand this dichotomy, we study the interaction between BBR and DASH video. Our investigation reveals that BBR under deep buffers and high network burstiness severely overestimates available bandwidth and does not converge to steady state, both of which results in BBR sending substantially more data into the network, causing a queue buildup. This elevated packet sending rate under BBR is ultimately caused by the router’s ability to absorb bursts in traffic, which destabilizes BBR’s bandwidth estimation and overrides BBR’s expected logic for exiting the startup phase. We design a new bandwidth estimation algorithm and apply it to BBR (and a still-unreleased, newer version of BBR called BBR2). Our modified BBR and BBR2 both see significantly improved video QoE even under deep buffers. Santiago Vargas, Rebecca Drucker, Aiswarya Renganathan, Aruna Balasubramanian, Anshul Gandhi |
WWW | 5 |
| 2020 | Analyzing the distribution fit for storage workload and Internet traffic traces
Muhammad Wajahat, Aditya Yele, Tyler Estro, Anshul Gandhi, Erez Zadok |
Perform. Evaluation | 4 |
| 2020 | Providing Performance Guarantees for Cloud-Deployed ApplicationsabstractApplications with a dynamic workload demand need access to a flexible infrastructure to meet performance guarantees and minimize resource costs. While cloud computing provides the elasticity to scale the infrastructure on demand, cloud service providers lack control and visibility of user space applications, making it difficult to accurately scale the infrastructure. Thus, the burden of scaling falls on the user. That is, the user must determine when to trigger scaling and how much to scale. Scaling becomes even more challenging when applications exhibit dynamic changes in their behavior. In this paper, we propose a new cloud service, Dependable Compute Cloud (DC2), that automatically scales the infrastructure to meet the user-specified performance requirements, even when multiple user requests execute concurrently. DC2 employs Kalman filtering to automatically learn the (possibly changing) system parameters for each application, allowing it to proactively scale the infrastructure to meet performance guarantees. Importantly, DC2 is designed for the cloud - it is application-agnostic and does not require any offline application profiling or benchmarking, training data, or expert knowledge about the application. We evaluate DC2 via implementation on OpenStack using a multi-tier application under a range of workload mixes and arrival traces. Our experimental results demonstrate the robustness and superiority of DC2 over existing rule-based approaches with respect to avoiding SLA violations and minimizing resource consumption. Anshul Gandhi, Parijat Dube, Alexei A. Karve, Andrzej Kochut, Li Zhang 0002 |
IEEE Trans. Cloud Comput. | 1 |
| 2019 | Scavenger: A Black-Box Batch Workload Resource Manager for Improving Utilization in Cloud EnvironmentsabstractResource under-utilization is common in cloud data centers. Prior works have proposed improving utilization by running provider workloads in the background, colocated with tenant workloads. However, an important challenge that has still not been addressed is considering the tenant workloads as a black-box. We present Scavenger, a batch workload manager that opportunistically runs containerized batch jobs next to black-box tenant VMs to improve utilization. Scavenger is designed to work without requiring any offline profiling or prior information about the tenant workload. To meet the tenant VMs' resource demand at all times, Scavenger dynamically regulates the resource usage of batch jobs, including processor usage, memory capacity, and network bandwidth. We experimentally evaluate Scavenger on two different testbeds using latency-sensitive tenant workloads colocated with Spark jobs in the background and show that Scavenger significantly increases resource usage without compromising the resource demands of tenant VMs. Seyyed Ahmad Javadi, Amoghavarsha Suresh, Muhammad Wajahat, Anshul Gandhi |
SoCC | 4 |
| 2019 | Towards Automated Patch Management in a Hybrid Cloud
Ubaid Ullah Hafeez, Alexei A. Karve, Braulio Dumba, Anshul Gandhi, Sai Zeng |
ICSOC | 4 |
| 2019 | When to use and when not to use BBR: An empirical analysis and evaluation studyabstractThis short paper presents a detailed empirical study of BBR's performance under different real-world and emulated testbeds across a range of network operating conditions. Our empirical results help to identify network conditions under which BBR outperforms, in terms of goodput, contemporary TCP congestion control algorithms. We find that BBR is well suited for networks with shallow buffers, despite its high retransmissions, whereas existing loss-based algorithms are better suited for deep buffers. Kriti Sharma, Aruna Balasubramanian, Anshul Gandhi |
Internet Measurement Conference | 5 |
| 2019 | ECON: Modeling the network to improve application performanceabstractGiven the growing significance of network performance, it is crucial to examine how to make the most of available network options and protocols. We propose ECON, a model that predicts performance of applications under different protocols and network conditions to scalably make better network choices. ECON is built on an analytical framework to predict TCP performance, and uses the TCP model as a building block for predicting application performance. ECON infers a relationship between loss and congestion using empirical data that drives an online model to predict TCP performance. ECON then builds on the TCP model to predict latency and HTTP performance. Across four wired and one wireless network, our model outperforms seven alternative TCP models. We demonstrate how ECON (i) can be used by a Web server application to choose between HTTP/1.1 and HTTP/2 for a given Web page and network condition, and (ii) can be used by a video application to choose the optimal bitrate that maximizes video quality without rebuffering. Javad Nejati, Aruna Balasubramanian, Anshul Gandhi |
Internet Measurement Conference | 4 |
| 2019 | Optimal Markovian Dynamic Control of Interference-Prone Server FarmsabstractInterference is a key performance challenge faced by cloud users, and can significantly degrade application performance on virtual machines (VMs). For load-balanced cloud applications, a key question is how to distribute the load among VMs in the presence of interference. Using a Markov decision process (MDP) model, we investigate dynamic control polices to assign jobs among a cluster of VMs that are prone to interference in a system with a central queue and an arbitrary number of VMs. We characterize the structural properties of the MDP optimality equation, and we prove that the optimal control policy is a threshold policy based on the queue length. The optimal policy is characterized by multiple thresholds depending on the current conditions of the VMs, including the number of busy under-interference VMs. We discuss the existence of an ordering among such thresholds, and we prove the ordering for a two-VM system. Our numerical results show that the optimal dynamic policy can significantly improve performance compared to the the commonly employed non-idling policy. For low utilization systems, we observe improvements on the order of around 20%. We further implement the optimal policy in a real-world testbed using the HAProxy load balancer, and show that it can reduce web server response times by as much as 40%–60%, even for time-varying request rates. Scott Votke, Jazeem Abdul Jaleel, Amoghavarsha Suresh, Mohammad Delasay, Sherwin Doroudi, Anshul Gandhi |
MASCOTS | 6 |
| 2019 | Distribution Fitting and Performance Modeling for Storage TracesabstractUnderstanding I/O workloads and modeling their performance is important for optimizing storage systems. A useful first step towards understanding the characteristics of storage workloads is to analyze their inter-arrival times and service requirements. If these characteristics are found to follow certain probability distributions, then corresponding stochastic models can be employed to efficiently estimate the performance of storage workloads. Such approaches have been explored in other domains using an assortment of distributions, including the Normal, Weibull, and Exponential. However, our analysis and others' past attempts revealed that none of those distributions provided a good fit for storage workloads. We analyzed over 200 traces across 4 different workload families using 20 widely used distributions, including ones seldom used for storage modeling. We found that the Hyper-exponential distribution with just two phases H_2 was superior in modeling the storage traces compared to other distributions under five diverse metrics of accuracy, including metrics that assess the risk of over-fitting. Based on these results, we developed a Markov-chain-based stochastic model that accurately estimates the storage system performance across several workload traces. To highlight the applicability of our model, we conducted what-if analyses to investigate the performance impact of workload variability and garbage collection under various scenarios. Muhammad Wajahat, Aditya Yele, Tyler Estro, Anshul Gandhi, Erez Zadok |
MASCOTS | 4 |
| 2019 | Using Variability as a Guiding Principle to Reduce Latency in Web Applications via OS ProfilingabstractRequest latency is a critical metric in determining the usability of web services. The latency of a request includes service time - the time when the request is being actively serviced - and waiting time - the time when the request is waiting to be served. Most existing works aim to reduce request latency by focusing on reducing the mean service time (that is, shortening the critical path). Amoghavarsha Suresh, Anshul Gandhi |
WWW | 2 |
| 2018 | Application-Agnostic Batch Workload Management in Cloud EnvironmentsabstractWe present Scavenger, a reactive batch workload manager that opportunistically runs containerized batch jobs next to customer Virtual Machines (VMs) in a public cloud like setting to improve utilization. Scavenger dynamically regulates the resource usage of batch jobs, including CPU usage, memory capacity, and LLC capacity, to ensure that the customer VMs' resource demand is met at all times. We experimentally evaluate Scavenger and show that it considerably increases resource usage without compromising on the resource demand of customer VMs. Importantly, Scavenger does so without requiring any offline profiling or prior information about the customer workloads. Seyyed Ahmad Javadi, Shalini Bhaskara, Rahul Doshi, Prashanth Soundarapandian, Muhammad Wajahat, Anshul Gandhi |
SoCC | 6 |
| 2018 | ElMem: Towards an Elastic Memcached SystemabstractMemory caches, such as Memcached, are a critical component of online applications as they help maintain low latencies by alleviating the load at the database. However, memory caches are expensive, both in terms of power and operating costs. It is thus important to dynamically scale such caches in response to workload variations. Unfortunately, stateful systems, such as Memcached, are not elastic in nature. The performance loss that follows a scaling action can severely impact latencies and lead to SLO violations. This paper proposes ElMem, an elastic Memcached system that mitigates post-scaling performance loss by proactively migration hot data between nodes. The key enabler of our work is an efficient algorithm, FuseCache, that migrates the optimal amount of hot data to minimize performance loss. Our experimental results on OpenStack, across several workload traces, show that ElMem elastically scales Memcached while reducing the postscaling performance degradation by about 90%. Ubaid Ullah Hafeez, Muhammad Wajahat, Anshul Gandhi |
ICDCS | 3 |
| 2018 | EASY: Efficient Segment Assignment Strategy for Reducing Tail Latencies in PinotabstractCustomer facing online services, such as LinkedIn and Uber, rely on scalable and low-latency data stores to maintain acceptable query tail latencies. An important challenge for managing the performance of these systems is the assignment of newly created data segments to data nodes to balance load. Given the rate at which these services are accessed (thus generating new data), the segment assignment problem is particularly important. This paper presents EASY, an efficient segment assignment strategy that leverages analytical modeling to predict the future load induced by data segments, thus allowing for long-term balancing of load across data nodes. Our implementation and evaluation of EASY on Pinot shows that we can significantly reduce query tail latencies in the presence of dynamically generated data segments. Seyyed Ahmad Javadi, Robin Manhas, Shweta Sahu, Anshul Gandhi |
ICDCS | 5 |
| 2018 | A FrameNet for Cancer Information in Clinical Narratives: Schema and Annotation
Kirk Roberts, Yuqi Si, Anshul Gandhi, Elmer V. Bernstam |
LREC | 3 |
| 2018 | A Model-Driven Graybox Approach to Rehoming Service ChainsabstractNetwork clouds are typically private clouds owned by the network provider, consisting of a large number of geo-distributed sites with heterogeneous capabilities and small capacities. Each of these small clouds often run specialized service chains of Virtual Network Functions (VNFs), which need to meet strict Service Level Objectives (SLOs), especially along the lines of availability (e.g., First responder services). Hence, VNFs in such thinly provisioned clouds may need to be moved (rehomed), both within and across sites, much more frequently than in traditional public clouds (like Amazon's EC2 cloud), in order to meet the performance SLOs, when reacting to various cloud events like hotspots, interference from co-located VMs, failures and upgrades. Rehoming is also required by the infrastructure (platform) providers for various other reasons such as consolidation of resources for saving energy and improving the platform utilization. In this paper, we propose a model-based approach to show that naive strategies for rehoming, applied uniformly across all VNFs of the service chain, are often sub-optimal when considering different metrics like user-perceived service disruption time and the time taken to complete the rehoming action. Our model leverages the transparency between the services and platforms on private clouds (grayness), and provides appropriate rehoming recommendations based on various factors including service characteristics and runtime platform dynamics. We validate our models using a simple, yet ubiquitously deployed service chain, and using out-of-the-box rehoming options provided by Openstack, the most commonly used open-source cloud. Our results show that our graybox approach is able to achieve significant reductions in service disruption times and time taken for the rehoming action. Muhammad Wajahat, Bharath Balasubramanian, Anshul Gandhi, Gueyoung Jung, Shankaranarayanan Puzhavakath Narayanan |
MASCOTS | 3 |
| 2018 | Model-driven optimal resource scaling in cloud
Anshul Gandhi, Parijat Dube, Alexei A. Karve, Andrzej Kochut, Li Zhang 0002 |
Softw. Syst. Model. | 1 |
| 2017 | Modeling and Analysis of Performance Under Interference in the CloudabstractOne of the key performance challenges in cloud computing is the problem of interference, or resource contention, among colocated VMs. While prior work has empirically analyzed interference for specific workloads under specific settings, there is a need for a generic approach to estimate application performance under any interference condition.In this paper, we present an analytical model to estimate performance as a function of various workload, system, and interference conditions, including the intensity and length ofinterference, for single- and multi-VM systems. Comparisons with empirical results under various scenarios show that our model can provide accurate latency estimations (less than 5% error). We employ our model to analyze systems under interference, and derive useful results to aid practitioners. Scott Votke, Seyyed Ahmad Javadi, Anshul Gandhi |
MASCOTS | 3 |
| 2016 | Autoscaling for Hadoop ClustersabstractUnforeseen events such as node failures and resource contention can have a severe impact on the performance of data processing frameworks, such as Hadoop, especially in cloud environments where such incidents are common. SLA compliance in the presence of such events requires the ability to quickly and dynamically resize infrastructure resources. Unfortunately, the distributed and stateful nature of data processing frameworks makes it challenging to accurately scale the system at run-time. In this paper, we present the design and implementation of a model-driven autoscaling solution for Hadoop clusters. We first develop novel gray-box performance models for Hadoop workloads that specifically relate job execution times to resource allocation and workload parameters. We then employ these models to dynamically determine the resources required to successfully complete the Hadoop jobs as per the user-specified SLA under various scenarios including node failures and multi-job executions. Our experimental results on three different Hadoop cloud clusters and across different workloads demonstrate the efficacy of our models and highlight their autoscaling capabilities. Anshul Gandhi, Sidhartha Thota, Parijat Dube, Andrzej Kochut, Li Zhang 0002 |
IC2E | 1 |
| 2016 | UIE: User-Centric Interference Estimation for Cloud ApplicationsabstractInterference is one of the key deterrents to cloud adoption, and is known to cause severe degradation in application performance, costing service providers in lost revenues. In this paper, we present UIE, a user-centric approach to detecting and, importantly, estimating the degree of interference experienced by user applications in the cloud. UIE employs queueing theory to model the impact of resource contention on application performance. By leveraging UIE, users can estimate the true amount of resources, including CPU, network, and I/O, allocated to their application at any given time, without any assistance from the cloud provider or hypervisor. Seyyed Ahmad Javadi, Sagar Mehra, Bharath Kumar Reddy Vangoor, Anshul Gandhi |
IC2E | 4 |
| 2016 | Analyzing the Power Consumption of the Mobile Page LoadabstractNo abstract available. Javad Nejati, Pavan Maguluri, Aruna Balasubramanian, Anshul Gandhi |
SIGMETRICS | 5 |
| 2016 | Using Predictions in Online Optimization: Looking Forward with an Eye on the PastabstractWe consider online convex optimization (OCO) problems with switching costs and noisy predictions. While the design of online algorithms for OCO problems has received considerable attention, the design of algorithms in the context of noisy predictions is largely open. To this point, two promising algorithms have been proposed: Receding Horizon Control (RHC) and Averaging Fixed Horizon Control (AFHC). The comparison of these policies is largely open. AFHC has been shown to provide better worst-case performance, while RHC outperforms AFHC in many realistic settings. In this paper, we introduce a new class of policies, Committed Horizon Control (CHC), that generalizes both RHC and AFHC. We provide average-case analysis and concentration results for CHC policies, yielding the first analysis of RHC for OCO problems with noisy predictions. Further, we provide explicit results characterizing the optimal CHC policy as a function of properties of the prediction noise, e.g., variance and correlation structure. Our results provide a characterization of when AFHC outperforms RHC and vice versa, as well as when other CHC policies outperform both RHC and AFHC. Niangjun Chen, Joshua Comden, Zhenhua Liu 0002, Anshul Gandhi, Adam Wierman |
SIGMETRICS | 4 |
| 2015 | HALO: Heterogeneity-Aware Load BalancingabstractLoad Balancers (LBs) play a critical role in managing the performance and resource utilization of distributed systems. However, developing efficient LBs for large, distributed clusters is challenging for several reasons: (i) large clusters require numerous scheduling decisions per second, (ii) such clusters typically consist of heterogeneous servers that widely differ in their computing power, and (iii) such clusters often experience significant changes in load. In this paper we propose HALO, a class of scalable, heterogeneity-aware LBs for cluster systems. HALO LBs are based on simple randomized algorithms that are analytically optimized for heterogeneity. We develop HALO for randomized, Round-Robin, and Power-of-D LBs. We illustrate the benefits of HALO and demonstrate its superiority over other comparable LBs using analytical, simulation, and (Apache-based) implementation results. Our results show that HALO LBs provide significantly lower response times without incurring additional overhead across a wide range of scenarios. Anshul Gandhi, Naman Mittal |
MASCOTS | 1 |
| 2014 | Modeling the Impact of Workload on Cloud Resource ScalingabstractCloud computing offers the flexibility to dynamically size the infrastructure in response to changes in workload demand. While both horizontal and vertical scaling of infrastructure is supported by major cloud providers, these scaling options differ significantly in terms of their cost, provisioning time, and their impact on workload performance. Importantly, the efficacy of horizontal and vertical scaling critically depends on the workload characteristics, such as the workload's parallelizability and its core scalability. In today's cloud systems, the scaling decision is left to the users, requiring them to fully understand the tradeoffs associated with the different scaling options. In this paper, we present our solution for optimizing the resource scaling of cloud deployments via implementation in OpenStack. The key component of our solution is the modelling engine that characterizes the workload and then quantitatively evaluates different scaling options for that workload. Our modelling engine leverages Amdahl's Law to model service time scaling in scaleup environments and queueing-theoretic concepts to model performance scaling in scale-out environments. We further employ Kalman filtering to account for inaccuracies in the model-based methodology, and to dynamically track changes in the workload and cloud environment. Anshul Gandhi, Parijat Dube, Alexei A. Karve, Andrzej Kochut, Li Zhang 0002 |
SBAC-PAD | 1 |
| 2013 | Exact analysis of the M/M/k/setup class of Markov chains via recursive renewal rewardabstractThe M/M/k/setup model, where there is a penalty for turning servers on, is common in data centers, call centers and manufacturing systems. Setup costs take the form of a time delay, and sometimes there is additionally a power penalty, as in the case of data centers. While the M/M/1/setup was exactly analyzed in 1964, no exact analysis exists to date for the M/M/k/setup with k>1. In this paper we provide the first exact, closed-form analysis for the M/M/k/setup and some of its important variants including systems in which idle servers delay for a period of time before turning off or can be put to sleep. Our analysis is made possible by our development of a new technique, Recursive Renewal Reward (RRR), for solving Markov chains with a repeating structure. RRR uses ideas from renewal reward theory and busy period analysis to obtain closed-form expressions for metrics of interest such as the transform of time in system and the transform of power consumed by the system. The simplicity, intuitiveness, and versatility of RRR makes it useful for analyzing Markov chains far beyond the M/M/k/setup. In general, RRR should be used to reduce the analysis of any 2-dimensional Markov chain which is infinite in at most one dimension and repeating to the problem of solving a system of polynomial equations. In the case where all transitions in the repeating portion of the Markov chain are skip-free and all up/down arrows are unidirectional, the resulting system of equations will yield a closed-form solution. Anshul Gandhi, Sherwin Doroudi, Mor Harchol-Balter, Alan Scheller-Wolf |
SIGMETRICS | 1 |
| 2012 | SOFTScale: Stealing Opportunistically for Transient Scaling
Anshul Gandhi, Timothy Zhu, Mor Harchol-Balter, Michael A. Kozuch |
Middleware | 1 |
| 2012 | AutoScale: Dynamic, Robust Capacity Management for Multi-Tier Data CentersabstractEnergy costs for data centers continue to rise, already exceeding $15 billion yearly. Sadly much of this power is wasted. Servers are only busy 10--30% of the time on average, but they are often left on, while idle, utilizing 60% or more of peak power when in the idle state. We introduce a dynamic capacity management policy, AutoScale , that greatly reduces the number of servers needed in data centers driven by unpredictable, time-varying load, while meeting response time SLAs. AutoScale scales the data center capacity, adding or removing servers as needed. AutoScale has two key features: (i) it autonomically maintains just the right amount of spare capacity to handle bursts in the request rate; and (ii) it is robust not just to changes in the request rate of real-world traces, but also request size and server efficiency. We evaluate our dynamic capacity management approach via implementation on a 38-server multi-tier data center, serving a web site of the type seen in Facebook or Amazon, with a key-value store workload. We demonstrate that AutoScale vastly improves upon existing dynamic capacity management policies with respect to meeting SLAs and robustness. Anshul Gandhi, Mor Harchol-Balter, Ram Raghunathan, Michael A. Kozuch |
ACM Trans. Comput. Syst. | 1 |
| 2010 | Optimality analysis of energy-performance trade-off for server farm management
Anshul Gandhi, Varun Gupta 0004, Mor Harchol-Balter, Michael A. Kozuch |
Perform. Evaluation | 1 |
| 2010 | Server farms with setup costs
Anshul Gandhi, Mor Harchol-Balter, Ivo J. B. F. Adan |
Perform. Evaluation | 1 |