Alok Gautam Kumbhare

dblp:10/10279 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
2since 2021 · last 2021
0000-0001-5433-4688ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Energy-efficient computing · 37% Cloud and datacenter computing · 34% Distributed systems · 22%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cluster resource management and scheduling
0.722021
Prediction-Based Power Oversubscription in Cloud Platforms · USENIX ATC 2021
Exploiting application dynamism and cloud elasticity for continuous dataflows · SC 2013
Energy-efficient computing
datacenter power management
0.512021
Flex: High-Availability Datacenters With Zero Reserved Power · ISCA 2021
Energy-efficient computing
power management
0.512021
Prediction-Based Power Oversubscription in Cloud Platforms · USENIX ATC 2021
Energy-efficient computing › datacenter power management
power oversubscription
0.512021
Prediction-Based Power Oversubscription in Cloud Platforms · USENIX ATC 2021
Cloud and datacenter computing › resource allocation
workload allocation
0.512021
Flex: High-Availability Datacenters With Zero Reserved Power · ISCA 2021
Memory systems › cache coherence
cache coherence protocol
0.312017
Geo-distribution of actor-based services · Proc. ACM Program. Lang. 2017
Distributed systems › distributed system architecture
geo-distributed systems
0.312017
Geo-distribution of actor-based services · Proc. ACM Program. Lang. 2017
Distributed systems
replication and consistency
0.312017
Geo-distribution of actor-based services · Proc. ACM Program. Lang. 2017
Distributed systems
distributed coordination
0.212013
Exploiting application dynamism and cloud elasticity for continuous dataflows · SC 2013
Distributed systems
fault tolerance
0.112021
Flex: High-Availability Datacenters With Zero Reserved Power · ISCA 2021
Cloud and datacenter computing › datacenter storage
datacenter replication
0.112017
Geo-distribution of actor-based services · Proc. ACM Program. Lang. 2017
Cloud and datacenter computing › elastic computing
cloud elasticity
0.012013
Exploiting application dynamism and cloud elasticity for continuous dataflows · SC 2013
Cloud and datacenter computing
virtualization
0.012013
Exploiting application dynamism and cloud elasticity for continuous dataflows · SC 2013

Methods — techniques the papers use, named apart from their topics

prediction · 0.5offline workload placement · 0.5failure monitoring · 0.5linearizable consistency · 0.3eventual consistency · 0.3distributed cache coherence · 0.3variable sized bin packing heuristics · 0.2optimization · 0.2
YearPublicationVenuePosition
2021 Flex: High-Availability Datacenters With Zero Reserved Power
abstract
Cloud providers, like Amazon and Microsoft, must guarantee high availability for a large fraction of their workloads. For this reason, they build datacenters with redundant infrastructures for power delivery and cooling. Typically, the redundant resources are reserved for use only during infrastructure failure or maintenance events, so that workload performance and availability do not suffer. Unfortunately, the reserved resources also produce lower power utilization and, consequently, require more datacenters to be built. To address these problems, in this paper we propose "zero-reserved-power" datacenters and the Flex system to ensure that workloads still receive their desired performance and availability. Flex leverages the existence of software-redundant workloads that can tolerate lower infrastructure availability, while imposing minimal (if any) performance degradation for those that require high infrastructure availability. Flex mainly comprises (1) a new offline workload placement policy that reduces stranded power while ensuring safety during failure or maintenance events, and (2) a distributed system that monitors for failures and quickly reduces the power draw while respecting the workloads’ requirements, when it detects a failure. Our evaluation shows that Flex produces less than 5% stranded power and increases the number of deployed servers by up to 33%, which translates to hundreds of millions of dollars in construction cost savings per datacenter site. We end the paper with lessons from our experience bringing Flex to production in Microsoft’s datacenters.
Chaojie Zhang 0001, Alok Gautam Kumbhare, Ioannis Manousakis, Deli Zhang, Pulkit A. Misra, Rod Assis, Kyle Woolcock, Nithish Mahalingam, Brijesh Warrier, David Gauthier, Lalu Kunnath, Steve Solomon, Osvaldo Morales, Marcus Fontoura, Ricardo Bianchini
ISCA2
2021 Prediction-Based Power Oversubscription in Cloud Platforms
Alok Gautam Kumbhare, Ioannis Manousakis, Anand Bonde, Felipe Vieira Frujeri, Nithish Mahalingam, Pulkit A. Misra, Seyyed Ahmad Javadi, Bianca Schroeder, Marcus Fontoura, Ricardo Bianchini
USENIX ATC1
2017 Geo-distribution of actor-based services
abstract
Many service applications use actors as a programming model for the middle tier, to simplify synchronization, fault-tolerance, and scalability. However, efficient operation of such actors in multiple, geographically distant datacenters is challenging, due to the very high communication latency. Caching and replication are essential to hide latency and exploit locality; but it is not a priori clear how to combine these techniques with the actor programming model. We present Geo, an open-source geo-distributed actor system that improves performance by caching actor states in one or more datacenters, yet guarantees the existence of a single latest version by virtue of a distributed cache coherence protocol. Geo's programming model supports both volatile and persistent actors, and supports updates with a choice of linearizable and eventual consistency. Our evaluation on several workloads shows substantial performance benefits, and confirms the advantage of supporting both replicated and single-instance coherence protocols as configuration choices. For example, replication can provide fast, always-available reads and updates globally, while batching of linearizable storage accesses at a single location can boost the throughput of an order processing workload by 7x.
Philip A. Bernstein, Sebastian Burckhardt, Sergey Bykov, Natacha Crooks, Jose M. Faleiro, Gabriel Kliot, Alok Gautam Kumbhare, Muntasir Raihan Rahman, Vivek Shah 0001, Adriana Szekeres, Jorgen Thelin
Proc. ACM Program. Lang.7
2015 Real-Time Analytics for Fast Evolving Social Graphs
abstract
Existing Big Data streams coming from social and other connected sensor networks exhibit intrinsic inter-dependency enabling unique challenges to scalable graph analytics. Data from these graphs is usually collected in different geographically located data servers making it suitable for distributed processing on clouds. While numerous solutions for large scale static graph analysis have been proposed, addressing in real-time the dynamics of social interactions requires novel approaches that leverage incremental stream processing and graph analytics on elastic clouds. We propose a scalable solution based on our stream processing engine, Floe, on top of which we perform real-time data processing and graph updates to enable low latency graph analytics on large evolving social networks. We demonstrate the platform on a large Twitter data set by performing several fast graph and non-graph analytics to extract in real-time the top k influential nodes, with different metrics, during key events such as the US NFL playoffs. This information allows advertisers to maximize their exposure to the public by always targeting the continuously changing set of most influential nodes. Its applicability spans multiple domains including surveillance, counter-terrorism, or disease spread monitoring. The evaluation will be performed on a combination our local cluster of 16 eight-core nodes running Eucalyptus fabric and 100s of virtual machines on the Amazon AWS public cloud. We will showcase the low latency in detecting changes in the graph under variable data streams, and also the efficiency of the platform to utilize resources and to elastically scale to meet demand.
Charith Wickramaarachchi, Alok Gautam Kumbhare, Marc Frîncu, Charalampos Chelmis, Viktor Prasanna 0001
CCGRID2
2015 Fault-Tolerant and Elastic Streaming MapReduce with Decentralized Coordination
abstract
The MapReduce programming model, due to its simplicity and scalability, has become an essential tool for processing large data volumes in distributed environments. Recent Stream Processing Systems (SPS) this model to provide low-latency analysis of high-velocity continuous data streams. However, integrating MapReduce with streaming poses challenges: first, the runtime variations in data characteristics such as data-rates and key-distribution cause resource overload, that in-turn leads to fluctuations in the Quality of the Service (QoS), and second, the stateful reducers, whose state depends on the complete tuple history, necessitates efficient fault-recovery mechanisms to maintain the desired QoS in the presence of resource failures. We propose an integrated streaming MapReduce architecture leveraging the concept of consistent hashing to support runtime elasticity along with locality-aware data and state replication to provide efficient load-balancing with low-overhead fault-tolerance and parallel fault-recovery from multiple simultaneous failures. Our evaluation on a private cloud shows up to 2.8× improvement in peak throughput compared to Apache Storm SPS, and a low recovery latency of 700 - 1500 ms from multiple failures.
Alok Gautam Kumbhare, Marc Frîncu, Yogesh L. Simmhan, Viktor Prasanna 0001
ICDCS1
2015 Distributed Programming over Time-Series Graphs
abstract
Graphs are a key form of Big Data, and performing scalable analytics over them is invaluable to many domains. There is an emerging class of inter-connected data which accumulates or varies over time, and on which novel algorithms both over the network structure and across the time-variant attribute values is necessary. We formalize the notion of time-series graphs and propose a Temporally Iterative BSP programming abstraction to develop algorithms on such datasets using several design patterns. Our abstractions leverage a sub-graph centric programming model and extend it to the temporal dimension. We present three time-series graph algorithms based on these design patterns and abstractions, and analyze their performance using the Offish distributed platform on Amazon AWS Cloud. Our results demonstrate the efficacy of the abstractions to develop practical time-series graph algorithms, and scale them on commodity hardware.
Yogesh L. Simmhan, Neel Choudhury, Charith Wickramaarachchi, Alok Gautam Kumbhare, Marc Frîncu, Cauligi S. Raghavendra, Viktor Prasanna 0001
IPDPS4
2015 Reactive Resource Provisioning Heuristics for Dynamic Dataflows on Cloud Infrastructure
abstract
The need for low latency analysis over high-velocity data streams motivates the need for distributed continuous dataflow systems. Contemporary stream processing systems use simple techniques to scale on elastic cloud resources to handle variable data rates. However, application QoS is also impacted by variability in resource performance exhibited by clouds and hence necessitates autonomic methods of provisioning elastic resources to support such applications on cloud infrastructure. We develop the concept of “dynamic dataflows” which utilize alternate tasks as additional control over the dataflow's cost and QoS. Further, we formalize an optimization problem to represent deployment and runtime resource provisioning that allows us to balance the application's QoS, value, and the resource cost. We propose two greedy heuristics, centralized and sharded, based on the variable-sized bin packing algorithm and compare against a Genetic Algorithm (GA) based heuristic that gives a near-optimal solution. A large-scale simulation study, using the linear road benchmark and VM performance traces from the AWS public cloud, shows that while GA-based heuristic provides a better quality schedule, the greedy heuristics are more practical, and can intelligently utilize cloud elasticity to mitigate the effect of variability, both in input data rates and cloud resource performance, to meet the QoS of fast data applications.
Alok Gautam Kumbhare, Yogesh L. Simmhan, Marc Frîncu, Viktor Prasanna 0001
IEEE Trans. Cloud Comput.1
2014 PLAStiCC: Predictive Look-Ahead Scheduling for Continuous Dataflows on Clouds
abstract
Scalable stream processing and continuous dataflow systems are gaining traction with the rise of big data due to the need for processing high velocity data in near real time. Unlike batch processing systems such as MapReduce and workflows, static scheduling strategies fall short for continuous data flows due to the variations in the input data rates and the need for sustained throughput. The elastic resource provisioning of cloud infrastructure is valuable to meet the changing resource needs of such continuous applications. However, multi-tenant cloud resources introduce yet another dimension of performance variability that impacts the application's throughput. In this paper we propose Plastic, an adaptive scheduling algorithm that balances resource cost and application throughput using a prediction-based look-ahead approach. It not only addresses variations in the input data rates but also the underlying cloud infrastructure. In addition, we also propose several simpler static scheduling heuristics that operate in the absence of accurate performance prediction model. These static and adaptive heuristics are evaluated through extensive simulations using performance traces obtained from Amazon AWS IaaS public cloud. Our results show an improvement of up to 20% in the overall profit as compared to the reactive adaptation algorithm.
Alok Gautam Kumbhare, Yogesh L. Simmhan, Viktor Prasanna 0001
CCGRID1
2014 GoFFish: A Sub-graph Centric Framework for Large-Scale Graph Analytics
Yogesh L. Simmhan, Alok Gautam Kumbhare, Charith Wickramaarachchi, Soonil Nagarkar, Santosh Ravi, Cauligi S. Raghavendra, Viktor Prasanna 0001
Euro-Par2
2013 Exploiting application dynamism and cloud elasticity for continuous dataflows
abstract
Contemporary continuous dataflow systems use elastic scaling on distributed cloud resources to handle variable data rates and to meet applications' needs while attempting to maximize resource utilization. However, virtualized clouds present an added challenge due to the variability in resource performance -- over time and space -- thereby impacting the application's QoS. Elastic use of cloud resources and their allocation to continuous dataflow tasks need to adapt to such infrastructure dynamism. In this paper, we develop the concept of "dynamic dataflows" as an extension to continuous dataflows that utilizes alternate tasks and allows additional control over the dataflow's cost and QoS. We formalize an optimization problem to perform both deployment and runtime cloud resource management for such dataflows, and define an objective function that allows trade-off between the application's value against resource cost. We present two novel heuristics, local and global, based on the variable sized bin packing heuristics to solve this NP-hard problem. We evaluate the heuristics against a static allocation policy for a dataflow with different data rate profiles that is simulated using VM performance traces from a private cloud data center. The results show that the heuristics are effective in intelligently utilizing cloud elasticity to mitigate the effect of both input data rate and cloud resource performance variabilities on QoS.
Alok Gautam Kumbhare, Yogesh L. Simmhan, Viktor Prasanna 0001
SC1
2012 Cryptonite: A Secure and Performant Data Repository on Public Clouds
abstract
Cloud storage has become immensely popular for maintaining synchronized copies of files and for sharing documents with collaborators. However, there is heightened concern about the security and privacy of Cloud-hosted data due to the shared infrastructure model and an implicit trust in the service providers. Emerging needs of secure data storage and sharing for domains like Smart Power Grids, which deal with sensitive consumer data, require the persistence and availability of Cloud storage but with client-controlled security and encryption, low key management overhead, and minimal performance costs. Cryptonite is a secure Cloud storage repository that addresses these requirements using a Strongbox model for shared key management. We describe the Cryptonite service and desktop client, discuss performance optimizations, and provide an empirical analysis of the improvements. Our experiments shows that Cryptonite clients achieve a 40% improvement in file upload bandwidth over plaintext storage using the Azure Storage Client API despite the added security benefits, while our file download performance is 5 times faster than the baseline for files greater than 100MB.
Alok Gautam Kumbhare, Yogesh L. Simmhan, Viktor Prasanna 0001
IEEE CLOUD1
2011 An Analysis of Security and Privacy Issues in Smart Grid Software Architectures on Clouds
abstract
Power utilities globally are increasingly upgrading to Smart Grids that use bi-directional communication with the consumer to enable an information-driven approach to distributed energy management. Clouds offer features well suited for Smart Grid software platforms and applications, such as elastic resources and shared services. However, the security and privacy concerns inherent in an information-rich Smart Grid environment are further exacerbated by their deployment on Clouds. Here, we present an analysis of security and privacy issues in a Smart Grids software architecture operating on different Cloud environments, in the form of a taxonomy. We use the Los Angeles Smart Grid Project that is underway in the largest U.S. municipal utility to drive this analysis that will benefit both Cloud practitioners targeting Smart Grid applications, and Cloud researchers investigating security and privacy.
Yogesh L. Simmhan, Alok Gautam Kumbhare, Baohua Cao, Viktor Prasanna 0001
IEEE CLOUD2