Ganesh Ananthanarayanan

dblp:08/5351 · DBLP profile ↗
← Back
45ranked-venue papers
13as first author
11since 2021 · last 2025
0000-0002-7479-1664ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 25 · 9 first-author · 7 since 2021Systems, architecture and hardware · 11 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 CIS: Checkpointed Inference for Data Drift-Resilient Model Serving at Edge Servers
abstract
Small deep learning models deployed at edge servers suffer from decreasing accuracy due to data drift and hence require continual learning, which leads to resource competition with inference execution, decreasing accuracy and latency service-level-objective (SLO) fulfillment. Previous methods fail to maximize both accuracy and SLO fulfillment. To address this problem, we propose a Checkpointed Inference-based system for accurate and SLO-guaranteed data drift-resilient model Serving (CIS). CIS incorporates checkpointed inference - maximizing accuracy by continuously switching to intermediately retrained models during inference. However, model switching introduces significant inference latency overhead. To mitigate this, first, CIS proposes a lightweight allocation and placement scheduler to minimize switching time. Second, during inference execution, CIS reorders retraining samples to reduce switching frequency. Third, it temporarily reallocates GPU space from retraining tasks to inference tasks to address request queuing issue caused by the switching, with minimal impact on accuracy. Trace-driven experiments show that CIS achieves up to 25.1% higher accuracy without affecting SLO fulfillment, and requires 4× lower GPU cost compared to the existing method in achieving similar or higher accuracy.
Sudipta Saha Shubha, Haiying Shen, Ganesh Ananthanarayanan
SoCC3
2025 METIS: Fast Quality-Aware RAG Systems with Configuration Adaptation
abstract
RAG (Retrieval Augmented Generation) allows LLMs (large language models) to generate better responses with external knowledge, but using more external knowledge causes higher response delay. Prior work focuses either on reducing the response delay (e.g., better scheduling of RAG queries) or on maximizing quality (e.g., tuning the RAG workflow), but they fall short in systematically balancing the tradeoff between the delay and quality of RAG responses. To balance both quality and response delay, this paper presents METIS, the first RAG system that jointly schedules queries and adapts the key RAG configurations of each query, such as the number of retrieved text chunks and synthesis methods. Using four popular RAG-QA datasets, we show that compared to the state-of-the-art RAG optimization schemes, METIS reduces the generation latency by 1.64 – 2.54× without sacrificing generation quality.
Siddhant Ray, Rui Pan 0003, Zhuohan Gu, Kuntai Du, Shaoting Feng, Ganesh Ananthanarayanan, Ravi Netravali, Junchen Jiang
SOSP6
2024 Vulcan: Automatic Query Planning for Live ML Analytics
Yiwen Zhang 0008, Xumiao Zhang, Ganesh Ananthanarayanan, Anand Padmanabha Iyer, Yuanchao Shu, Paramvir Bahl, Z. Morley Mao, Mosharaf Chowdhury
NSDI3
2024 CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
abstract
As large language models (LLMs) take on complex tasks, their inputs are supplemented with longer contexts that incorporate domain knowledge. Yet using long contexts is challenging as nothing can be generated until the whole context is processed by the LLM. While the context-processing delay can be reduced by reusing the KV cache of a context across different inputs, fetching the KV cache, which contains large tensors, over the network can cause high extra network delays.
Yuhan Liu 0004, Hanchen Li, Yihua Cheng, Siddhant Ray, Qizheng Zhang, Kuntai Du, Shan Lu 0001, Ganesh Ananthanarayanan, Michael Maire, Henry Hoffmann, Ari Holtzman, Junchen Jiang
SIGCOMM10
2023 OneAdapt: Fast Adaptation for Deep Learning Applications via Backpropagation
abstract
Deep learning inference on streaming media data, such as object detection in video or LiDAR feeds and text extraction from audio waves, is now ubiquitous. To achieve high inference accuracy, these applications typically require significant network bandwidth to gather high-fidelity data and extensive GPU resources to run deep neural networks (DNNs). While the high demand for network bandwidth and GPU resources could be substantially reduced by optimally adapting the configuration knobs, such as video resolution and frame rate, current adaptation techniques fail to meet three requirements simultaneously: adapt configurations (i) with minimum extra GPU or bandwidth overhead (ii) to reach near-optimal decisions based on how the data affects the final DNN's accuracy, and (iii) do so for a range of configuration knobs. This paper presents OneAdapt, which meets these requirements by leveraging a gradient-ascent strategy to adapt configuration knobs. The key idea is to embrace DNNs' differentiability to quickly estimate the accuracy's gradient to each configuration knob, called AccGrad. Specifically, OneAdapt estimates AccGrad by multiplying two gradients: InputGrad (i.e., how each configuration knob affects the input to the DNN) and DNNGrad (i.e., how the DNN input affects the DNN inference output). We evaluate OneAdapt across five types of configurations, four analytic tasks, and five types of input data. Compared to state-of-the-art adaptation schemes, OneAdapt cuts bandwidth usage and GPU usage by 15-59% while maintaining comparable accuracy or improves accuracy by 1-5% while using equal or fewer resources.
Kuntai Du, Yuhan Liu 0004, Yitian Hao, Qizheng Zhang, Ganesh Ananthanarayanan, Junchen Jiang
SoCC7
2023 RECL: Responsive Resource-Efficient Continuous Learning for Video Analytics
Mehrdad Khani Shirkoohi, Ganesh Ananthanarayanan, Kevin Hsieh, Junchen Jiang, Ravi Netravali, Yuanchao Shu, Mohammad Alizadeh, Paramvir Bahl
NSDI2
2023 Gemel: Model Merging for Memory-Efficient, Real-Time Video Analytics at the Edge
Arthi Padmanabhan, Neil Agarwal, Anand Padmanabha Iyer, Ganesh Ananthanarayanan, Yuanchao Shu, Nikolaos Karianakis, Guoqing Harry Xu, Ravi Netravali
NSDI4
2023 Tambur: Efficient loss recovery for videoconferencing via streaming codes
Michael Rudow, Francis Y. Yan, Ganesh Ananthanarayanan, Martin Ellis, K. V. Rashmi
NSDI4
2022 Ekya: Continuous Learning of Video Analytics Models on Edge Compute Servers
Romil Bhardwaj, Zhengxu Xia, Ganesh Ananthanarayanan, Junchen Jiang, Yuanchao Shu, Nikolaos Karianakis, Kevin Hsieh, Paramvir Bahl, Ion Stoica
NSDI3
2021 Spider: A Multi-Hop Millimeter-Wave Network for Live Video Analytics
Zhuqi Li, Yuanchao Shu, Ganesh Ananthanarayanan, Longfei Shangguan, Kyle Jamieson, Paramvir Bahl
SEC3
2021 PECAM: privacy-enhanced video streaming and analytics via securely-reversible transformation
abstract
As Video Streaming and Analytics (VSA) systems become increasingly popular, serious privacy concerns have risen on exposing too much unnecessary private information to the VSA providers. Yet, it is challenging to protect privacy while still preserving desired VSA features, i.e., effective analytics, forensic support, resource efficiency, and real-time execution. In this paper, we present a VSA privacy enhancement system (PECAM), which addresses the above challenge with no change in the VSA back-end. PECAM leverages a novel Generative Adversarial Network to perform the privacy-enhanced securely-reversible video transformation. PECAM also incorporates a couple of system optimizations into its VSA workflow to reduce network bandwidth usage and enable real-time processing on cameras. We implement our PECAM prototype on commodity hardware and evaluate its performance via both security study and extensive experiments. Results demonstrate that PECAM can effectively enhance the visual privacy of VSA in the presence of an adversary, and its transformed videos, when taken as input for various VSA back-end tasks, maintain a 96% accuracy of corresponding original videos. Additionally, it performs 12.3× and 1.8× better than baseline methods in terms of the computing cost and network bandwidth usage, respectively.
Hao Wu 0067, Xuejin Tian, Minghao Li 0003, Yunxin Liu 0001, Ganesh Ananthanarayanan, Fengyuan Xu, Sheng Zhong 0002
MobiCom5
2020 On the Future of Congestion Control for the Public Internet
abstract
The conventional wisdom requires that all congestion control algorithms deployed on the public Internet be TCP-friendly. If universally obeyed, this requirement would greatly constrain the future of such congestion control algorithms. If partially ignored, as is increasingly likely, then there could be significant inequities in the bandwidth received by different flows. To avoid this dilemma, we propose an alternative to the TCP-friendly paradigm that can accommodate innovation, is consistent with the Internet's current economic model, and is feasible to deploy given current usage trends.
Lloyd Brown, Ganesh Ananthanarayanan, Ethan Katz-Bassett, Arvind Krishnamurthy, Sylvia Ratnasamy, Michael Schapira, Scott Shenker
HotNets2
2020 Collage Inference: Using Coded Redundancy for Lowering Latency Variation in Distributed Image Classification Systems
abstract
MLaaS (ML-as-a-Service) offerings by cloud computing platforms are becoming increasingly popular. Hosting pre-trained machine learning models in the cloud enables elastic scalability as the demand grows. But providing low latency and reducing the latency variance is a key requirement. Variance is harder to control in a cloud deployment due to uncertain-ties in resource allocations across many virtual instances. We propose the collage inference technique, which uses a novel convolutional neural network model, collage-cnn, to provide low-cost redundancy. A collage-cnn model takes a collage image formed by combining multiple images and performs multi-image classification in one shot, albeit at slightly lower accuracy. We augment a collection of traditional single image classifier models with a single collage-cnn classifier, which acts as their low-cost redundant backup. Collage-cnn provides backup classification results if any single image classification requests experience a slowdown. Deploying the collage-cnn models in the cloud, we demonstrate that the 99th percentile tail latency of inference can be reduced by 1.2x to 2x compared to replication-based approaches while providing high accuracy. Variation in inference latency can be reduced by 1.8x to 15x.
Krishna Narra, Zhifeng Lin, Ganesh Ananthanarayanan, Amir Salman Avestimehr, Murali Annavaram
ICDCS3
2020 Spatula: Efficient cross-camera video analytics on large camera networks
abstract
Cameras are deployed at scale with the purpose of searching and tracking objects of interest (e.g., a suspected person) through the camera network on live videos. Such cross-camera analytics is data and compute intensive, whose costs grow with the number of cameras and time. We present Spatula, a cost-efficient system that enables scaling cross-camera analytics on edge compute boxes to large camera networks by leveraging the spatial and temporal cross-camera correlations. While such correlations have been used in computer vision community, Spatula uses them to drastically reduce the communication and computation costs by pruning search space of a query identity (e.g., ignoring frames not correlated with the query identity’s current position). Spatula provides the first system substrate on which cross-camera analytics applications can be built to efficiently harness the cross-camera correlations that are abundant in large camera deployments. Spatula reduces compute load by $8.3\times$ on an 8-camera dataset, and by $23\times-86\times$ on two datasets with hundreds of cameras (simulated from real vehicle/pedestrian traces). We have also implemented Spatula on a testbed of 5 AWS DeepLens cameras.
Samvit Jain, Ganesh Ananthanarayanan, Junchen Jiang, Yuanchao Shu, Paramvir Bahl, Joseph Gonzalez 0001
SEC4
2020 Visor: Privacy-Preserving Video Analytics as a Cloud Service
Rishabh Poddar, Ganesh Ananthanarayanan, Srinath Setty, Stavros Volos, Raluca A. Popa
USENIX Security Symposium2
2019 HotEdgeVideo'19: Workshop on Hot Topics in Video Analytics and Intelligent Edges
abstract
No abstract available.
Ganesh Ananthanarayanan, Yunxin Liu 0001, Yuanchao Shu
MobiCom1
2019 Video Analytics - Killer App for Edge Computing
abstract
The world is witnessing an unprecedented increase in camera deployment. The USA and UK, for instance, have one camera for every 8 people. Video analytics from these cameras are becoming more and more pervasive, exerting important functions on a wide range of verticals including manufacturing, transportation, and retails. While vision techniques have seen considerable advancement, they have come at the expense of compute and network cost.
Ganesh Ananthanarayanan, Paramvir Bahl, Landon P. Cox, Alex Crown, Shadi A. Noghabi, Yuanchao Shu
MobiSys1
2019 Zooming in on wide-area latencies to a global cloud provider
abstract
The network communications between the cloud and the client have become the weak link for global cloud services that aim to provide low latency services to their clients. In this paper, we first characterize WAN latency from the viewpoint of a large cloud provider Azure, whose network edges serve hundreds of billions of TCP connections a day across hundreds of locations worldwide. In particular, we focus on instances of latency degradation and design a tool, BlameIt, that enables cloud operators to localize the cause (i.e., faulty AS) of such degradation. BlameIt uses passive diagnosis, using measurements of existing connections between clients and the cloud locations, to localize the cause to one of cloud, middle, or client segments. Then it invokes selective active probing (within a probing budget) to localize the cause more precisely. We validate BlameIt by comparing its automatic fault localization results with that arrived at by network engineers manually, and observe that BlameIt correctly localized the problem in all the 88 incidents. Further, BlameIt issues 72X fewer active probes than a solution relying on active probing alone, and is deployed in production at Azure.
Sundararajan Renganathan, Ganesh Ananthanarayanan, Junchen Jiang, Venkat N. Padmanabhan, Manuel Schröder, Matt Calder, Arvind Krishnamurthy
SIGCOMM3
2018 Wide-area analytics with multiple resources
abstract
Running data-parallel jobs across geo-distributed sites has emerged as a promising direction due to the growing need for geo-distributed cluster deployment. A key difference between geo-distributed and intra-cluster jobs is the heterogeneous (and often constrained) nature of compute and network resources across the sites. We propose Tetrium, a system for multi-resource allocation in geo-distributed clusters, that jointly considers both compute and network resources for task placement and job scheduling. Tetrium significantly reduces job response time, while incorporating several other performance goals with simple control knobs. Our EC2 deployment and trace-driven simulations suggest that Tetrium improves the average job response time by up to 78% compared to existing data-locality-based solutions, and up to 55% compared to Iridium, the recently proposed geo-distributed analytics system.
Chien-Chun Hung, Ganesh Ananthanarayanan, Leana Golubchik, Minlan Yu, Mingyang Zhang 0005
EuroSys2
2018 Odin: Microsoft's Scalable Fault-Tolerant CDN Measurement System
Matt Calder, Ryan Gao, Manuel Schröder, Ryan Stewart, Jitendra Padhye, Ratul Mahajan, Ganesh Ananthanarayanan, Ethan Katz-Bassett
NSDI7
2018 Focus: Querying Large Video Datasets with Low Latency and Low Cost
Kevin Hsieh, Ganesh Ananthanarayanan, Peter Bodík, Shivaram Venkataraman, Paramvir Bahl, Matthai Philipose, Phillip B. Gibbons, Onur Mutlu
OSDI2
2018 Chameleon: scalable adaptation of video analytics
abstract
Applying deep convolutional neural networks (NN) to video data at scale poses a substantial systems challenge, as improving inference accuracy often requires a prohibitive cost in computational resources. While it is promising to balance resource and accuracy by selecting a suitable NN configuration (e.g., the resolution and frame rate of the input video), one must also address the significant dynamics of the NN configuration's impact on video analytics accuracy. We present Chameleon, a controller that dynamically picks the best configurations for existing NN-based video analytics pipelines. The key challenge in Chameleon is that in theory, adapting configurations frequently can reduce resource consumption with little degradation in accuracy, but searching a large space of configurations periodically incurs an overwhelming resource overhead that negates the gains of adaptation. The insight behind Chameleon is that the underlying characteristics (e.g., the velocity and sizes of objects) that affect the best configuration have enough temporal and spatial correlation to allow the search cost to be amortized over time and across multiple video feeds. For example, using the video feeds of five traffic cameras, we demonstrate that compared to a baseline that picks a single optimal configuration offline, Chameleon can achieve 20-50% higher accuracy with the same amount of resources, or achieve the same accuracy with only 30--50% of the resources (a 2-3X speedup).
Junchen Jiang, Ganesh Ananthanarayanan, Peter Bodík, Siddhartha Sen 0001, Ion Stoica
SIGCOMM2
2018 Deep recurrent neural networks based binaural speech segregation for the selection of closest target of interest
Ganesh Ananthanarayanan
Multim. Tools Appl.2
2017 Live Video Analytics at Scale with Approximation and Delay-Tolerance
Ganesh Ananthanarayanan, Peter Bodík, Matthai Philipose, Paramvir Bahl, Michael J. Freedman
NSDI2
2016 Hold 'em or fold 'em?: aggregation queries under performance variations
abstract
Systems are increasingly required to provide responses to queries, even if not exact, within stringent time deadlines. These systems parallelize computations over many processes and aggregate them hierarchically to get the final response (e.g., search engines and data analytics). Due to large performance variations in clusters, some processes are slower. Therefore, aggregators are faced with the question of how long to wait for outputs from processes before combining and sending them upstream. Longer waits increase the response quality as it would include outputs from more processes. However, it also increases the risk of the aggregator failing to provide its result by the deadline. This leads to all its results being ignored, degrading response quality. Our algorithm, Cedar, proposes a solution to this quandary of deciding wait durations at aggregators. It uses an online algorithm to learn distributions of durations at each level in the hierarchy and collectively optimizes the wait duration. Cedar's solution is theoretically sound, fully distributed, and generically applicable across systems that use aggregation trees since it is agnostic to the causes of performance variations. Evaluation using production latency distributions from Google, Microsoft and Facebook using deployment and simulation shows that Cedar improves average response quality by over 100%.
Gautam Kumar 0001, Ganesh Ananthanarayanan, Sylvia Ratnasamy, Ion Stoica
EuroSys2
2016 Altruistic Scheduling in Multi-Resource Clusters
Robert Grandl, Mosharaf Chowdhury, Aditya Akella, Ganesh Ananthanarayanan
OSDI4
2016 CLARINET: WAN-Aware Optimization for Analytics Queries
Raajay Viswanathan, Ganesh Ananthanarayanan, Aditya Akella
OSDI2
2016 Via: Improving Internet Telephony Call Quality Using Predictive Relay Selection
abstract
Interactive real-time streaming applications such as audio-video conferencing, online gaming and app streaming, place stringent requirements on the network in terms of delay, jitter, and packet loss. Many of these applications inherently involve client-to-client communication, which is particularly challenging since the performance requirements need to be met while traversing the public wide-area network (WAN). This is different from the typical situation of cloud-to-client communication, where the WAN can often be bypassed by moving a communication end-point to a cloud “edge”, close to the client. Can we nevertheless take advantage of cloud resources to improve the performance of real-time client-to-client streaming over the WAN?
Junchen Jiang, Rajdeep Das, Ganesh Ananthanarayanan, Philip A. Chou, Venkat N. Padmanabhan, Vyas Sekar, Esbjorn Dominique, Marcin Goliszewski, Dalibor Kukoleca, Renat Vafin, Hui Zhang 0001
SIGCOMM3
2015 FastLane: making short flows shorter with agile drop notification
abstract
The drive towards richer and more interactive web content places increasingly stringent requirements on datacenter network performance. Applications running atop these networks typically partition an incoming query into multiple subqueries, and generate the final result by aggregating the responses for these subqueries. As a result, a large fraction --- as high as 80% --- of the network flows in such workloads are short and latency-sensitive. The speed with which existing networks respond to packet drops limits their ability to meet high-percentile flow completion time SLOs. Indirect notifications indicating packet drops (e.g., duplicates in an end-to-end acknowledgement sequence) are an important limitation to the agility of response to packet drops.
David Zats, Anand Padmanabha Iyer, Ganesh Ananthanarayanan, Rachit Agarwal 0001, Randy H. Katz, Ion Stoica, Amin Vahdat
SoCC3
2015 Low Latency Geo-distributed Data Analytics
abstract
Low latency analytics on geographically distributed datasets (across datacenters, edge clusters) is an upcoming and increasingly important challenge. The dominant approach of aggregating all the data to a single datacenter significantly inflates the timeliness of analytics. At the same time, running queries over geo-distributed inputs using the current intra-DC analytics frameworks also leads to high query response times because these frameworks cannot cope with the relatively low and variable capacity of WAN links. We present Iridium, a system for low latency geo-distributed analytics. Iridium achieves low query response times by optimizing placement of both data and tasks of the queries. The joint data and task placement optimization, however, is intractable. Therefore, Iridium uses an online heuristic to redistribute datasets among the sites prior to queries' arrivals, and places the tasks to reduce network bottlenecks during the query's execution. Finally, it also contains a knob to budget WAN usage. Evaluation across eight worldwide EC2 regions using production queries show that Iridium speeds up queries by 3× -- 19× and lowers WAN usage by 15% -- 64% compared to existing baselines.
Qifan Pu, Ganesh Ananthanarayanan, Peter Bodík, Srikanth Kandula, Aditya Akella, Paramvir Bahl, Ion Stoica
SIGCOMM2
2015 Hopper: Decentralized Speculation-aware Cluster Scheduling at Scale
abstract
As clusters continue to grow in size and complexity, providing scalable and predictable performance is an increasingly important challenge. A crucial roadblock to achieving predictable performance is stragglers, i.e., tasks that take significantly longer than expected to run. At this point, speculative execution has been widely adopted to mitigate the impact of stragglers. However, speculation mechanisms are designed and operated independently of job scheduling when, in fact, scheduling a speculative copy of a task has a direct impact on the resources available for other jobs. In this work, we present Hopper, a job scheduler that is speculation-aware, i.e., that integrates the tradeoffs associated with speculation into job scheduling decisions. We implement both centralized and decentralized prototypes of the Hopper scheduler and show that 50% (66%) improvements over state-of-the-art centralized (decentralized) schedulers and speculation strategies can be achieved through the coordination of scheduling and speculation.
Xiaoqi Ren, Ganesh Ananthanarayanan, Adam Wierman, Minlan Yu
SIGCOMM2
2014 Wrangler: Predictable and Faster Jobs using Fewer Resources
abstract
Straggler tasks continue to be a major hurdle in achieving faster completion of data intensive applications running on modern data-processing frameworks. Existing straggler mitigation techniques are inefficient due to their reactive and replicative nature -- they rely on a wait-speculate-re-execute mechanism, thus leading to delayed straggler detection and inefficient resource utilization. Existing proactive techniques also over-utilize resources due to replication. Existing modeling-based approaches are hard to rely on for production-level adoption due to modeling errors. We present Wrangler, a system that proactively avoids situations that cause stragglers. Wrangler automatically learns to predict such situations using a statistical learning technique based on cluster resource utilization counters. Furthermore, Wrangler introduces a notion of a confidence measure with these predictions to overcome the modeling error problems; this confidence measure is then exploited to achieve a reliable task scheduling. In particular, by using these predictions to balance delay in task scheduling against the potential for idling of resources, Wrangler achieves a speed up in the overall job completion time. For production-level workloads from Facebook and Cloudera's customers, Wrangler improves the 99th percentile job completion time by up to 61% as compared to speculative execution, a widely used straggler mitigation technique. Moreover, Wrangler achieves this speed-up while significantly improving the resource consumption (by up to 55%).
Neeraja J. Yadwadkar, Ganesh Ananthanarayanan, Randy H. Katz
SoCC2
2014 GRASS: Trimming Stragglers in Approximation Analytics
Ganesh Ananthanarayanan, Chien-Chun Hung, Xiaoqi Ren, Ion Stoica, Adam Wierman, Minlan Yu
NSDI1
2014 The Power of Choice in Data-Aware Cluster Scheduling
Shivaram Venkataraman, Aurojit Panda, Ganesh Ananthanarayanan, Michael J. Franklin, Ion Stoica
OSDI3
2014 Multi-resource packing for cluster schedulers
abstract
Tasks in modern data parallel clusters have highly diverse resource requirements, along CPU, memory, disk and network. Any of these resources may become bottlenecks and hence, the likelihood of wasting resources due to fragmentation is now larger. Today's schedulers do not explicitly reduce fragmentation. Worse, since they only allocate cores and memory, the resources that they ignore (disk and network) can be over-allocated leading to interference, failures and hogging of cores or memory that could have been used by other tasks. We present Tetris, a cluster scheduler that packs, i.e., matches multi-resource task requirements with resource availabilities of machines so as to increase cluster efficiency (makespan). Further, Tetris uses an analog of shortest-running-time-first to trade-off cluster efficiency for speeding up individual jobs. Tetris' packing heuristics seamlessly work alongside a large class of fairness policies. Trace-driven simulations and deployment of our prototype on a 250 node cluster shows median gains of 30% in job completion time while achieving nearly perfect fairness.
Robert Grandl, Ganesh Ananthanarayanan, Srikanth Kandula, Sriram Rao, Aditya Akella
SIGCOMM2
2013 Effective Straggler Mitigation: Attack of the Clones
Ganesh Ananthanarayanan, Ali Ghodsi 0002, Scott Shenker, Ion Stoica
NSDI1
2012 True elasticity in multi-tenant data-intensive compute clusters
abstract
Data-intensive computing (DISC) frameworks scale by partitioning a job across a set of fault-tolerant tasks, then diffusing those tasks across large clusters. Multi-tenanted clusters must accommodate service-level objectives (SLO) in their resource model, often expressed as a maximum latency for allocating the desired set of resources to every job. When jobs are partitioned into tasks statically, a cluster cannot meet its SLOs while maintaining both high utilization and efficiency. Ideally, we want to give resources to jobs when they are free but would expect to reclaim them instantaneously when new jobs arrive, without losing work. DISC frameworks do not support such elasticity because interrupting running tasks incurs high overheads. Amoeba enables lightweight elasticity in DISC frameworks by identifying points at which running tasks of over-provisioned jobs can be safely exited, committing their outputs, and spawning new tasks for the remaining work. Effectively, tasks of DISC jobs are now sized dynamically in response to global resource scarcity or abundance. Simulation and deployment of our prototype shows that Amoeba speeds up jobs by 32% without compromising utilization or efficiency.
Ganesh Ananthanarayanan, Chris Douglas, Raghu Ramakrishnan 0001, Sriram Rao, Ion Stoica
SoCC1
2012 PACMan: Coordinated Memory Caching for Parallel Jobs
Ganesh Ananthanarayanan, Ali Ghodsi 0002, Andy Warfield, Dhruba Borthakur, Srikanth Kandula, Scott Shenker, Ion Stoica
NSDI1
2011 Scarlett: coping with skewed content popularity in mapreduce clusters
abstract
To improve data availability and resilience MapReduce frameworks use file systems that replicate data uniformly. However, analysis of job logs from a large production cluster shows wide disparity in data popularity. Machines and racks storing popular content become bottlenecks; thereby increasing the completion times of jobs accessing this data even when there are machines with spare cycles in the cluster. To address this problem, we present Scarlett, a system that replicates blocks based on their popularity. By accurately predicting file popularity and working within hard bounds on additional storage, Scarlett causes minimal interference to running jobs. Trace driven simulations and experiments in two popular MapReduce frameworks (Hadoop, Dryad) show that Scarlett effectively alleviates hotspots and can speed up jobs by 20.2%.
Ganesh Ananthanarayanan, Sameer Agarwal 0002, Srikanth Kandula, Albert G. Greenberg, Ion Stoica, Duke Harlan
EuroSys1
2011 Disk-Locality in Datacenter Computing Considered Irrelevant
Ganesh Ananthanarayanan, Ali Ghodsi 0002, Scott Shenker, Ion Stoica
HotOS1
2010 Reining in the Outliers in Map-Reduce Clusters using Mantri
Ganesh Ananthanarayanan, Srikanth Kandula, Albert G. Greenberg, Ion Stoica, Yi Lu 0001, Bikas Saha
OSDI1
2009 StarTrack: a framework for enabling track-based applications
abstract
Mobile devices are increasingly equipped with hardware and software services allowing them to determine their locations, but support for building location-aware applications remains rudimentary. This paper proposes tracks of location coordinates as a high-level abstraction for a new class of mobile applications including ride sharing, location-based collaboration, and health monitoring. Each track is a sequence of entries recording a person's time, location, and application-specific data. StarTrack provides applications with a comprehensive set of operations for recording, comparing, clustering and querying tracks. StarTrack can efficiently operate on thousands of tracks.
Ganesh Ananthanarayanan, Maya Haridasan, Iqbal Mohomed, Douglas B. Terry, Chandramohan A. Thekkath
MobiSys1
2009 Blue-Fi: enhancing Wi-Fi performance using bluetooth signals
abstract
Mobile devices are increasingly equipped with multiple network interfaces with complementary characteristics. In particular, the Wi-Fi interface has high throughput and transfer power efficiency, but its idle power consumption is prohibitive. In this paper we present, Blue-Fi, a sytem that predicts the availability of the Wi-Fi connectivity by using a combination of bluetooth contact-patterns and cell-tower information. This allows the device to intelligently switch the Wi-Fi interface on only when there is Wi-Fi connectivity available, thus avoiding the long periods in idle state and significantly reducing the the number of scans for discovery.
Ganesh Ananthanarayanan, Ion Stoica
MobiSys1
2007 COMBINE: leveraging the power of wireless peers through collaborative downloading
abstract
Mobile devices are increasingly equipped with multiple network interfaces: Wireless Local Area Network (WLAN) interfaces for local connectivity and Wireless Wide Area Network (WWAN) interfaces for wide-area connectivity. The WWAN typically provides much wider coverage but much lower speeds than the WLAN. To address this dichotomy, we present COMBINE, a system for collaborative downloading wherein devices that are within WLAN range pool together their WWAN links, significantly increasing the effective speed available to them.
Ganesh Ananthanarayanan, Venkat N. Padmanabhan, Lenin Ravindranath, Chandramohan A. Thekkath
MobiSys1
2006 SPACE: Secure Protocol for Address Book based Connection Establishment
Ganesh Ananthanarayanan, Ramarathnam Venkatesan, Prasad Naldurg, Sean Olin Blagsvedt, Adithya Hemakumar
HotNets1