David Irwin 0001

dblp:28/760-1 · also David E. Irwin 0001, David Emory Irwin · DBLP profile ↗
← Back
80ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0003-1722-4927ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 38 · 2 first-author · 21 since 2021Computer networks · 13 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 9 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Databases, data management, data science and information retrieval · 4Graphics, computer vision, multimedia, augmented reality and games · 3
YearPublicationVenuePosition
2026 To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
abstract
Computational offloading is a promising approach to overcome client device resource constraints by moving application computations to remote servers. With the advent of specialized hardware accelerators, client devices can now perform fast local processing of tasks such as machine learning inference, reducing the need for offloading. However, edge servers with accelerators also offer faster offloading performance than was previously possible. In this paper, we present an analytic and experimental comparison of on-device processing and edge offloading across accelerator, multi-tenant, and workload scenarios to understand when to use local processing versus offloading. We present models that leverage queuing theory to derive explainable closed-form equations for end-to-end latencies, yielding quantitative performance crossover predictions to guide adaptive offloading. We validate our models across various settings and show that they achieve a mean absolute percentage error of 2.2% compared to observed latencies. We further use these models to develop a resource manager for adaptive offloading and demonstrate its effectiveness in dynamic multi-tenant edge environments.
Nathan Ng 0002, David Irwin 0001, Ananthram Swami, Don Towsley, Prashant J. Shenoy
ICPE2
2026 CarbonShare: Carbon-Fair Allocation for Shared Clusters
abstract
Computing's energy demand, and thus its carbon emissions, are rapidly accelerating with the emergence of a wide range of useful, but computationally-intensive, AI-driven applications. At the same time, computing, along with the rest of society, must rapidly reduce its emissions to avoid the worst consequences of climate change, e.g., population displacement, agricultural collapse, mass extinction, extreme weather, etc. To do so, computing and other industries will eventually need to limit their carbon emissions. Such a limit imposes a new allocation problem: how should datacenters allocate their limited carbon emissions across multiple applications?
John Thiede, David Irwin 0001, Prashant J. Shenoy
ICPE2
2026 Untangling the Carbon-Cost Tradeoffs and Stampede Effect Challenges in Cloud Computing
abstract
As society explores new computing applications, data centers have seen a sharp rise in capacity and energy use, driving up data centers’ carbon footprint and raising concerns about their sustainability. In response, efforts to reduce emissions now complement the traditional goals of cutting energy costs and boosting performance. Given the spatiotemporal variability of the grid carbon intensity, researchers are increasingly adopting workload shifting strategies to minimize the operational carbon footprint of cloud workloads. However, shifting from cost- and performance-centric operations to carbon-aware operations introduces performance penalties and cost overheads. Moreover, large-scale spatial and temporal shifts can cause stampede effects, with synchronized migration to the same low-carbon regions or times, thereby straining resources and creating imbalances. In this paper, we quantify the carbon-cost-performance trade-offs of carbon-aware scheduling and evaluate the risk of such stampede effects. We provide real-world examples illustrating these tradeoffs for application and cloud providers. Lastly, we offer practical insights to help mitigate these challenges and guide sustainable scheduling decisions.
Walid A. Hanafy, Thanathorn Sukprasert, Abel Souza, David Irwin 0001, Prashant J. Shenoy
IEEE Trans. Computers4
2025 PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
abstract
The exponential growth of large-scale AI models has led to computational and power demands that can exceed the capacity of a single data center. This is due to the limited power supplied by regional grids that leads to limited regional computational power. Consequently, distributing training workloads across geographically distributed sites has become essential. However, this approach introduces a significant challenge in the form of communication overhead, creating a fundamental trade-off between the performance gains from accessing greater aggregate power and the performance losses from increased network latency. Although prior work has focused on reducing communication volume or using heuristics for distribution, these methods assume constant homogeneous power supplies and ignore the challenge of heterogeneous power availability between sites.
Talha Mehboob, Luanzheng Guo, Nathan R. Tallent, Michael Zink, David Irwin 0001
SoCC5
2025 FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
abstract
Model serving systems have become popular for deploying deep learning models for various latency-sensitive inference tasks. While traditional replication-based methods have been used for failure-resilient model serving in the cloud, such methods are often infeasible in edge environments due to significant resource constraints that preclude full replication. To address this problem, this paper presents FailLite, a failure-resilient model serving system that employs (i) a heterogeneous replication where the failover model is a smaller variant of the original one, (ii) an intelligent approach that uses warm replicas to ensure quick failover for critical applications while using cold replicas, and (iii) progressive failover to provide low mean time to recovery (MTTR) for the remaining applications. We implement a full prototype of our system and demonstrate its efficacy on an experimental edge testbed and large-scale simulations. Our results using 27 models show that FailLite can recover all failed applications with 2× lower MTTR and only a 0.6% reduction in accuracy. Under extreme failure scenarios, where 50% of edge sites fail simultaneously, FailLite improves recovery rate by at least 39.3% compared to the baseline methods.
Walid A. Hanafy, Tarek F. Abdelzaher, David Irwin 0001, Jesse Milzman, Prashant J. Shenoy
SoCC4
2025 CarbonEdge: Leveraging Mesoscale Spatial Carbon-Intensity Variations for Low Carbon Edge Computing
abstract
The proliferation of latency-critical and compute-intensive edge applications is driving increases in computing demand and carbon emissions at the edge. To better understand carbon emissions at the edge, we analyze granular carbon intensity traces at intermediate "mesoscales," such as within a single US state or among neighboring countries in Europe, and observe significant variations in carbon intensity at these spatial scales. Importantly, our analysis shows that carbon intensity variations, which are known to occur at large continental scales (e.g., cloud regions), also occur at much finer spatial scales, making it feasible to exploit geographic workload shifting in the edge computing context. Motivated by these findings, we propose CarbonEdge, a carbon-aware framework for edge computing that optimizes the placement of edge workloads across mesoscale edge data centers to reduce carbon emissions while meeting latency SLOs. We implement CarbonEdge and evaluate it on a real edge computing testbed and through large-scale simulations for multiple edge workloads and settings. Our experimental results on a real testbed demonstrate that CarbonEdge can reduce emissions by up to 78.7% for a regional edge deployment in central Europe. Moreover, our CDN-scale experiments show potential savings of 49.5% and 67.8% in the US and Europe, respectively, while limiting the one-way latency increase to less than 5.5 ms.
Walid A. Hanafy, Abel Souza, Jan Harkes, David Irwin 0001, Mahadev Satyanarayanan, Prashant J. Shenoy
HPDC6
2025 Ahead of the Curve: Leveraging Periodicity to Improve Job Placement in Data Centers
Xiaoding Guan, Noman Bashir, David Irwin 0001, Prashant J. Shenoy
IC2E3
2025 EcoLearn: Optimizing the Carbon Footprint of Federated Learning
abstract
Federated Learning (FL) distributes machine learning (ML) training across edge devices to reduce data transfer overhead and protect data privacy. Since FL model training may span hundreds of devices and is thus resource- and energy-intensive, it has a significant carbon footprint. Importantly, since energy's carbon-intensity differs substantially (by up to 60×) across locations, training on the same device using the same amount of energy, but at different locations, can incur widely different carbon emissions. While prior work has focused on improving FL's resource- and energy-efficiency by optimizing time-to-accuracy, it implicitly assumes all energy has the same carbon intensity and thus does not optimize carbon efficiency, i.e., work done per unit of carbon emitted.
Talha Mehboob, Noman Bashir, Jesus Omaña Iglesias, Michael Zink, David Irwin 0001
SEC5
2025 LLM-Driven Auto Configuration for Transient IoT Device Collaboration
abstract
Today's Internet of Things (IoT) has evolved from simple sensing and actuation devices to those with embedded processing and intelligent services, enabling rich collaborations between users and their devices. However, enabling such collaboration becomes challenging when transient devices need to interact with host devices in temporarily visited environments. In such cases, fine-grained access control policies are necessary to ensure secure interactions; however, manually implementing them is often impractical for non-expert users. Moreover, at run-time, the system must automatically configure the devices and enforce such fine-grained access control rules. Additionally, the system must address the heterogeneity of devices.
Hetvi Shastri, Walid A. Hanafy, David Irwin 0001, Mani Srivastava 0001, Prashant J. Shenoy
SEC4
2025 Poster Abstract: Rethinking Collaboration Among Mobile Devices in IoT Environments
abstract
Many emerging IoT devices are mobile, enabling them to visit new environments and networks beyond their home networks. Mobile devices often have to interact and collaborate with users and their devices, which belong to the different administrative environments they are temporarily visiting. In this paper, we envision a system for seamless collaboration among transient devices in IoT environments. The system is based on zero-conf collaboration and allows for fine-grained access control. Our proposed design supports hardware-independent interfaces and supports a large number of devices.
Hetvi Shastri, Walid A. Hanafy, David Irwin 0001, Mani Srivastava 0001, Prashant J. Shenoy
SenSys4
2024 Going Green for Less Green: Optimizing the Cost of Reducing Cloud Carbon Emissions
abstract
The continued exponential growth of cloud datacenter capacity has increased awareness of the carbon emissions when executing large compute-intensive workloads. To reduce carbon emissions, cloud users often temporally shift their batch workloads to periods with low carbon intensity. While such time shifting can increase job completion times due to their delayed execution, the cost savings from cloud purchase options, such as reserved instances, also decrease when users operate in a carbon-aware manner. This happens because carbon-aware adjustments change the demand pattern by periodically leaving resources idle, which creates a trade-off between carbon emissions and cost. In this paper, we present GAIA, a carbon-aware scheduler that enables users to address the three-way trade-off between carbon, performance, and cost in cloud-based batch schedulers. Our results quantify the carbon-performance-cost trade-off in cloud platforms and show that compared to existing carbon-aware scheduling policies, our proposed policies can double the amount of carbon savings per percentage increase in cost, while decreasing the performance overhead by 26%.
Walid A. Hanafy, Qianlin Liang, Noman Bashir, Abel Souza, David Irwin 0001, Prashant J. Shenoy
ASPLOS (3)5
2024 SLO-Power: SLO and Power-aware Elastic Scaling for Web Services
abstract
Managing the performance of online web services in cloud data centers while optimizing resource allocation and power consumption is a multifaceted challenge. Often, resource and power management techniques, such as elastic scaling and power capping, are handled independently, leading to conflicts and sub-optimal power-performance trade-offs. To tackle this issue, we introduce SLO-Power, a system that coordinates the resource and power scaling techniques to achieve power savings while adhering to service level objectives (SLOs), such as tail latency constraints. Our approach employs a combination of analytic queuing models and feedback-driven techniques to jointly allocate resources and power to cloud applications in an SLO and power-aware manner. We implement a prototype of our system and evaluate it using realistic workloads to demonstrate its ability to harmonize elastic and power scaling, enabling enhanced resource utilization and reduced power consumption while ensuring the application performance. Our findings indicate that SLO-Power achieves exceptional power and resource efficiency, approaching near-optimal power-efficiency levels at 90%, all while preventing SLO violations. Furthermore, compared to state-of-the-art solutions, SLO-Power demonstrates lower P95 latency, accompanied by a 12% reduction in resource usage.
Mehmet Savasci, Abel Souza, David Irwin 0001, Ahmed Ali-Eldin, Prashant J. Shenoy
CCGrid4
2024 The Sunk Carbon Fallacy: Rethinking Carbon Footprint Metrics for Effective Carbon-Aware Scheduling
abstract
The rapid increase in computing demand and corresponding energy consumption have focused attention on computing's impact on the climate and sustainability. Prior work proposes metrics that quantify computing's carbon footprint across several lifecycle phases, including its supply chain, operation, and end-of-life. Industry uses these metrics to optimize the carbon footprint of manufacturing hardware and running computing applications. Unfortunately, prior work on optimizing datacenters' carbon footprint often succumbs to the sunk cost fallacy by considering embodied carbon emissions (a sunk cost) when making operational decisions (i.e., job scheduling and placement), which leads to operational decisions that do not always reduce the total carbon footprint.
Noman Bashir, Varun Gohil, Anagha Belavadi Subramanya, Mohammad Shahrad, David Irwin 0001, Elsa Olivetti, Christina Delimitrou
SoCC5
2024 CDN-Shifter: Leveraging Spatial Workload Shifting to Decarbonize Content Delivery Networks
abstract
Content Delivery Networks (CDNs) are Internet-scale systems that deliver streaming and web content to users from many geographically distributed edge data centers. Since large CDNs can comprise hundreds of thousands of servers deployed in thousands of global data centers, they can consume a large amount of energy for their operations and thus are responsible for large amounts of Green House Gas (GHG) emissions. As these networks scale to cope with increased demand for bandwidth-intensive content, their emissions are expected to rise further, making sustainable design and operation an important goal for the future. Since different geographic regions vary in the carbon intensity and cost of their electricity supply, in this paper, we consider spatial shifting as a key technique to jointly optimize the carbon emissions and energy costs of a CDN. We present two forms of shifting: spatial load shifting, which operates within the time scale of minutes, and VM capacity shifting, which operates at a coarse time scale of days or weeks. The proposed techniques jointly reduce carbon and electricity costs while considering the performance impact of increased request latency from such optimizations. Using real-world traces from a large CDN and carbon intensity and energy prices data from electric grids in different regions, we show that increasing the latency by 60ms can reduce carbon emissions by up to 35.5%, 78.6%, and 61.7% across the US, Europe, and worldwide, respectively. In addition, we show that capacity shifting can increase carbon savings by up to 61.2%. Finally, we analyze the benefits of spatial shifting and show that it increases carbon savings from added solar energy by 68% and 130% in the US and Europe, respectively.
Jorge Murillo, Walid A. Hanafy, David Irwin 0001, Ramesh K. Sitaraman, Prashant J. Shenoy
SoCC3
2024 TailClipper: Reducing Tail Response Time of Distributed Services Through System-Wide Scheduling
abstract
Reducing tail latency has become a crucial issue for optimizing the performance of online cloud services and distributed applications. In distributed applications, there are many causes of high end-to-end tail latency, including operating system delays, request re-ordering due to fan-out/fanin, and network congestion. Although recent research has focused on reducing tail latency for individual application components, such as by replicating requests and scheduling, in this paper, we argue for a holistic approach for reducing the end-to-end tail latency across application components. We propose TailClipper, a distributed scheduler that tags each arriving request with an arrival timestamp, and propagates it across the microservices' call chain. TailClipper then uses arrival timestamps to implement an oldest request first scheduler that combines global first-come first serve with a limited form of processor sharing to reduce end-to-end tail latency. In doing so, TailClipper can counter the performance degradation caused by request reordering in multi-tiered and microservices-based applications. We implement TailClipper as a userspace Linux scheduler and evaluate it using cloud workload traces and a real-world microservices application. Compared to state-of-the-art schedulers, our experiments reveal that TailClipper improves the 99th percentile response time by up to 81%, while also improving the mean response time and the system throughput by up to 54% and 29% respectively under high loads.
Nathan Ng 0002, Abel Souza, Ahmed Ali-Eldin, David Irwin 0001, Don Towsley, Prashant J. Shenoy
SoCC4
2024 On the Limitations of Carbon-Aware Temporal and Spatial Workload Shifting in the Cloud
abstract
Cloud platforms have been focusing on reducing their carbon emissions by shifting workloads across time and locations to when and where low-carbon energy is available. Despite the prominence of this idea, prior work has only quantified the potential of spatiotemporal workload shifting in narrow settings, i.e., for specific workloads in select regions. In particular, there has been limited work on quantifying an upper bound on the ideal and practical benefits of carbon-aware spatiotemporal workload shifting for a wide range of cloud workloads. To address the problem, we conduct a detailed data-driven analysis to understand the benefits and limitations of carbon-aware spatiotemporal scheduling for cloud workloads. We utilize carbon intensity data from 123 regions, encompassing most major cloud sites, to analyze two broad classes of workloads---batch and interactive---and their various characteristics, e.g., job duration, deadlines, and SLOs. Our findings show that while spatiotemporal workload shifting can reduce workloads' carbon emissions, the practical upper bounds of these carbon reductions are currently limited and far from ideal. We also show that simple scheduling policies often yield most of these reductions, with more sophisticated techniques yielding little additional benefit. Notably, we also find that the benefit of carbon-aware workload scheduling relative to carbon-agnostic scheduling will decrease as the energy supply becomes "greener."
Thanathorn Sukprasert, Abel Souza, Noman Bashir, David Irwin 0001, Prashant J. Shenoy
EuroSys4
2024 INVAR: Inversion Aware Resource Provisioning and Workload Scheduling for Edge Computing
abstract
Edge computing is emerging as a complementary architecture to cloud computing to address some of its associated issues. One of the major advantages of edge computing is that edge data centers are usually much closer to users compared to traditional cloud data centers. Therefore, it is commonly believed that for developers of latency-sensitive applications, they can effectively reduce the overall end-to-end latency by simply transitioning from a cloud deployment to an edge deployment. However, as recent work has shown, the performance of an edge deployment is vulnerable to a couple of factors which under many practical scenarios can lead to edge servers providing worse end-to-end response time than cloud servers. This phenomenon is referred to as edge performance inversion. In this paper, we propose resource allocation and workload scheduling algorithms that actively prevent edge performance inversion. Our algorithms, named INVAR, are based on queueing theory results and optimization techniques. Evaluation results show that INVAR can find a near-optimal solution that outperforms the performance of a cloud deployment by an adjustable margin. Simulation results based on production workloads from Akamai data centers show that INVAR can outperform common heuristic-based edge deployment by 11% to 24% in real-world scenarios.
David Irwin 0001, Prashant J. Shenoy, Don Towsley
INFOCOM2
2023 Ecovisor: A Virtual Energy System for Carbon-Efficient Applications
abstract
Cloud platforms' rapid growth is raising significant concerns about their carbon emissions. To reduce carbon emissions, future cloud platforms will need to increase their reliance on renewable energy sources, such as solar and wind, which have zero emissions but are highly unreliable. Unfortunately, today's energy systems effectively mask this unreliability in hardware, which prevents applications from optimizing their carbon-efficiency, or work done per kilogram of carbon emitted. To address the problem, we design an "ecovisor", which virtualizes the energy system and exposes software-defined control of it to applications. An ecovisor enables each application to handle clean energy's unreliability in software based on its own specific requirements. We implement a small-scale ecovisor prototype that virtualizes a physical energy system to enable software-based application-level i) visibility into variable grid carbon-intensity and local renewable generation and ii) control of server power usage and battery charging and discharging. We evaluate the ecovisor approach by showing how multiple applications can concurrently exercise their virtual energy system in different ways to better optimize carbon-efficiency based on their specific requirements compared to general system-wide policies.
Abel Souza, Noman Bashir, Jorge Murillo, Walid A. Hanafy, Qianlin Liang, David Irwin 0001, Prashant J. Shenoy
ASPLOS (2)6
2023 Carbon Containers: A System-level Facility for Managing Application-level Carbon Emissions
abstract
To reduce their environmental impact, cloud datacenters' are increasingly focused on optimizing applications' carbon-efficiency, or work done per mass of carbon emitted. To facilitate such optimizations, we present Carbon Containers, a simple system-level facility, which extends prior work on power containers, that automatically regulates applications' carbon emissions in response to variations in both their work-load's intensity and their energy's carbon-intensity. Specifically, Carbon Containers enable applications to specify a maximum carbon emissions rate (in g.CO2e/hr), and then transparently enforce this rate via a combination of vertical scaling, container migration, and suspend/resume while maximizing either energy-efficiency or performance.
John Thiede, Noman Bashir, David Irwin 0001, Prashant J. Shenoy
SoCC3
2023 Energy Time Fairness: Balancing Fair Allocation of Energy and Time for GPU Workloads
abstract
Traditionally, multi-tenant cloud and edge platforms use fair-share schedulers to fairly multiplex resources across applications. These schedulers ensure applications receive processing time proportional to a configurable share of the total time. Unfortunately, enforcing time-fairness across applications often violates energy-fairness, such that some applications consume more than their fair share of energy. This occurs because applications either do not fully utilize their resources or operate at a reduced frequency/voltage during their time-slice. The problem is particularly acute for machine learning (ML) applications using GPUs, where model size largely dictates utilization and energy usage. Enforcing energy-fairness is also important since energy is a costly and limited resource. For example, in cloud platforms, energy dominates operating costs and is limited by the power delivery infrastructure, while in edge platforms, energy is often scarce and limited by energy harvesting and battery constraints.
Qianlin Liang, Walid A. Hanafy, Noman Bashir, David Irwin 0001, Prashant J. Shenoy
SEC4
2023 Is Sharing Caring? Analyzing the Incentives for Shared Cloud Clusters
abstract
Many organizations maintain and operate large shared computing clusters, since they can substantially reduce computing costs by leveraging statistical multiplexing to amortize it across all users. Importantly, such shared clusters are generally not free to use, but have an internal pricing model that funds their operation. Since employees at many large organizations, especially Universities, have some budgetary autonomy over purchase decisions, internal shared clusters are increasingly competing for users with cloud platforms, which may offer lower costs and better performance. As a result, many organizations are shifting their shared clusters to operate on cloud resources. This paper empirically analyzes the user incentives for shared cloud clusters under two different pricing models using an 8-year job trace from a large shared cluster for a large University system.
Talha Mehboob, Noman Bashir, Michael Zink, David Irwin 0001
ICPE4
2023 WattScope: Non-intrusive application-level power disaggregation in datacenters
Xiaoding Guan, Noman Bashir, David Irwin 0001, Prashant J. Shenoy
Perform. Evaluation3
2022 PeakTK: An Open Source Toolkit for Peak Forecasting in Energy Systems
abstract
As the electric grid undergoes the transition to a carbon free future, many new techniques for optimizing the grid’s energy usage and carbon footprint are being designed. A common technique used by many approaches is to reduce the energy usage of the grid’s peak demand periods since doing so is beneficial for reducing the carbon usage of the grid. Consequently, the design of peak forecasting methods that predict when and how much peak demand will be seen is at the heart of many energy optimization approaches. In this paper, we present PeakTK, an open-source toolkit and reference datasets for peak forecasting in energy systems. PeakTK implements a range of peak forecasting methods that have been proposed recently and exposes them through well-defined interfaces and library modules. Our goal is to improve reproducibility of energy systems research by providing a common framework for evaluating and comparing new peak forecasting algorithms. Further, PeakTK provides libraries to enable researchers and practitioners to easily incorporate peak forecasting methods into their research when implementing higher level grid optimizations. We discuss the design and implementation of PeakTK and present case studies to demonstrate how PeakTK can be used for forecasting or quantitative comparisons of energy optimization methods.
Phuthipong Bovornkeeratiroj, John Wamburu, David Irwin 0001, Prashant J. Shenoy
COMPASS3
2022 Guest Editorial: Special Section on Intersection of Computing and Communication Technologies With Energy Systems
abstract
The papers in this special section focus on the intersection of computing and communication technologies with energy systems. Computing and communication technologies impact energy systems in two distinct ways. The exponential growth of these technologies has made them large energy consumers. Therefore, new architectures, technologies and systems are being developed and deployed to make computing and networked systems more energy efficient. Additionally, these technologies will play a central role in the ongoing transformation of our energy systems. They help measure, monitor and control energy resources, inform and shape human demand, and determine how utilities, generators, regulators, and consumers interact. Recently, there have been vibrant developments in the research community at the intersection of computing and communication technologies with energy systems. Diverse applications of computing and networked systems have made legacy systems more energy-efficient, as well as improved the design, analysis, and development of innovative new energy systems.
Sid Chi-Kin Chau, David Irwin 0001, Minghua Chen 0001, Gopal Ramchurn
IEEE Trans. Sustain. Comput.2
2021 Good Things Come to Those Who Wait: Optimizing Job Waiting in the Cloud
abstract
Cloud-enabled schedulers execute jobs on either fixed resources or those acquired on demand from cloud platforms. Thus, these schedulers must define not only a scheduling policy, which selects which jobs run when fixed resources become available, but also a waiting policy, which selects which jobs wait for fixed resources when they are not available, rather than run on on-demand resources. As with scheduling policies, optimizing waiting policies requires a priori knowledge of job runtime. Unfortunately, prior work has shown that accurately predicting job runtime is challenging. In this paper, we show that optimizing job waiting in the cloud is possible without accurate job runtime predictions. To do so, we i) speculatively execute jobs on on-demand resources for a small time and cost to learn more about job runtime, and ii) develop a ML model to predict wait time from cluster state, which is more accurate and has less overhead than prior approaches that use job runtime predictions. We evaluate our approach on a year-long batch workload consisting of 14 million jobs, and show that it yields a cost and average wait time within 4% and 13%, respectively, of the optimal.
Lurdh Pradeep Reddy Ambati, Noman Bashir, David Irwin 0001, Prashant J. Shenoy
SoCC3
2021 Enabling Sustainable Clouds: The Case for Virtualizing the Energy System
abstract
Cloud platforms' growing energy demand and carbon emissions are raising concern about their environmental sustainability. The current approach to enabling sustainable clouds focuses on improving energy-efficiency and purchasing carbon offsets. These approaches have limits: many cloud data centers already operate near peak efficiency, and carbon offsets cannot scale to near zero carbon where there is little carbon left to offset. Instead, enabling sustainable clouds will require applications to adapt to when and where unreliable low-carbon energy is available. Applications cannot do this today because their energy use and carbon emissions are not visible to them, as the energy system provides the rigid abstraction of a continuous, reliable energy supply. This vision paper instead advocates for a "carbon first" approach to cloud design that elevates carbon-efficiency to a firs--class metric. To do so, we argue that cloud platforms should virtualize the energy system by exposing visibility into, and software-defined control of, it to applications, enabling them to define their own abstractions for managing energy and carbon emissions based on their own requirements.
Noman Bashir, Tian Guo 0001, Mohammad Hajiesmaili, David Irwin 0001, Prashant J. Shenoy, Ramesh K. Sitaraman, Abel Souza, Adam Wierman
SoCC4
2021 Take it to the limit: peak prediction-driven resource overcommitment in datacenters
abstract
To increase utilization, datacenter schedulers often overcommit resources where the sum of resources allocated to the tasks on a machine exceeds its physical capacity. Setting the right level of overcommitment is a challenging problem: low overcommitment leads to wasted resources, while high overcommitment leads to task performance degradation. In this paper, we take a first principles approach to designing and evaluating overcommit policies by asking a basic question: assuming complete knowledge of each task's future resource usage, what is the safest overcommit policy that yields the highest utilization? We call this policy the peak oracle. We then devise practical overcommit policies that mimic this peak oracle by predicting future machine resource usage. We simulate our overcommit policies using the recently-released Google cluster trace, and show that they result in higher utilization and less overcommit errors than policies based on per-task allocations. We also deploy these policies to machines inside Google's datacenters serving its internal production workload. We show that our overcommit policies increase these machines' usable CPU capacity by 10-16% compared to no overcommitment.
Noman Bashir, Krzysztof Rzadca, David Irwin 0001, Sree Kodak, Rohit Jnagal
EuroSys4
2021 Model-driven Per-panel Solar Anomaly Detection for Residential Arrays
abstract
There has been significant growth in both utility-scale and residential-scale solar installations in recent years, driven by rapid technology improvements and falling prices. Unlike utility-scale solar farms that are professionally managed and maintained, smaller residential-scale installations often lack sensing and instrumentation for performance monitoring and fault detection. As a result, faults may go undetected for long periods of time, resulting in generation and revenue losses for the homeowner. In this article, we present SunDown, a sensorless approach designed to detect per-panel faults in residential solar arrays. SunDown does not require any new sensors for its fault detection and instead uses a model-driven approach that leverages correlations between the power produced by adjacent panels to detect deviations from expected behavior. SunDown can handle concurrent faults in multiple panels and perform anomaly classification to determine probable causes. Using two years of solar generation data from a real home and a manually generated dataset of multiple solar faults, we show that SunDown has a Mean Absolute Percentage Error of 2.98% when predicting per-panel output. Our results show that SunDown is able to detect and classify faults, including from snow cover, leaves and debris, and electrical failures with 99.13% accuracy, and can detect multiple concurrent faults with 97.2% accuracy.
Menghong Feng, Noman Bashir, Prashant J. Shenoy, David Irwin 0001, Beka Kosanovic
ACM Trans. Cyber Phys. Syst.4
2021 Modeling and Analyzing Waiting Policies for Cloud-Enabled Schedulers
abstract
Cloud platforms have popularized the Infrastructure-as-a-Service (IaaS) purchasing model, which enables users to rent computing resources on demand to execute their jobs. However, buying fixed resources is still much cheaper than renting if their resource utilization is high. Thus, to optimize cost, users must decide how many fixed resources to provision versus rent “on demand” based on their workload. In this article, we introduce the concept of a waiting policy for cloud-enabled schedulers and show that the optimal cost depends on it. The waiting policy explicitly controls how long jobs wait for resources, as jobs never need to wait, since cloud platforms provide the illusion of infinite scalability. A waiting policy is the dual of a scheduling policy: while a scheduling policy determines which jobs should run when fixed resources are available, a waiting policy determines which jobs should wait when fixed resources are not available. We define multiple waiting policies and develop simple and general analytical models to reveal their tradeoff between fixed resource provisioning, cost, and job waiting time. We evaluate the impact of different waiting policies on a real year-long batch workload consisting of 14M jobs run on a 14.3k-core cluster. We show that a compound waiting policy, which forces jobs with long running times or short waiting times to wait for fixed resources, offers the best tradeoff. The policy decreases both the cost (by 5 percent) and mean job waiting time (by 7×) compared to the current cluster, and also decreases the cost (by 43 percent) compared to renting on-demand resources for a modest increase in mean job waiting time (at 1.74 hours).
Lurdh Pradeep Reddy Ambati, Noman Bashir, David Irwin 0001, Prashant J. Shenoy
IEEE Trans. Parallel Distributed Syst.3
2020 Extend: A Framework for Increasing Energy Access by Interconnecting Solar Home Systems
abstract
The means of electrifying households and the resulting electricity networks are rapidly evolving. Traditionally, an extension of existing centralized grids was the only prominent technique, but now electrification is seeing massive expansion via decentralized solar home systems (SHSs). These systems consist of a low-wattage photovoltaic (PV) panel (typically 5-100W), a battery, a collection of energy-efficient DC appliances, and a charge controller. Spurred by significant advances and reduced costs in solar, batteries, energyefficient appliances, and mobile money-driven business models, SHSs have proliferated rapidly, with tens of millions of systems now deployed, primarily in regions with otherwise low rates of electricity access.
Santiago Correa, Noman Bashir, Andrew Tran, David Irwin 0001, Jay Taneja
COMPASS4
2020 SunDown: Model-driven Per-Panel Solar Anomaly Detection for Residential Arrays
abstract
Solar arrays often experience faults that go undetected for long periods of time, resulting in generation and revenue losses. In this paper, we present SunDown, a sensorless approach for detecting per-panel faults in solar arrays. SunDown's model-driven approach leverages correlations between the power produced by adjacent panels to detect deviations from expected behavior, can handle concurrent faults in multiple panels, and performs anomaly classification to determine probable causes. Using two years of solar data from a real home and a manually generated dataset of solar faults, we show that our approach is able to detect and classify faults, including from snow, leaves and debris, and electrical failures with 99.13% accuracy, and can detect concurrent faults with 97.2% accuracy.
Menghong Feng, Noman Bashir, Prashant J. Shenoy, David Irwin 0001, Dragoljub Kosanovic
COMPASS4
2020 Hedge Your Bets: Optimizing Long-term Cloud Costs by Mixing VM Purchasing Options
abstract
Cloud platforms offer the same VMs under many purchasing options that specify different costs and time commitments, such as on-demand, reserved, sustained-use, scheduled reserve, transient, and spot block. In general, the stronger the commitment, i.e., longer and less flexible, the lower the price. However, longer and less flexible time commitments can increase cloud costs for users if future workloads cannot utilize the VMs they committed to buying. Large cloud customers often find it challenging to choose the right mix of purchasing options to reduce their long-term costs, while retaining the ability to adjust capacity up and down in response to workload variations.To address the problem, we design policies to optimize long-term cloud costs by selecting a mix of VM purchasing options based on short- and long-term expectations of workload utilization. We consider a batch trace spanning 4 years from a large shared cluster for a major state University system that includes 14k cores and 60 million job submissions, and evaluate how these jobs could be judiciously executed using cloud servers using our approach. Our results show that our policies incur a cost within 41% of an optimistic optimal offline approach, and 50% less than solely using on-demand VMs.
Lurdh Pradeep Reddy Ambati, Noman Bashir, David Irwin 0001, Mohammad Hajiesmaili, Prashant J. Shenoy
IC2E3
2020 Waiting game: optimally provisioning fixed resources for cloud-enabled schedulers
abstract
While cloud platforms enable users to rent computing resources on demand to execute their jobs, buying fixed resources is still much cheaper than renting if their utilization is high. Thus, optimizing cloud costs requires users to determine how many fixed resources to buy versus rent based on their workload. In this paper, we introduce the concept of a waiting policy for cloud-enabled schedulers, which is the dual of a scheduling policy, and show that the optimal cost depends on it. We define multiple waiting policies and develop simple analytical models to reveal their tradeoff between fixed resource provisioning, cost, and job waiting time. We evaluate the impact of these waiting policies on a year-long production batch workload consisting of 14Mjobs run on a 14.3k-core cluster, and show that a compound waiting policy decreases the cost (by 5%) and mean job waiting time (by 7×) compared to a fixed cluster of the current size.
Lurdh Pradeep Reddy Ambati, Noman Bashir, David Irwin 0001, Prashant J. Shenoy
SC3
2019 Understanding Synchronization Costs for Distributed ML on Transient Cloud Resources
abstract
Cloud platforms often execute parallel batch applications, such as distributed machine learning (ML), that include numerous synchronization barriers. These barriers, which prevent any task from advancing beyond a specified point until all tasks have reached that point, significantly degrade application performance by reducing it to that of the slowest "straggler" task. To address the problem, researchers have proposed numerous straggler mitigation techniques, including speculatively re-executing straggler tasks and various relaxations of strict barrier semantics. While these techniques improve parallel application performance, they incur a cost in terms of the resources wasted re-executing tasks or waiting. Importantly, these costs, which are often implicit in prior work that targets dedicated resources, become explicit in the cloud, which charges for resources at fine-grained intervals. In addition, the cost difference between techniques is exacerbated in cloud platforms, since they charge substantially less for transient resources that effectively yield a probabilistic performance across a wide range. While transient resources' low list price is attractive, revocations increase the frequency and severity of stragglers, which decreases parallel job performance and increases overall execution cost. To better understand the cost of synchronization, we develop simple analytical models of different straggler mitigation techniques and compare their cost and performance on on-demand and transient resources. Our analysis shows that i) transient servers offer complex tradeoffs compared to on-demand servers, and can result in higher overall costs despite their highly discounted price due to their probabilistic performance; ii) common approaches to straggler mitigation, which is a well-studied problem, are less effective using transient servers that cause frequent and severe stragglers; and iii) a recent approach to flexible synchronization offers the best cost and performance.
Lurdh Pradeep Reddy Ambati, David Irwin 0001, Prashant J. Shenoy, Lixin Gao 0001, Ahmed Ali-Eldin, Jeannie R. Albrecht
IC2E2
2019 The Price Is (Not) Right: Reflections on Pricing for Transient Cloud Servers
abstract
Amazon introduced spot instances in December 2009, enabling "customers to bid on unused Amazon EC2 capacity and run those instances for as long as their bid exceeds the current Spot Price.'' Amazon's real-time computational spot market was novel in multiple respects. For example, it was the first (and to date only) large-scale public implementation of market-based resource allocation based on dynamic pricing after decades of research, and it provided users with useful information, control knobs, and options for optimizing the cost of running cloud applications. Spot instances also introduced the concept of transient cloud servers derived from variable idle capacity that cloud platforms could revoke at any time. Transient servers have since become central to efficient resource management of modern clusters and clouds. As a result, Amazon's spot market was the motivation for substantial research over the past decade. Yet, in November 2017, Amazon effectively ended its realtime spot market by announcing that users no longer needed to place bids and that spot prices will "...adjust more gradually, based on longer-term trends in supply and demand.'' The changes made spot instances more similar to the fixed-price transient servers offered by other cloud platforms. Unfortunately, while these changes made spot instances less complex, they eliminated many benefits to sophisticated users in optimizing their applications. This paper provides a retrospective on Amazon's real-time spot market, including its advantages and disadvantages for allocating transient servers compared to current fixed-price approaches. We also discuss some fundamental problems with Amazon's spot market, which we identified in prior work (from 2016), that predicted its eventual end. We then discuss potential options for allocating transient servers that combine the advantages of Amazon's real-time spot market, while also addressing the problems that likely led to its elimination.
David Irwin 0001, Prashant J. Shenoy, Lurdh Pradeep Reddy Ambati, Prateek Sharma 0001, Supreeth Shastri, Ahmed Ali-Eldin
ICCCN1
2019 Solar-TK: A Data-Driven Toolkit for Solar PV Performance Modeling and Forecasting
abstract
Solar energy capacity is continuing to increase. The key challenge with integrating solar into buildings and the electric grid is its high power generation variability, which is a function of many factors, including a site's location, time, weather, and numerous physical attributes. There has been significant prior work on solar performance modeling and forecasting that infers a site's current and future solar generation based on these factors. Accurate solar performance models and forecasts are also a pre-requisite for conducting a wide range of building and grid energy-efficiency research. Unfortunately, much of the prior work is not accessible to researchers, either because it has not been released as open source, is time-consuming to re-implement, or requires access to proprietary data sources. To address the problem, we present Solar-TK, a data-driven toolkit for solar performance modeling and forecasting that is simple, extensible, and publicly accessible. Solar-TK's simple approach models and forecasts a site's solar output given only its location and a small amount of historical generation data. Solar-TK's extensible design includes a small collection of independent modules that connect together to implement basic modeling and forecasting, while also enabling users to implement new energy analytics. We plan to release Solar-TK as open source to enable research that requires realistic solar models and forecasts, and to serve as a baseline for comparing new solar modeling and forecasting techniques. We compare Solar-TK's simple approach with PVlib and show that it yields comparable accuracy. We present three case studies showing how Solar-TK can advance energy-efficiency research.
Noman Bashir, Dong Chen 0010, David Irwin 0001, Prashant J. Shenoy
MASS3
2019 Building Virtual Power Meters for Online Load Tracking
abstract
Many energy optimizations require fine-grained, load-level energy data collected in real time, most typically by a plug-level energy meter. Online load tracking is the problem of monitoring an individual electrical load’s energy usage in software by analyzing the building’s aggregate smart meter data. Load tracking differs from the well-studied problem of load disaggregation in that it emphasizes per-load accuracy and efficient, online operation rather than accurate disaggregation of every building load via offline analysis. In essence, tracking a particular load creates a virtual power meter for it, which mimics having a networked-connected power meter attached to the load, but notably does not require tracking every other load as well. We propose PowerPlay , a model-driven system for performing accurate, high-performance online load tracking. Our results from applying the system to real-world energy data demonstrate that PowerPlay (i) enables efficient online tracking on low-power embedded platforms, (ii) scales to thousands of loads (across many buildings) on server platforms, and (iii) improves per-load accuracy by more than a factor of two compared to a state-of-the-art load disaggregation algorithm. Our results point to the potential of replacing physical energy meters by “virtual” power meters using a system like PowerPlay.
Sean Kenneth Barker, Sandeep Kalra, David Irwin 0001, Prashant J. Shenoy
ACM Trans. Cyber Phys. Syst.3
2019 Inferring Smart Schedules for Dumb Thermostats
abstract
Heating, ventilation, and air conditioning (HVAC) accounts for over 50% of a typical home’s energy usage. A thermostat generally controls HVAC usage in a home to ensure user comfort. In this article, we focus on making existing “dumb” programmable thermostats smart by applying energy analytics on smart meter data to infer home occupancy patterns and compute an optimized thermostat schedule. Utilities with smart meter deployments are capable of immediately applying our approach, called iProgram, to homes across their customer base. iProgram addresses new challenges in inferring home occupancy from smart meter data where (i) training data is not available and (ii) the thermostat schedule may be misaligned with occupancy, frequently resulting in high power usage during unoccupied periods. iProgram translates occupancy patterns inferred from opaque smart meter data into a custom schedule for existing types of programmable thermostats, e.g., 1-day, 7-day, and so on. We implement iProgram as a web service and show that it reduces the mismatch time between the occupancy pattern and the thermostat schedule by a median value of 44.28min (out of 100 homes) when compared to a default 8am-6pm weekday schedule, with a median deviation of 30.76min off the optimal schedule. Further, iProgram yields a daily energy savings of 0.42kWh on average across the 100 homes. Moreover, the schedules generated from iProgram converge to optimal schedules within a couple of weeks for most homes. We also show that homeowners having multiple HVAC zones can utilize iProgram and potentially increase unconditioned times of less occupied parts of their homes by 70%. Utilities may use iProgram to recommend thermostat schedules to customers and provide them estimates of potential energy savings in their energy bills.
Srinivasan Iyengar, Sandeep Kalra, Anushree Ghosh, David Irwin 0001, Prashant J. Shenoy, Benjamin M. Marlin
ACM Trans. Cyber Phys. Syst.4
2018 Sync-on-the-fly: A Parallel Framework for Gradient Descent Algorithms on Transient Resources
abstract
Many cloud service providers offer transient resources (i.e., spare servers) for a fraction of the cost of on-demand servers. Many big data analytics tasks composed of iterative computations are ideal to run on such transient resources. However, modern distributed data processing systems, such as MapReduce and Spark, provide little support for running iterative computation on transient resources. The fault-tolerant mechanism provided in MapReduce and Spark typically leads to cascading re-computations after revocations of transiently available resources. To address the problem, we propose a distributed framework, called Sync-on-the-fly, that takes advantage of the fact that many machine learning algorithms do not require fixed synchronization barriers. These synchronization barriers can be established at any time, such as immediately before workers running on transient servers are revoked. We adapt and implement widely used algorithms based on gradient descent, such as Logistic Regression and Matrix Factorization, as examples to illustrate Sync-on-the-fly's approach. Our evaluation shows that Sync-on-the-fly can achieve up to 5× speedup over Spark and reduce 85% of the costs.
Guoyi Zhao, Lixin Gao 0001, David Irwin 0001
IEEE BigData3
2018 Cloud Index Tracking: Enabling Predictable Costs in Cloud Spot Markets
abstract
Cloud spot markets rent VMs for a variable price that is typically much lower than the price of on-demand VMs, which makes them attractive for a wide range of large-scale applications. However, applications that run on spot VMs suffer from cost uncertainty, since spot prices fluctuate, in part, based on supply, demand, or both. The difficulty in predicting spot prices affects users and applications: the former cannot effectively plan their IT expenditures, while the latter cannot infer the availability and performance of spot VMs, which are a function of their variable price. Prior work attempts to address this uncertainty by modeling and predicting individual spot prices based on historical data. However, a single model likely does not apply to different spot VMs, since they may have different levels of supply and demand. In addition, cloud providers may unilaterally change spot pricing algorithms, as EC2 has done multiple times, which can invalidate existing price models and prediction methods.
Supreeth Shastri, David Irwin 0001
SoCC2
2018 Private Memoirs of IoT Devices: Safeguarding User Privacy in the IoT Era
abstract
The rise of the Internet-of-Things (IoT) holds great promise to transform people's lives by making society more efficient in many areas, including energy, transportation, healthcare, commerce, manufacturing, etc. At their core, IoT devices use sensors to collect data on real-world physical processes and then transmit it over the Internet to cloud servers, which store, process, and learn from the data to better optimize these processes, either directly (by issuing remote commands that actuate IoT devices) or indirectly (by issuing notifications that direct users to take some action). Unfortunately, IoT devices also expose users to multiple new types of privacy attacks. In particular, the sensor data collected from IoT devices can indirectly reveal a variety of sensitive private information. In addition, users generally connect IoT devices to local networks, which they implicitly trust, with little understanding of what the IoT device is doing on the network. In this visionpaper, we discuss recent work on sensor data privacy in the context of smart energy systems to provide examples of i) the surprising types of private information we can glean from seemingly innocuous IoT data and ii) the different types of defenses we have developed to preserve IoT data privacy for smart energy systems. These defenses lie at different discrete points in the tradeoff between user privacy and IoT functionality, which motivates ongoing work on developing defenses that provide a more tunable tradeoff. We also discuss the privacy implications of connecting tens-to-hundreds of untrusted IoT devices to implicitly trusted local networks, and avenues for research to mitigate these concerns.
Dong Chen 0010, Phuthipong Bovornkeeratiroj, David Irwin 0001, Prashant J. Shenoy
ICDCS3
2018 WattHome: A Data-driven Approach for Energy Efficiency Analytics at City-scale
abstract
Buildings consume over 40% of the total energy in modern societies and improving their energy efficiency can significantly reduce our energy footprint. In this paper, we present WattHome, a data-driven approach to identify the least energy efficient buildings from a large population of buildings in a city or a region. Unlike previous approaches such as least squares that use point estimates, WattHome uses Bayesian inference to capture the stochasticity in the daily energy usage by estimating the parameter distribution of a building. Further, it compares them with similar homes in a given population using widely available datasets. WattHome also incorporates a fault detection algorithm to identify the underlying causes of energy inefficiency. We validate our approach using ground truth data from different geographical locations, which showcases its applicability in different settings. Moreover, we present results from a case study from a city containing >10,000 buildings and show that more than half of the buildings are inefficient in one way or another indicating a significant potential from energy improvement measures. Additionally, we provide probable cause of inefficiency and find that 41%, 23.73%, and 0.51% homes have poor building envelope, heating, and cooling system faults respectively.
Srinivasan Iyengar, Stephen Lee, David Irwin 0001, Prashant J. Shenoy, Benjamin Weil
KDD3
2018 Mechanisms and Policies for Controlling Distributed Solar Capacity
abstract
The rapid expansion of intermittent grid-tied solar capacity is making the job of balancing electricity’s real-time supply and demand increasingly challenging. Recent work proposes mechanisms for actively controlling solar power in the grid at individual sites by enabling software to cap it as a fraction of its time-varying maximum output. However, while enforcing an equal fraction of each solar site’s time-varying maximum output results in “fair” short-term contributions of solar power across all sites, it does not result in “fair” long-term contributions of solar energy. Enforcing fair long-term energy access is important when controlling distributed solar capacity, since limits on solar output impact the compensation users receive for net metering and the battery capacity required to store excess solar energy. This discrepancy arises from fundamental differences in enforcing “fair” access to the grid to contribute solar energy, compared to analogous fair sharing in networks and processors. To address the problem, we first present both a centralized and distributed algorithm to enable control of distributed solar capacity that enforces fair grid energy access. We then present multiple policies that show how utilities can leverage this new distributed rate-limiting mechanism to reduce variations in grid demand from intermittent solar generation.
Noman Bashir, David Irwin 0001, Prashant J. Shenoy, Jay Taneja
ACM Trans. Sens. Networks2
2018 Managing Risk in a Derivative IaaS Cloud
abstract
Infrastructure-as-a-Service (IaaS) cloud platforms rent computing resources with different cost and availability tradeoffs. For example, users may acquire virtual machines (VMs) in the spot market-that are cheap, but can be unilaterally terminated by the cloud operator. Because of this revocation risk, spot servers have been conventionally used for delay and risk tolerant batch jobs. In this paper, we develop risk mitigation policies which allow even interactive applications to run on spot servers. Our System, SpotCheck is a derivative cloud platform, and provides the illusion of an IaaS platform that offers always-available VMs on demand for a cost near that of spot servers, and supports unmodified applications. SpotCheck's design combines virtualization-based mechanisms for fault-tolerance, and bidding and server selection policies for managing the risk and cost. We implement SpotCheck on EC2 and show that it i) provides nested VMs with 99.9989 percent availability, ii) achieves upto 2-5x cost savings compared to using on-demand VMs, and iii) eliminates any risk of losing VM state.
Prateek Sharma 0001, Stephen Lee, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy
IEEE Trans. Parallel Distributed Syst.4
2017 Weatherman: Exposing weather-based privacy threats in big energy data
abstract
Smart energy meters record electricity consumption and generation at fine-grained intervals, and are among the most widely deployed sensors in the world. Energy data embeds detailed information about a building's energy-efficiency, as well as the behavior of its occupants, which academia and industry are actively working to extract. In many cases, either inadvertently or by design, these third-parties only have access to anonymous energy data without an associated location. The location of energy data is highly useful and highly sensitive information: it can provide important contextual information to improve big data analytics or interpret their results, but it can also enable third-parties to link private behavior derived from energy data with a particular location. In this paper, we present Weatherman, which leverages a suite of analytics techniques to localize the source of anonymous energy data. Our key insight is that energy consumption data, as well as wind and solar generation data, largely correlates with weather, e.g., temperature, wind speed, and cloud cover, and that every location on Earth has a distinct weather signature that uniquely identifies it. Weatherman represents a serious privacy threat, but also a potentially useful tool for researchers working with anonymous smart meter data. We evaluate Weatherman's potential in both areas by localizing data from over one hundred smart meters using a weather database that includes data from over 35,000 locations. Our results show that Weatherman localizes coarse (one-hour resolution) energy consumption, wind, and solar data to within 16.68km, 9.84km, and 5.12km, respectively, on average, which is more accurate using much coarser resolution data than prior work on localizing only anonymous solar data using solar signatures.
Dong Chen 0010, David Irwin 0001
IEEE BigData2
2017 HotSpot: automated server hopping in cloud spot markets
abstract
Cloud spot markets offer virtual machines (VMs) for a dynamic price that is much lower than the fixed price of on-demand VMs. In exchange, spot VMs expose applications to multiple forms of risk, including price risk, or the risk that a VM's price will increase relative to others. Since spot prices vary continuously across hundreds of different types of VMs, flexible applications can mitigate price risk by moving to the VM that currently offers the lowest cost. To enable this flexibility, we present HotSpot, a resource container that "hops" VMs---by dynamically selecting and self-migrating to new VMs---as spot prices change. HotSpot containers define a migration policy that lowers cost by determining when to hop VMs based on the transaction costs (from vacating a VM early and briefly double paying for it) and benefits (the expected cost savings). As a side effect of migrating to minimize cost, HotSpot is also able to reduce the number of revocations without degrading performance. HotSpot is simple and transparent: since it operates at the systems-level on each host VM, users need only run an HotSpot-enabled VM image to use it. We implement a HotSpot prototype on EC2, and evaluate it using job traces from a production Google cluster. We then compare HotSpot to using on-demand VMs and spot VMs (with and without fault-tolerance) in EC2, and show that it is able to lower cost and reduce the number of revocations without degrading performance.
Supreeth Shastri, David Irwin 0001
SoCC2
2017 The Financialization of Cloud Computing: Opportunities and Challenges
abstract
Under competitive pressure to maximize their infrastructure's utilization and revenue, modern cloud platforms are quickly evolving into server markets that offer increasingly sophisticated contracts beyond simple on-demand servers, such as spot, preemptible, burstable, and reserved servers. In parallel, continuing advances in system and network virtualization are making server-time a more fungible commodity. These trends have motivated calls for open cloud commodity markets akin to other commodity markets, e.g., for oil, gold, corn, etc. However, such open cloud markets have not yet materialized due to key differences between cloud resources and other commodities. In particular, the relationship between applications and their underlying server resources is fundamentally different and more complex than other commodities. Unfortunately, software developers generally do not have the necessary background to effectively manage this complexity as part of their applications. Financial cloud computing is an emerging area that focuses on adapting and extending concepts from economics and finance to explicitly manage applications' tradeoffs between cost, risk, availability, and performance in cloud markets. A key goal of financial cloud computing is to develop systems-level abstractions and mechanisms that manage the market's complexity. This paper introduces this emerging area and its potential benefits, surveys related work, discusses challenges to realizing a cloud commodity market, and then outlines future research directions.
David Irwin 0001, Prateek Sharma 0001, Supreeth Shastri, Prashant J. Shenoy
ICCCN1
2017 Pervasive Energy Monitoring and Control Through Low-Bandwidth Power Line Communication
abstract
The Internet of Things (IoT) is growing rapidly, with increasingly sophisticated networking, sensing, and actuation functions embedded into everyday devices. One important IoT application is managing a building's energy usage by monitoring and controlling its electrical devices. Many existing IoT-enabled devices operate through low-cost and convenient power line networks, using protocols such as X10 and Insteon for communication. However, as these technologies have traditionally targeted low-bandwidth device control, they are often not readily suited to higher bandwidth uses such as continuous energy monitoring. In this paper, we consider the challenge of leveraging existing low-bandwidth power line communication networks for energy monitoring, and present several techniques that enable reliable and high-resolution monitoring in such networks. As a case study, we consider the popular Insteon protocol and show that intelligent polling and event detection methods can reduce the bandwidth requirements and undetected power events in a realworld Insteon network by 50% or more versus naive methods. Our techniques have been employed in a real IoT-enabled smart home, which has collected much of the data publicly released in the UMass Smart* energy dataset.
Sean Kenneth Barker, David Irwin 0001, Prashant J. Shenoy
IEEE Internet Things J.2
2017 Minimizing Transmission Loss in Smart Microgrids by Sharing Renewable Energy
abstract
Renewable energy (e.g., solar energy) is an attractive option to provide green energy to homes. Unfortunately, the intermittent nature of renewable energy results in a mismatch between when these sources generate energy and when homes demand it. This mismatch reduces the efficiency of using harvested energy by either (i) requiring batteries to store surplus energy, which typically incurs ∼ 20% energy conversion losses, or (ii) using net metering to transmit surplus energy via the electric grid’s AC lines, which severely limits the maximum percentage of renewable penetration possible. In this article, we propose an alternative structure where nearby homes explicitly share energy with each other to balance local energy harvesting and demand in microgrids. We develop a novel energy sharing approach to determine which homes should share energy, and when to minimize system-wide energy transmission losses in the microgrid. We evaluate our approach in simulation using real traces of solar energy harvesting and home consumption data from a deployment in Amherst, MA. We show that our system (i) reduces the energy loss on the AC line by 64% without requiring large batteries, (ii) performance scales up with larger battery capacities, and (iii) is robust to different energy consumption patterns and energy prediction accuracy in the microgrid.
Zhichuan Huang, Ting Zhu 0001, David Irwin 0001, Aditya Kumar Mishra, Daniel Sadoc Menasché, Prashant J. Shenoy
ACM Trans. Cyber Phys. Syst.3
2017 Enabling Distributed Energy Storage by Incentivizing Small Load Shifts
abstract
Reducing peak demands and achieving a high penetration of renewable energy sources are important goals in achieving a smarter grid. To reduce peak demand, utilities are introducing variable rate electricity prices to incentivize consumers to manually shift their demand to low-price periods. Consumers may also use energy storage to automatically shift their demand by storing energy during low-price periods for use during high-price periods. Unfortunately, variable rate pricing provides only a weak incentive for distributed energy storage and does not promote its adoption at large scales. In this article, we present the storage adoption dilemma to capture the problems with incentivizing energy storage using variable rate prices. To address the problem, we propose a simple pricing scheme, called flat-power pricing , which incentivizes consumers to shift small amounts of load to flatten their demand rather than shift as much of their power usage as possible to low-price, off-peak periods. We show that compared to variable rate pricing, flat-power pricing (i) reduces consumers’ upfront capital costs, as it requires significantly less storage capacity per consumer; (ii) increases energy storage’s return on investment, as it mitigates free riding and maintains the incentive to use energy storage at large scales; and (iii) uses aggregate storage capacity within 31% of an optimal centralized approach. In addition, unlike variable rate pricing, we also show that flat-power pricing incentivizes the scheduling of elastic background loads, such as air conditioners and heaters, to reduce peak demand. We evaluate our approach using real smart meter data from 14,000 homes in a small town.
David Irwin 0001, Srinivasan Iyengar, Stephen Lee, Aditya Kumar Mishra, Prashant J. Shenoy
ACM Trans. Cyber Phys. Syst.1
2017 A Cloud-Based Black-Box Solar Predictor for Smart Homes
abstract
The popularity of rooftop solar for homes is rapidly growing. However, accurately forecasting solar generation is critical to fully exploiting the benefits of locally generated solar energy. In this article, we present two machine-learning techniques to predict solar power from publicly available weather forecasts. We use these techniques to develop SolarCast, a cloud-based web service that automatically generates models that provide customized site-specific predictions of solar generation. SolarCast utilizes a “black box” approach that requires only (1) a site’s geographic location and (2) a minimal amount of historical generation data. Since we intend SolarCast for small rooftop deployments, it does not require detailed site- and panel-specific information, which owners may not know, but instead automatically learns these parameters for each site. We evaluate the accuracy of SolarCast’s different algorithms on two publicly available datasets, each containing over 100 rooftop deployments with a variety of attributes (e.g., climate, tilt, orientation, etc.). We show that SolarCast learns a more accurate model using much less data (∼1 month) than prior SVM-based approaches, which require ∼3 months of data. SolarCast also provides a programmatic API, enabling developers to integrate its predictions into energy efficiency applications. Finally, we present two case studies of using SolarCast to demonstrate how real-world applications can leverage its predictions. We first evaluate a “sunny” load scheduler, which schedules a dryer’s energy usage to maximally align with a home’s solar generation. We then evaluate a smart solar-powered charging station, which can optimally charge the maximum number of electric vehicles (EVs) on a given day. Our results indicate that a representative home is capable of reducing its grid demand up to 40% by providing a modest amount of flexibility (of ∼5 hours) in the dryer’s start time with opportunistic load scheduling. Further, our charging station uses SolarCast to provide EV owners the amount of energy they can expect to receive from solar energy sources.
Srinivasan Iyengar, Navin Sharma, David Irwin 0001, Prashant J. Shenoy, Krithi Ramamritham
ACM Trans. Cyber Phys. Syst.3
2016 Flint: batch-interactive data-intensive processing on transient servers
abstract
Cloud providers now offer transient servers, which they may revoke at anytime, for significantly lower prices than on-demand servers, which they cannot revoke. The low price of transient servers is particularly attractive for executing an emerging class of workload, which we call Batch-Interactive Data-Intensive (BIDI), that is becoming increasingly important for data analytics. BIDI workloads require large sets of servers to cache massive datasets in memory to enable low latency operation. In this paper, we illustrate the challenges of executing BIDI workloads on transient servers, where revocations (akin to failures) are the common case. To address these challenges, we design Flint, which is based on Spark and includes automated checkpointing and server selection policies that i) support batch and interactive applications and ii) dynamically adapt to application characteristics. We evaluate a prototype of Flint using EC2 spot instances, and show that it yields cost savings of up to 90% compared to using on-demand servers, while increasing running time by < 2%.
Prateek Sharma 0001, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy
EuroSys4
2016 SpotLight: An Information Service for the Cloud
abstract
Infrastructure-as-a-Service cloud platforms are incredibly complex: they rent hundreds of different types of servers across multiple geographical regions under a wide range of contract types that offer varying tradeoffs between risk and cost. Unfortunately, the internal dynamics of cloud platforms are opaque along several dimensions. For example, while the risk of servers not being available when requested is critical in optimizing the cloud's risk-cost tradeoffs, it is not typically made visible to users. Thus, inspired by prior work on Internet bandwidth probing, we propose actively probing cloud platforms to explicitly learn such information, where each "probe" is a request for a particular type of server. We model the relationships between different contracts types to develop a market-based probing policy, which leverages the insight that real-time prices in cloud spot markets loosely correlate with the supply (and availability) of fixed-price on-demand servers. That is, the higher the spot price for a server, the more likely the corresponding fixed-price on-demand server is not available. We incorporate market-based probing into SpotLight, an information service that enables cloud applications to query this and other data, and use it to monitor the availability of more than 4500 distinct server types across 9 geographical regions in Amazon's Elastic Compute Cloud over a 3 month period. We analyze this data to reveal interesting observations about the platform's internal dynamics. We then show how SpotLight enables two recently proposed derivative cloud services to select a better mix of servers to host applications, which improves their availability from ~70-90% to near 100% in practice.
David Irwin 0001, Prashant J. Shenoy
ICDCS2
2016 Transient guarantees: maximizing the value of idle cloud capacity
abstract
To prevent rejecting requests, cloud platforms typically provision for their peak demand. Thus, a platform's idle capacity can be significant, as demand varies widely over multiple time scales, e.g., daily and seasonally. To reduce waste, platforms have begun to offer this idle capacity in the form of transient servers, which they may unilaterally revoke, for much lower prices - ~50-90% less - than on-demand servers, which they cannot revoke. However, transient servers' revocation characteristics - their volatility and predictability - influence their performance, since they affect the overhead of fault-tolerance mechanisms applications use to handle revocations. Unfortunately, current cloud platforms offer no guarantees on revocation characteristics, which makes it difficult for users to optimally configure (and correctly value) transient servers. To address the problem, we propose the abstraction of a transient guarantee, which offers probabilistic assurances on revocation characteristics. Transient guarantees have numerous benefits: they increase the performance of transient servers, enable users to optimally use and correctly value them, and permit platforms to control their freedom to revoke them. We present policies for partitioning a variable amount of idle capacity into classes with different transient guarantees to maximize performance and value. We then implement and evaluate these policies on job traces from a production Google cluster. We show that our approach can increase the aggregate revenue from idle server capacity by up to ~6.5× compared to existing approaches.
Supreeth Shastri, Amr Rizk, David Irwin 0001
SC3
2016 Analyzing the Efficiency of a Green University Data Center
abstract
Data centers are an indispensable part of today's IT infrastructure. To keep pace with modern computing needs, data centers continue to grow in scale and consume increasing amounts of power. While prior work on data centers has led to significant improvements in their energy-efficiency, detailed measurements from these facilities' operations are not widely available, as data center design is often considered part of a company's competitive advantage. However, such detailed measurements are critical to the research community in motivating and evaluating new energy-efficiency optimizations. In this paper, we present a detailed analysis of a state-of-the-art 15MW green multi-tenant data center that incorporates many of the technological advances used in commercial data centers. We analyze the data center's computing load and its impact on power, water, and carbon usage using standard effectiveness metrics, including PUE, WUE, and CUE. Our results reveal the benefits of optimizations, such as free cooling, and provide insights into how the various effectiveness metrics change with the seasons and increasing capacity usage. More broadly, our PUE, WUE, and CUE analysis validate the green design of this LEED Platinum data center.
Patrick Pegus II, Benoy Varghese, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy, Anirban Mahanti, James Culbert, John Goodhue, Chris Hill
ICPE4
2016 Beyond Energy-Efficiency: Evaluating Green Datacenter Applications for Energy-Agility
abstract
Computing researchers have long focused on improving energy-efficiency under the implicit assumption that all energy is created equal. Yet, this assumption is actually incorrect: energy's cost and carbon footprint vary substantially over time. As a result, consuming energy inefficiently when it is cheap and clean may sometimes be preferable to consuming it efficiently when it is expensive and dirty. Green datacenters adapt their energy usage to optimize for such variations, as reflected in changing electricity prices or renewable energy output. Thus, we introduce energy-agility as a new metric to evaluate green datacenter applications. To illustrate fundamental tradeoffs in energy-agile design, we develop GreenSort, a distributed sorting system optimized for energy-agility. GreenSort is representative of the long-running, massively-parallel, data-intensive tasks that are common in datacenters and amenable to delays from power variations. Our results demonstrate the importance of energy-agile design when considering the benefits of using variable power. For example, we show that GreenSort requires 31% more time and energy to complete when power varies based on real-time electricity prices versus when it is constant. Thus, in this case, real-time prices should be at least 31% lower than fixed prices to warrant using them.
Supreeth Subramanya, Zain Mustafa, David Irwin 0001, Prashant J. Shenoy
ICPE3
2015 SpotOn: a batch computing service for the spot market
abstract
Cloud spot markets enable users to bid for compute resources, such that the cloud platform may revoke them if the market price rises too high. Due to their increased risk, revocable resources in the spot market are often significantly cheaper (by as much as 10×) than the equivalent non-revocable on-demand resources. One way to mitigate spot market risk is to use various fault-tolerance mechanisms, such as checkpointing or replication, to limit the work lost on revocation. However, the additional performance overhead and cost for a particular fault-tolerance mechanism is a complex function of both an application's resource usage and the magnitude and volatility of spot market prices.
Supreeth Subramanya, Tian Guo 0001, Prateek Sharma 0001, David Irwin 0001, Prashant J. Shenoy
SoCC4
2015 SpotCheck: designing a derivative IaaS cloud on the spot market
abstract
Infrastructure-as-a-Service (IaaS) cloud platforms rent resources, in the form of virtual machines (VMs), under a variety of contract terms that offer different levels of risk and cost. For example, users may acquire VMs in the spot market that are often cheap but entail significant risk, since their price varies over time based on market supply and demand and they may terminate at any time if the price rises too high. Currently, users must manage all the risks associated with using spot servers. As a result, conventional wisdom holds that spot servers are only appropriate for delay-tolerant batch applications. In this paper, we propose a derivative cloud platform, called SpotCheck, that transparently manages the risks associated with using spot servers for users.
Prateek Sharma 0001, Stephen Lee, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy
EuroSys4
2015 Cutting the Cost of Hosting Online Services Using Cloud Spot Markets
abstract
The use of cloud servers to host modern Internet-based services is becoming increasingly common. Today's cloud platforms offer a choice of server types, including non-revocable on-demand servers and cheaper but revocable spot servers. A service provider requiring servers can bid in the spot market where the price of a spot server changes dynamically according to the current supply and demand for cloud resources. Spot servers are usually cheap, but can be revoked by the cloud provider when the cloud resources are scarce. While it is well-known that spot servers can reduce the cost of performing time-flexible interruption-tolerant tasks, we explore the novel possibility of using spot servers for reducing the cost of hosting an Internet-based service such as an e-commerce site that must {\em always} be on and the penalty for service unavailability is high.
Prashant J. Shenoy, Ramesh K. Sitaraman, David Irwin 0001
HPDC4
2014 Combined heat and privacy: Preventing occupancy detection from smart meters
abstract
Electric utilities are rapidly deploying smart meters that record and transmit electricity usage in real-time. As prior research shows, smart meter data indirectly leaks sensitive, and potentially valuable, information about a home's activities. An important example of the sensitive information smart meters reveal is occupancy-whether or not someone is home and when. As prior work also shows, occupancy is surprisingly easy to detect, since it highly correlates with simple statistical metrics, such as power's mean, variance, and range. Unfortunately, prior research that uses chemical energy storage, e.g., batteries, to prevent appliance power signature detection is prohibitively expensive when applied to occupancy detection. To address this problem, we propose preventing occupancy detection using the thermal energy storage of large elastic heating loads already present in many homes, such as electric water and space heaters. In essence, our approach, which we call Combined Heat and Privacy (CHPr), controls the power usage of these large loads to make it look like someone is always home. We design a CHPr-enabled water heater that regulates its energy usage to mask occupancy without violating its objective, e.g., to provide hot water on demand, and evaluate it in simulation and using a prototype. Our results show that a 50-gallon CHPr-enabled water heater decreases the Matthews Correlation Coefficient (a standard measure of a binary classifier's performance) of a threshold-based occupancy detection attack in a representative home by 10x (from 0.44 to 0.045), effectively preventing occupancy detection at no extra cost.
Dong Chen 0010, David Irwin 0001, Prashant J. Shenoy, Jeannie R. Albrecht
PerCom2
2014 Empirical Characterization, Modeling, and Analysis of Smart Meter Data
abstract
Smart meter deployments are spurring renewed interest in analysis techniques for electricity usage data. However, an important prerequisite for data analysis is characterizing and modeling how electrical loads use power. While prior work has made significant progress in deriving insights from electricity data, one issue that limits accuracy is the use of general and often simplistic load models. Prior models often associate a fixed power level with an “on” state and either no power, or some minimal amount, with an “off” state. This paper's goal is to develop a new methodology for modeling electric loads that is both simple and accurate. Our approach is empirical in nature: we monitor a wide variety of common loads to distill a small number of common usage characteristics, which we then leverage to construct accurate load-specific models. We show that our models are significantly more accurate than binary on-off models, decreasing the root mean square error by as much as 8× for representative loads. Finally, we demonstrate three novel applications that use our empirical load models to analyze and derive insights from smart meter data, including i) generating device-accurate synthetic traces of building electricity usage; ii) filtering out loads that generate rapid and random power variations in smart meter data; and iii) detecting the presence of specific load models in time-series power data.
Sean Kenneth Barker, Sandeep Kalra, David Irwin 0001, Prashant J. Shenoy
IEEE J. Sel. Areas Commun.3
2013 GreenCache: augmenting off-the-grid cellular towers with multimedia caches
abstract
The growth of smartphones combined with advances in mobile networking have revolutionized the way people consume multimedia data. In particular, users in developing countries primarily rely on smartphones since they often do not have access to more powerful (and more expensive) computing devices. Unfortunately, cellular networks in developing countries have historically had low reliability, due to grid instability and lack of infrastructure. The situation has led network operators to experiment with running cellular towers "off the grid" using intermittent renewable energy sources. In parallel, network operators are also experimenting with co-locating server caches close to cell towers to reduce access latency and back-haul bandwidth. In this paper, we study techniques for optimizing multimedia caches for intermittent renewable energy sources. Specifically, we examine how to apply a blinking abstraction proposed in prior work, which rapidly transitions servers between an active and inactive state, to improve the performance of a multimedia cache powered by renewables, called GreenCache. Our results show that GreenCache's staggered load-proportional blinking policy, which coordinates when servers are active over brief intervals, results in 3X less buffering (or pause) time by the client compared to an activation blinking policy, which simply activates and deactivates servers over long periods as power fluctuates, for realistic power variations from renewable energy sources.
Navin Sharma, Dilip Kumar Krishnappa, David Irwin 0001, Michael Zink, Prashant J. Shenoy
MMSys3
2013 Yank: Enabling Green Data Centers to Pull the Plug
David Irwin 0001, Prashant J. Shenoy, K. K. Ramakrishnan
NSDI2
2013 GreenCharge: Managing RenewableEnergy in Smart Buildings
abstract
Distributed generation (DG) uses many small on-site energy harvesting deployments at individual buildings to generate electricity. DG has the potential to make generation more efficient by reducing transmission and distribution losses, carbon emissions, and demand peaks. However, since renewables are intermittent and uncontrollable, buildings must still rely, in part, on the electric grid for power. While DG deployments today use net metering to offset costs and balance local supply and demand, scaling net metering for intermittent renewables to a large fraction of buildings is challenging. In this paper, we explore an alternative approach that combines market-based electricity pricing models with on-site renewables and modest energy storage (in the form of batteries) to incentivize DG. We propose a system architecture and optimization algorithm, called GreenCharge, to efficiently manage the renewable energy and storage to reduce a building's electric bill. To determine when to charge and discharge the battery each day, the algorithm leverages prediction models for forecasting both future energy demand and future energy harvesting. We evaluate GreenCharge in simulation using a collection of real-world data sets, and compare with an oracle that has perfect knowledge of future energy demand/harvesting and a system that only leverages a battery to lower costs (without any renewables). We show that GreenCharge's savings for a typical home today are near 20%, which are greater than the savings from using only net metering.
Aditya Kumar Mishra, David Irwin 0001, Prashant J. Shenoy, James F. Kurose, Ting Zhu 0001
IEEE J. Sel. Areas Commun.2
2012 Compute cloud based weather detection and warning system
abstract
Compute cloud platforms pay-as-you-use model suits applications which require resources sporadically. Severe weather detection and prediction is one such application. Since severe weather events are rare, dedicating servers for such application wastes resources. In this paper, we present the feasibility of using commercial cloud services for severe weather detection and prediction. We show that commercial cloud services provide the required network capability to perform the real-time operation of weather detection and prediction from the radars to the cloud service instance. We automate the process of weather prediction on the cloud based on the results of our weather detection algorithms.
Dilip Kumar Krishnappa, Eric Lyons 0001, David Irwin 0001, Michael Zink
IGARSS3
2012 CloudCast: Cloud computing for short-term mobile weather forecasts
abstract
Since today's weather forecasts only cover large regions every few hours, their use in severe weather is limited. In this paper, we present CloudCast, an application that provides short-term weather forecasts depending on users current location. Since severe weather is rare, CloudCast leverages pay-as-you-go cloud platforms to eliminate dedicated computing infrastructure. CloudCast has two components: 1) an architecture linking weather radars to cloud resources, and 2) a Nowcasting algorithm for generating accurate short-term weather forecasts. We study CloudCast's design space, which requires significant data staging to the cloud. Our results indicate that serial transfers achieve tolerable throughput, while parallel transfers represent a bottleneck for real-time mobile Nowcasting. We also analyze forecast accuracy and show high accuracy for ten minutes in the future. Finally, we execute CloudCast live using an on-campus radar, and show that it delivers a 15-minute Nowcast to a mobile client in less than 2 minutes after data sampling started.
Dilip Kumar Krishnappa, David Irwin 0001, Eric Lyons 0001, Michael Zink
IPCCC2
2012 Network capabilities of cloud services for a real time scientific application
abstract
Dedicating high-end servers for executing scientific applications that run intermittently, such as severe weather detection or generalized weather forecasting, wastes resources. While the Infrastructure-as-a-Service (IaaS) model used by today's cloud platforms is well-suited for the bursty computational demands of these applications, it is unclear if the network capabilities of today's cloud platforms are sufficient. In this paper, we analyze the networking capabilities of multiple commercial (Amazon's EC2 and Rackspace) and research (GENICloud and ExoGENI cloud) platforms in the context of a Nowcasting application, a forecasting algorithm for highly accurate, near-term, e.g., 5-20 minutes, weather predictions. The application has both computational and network requirements. While it executes rarely, whenever severe weather approaches, it benefits from an IaaS model; However, since its results are time-critical, enough bandwidth must be available to transmit radar data to cloud platforms before it becomes stale. We conduct network capacity measurements between radar sites and cloud platforms throughout the country. Our results indicate that ExoGENI cloud performs the best for both serial and parallel data transfer with an average throughput of 110.22 Mbps and 17.2 Mbps, respectively. We also found that the cloud services perform better in the distributed data transfer case, where a subset of nodes transmit data in parallel to a cloud instance. Ultimately, we conclude that commercial and research clouds are capable of providing sufficient bandwidth for our real-time Nowcasting application.
Dilip Kumar Krishnappa, Eric Lyons 0001, David Irwin 0001, Michael Zink
LCN3
2012 SmartCap: Flattening peak electricity demand in smart homes
abstract
Flattening household electricity demand reduces generation costs, since costs are disproportionately affected by peak demands. While the vast majority of household electrical loads are interactive and have little scheduling flexibility (TVs, microwaves, etc.), a substantial fraction of home energy use derives from background loads with some, albeit limited, flexibility. Examples of such devices include A/Cs, refrigerators, and dehumidifiers. In this paper, we study the extent to which a home is able to transparently flatten its electricity demand by scheduling only background loads with such flexibility. We propose a Least Slack First (LSF) scheduling algorithm for household loads, inspired by the well-known Earliest Deadline First algorithm. We then integrate the algorithm into Smart-Cap, a system we have built for monitoring and controlling electric loads in homes. To evaluate LSF, we collected power data at outlets, panels, and switches from a real home for 82 days. We use this data to drive simulations, as well as experiment with a real testbed implementation that uses similar background loads as our home. Our results indicate that LSF is most useful during peak usage periods that exhibit “peaky” behavior, where power deviates frequently and significantly from the average. For example, LSF decreases the average deviation from the mean power by over 20% across all 4-hour periods where the deviation is at least 400 watts.
Sean Kenneth Barker, Aditya Kumar Mishra, David Irwin 0001, Prashant J. Shenoy, Jeannie R. Albrecht
PerCom3
2012 MultiSense: proportional-share for mechanically steerable sensor networks
Navin Sharma, David Irwin 0001, Michael Zink, Prashant J. Shenoy
Multim. Syst.2
2011 Blink: managing server clusters on intermittent power
abstract
Reducing the energy footprint of data centers continues to receive significant attention due to both its financial and environmental impact. There are numerous methods that limit the impact of both factors, such as expanding the use of renewable energy or participating in automated demand-response programs. To take advantage of these methods, servers and applications must gracefully handle intermittent constraints in their power supply. In this paper, we propose blinking---metered transitions between a high-power active state and a low-power inactive state---as the primary abstraction for conforming to intermittent power constraints. We design Blink, an application-independent hardware-software platform for developing and evaluating blinking applications, and define multiple types of blinking policies. We then use Blink to design BlinkCache, a blinking version of memcached, to demonstrate the effect of blinking on an example application. Our results show that a load-proportional blinking policy combines the advantages of both activation and synchronous blinking for realistic Zipf-like popularity distributions and wind/solar power signals by achieving near optimal hit rates (within 15% of an activation policy), while also providing fairer access to the cache (within 2% of a syn- chronous policy) for equally popular objects.
Navin Sharma, Sean Kenneth Barker, David Irwin 0001, Prashant J. Shenoy
ASPLOS3
2011 MultiSense: fine-grained multiplexing for steerable camera sensor networks
abstract
Steerable sensors, such as pan-tilt-zoom video cameras, expose programmable actuators to applications, which steer them in different directions based on their goals. Despite being expensive to deploy and maintain, existing steerable sensor networks allow only a single application to control them due to the slow speed of their mechanical actuators. To address the problem, we design MultiSense to enable fine-grained multiplexing by (i) exposing a virtual sensor to each application and (ii) optimizing the time to context-switch between virtual sensors and satisfy requests.
Navin Sharma, David Irwin 0001, Prashant J. Shenoy, Michael Zink
MMSys2
2010 Resource management in data-intensive clouds: Opportunities and challenges
abstract
Today's cloud computing platforms have seen much success in running compute-bound applications with time-varying or one-time needs. In this position paper, we will argue that the cloud paradigm is also well suited for handling data-intensive applications, characterized by the processing and storage of data produced by high-bandwidth sensors or streaming applications. The data rates and the processing demands vary over time for many such applications, making the on-demand cloud paradigm a good match for their needs. However, today's cloud platforms need to evolve to meet the storage, communication, and processing demands of data-intensive applications. We present an ongoing GENI project to connect high-bandwidth radar sensor networks with computational and storage resources in the cloud and use this example to highlight the opportunities and challenges in designing end-to-end data-intensive cloud systems.
David Irwin 0001, Prashant J. Shenoy, Emmanuel Cecchet, Michael Zink
LANMAN1
2010 Cloudy Computing: Leveraging Weather Forecasts in Energy Harvesting Sensor Systems
abstract
To sustain perpetual operation, systems that harvest environmental energy must carefully regulate their usage to satisfy their demand. Regulating energy usage is challenging if a system's demands are not elastic and its hardware components are not energy-proportional, since it cannot precisely scale its usage to match its supply. Instead, the system must choose when to satisfy its energy demands based on its current energy reserves and predictions of its future energy supply. In this paper, we explore the use of weather forecasts to improve a system's ability to satisfy demand by improving its predictions. We analyze weather forecast, observational, and energy harvesting data to formulate a model that translates a weather forecast to a wind or solar energy harvesting prediction, and quantify its accuracy. We evaluate our model for both energy sources in the context of two different energy harvesting sensor systems with inelastic demands: a sensor testbed that leases sensors to external users and a lexicographically fair sensor network that maintains steady node sensing rates. We show that using weather forecasts in both wind- and solar-powered sensor systems increases each system's ability to satisfy its demands compared with existing prediction strategies.
Navin Sharma, Jeremy Gummeson, David Irwin 0001, Prashant J. Shenoy
SECON3
2009 SRCP: Simple Remote Control for Perpetual High-Power Sensor Networks
Navin Sharma, Jeremy Gummeson, David Irwin 0001, Prashant J. Shenoy
EWSN3
2007 Automated and on-demand provisioning of virtual machines for database applications
abstract
Utility computing delivers compute and storage resources to applications as an 'on-demand utility', much like electricity, from a distributed collection of computing resources. There is great interest in running database applications on utility resources (e.g., Oracle's Grid initiative) due to reduced infrastructure and management costs, higher resource utilization, and the ability to handle sudden load surges. Virtual Machine (VM) technology offers powerful mechanisms to manage a utility resource infrastructure. However, provisioning VMs for applications to meet system performance goals, e.g., to meet service level agreements (SLAs), is an open problem. We are building two systems at Duke - Shirako and NIMO - that collectively address this problem.
Piyush Shivam, Azbayar Demberel, Pradeep Gunda, David Irwin 0001, Laura E. Grit, Aydan R. Yumerefendi, Shivnath Babu, Jeffrey S. Chase
SIGMOD Conference4
2006 Ensemble-level Power Management for Dense Blade Servers
abstract
One of the key challenges for high-density servers (e.g., blades) is the increased costs in addressing the power and heat density associated with compaction. Prior approaches have mainly focused on reducing the heat generated at the level of an individual server. In contrast, this work proposes power efficiencies at a larger scale by leveraging statistical properties of concurrent resource usage across a collection of systems ("ensemble"). Specifically, we discuss an implementation of this approach at the blade enclosure level to monitor and manage the power across the individual blades in a chassis. Our approach requires low-cost hardware modifications and relatively simple software support. We evaluate our architecture through both prototyping and simulation. For workloads representing 132 servers from nine different enterprise deployments, we show significant power budget reductions at performances comparable to conventional systems.
Parthasarathy Ranganathan, Phil Leech, David Irwin 0001, Jeffrey S. Chase
ISCA3
2006 Grid allocation and reservation - Toward a doctrine of containment: grid hosting with adaptive resource control
abstract
Grid computing environments need secure resource control and predictable service quality in order to be sustainable. We propose a grid hosting model in which independent, self-contained grid deployments run within isolated containers on shared resource provider sites. Sites and hosted grids interact via an underlying resource control plane to manage a dynamic binding of computational resources to containers. We present a prototype grid hosting system, in which a set of independent Globus grids share a network of cluster sites. Each grid instance runs a coordinator that leases and configures cluster resources for its grid on demand. Experiments demonstrate adaptive provisioning of cluster resources and contrast job-level and container-level resource management in the context of two grid application managers.
Lavanya Ramakrishnan, David Irwin 0001, Laura E. Grit, Aydan R. Yumerefendi, Adriana Iamnitchi, Jeffrey S. Chase
SC2
2006 Sharing Networked Resources with Brokered Leases
David Irwin 0001, Jeffrey S. Chase, Laura E. Grit, Aydan R. Yumerefendi, Ken Yocum
USENIX ATC, General Track1
2004 Balancing Risk and Reward in a Market-Based Task Service
David Irwin 0001, Laura E. Grit, Jeffrey S. Chase
HPDC1
2003 Dynamic Virtual Clusters in a Grid Site Manager
abstract
This paper presents new mechanisms for dynamic resource management in a cluster manager called Cluster-on-Demand (COD). COD allocates servers from a common pool to multiple virtual clusters (vclusters), with independently configured software environments, name spaces, user access controls, and network storage volumes. We present experiments using the popular Sun GridEngine batch scheduler to demonstrate that dynamic virtual clusters are an enabling abstraction for advanced resource management in computing utilities and grids. In particular, they support dynamic, policy-based cluster sharing between local users and hosted Grid services, resource reservation and adaptive provisioning, scavenging of the idle resources, and dynamic instantiation of Grid services. These goals are achieved in a direct and general way through a new set of fundamental cluster management functions, with minimal impact on the Grid middleware itself.
Jeffrey S. Chase, David Irwin 0001, Laura E. Grit, Justin D. Moore, Sara Sprenkle
HPDC2