EDBT 2026 Demo / reviewers in the wild / expert
Richard Wolski
dblp:w/RichardWolski · also Rich Wolski
· DBLP profile ↗
97ranked-venue papers
15as first author
12since 2021 · last 2024
0000-0003-3722-473XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 65 · 13 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Computer networks · 4Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Distributed Dataflow Across the Edge-Cloud ContinuumabstractInternet of Things (IoT) applications span the edge-cloud continuum to form multiscale distributed systems. The heterogeneity that defines this architecture, coupled with the asynchronous, event-triggered and failure-prone nature of these deployments create significant programming and maintenance challenges for developers of IoT applications. To address this impediment to innovation, we present Lam-1nar’a dataflow programming model for IoT applications implemented using a novel log-based and concurrent runtime system that spans all resource scales. We describe the properties that underpin Laminar'sdesign and compare it to a lower-level event-based approach. We show that Laminar'sdataflow model hides many of the complexities of “lock-free” event-driven programming. Through an empirical evaluation of Laminar,we find its design and implementation are both more straightforward for developers and more performant. Tyler Ekaireb, Lukas Brand, Nagarjun Avaraddy, Markus Mock, Chandra Krintz, Richard Wolski |
CLOUD | 6 |
| 2024 | Energy-Aware IoT Deployment PlanningabstractIncreasingly, the Internet of Things (IoT) is evolving toward an architecture consisting of sensing and actuation devices communicating with edge computers and storage systems. These "edge deployments" localize communication, computation, and storage for security, increased efficiencies (e.g. lower latency response), and reliability. In settings where electrical power infrastructure is lacking, however, these deployments typically rely on renewable energy and battery storage for power. Peiyuan Guan, Animesh Dangwal, Amirhosein Taherkordi, Richard Wolski, Chandra Krintz |
CF | 4 |
| 2023 | Depot: Dependency-Eager Platform of TransformationsabstractThis paper presents a new model for a data management system specifically designed to enable community-curated data repositories and collaboration. Depot (a Dependence-Eager Platform of Transformations) is based on a data-lake approach that eases the technological burdens associated with data contribution while providing an interactive programming environment for developing transformations that result in structured tables supporting SQL database operations. Crucially, Depot implements lazy evaluation of these transformations so that only the structured data that is demanded by a data consumer is generated. Until the structured data is "materialized," Depot tracks and maintains the dependencies that are required to perform the eventual materialization. This lazy approach to creating structured data allows Depot to maintain a smaller resource footprint compared to a typical data warehouse approach while maintaining the flexibility of the data lake model. Furthermore, Depot is designed as a community-sustainable platform. The initial prototype is implemented for cloud deployment and it distributes the storage and ETL workload cost among the data consumers. Performance results of the early prototype are encouraging, making Depot a new infrastructure for creating data lakes that foster contributed-consumer collaboration. Kerem Çelik, Samridhi Maheshwari, Shereen ElSayed, Markus Mock, Chandra Krintz, Richard Wolski |
CloudCom | 6 |
| 2023 | GreenCoin: A Renewable Energy-Aware CryptocurrencyabstractIn this paper, we propose GreenCoin – an energy-efficient cryptocurrency system with mining protocols designed to favor locations with relatively higher availability of renewable energy. Traditionally, crypto coin mining involves solving complex mathematical problems by high-end computing devices consuming an enormous amount of electricity, thus adversely affecting net carbon emissions. To reduce cost and emissions, GreenCoin uses a modified proof of stake (PoS) consensus algorithm, which itself is more energy efficient compared to other state-of-the-art methods. Our modified PoS algorithm, called Green PoS (GPoS), allows GreenCoin to favor nodes (with reward and privilege) located in regions with higher availability of renewable energy. We present a detailed system architecture of GreenCoin and explain the operating method of GPoS. We also provide results from empirical studies demonstrating the renewable energy-aware approach of GreenCoin. Shivaansh Kapoor, Chandra Krintz, Richard Wolski, Markus Mock |
IC2E | 4 |
| 2023 | Data Acquisition and Analysis for Improving the Utility of Low Cost Soil Moisture SensorsabstractTo cultivate healthy plants and high crop yields, growers must be able to measure soil moisture and irrigate accordingly. Errors in soil moisture measurements can lead to irrigation mismanagement with costly consequences. In this paper, we present a new approach to smart computing for irrigation management to address these challenges at a lower cost. We calibrate low cost, low precision soil moisture sensors to more accurately distinguish wet from dry soils using high cost, high precision Davis Instrument sensors. We investigate different modeling techniques including the natural log of the odds ratio (Log-odds), Monte Carlo simulation, and linear regression to distinguish between wet and moist soils and to establish a trustworthy threshold between these two moisture states. We have also developed a new smartphone application that simplifies the process of data collection and implements our analysis approach. The application is extensible by others and provides growers with low cost, data-driven decision support for irrigation. We implement our approach for UCSB’s Edible Campus student farm and empirically evaluate it using multiple test beds. Our results show an accuracy rate of 91% and lowers costs by 4x per deployment, making it useful for gardeners and farmers alike. Gautam Mundewadi, Richard Wolski, Chandra Krintz |
SMARTCOMP | 2 |
| 2023 | Replicated Versioned Data Structures for Wide-Area Distributed SystemsabstractIn this work, we investigate the integration of replicated versioned data structures and append-only distributed storage systems. Doing so facilitates high availability and scalability while providing developer access to different versions of program data structures across program executions. Modern distributed systems such as the Internet of Things (IoT) often employ multi-tiered (cloud/edge/sensors) architectures consisting of a wide array of heterogeneous devices generating data frequently. Hence system availability is imperative to avoid data loss, while scalability is required for the efficient operation of the system not only within the same tier but across different tiers as well. Our proposed approach replicates, persists, and versionsprogram data structuressuch as binary search trees and linked lists for use in distributed IoT applications. The versioning and persistence of these structures aid failure recovery and facilitate system debugging from its inception instead of making such considerations an afterthought. Moreover, our experiments suggest versioned data structures can perform better in applications performing high volumes of temporal queries versus traditional methods of persisting data (e.g., in a database). We empirically evaluate the overheads associated with versioning and storage persistence of program data structures, present experimental results for multiple end-to-end applications, and demonstrate the scalability of this approach. Chandra Krintz, Richard Wolski |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Log-Based CRDT for Edge ApplicationsabstractIn this paper, we investigate extensions for Conflict-Free Replicated Data Types (CRDTs) that permit their use in failure-prone, heterogeneous, resource-constrained, distributed, multi-tier (cloud/edge/device) cloud deployments such as the Internet-of-Things (IoT), while addressing multiple CRDT limitations. Specifically, we employ distributed logging to implement robust, strong eventual consistency of replicas. Our approach also enables uniform reversal of operations and precludes the requirement of exactly-once delivery and idempotence imposed by operation-based CRDTs. Moreover, it exposes CRDT versions for use in debugging and history-based programming. We evaluate our approach for commonly used CRDTs and show that it enables higher operation throughput (up to 1.8x) versus conventional CRDTs for the workloads we consider. Chandra Krintz, Richard Wolski |
IC2E | 3 |
| 2022 | Collaborative experience between scientific software projects using Agile Scrum developmentabstractDeveloping sustainable software for the scientific community requires expertise in software engineering and domain science. This can be challenging due to the unique needs of scientific software, the insufficient resources for software engineering practices in the scientific community, and the complexity of developing for evolving scientific contexts. While open-source software can partially address these concerns, it can introduce complicating dependencies and delay development. These issues can be reduced if scientists and software developers collaborate. We present a case study wherein scientists from the SuperNova Early Warning System collaborated with software developers from the Scalable Cyberinfrastructure for Multi-Messenger Astrophysics project. The collaboration addressed the difficulties of open-source software development, but presented additional risks to each team. For the scientists, there was a concern of relying on external systems and lacking control in the development process. For the developers, there was a risk in supporting a user-group while maintaining core development. These issues were mitigated by creating a second Agile Scrum framework in parallel with the developers' ongoing Agile Scrum process. This Agile collaboration promoted communication, ensured that the scientists had an active role in development, and allowed the developers to evaluate and implement the scientists' software requirements. The collaboration provided benefits for each group: the scientists actuated their development by using an existing platform, and the developers utilized the scientists' use-case to improve their systems. This case study suggests that scientists and software developers can avoid scientific computing issues by collaborating and that Agile Scrum methods can address emergent concerns. Amanda L. Baxter, Segev Y. BenZvi, Walter M. Bonivento, Adam Brazier, Michael Clark, Alexis Coleiro, David Collom, Marta Colomer-Molla, Bryce Cousins, Aliwen Delgado Orellana, Damien Dornic, Vladislav Ekimtcov, Shereen ElSayed, Andrea Gallo Rosso, Patrick Godwin, Spencer Griswold, Alec Habig, Remington Hill, Shunsaku Horiuchi, D. Andrew Howell, Margaret W. G. Johnson, Mario Juric, James P. Kneller, Abigail Kopec, Claudio Kopper, Vladimir Kulikovskiy, Mathieu Lamoureux, Rafael F. Lang, Shengchao Li, Massimiliano Lincetto, Lindy Lindstrom, Mark W. Linvill, Curtis McCully, Jost Migenda, Danny Milisavljevic, Spencer Nelson, Rita Novoseltseva, Erin O'Sullivan, Donald Petravick, Barry W. Pointon, Nirmal Raj, Andrew Renshaw, Janet Rumleskie, Tom Sonley, Ron Tapia, Jeffrey C. L. Tseng, Christopher D. Tunnell, Godefroy Vannoye, Carlo F. Vigorito, Clarence J. Virtue, Christopher Weaver 0001, Kathryn E. Weil, Lindley Winslow, Richard Wolski, Xun-Jie Xu |
Softw. Pract. Exp. | 54 |
| 2021 | On the Future of Cloud EngineeringabstractEver since the commercial offerings of the Cloud started appearing in 2006, the landscape of cloud computing has been undergoing remarkable changes with the emergence of many different types of service offerings, developer productivity enhancement tools, and new application classes as well as the manifestation of cloud functionality closer to the user at the edge. The notion of utility computing, however, has remained constant throughout its evolution, which means that cloud users always seek to save costs of leasing cloud resources while maximizing their use. On the other hand, cloud providers try to maximize their profits while assuring service-level objectives of the cloud-hosted applications and keeping operational costs low. All these outcomes require systematic and sound cloud engineering principles. The aim of this paper is to highlight the importance of cloud engineering, survey the landscape of best practices in cloud engineering and its evolution, discuss many of the existing cloud engineering advances, and identify both the inherent technical challenges and research opportunities for the future of cloud computing in general and cloud engineering in particular. David Bermbach, Abhishek Chandra, Chandra Krintz, Aniruddha S. Gokhale, Aleksander Slominski, Lauritz Thamsen, Everton Cavalcante, Tian Guo 0001, Ivona Brandic, Richard Wolski |
IC2E | 10 |
| 2021 | PEDaLS: Persisting Versioned Data StructuresabstractIn this paper, we investigate how to automatically persist versioned data structures in distributed settings (e.g. cloud + edge) using append-only storage. By doing so, we facilitate resiliency by enabling program state to survive program activations and termination, and program-level data structures and their version information to be accessed programmatically by multiple clients (for replay, provenance tracking, debugging, and coordination avoidance, and more). These features are useful in distributed, failure-prone contexts such as those for heterogeneous and pervasive Internet of Things (IoT) deployments. We prototype our approach within an open-source, distributed operating system for IoT. Our results show that it is possible to achieve algorithmic complexities similar to those of in-memory versioning but in a distributed setting. Chandra Krintz, Richard Wolski |
IC2E | 3 |
| 2021 | CAPLets: Resource Aware, Capability-Based Access Control for IoT
Fatih Bakir, Chandra Krintz, Richard Wolski |
SEC | 3 |
| 2021 | Edge-adaptable serverless acceleration for machine learning Internet of Things applicationsabstractAbstract Serverless computing is an emerging event‐driven programming model that accelerates the development and deployment of scalable web services on cloud computing systems. Though widely integrated with the public cloud, serverless computing use is nascent for edge‐based, Internet of Things (IoT) deployments. In this work, we present STOIC (serverless teleoperable hybrid cloud), an IoT application deployment and offloading system that extends the serverless model in three ways. First, STOIC adopts a dynamic feedback control mechanism to precisely predict latency and dispatch workloads uniformly across edge and cloud systems using a distributed serverless framework. Second, STOIC leverages hardware acceleration (e.g., GPU resources) for serverless function execution when available from the underlying cloud system. Third, STOIC can be configured in multiple ways to overcome deployment variability associated with public cloud use. We overview the design and implementation of STOIC and empirically evaluate it using real‐world machine learning applications and multitier IoT deployments (edge and cloud). Specifically, we show that STOIC can be used fortrainingimage processing workloads (for object recognition)—once thought too resource‐intensive for edge deployments. We find that STOIC reduces overall execution time (response latency) and achieves placement accuracy that ranges from 92% to 97%. Chandra Krintz, Richard Wolski |
Softw. Pract. Exp. | 3 |
| 2020 | NanoLambda: Implementing Functions as a Service at All Resource Scales for the Internet of ThingsabstractInternet of Things (IoT) devices are becoming increasingly prevalent in our environment, yet the process of programming these devices and processing the data they produce remains difficult. Typically, data is processed on device, involving arduous work in low level languages, or data is moved to the cloud, where abundant resources are available for Functions as a Service (FaaS) or other handlers. FaaS is an emerging category of flexible computing services, where developers deploy self-contained functions to be run in portable and secure containerized environments; however, at the moment, these functions are limited to running in the cloud or in some cases at the “edge” of the network using resource rich, Linux-based systems.In this paper, we present NanoLambda, a portable platform that brings FaaS, high-level language programming, and familiar cloud service APIs to non-Linux and microcontroller-based IoT devices. To enable this, NanoLambda couples a new, minimal Python runtime system that we have designed for the least capable end of the IoT device spectrum, with API compatibility for AWS Lambda and S3. NanoLambda transfers functions between IoT devices (sensors, edge, cloud), providing power and latency savings while retaining the programmer productivity benefits of high-level languages and FaaS. A key feature of NanoLambda is a scheduler that intelligently places function executions across multi-scale IoT deployments according to resource availability and power constraints. We evaluate a range of applications that use NanoLambda to run on devices as small as the ESP8266 with 64KB of ram and 512KB flash storage. Gareth George, Fatih Bakir, Richard Wolski, Chandra Krintz |
SEC | 3 |
| 2020 | Detecting Performance Anomalies in Cloud Platform ApplicationsabstractWe present Roots, a full-stack monitoring and analysis system for performance anomaly detection and bottleneck identification in cloud platform-as-a-service (PaaS) systems. Roots facilitates application performance monitoring as a core capability of PaaS clouds, and relieves the developers from having to instrument application code. Roots tracks HTTP/S requests to hosted cloud applications and their use of PaaS services. To do so it employs lightweight monitoring of PaaS service interfaces. Roots processes this data in the background using multiple statistical techniques that in combination detect performance anomalies (i.e. violations of service-level objectives). For each anomaly, Roots determines whether the event was caused by a change in the request workload or by a performance bottleneck in a PaaS service. By correlating data collected across different layers of the PaaS, Roots is able to trace high-level performance anomalies to bottlenecks in specific components in the cloud platform. We implement Roots using the AppScale PaaS and evaluate its overhead and accuracy. Hiranya Jayathilaka, Chandra Krintz, Richard Wolski |
IEEE Trans. Cloud Comput. | 3 |
| 2019 | Seneca: Fast and Low Cost Hyperparameter Search for Machine Learning ModelsabstractThe goal of our work is to simplify and expedite the construction and evaluation of machine learning models using autoscaled cloud computing resources. To enable this, we develop an open source system called Seneca, which leverages the serverless programming model and its implementation in Amazon Web Services (AWS) Lambda. Seneca takes a machine learning application, dataset, and a list of possible hyperparameter options as input and automatically constructs an AWS Lambda function. The function ingresses and splits the input dataset into training and testing subsets and constructs, tests, and evaluates (i.e. scores) a machine learning model for a given set of hyperparameter values. Seneca concurrently invokes functions for all combinations of the hyperparameters specified. It then returns the configuration (or model) that results in the best score to the user. In this paper, we overview the design and implementation of Seneca, and empirically evaluate its performance for a popular classification application. Chandra Krintz, Markus Mock, Richard Wolski |
CLOUD | 4 |
| 2019 | Analyzing AWS Spot Instance PricingabstractMany cloud computing vendors offer a preemptible class of service for rented virtual machines. In November 2017, Amazon.com changed the pricing mechanism for its preemptible "spot instances" so that prices would change more "smoothly." This paper analyzes the effect of this change on spot instance prices. It examines the prices immediately before and after the mechanism change to determine the extent to which prices themselves changed. It then compares the 90-day period immediately after the change in mechanism to the next 90-day period. Finally, it compares the two most recent 90-day periods (ending on October 15, 2018). Our results indicate that in addition to smoothing prices, the mechanism change introduced generally higher prices which is a trend that continues. Gareth George, Richard Wolski, Chandra Krintz, John Brevik |
IC2E | 2 |
| 2018 | Tracing Function Dependencies across CloudsabstractIn this paper, we present Lowgo, a crosscloud tracing tool for capturing causal relationships in serverless applications. To do so, Lowgo records dependencies between functions, through cloud services, and across regions to facilitate debugging and reasoning about highly concurrent, multi-cloud applications. We empirically evaluate Lowgo using microbenchmarks and multi-function and multi-cloud applications. We find that Lowgo is able to capture causal dependencies with overhead that ranges from 2-12%, which is less than half that of the best-performing, cloud-specific approach. Wei-Tsung Lin, Chandra Krintz, Richard Wolski |
IEEE CLOUD | 3 |
| 2018 | Tracking Causal Order in AWS Lambda ApplicationsabstractServerless computing is a new cloud programming and deployment paradigm that is receiving wide-spread uptake. Serverless offerings such as Amazon Web Services (AWS) Lambda, Google Functions, and Azure Functions automatically execute simple functions uploaded by developers, in response to cloud-based event triggers. The serverless abstraction greatly simplifies integration of concurrency and parallelism into cloud applications, and enables deployment of scalable distributed systems and services at very low cost. Although a significant first step, the serverless abstraction requires tools that software engineers can use to reason about, debug, and optimize their increasingly complex, asynchronous applications. Toward this end, we investigate the design and implementation of GammaRay, a cloud service that extracts causal dependencies across functions and through cloud services, without programmer intervention. We implement GammaRay for AWS Lambda and evaluate the overheads that it introduces for serverless micro-benchmarks and applications written in Python. Wei-Tsung Lin, Chandra Krintz, Richard Wolski, Xiaogang Cai, Tongjun Li, Weijin Xu |
IC2E | 3 |
| 2017 | PYTHIA: Admission Control for Multi-Framework, Deadline- Driven, Big Data WorkloadsabstractIn this paper, we present PYTHIA, deadline-aware admission control for systems that execute jobs from multiple big data (batch) frameworks using shared resources. PYTHIA adds support for deadline-driven workloads in resource-constrained cloud settings, for use by resource negotiators such as Apache Mesos or YARN. PYTHIA uses histories of job statistics to estimate the minimum number of CPUs to allocate to a job in order for it to meet its deadline. PYTHIA admits jobs when these resources are available. Any job not admitted “fails fast”and wastes no resources. We implement a PYTHIA prototype and empirically evaluate it using production YARN traces under different resource constraints and deadline assignments. Our results show that PYTHIA is able to meet significantly more deadlines than fair share approaches and wastes fewer cloud resources in resource-limited scenarios, for the workloads, cluster sizes, and deadline assignments that we consider Stratos Dimopoulos, Chandra Krintz, Richard Wolski |
CLOUD | 3 |
| 2017 | QPRED: Using Quantile Predictions to Improve Power Usage for Private CloudsabstractIn this paper we describe a new, efficient predictive scheduling methodology for implementing computing infrastructure power savings using private clouds. Our approach, termed "QPRED," estimates the quantiles on the distribution of future machine usage so that unneeded machines may be powered down to save power. A cloud administrator sets a bound on the probability that all available machines will be powered down when a cloud request arrives. This target probability is the basis of a Service Level Agreement between the cloud administrator and all cloud users covering start-up delay resulting from power savings. Our results, validated using activity traces from several private clouds used in commercial production, indicate that QPRED successfully reduces power consumption substantially while maintaining the SLAs specified by the cloud administrator. Richard Wolski, John Brevik |
CLOUD | 1 |
| 2017 | Justice: A Deadline-Aware, Fair-Share Resource Allocator for Implementing Multi-AnalyticsabstractIn this paper, we present Justice, a fair-share deadline-aware resource allocator for big data cluster managers. In resource constrained environments, where resource contention introduces significant execution delays, Justice outperforms the popular existing fair-share allocator that is implemented as part of Mesos and YARN. Justice uses deadline information supplied with each job and historical job execution logs to implement admission control. It automatically adapts to changing workload conditions to assign enough resources for each job to meet its deadline "just in time." We use trace-based simulation of production YARN workloads to evaluate Justice under different deadline formulations. We compare Justice to the existing fair-share allocation policy deployed on cluster managers like YARN and Mesos and find that in resource-constrained settings, Justice improves fairness, satisfies significantly more deadlines, and utilizes resources more efficiently. Stratos Dimopoulos, Chandra Krintz, Richard Wolski |
CLUSTER | 3 |
| 2017 | EXFed: Efficient Cross-Federation with Availability SLAs on Preemptible IaaS InstancesabstractPrivate IaaS clouds offer the benefits of cloud computing on-site but their efficiency is limited by capacity constraints during peak times. We present EXFed, an efficient cross-federation system for IaaS clouds that "ships" jobs between clouds and provides ahead-of-time certainty about resource availability despite retaining individual clouds' ability to preempt foreign workload after admission. Clouds participating in the federation remain in control of their local resources at all times and exclusively use a predictable tier of preemptible instances to run federated jobs. This predictable tier is enabled through a new method that provides an SLA on the preemption probability of groups of instances. The SLA is learned statistically from cloud utilization and data transfer rates in the recent past. We deploy EXFed across multiple data centers and evaluate its robustness under realistic and adverse scenarios with production traces recorded from industrial "big data" clouds. Alexander Pucher, Richard Wolski, Chandra Krintz |
IC2E | 2 |
| 2017 | Probabilistic guarantees of execution duration for Amazon spot instancesabstractIn this paper we propose DrAFTS - a methodology for implementing probabilistic guarantees of instance reliability in the Amazon Spot tier. Amazon offers "unreliable" virtual machine instances (ones that may be terminated at any time) at a potentially large discount relative to "reliable" On-demand and Reserved instances. Our method predicts the "bid values" that users can specify to provision Spot instances which ensure at least a fixed duration of execution with a given probability. We illustrate the method and test its validity using Spot pricing data post facto, both randomly and using real-world workload traces. We also test the efficacy of the method experimentally by using it to launch Spot instances and then observing the instance termination rate. Our results indicate that it is possible to obtain the same level of reliability from unreliable instances that the Amazon service level agreement guarantees for reliable instances with a greatly reduced cost. Richard Wolski, John Brevik, Ryan Chard, Kyle Chard |
SC | 1 |
| 2017 | Performance Monitoring and Root Cause Analysis for Cloud-hosted Web ApplicationsabstractIn this paper, we describe Roots - a system for automatically identifying the "root cause" of performance anomalies in web applications deployed in Platform-as-a-Service (PaaS) clouds. Roots does not require application-level instrumentation. Instead, it tracks events within the PaaS cloud that are triggered by application requests using a combination of metadata injection and platform-level instrumentation. Hiranya Jayathilaka, Chandra Krintz, Richard Wolski |
WWW | 3 |
| 2016 | Big data framework interference in restricted private cloud settingsabstractIn this paper, we characterize the behavior of “big” and “fast” data analysis frameworks, in multi-tenant, shared settings for which computing resources (CPU and memory) are limited, an increasingly common scenario used to increase utilization and lower cost. We study how popular analytics frameworks behave and interfere with each other under such constraints. We empirically evaluate Hadoop, Spark, and Storm multi-tenant workloads managed by Mesos. Our results show that in constrained environments, there is significant performance interference that manifests in failed fair sharing, performance variability, and deadlock of resources. Stratos Dimopoulos, Chandra Krintz, Richard Wolski |
IEEE BigData | 3 |
| 2015 | Response time service level agreements for cloud-hosted web applicationsabstractCloud computing is a successful model for hosting web-facing applications that are accessed by their users as services. While clouds currently offer Service Level Agreements (SLAs) containing guarantees of availability, they do not make performance guarantees for deployed applications. Hiranya Jayathilaka, Chandra Krintz, Richard Wolski |
SoCC | 3 |
| 2015 | Service-Level Agreement Durability for Web Service Response TimeabstractCloud computing is an attractive model for deploying web services in a highly scalable manner. Users access such cloud-hosted services via their web-facing application programming interfaces (APIs). Prior work has shown that it is possible to use a combined approach of static analysis and cloud platform monitoring to predict the response time upper bounds of such web APIs. This technique can be employed to automatically generate service level agreements (SLAs) concerning the performance of cloud-hosted web APIs. In this work, we explore the validity period of auto-generated SLAs in cloud settings. We discuss a simple model by which API consumers can establish a response time SLA with the cloud platform, and renegotiate it when/if the SLA becomes invalid due to the dynamic nature of the cloud. Using empirical methods and simulations on a real world public cloud platform, we show that it is possible to auto-generatehighly durable response time SLAs for cloud-hosted web APIs, thereby keeping the number of SLA invalidations and renegotiations very low, over long periods. Hiranya Jayathilaka, Chandra Krintz, Richard Wolski |
CloudCom | 3 |
| 2015 | SuperContra: Cross-Language, Cross-Runtime Contracts as a ServiceabstractThis paper presents SuperContra - a Design-by-Contract (DbC) framework that can ship with future PaaS offerings to enforce lightweight contracts across different programming systems, as-a-service. SuperContra is unique in that developers employ a familiar, high-level language to write contracts regardless of the programming language used to implement the component under test. We evaluate SuperContra using widely used, open-source software and compare its performance against existing DbC frameworks. Our results show that SuperContra performs on par with non-service-based DbC approaches and in some cases similarly to code running without contracts. Stratos Dimopoulos, Chandra Krintz, Richard Wolski, Anand Gupta |
IC2E | 3 |
| 2015 | EAGER: Deployment-Time API Governance for Modern PaaS CloudsabstractTo track, control, and compel reuse of web APIs, we investigate a new approach to API governance -- combined policy, implementation, and deployment control of web APIs. Our approach, called EAGER, provides a software architecture that integrates into PaaS platforms to support systemwide, deployment-time enforcement of governance policies. Specifically, EAGER checks for and prevents backward incompatible API changes from being deployed into production PaaS clouds, enforces service reuse, and facilitates enforcement of other best practices in software maintenance via policies. Our experiments with an EAGER prototype show that enforcing API governance at deployment-time in PaaS clouds is efficient and scalable to thousands of APIs and policies. Hiranya Jayathilaka, Chandra Krintz, Richard Wolski |
IC2E | 3 |
| 2015 | Using Trustworthy Simulation to Engineer Cloud SchedulersabstractIn recent years, researchers have contributed promising new techniques for allocating cloud resources in more robust, efficient, and ecologically sustainable ways. Unfortunately, the wide-spread use of these techniques in production systems has, to date, remained elusive. One reason for this is that the state of the art for investigating these innovations at scale often relies solely on model-driven simulation. Production-grade cloud software, however, demands certainty and precision for development and business planning that only comes from validating simulation against empirical observation. In this work, we take an alternative approach to facilitating cloud research and engineering in order to transition innovations to production deployment faster. In particular, we present a new methodology that complements existing model-driven simulation with platform-specific and statistically trustworthy results. We simulate systems at scales and on time frames that are testable, and then, based on the statistical validation of these simulations, investigate scenarios beyond those feasibly observable in practice. We demonstrate the approach by developing an energy-aware cloud scheduler and evaluating it using production and synthetic traces in faster than real time. Our results show that we can accurately simulate a production IaaS system, ease capacity planning, and expedite the reliable development of its components and extensions. Alexander Pucher, Emre Gul, Richard Wolski, Chandra Krintz |
IC2E | 3 |
| 2015 | VM-centric snapshot deduplication for cloud data backupabstractData deduplication is important for snapshot backup of virtual machines (VMs) because of excessive redundant content. Fingerprint search for source-side duplicate detection is resource intensive when the backup service for VMs is co-located with other cloud services. This paper presents the design and analysis of a fast VM-centric backup service with a tradeoff for a competitive deduplication efficiency while using small computing resources, suitable for running on a converged cloud architecture that cohosts many other services. The design consideration includes VM-centric file system block management for the increased VM snapshot availability. This paper describes an evaluation of this VM-centric scheme to assess its deduplication efficiency, resource usage, and fault tolerance. Wei Zhang 0118, Daniel Agun, Tao Yang 0009, Richard Wolski, Hong Tang 0004 |
MSST | 4 |
| 2014 | Cloud Platform Support for API GovernanceabstractAs scalable information technology evolves to a more cloud-like model, digital assets (code, data and software environments) increasingly require curation as web-accessible services. "Service-izing" digital assets consists of encapsulating assets in software that exposes them to web and mobile applications via well-defined yet flexible, network accessible, application programming interfaces (APIs). In this paper, we postulate that recent advances in cloud computing make cloud platforms as-a-service (PaaS) ideal for deployment, lifecycle management, and policy-based control i.e. API governance - for extant and future digital assets. Toward this end, we overview API governance as a PaaS technology and outline some early results generated by our investigation of a prototype we are developing, called EAGER, for implementing API governance at scale. Chandra Krintz, Hiranya Jayathilaka, Stratos Dimopoulos, Alexander Pucher, Richard Wolski, Tevfik Bultan |
IC2E | 5 |
| 2014 | Using Parametric Models to Represent Private Cloud WorkloadsabstractCloud computing has become a popular metaphor for dynamic and secure self-service access to computational and storage capabilities. In this study, we analyze and model workloads gathered from enterprise-operated commercial private clouds that implement “Infrastructure as a Service.” Our results show that 3-phase hyperexponential distributions fit using the Estimation Maximization (E-M) algorithm capture workload attributes accurately. In addition, these models of individual attributes compose to produce estimates of overall cloud performance that our results verify to be accurate. As an early study of commercial enterprise private clouds, this work provides guidance to those researching, designing, or maintaining such installations. In particular, the cloud workloads under study do not exhibit “heavy-tailed” distributional properties in the same way that “bare metal” operating systems do, potentially leading to different design and engineering tradeoffs. Richard Wolski, John Brevik |
IEEE Trans. Serv. Comput. | 1 |
| 2011 | Deadline-sensitive workflow orchestration without explicit resource control
Lavanya Ramakrishnan, Jeffrey S. Chase, Dennis Gannon, Daniel Nurmi, Richard Wolski |
J. Parallel Distributed Comput. | 5 |
| 2009 | The Eucalyptus Open-Source Cloud-Computing SystemabstractCloud computing systems fundamentally provide access to large pools of data and computational resources through a variety of interfaces similar in spirit to existing grid and HPC resource management and programming systems. These types of systems offer a new programming target for scalable application developers and have gained popularity over the past few years. However, most cloud computing systems in operation today are proprietary, rely upon infrastructure that is invisible to the research community, or are not explicitly designed to be instrumented and modified by systems researchers. In this work, we present Eucalyptus - an open-source software framework for cloud computing that implements what is commonly referred to as infrastructure as a service (IaaS); systems that give users the ability to run and control entire virtual machine instances deployed across a variety physical resources. We outline the basic principles of the Eucalyptus design, detail important operational aspects of the system, and discuss architectural trade-offs that we have made in order to allow EUCALYPTUS to be portable, modular and simple to use on infrastructure commonly found within academic settings. Finally, we provide evidence that EUCALYPTUS enables users familiar with existing grid and HPC systems to explore new cloud computing functionality while maintaining access to existing, familiar application development software and grid middleware. Daniel Nurmi, Richard Wolski, Chris Grzegorczyk, Graziano Obertelli, Sunil Soman, Lamia Youseff, Dmitrii Zagorodnov |
CCGRID | 2 |
| 2009 | VGrADS: enabling e-Science workflows on grids and clouds with fault toleranceabstractToday's scientific workflows use distributed heterogeneous resources through diverse grid and cloud interfaces that are often hard to program. In addition, especially for time-sensitive critical applications, predictable quality of service is necessary across these distributed resources. VGrADS' virtual grid execution system (vgES) provides an uniform qualitative resource abstraction over grid and cloud systems. We apply vgES for scheduling a set of deadline sensitive weather forecasting workflows. Specifically, this paper reports on our experiences with (1) virtualized reservations for batchqueue systems, (2) coordinated usage of TeraGrid (batch queue), Amazon EC2 (cloud), our own clusters (batch queue) and Eucalyptus (cloud) resources, and (3) fault tolerance through automated task replication. The combined effect of these techniques was to enable a new workflow planning method to balance performance, reliability and cost considerations. The results point toward improved resource selection and execution management support for a variety of e-Science applications over grids and cloud systems. Lavanya Ramakrishnan, Charles Koelbel, Yang-Suk Kee, Richard Wolski, Daniel Nurmi, Dennis Gannon, Graziano Obertelli, Asim YarKhan, Anirban Mandal, T. Mark Huang, Kiran Thyagaraja, Dmitrii Zagorodnov |
SC | 4 |
| 2008 | VARQ: virtual advance reservations for queuesabstractIn high-performance computing (HPC) settings, in which multiprocessor machines are shared among users with potentially competing resource demands, processors are allocated to user workload using space sharing. Typically, users interact with a given machine by submitting their jobs to a centralized batch scheduler that implements a site-specific policy designed to maximize machine utilization while providing tolerable turn-around times. To these users, the functioning of the batch scheduler and the policies it implements are both critical operating system components since they control how each job is serviced. In practice, while most HPC systems experience good utilization levels, the amount of time experienced by individual jobs waiting to begin execution has been shown to be highly variable and difficult to predict, leading to user confusion and/or frustration. Daniel Nurmi, Richard Wolski, John Brevik |
HPDC | 2 |
| 2008 | The impact of paravirtualized memory hierarchy on linear algebra computational kernels and softwareabstractPrevious studies have revealed that paravirtualization imposes minimal performance overhead on High Performance Computing (HPC) workloads, while exposing numerous benefits for this field. In this study, we are investigating the memory hierarchy characteristics of paravirtualized systems and their impact on automatically-tuned software systems. We are presenting an accurate characterization of memory attributes using hardware counters and user-process accounting. For that, we examine the proficiency of ATLAS, a quintessential example of an autotuning software system, in tuning the BLAS library routines for paravirtualized systems. In addition, we examine the effects of paravirtualization on the performance boundary. Our results show that the combination of ATLAS and Xen paravirtualization delivers native execution performance and nearly identical memory hierarchy performance profiles. Our research thus exposes new benefits to memory-intensive applications arising from the ability to slim down the guest OS without influencing the system performance. In addition, our findings support a novel and very attractive deployment scenario for computational science and engineering codes on virtual clusters and computational clouds. Lamia Youseff, Keith Seymour, Haihang You, Jack J. Dongarra, Richard Wolski |
HPDC | 5 |
| 2008 | Enabling personal clusters on demand for batch resources using commodity softwareabstractProviding QoS (quality of service) in batch resources against the uncertainty of resource availability due to the space-sharing nature of scheduling policies is a critical capability required for high-performance computing. This paper introduces a technique called personal cluster which reserves a partition of batch resources on user's demand in a best-effort manner. A personal cluster provides a private cluster dedicated to the user during a user-specified time period by installing a user-level resource manager on the resource partition. This technique not only enables cost-effective resource utilization and efficient task management but also provides the user a uniform interface to heterogeneous resources regardless of local resource management software. A prototype implementation using a PBS batch resource manager and Globus Toolkits based on Web services shows that the overhead of instantiating a personal cluster of medium size is small, which is just about 1 minute for a personal cluster having 32 processors. Yang-Suk Kee, Carl Kesselman, Daniel Nurmi, Richard Wolski |
IPDPS | 4 |
| 2008 | Using bandwidth data to make computation offloading decisionsabstractWe present a framework for making computation offloading decisions in computational grid settings in which schedulers determine when to move parts of a computation to more capable resources to improve performance. Such schedulers must predict when an offloaded computation will outperform one that is local by forecasting the local cost (execution time for computing locally) and remote cost (execution time for computing remotely and transmission time for the input/output of the computation to/from the remote system). Typically, this decision amounts to predicting the bandwidth between the local and remote systems to estimate these costs. Our framework unifies such decision models by formulating the problem as a statistical decision problem that can either be treated "classically" or using a Bayesian approach. Using an implementation of this framework, we evaluate the efficacy of a number of different decision strategies (several of which have been employed by previous systems). Our results indicate that a Bayesian approach employing automatic change-point detection when estimating the prior distribution is the best-performing approach. Richard Wolski, Selim Gurun, Chandra Krintz, Daniel Nurmi |
IPDPS | 1 |
| 2008 | Probabilistic advanced reservations for batch-scheduled parallel machinesabstractNo abstract available. Daniel Nurmi, Richard Wolski, John Brevik |
PPoPP | 2 |
| 2008 | Efficient auction-based grid reservations using dynamic programmingabstractAuction mechanisms have been proposed as a means to efficiently and fairly schedule jobs in high-performance computing environments. The generalized vickrey auction has long been known to produce efficient allocations while exposing users to truth-revealing incentives, but the algorithms used to compute its payments can be computationally intractable. In this paper we present a novel implementation of the generalized vickrey auction that uses dynamic programming to schedule jobs and compute payments in pseudo-polynomial time. Additionally, we have built a version of the PBS scheduler that uses this algorithm to schedule jobs, and in this paper we present the results of our tests using this scheduler. Andrew Mutz, Richard Wolski |
SC | 2 |
| 2008 | NWSLite: A general-purpose, nonparametric prediction utility for embedded systemsabstractTime series-based prediction methods have a wide range of uses in embedded systems. Many OS algorithms and applications require accurate prediction of demand and supply of resources. However, configuring prediction algorithms is not easy, since the dynamics of the underlying data requires continuous observation of the prediction error and dynamic adaptation of the parameters to achieve high accuracy. Current prediction methods are either too costly to implement on resource-constrained devices or their parameterization is static, making them inappropriate and inaccurate for a wide range of datasets. This paper presents NWSLite, a prediction utility that addresses these shortcomings on resource-restricted platforms. Selim Gurun, Chandra Krintz, Richard Wolski |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2007 | Isla Vista Heap Sizing: Using Feedback to Avoid PagingabstractManaged runtime environments (MREs) employ garbage collection (GC) for automatic memory management. However, GC induces pressure on the virtual memory (VM) manager, since it may touch pages that are not related to the working set of the application. Paging due to GC can significantly hurt performance, even when the application's working set fits into physical memory. We present a feedback-directed heap resizing mechanism to avoid GC-induced paging, using information from the operating system (OS). We avoid costly GCs when there is physical memory available, and trade off GC for paging when memory is constrained. Our mechanism is simple and uses allocation stall events during GC alone to trigger heap resizing, without user participation or OS kernel modification. Our system enables significant performance improvements when real memory is restricted and similar to, or better performance than, the current state-of-the-art MRE, when memory is unconstrained Chris Grzegorczyk, Sunil Soman, Chandra Krintz, Richard Wolski |
CGO | 4 |
| 2007 | Deploying Video-on-Demand Services on Cable NetworksabstractEfficient video-on-demand (VoD) is a highly desired service for media and telecom providers. VoD allows subscribers to view any item in a large media catalog nearly-instantaneously. However, systems that provide this services currently require large amounts of centralized resources and significant bandwidth to accommodate their subscribers. Hardware requirements become more substantial as the service providers increase the catalog size or number of subscribers. In this paper, we describe how cable companies can leverage deployed hardware in a peer- to-peer architecture to provide an efficient alternative We propose a distributed VoD system, and use real measurements from a deployed VoD system to evaluate different design decisions. Our results show that with minor changes, currently deployed cable infrastructures can support a video-on-demand system that scales to a large number of users and catalog size with low centralized resources. Matthew S. Allen, Ben Y. Zhao, Richard Wolski |
ICDCS | 3 |
| 2007 | VIProf: Vertically Integrated Full-System Performance ProfilerabstractIn this paper, we present VIProf, a full-system, performance sampling system capable of extracting runtime behavior across an entire software stack. Our long-term goal is to employ VIProf profiles to guide online optimization of programs and their execution environments according to the dynamically changing execution behavior and resource availability. VIProf thus, must be transparent while producing accurate and useful performance profiles. We overview the design and implementation of VIProf and empirically evaluate the system using a popular software stack - one that includes a Linux operating system, a Java virtual machine, and a set of applications. This composition is commonly employed and important for high-end systems such as application and Web servers as well as computational grid services. We show that VIProf introduces little overhead and is able to capture accurate (function-level) full-system performance data that previously required multiple profiles and extensive, manual, and offline post-processing of profile data. Hussam Mousa, Chandra Krintz, Lamia Youseff, Richard Wolski |
IPDPS | 4 |
| 2007 | An Analysis of Availability Distributions in CondorabstractIn this article, we investigate the dynamics exhibited by the production Condor pool at the University of Wisconsin with the goal of understanding its distributional properties. Condor is a cycle-harvesting service originally designed to launch and control "guest" user jobs (in batch mode) on idle workstations. Since its inception in 1985, however, it has expanded to include the ability to run in dedicated mode on clusters, to "glide in" to systems that are not strictly dedicated to Condor, and to "flock" jobs from one site to another based on pre-determined service level agreements (SLAs). Thus it has developed from an enterprise-wide desktop system into a full-fledged global computing infrastructure over its lifetime. Richard Wolski, Daniel Nurmi, John Brevik |
IPDPS | 1 |
| 2007 | QBETS: Queue Bounds Estimation from Time Series
Daniel Nurmi, John Brevik, Richard Wolski |
JSSPP | 3 |
| 2007 | Disens: scalable distributed sensor network simulationabstractSimulation is widely used for developing, evaluating and analyzing sensor network applications, especially when deploying a large scale sensor network remains expensive and labor intensive. However, due to its computation intensive nature, existent simulation tools have to make trade-offs between fidelity and scalability and thus offer limited capabilities as design and analysis tools. In this paper, we introduce DiSenS (DIstributed SENsor network Simulation) -- a highly scalable distributed simulation system for sensor networks. DiSenS does not only faithfully emulates an extensive set of sensor hardware and supports extensible radio/power models, so that sensor network applications can be simulated transparently with high fidelity, but also employs distributed-memory parallel cluster system to attack the complex simulation problem. Combining an efficient distributed synchronization protocol and a sophisticated node partitioning algorithm (based on existent research), DiSenS achieves greater scalability than even many discrete event simulators. On a small to medium size cluster (16-64 nodes), DiSenS is able to simulate hundreds of motes in realtime speed and scale to thousands in sub-realtime speed. To our knowledge, DiSenS is the first full-system sensor network simulator with such scalability. Ye Wen, Richard Wolski, Gregory Moore |
PPoPP | 2 |
| 2007 | Simulation-based augmented reality for sensor network developmentabstractSoftware development for sensor network is made difficult by resource constrained sensor devices, distributed system complexity, communication unreliability, and high labor cost. Simulation, as a useful tool, provides an affordable way to study algorithmic problems with flexibility and controllability. However, in exchange for speed simulation often trades detail that ultimately limits its utility. In this paper, we propose a new development paradigm, simulation-based augmented reality, in which simulation is used to enhance development on physical hardware by seamlessly integrating a running simulated network with a physical deployment in a way that is transparent to each. The advantages of such an augmented network include the ability to study a large sensor network with limited hardware and the convenience of studying a part of the physical network with simulation's debugging, profiling and tracing capabilities. We implement the augmented reality system based on a sensor network simulator with high fidelity and high scalability. Key to the design are "super" sensor nodes which are half virtual and half physical that interconnect simulation and physical network with fine-grained traffic forwarding and accurate time synchronization. Our results detail the overhead associated with integrating live and simulated networks and the timing accuracy between virtual and physical parts of the network. We also discuss various application scenarios for our system. Ye Wen, Wei Zhang 0118, Richard Wolski, Navraj Chohan |
SenSys | 3 |
| 2007 | QBETS: queue bounds estimation from time seriesabstractNo abstract available. Daniel Nurmi, John Brevik, Richard Wolski |
SIGMETRICS | 3 |
| 2007 | The GridSAT portal: a Grid Web-based portal for solving satisfiability problems using the national cyberinfrastructureabstractAbstract We present a Grid portal problem (which is accessible through http://orca.cs.ucsb.edu/sat_portal ) for solving Boolean satisfiability. The portal provides a simple and public interface to a sophisticated and complex Grid application—GridSAT—running on a large set of distributed computational resources hosted in different large‐scale national computing centers (i.e. the national cyberinfrastructure circa 2005). In this paper we describe the design goals of the portal and how it has influenced some of the application features. We also describe how the adaptive and self‐tuning features of Grid applications (written from first principles) make the portal simpler and easier to implement. Copyright © 2006 John Wiley & Sons, Ltd. Wahid Chrabakh, Richard Wolski |
Concurr. Comput. Pract. Exp. | 2 |
| 2007 | Special Issue Featuring Selected Papers from HPDC-15
Richard Wolski, Henri E. Bal |
J. Grid Comput. | 1 |
| 2006 | S2DB: a novel simulation-based debugger for sensor network applicationsabstractSensor network computing can be characterized as resource-constrained distributed computing using unreliable, low bandwidth communication. This combination of characteristics poses significant software development and maintenance challenges. Effective and efficient debugging tools for sensor network are thus critical. Existent development tools, such as TOSSIM, EmStar, ATEMU and Avrora, provide useful debugging support, but not with the fidelity, scale and functionality that we believe are sufficient to meet the needs of the next generation of applications.In this paper, we propose a debugger, called S2DB, based on a distributed full system sensor network simulator with high fidelity and scalable performance, DiSenS. By exploiting the potential of DiSenS as a scalable full system simulator, S2DB extends conventional debugging methods by adding novel device level, program source level, group level, and network level debugging abstractions. The performance evaluation shows that all these debugging features introduce overhead that is generally less than 10% into the simulator and thus making S2DB an efficient and effective debugging tool for sensor networks. Ye Wen, Richard Wolski, Selim Gurun |
EMSOFT | 2 |
| 2006 | Predicting bounds on queuing delay for batch-scheduled parallel machinesabstractMost space-sharing parallel computers presently operated by high-performance computing centers use batch-queuing systems to manage processor allocation. In many cases, users wishing to use these batch-queued resources have accounts at multiple sites and have the option of choosing at which site or sites to submit a parallel job. In such a situation, the amount of time a user's job will wait in any one batch queue can significantly impact the overall time a user waits from job submission to job completion. In this work, we explore a new method for providing end-users with predictions for the bounds on the queuing delay individual jobs will experience. We evaluate this method using batch scheduler logs for distributed-memory parallel machines that cover a 9-year period at 7 large HPC centers.Our results show that it is possible to predict delay bounds reliably for jobs in different queues, and for jobs requesting different ranges of processor counts. Using this information, scientific application developers can intelligently decide where to submit their parallel codes in order to minimize overall turnaround time. John Brevik, Daniel Nurmi, Richard Wolski |
PPoPP | 3 |
| 2006 | Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time predictionabstractLarge-scale distributed systems offer computational power at unprecedented levels. In the past, HPC users typically had access to relatively few individual supercomputers and, in general, would assign a one-to-one mapping of applications to machines. Modern HPC users have simultaneous access to a large number of individual machines and are beginning to make use of all of them for single-application execution cycles. One method that application developers have devised in order to take advantage of such systems is to organize an entire application execution cycle as a workflow. The scheduling of such workflows has been the topic of a great deal of research in the past few years and, although very sophisticated algorithms have been devised, a very specific aspect of these distributed systems, namely that most supercomputing resources employ batch queue scheduling software, has heretofore been omitted from consideration, presumably because it is difficult to model accurately. In this work, we augment an existing workflow scheduler through the introduction of methods which make accurate predictions of both the performance of the application on specific hardware, and the amount of time individual workflow tasks will spend waiting in batch queues. Our results show that although a workflow scheduler alone may choose correct task placement based on data locality or network connectivity, this benefit is often compromised by the fact that most jobs submitted to current systems must wait in overcommited batch queues for a significant portion of time. However, incorporating the enhancements we describe improves workflow execution time in settings where batch queues impose significant delays on constituent workflow tasks. Daniel Nurmi, Anirban Mandal, John Brevik, Charles Koelbel, Richard Wolski, Ken Kennedy |
SC | 5 |
| 2006 | GridSAT: Design and Implementation of a Computational Grid Application
Wahid Chrabakh, Richard Wolski |
J. Grid Comput. | 2 |
| 2006 | GridSAT: a system for solving satisfiability problems using a computational grid
Wahid Chrabakh, Richard Wolski |
Parallel Comput. | 2 |
| 2005 | Minimizing the Network Overhead of Checkpointing in Cycle-harvesting Cluster EnvironmentsabstractCycle-harvesting systems such as Condor have been developed to make desktop machines in a local area (which are often similar to clusters in hardware configuration) available as a compute platform. To provide a dual-use capability, opportunistic jobs harvesting cycles from the desktop must be checkpointed before the desktop resources are reclaimed by their owners and the job is evacuated. In this paper, we investigate a new system for computing efficient checkpoint schedules in cycle-harvesting environments. Our system records the historical availability from each resource and fits a statistical model to the observations. Because checkpointing must often traverse the network (i.e. the desktop hosts do not provide sufficient persistent storage for checkpoints), we combine this model with predictions of network performance to the storage site to compute a checkpoint schedule. When an application is initiated on a particular resource, the system uses the computed distribution to parameterize a Markov state-transition model for the application's execution, evaluates the expected time and network overhead as a function of the checkpoint interval, and numerically optimizes with respect to time. We report on the performance of and implementation of this system using the Condor cycle-harvesting environment at the University of Wisconsin. We also evaluate the efficiencies we achieve for a variety of network overheads using trace-based simulation. Finally, we validate our simulations against the observed performance with Condor. Our results indicate that while the choice of model distribution has a relatively small but positive effect on time efficiency, it has a substantial impact on network utilization Daniel Nurmi, John Brevik, Richard Wolski |
CLUSTER | 3 |
| 2005 | Modeling Machine Availability in Enterprise and Wide-Area Distributed Computing Environments
Daniel Nurmi, John Brevik, Richard Wolski |
Euro-Par | 3 |
| 2005 | Quorum: Flexible Quality of Service for Internet Services
Josep M. Blanquer, Antoni Batchelli, Klaus E. Schauser, Richard Wolski |
NSDI | 4 |
| 2004 | Automatic methods for predicting machine availability in desktop Grid and peer-to-peer systemsabstractIn this paper we examine the problem of predicting machine availability in desktop and enterprise computing environments. Predicting the duration that a machine will run until it restarts (availability duration) is critically useful to application scheduling and resource characterization in federated systems. We describe one parametric model fitting technique and two nonparametric prediction techniques, comparing their accuracy in predicting the quantiles of empirically observed machine availability distributions. We describe each method analytically and evaluate its precision using a synthetic trace of machine availability constructed from a known distribution. To detail their practical efficacy, we apply them to machine availability traces from three separate desktop and enterprise computing environments, and evaluate each method in terms of the accuracy with which it predicts availability in a trace driven simulation. Our results indicate that availability duration can be predicted with quantifiable confidence bounds and that these bounds can he used as conservative bounds on lifetime predictions. Moreover a nonparametric method based on a binomial approach generates the most accurate estimates. John Brevik, Daniel Nurmi, Richard Wolski |
CCGRID | 3 |
| 2004 | Application-level prediction of battery dissipationabstractMobile, battery-powered devices such as personal digital assistants and web-enabled mobile phones have successfully emerged as new access points to the world's digital infrastructure. However, the growing gap between device capabilities and battery technology requires novel techniques that extend battery life. Key to the success of such techniques, is our ability to accurately predict the power consumption of a program.In this paper, we investigate the degree to which battery dissipation induced by program execution can be measured by applica-tion-level software tools and predicted by a compiler and runtime system. We present a novel technique with which we can accurately estimate whole-program power-consumption for an arbitrary program by composing battery dissipation rates of benchmarks. We empirically evaluate our technique using an iPAQ hand-held device and a number of MiBench and other programs. Chandra Krintz, Ye Wen, Richard Wolski |
ISLPED | 3 |
| 2004 | NWSLite: A Light-Weight Prediction Utility for Mobile DevicesabstractComputation off-loading, i.e., remote execution, has been shown to be effective for extending the computational power and battery life of resource-restricted devices, e.g., hand-held, wearable, and pervasive computers. Remote execution systems must predict the cost of executing both locally and remotely to determine when off-loading will be most beneficial. These costs however, are dependent upon the execution behavior of the task being considered and the highly-variable performance of the underlying resources, e.g., CPU (local and remote), bandwidth, and network latency. As such, remote execution systems must employ sophisticated, prediction techniques that accurately guide computation off-loading. Moreover, these techniques must be efficient, i.e., they cannot consume significant resources, e.g., energy, execution time, etc., since they are performed on the mobile device.In this paper, we present NWSLite, a computationally efficient, highly accurate prediction utility for mobile devices. NWSLite is an extension to the Network Weather Service (NWS), a dynamic forecasting toolkit for adaptive scheduling of high-performance Computational Grid applications. We significantly scaled down the NWS to reduce its resource consumption yet still achieve accuracy that exceeds that of extant remote execution prediction methods. We empirically analyze and compare both the prediction accuracy and the cost of NWSLite and a number of different forecasting methods from existing remote execution systems. We evaluate the efficacy of the different methods using a wide range of mobile applications and resources. Selim Gurun, Chandra Krintz, Richard Wolski |
MobiSys | 3 |
| 2003 | The Livny and Plank-Beck Problems: Studies in Data Movement on the Computational GridabstractOver the last few years the Grid Computing research community has become interested in developing data intensive applications for the Grid. These applications face significant challenges because their widely distributed nature makes it difficult to access data with reasonable speed. In order to address this problem, we feel that the Grid community needs to develop and explore data movement challenges that represent problems encountered in these applications. In this paper, we will identify two such problems that we have dubbed the Livny Problem and the Plank-Beck Problem. We will also present data movement scheduling techniques that we have developed to address these problems. Matthew S. Allen, Richard Wolski |
SC | 2 |
| 2003 | GridSAT: A Chaff-based Distributed SAT Solver for the GridabstractWe present GridSAT, a parallel and complete satisfiability solver designed to solve non-trivial SAT problem instances using a large number of widely distributed and heterogeneous resources. The GridSAT parallel algorithm uses intelligent backtracking, distributed and carefully scheduled sharing of learned clauses, and clause reduction. Our implementation focuses on dynamic resource acquisition and release to optimize application execution. We show how the large number of computational resources that are available from a Grid can be managed effectively for the application by an automatic scheduler and effective implementation. GridSAT execution speed is compared against the best sequential solver as rated by the SAT2002 competition using a wide variety of problem instances. The results show that GridSAT delivers speed-up for all but one of the test problem instances that are of significant size. In addition, we describe how GridSAT has solved previously unsolved satisfiability problems and the domain science contribution these results make. Wahid Chrabakh, Richard Wolski |
SC | 2 |
| 2003 | The Internet Backplane Protocol: a study in resource sharing
Alessandro Bassi, Micah D. Beck, Terry Moore, James S. Plank, D. Martin Swany, Richard Wolski, Graham E. Fagg |
Future Gener. Comput. Syst. | 6 |
| 2003 | Guest editor introduction: special issue on Computational Grids
Jon B. Weissman, Richard Wolski |
J. Parallel Distributed Comput. | 2 |
| 2003 | Adaptive Computing on the Grid Using AppLeSabstractEnsembles of distributed, heterogeneous resources, also known as computational grids, have emerged as critical platforms for high-performance and resource-intensive applications. Such platforms provide the potential for applications to aggregate enormous bandwidth, computational power, memory, secondary storage, and other resources during a single execution. However, achieving this performance potential in dynamic, heterogeneous environments is challenging. Recent experience with distributed applications indicates that adaptivity is fundamental to achieving application performance in dynamic grid environments. The AppLeS (Application Level Scheduling) project provides a methodology, application software, and software environments for adaptively scheduling and deploying applications in heterogeneous, multiuser grid environments. We discuss the AppLeS project and outline our findings. Francine Berman, Richard Wolski, Henri Casanova, Walfredo Cirne, Holly Dail, Marcio Faerman, Silvia M. Figueira, Jim Hayes, Graziano Obertelli, Jennifer M. Schopf, Gary Shao, Shava Smallen, Neil Spring, Alan Su 0001, Dmitrii Zagorodnov |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2002 | The Internet Backplane Protocol: A Study in Resource SharingabstractIn this work we present the Internet Backplane Protocol (IBP), a middleware created to allow the sharing of storage resources, implemented as part of the network fabric. IBP allows an application to control intermediate data staging operations explicitly. As IBP follows a very simple philosophy, very similar to the Internet Protocol, and the resulting semantic might be too weak for some applications, we introduce the exNode, a data structure that aggregates storage allocations on the Internet. Alessandro Bassi, Micah D. Beck, Graham E. Fagg, Terry Moore, James S. Plank, D. Martin Swany, Richard Wolski |
CCGRID | 7 |
| 2002 | Representing Dynamic Performance Information in Grid Environments with the Network Weather ServiceabstractIn this paper, we discuss requirements for integrating dynamic performance information from the Network Weather Service (NWS) into the Grid Information Service infrastructure (GIS). We describe the object model that NWS uses internally and provide some rationale for its structure. Finally, we present the NWS 's implementation of a caching LDAP daemon that integrates NWS information into the reference GIS -the Glob s MDS. D. Martin Swany, Richard Wolski |
CCGRID | 2 |
| 2002 | Adaptive Timeout Discovery Using the Network Weather ServiceabstractIn this paper we present a novel methodology for improving the performance and dependability of application-level messaging in Grid systems. Based on the Network Weather Service, our system uses nonparametric statistical forecasts of request-response times to automatically determine message timeouts. By choosing a timeout based on predicted network performance, the methodology improves application and Grid service performance as extraneous and overly-long timeouts are avoided. We describe the technique, the additional execution and programming overhead it introduces, and demonstrate the effectiveness using a wide-area test application. Matthew S. Allen, Richard Wolski, James S. Plank |
HPDC | 2 |
| 2002 | Multivariate resource performance forecasting in the network weather serviceabstractThis paper describes a new technique in the Network Weather Service for producing multi-variate forecasts. The new technique uses the NWS’s univariate forecasters and emprically gathered Cumulative Distribution Functions (CDFs) to make predictions from correlated measurement streams. Experimental results are shown in which throughput is predicted for long TCP/IP transfers from short NWS network probes. D. Martin Swany, Richard Wolski |
SC | 2 |
| 2002 | Middleware for the use of storage in communication
Micah D. Beck, Dorian C. Arnold, Alessandro Bassi, Francine Berman, Henri Casanova, Jack J. Dongarra, Terry Moore, Graziano Obertelli, James S. Plank, D. Martin Swany, Sathish S. Vadhiyar, Richard Wolski |
Parallel Comput. | 12 |
| 2001 | Data Staging Effects in Wide Area Task Farming ApplicationsabstractRecent advances in computing and communication have given rise to the computational grid notion (I. Foster and C. Kesselman, 1998). The core of this computing paradigm is the design of a system for drawing compute power from a confederation of geographically dispersed heterogeneous resources, seamlessly and ubiquitously. If high-performance levels are to be achieved, data locality must be identified and managed. We consider the effect of server side staging on the behavior of a class of wide area "task farming" or parameter sweep applications. We show that staging improves task throughput mainly through the increased parallelism rather than the reduction in overall turnaround time per task. We derive a model for farming applications with and without server side staging and verify the model through live experiments as well as simulations. Wael R. Elwasif, James S. Plank, Richard Wolski |
CCGRID | 3 |
| 2001 | NwsAlarm: A Tool for Accurately Detecting Resource Performance DegradationabstractEnd-users of high-performance computing resources have come to expect that consistent levels of performance be delivered to their applications. The advancement of the computational grid enables the seamless use of a multitude of computing resources by these users. The combination of these developments has generated a need for users to monitor the end-to-end-performance available to an application. In addition, tools are needed to alert users of degradation in expected performance. We present the NwsAlarm, a Java-based utility that enables users to monitor performance levels of any resource being monitored by the Network Weather Service. The NwsAlarm is invoked by a user without special privileges with a simple click on a Web page link. More importantly the NwsAlarm allows any user of the NwsAlarm to register and set expected performance levels. When performance levels fall below these thresholds, the registered administrators are immediately notified via email. The NwsAlarm uses prediction of performance measurements to filter false alarm values. We exemplify the importance of and accuracy achieved by the NwsAlarm with real examples of performance degradation caused by routing table changes and loss of service on the Abilene, Internet-2 research network used for experimentation with evolving Grid software technology. On average, 92% fewer false alarms are raised by the NwsAlarm than if raw measurements are used. Chandra Krintz, Richard Wolski |
CCGRID | 2 |
| 2001 | The Logistical Session LayerabstractThe Logistical Session Layer is a system to enable enhanced functionality to distributed programming systems. The term Logistical refers to the fact that we enhance the traditional client-server model to allow for intermediate systems which are neither. This system generalizes the notion of caches but represents a cleaner architecture in that it explicitly declares itself to be a session layer protocol. D. Martin Swany, Richard Wolski |
HPDC | 2 |
| 2001 | G-commerce: Market Formulations Controlling Resource Allocation on the Computational GridabstractIn this paper we investigate G-commerce-computational economies for controlling resource allocation in Computational Grid settings. We define hypothetical resource consumers (representing users and Grid-aware applications) and resource producers (representing resource owners who "sell" their resources to the Grid). We then measure the efficiency of resource allocation under two different market conditions: commodities markets and auctions. We compare both market strategies in terms of price stability, market equilibrium, consumer efficiency, and producer efficiency. Our results indicate that commodities markets are a better choice for controlling Grid resources than previously defined auction strategies. Richard Wolski, James S. Plank, John Brevik, Todd Bryan |
IPDPS | 1 |
| 2001 | The Effect of Timeout Prediction and Selection on Wide Area Collective OperationsabstractFailure identification is a fundamental operation concerning exceptional conditions that network programs must be able to perform. In this paper, we explore the use of timeouts to perform failure identification at the application level. We evaluate the use of static timeouts and of dynamic timeouts based on forecasts using the Network Weather Service. For this evaluation, we perform experiments on a wide-area collection of 31 machines distributed in eight institutions. Though the conclusions are limited to the collection of machines used, we observe that a single static timeout is not reasonable, even for a collection of similar machines over time. Dynamic timeouts perform roughly as well as the best static timeouts and, more importantly, they provide a single methodology for timeout determination that should be effective for wide-area applications. James S. Plank, Matthew S. Allen, Richard Wolski |
NCA | 3 |
| 2001 | Data Logistics in Network Computing: The Logistical Session LayerabstractPresents a strategy for optimizing end-to-end TCP/IP throughput over long-haul networks (i.e. those where the product of the bandwidth and the delay is high.) Our approach defines a logistical session layer (LSL) that uses intermediate process-level "depots" along the network route from source to sink to implement an end-to-end communication session. Despite the additional processing overhead resulting from TCP/IP protocol stack-Unix kernel boundary traversals at each depot, our experiments show that dramatic end-to-end bandwidth improvements are possible. We also describe a prototype implementation of LSL that does not require any Unix kernel modification or root access privilege, which we used to generate the results, and we discuss its utility in the context of extant TCP/IP tuning methodologies. D. Martin Swany, Richard Wolski |
NCA | 2 |
| 2001 | Using JavaNws to compare C and Java TCP-Socket performanceabstractAbstract As research and implementation continue to facilitate high‐performance computing in Java, applications can benefit from resource management and prediction tools. In this work, we present such a tool for network round‐trip time and bandwidth between a user's desktop and any machine running a Web server (this assumes that the user's browser is capable of supporting Java 1.1 and above). JavaNws is a Java implementation and extension of a powerful subset of the Network Weather Service (NWS), a performance prediction toolkit that dynamically characterizes and forecasts the performance available to an application. However, due to the Java language implementation and functionality (portability, security, etc.), it is unclear whether a Java program is able to measure and predict the network performance experienced by C‐applications with the same accuracy as an equivalent C program. We provide a quantitative equivalence study of the Java and C TCP‐socket interface and show that the data collected by the JavaNws is as predictable as that collected by the NWS (using C). Copyright © 2001 John Wiley & Sons, Ltd. Chandra Krintz, Richard Wolski |
Concurr. Comput. Pract. Exp. | 2 |
| 2001 | Writing Programs that Run EveryWare on the Computational GridabstractThe Computational Grid has been proposed, for the implementation of high-performance applications using widely dispersed computational resources. The goal of a Computational Grid is to aggregate ensembles of shared, heterogeneous, and distributed resources (potentially controlled by separate organizations) to provide computational, "power" to an application program. We provide a toolkit for the development of globally deployable Grid applications. The toolkit, called EveryWare, enables an application to draw computational power transparently from the Grid. It consists of a portable set of processes and libraries that can be incorporated into an application so that a wide variety of dynamically changing distributed infrastructures and resources can be used together to achieve supercomputer-like performance. We provide our experiences gained while building the EveryWare toolkit prototype and an explanation of its use in implementing a large-scale Grid application. Richard Wolski, John Brevik, Graziano Obertelli, Neil Spring, Alan Su 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2000 | Synchronizing Network Probes to Avoid Measurement Intrusiveness with the Network Weather ServiceabstractPresents a scalable protocol for conducting periodic probes of network performance in a way that minimizes collisions between separate probes. The goal of the protocol is to enable active performance monitoring of large-scale distributed computational systems and networks. We use the protocol to generate time series of measurement data that are then exposed to numerical forecasting models when a prediction of network performance is required. We present the protocol and demonstrate its effectiveness using the Network Weather Service -a tool for dynamically predicting network, CPU, memory and storage performance. Richard Wolski, Benjamin Gaidioz, Bernard Tourancheau |
HPDC | 1 |
| 2000 | The AppLeS Parameter Sweep Template: User-Level Middleware for the GridabstractThe Computational Grid is a promising platform for the efficient execution of parameter sweep applications over large parameter spaces. To achieve performance on the Grid, such applications must be scheduled so that shared data files are strategically placed to maximize reuse, and so that the application execution can adapt to the deliverable performance potential of target heterogeneous, distributed and shared resources. Parameter sweep applications are an important class of applications and would greatly benefit from the development of Grid middleware that embeds a scheduler for performance and targets Grid resources transparently. In this paper we describe a user-level Grid middleware project, the AppLeS Parameter Sweep Template (APST), that uses application-level scheduling techniques [1] and various Grid technologies to allow the efficient deployment of parameter sweep applications over the Grid. We discuss several possible scheduling algorithms and detail our software design. We then describe our current implementation of APST using systems like Globus [2], NetSolve [3] and the Network Weather Service [4], and present experimental results. Henri Casanova, Graziano Obertelli, Francine Berman, Richard Wolski |
SC | 4 |
| 1999 | Predicting the CPU Availability of Time-shared Unix Systems on the Computational GridabstractFocuses on the problem of making short- and medium-term forecasts of CPU availability on time-shared Unix systems. We evaluate the accuracy with which availability can be measured using the Unix load average, the Unix utility "vmstat" and the Network Weather Service (NWS) CPU sensor that uses both. We also examine the autocorrelation between successive CPU measurements to determine their degree of self-similarity. While our observations show a long-range autocorrelation dependence, we demonstrate how this dependence manifests itself in the short- and medium-term predictability of the CPU resources in our study. Richard Wolski, Neil Spring, Jim Hayes |
HPDC | 1 |
| 1999 | Adaptive Performance Prediction for Distributed Data-Intensive ApplicationsabstractThe computational grid is becoming the platform of choice for large-scale distributed data-intensive applications. Accurately predicting the transfer times of remote data les, a fundamental component of such applications, is critical to achieving application performance. In this paper, we introduce a performance prediction method, ARM (Adaptive Regression Modeling), to determine data transfer times for network-bound distributed dataintensive applications. We demonstrate the eectiveness of the ARM method on two distributed data applications, SARA (Synthetic Aperture Radar Atlas) and SRB (Storage Resource Broker) , and discuss how it can be used for application scheduling. Our experiments demonstrate that applying the ARM method to these applications predicted data transfer times in wide-area multi-user grid environments with accuracy of 88% or better. 1 Introduction Ensembles of distributed computational, storage, and other resources, also known as computational grids [12, 14], are... Marcio Faerman, Alan Su 0001, Richard Wolski, Francine Berman |
SC | 3 |
| 1999 | A Network Performance Tool for Grid EnvironmentsabstractIn grid computing environments, network bandwidth discovery and allocation is a serious issue.Before their applications are running, grid users will need to choose hosts based on available bandwidth.Running applications may need to adapt to a changing set of hosts.Hence, a tool is needed for monitoring network performance that is integral to the grid environment.To address this need, Gloperf was developed as part of the Globus grid computing toolkit.Gloperf is designed for ease of deployment and makes simple, end-to-end TCP measurements requiring no special host permissions.Scalability is addressed by a hierarchy of measurements based on group membership and by limiting overhead to a small, acceptable, fixed percentage of the available bandwidth.Since this fixed overhead may push host-pair revisit time into the tens-of-hours, we also quantitatively examine the "trajectory" of the cost-error trade-off for measurement frequency. Craig A. Lee, James Stepanek, Richard Wolski, Carl Kesselman, Ian T. Foster |
SC | 3 |
| 1999 | Running EveryWare on the Computational GridabstractThe Computational Grid [10] has recently been proposed for the implementation of high-performance applications using widely dispersed computational resources. The goal of a Computational Grid is to aggregate ensembles of shared, heterogeneous, and distributed resources (potentially controlled by separate organizations) to provide computational "power" to an application program. In this paper, we provide a toolkit for the development of Grid applications. The toolkit, called EveryWare, enables an application to draw computational power transparently from the Grid. The toolkit consists of a portable set of processes and libraries that can be incorporated into an application so that a wide variety of dynamically changing distributed infrastructures and resources can be used together to achieve supercomputer-like performance. We provide our experiences gained while building the EveryWare toolkit prototype and the first true Grid application. 1 Introduction Increasingly, the high-perform... Richard Wolski, John Brevik, Chandra Krintz, Graziano Obertelli, Neil Spring, Alan Su 0001 |
SC | 1 |
| 1999 | Logistical quality of service in NetSolve
Micah D. Beck, Henri Casanova, Jack J. Dongarra, Terry Moore, James S. Plank, Francine Berman, Richard Wolski |
Comput. Commun. | 7 |
| 1999 | The network weather service: a distributed resource performance forecasting service for metacomputing
Richard Wolski, Neil Spring, Jim Hayes |
Future Gener. Comput. Syst. | 1 |
| 1998 | Application Level Scheduling of Gene Sequence Comparison on MetacomputersabstractArticle Application level scheduling of gene sequence comparison on metacomputers Share on Authors: Neil Spring View Profile , Rich Wolski View Profile Authors Info & Claims ICS '98: Proceedings of the 12th international conference on SupercomputingJuly 1998 Pages 141–148https://doi.org/10.1145/277830.277860Published:13 July 1998 41citation266DownloadsMetricsTotal Citations41Total Downloads266Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Neil Spring, Richard Wolski |
International Conference on Supercomputing | 2 |
| 1997 | Forecasting Network Performance to Support Dynamic Scheduling using the Network Weather ServiceabstractThe Network Weather Service is a generalizable and extensible facility designed to provide dynamic resource performance forecasts in metacomputing environments. In this paper, we outline its design and detail the predictive performance of the forecasts it generates. While the forecasting methods are general, we focus on their ability to predict the TCP/IP end-to-end throughput and latency that is attainable by an application using systems located at different sites. Such network forecasts are needed both to support scheduling, and by the metacomputing software infrastructure to develop quality-of-service guarantees. Richard Wolski |
HPDC | 1 |
| 1997 | Implementing a Performance Forecasting System for Metacomputing The Network Weather ServiceabstractIn this paper we describe the design and implementation of a system called the Network Weather Service (NWS) that takes periodic measurements of deliverable resource performance from distributed networked resources, and uses numerical models to dynamically generate forecasts of future performance levels. These performance forecasts, along with measures of performance fluctuation (e.g the mean square prediction error) and forecast lifetime that the NWS generates, are made available to schedulers and other resource management mechanisms at runtime so that they may determine the quality-of-service that will be available from each resource. We describe the architecture of the NWS and implementations that we have developed and are currently deploying for the Legion [13] and Globus/Nexus [7] metacomputing infrastructures. We also detail NWS forecasts of resource performance using both the Legion and Globus/Nexus implementations. Our results show that simple forecasting techniques substantially outperform measurements of current conditions (commonly used to gauge resource availability and load) in terms of prediction accuracy. In addition, the techniques we have employed are almost as accurate as substantially more complex modeling methods. We compare our techniques to a sophisticated time-series analysis system in terms of forecasting accuracy and computational complexity. Richard Wolski, Neil Spring, Chris Peterson 0001 |
SC | 1 |
| 1996 | Scheduling from the Perspective of the ApplicationabstractMetacomputing is the aggregation of distributed and high-performance resources on coordinated networks. With careful scheduling, resource-intensive applications can be implemented efficiently on metacomputing systems at the sizes of interest to developers and users. In this paper, we focus on the problem of scheduling applications on metacomputing systems. We introduce the concept of application-centric scheduling in which everything about the system is evaluated in terms of its impact on the application. Application-centric scheduling is used by virtually all metacomputer programmers to achieve performance on metacomputing systems. We describe two successful metacomputing applications to illustrate this approach, and describe AppLeS (Application-Level Scheduling) agents which generalize the application-centric scheduling approach. Finally, we show preliminary results which compare AppLeS-derived schedules with conventional strip and blocked schedules for a 2D Jacobi code. Francine Berman, Richard Wolski |
HPDC | 2 |
| 1996 | Application-Level Scheduling on Distributed Heterogeneous NetworksabstractHeterogeneous networks are increasingly being used as platforms for resource-intensive distributed parallel applications. A critical contributor to the performance of such applications is the scheduling of constituent application tasks on the network. Since often the distributed resources cannot be brought under the control of a single global scheduler, the application must be scheduled by the user. To obtain the best performance, the user must take into account both application-specific and dynamic system information in developing a schedule which meets his or her performance criteria. In this paper, we define a set of principles underlying application-level scheduling and describe our work-in-progress building AppLeS (application-level scheduling) agents. We illustrate the application-level scheduling approach with a detailed description and results for a distributed 2D Jacobi application on two production heterogeneous platforms. Francine Berman, Richard Wolski, Silvia M. Figueira, Jennifer M. Schopf, Gary Shao |
SC | 2 |
| 1995 | Time Sharing Massively Parallel Machines
Brent C. Gorda, Richard Wolski |
ICPP (2) | 2 |
| 1993 | Program Partitioning for NUMA Multiprocessor Computer Systems
Richard Wolski, John Feo |
J. Parallel Distributed Comput. | 1 |