John Brevik

dblp:89/4534 · also John O. Brevik · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
0since 2021 · last 2019
0000-0001-5468-2463ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 2 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Cloud and datacenter computing · 69% Performance modeling and evaluation · 16% High-performance computing · 10%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cluster resource management and scheduling
0.542017
Probabilistic guarantees of execution duration for Amazon spot instances · SC 2017
VARQ: virtual advance reservations for queues · HPDC 2008
QBETS: queue bounds estimation from time series · SIGMETRICS 2007
Cloud and datacenter computing › job scheduling
batch scheduling
0.232008
Probabilistic advanced reservations for batch-scheduled parallel machines · PPoPP 2008
VARQ: virtual advance reservations for queues · HPDC 2008
Predicting bounds on queuing delay for batch-scheduled parallel machines · PPoPP 2006
High-performance computing
queue wait time prediction
0.232008
VARQ: virtual advance reservations for queues · HPDC 2008
QBETS: queue bounds estimation from time series · SIGMETRICS 2007
Predicting bounds on queuing delay for batch-scheduled parallel machines · PPoPP 2006
Performance modeling and evaluation › workload characterization
workload modeling
0.212014
Using Parametric Models to Represent Private Cloud Workloads · IEEE Trans. Serv. Comput. 2014
Cloud and datacenter computing › resource allocation › resource allocation policy
advance reservation
0.222008
Probabilistic advanced reservations for batch-scheduled parallel machines · PPoPP 2008
VARQ: virtual advance reservations for queues · HPDC 2008
Performance modeling and evaluation
performance prediction
0.122007
Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time prediction · SC 2006
QBETS: queue bounds estimation from time series · SIGMETRICS 2007
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.112008
Probabilistic advanced reservations for batch-scheduled parallel machines · PPoPP 2008
Performance modeling and evaluation
queueing models
0.112006
Predicting bounds on queuing delay for batch-scheduled parallel machines · PPoPP 2006
Cloud and datacenter computing
workflow scheduling
0.112006
Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time prediction · SC 2006
Cloud and datacenter computing › cloud deployment
private clouds
0.112014
Using Parametric Models to Represent Private Cloud Workloads · IEEE Trans. Serv. Comput. 2014
Distributed systems
grid computing
0.122001
Writing Programs that Run EveryWare on the Computational Grid · IEEE Trans. Parallel Distributed Syst. 2001
Running EveryWare on the Computational Grid · SC 1999
Parallel and multicore computing
parallel scheduling
0.012008
Probabilistic advanced reservations for batch-scheduled parallel machines · PPoPP 2008
Distributed systems › distributed resource management
distributed resource aggregation
0.011999
Running EveryWare on the Computational Grid · SC 1999
Performance modeling and evaluation › statistical analysis
time series analysis
0.012007
QBETS: queue bounds estimation from time series · SIGMETRICS 2007
Parallel and multicore computing
parallel programming models
0.012001
Writing Programs that Run EveryWare on the Computational Grid · IEEE Trans. Parallel Distributed Syst. 2001
Distributed systems
resource sharing
0.012001
Writing Programs that Run EveryWare on the Computational Grid · IEEE Trans. Parallel Distributed Syst. 2001

Methods — techniques the papers use, named apart from their topics

probabilistic modeling · 0.4workload trace analysis · 0.3hyperexponential distribution · 0.2estimation maximization · 0.2virtual advance reservations · 0.1time series analysis · 0.1statistical prediction from scheduler logs · 0.1performance modeling · 0.1middleware toolkit · 0.0portable processes and libraries · 0.0
YearPublicationVenuePosition
2019 Analyzing AWS Spot Instance Pricing
abstract
Many cloud computing vendors offer a preemptible class of service for rented virtual machines. In November 2017, Amazon.com changed the pricing mechanism for its preemptible "spot instances" so that prices would change more "smoothly." This paper analyzes the effect of this change on spot instance prices. It examines the prices immediately before and after the mechanism change to determine the extent to which prices themselves changed. It then compares the 90-day period immediately after the change in mechanism to the next 90-day period. Finally, it compares the two most recent 90-day periods (ending on October 15, 2018). Our results indicate that in addition to smoothing prices, the mechanism change introduced generally higher prices which is a trend that continues.
Gareth George, Richard Wolski, Chandra Krintz, John Brevik
IC2E4
2017 QPRED: Using Quantile Predictions to Improve Power Usage for Private Clouds
abstract
In this paper we describe a new, efficient predictive scheduling methodology for implementing computing infrastructure power savings using private clouds. Our approach, termed "QPRED," estimates the quantiles on the distribution of future machine usage so that unneeded machines may be powered down to save power. A cloud administrator sets a bound on the probability that all available machines will be powered down when a cloud request arrives. This target probability is the basis of a Service Level Agreement between the cloud administrator and all cloud users covering start-up delay resulting from power savings. Our results, validated using activity traces from several private clouds used in commercial production, indicate that QPRED successfully reduces power consumption substantially while maintaining the SLAs specified by the cloud administrator.
Richard Wolski, John Brevik
CLOUD2
2017 Probabilistic guarantees of execution duration for Amazon spot instances
abstract
In this paper we propose DrAFTS - a methodology for implementing probabilistic guarantees of instance reliability in the Amazon Spot tier. Amazon offers "unreliable" virtual machine instances (ones that may be terminated at any time) at a potentially large discount relative to "reliable" On-demand and Reserved instances. Our method predicts the "bid values" that users can specify to provision Spot instances which ensure at least a fixed duration of execution with a given probability. We illustrate the method and test its validity using Spot pricing data post facto, both randomly and using real-world workload traces. We also test the efficacy of the method experimentally by using it to launch Spot instances and then observing the instance termination rate. Our results indicate that it is possible to obtain the same level of reliability from unreliable instances that the Amazon service level agreement guarantees for reliable instances with a greatly reduced cost.
Richard Wolski, John Brevik, Ryan Chard, Kyle Chard
SC2
2014 Using Parametric Models to Represent Private Cloud Workloads
abstract
Cloud computing has become a popular metaphor for dynamic and secure self-service access to computational and storage capabilities. In this study, we analyze and model workloads gathered from enterprise-operated commercial private clouds that implement “Infrastructure as a Service.” Our results show that 3-phase hyperexponential distributions fit using the Estimation Maximization (E-M) algorithm capture workload attributes accurately. In addition, these models of individual attributes compose to produce estimates of overall cloud performance that our results verify to be accurate. As an early study of commercial enterprise private clouds, this work provides guidance to those researching, designing, or maintaining such installations. In particular, the cloud workloads under study do not exhibit “heavy-tailed” distributional properties in the same way that “bare metal” operating systems do, potentially leading to different design and engineering tradeoffs.
Richard Wolski, John Brevik
IEEE Trans. Serv. Comput.2
2008 VARQ: virtual advance reservations for queues
abstract
In high-performance computing (HPC) settings, in which multiprocessor machines are shared among users with potentially competing resource demands, processors are allocated to user workload using space sharing. Typically, users interact with a given machine by submitting their jobs to a centralized batch scheduler that implements a site-specific policy designed to maximize machine utilization while providing tolerable turn-around times. To these users, the functioning of the batch scheduler and the policies it implements are both critical operating system components since they control how each job is serviced. In practice, while most HPC systems experience good utilization levels, the amount of time experienced by individual jobs waiting to begin execution has been shown to be highly variable and difficult to predict, leading to user confusion and/or frustration.
Daniel Nurmi, Richard Wolski, John Brevik
HPDC3
2008 Probabilistic advanced reservations for batch-scheduled parallel machines
abstract
No abstract available.
Daniel Nurmi, Richard Wolski, John Brevik
PPoPP3
2007 An Analysis of Availability Distributions in Condor
abstract
In this article, we investigate the dynamics exhibited by the production Condor pool at the University of Wisconsin with the goal of understanding its distributional properties. Condor is a cycle-harvesting service originally designed to launch and control "guest" user jobs (in batch mode) on idle workstations. Since its inception in 1985, however, it has expanded to include the ability to run in dedicated mode on clusters, to "glide in" to systems that are not strictly dedicated to Condor, and to "flock" jobs from one site to another based on pre-determined service level agreements (SLAs). Thus it has developed from an enterprise-wide desktop system into a full-fledged global computing infrastructure over its lifetime.
Richard Wolski, Daniel Nurmi, John Brevik
IPDPS3
2007 QBETS: Queue Bounds Estimation from Time Series
Daniel Nurmi, John Brevik, Richard Wolski
JSSPP2
2007 QBETS: queue bounds estimation from time series
abstract
No abstract available.
Daniel Nurmi, John Brevik, Richard Wolski
SIGMETRICS2
2006 Predicting bounds on queuing delay for batch-scheduled parallel machines
abstract
Most space-sharing parallel computers presently operated by high-performance computing centers use batch-queuing systems to manage processor allocation. In many cases, users wishing to use these batch-queued resources have accounts at multiple sites and have the option of choosing at which site or sites to submit a parallel job. In such a situation, the amount of time a user's job will wait in any one batch queue can significantly impact the overall time a user waits from job submission to job completion. In this work, we explore a new method for providing end-users with predictions for the bounds on the queuing delay individual jobs will experience. We evaluate this method using batch scheduler logs for distributed-memory parallel machines that cover a 9-year period at 7 large HPC centers.Our results show that it is possible to predict delay bounds reliably for jobs in different queues, and for jobs requesting different ranges of processor counts. Using this information, scientific application developers can intelligently decide where to submit their parallel codes in order to minimize overall turnaround time.
John Brevik, Daniel Nurmi, Richard Wolski
PPoPP1
2006 Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time prediction
abstract
Large-scale distributed systems offer computational power at unprecedented levels. In the past, HPC users typically had access to relatively few individual supercomputers and, in general, would assign a one-to-one mapping of applications to machines. Modern HPC users have simultaneous access to a large number of individual machines and are beginning to make use of all of them for single-application execution cycles. One method that application developers have devised in order to take advantage of such systems is to organize an entire application execution cycle as a workflow. The scheduling of such workflows has been the topic of a great deal of research in the past few years and, although very sophisticated algorithms have been devised, a very specific aspect of these distributed systems, namely that most supercomputing resources employ batch queue scheduling software, has heretofore been omitted from consideration, presumably because it is difficult to model accurately. In this work, we augment an existing workflow scheduler through the introduction of methods which make accurate predictions of both the performance of the application on specific hardware, and the amount of time individual workflow tasks will spend waiting in batch queues. Our results show that although a workflow scheduler alone may choose correct task placement based on data locality or network connectivity, this benefit is often compromised by the fact that most jobs submitted to current systems must wait in overcommited batch queues for a significant portion of time. However, incorporating the enhancements we describe improves workflow execution time in settings where batch queues impose significant delays on constituent workflow tasks.
Daniel Nurmi, Anirban Mandal, John Brevik, Charles Koelbel, Richard Wolski, Ken Kennedy
SC3
2005 Minimizing the Network Overhead of Checkpointing in Cycle-harvesting Cluster Environments
abstract
Cycle-harvesting systems such as Condor have been developed to make desktop machines in a local area (which are often similar to clusters in hardware configuration) available as a compute platform. To provide a dual-use capability, opportunistic jobs harvesting cycles from the desktop must be checkpointed before the desktop resources are reclaimed by their owners and the job is evacuated. In this paper, we investigate a new system for computing efficient checkpoint schedules in cycle-harvesting environments. Our system records the historical availability from each resource and fits a statistical model to the observations. Because checkpointing must often traverse the network (i.e. the desktop hosts do not provide sufficient persistent storage for checkpoints), we combine this model with predictions of network performance to the storage site to compute a checkpoint schedule. When an application is initiated on a particular resource, the system uses the computed distribution to parameterize a Markov state-transition model for the application's execution, evaluates the expected time and network overhead as a function of the checkpoint interval, and numerically optimizes with respect to time. We report on the performance of and implementation of this system using the Condor cycle-harvesting environment at the University of Wisconsin. We also evaluate the efficiencies we achieve for a variety of network overheads using trace-based simulation. Finally, we validate our simulations against the observed performance with Condor. Our results indicate that while the choice of model distribution has a relatively small but positive effect on time efficiency, it has a substantial impact on network utilization
Daniel Nurmi, John Brevik, Richard Wolski
CLUSTER2
2005 Modeling Machine Availability in Enterprise and Wide-Area Distributed Computing Environments
Daniel Nurmi, John Brevik, Richard Wolski
Euro-Par2
2004 Automatic methods for predicting machine availability in desktop Grid and peer-to-peer systems
abstract
In this paper we examine the problem of predicting machine availability in desktop and enterprise computing environments. Predicting the duration that a machine will run until it restarts (availability duration) is critically useful to application scheduling and resource characterization in federated systems. We describe one parametric model fitting technique and two nonparametric prediction techniques, comparing their accuracy in predicting the quantiles of empirically observed machine availability distributions. We describe each method analytically and evaluate its precision using a synthetic trace of machine availability constructed from a known distribution. To detail their practical efficacy, we apply them to machine availability traces from three separate desktop and enterprise computing environments, and evaluate each method in terms of the accuracy with which it predicts availability in a trace driven simulation. Our results indicate that availability duration can be predicted with quantifiable confidence bounds and that these bounds can he used as conservative bounds on lifetime predictions. Moreover a nonparametric method based on a binomial approach generates the most accurate estimates.
John Brevik, Daniel Nurmi, Richard Wolski
CCGRID1
2001 G-commerce: Market Formulations Controlling Resource Allocation on the Computational Grid
abstract
In this paper we investigate G-commerce-computational economies for controlling resource allocation in Computational Grid settings. We define hypothetical resource consumers (representing users and Grid-aware applications) and resource producers (representing resource owners who "sell" their resources to the Grid). We then measure the efficiency of resource allocation under two different market conditions: commodities markets and auctions. We compare both market strategies in terms of price stability, market equilibrium, consumer efficiency, and producer efficiency. Our results indicate that commodities markets are a better choice for controlling Grid resources than previously defined auction strategies.
Richard Wolski, James S. Plank, John Brevik, Todd Bryan
IPDPS3
2001 Writing Programs that Run EveryWare on the Computational Grid
abstract
The Computational Grid has been proposed, for the implementation of high-performance applications using widely dispersed computational resources. The goal of a Computational Grid is to aggregate ensembles of shared, heterogeneous, and distributed resources (potentially controlled by separate organizations) to provide computational, "power" to an application program. We provide a toolkit for the development of globally deployable Grid applications. The toolkit, called EveryWare, enables an application to draw computational power transparently from the Grid. It consists of a portable set of processes and libraries that can be incorporated into an application so that a wide variety of dynamically changing distributed infrastructures and resources can be used together to achieve supercomputer-like performance. We provide our experiences gained while building the EveryWare toolkit prototype and an explanation of its use in implementing a large-scale Grid application.
Richard Wolski, John Brevik, Graziano Obertelli, Neil Spring, Alan Su 0001
IEEE Trans. Parallel Distributed Syst.2
1999 Running EveryWare on the Computational Grid
abstract
The Computational Grid [10] has recently been proposed for the implementation of high-performance applications using widely dispersed computational resources. The goal of a Computational Grid is to aggregate ensembles of shared, heterogeneous, and distributed resources (potentially controlled by separate organizations) to provide computational "power" to an application program. In this paper, we provide a toolkit for the development of Grid applications. The toolkit, called EveryWare, enables an application to draw computational power transparently from the Grid. The toolkit consists of a portable set of processes and libraries that can be incorporated into an application so that a wide variety of dynamically changing distributed infrastructures and resources can be used together to achieve supercomputer-like performance. We provide our experiences gained while building the EveryWare toolkit prototype and the first true Grid application. 1 Introduction Increasingly, the high-perform...
Richard Wolski, John Brevik, Chandra Krintz, Graziano Obertelli, Neil Spring, Alan Su 0001
SC2