Jerome A. Rolia

dblp:r/JerryRolia · also Jerry Rolia · DBLP profile ↗
← Back
45ranked-venue papers
6as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 15 · 1 first-authorSystems, architecture and hardware · 11 · 2 first-authorSoftware engineering, systems software and programming languages · 10 · 2 first-authorSecurity and privacy · 3Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Performance modeling and evaluation · 64% Parallel and multicore computing · 17% Cloud and datacenter computing · 9%
Computer networks
1 paper
Routing and switching · 61% Network optimization and economics · 30% Transport protocols and congestion control · 9%

Topics — the 27 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
workload characterization
0.452012
DEC: Service Demand Estimation with Confidence · IEEE Trans. Software Eng. 2012
BURN: Enabling Workload Burstiness in Customized Service Benchmarks · IEEE Trans. Software Eng. 2012
A Synthetic Workload Generation Technique for Stress Testing Session-Based Systems · IEEE Trans. Software Eng. 2006
Parallel and multicore computing › data-parallel programming
mapreduce
0.222012
SkewTune in Action: Mitigating Skew in MapReduce Applications · Proc. VLDB Endow. 2012
SkewTune: mitigating skew in mapreduce applications · SIGMOD Conference 2012
Performance modeling and evaluation
benchmarking
0.222012
BURN: Enabling Workload Burstiness in Customized Service Benchmarks · IEEE Trans. Software Eng. 2012
A Synthetic Workload Generation Technique for Stress Testing Session-Based Systems · IEEE Trans. Software Eng. 2006
Performance modeling and evaluation
queueing models
0.232011
WAM - The Weighted Average Method for Predicting the Performance of Systems with Bursts of Customer Sessions · IEEE Trans. Software Eng. 2011
Trace-Based Load Characterization for Gernerating Performance Software Models · IEEE Trans. Software Eng. 1999
The Method of Layers · IEEE Trans. Software Eng. 1995
High-performance computing
cluster computing
0.112012
SkewTune in Action: Mitigating Skew in MapReduce Applications · Proc. VLDB Endow. 2012
Cloud and datacenter computing
cluster resource management and scheduling
0.112012
SkewTune: mitigating skew in mapreduce applications · SIGMOD Conference 2012
Performance modeling and evaluation › statistical analysis
confidence interval estimation
0.112012
DEC: Service Demand Estimation with Confidence · IEEE Trans. Software Eng. 2012
Parallel and multicore computing
skew mitigation
0.112012
SkewTune in Action: Mitigating Skew in MapReduce Applications · Proc. VLDB Endow. 2012
Performance modeling and evaluation › workload characterization
workload generation
0.112012
BURN: Enabling Workload Burstiness in Customized Service Benchmarks · IEEE Trans. Software Eng. 2012
Performance modeling and evaluation
performance prediction
0.112011
WAM - The Weighted Average Method for Predicting the Performance of Systems with Bursts of Customer Sessions · IEEE Trans. Software Eng. 2011
Routing and switching
inter-domain routing
0.112007
Load-Balancing Data Traffic Among Inter-Domain Links · IEEE J. Sel. Areas Commun. 2007
Network optimization and economics
resource allocation
0.112007
Load-Balancing Data Traffic Among Inter-Domain Links · IEEE J. Sel. Areas Commun. 2007
Routing and switching › traffic engineering
traffic load balancing
0.112007
Load-Balancing Data Traffic Among Inter-Domain Links · IEEE J. Sel. Areas Commun. 2007
Performance modeling and evaluation › workload characterization › workload modeling
synthetic workload generation
0.112006
A Synthetic Workload Generation Technique for Stress Testing Session-Based Systems · IEEE Trans. Software Eng. 2006
Performance modeling and evaluation › queueing models › queueing network model
layered queueing networks
0.122011
WAM - The Weighted Average Method for Predicting the Performance of Systems with Bursts of Customer Sessions · IEEE Trans. Software Eng. 2011
Trace-Based Load Characterization for Gernerating Performance Software Models · IEEE Trans. Software Eng. 1999
Performance modeling and evaluation
capacity planning
0.012012
DEC: Service Demand Estimation with Confidence · IEEE Trans. Software Eng. 2012
Distributed systems › distributed system architecture
multi-tier application
0.012012
BURN: Enabling Workload Burstiness in Customized Service Benchmarks · IEEE Trans. Software Eng. 2012
Parallel and multicore computing
parallel programming models and runtimes
0.012012
SkewTune: mitigating skew in mapreduce applications · SIGMOD Conference 2012
Cloud and datacenter computing
resource management
0.012012
DEC: Service Demand Estimation with Confidence · IEEE Trans. Software Eng. 2012
Distributed systems › replication
process replication
0.012000
Designing Process Replication and Activation: A Quantitative Approach · IEEE Trans. Software Eng. 2000
Transport protocols and congestion control › TCP congestion control
congestion avoidance
0.012007
Load-Balancing Data Traffic Among Inter-Domain Links · IEEE J. Sel. Areas Commun. 2007
Performance modeling and evaluation › performance evaluation methodology
stress testing
0.012006
A Synthetic Workload Generation Technique for Stress Testing Session-Based Systems · IEEE Trans. Software Eng. 2006
Performance modeling and evaluation › queueing models
mean value analysis
0.011995
The Method of Layers · IEEE Trans. Software Eng. 1995
Performance modeling and evaluation › software performance engineering
software performance modeling
0.011995
The Method of Layers · IEEE Trans. Software Eng. 1995
Distributed systems
remote procedure call
0.011994
Modeling RPC Performance · SIGMETRICS 1994
Performance modeling and evaluation › analytical modeling
analytical performance evaluation
0.012000
Designing Process Replication and Activation: A Quantitative Approach · IEEE Trans. Software Eng. 2000
Software maintenance and evolution
software performance engineering
0.011999
Trace-Based Load Characterization for Gernerating Performance Software Models · IEEE Trans. Software Eng. 1999

Methods — techniques the papers use, named apart from their topics

task repartitioning · 0.1straggler mitigation · 0.1runtime skew detection · 0.1regression · 0.1queueing model · 0.1proactive task repartitioning · 0.1overdemand metric · 0.1optimization · 0.1markov model · 0.1markov chain · 0.1simulation · 0.1
YearPublicationVenuePosition
2015 Resource Contention Detection in Virtualized Environments
abstract
Public and private cloud computing environments employ virtualization methods to consolidate application workloads onto shared servers. Modern servers typically have one or more sockets each with one or more computing cores, a multi-level caching hierarchy, a memory subsystem, and an interconnect to the memory of other sockets. While resource management methods may manage application performance by controlling the sharing of processing time and input-output rates, there is generally no management of contention for virtualization kernel resources or for the memory hierarchy and subsystems. Yet such contention can have a significant impact on application performance. Hardware platform specific counters have been proposed for detecting such contention. We show that such counters alone are not always sufficient for detecting contention. We propose a software probe based approach for detecting contention for shared platform resources and demonstrate its effectiveness. We show that the probe imposes low overhead and is remarkably effective at detecting performance degradations due to inter-VM interference over a wide variety of workload scenarios and on two different server architectures. The probe successfully detected virtualization-induced software bottleneck and memory contention on both server architectures. Our approach supports the management of workload placement on shared servers and pools of shared servers.
Joydeep Mukherjee, Diwakar Krishnamurthy, Jerome A. Rolia
IEEE Trans. Netw. Serv. Manag.3
2013 Workload analysis and demand prediction for the HP ePrint Service
Vipul Garg, Ludmila Cherkasova, Swaminathan Packirisami, Jerome A. Rolia
IM4
2013 ACE: Automated capacity evaluation for HP ePrint
Vipul Garg, Ludmila Cherkasova, Swaminathan Packirisami, Jerome A. Rolia
IM4
2013 Resource contention detection and management for consolidated workloads
Joydeep Mukherjee, Diwakar Krishnamurthy, Jerome A. Rolia, Chris Hyser
IM3
2012 Selling T-shirts and Time Shares in the Cloud
abstract
Cloud computing has emerged as a new and alternative approach for providing computing services. Customers acquire and release resources by requesting and returning virtual machines to the cloud. Different service models and pricing schemes are offered by cloud service providers. This can make it difficult for customers to compare cloud services and select an appropriate solution. Cloud Infrastructure-as-a-Service vendors offer a t-shirt approach for Virtual Machines (VMs) on demand. Customers can select from a set of fixed size VMs and vary the number of VMs as their demands change. Private clouds often offer another alternative, called time-sharing, where the capacity of each VM is permitted to change dynamically. With this approach each virtual machine is allocated a dynamic amount of CPU and memory resources over time to better utilize available resources. We present a tool that can help customers make informed decisions about which approach works most efficiently for their workloads in aggregate and for each workload separately. A case study using data from an enterprise customer with 312 workloads demonstrates the use of the tool. It shows that for the given set of workloads the t-shirt model requires almost twice the number of physical servers as the time share model. The costs for such infrastructure must ultimately be passed on to the customer in terms of monetary costs or performance risks. We conclude that private and public clouds should consider offering both resource sharing models to meet the needs of customers.
Daniel Gmach, Jerome A. Rolia, Ludmila Cherkasova
CCGRID2
2012 A Context-Aware framework for patient Navigation and Engagement (CANE)
abstract
Engaging patients in the management of their health care can improve the quality of their care and enhance their experience while making more efficient use of care provider resources, especially for chronic diseases. However, health care system complexity and the challenge of consumer health literac
Jerome A. Rolia, Sujoy Basu, Sharad Singhal, Akhil Kumar 0001
CollaborateCom2
2012 Passive crowd-based monitoring of World Wide Web infrastructure and its performance
abstract
The World Wide Web and the services it provides are continually evolving. Even for a single time instant, it is a complex task to methodologically determine the infrastructure over which these services are provided and the corresponding effect on user perceived performance. For such tasks, researchers typically rely on active measurements or large numbers of volunteer users. In this paper, we consider an alternative approach, which we refer to as passive crowd-based monitoring. More specifically, we use passively collected proxy logs from a global enterprise to observe differences in the quality of service (QoS) experienced by users on different continents. We also show how this technique can measure properties of the underlying infrastructures of different Web content providers. While some of these properties have been observed using active measurements, we are the first to show that many of these properties (such as location of servers) can be obtained using passive measurements of actual user activity. Passive crowd-based monitoring has the advantages that it does not add any overhead on Web infrastructure, it does not require any specific software on the clients, but still captures the performance and infrastructure observed by actual Web usage.
Martin F. Arlitt, Niklas Carlsson, Carey L. Williamson, Jerome A. Rolia
ICC4
2012 Comparing efficiency and costs of cloud computing models
abstract
Public and private clouds are being adopted as a cost-effective approach for sharing IT resources. Customers acquire and release resources by requesting and returning virtual machines to the cloud. Different service models are proposed for virtual machine resource management. Some public cloud providers follow a t-shirt model for VM resource sizing. A second approach for resource management is based on a time share model. This paper compares the two approaches from the perspective of resource usage for both the service provider and workload owner. Using data from 312 customer applications, we show that the t-shirt model requires 40% more infrastructure than when a finer degree of resource sharing based on time varying resource shares is permitted.
Daniel Gmach, Jerome A. Rolia, Ludmila Cherkasova
NOMS2
2012 SkewTune: mitigating skew in mapreduce applications
abstract
We present an automatic skew mitigation approach for user-defined MapReduce programs and present SkewTune, a system that implements this approach as a drop-in replacement for an existing MapReduce implementation. There are three key challenges: (a) require no extra input from the user yet work for all MapReduce applications, (b) be completely transparent, and (c) impose minimal overhead if there is no skew. The SkewTune approach addresses these challenges and works as follows: When a node in the cluster becomes idle, SkewTune identifies the task with the greatest expected remaining processing time. The unprocessed input data of this straggling task is then proactively repartitioned in a way that fully utilizes the nodes in the cluster and preserves the ordering of the input data so that the original output can be reconstructed by concatenation. We implement SkewTune as an extension to Hadoop and evaluate its effectiveness using several real applications. The results show that SkewTune can significantly reduce job runtime in the presence of skew and adds little to no overhead in the absence of skew.
YongChul Kwon, Magdalena Balazinska, Bill Howe, Jerome A. Rolia
SIGMOD Conference4
2012 SkewTune in Action: Mitigating Skew in MapReduce Applications
abstract
We demonstrate SkewTune, a system that automatically mitigates skew in user-defined MapReduce programs and is a drop-in replacement for Hadoop. The demonstration has two parts. First, we demonstrate how SkewTune mitigates skew in real MapReduce applications at runtime by running a real application in a public cloud. Second, through an interactive graphical interface, we demonstrate the details of the skew mitigation process using both real and synthetic workloads that represent various skew configurations.
YongChul Kwon, Magdalena Balazinska, Bill Howe, Jerome A. Rolia
Proc. VLDB Endow.4
2012 BURN: Enabling Workload Burstiness in Customized Service Benchmarks
abstract
We introduce BURN, a methodology to create customized benchmarks for testing multitier applications under time-varying resource usage conditions. Starting from a set of preexisting test workloads, BURN finds a policy that interleaves their execution to stress the multitier application and generate controlled burstiness in resource consumption. This is useful to study, in a controlled way, the robustness of software services to sudden changes in the workload characteristics and in the usage levels of the resources. The problem is tackled by a model-based technique which first generates Markov models to describe resource consumption patterns of each test workload. Then, a policy is generated using an optimization program which sets as constraints a target request mix and user-specified levels of burstiness at the different resources in the system. Burstiness is quantified using a novel metric called overdemand, which describes in a natural way the tendency of a workload to keep a resource congested for long periods of time and across multiple requests. A case study based on a three-tier application testbed shows that our method is able to control and predict burstiness for session service demands at a fine-grained scale. Furthermore, experiments demonstrate that for any given request mix our approach can expose latency and throughput degradations not found with nonbursty workloads having the same request mix.
Giuliano Casale, Amir S. Kalbasi, Diwakar Krishnamurthy, Jerome A. Rolia
IEEE Trans. Software Eng.4
2012 DEC: Service Demand Estimation with Confidence
abstract
We present a new technique for predicting the resource demand requirements of services implemented by multitier systems. Accurate demand estimates are essential to ensure the efficient provisioning of services in an increasingly service-oriented world. The demand estimation technique proposed in this paper has several advantages compared with regression-based demand estimation techniques, which many practitioners employ today. In contrast to regression, it does not suffer from the problem of multicollinearity, it provides more reliable aggregate resource demand and confidence interval predictions, and it offers a measurement-based validation test. The technique can be used to support system sizing and capacity planning exercises, costing and pricing exercises, and to predict the impact of changes to a service upon different service customers.
Amir S. Kalbasi, Diwakar Krishnamurthy, Jerome A. Rolia, Stephen Dawson
IEEE Trans. Software Eng.3
2011 MODE: Mix Driven On-line Resource Demand Estimation
Amir S. Kalbasi, Diwakar Krishnamurthy, Jerome A. Rolia
CNSM3
2011 Resource and virtualization costs up in the cloud: Models and design choices
abstract
Virtualization offers the potential for cost-effective service provisioning. For service providers who make significant investments in new virtualized data centers in support of private or public clouds, one of the serious challenges is the problem of recovering costs for new server hardware, software, network, storage, management, etc. Gaining visibility and accurately determining the cost of shared resources used by collocated services is essential for implementing a proper chargeback approach in cloud environments. We introduce and compare three different models for apportioning cost and champion the one that is least sensitive to workload placement decisions and provides the most robust and repeatable cost estimates. A detailed study involving 312 workloads from an HP customer environment demonstrates the result. Finally, we employ the cost model in a case study that evaluates the impact on the cost of exploiting different virtualization platform alternatives for the 312 workloads. For example, some workloads may cost more to host using certain virtualization platforms than on others or on standalone hosts. We demonstrate different decision points with potential cost savings of nearly 20% by “right-virtualizing” the workloads.
Daniel Gmach, Jerome A. Rolia, Ludmila Cherkasova
DSN2
2011 Chargeback model for resource pools in the cloud
abstract
This paper presents three methods for apportioning server costs among workloads in shared resource environments such as computing clouds. We consider a fine sharing of resources, the impact of time varying resource usage, large ratios for peak to mean workload demands, and the influence of random choices for the co-placement of workloads on shared servers. These features can affect the quantity of servers needed to support workloads as well as the robustness of the cost values assigned to each workload. We compare the three methods for apportioning costs and recommend the method that assigns costs in the most repeatable manner. A detailed study involving 312 workloads from an HP customer environment demonstrates the result.
Daniel Gmach, Jerome A. Rolia, Ludmila Cherkasova
Integrated Network Management2
2011 WAM - The Weighted Average Method for Predicting the Performance of Systems with Bursts of Customer Sessions
abstract
Predictive performance models are important tools that support system sizing, capacity planning, and systems management exercises. We introduce the Weighted Average Method (WAM) to improve the accuracy of analytic predictive performance models for systems with bursts of concurrent customers. WAM considers the customer population distribution at a system to reflect the impact of bursts. The WAM approach is robust with respect to distribution functions, including heavy-tail-like distributions, for workload parameters. We demonstrate the effectiveness of WAM using a case study involving a multitier TPC-W benchmark system. To demonstrate the utility of WAM with multiple performance modeling approaches, we developed both Queuing Network Models and Layered Queuing Models for the system. Results indicate that WAM improves prediction accuracy for bursty workloads for QNMs and LQMs by 10 and 12 percent, respectively, with respect to a Markov Chain approach reported in the literature.
Diwakar Krishnamurthy, Jerome A. Rolia
IEEE Trans. Software Eng.2
2010 Skew-resistant parallel processing of feature-extracting scientific user-defined functions
abstract
Scientists today have the ability to generate data at an unprecedented scale and rate and, as a result, they must increasingly turn to parallel data processing engines to perform their analyses. However, the simple execution model of these engines can make it difficult to implement efficient algorithms for scientific analytics. In particular, many scientific analytics require the extraction of features from data represented as either a multidimensional array or points in a multidimensional space. These applications exhibit significant computational skew, where the runtime of different partitions depends on more than just input size and can therefore vary dramatically and unpredictably. In this paper, we present SkewReduce, a new system implemented on top of Hadoop that enables users to easily express feature extraction analyses and execute them efficiently. At the heart of the SkewReduce system is an optimizer, parameterized by user-defined cost functions, that determines how best to partition the input data to minimize computational skew. Experiments on real data from two different science domains demonstrate that our approach can improve execution times by a factor of up to 8 compared to a naive implementation.
YongChul Kwon, Magdalena Balazinska, Bill Howe, Jerome A. Rolia
SoCC4
2010 Capacity planning and power management to exploit sustainable energy
abstract
This paper describes an approach for designing a power management plan that matches the supply of power with the demand for power in data centers. Power may come from the grid, from local renewable sources, and possibly from energy storage subsystems. The supply of renewable power is often time-varying in a manner that depends on the source that provides the power, the location of power generators, and the weather conditions. The demand for power is mainly determined by the time-varying workloads hosted in the data center and the power management policies implemented by the data center. A case study demonstrates how our approach can be used to design a plan for realistic and complex data center workloads. The study considers a data center's deployment in two geographic locations with different supplies of power. Our approach offers greater precision than other planning methods that do not take into account time-varying power supply and demand and data center power management policies.
Daniel Gmach, Jerome A. Rolia, Cullen E. Bash, Yuan Chen 0001, Tom Christian, Amip Shah, Ratnesh K. Sharma, Zhikui Wang
CNSM2
2009 Automatic Stress Testing of Multi-tier Systems by Dynamic Bottleneck Switch Generation
Giuliano Casale, Amir S. Kalbasi, Diwakar Krishnamurthy, Jerome A. Rolia
Middleware4
2009 Resource pool management: Reactive versus proactive or let's be friends
Daniel Gmach, Jerome A. Rolia, Ludmila Cherkasova, Alfons Kemper
Comput. Networks2
2008 An integrated approach to resource pool management: Policies, efficiency and quality metrics
abstract
The consolidation of multiple servers and their workloads aims to minimize the number of servers needed thereby enabling the efficient use of server and power resources. At the same time, applications participating in consolidation scenarios often have specific quality of service requirements that need to be supported. To evaluate which workloads can be consolidated to which servers we employ a trace-based approach that determines a near optimal workload placement that provides specific qualities of service. However, the chosen workload placement is based on past demands that may not perfectly predict future demands. To further improve efficiency and application quality of service we apply the trace-based technique repeatedly, as a workload placement controller. We integrate the workload placement controller with a reactive controller that observes current behavior to i) migrate workloads off of overloaded servers and ii) free and shut down lightly-loaded servers. To evaluate the effectiveness of the approach, we developed a new host load emulation environment that simulates different management policies in a time effective manner. A case study involving three months of data for 138 SAP applications compares our integrated controller approach with the use of each controller separately. The study considers trade-offs between i) required capacity and power usage, ii) resource access quality of service for CPU and memory resources, and iii) the number of migrations. We consider two typical enterprise environments: blade and server based resource pool infrastructures. The results show that the integrated controller approach outperforms the use of either controller separately for the enterprise application workloads in our study. We show the influence of the blade and server pool infrastructures on the effectiveness of the management policies.
Daniel Gmach, Jerome A. Rolia, Ludmila Cherkasova, Guillaume Belrose, Tom Turicchi, Alfons Kemper
DSN2
2007 Capacity Management and Demand Prediction for Next Generation Data Centers
abstract
Advances in server, network, and storage virtualization are enabling the creation of resource pools of servers that permit multiple application workloads to share each server in the pool. This paper proposes and evaluates aspects of a capacity management process for automating the efficient use of such pools when hosting large numbers of services. We use a trace based approach to capacity management that relies on i) a definition for required capacity, ii) the characterization of workload demand patterns, iii) the generation of synthetic workloads that predict future demands based on the patterns, and iv) a workload placement recommendation service. A case study with 6 months of data representing the resource usage of 139 workloads in an enterprise data center demonstrates the effectiveness of the proposed capacity management process. Our results show that when consolidating to 8 processor systems, we predicted future per-server required capacity to within one processor 95% of the time. The approach enabled a 35% reduction in processor usage as compared to today's current best practice for workload placement.
Daniel Gmach, Jerome A. Rolia, Ludmila Cherkasova, Alfons Kemper
ICWS2
2007 Load-Balancing Data Traffic Among Inter-Domain Links
abstract
The Internet has evolved into a multi-service infrastructure for the telecom and computer industries. Internet services are affected by congestions caused by the operations of Internet routing protocols. We propose a novel information service that guides the operations of Internet routing protocols to avoid such congestion. The information service maintains information about utilization of inter-domain links. The proposed service minimizes the maximum of utilization of Internet links by selecting potential network domains to be traversed by Internet services. Simulation results show better balanced distribution of traffic workloads among network links using the proposed service. We show that the proposed service scales well as network size grows. This comes at the cost of greater control messaging overhead which suggests using the proposed service with long-lived and higher bandwidth services
Mohamed El-Darieby, Dorina C. Petriu, Jerome A. Rolia
IEEE J. Sel. Areas Commun.3
2006 Replay: A Model-Based Service for Supporting Transparent Cluster Analysis Tools
abstract
Grid computing environments typically federate heterogeneous resource clusters belonging to several organizations. To fully realize the promise of a grid environment, it is necessary to support tools that help obtain insights into the behaviour of individual clusters. This paper describes a cluster service called Replay that simplifies the development and maintenance of such tools. The service provides a common model-based interface for obtaining current and historical information about a cluster. Replay can manage multiple views which allows tools to obtain information about an existing cluster as well as information that shows how a cluster might have behaved under alternate configurations and workloads. The model-based interface allows tools to be ported to different clusters with little effort. Furthermore, Replay uses different mechanisms to manage information that typically changes infrequently and information that can change in a more dynamic, continuous manner allowing it to handle information more efficiently than existing services that provide similar functionality. The paper presents a job analysis tool to illustrate the utility of the service
Diwakar Krishnamurthy, Cameron Kiddle, Jerome A. Rolia, Rob Simmonds
CLUSTER3
2006 R-Opus: A Composite Framework for Application Performability and QoS in Shared Resource Pools
abstract
We consider shared resource pool management taking into account per-application quality of service (QoS) requirements and server failures. Application QoS requirements are defined by complementary specifications for acceptable and time-limited degraded performance. Furthermore, a requirement specification is provided for both the normal case and for the case where an application server fails in the pool. Independently, the resource pool operator provides a resource access QoS commitment for two classes of service (CoS). These govern statistical multiplexing within the pool. A QoS translation automatically maps application demands onto the resource pool's CoS to best enable sharing. A workload placement service consolidates applications to a small number of servers while satisfying application QoS requirements. The service reports whether a spare server is needed or how applications affected by a single failure can operate according to failure QoS constraints using remaining servers until the failure can be repaired. A case study demonstrates the approach
Ludmila Cherkasova, Jerome A. Rolia
DSN2
2006 Hierarchical Creation of Virtual Networks
abstract
This paper describes a scalable and hierarchical virtual network creation (HVNC) protocol. The protocol encapsulates network control operations such as signaling and routing. It relies on a hierarchical organization of network resources and their managers. The network hierarchy of HVNC provides network-wide views of network resource status. This enables network-wide decisions such as balancing traffic loads among network domains and traffic isolation in accordance to policies. HVNC enables flexible and dynamic virtual network management including reconfiguration and failure recovery. These advantages come at the cost of higher setup and information storage costs when compared to standard service creation protocols. We characterize the advantages and costs of HVNC by simulation. The relatively high costs of HVNC make it more suitable to create long-lived high-bandwidth virtual networks. This type of virtual networks supports emerging technologies and applications such as grid and peer-to-peer computing and data-intensive e-science applications
Mohamed El-Darieby, Jerome A. Rolia
NOMS2
2006 Configuring Workload Manager Control Parameters for Resource Pools
abstract
Resource pools are computing environments that offer virtualized access to shared resources. When used effectively they can align the use of capacity with business needs (flexibility), lower infrastructure costs (via resource sharing), and lower operating costs (via automation). Using resources effectively can rely on a combination of workload placement and workload management technologies. Workload placement decides which workloads will share resources. Workload management governs short term access to resource capacity. It provides performance isolation within resource pools to ensure resource sharing even under high loads. A workload manager can have a direct impact both on an application's overall resource access quality of service and on the number of workloads that can be assigned to a pool. In this paper we take a detailed look at an application workload's demands. We explore tradeoffs in resource access quality of service received by the application and the minimum allocation of resources for the workload. We show that by careful selection of workload scheduling parameters along with a proposed fast allocation policy we can sometimes more than triple the number of workloads that can be assigned to a pool without sacrificing application workload quality of service or the efficiency of the resource pool.
Jerome A. Rolia, Ludmila Cherkasova, Clifford McCarthy
NOMS1
2006 A Synthetic Workload Generation Technique for Stress Testing Session-Based Systems
abstract
Enterprise applications are often business critical but lack effective synthetic workload generation techniques to evaluate performance. These workloads are characterized by sessions of interdependent requests that often cause and exploit dynamically generated responses. Interrequest dependencies must be reflected in synthetic workloads for these systems to exercise application functions correctly. This poses significant challenges for automating the construction of representative synthetic workloads and manipulating workload characteristics for sensitivity analyses. This paper presents a technique to overcome these problems. Given request logs for a system under study, the technique automatically creates a synthetic workload that has specified characteristics and maintains the correct interrequest dependencies. The technique is demonstrated through a case study involving a TPC-W e-commerce system. Results show that incorrect performance results can be obtained by neglecting interrequest dependencies, thereby highlighting the value of our technique. The study also exploits our technique to investigate the impact of several workload characteristics on system performance. Results establish that high variability in the distributions of session length, session idle times, and request service times can cause increased contention among sessions, leading to poor system responsiveness. To the best of our knowledge, these are the first results of this kind for a session-based system. We believe our technique is of value for studies where fine control over workload is essential
Diwakar Krishnamurthy, Jerome A. Rolia, Shikharesh Majumdar
IEEE Trans. Software Eng.2
2005 Quartermaster - a resource utility system
abstract
Utility computing is envisioned as the future of enterprise IT environments. Achieving utility computing is a daunting task, because enterprise users have diverse and complex needs. In this paper we describe quartermaster, an integrated set of tools that addresses some of these needs. Quartermaster supports the entire lifecycle of computing tasks - including design, deployment, operation, and decommissioning of each task. Although individual components of this lifecycle have been addressed in earlier work, quartermaster integrates them in a unified framework using model-based automation. All tools within quartermaster are integrated using models based on the common information model (CIM), an industry-standard model from the distributed management task force (DMTF). The paper discusses the quartermaster implementation, and describes two case studies using quartermaster.
Sharad Singhal, Martin F. Arlitt, Dirk Beyer 0002, Sven Graupner, Vijay Machiraju, Jim Pruyne, Jerome A. Rolia, Akhil Sahai, Cipriano A. Santos, Julie Ward, Xiaoyun Zhu
Integrated Network Management7
2004 Statistical service assurances for applications in utility grid environments
Jerome A. Rolia, Xiaoyun Zhu, Martin F. Arlitt, Artur Andrzejak 0001
Perform. Evaluation1
2003 HDRA: hierarchical distributed routing algorithm
abstract
The paper presents a novel scalable end-to-end routing protocol, HDRA, which combines the characteristics of both hierarchical and distributed routing protocols. HDRA selects a number of network domains to be traversed by a requested route. Only the selected domains are flooded with probing messages. This enables the deployment of flooding-based routing algorithms in large networks. The protocol has smaller setup time because it performs no central routing computation. HDRA does not require global state maintenance. However, it has higher message overheads when compared to PNNI (private network-network interface).
Mohamed El-Darieby, Dorina C. Petriu, Jerome A. Rolia
GLOBECOM3
2003 Hierarchical End-to-End Service Recovery
Mohamed El-Darieby, Dorina C. Petriu, Jerome A. Rolia
Integrated Network Management3
2003 Resource Access Management for a Utility Hosting Enterprise Applications
Jerome A. Rolia, Xiaoyun Zhu, Martin F. Arlitt
Integrated Network Management1
2003 Grids for Enterprise Applications
Jerome A. Rolia, Jim Pruyne, Xiaoyun Zhu, Martin F. Arlitt
JSSPP1
2002 A hierarchical distributed protocol for MPLS path creation
abstract
Network service provisioning involves the control of network resources through signaling, routing and management protocols that achieve quality of service and traffic engineering objectives. Network service providers are typically faced with scalability problems due to the explosion in network size and demand. In this paper we propose a new hierarchical distributed protocol for the provisioning of MPLS-based tunnels. The protocol encapsulates the required signaling and routing. The protocol alleviates the scalability problem through extending the current two-tier architecture of the Internet to a multi-level hierarchical one. It enables mechanisms for quality of service, and intra- and inter-domain traffic engineering. Analytical analysis and simulation results show that the protocol exploits parallelism in routing computation to reduce the setup time of an MPLS path. This comes at the expense of an increased message complexity relative to other hierarchical protocols.
Mohamed El-Darieby, Dorina C. Petriu, Jerome A. Rolia
ISCC3
2001 Performance Modeling for Virtual Network based Service Provisioning
abstract
One of the goals of network service providers (SP) is to provision efficiently value-added networking services to their end users. The services may have different duration, connectivity, and quality of service (QoS) requirements. An SP meets such requirements by exploiting various routing and management mechanisms for edge and core networks. We focus on network services that require the establishment of virtual networks (VNs). This paper considers a VN creation architecture based on multi-protocol label switching (MPLS). MPLS is one approach for implementing VNs that enables programmability for network infrastructure and hence the creation of value-added services. An analytic performance model is developed to assess the scalability of the architecture as the underlying physical network infrastructure evolves. The model we present describes only coarse features of the architecture, yet offers insights regarding its performance and scalability. In particular we assess the impact of a flat versus hierarchical MPLS mechanism on the response time of service creation.
Mohamed El-Darieby, Jerome A. Rolia, Dorina C. Petriu
Integrated Network Management2
2001 Performance validation tools for software/hardware systems
Ulrich Herzog, Jerome A. Rolia
Perform. Evaluation2
2001 Characterizing the scalability of a large web-based shopping system
abstract
This article presents an analysis of five days of workload data from a large Web-based shopping system. The multitier environment of this Web-based shopping system includes Web servers, application servers, database servers, and an assortment of load-balancing and firewall appliances. We characterize user requests and sessions and determine their impact on system performance scalability. The purpose of our study is to assess scalability and support capacity planning exercises for the multitier system. We find that horizontal scalability is not always an adequate mechanism for supporting increased workloads and that personalization and robots can have a significant impact on system scalability.
Martin F. Arlitt, Diwakar Krishnamurthy, Jerome A. Rolia
ACM Trans. Internet Techn.3
2000 Designing Process Replication and Activation: A Quantitative Approach
abstract
Distributed application systems are composed of classes of objects with instances that interact to accomplish common goals. Such systems can have many classes of users with many types of requests. Furthermore, the relative load of these classes can shift throughout the day, causing changes to system behavior and bottlenecks. When designing and deploying such systems, it is necessary to determine a process replication and threading policy for the server processes that contain the objects, as well as process activation policies. To avoid bottlenecks, the policy must support all possible workload conditions. Licensing, implementation or resource constraints can limit the number of permitted replicas or threads of a server process. Process activation policies determine whether a server is persistent or should be created and terminated with each call. This paper describes quantitative techniques for choosing process replication or threading levels and process activation policies. Inappropriate policies can lead to unnecessary queuing delays for callers or unnecessarily high consumption of memory resources. The algorithms presented consider all workload conditions, are iterative in nature and are hybrid mathematical programming and analytic performance evaluation methods. An example is given to demonstrate the technique and describe how the results can be applied during software design and deployment.
Marin Litoiu, Jerome A. Rolia, Giuseppe Serazzi
IEEE Trans. Software Eng.2
1999 Trace-Based Load Characterization for Gernerating Performance Software Models
abstract
Performance models of software designs can give early warnings of problems such as resource saturation or excessive delays. However models are seldom used because of the considerable effort needed to construct them. The ANGIOTRACE/sup TM/ was developed to gather the necessary information from an executable design and develop a model in an automated fashion. It applies to distributed and concurrent software with synchronous (send-reply or RPC) communications, developing a layered queuing network model. The trace-based load characterization (TLC) technique presented here extends the ANGIOTRACE/sup TM/ to handle software with both synchronous and asynchronous interactions. TLC also detects interactions which are effectively synchronous or partly-synchronous (forwarding) but are built up from asynchronous messages. These patterns occur in telephony software and in other systems. The TLC technique can be applied throughout the software life-cycle, even after deployment.
Curtis E. Hrischuk, C. Murray Woodside, Jerome A. Rolia, Rod Iversen
IEEE Trans. Software Eng.3
1998 Web Server Performance Measurement and Modeling Techniques
John Dilley, Rich Friedrich, Tai Jin, Jerome A. Rolia
Perform. Evaluation4
1997 More Manageable Management Applications
Asham El Rayess, Jerome A. Rolia
DAIS2
1995 A Toolset for Performance Engineering and Software Design of Client-Server Systems
Greg Franks, Alex Hubbard, Shikharesh Majumdar, John E. Neilson, Dorina C. Petriu, Jerome A. Rolia, C. Murray Woodside
Perform. Evaluation6
1995 The Method of Layers
abstract
Distributed applications are being developed that contain one or more layers of software servers. Software processes within such systems suffer contention delays both for shared hardware and at the software servers. The responsiveness of these systems is affected by the software design, the threading level and number of instances of software processes, and the allocation of processes to processors. The Method of Layers (MOL) is proposed to provide performance estimates for such systems. The MOL uses the mean value analysis (MVA) linearizer algorithm as a subprogram to assist in predicting model performance measures.>
Jerome A. Rolia, Kenneth C. Sevcik
IEEE Trans. Software Eng.1
1994 Modeling RPC Performance
abstract
Distributed computing applications are collections of processes allocated across a network that cooperate to accomplish common goals. The applications require the support of a distributed computing runtime environment that provides services to help manage process concurrency and interprocess communication. This support helps to hide much of the inherent complexity of distributed environments via industry standard interfaces and permits developers to create more portable applications. The resource requirements of the runtime services can be significant and may impact application performance and system throughput. This paper describes work done to study the potential benefits of redesigning some aspects of the DCE RPC and its current implementation on a specific platform.
Jerome A. Rolia, M. Starkey, Gerald Boersma
SIGMETRICS1