EDBT 2026 Demo / reviewers in the wild / expert
Malgorzata Steinder
dblp:28/980 · also Malgorzata Gosia Steinder
· DBLP profile ↗
41ranked-venue papers
11as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 7 first-authorSystems, architecture and hardware · 11 · 1 first-authorSoftware engineering, systems software and programming languages · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Cloud and datacenter computing · 70% GPUs and heterogeneous computing · 25% Performance modeling and evaluation · 3% | |
| Computer networks
3 papers |
Network management and operations · 95% Routing and switching · 5% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing
cluster resource management and scheduling |
0.4 | 3 | 2017 | Topology-aware GPU scheduling for learning workloads in cloud environments · SC 2017 A scalable application placement controller for enterprise data centers · WWW 2007 Dynamic placement for clustered web applications · WWW 2006 |
GPUs and heterogeneous computing
GPU scheduling |
0.3 | 1 | 2017 | Topology-aware GPU scheduling for learning workloads in cloud environments · SC 2017 |
GPUs and heterogeneous computing
multi-GPU computing |
0.3 | 1 | 2017 | Topology-aware GPU scheduling for learning workloads in cloud environments · SC 2017 |
Cloud and datacenter computing › job scheduling › network-aware scheduling
topology-aware scheduling |
0.3 | 1 | 2017 | Topology-aware GPU scheduling for learning workloads in cloud environments · SC 2017 |
Cloud and datacenter computing › cloud service management
cloud service selection |
0.2 | 1 | 2015 | Selecting Optimum Cloud Availability Zones by Learning User Satisfaction Levels · IEEE Trans. Serv. Comput. 2015 |
Network management and operations › fault management
fault diagnosis |
0.2 | 3 | 2007 | Multidomain Diagnosis of End-to-End Service Failures in Hierarchically Routed Networks · IEEE Trans. Parallel Distributed Syst. 2007 Probabilistic fault localization in communication systems using belief networks · IEEE/ACM Trans. Netw. 2004 Increasing robustness of fault localization through analysis of lost, spurious, and positive symptoms · INFOCOM 2002 |
Network management and operations › fault management › fault diagnosis
fault localization |
0.2 | 3 | 2007 | Multidomain Diagnosis of End-to-End Service Failures in Hierarchically Routed Networks · IEEE Trans. Parallel Distributed Syst. 2007 Probabilistic fault localization in communication systems using belief networks · IEEE/ACM Trans. Netw. 2004 Increasing robustness of fault localization through analysis of lost, spurious, and positive symptoms · INFOCOM 2002 |
Cloud and datacenter computing › datacenter operations
datacenter workload management |
0.1 | 1 | 2012 | Autonomic Placement of Mixed Batch and Transactional Workloads · IEEE Trans. Parallel Distributed Syst. 2012 |
Machine learning › Efficient and distributed learning
distributed training |
0.1 | 1 | 2017 | Topology-aware GPU scheduling for learning workloads in cloud environments · SC 2017 |
Cloud and datacenter computing › resource allocation › workload allocation
dynamic placement |
0.1 | 1 | 2008 | Managing SLAs of heterogeneous workloads using dynamic application placement · HPDC 2008 |
Cloud and datacenter computing
resource management |
0.1 | 1 | 2008 | Managing SLAs of heterogeneous workloads using dynamic application placement · HPDC 2008 |
Cloud and datacenter computing › datacenter architecture
virtualized datacenter |
0.1 | 1 | 2008 | Managing SLAs of heterogeneous workloads using dynamic application placement · HPDC 2008 |
Cloud and datacenter computing › cluster resource management and scheduling › resource scheduling
application placement |
0.1 | 1 | 2007 | A scalable application placement controller for enterprise data centers · WWW 2007 |
Cloud and datacenter computing › resource provisioning
dynamic resource provisioning |
0.1 | 1 | 2007 | A scalable application placement controller for enterprise data centers · WWW 2007 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2015 | Selecting Optimum Cloud Availability Zones by Learning User Satisfaction Levels · IEEE Trans. Serv. Comput. 2015 |
Cloud and datacenter computing › resource allocation
dynamic resource allocation |
0.1 | 1 | 2006 | Dynamic placement for clustered web applications · WWW 2006 |
Distributed systems
live reconfiguration |
0.0 | 1 | 2012 | Autonomic Placement of Mixed Batch and Transactional Workloads · IEEE Trans. Parallel Distributed Syst. 2012 |
Cloud and datacenter computing
virtualization |
0.0 | 1 | 2012 | Autonomic Placement of Mixed Batch and Transactional Workloads · IEEE Trans. Parallel Distributed Syst. 2012 |
Cloud and datacenter computing › cloud service management
service level agreement |
0.0 | 1 | 2008 | Managing SLAs of heterogeneous workloads using dynamic application placement · HPDC 2008 |
Routing and switching › routing
hierarchical routing |
0.0 | 1 | 2007 | Multidomain Diagnosis of End-to-End Service Failures in Hierarchically Routed Networks · IEEE Trans. Parallel Distributed Syst. 2007 |
Cloud and datacenter computing
resource allocation |
0.0 | 1 | 2007 | A scalable application placement controller for enterprise data centers · WWW 2007 |
Methods — techniques the papers use, named apart from their topics
topology-aware scheduling · 0.6simulation · 0.3predictive modeling · 0.2historical usage data · 0.2prototype · 0.1autonomic placement · 0.1utility function · 0.1suspension · 0.1migration · 0.1probabilistic inference · 0.1approximation algorithm · 0.1iterative belief updating · 0.0belief propagation · 0.0bayesian reasoning · 0.0belief network · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Resource Profile Advisor for Containers in Cognitive PlatformabstractContainers have transformed the cluster management into an application oriented endeavor, thus being widely used as the deployment units (i.e., micro-services) of large scale cloud services. As opposed to VMs, containers allow for resource provisioning with fine granularity and their resource usage directly reflects the micro-service behaviors. Container management systems like Kubernetes and Mesos provision resources to containers according to the capacity requested by the developers. Resource usages estimated by the developers are grossly inaccurate. They tend to be risk-averse and over provision resources, as under-provisioning would cause poor runtime performance or failures. Mehmet Fatih Aktas, Chen Wang 0039, Alaa Youssef, Malgorzata Steinder |
SoCC | 4 |
| 2017 | Enabling Distributed Software-Defined Environments Using Dynamic Infrastructure Service CompositionabstractService-based access models coupled with emerging application deployment technologies are enabling opportunities for realizing highly customized software-defined environments, which can support dynamic and data-driven applications. However, this requires rethinking traditional resource federation models to support dynamic resource compositions, which can adapt to evolving application needs and the dynamic state of underlying resources. In this paper, we present a programmable approach that leverages software-defined techniques to create a dynamic space-time infrastructure service composition. We propose the use of Constraint Programming as a formal language to allow users, applications, and service providers to define the desired state of the execution environment. The resulting distributed software-defined environment continually adapts to meet objectives/constraints set by the users, applications, and/or resource providers. We present the design and prototype implementation of such distributed software-defined environment. We use a cancer informatics workflow to demonstrate the operation of our framework using resources from five different cloud providers, which are aggregated on-demand based on dynamic user and resource provider constraints. Moustafa AbdelBaky, Javier Diaz Montes, Merve Unuvar, Melissa Romanus, Ivan Rodero, Malgorzata Steinder, Manish Parashar |
CCGrid | 6 |
| 2017 | Batch spot market for data analytics cloud providersabstractHosting data analytics services is challenging as their workload is often composed of on-line (e.g., interactive or streaming), requiring fast on-demand provisioning, and batch jobs. As workload demand fluctuations lead to varying idle capacity, efficient resource management is difficult, in particular given different provider objectives, e.g., utilization, revenue. Stefania Costache 0002, Tommaso Madonia, Asser N. Tantawi, Malgorzata Steinder |
SoCC | 4 |
| 2017 | Topology-aware GPU scheduling for learning workloads in cloud environmentsabstractRecent advances in hardware, such as systems with multiple GPUs and their availability in the cloud, are enabling deep learning in various domains including health care, autonomous vehicles, and Internet of Things. Multi-GPU systems exhibit complex connectivity among GPUs and between GPUs and CPUs. Workload schedulers must consider hardware topology and workload communication requirements in order to allocate CPU and GPU resources for optimal execution time and improved utilization in shared cloud environments. Marcelo Amaral, Jorda Polo, David Carrera 0001, Seetharami R. Seelam, Malgorzata Steinder |
SC | 5 |
| 2015 | Performance Evaluation of Microservices Architectures Using ContainersabstractMicro services architecture has started a new trend for application development for a number of reasons: (1) to reduce complexity by using tiny services, (2) to scale, remove and deploy parts of the system easily, (3) to improve flexibility to use different frameworks and tools, (4) to increase the overall scalability, and (5) to improve the resilience of the system. Containers have empowered the usage of micro services architectures by being lightweight, providing fast start-up times, and having a low overhead. Containers can be used to develop applications based on monolithic architectures where the whole system runs inside a single container or inside a micro services architecture where one or few processes run inside the containers. Two models can be used to implement a micro services architecture using containers: master-slave, or nested-container. The goal of this work is to compare the performance of CPU and network running benchmarks in the two aforementioned models of micro services architecture hence provide a benchmark analysis guidance for system designers. Marcelo Amaral, Jorda Polo, David Carrera 0001, Iqbal Mohomed, Merve Unuvar, Malgorzata Steinder |
NCA | 6 |
| 2015 | Selecting Optimum Cloud Availability Zones by Learning User Satisfaction LevelsabstractCloud service providers enable enterprises with the ability to place their business applications into availability zones across multiple locations worldwide. While this capability helps achieve higher availability with smaller failure rates, business applications deployed across these independent zones may experience different quality of service (QoS) due to heterogeneous physical infrastructures. Since the perceived QoS against specific requirements are not usually advertised by cloud providers, selecting an availability zone that would best satisfy the user requirements is a challenge. In this paper, we introduce a predictive approach to identify the cloud availability zone that maximizes satisfaction of an incoming request against a set of requirements. The prediction models are built from historical usage data for each availability zone and are updated as the nature of the zones and requests change. Simulation results show that our method successfully predicts the unpublished zone behavior from historical data and identifies the availability zone that maximizes user satisfaction against specific requirements. Merve Unuvar, Stefania Tosi, Yurdaer N. Doganata, Malgorzata Steinder, Asser N. Tantawi |
IEEE Trans. Serv. Comput. | 4 |
| 2014 | A Predictive Method for Identifying Optimum Cloud Availability ZonesabstractCloud service providers enable enterprises with the ability to place their business applications into availability zones across multiple locations worldwide. While this capability helps achieve higher availability with smaller failure rates, business applications deployed across these independent zones may experience different Quality of Service (QoS) due to heterogeneous physical infrastructures. Since the perceived QoS against specific requirements are not usually advertised by cloud providers, selecting an availability zone that would best satisfy the user requirements is a challenge. In this paper, we introduce a predictive approach to identify the cloud availability zone that maximizes satisfaction of an incoming request against a set of requirements. The predictive models are built from historical usage data for each availability zone and are updated as the nature of the zones and requests change. Simulation results show that our method successfully predicts the unpublished zone behavior from historical data and identifies the availability zone that maximizes user satisfaction against specific requirements. Merve Unuvar, Yurdaer N. Doganata, Malgorzata Steinder, Asser N. Tantawi, Stefania Tosi |
IEEE CLOUD | 3 |
| 2014 | Adaptive MapReduce Scheduling in Shared EnvironmentsabstractIn this paper we present a MapReduce task scheduler for shared environments in which MapReduce is executed along with other resource-consuming workloads, such as transactional applications. All workloads may potentially share the same data store, some of them consuming data for analytics purposes while others acting as data generators. This kind of scenario is becoming increasingly important in data centers where improved resource utilization can be achieved through workload consolidation, and is specially challenging due to the interaction between workloads of different nature that compete for limited resources. The proposed scheduler aims to improve resource utilization across machines while observing completion time goals. Unlike other MapReduce schedulers, our approach also takes into account the resource demands for non-MapReduce workloads, and assumes that the amount of resources made available to the MapReduce applications is variable over time. As shown in our experiments, our proposal improves the management of MapReduce jobs in the presence of variable resource availability, increasing the accuracy of the estimations made by the scheduler, thus improving completion time goals without an impact on the fairness of the scheduler. Jorda Polo, Yolanda Becerra 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé, Malgorzata Steinder |
CCGRID | 6 |
| 2014 | Monitoring applications and services to improve the Cloud Foundry PaaSabstractPlatform as a Service (PaaS) systems fully exploit the potential of elastic Cloud computing Infrastructure as a Service (IaaS) layer, by providing computational platforms for the developers characterized by a set of frameworks and runtimes. In these scenarios, the developers could focusing only on the implementation side of web applications without having to deal with configuration of the environment the web apps require for running properly. Services represent a central point of PaaS systems, providing external features for web applications such as SQL databases, messaging systems, and any kind of external software required by the developer. The monitoring of the availability and the performances of these Services plays an essential role in PaaS environments. The paper tackles above issues focusing on the real use case of the Cloud Foundry PaaS; collected results assess the effectiveness of the proposed monitoring function and confirm its feasibility and low overhead. Antonio Corradi, Luca Foschini 0001, Sebastiano Fraternale, Diana J. Arrojo, Malgorzata Steinder |
ISCC | 5 |
| 2014 | Hybrid Cloud Placement AlgorithmabstractA fully functional hybrid cloud solution requires a placement service to automatically decide whether an application should be deployed on premise, in a public cloud, or across private and public clouds. Such a service must consider application structure and communication patterns, application affinity requirements, which usually result from data protection rules, and deployment costs. In this paper, we propose a hybrid cloud placement approach which addresses these challenges. Our approach considers application requirements, cost, and private cloud capacity. Further, it is tunable to allow for changing application patterns and business objectives and offers a useful trade-off between application QoS and its deployment cost. Merve Unuvar, Malgorzata Steinder, Asser N. Tantawi |
MASCOTS | 2 |
| 2013 | Ripple: Improved Architecture and Programming Model for Bulk Synchronous Parallel Style of AnalyticsabstractWe present Ripple, an architecture and a programming model for a broad set of data analytics. Ripple builds on the ideas of iterated MapReduce and adds two innovations. First it has a richer programming model, including more ideas from the Bulk Synchronous Parallel (BSP) model of computation and others. By doing so, Ripple creates a flexible and higher-level platform that is easier for both application programmers and platform implementors. Second, Ripple is based on a limited interface for key/value storage making it portable among many different key/value store implementations. By building on these two ideas Ripple improves the scope, performance, and openness of the data analytics platform. We evaluate Ripple using three representative, and non-trivial, data analysis scenarios requiring iterative computation. Using these examples, we show how Ripple achieves clear performance advantages over iterated MapReduce. Mike Spreitzer, Malgorzata Steinder, Ian Whalley |
ICDCS | 2 |
| 2013 | Enabling Distributed Key-Value Stores with Low Latency-Impact Snapshot SupportabstractCurrent distributed key-value stores generally provide greater scalability at the expense of weaker consistency and isolation. However, additional isolation support is becoming increasingly important in the environments in which these stores are deployed, where different kinds of applications with different needs are executed, from transactional workloads to data analytics. While fully-fledged ACID support may not be feasible, it is still possible to take advantage of the design of these data stores, which often include the notion of multiversion concurrency control, to enable them with additional features at a much lower performance cost and maintaining its scalability and availability. In this paper we explore the effects that additional consistency guarantees and isolation capabilities may have on a state of the art key-value store: Apache Cassandra. We propose and implement a new multiversioned isolation level that provides stronger guarantees without compromising Cassandra's scalability and availability. As shown in our experiments, our version of Cassandra allows Snapshot Isolation-like transactions, preserving the overall performance and scalability of the system. Jorda Polo, Yolanda Becerra 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé, Mike Spreitzer, Malgorzata Steinder |
NCA | 7 |
| 2013 | Deadline-Based MapReduce Workload ManagementabstractThis paper presents a scheduling technique for multi-job MapReduce workloads that is able to dynamically build performance models of the executing workloads, and then use these models for scheduling purposes. This ability is leveraged to adaptively manage workload performance while observing and taking advantage of the particulars of the execution environment of modern data analytics applications, such as hardware heterogeneity and distributed storage. The technique targets a highly dynamic environment in which new jobs can be submitted at any time, and in which MapReduce workloads share physical resources with other workloads. Thus the actual amount of resources available for applications can vary over time. Beyond the formulation of the problem and the description of the algorithm and technique, a working prototype (called Adaptive Scheduler) has been implemented. Using the prototype and medium-sized clusters (of the order of tens of nodes), the following aspects have been studied separately: the scheduler's ability to meet high-level performance goals guided only by user-defined completion time goals; the scheduler's ability to favor data-locality in the scheduling algorithm; and the scheduler's ability to deal with hardware heterogeneity, which introduces hardware affinity and relative performance characterization for those applications that can benefit from executing on specialized processors. Jorda Polo, Yolanda Becerra 0001, David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2012 | Enabling Efficient Placement of Virtual Infrastructures in the Cloud
Ioana Giurgiu, Claris Castillo, Asser N. Tantawi, Malgorzata Steinder |
Middleware | 4 |
| 2012 | Cost-aware replication for dataflowsabstractIn this work we are concerned with the cost associated with replicating intermediate data for dataflows in Cloud environments. This cost is attributed to the extra resources required to create and maintain the additional replicas for a given data set. Existing data-analytic platforms such as Hadoop provide for fault-tolerance guarantee by relying on aggressive replication of intermediate data. We argue that the decision to replicate along with the number of replicas should be a function of the resource usage and utility of the data in order to minimize the cost of reliability. Furthermore, the utility of the data is determined by the structure of the dataflow and the reliability of the system. We propose a replication technique, which takes into account resource usage, system reliability and the characteristic of the dataflow to decide what data to replicate and when to replicate. The replication decision is obtained by solving a constrained integer programming problem given information about the dataflow up to a decision point. In addition, we built a working prototype, CARDIO of our technique which shows through experimental evaluation using a real testbed that finds an optimal solution. Claris Castillo, Asser N. Tantawi, Diana Arroyo, Malgorzata Steinder |
NOMS | 4 |
| 2012 | Autonomic Placement of Mixed Batch and Transactional WorkloadsabstractTo reduce the cost of infrastructure and electrical energy, enterprise datacenters consolidate workloads on the same physical hardware. Often, these workloads comprise both transactional and long-running analytic computations. Such consolidation brings new performance management challenges due to the intrinsically different nature of a heterogeneous set of mixed workloads, ranging from scientific simulations to multitier transactional applications. The fact that such different workloads have different natures imposes the need for new scheduling mechanisms to manage collocated heterogeneous sets of applications, such as running a web application and a batch job on the same physical server, with differentiated performance goals. In this paper, we present a technique that enables existing middleware to fairly manage mixed workloads: long running jobs and transactional applications. Our technique permits collocation of the workload types on the same physical hardware, and leverages virtualization control mechanisms to perform online system reconfiguration. In our experiments, including simulations as well as a prototype system built on top of state-of-the-art commercial middleware, we demonstrate that our technique maximizes mixed workload performance while providing service differentiation based on high-level performance goals. David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2011 | Towards efficient resource management for data-analytic platformsabstractWe present architectural and experimental work exploring the role of intermediate data handling in the performance of MapReduce workloads. Our findings show that: (a) certain jobs are more sensitive to disk cache size than others and (b) this sensitivity is mostly due to the local file I/O for the intermediate data. We also show that a small amount of memory is sufficient for the normal needs of map workers to hold their intermediate data until it is read. We introduce Hannibal, which exploits the modesty of that need in a simple and direct way — holding the intermediate data in application-level memory for precisely the needed time — to improve performance when the disk cache is stressed. We have implemented Hannibal and show through experimental evaluation that Hannibal can make MapReduce jobs run faster than Hadoop when little memory is available to the disk cache. This provides better performance insulation between concurrent jobs. Claris Castillo, Mike Spreitzer, Malgorzata Steinder |
Integrated Network Management | 3 |
| 2011 | Licence-aware management of virtual machinesabstractWhen managing dynamic virtual environments, a consideration that is often overlooked is that of licences. This consideration is particularly important when the environment contains expensive enterprise software with complex licensing terms: if the management system does not understand those terms, it is very likely that too many licences will be used, placing the administrator in an awkward position with the software vendors. We describe a licence management technique integrated with an existing virtual machine management system, and demonstrate its ability to adhere to a variety of licence rules. We also describe two sets of licence rules in sufficient detail to motivate the design and implementation of the system. We evaluate our technique through simulation. Ian Whalley, Malgorzata Steinder |
Integrated Network Management | 2 |
| 2011 | Resource-Aware Adaptive Scheduling for MapReduce Clusters
Jorda Polo, Claris Castillo, David Carrera 0001, Yolanda Becerra 0001, Ian Whalley, Malgorzata Steinder, Jordi Torres, Eduard Ayguadé |
Middleware | 6 |
| 2010 | Multi-aspect hardware management in enterprise server consolidationabstractAn autonomic manager for enterprise server hardware management, called AMP, is described. AMP is designed to handle multiple aspects of hardware management and to work in conjunction with other management components, in particular application managers, in a way that reduces energy waste, protects server health, and preserves a high degree of autonomy both for itself and for the managers with which it works. AMP interacts with other managers in two ways: (1) exchange of nominal control over individual servers; and (2) provision of a synthetic cost function giving AMP's assessment of relative desirability of using different servers. The high-level architecture of AMP is discussed, with particular focus on the way it effects a natural decomposition of the combined hardware-and-application management problem, and on initial versions of the algorithms it uses to manage server power states and determine the cost function. AMP's viability in practice is demonstrated via prototype implementation in which it operates on real servers in collaboration with a state-of-the-art application manager. The overall system behavior is investigated via simulation. James E. Hanson, Ian Whalley, Malgorzata Steinder, Jeffrey O. Kephart |
NOMS | 3 |
| 2010 | Runtime Demand Estimation for effective dynamic resource managementabstractSystems management techniques that allocate resources to running entities, such as processes and virtual machines (VMs), often require estimates of the resources required by each of these resource consumers. For example, many proposed virtual machine placement algorithms attempt to allocate VMs to physical hosts in such a way as to minimize the number of physical hosts that are occupied, while ensuring that each VM receives the CPU required to do its task adequately. The common practice is to assume that the CPU requirement is equal to the current CPU utilization, or to use a prediction of it over an appropriate time horizon. In this paper, we demonstrate that, when multiple VMs or processes co-reside on a physical host, the measured CPU utilization may provide a poor estimate of the actual requirement. We derive a simple, much more accurate alternative estimate of CPU demand, implement it, and demonstrate its superiority experimentally. Furthermore, we demonstrate that using our demand estimation framework in conjunction with dynamic resource allocation in a virtualized environment greatly improves the effectiveness of dynamic placement, resulting in one-shot convergence to optimal placement and significant improvements in the overall performance of the individual VMs. Canturk Isci, James E. Hanson, Ian Whalley, Malgorzata Steinder, Jeffrey O. Kephart |
NOMS | 4 |
| 2010 | Performance-driven task co-scheduling for MapReduce environmentsabstractMapReduce is a data-driven programming model proposed by Google in 2004 which is especially well suited for distributed data analytics applications. We consider the management of MapReduce applications in an environment where multiple applications share the same physical resources. Such sharing is in line with recent trends in data center management which aim to consolidate workloads in order to achieve cost and energy savings. In a shared environment, it is necessary to predict and manage the performance of workloads given a set of performance goals defined for them. In this paper, we address this problem by introducing a new task scheduler for a MapReduce framework that allows performance-driven management of MapReduce tasks. The proposed task scheduler dynamically predicts the performance of concurrent MapReduce jobs and adjusts the resource allocation for the jobs. It allows applications to meet their performance objectives without over-provisioning of physical resources. Jorda Polo, David Carrera 0001, Yolanda Becerra 0001, Malgorzata Steinder, Ian Whalley |
NOMS | 4 |
| 2010 | Decentralized allocation of CPU computation power for web applications
Shrutivandana Sharma, Asser N. Tantawi, Mike Spreitzer, Malgorzata Steinder |
Perform. Evaluation | 4 |
| 2008 | Managing SLAs of heterogeneous workloads using dynamic application placementabstractIn this paper we address the problem of managing heterogeneous workloads in a virtualized data center. We consider two different workloads: transactional applications and long-running jobs. We present a technique that permits collocation of these workload types on the same physical hardware. Our technique dynamically modifies workload placement by leveraging control mechanisms such as suspension and migration, and strives to optimally trade off resource allocation among these workloads in spite of their differing characteristics and performance objectives. Our approach builds upon our previous work on dynamically placing transactional workloads. This paper extends our framework with the capability to manage long-running workloads. We achieve this goal by using utility functions, which permit us to compare the performance of various workloads, and which are used to drive allocation decisions. We demonstrate that our technique maximizes heterogeneous workload performance while providing service differentiation based on high-level performance goals. David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
HPDC | 2 |
| 2008 | Enabling Resource Sharing between Transactional and Batch Workloads Using Dynamic Application Placement
David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
Middleware | 2 |
| 2008 | Utility-based placement of dynamic Web applications with fairness goalsabstractWe study the problem of dynamic resource allocation to clustered Web applications. We extend application server middleware with the ability to automatically decide the size of application clusters and their placement on physical machines. Unlike existing solutions, which focus on maximizing resource utilization and may unfairly treat some applications, the approach introduced in this paper considers the satisfaction of each application with a particular resource allocation and attempts to at least equally satisfy all applications. We model satisfaction using utility functions, mapping CPU resource allocation to the performance of an application relative to its objective. The demonstrated online placement technique aims at equalizing the utility value across all applications while also satisfying operational constraints, preventing the over-allocation of memory, and minimizing the number of placement changes. We have implemented our technique in a leading commercial middleware product. Using this real-life testbed and a simulation we demonstrate the benefit of the utility-driven technique as compared to other state-of-the-art techniques. David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
NOMS | 2 |
| 2008 | Coordinated management of power usage and runtime performanceabstractWith the continued growth of computing power and reduction in physical size of enterprise servers, the need for actively managing electrical power usage in large datacenters is becoming ever more pressing. By far the greatest savings in electrical power can be effected by dynamically consolidating workload onto the minimum number of servers needed at a given time and powering off the remainder. However, simple schemes for achieving this goal fail to cope with the complexities of realistic usage scenarios. In this paper we present a combined power-and performance-management system that builds on a state-of-the-art performance manager to achieve significant power savings without unacceptable loss of performance. In our system, the degree to which performance may be traded off against power is itself adjustable using a small number of easily-understood parameters, permitting administrators in different facilities to select the optimal tradeoff for their needs. We characterize the power saved, the effects of the tradeoff between power and performance, and the changes in behavior as the tradeoff parameters are adjusted, both in simulation and in a sample deployment of the real system. Malgorzata Steinder, Ian Whalley, James E. Hanson, Jeffrey O. Kephart |
NOMS | 1 |
| 2007 | A Service Middleware that Scales in System Size and ApplicationsabstractWe present a peer-to-peer service management middleware that dynamically allocates system resources to a large set of applications. The system achieves scalability in number of nodes (1000s or more) through three decentralized mechanisms that run on different time scales. First, overlay construction interconnects all nodes in the system for exchanging control and state information. Second, request routing directs requests to nodes that offer the corresponding applications. Third, application placement controls the set of offered applications on each node, in order to achieve efficient operation and service differentiation. The design supports a large number of applications (100s or more) through selective propagation of configuration information needed for request routing. The control load on a node increases linearly with the number of applications in the system. Service differentiation is achieved through assigning a utility to each application, which influences the application placement process. Simulation studies show that the system operates efficiently for different sizes, adapts fast to load changes and failures and effectively differentiates between different applications under overload. Constantin Adam, Rolf Stadler, Chunqiang Tang, Malgorzata Steinder, Mike Spreitzer |
Integrated Network Management | 4 |
| 2007 | Server virtualization in autonomic management of heterogeneous workloadsabstractServer virtualization opens up a range of new possibilities for autonomic datacenter management, through the availability of new automation mechanisms that can be exploited to control and monitor tasks running within virtual machines. This offers not only new and more flexible control to the operator using a management console, but also more powerful and flexible autonomic control, through management software that maintains the system in a desired state in the face of changing workload and demand. This paper explores in particular the use of server virtualization technology in the autonomic management of data centers running a heterogeneous mix of workloads. We present a system that manages heterogeneous workloads to their performance goals and demonstrate its effectiveness via real-system experiments and simulation. We also present some of the significant challenges to wider usage of virtual servers in autonomic datacenter management. Malgorzata Steinder, Ian Whalley, David Carrera 0001, Ilona Gaweda, David M. Chess |
Integrated Network Management | 1 |
| 2007 | A scalable application placement controller for enterprise data centersabstractGiven a set of machines and a set of Web applications with dynamically changing demands, an online application placement controller decides how many instances to run for each application and where to put them, while observing all kinds of resource constraints. This NP hard problem has real usage in commercial middleware products. Existing approximation algorithms for this problem can scale to at most a few hundred machines, and may produce placement solutions that are far from optimal when system resources are tight. In this paper, we propose a new algorithm that can produce within 30 seconds high-quality solutions for hard placement problems with thousands of machines and thousands of applications. This scalability is crucial for dynamic resource provisioning in large-scale enterprise data centers. Our algorithm allows multiple applications to share a single machine, and strives to maximize the total satisfied application demand, to minimize the number of application starts and stops, and to balance the load across machines. Compared with existing state-of-the-art algorithms, for systems with 100 machines or less, our algorithm is up to 134 times faster, reduces application starts and stops by up to 97%, and produces placement solutions that satisfy up to 25% more application demands. Our algorithm has been implemented and adopted in a leading commercial middleware product for managing the performance of Web applications. Chunqiang Tang, Malgorzata Steinder, Mike Spreitzer, Giovanni Pacifici |
WWW | 2 |
| 2007 | Multidomain Diagnosis of End-to-End Service Failures in Hierarchically Routed NetworksabstractProbabilistic inference was shown effective in the nondeterministic diagnosis of end-to-end service failures when applied in a centralized management system where the manager possesses a global knowledge of the system structure and state. Since many networks are organized into multiple administrative domains that may be unable to share configuration and state information, these centralized techniques are not applicable to them. This paper proposes a fault localization technique suitable for multidomain networks with hierarchical routing. The proposed technique divides the computational effort and system knowledge among multiple, hierarchically organized managers. Each manager performs fault localization in the domain it manages and requires only the knowledge of its own domain. We show through simulation that the proposed approach not only improves the feasibility of fault localization in multidomain networks, but also increases the effectiveness of probabilistic diagnosis and makes it realizable in networks of considerable size Malgorzata Steinder, Adarshpal S. Sethi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2006 | Dynamic placement for clustered web applicationsabstractWe introduce and evaluate a middleware clustering technology capable of allocating resources to web applications through dynamic application instance placement. We define application instance placement as the problem of placing application instances on a given set of server machines to adjust the amount of resources available to applications in response to varying resource demands of application clusters. The objective is to maximize the amount of demand that may be satisfied using a configured placement. To limit the disturbance to the system caused by starting and stopping application instances, the placement algorithm attempts to minimize the number of placement changes. It also strives to keep resource utilization balanced across all server machines. Two types of resources are managed, one load-dependent and one load-independent. When putting the chosen placement in effect our controller schedules placement changes in a manner that limits the disruption to the system. Alexei A. Karve, Tracy Kimbrel, Giovanni Pacifici, Mike Spreitzer, Malgorzata Steinder, Maxim Sviridenko, Asser N. Tantawi |
WWW | 5 |
| 2004 | Multi-domain Diagnosis of End-to-End Service Failures in Hierarchically Routed NetworksabstractThis paper investigates an approach to improving the scalability and feasibility of probabilistic fault localization in communication systems by exploiting the domain semantics of computer networks. The proposed technique divides the computational effort and system knowledge among multiple, hierarchically organized managers. Each manager performs fault localization in the domain it manages and requires only the knowledge of its own domain. Since failures propagate among domains, domain managers cooperate with each other to find a consensus explanation of the observed disorder. We show through simulation that the proposed approach increases the effectiveness of probabilistic diagnosis and makes it feasible in networks of considerable size 1. Malgorzata Steinder, Adarshpal S. Sethi |
NETWORKING | 1 |
| 2004 | Probabilistic fault diagnosis in communication systems through incremental hypothesis updating
Malgorzata Steinder, Adarshpal S. Sethi |
Comput. Networks | 1 |
| 2004 | A survey of fault localization techniques in computer networks
Malgorzata Steinder, Adarshpal S. Sethi |
Sci. Comput. Program. | 1 |
| 2004 | Probabilistic fault localization in communication systems using belief networksabstractWe apply Bayesian reasoning techniques to perform fault localization in complex communication systems while using dynamic, ambiguous, uncertain, or incorrect information about the system structure and state. We introduce adaptations of two Bayesian reasoning techniques for polytrees, iterative belief updating, and iterative most probable explanation. We show that these approximate schemes can be applied to belief networks of arbitrary shape and overcome the inherent exponential complexity associated with exact Bayesian reasoning. We show through simulation that our approximate schemes are almost optimally accurate, can identify multiple simultaneous faults in an event driven manner, and incorporate both positive and negative information into the reasoning process. We show that fault localization through iterative belief updating is resilient to noise in the observed symptoms and prove that Bayesian reasoning can now be used in practice to provide effective fault localization. Malgorzata Steinder, Adarshpal S. Sethi |
IEEE/ACM Trans. Netw. | 1 |
| 2003 | Probabilistic Event-driven Fault Diagnosis Through Incremental Hypothesis Updating
Malgorzata Steinder, Adarshpal S. Sethi |
Integrated Network Management | 1 |
| 2002 | Increasing robustness of fault localization through analysis of lost, spurious, and positive symptomsabstractThis paper utilizes belief networks to implement fault localization in communication systems taking into account comprehensive information about the system behavior. Most previous work on this subject performs fault localization based solely on the information about malfunctioning system components (i.e., negative symptoms). We show that positive information, i.e., the lack of any disorder in some system components, may be used to improve the accuracy of this process. The technique presented allows lost and spurious symptoms to be incorporated in the analysis. We show through simulation that in a noisy network environment the analysis of lost and spurious symptoms increases the robustness of fault localization with belief networks. We also demonstrate that belief networks yield high accuracy even for approximate probability input data and therefore are a promising model for non-deterministic fault localization. Malgorzata Steinder, Adarshpal S. Sethi |
INFOCOM | 1 |
| 2002 | End-to-end service failure diagnosis using belief networksabstractWe present fault localization techniques suitable for diagnosing end-to-end service problems in communication systems with complex topologies. We refine a layered system model that represents relationships between services and functions offered between neighboring protocol layers. In a given layer, an end-to-end service between two hosts may be provided using multiple host-to-host services offered in this layer between two hosts on the end-to-end path. Relationships among end-to-end and host-to-host services form a bipartite probabilistic dependency graph whose structure depends on the network topology in the corresponding protocol layer. When an end-to-end service fails or experiences performance problems it is important to efficiently find the responsible host-to-host services. Finding the most probable explanation (MPE) of the observed symptoms is NP-hard. We propose two fault localization techniques based on Pearl's (1988) iterative algorithms for singly connected belief networks. The probabilistic dependency graph is transformed into a belief network, and then the approximations based on Pearl's algorithms and exact bucket tree elimination algorithm are designed and evaluated through extensive simulation study. Malgorzata Steinder, Adarshpal S. Sethi |
NOMS | 1 |
| 2001 | Non-deterministic diagnosis of end-to-end service failures in a multi-layer communication systemabstractFault localization is a process of isolating faults responsible for the observable malfunctioning of the managed system. Previously, fault localization efforts concentrated mostly on diagnosing faults related to the availability of network resources in the lowest layers of the protocol stack. This paper focuses on end-to-end service failure diagnosis as a critical step towards multi-layer fault localization in an enterprise environment. By refining a previously proposed modeling technique, we present a universal method of modeling both availability and performance related problems associated with end-to-end services in a non-deterministic fashion. We introduce and evaluate a novel algorithm that allows an event-driven, incremental diagnosis of end-to-end service failures. Malgorzata Steinder, Adarshpal S. Sethi |
ICCCN | 1 |
| 2001 | Yemanja - A Layered Event Correlation Engine for Multi-domain Server FarmsabstractYemanja is a model-based event correlation engine for multi-layer fault diagnosis. It targets complex propagating fault scenarios, and can smoothly correlate low-level network events with high-level application performance alerts related to quality of service violations. Entity-models that represent devices or abstract components encapsulate entity behavior. Distantly associated entities are not explicitly aware of each other, and communicate through event propagation chains. Yemanja's state-based engine supports generic scenario definitions, prioritization of alternate solutions, integrated problem-state and device testing, and simultaneous analysis of overlapping problems. The system of correlation rules was developed based on device, layer, and dependency analysis, and reveals the layered structure of computer networks. The primary objectives of this research include the development of reusable, configuration independent, correlation scenarios; adaptability and the extensibility of the engine to match the constantly changing topology of a multi-domain server farm; and the development of a concise specification language that is relatively simple yet powerful. Karen Appleby, Germán S. Goldszmidt, Malgorzata Steinder |
Integrated Network Management | 3 |