VLDB 2026 Research / reviewers in the wild / expert
Evgenia Smirni
dblp:s/EvgeniaSmirni
· DBLP profile ↗
120ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0001-8754-581XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 72 · 7 first-author · 12 since 2021Software engineering, systems software and programming languages · 23 · 2 first-author · 3 since 2021Security and privacy · 17 · 5 since 2021Computer networks · 10 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Green or Fast? Learning to Balance Cold Starts and Idle Carbon in Serverless Computing
Christos D. Antonopoulos, Evgenia Smirni, Bin Ren 0002, Nikolaos Bellas, Spyros Lalis |
CCGrid | 3 |
| 2026 | PeakLife: Proactive VM Management via Joint Forecasting
Christos D. Antonopoulos, Evgenia Smirni, Bin Ren 0002, Nikolaos Bellas, Spyros Lalis |
Euro-Par (2) | 3 |
| 2026 | Mantis: Decoding HPC Telemetry Data for Robust System Prediction
Yiyang Lu 0001, Jie Ren 0015, Evgenia Smirni |
ICS | 3 |
| 2025 | Safety Interventions against Adversarial Patches in an Open-Source Driver Assistance SystemabstractDrivers are becoming increasingly reliant on advanced driver assistance systems (ADAS) as autonomous driving technology becomes more popular and developed with advanced safety features to enhance road safety. However, the increasing complexity of the ADAS makes autonomous vehicles (AVs) more exposed to attacks and accidental faults. In this paper, we evaluate the resilience of a widely used ADAS against safety-critical attacks that target perception inputs. Various safety mechanisms are simulated to assess their impact on mitigating attacks and enhancing ADAS resilience. Experimental results highlight the importance of timely intervention by human drivers and automated safety mechanisms in preventing accidents in both driving and lateral directions and the need to resolve conflicts among safety interventions to enhance system resilience and reliability. [Code Available at https://doi.org/10.6084/m9.figshare.28691090] Grant Xiao, Daehyun Lee, Lishan Yang 0001, Evgenia Smirni, Homa Alemzadeh, Xugui Zhou |
DSN | 5 |
| 2025 | TMModel: Modeling Texture Memory and Mobile GPU Performance to Accelerate DNN ComputationsabstractThe demand for Deep Neural Network (DNN) execution (including both inference and training) on mobile system-ona-chip (SoCs) has surged, driven by factors like the need for real-time latency, privacy, and reducing vendors' costs.Mainstream mobile GPUs (e.g., Qualcomm Adreno GPUs) usually have a 2.5D L1 texture cache that offers throughput superior to that of on-chip memory.However, to date, there is limited understanding of the performance features of such a 2.5D cache, which limits the optimization potential.This paper introduces TMModel, a framework with three components: 1) a set of micro-benchmarks and a novel performance assessment methodology to characterize a non-well-documented architecture with 2D memory, 2) a complete analytical performance model configurable for different data access pattern(s), tiling size(s), and other GPU execution parameters for a given operator (and associated size and shape), and 3) a compilation framework incorporating this model and generating optimized code with low overhead.TMModel is Jiexiong Guan, Zhenqing Hu, Christos D. Antonopoulos, Nikolaos Bellas, Spyros Lalis, Evgenia Smirni, Gang Zhou 0002, Gagan Agrawal, Bin Ren 0002 |
ICS | 6 |
| 2025 | DeepBAT: Performance and Cost Optimization of Serverless Inference Using TransformersabstractServerless computing is an autoscaling pay-as-yougo paradigm that can efficiently support machine learning inference especially under bursty workload conditions. Within the serverless paradigm, batching ML inference requests before serving them is widely adopted. Thanks to its parallelism properties, batching can highly improve inference performance while reducing the monetary cost of serverless. Identifying the correct serverless parameterization to simultaneously meet conflicting targets (i.e., keep monetary cost at a minimum while meeting pre-defined service level objectives, SLO) may be cast as a resource allocation problem. In this paper, we illustrate that a deep surrogate model can quickly discover optimized serverless configurations by learning the relationship among the workload patterns and achieve performance measures. We develop DeepBAT, an SLO-aware framework that leverages the Transformer encoder and multi-head attention mechanism to optimize the performance of serverless inference subject to bursty and previously unobserved workloads. We illustrate the effectiveness of DeepBAT on a set of case studies and show that for the problem of inference serving on AWS Lambda, DeepBAT can speed up the solution time of state-of-the-art analytic solutions by over 55 times while generalizing remarkably well on unseen workloads. Riccardo Pinciroli, Giuliano Casale, Evgenia Smirni |
IPDPS | 4 |
| 2025 | Serverless Computing for Next-generation Application DevelopmentabstractServerless computing is a cloud computing model that abstracts server management, allowing developers to focus solely on writing code without concerns about the underlying infrastructure. This paradigm shift is transforming application development by reducing time to market, lowering costs, and enhancing scalability. In serverless computing, functions are event-driven and automatically scale in response to events such as data changes or user requests. Despite its advantages, serverless computing presents several research challenges, including managing state for ephemeral functions, mitigating cold start delays, optimizing function composition, debugging, efficient auto-scaling, resource management, and ensuring security and compliance. This special issue focused on addressing these challenges by promoting research on innovative solutions and exploring the potential of serverless computing in new application domains. Adel Nadjaran Toosi, Bahman Javadi, Alexandru Iosup, Evgenia Smirni, Schahram Dustdar |
Future Gener. Comput. Syst. | 4 |
| 2024 | GPU Reliability Assessment: Insights Across the Abstraction LayersabstractGraphics Processing Units (GPUs) are widely de-ployed and utilized across various computing domains including cloud and high-performance computing. Considering its extensive usage and increasing popularity, ensuring GPU reliability is cru-cial. Software-based reliability evaluation methodologies, though fast, often neglect the complex hardware details of modern GPU designs. This oversight could lead to misleading measurements and misguided decisions regarding protection strategies. This paper breaks new ground by conducting an in-depth examination of well-established vulnerability assessment methods for modern GPU architectures, from the microarchitecture all the way to the software layers. It highlights divergences between popular software-based vulnerability evaluation methods and the ground truth cross-layer evaluation, which persist even under strong protections like triple modular redundancy. Accurate evaluation requires considering fault distribution from hardware to software. Our comprehensive measurements offer valuable insights into the accurate assessment of GPU reliability. Lishan Yang 0001, George Papadimitriou 0001, Dimitris Sartzetakis, Adwait Jog, Evgenia Smirni, Dimitris Gizopoulos |
CLUSTER | 5 |
| 2024 | Understanding GPU Memory Corruption at Extreme Scale: The Summit Case StudyabstractGPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted computations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environmental footprint, the efficiency of HPC operations becomes both an imperative and a challenge. Vladyslav Oles, Anna Schmedding, George Ostrouchov, Woong Shin, Evgenia Smirni, Christian Engelmann |
ICS | 5 |
| 2024 | Probing Weaknesses in GPU Reliability Assessment: A Cross-Layer ApproachabstractDue to extensive deployment and heavy usage of GPUs, ensuring the reliability of such devices is crucial. Current software-based reliability evaluation methodologies, albeit fast, often neglect the intricate hardware complexities of modern GPU designs. This oversight could result in misleading measurements and misguided decisions regarding protection strategies. This work breaks new ground by examining well-established vulnerability assessment methods for modern GPU architectures, from the microarchitecture all the way to the software layers. It highlights divergences between popular software-based vulnerability evaluation methods and the ground truth cross-layer evaluation (which, as we show, holds even when strong protection like triple modular redundancy is employed); accurate evaluation requires considering fault distribution from hardware to software. Our comprehensive measurements offer valuable insights into accurately assessing GPU reliability. Lishan Yang 0001, George Papadimitriou 0001, Dimitris Sartzetakis, Adwait Jog, Evgenia Smirni, Dimitris Gizopoulos |
ISPASS | 5 |
| 2024 | Aspis: Lightweight Neural Network Protection Against Soft ErrorsabstractConvolutional neural networks (CNN) are incorporated into many image-based tasks across a variety of domains. Some of these are safety critical tasks such as object classification/detection and lane detection for self-driving cars. These applications have strict safety requirements and must guarantee the reliable operation of the neural networks in the presence of soft errors (i.e., transient faults) in DRAM. Standard safety mechanisms (e.g., triplication of data/computation) provide high resilience, but introduce intolerable overhead. We perform detailed characterization and propose an efficient methodology for pinpointing critical weights by using an efficient proxy, the Taylor criterion. Using this characterization, we design Aspis, an efficient software protection scheme that does selective weight hardening and offers a performance/reliability tradeoff. Aspis provides higher resilience comparing to state-of-the-art methods and is integrated into PyTorch as a fully-automated library. Anna Schmedding, Lishan Yang 0001, Adwait Jog, Evgenia Smirni |
ISSRE | 4 |
| 2024 | Strategic Resilience Evaluation of Neural Networks Within Autonomous Vehicle Software
Anna Schmedding, Philip Schowitz, Xugui Zhou, Yiyang Lu 0001, Lishan Yang 0001, Homa Alemzadeh, Evgenia Smirni |
SAFECOMP | 7 |
| 2023 | Epidemic Spread Modeling for COVID-19 Using Cross-Fertilization of Mobility DataabstractWe present an individual-centric model for COVID-19 spread in an urban setting. We first analyze patient and route data of infected patients from January 20, 2020, to May 31, 2020, collected by the Korean Center for Disease Control & Prevention (KCDC) and discover how infection clusters develop as a function of time. This analysis offers a statistical characterization of mobility habits and patterns of individuals at the beginning of the pandemic. While the KCDC data offer a wealth of information, they are also by their nature limited. To compensate for their limitations, we use detailed mobility data from Berlin, Germany after observing that mobility of individuals is surprisingly similar in both Berlin and Seoul. Using information from the Berlin mobility data, we cross-fertilize the KCDC Seoul data set and use it to parameterize an agent-based simulation that models the spread of the disease in an urban environment. After validating the simulation predictions with ground truth infection spread in Seoul, we study the importance of each input parameter on the prediction accuracy, compare the performance of our model to state-of-the-art approaches, and show how to use the proposed model to evaluate differentwhat-ifcounter-measure scenarios. Anna Schmedding, Riccardo Pinciroli, Lishan Yang 0001, Evgenia Smirni |
IEEE Trans. Big Data | 4 |
| 2023 | Lifespan and Failures of SSDs and HDDs: Similarities, Differences, and Prediction ModelsabstractData center downtime typically centers around IT equipment failure. Storage devices are the most frequently failing components in data centers. We present a comparative study of hard disk drives (HDDs) and solid state drives (SSDs) that constitute the typical storage in data centers. Using six-year field data of 100,000 HDDs of different models from the same manufacturer from the Backblaze dataset and six-year field data of 30,000 SSDs of three models from a Google data center, we characterize the workload conditions that lead to failures. We illustrate that their root failure causes differ from common expectations and that they remain difficult to discern. For the case of HDDs we observe that young and old drives do not present many differences in their failures. Instead, failures may be distinguished by discriminating drives based on the time spent for head positioning. For SSDs, we observe high levels of infant mortality and characterize the differences between infant and non-infant failures. We develop several machine learning failure prediction models that are shown to be surprisingly accurate, achieving high recall and low false positive rates. These models are used beyond simple prediction as they aid us to untangle the complex interaction of workload characteristics that lead to failures and identify failure root causes from monitored symptoms. Riccardo Pinciroli, Lishan Yang 0001, Jacob Alter, Evgenia Smirni |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2022 | Strategic Safety-Critical Attacks Against an Advanced Driver Assistance SystemabstractA growing number of vehicles are being transformed into semi-autonomous vehicles (Level 2 autonomy) by relying on advanced driver assistance systems (ADAS) to improve the driving experience. However, the increasing complexity and connectivity of ADAS expose the vehicles to safety-critical faults and attacks. This paper investigates the resilience of a widely-used ADAS against safety-critical attacks that target the control system at opportune times during different driving scenarios and cause accidents. Experimental results show that our proposed Context-Aware attacks can achieve an 83.4% success rate in causing hazards, 99.7% of which occur without any warnings. These results highlight the intolerance of ADAS to safety-critical attacks and the importance of timely interventions by human drivers or automated recovery mechanisms to prevent accidents. Xugui Zhou, Anna Schmedding, Haotian Ren, Lishan Yang 0001, Philip Schowitz, Evgenia Smirni, Homa Alemzadeh |
DSN | 6 |
| 2022 | Optimizing Inference Serving on Serverless PlatformsabstractServerless computing is gaining popularity for machine learning (ML) serving workload due to its autonomous resource scaling, easy to use and pay-per-use cost model. Existing serverless platforms work well for image-based ML inference, where requests are homogeneous in service demands. That said, recent advances in natural language processing could not fully benefit from existing serverless platforms as their requests are intrinsically heterogeneous. Batching requests for processing can significantly increase ML serving efficiency while reducing monetary cost, thanks to the pay-per-use pricing model adopted by serverless platforms. Yet, batching heterogeneous ML requests leads to additional computation overhead as small requests need to be "padded" to the same size as large requests within the same batch. Reaching effective batching decisions (i.e., which requests should be batched together and why) is non-trivial: the padding overhead coupled with the serverless auto-scaling forms a complex optimization problem. To address this, we develop Multi-Buffer Serving (MBS), a framework that optimizes the batching of heterogeneous ML inference serving requests to minimize their monetary cost while meeting their service level objectives (SLOs). The core of MBS is a performance and cost estimator driven by analytical models supercharged by a Bayesian optimizer. MBS is prototyped and evaluated on AWS using bursty workloads. Experimental results show that MBS preserves SLOs while outperforming the state-of-the-art by up to 8 x in terms of cost savings while minimizing the padding overhead by up to 37 x with 3 x less number of serverless function invocations. Riccardo Pinciroli, Feng Yan 0001, Evgenia Smirni |
Proc. VLDB Endow. | 4 |
| 2021 | Data-centric Reliability Management in GPUsabstractGraphics Processing Units (GPUs) have become the default choice of acceleration in a wide range of application domains. To keep up with computational demands, the GPU memory system is constantly being innovated from both the cache and DRAM perspectives. Such innovations can adversely affect GPU reliability and in fact, can lead to an increase in the number of multi-bit faults. To address this problem, we systematically study a wide range of GPGPU applications and find that usually, only a small percentage of data needs protection to increase application resilience. This data is highly accessed and shared (constitutes hot memory), which implies that faults in this space can often lead to incorrect application output. An in-depth analysis of application code shows that information of such data can be passed on to the hardware to guide low-overhead detection/correction schemes. In this vein, we developed low-overhead partial data replication schemes that exploit latency tolerance in GPUs. Overall, this data-centric approach dramatically improves GPGPU application resilience, with a minimal additional average performance overhead of 1.2% for detection-only and 3.4% for detection-and-correction. Gurunath Kadam, Evgenia Smirni, Adwait Jog |
DSN | 2 |
| 2021 | Enabling Software Resilience in GPGPU Applications via Partial Thread ProtectionabstractGraphics Processing Units (GPUs) are widely used by various applications in a broad variety of fields to accelerate their computation but remain susceptible to transient hardware faults (soft errors) that can easily compromise application output. By taking advantage of a general purpose GPU application hierarchical organization in threads, warps, and cooperative thread arrays, we propose a methodology that identifies the resilience of threads and aims to map threads with the same resilience characteristics to the same warp. This allows to engage partial replication mechanisms for error detection/correction at the warp level. By exploring 12 benchmarks (17 kernels) from 4 benchmark suites, we illustrate that threads can be remapped into reliable or unreliable warps with only 1.63% introduced overhead (on average), and then enable selective protection via replication to those groups of threads that truly need it. Furthermore, we show that thread remapping to different warps does not sacrifice application performance. We show how this remapping facilitates warp replication for error detection and/or correction and achieves average reduction of 20.61% and 27.15% execution cycles, respectively comparing to standard duplication/triplication. Lishan Yang 0001, Bin Nie, Adwait Jog, Evgenia Smirni |
ICSE | 4 |
| 2021 | Practical Resilience Analysis of GPGPU Applications in the Presence of Single- and Multi-Bit FaultsabstractGraphics Processing Units (GPUs) have rapidly evolved to enable energy-efficient data-parallel computing for a broad range of scientific areas. While GPUs achieve exascale performance at a stringent power budget, they are also susceptible to soft errors, often caused by high-energy particle strikes, that can significantly affect the application output quality. Understanding the resilience of general purpose GPU (GPGPU) applications is especially challenging because unlike CPU applications, which are mostly single-threaded, GPGPU applications can contain hundreds to thousands of threads, resulting in a tremendously large fault site space in the order of billions, even for some simple applications and even when considering the occurrence of just a single-bit fault. We present a systematic way to progressively prune the fault site space aiming to dramatically reduce the number of fault injections such that assessment for GPGPU application error resilience becomes practical. The key insight behind our proposed methodology stems from the fact that while GPGPU applications spawn a lot of threads, many of them execute the same set of instructions. Therefore, several fault sites are redundant and can be pruned by careful analysis. We identify important features across a set of 10 applications (16 kernels) from Rodinia and Polybench suites and conclude that threads can be primarily classified based on the number of the dynamic instructions they execute. We therefore achieve significant fault site reduction by analyzing only a small subset of threads that are representative of the dynamic instruction behavior (and therefore error resilience behavior) of the GPGPU applications. Further pruning is achieved by identifying the dynamic instruction commonalities (and differences) across code blocks within this representative set of threads, a subset of loop iterations within the representative threads, and a subset of destination register bit positions. The above steps result in a tremendous reduction of fault sites by up to seven orders of magnitude. Yet, this reduced fault site space accurately captures the error resilience profile of GPGPU applications. We show the effectiveness of the proposed progressive pruning technique for a single-bit model and illustrate its application to even more challenging cases with three distinct multi-bit fault models. Lishan Yang 0001, Bin Nie, Adwait Jog, Evgenia Smirni |
IEEE Trans. Computers | 4 |
| 2021 | CEDULE+: Resource Management for Burstable Cloud Instances Using Predictive AnalyticsabstractNearly all principal cloud providers now provide burstable instances in their offerings. The main attraction of this type of instance is that it can boost its performance for a limited time to cope with workload variations. Although burstable instances are widely adopted, it is not clear how to efficiently manage them to avoid waste of resources. In this article, we use predictive data analytics to optimize the management of burstable instances. We design CEDULE+, a data-driven framework that enables efficient resource management for burstable cloud instances by analyzing the system workload and latency data. CEDULE+ selects the most profitable instance type to process incoming requests and controls CPU, I/O, and network usage to minimize the resource waste without violating Service Level Objectives (SLOs). CEDULE+ uses lightweight profiling and quantile regression to build a data-driven prediction model that estimates system performance for all combinations of instance type, resource type, and system workload. CEDULE+ is evaluated on Amazon EC2, and its efficiency and high accuracy are assessed through real-case scenarios. CEDULE+ predicts application latency with errors less than 10%, extends the maximum performance period of a burstable instance up to 2.4 times, and decreases deployment costs by more than 50%. Riccardo Pinciroli, Feng Yan 0001, Evgenia Smirni |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2020 | Characterizing Accuracy-Aware Resilience of GPGPU ApplicationsabstractGraphics Processing Units (GPUs) have rapidly evolved to enable energy-efficient data-parallel computing. In addition to achieving exascale performance at a stringent power budget, it is imperative for GPUs to provide reliable computing guarantees to the end user. In current commodity systems, such guarantees are often achieved by incurring high protection cost in terms of performance, power, and hardware resources. However, we argue that these strict guarantees are often not required (and that the associated protected overheads can be significantly reduced) because several GPGPU applications are either fault-tolerant or can accept a quantifiable loss in output quality. To this end, this paper characterizes in a hierarchical manner the accuracy-aware resilience of GPGPU applications consisting of thousands of threads. This characterization study shows that accuracy-aware error resilience exhibits several interesting patterns across threads at different hierarchies (i.e., kernel/thread-block/warp). The insights from this characterization study can be used to reduce the overheads of expensive protection or recovery mechanisms that are typically used by GPUs to ensure application reliability. Bin Nie, Adwait Jog, Evgenia Smirni |
CCGRID | 3 |
| 2020 | Mining Multivariate Discrete Event Sequences for Knowledge Discovery and Anomaly DetectionabstractModern physical systems deploy large numbers of sensors to record at different time-stamps the status of different systems components via measurements such as temperature, pressure, speed, but also the component's categorical state. Depending on the measurement values, there are two kinds of sequences: continuous and discrete. For continuous sequences, there is a host of state-of-the-art algorithms for anomaly detection based on time-series analysis, but there is a lack of effective methodologies that are tailored specifically to discrete event sequences. This paper proposes an analytics framework for discrete event sequences for knowledge discovery and anomaly detection. During the training phase, the framework extracts pairwise relationships among discrete event sequences using a neural machine translation model by viewing each discrete event sequence as a "natural language". The relationship between sequences is quantified by how well one discrete event sequence is "translated" into another sequence. These pairwise relationships among sequences are aggregated into a multivariate relationship graph that clusters the structural knowledge of the underlying system and essentially discovers the hidden relationships among discrete sequences. This graph quantifies system behavior during normal operation. During testing, if one or more pairwise relationships are violated, an anomaly is detected. The proposed framework is evaluated on two real-world datasets: a proprietary dataset collected from a physical plant where it is shown to be effective in extracting sensor pairwise relationships for knowledge discovery and anomaly detection, and a public hard disk drive dataset where its ability to effectively predict upcoming disk failures is illustrated. Bin Nie, Jianwu Xu, Jacob Alter, Evgenia Smirni |
DSN | 5 |
| 2020 | Batch: machine learning inference serving on serverless platforms with adaptive batchingabstractServerless computing is a new pay-per-use cloud service paradigm that automates resource scaling for stateless functions and can potentially facilitate bursty machine learning serving. Batching is critical for latency performance and cost-effectiveness of machine learning inference, but unfortunately it is not supported by existing serverless platforms due to their stateless design. Our experiments show that without batching, machine learning serving cannot reap the benefits of serverless computing. In this paper, we present BATCH, a framework for supporting efficient machine learning serving on serverless platforms. BATCH uses an optimizer to provide inference tail latency guarantees and cost optimization and to enable adaptive batching support. We prototype BATCH atop of AWS Lambda and popular machine learning inference systems. The evaluation verifies the accuracy of the analytic optimizer and demonstrates performance and cost advantages over the state-of-the-art method MArk and the state-of-the-practice tool SageMaker. Riccardo Pinciroli, Feng Yan 0001, Evgenia Smirni |
SC | 4 |
| 2020 | Simulating COVID-19 containment measures using the South Korean patient data: poster abstractabstractAs the COVID-19 outbreak evolves around the world, the World Health Organization (WHO) and its Member States have been heavily relying on staying at home and lock down measures to control the spread of the virus. In last months, various signs showed that the COVID-19 curve was flattening, but the premature lifting of some containment measures (e.g., school closures and telecommuting) are favouring a second wave of the disease. The accurate evaluation of possible countermeasures and their well-timed revocation are therefore crucial to avoid future waves or reduce their duration. In this paper, we analyze patient and route data collected by the Korea Centers for Disease Control & Prevention (KCDC). We extract information from real-world data sets and use them to parameterize simulations and evaluate different what-if scenarios. Lishan Yang 0001, Anna Schmedding, Riccardo Pinciroli, Evgenia Smirni |
SenSys | 4 |
| 2019 | Session details: High Performance Distributed Systems (Best Paper Nominees)abstractNo abstract available. Evgenia Smirni, Ali Raza Butt |
HPDC | 1 |
| 2019 | SSD failures in the field: symptoms, causes, and prediction modelsabstractIn recent years, solid state drives (SSDs) have become a staple of high-performance data centers for their speed and energy efficiency. In this work, we study the failure characteristics of 30,000 drives from a Google data center spanning six years. We characterize the workload conditions that lead to failures and illustrate that their root causes differ from common expectation but remain difficult to discern. Particularly, we study failure incidents that result in manual intervention from the repair process. We observe high levels of infant mortality and characterize the differences between infant and non-infant failures. We develop several machine learning failure prediction models that are shown to be surprisingly accurate, achieving high recall and low false positive rates. These models are used beyond simple prediction as they aid us to untangle the complex interaction of workload characteristics that lead to failures and identify failure root causes from monitored symptoms. Jacob Alter, Ji Xue, Alma Dimnaku, Evgenia Smirni |
SC | 4 |
| 2019 | Practical Reliability Analysis of GPGPUs in the Wild: From Systems to ApplicationsabstractGeneral Purpose Graphics Processing Units (GPGPUs) have rapidly evolved to enable energy-efficient data-parallel computing for a broad range of scientific areas. While GPUs achieve exascale performance at a stringent power budget, they are also susceptible to soft errors (faults), often caused by high-energy particle strikes, that can significantly affect application output quality. As those applications are normally long-running, investigating the characteristics of GPU errors becomes imperative to better understand the reliability of such systems. In this talk, I will present a study of the system conditions that trigger GPU soft errors using a six-month trace data collected from a large-scale, operational HPC system from Oak Ridge National Lab. Workload characteristics, certain GPU cards, temperature and power consumption could be indicative of GPU faults, but it is non-trivial to exploit them for error prediction. Motivated by these observations and challenges, I will show how machine-learning-based error prediction models can capture the hidden interactions among system and workload properties. The above findings beg the question: how can one better understand the resilience of applications if faults are bound to happen? To this end, I will illustrate the challenges of comprehensive fault injection in GPGPU applications and outline a novel fault injection solution that captures the error resilience profile of GPGPU applications. Evgenia Smirni |
ICPE | 1 |
| 2018 | Machine Learning Models for GPU Error Prediction in a Large Scale HPC SystemabstractGPUs are widely deployed on large-scale HPC systems to provide powerful computational capability for scientific applications from various domains. As those applications are normally long-running, investigating the characteristics of GPU errors becomes imperative for reliability. In this paper, we first study the system conditions that trigger GPU errors using six-month trace data collected from a large-scale, operational HPC system. Then, we use machine learning to predict the occurrence of GPU errors, by taking advantage of temporal and spatial dependencies of the trace data. The resulting machine learning prediction framework is robust and accurate under different workloads. Bin Nie, Ji Xue, Saurabh Gupta 0002, Tirthak Patel, Christian Engelmann, Evgenia Smirni, Devesh Tiwari |
DSN | 6 |
| 2018 | Fault Site Pruning for Practical Reliability Analysis of GPGPU ApplicationsabstractGraphics Processing Units (GPUs) have rapidly evolved to enable energy-efficient data-parallel computing for a broad range of scientific areas. While GPUs achieve exascale performance at a stringent power budget, they are also susceptible to soft errors, often caused by high-energy particle strikes, that can significantly affect the application output quality. Understanding the resilience of general purpose GPU applications is the purpose of this study. To this end, it is imperative to explore the range of application output by injecting faults at all the potential fault sites. This problem is especially challenging because unlike CPU applications, which are mostly single-threaded, GPGPU applications can contain hundreds to thousands of threads, resulting in a tremendously large fault site space – in the order of billions even for some simple applications. In this paper, we present a systematic way to progressively prune the fault site space aiming to dramatically reduce the number of fault injections such that assessment for GPGPU application error resilience can be practical. The key insight behind our proposed methodology stems from the fact that GPGPU applications spawn a lot of threads, however, many of them execute the same set of instructions. Therefore, several fault sites are redundant and can be pruned by a careful analysis of faults across threads and instructions. We identify important features across a set of 10 applications (16 kernels) from Rodinia and Polybench suites and conclude that threads can be first classified based on the number of the dynamic instructions they execute. We achieve significant fault site reduction by analyzing only a small subset of threads that are representative of the dynamic instruction behavior (and therefore error resilience behavior) of the GPGPU applications. Further pruning is achieved by identifying and analyzing: a) the dynamic instruction commonalities (and differences) across code blocks within this representative set of threads, b) a subset of loop iterations within the representative threads, and c) a subset of destination register bit positions. The above steps result in a tremendous reduction of fault sites by up to seven orders of magnitude. Yet, this reduced fault site space accurately captures the error resilience profile of GPGPU applications. Bin Nie, Lishan Yang 0001, Adwait Jog, Evgenia Smirni |
MICRO | 4 |
| 2018 | Evaluating Scalability and Performance of a Security Management Solution in Large Virtualized EnvironmentsabstractVirtualized infrastructure is a key capability of modern enterprise data centers and cloud computing, enabling a more agile and dynamic IT infrastructure with fast IT provisioning, simplified, automated management, and flexible resource allocation to handle a broad set of workloads. However, at the same time, virtualization introduces new challenges, since securing virtual servers is more difficult than physical machines. HyTrust Inc. has developed an innovative security solution, called HyTrust Cloud Control (HTCC), to mitigate risks associated with virtualization and cloud technologies. HTCC is a virtual appliance deployed as a transparent proxy in front of a VMware-based virtualized environment. Since HTCC serves as a gateway to a customer virtualized environment, it is important to carefully assess its performance and scalability as well as provide its accurate resource sizing. In this work, we introduce a novel approach for accomplishing this goal. First, we describe a special framework, based on a nested virtualization technique, which enables the creation and deployment of a large scale virtualized environment (with 30,000 VMs) using a limited number of physical servers (4 servers in our experiments). Second, we introduce a design and implementation of a novel, extensible benchmark, called HT-vmbench, that allows to mimic the session-based activities of different system administrators and users in virtualized environments. The benchmark is implemented using VMware Web Service SDK. By executing HT-vmbench in the emulated large-scale virtualized environments, we can support an efficient performance assessment of management and security solutions (such as HTCC), their overhead, and provide capacity planning rules and resource sizing recommendations. Lishan Yang 0001, Ludmila Cherkasova, Rajeev Badgujar, Jack Blancaflor, Rahul Konde, Jason Mills, Evgenia Smirni |
ICPE | 7 |
| 2018 | Spatial-Temporal Prediction Models for Active Ticket Managing in Data CentersabstractPerformance ticket handling is an expensive operation in data centers, where physical boxes host multiple virtual machines (VMs). A large body of tickets arise from resource usage warnings, e.g., CPU and RAM usages that exceed predefined thresholds. The transient nature of CPU and RAM usage as well as their strong correlation across time among co-located VMs within boxes drastically increase the complexity of ticket management. Based on large resource usage data collected from production data centers, with 6K physical boxes and more than 80K VMs, we first discover patterns of spatial and temporal dependencies among/within the usage series of co-located resources. Leveraging our key findings, we develop an active ticket managing (ATM) system that aims to drastically reduce usage tickets. ATM consists of: 1) a spatial-temporal dependency-based time series prediction methodology and 2) a proactive capacity planning policy for CPU and RAM resources for VMs co-located within a box and boxes within a single data center client, that aims to drastically reduce usage tickets. ATM exploits the spatial-temporal dependency across/within multiple resources of co-located VMs and single-client boxes for usage prediction, and then actuates proactive capacity planning. Evaluation results on traces of 6K physical boxes from operating data centers show that ATM is able to provide accurate prediction of usage series in cloud data centers with low computational overhead. At the same time ATM achieves significant ticket reduction up to 60% for both VM and box usage series. Ji Xue, Robert Birke, Lydia Y. Chen, Evgenia Smirni |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2018 | Efficient Deep Neural Network Serving: Fast and FuriousabstractThe emergence of deep neural networks (DNNs) as a state-of-the-art machine learning technique has enabled a variety of artificial intelligence applications for image recognition, speech recognition and translation, drug discovery, and machine vision. These applications are backed by large DNN models running in serving mode on a cloud computing infrastructure to process client inputs such as images, speech segments, and text segments. Given the compute-intensive nature of large DNN models, a key challenge for DNN serving systems is to minimize the request response latencies. This paper characterizes the behavior of different parallelism techniques for supporting scalable and responsive serving systems for large DNNs. We identify and model two important properties of DNN workloads: 1) homogeneous request service demand and 2) interference among requests running concurrently due to cache/memory contention. These properties motivate the design of serving deep learning systems fast (SERF), a dynamic scheduling framework that is powered by an interference-aware queueing-based analytical model. To minimize response latency for DNN serving, SERF quickly identifies and switches to the optimal parallel configuration of the serving system by using both empirical and analytical methods. Our evaluation of SERF using several well-known benchmarks demonstrates its good latency prediction accuracy, its ability to correctly identify optimal parallel configurations for each benchmark, its ability to adapt to changing load conditions, and its efficiency advantage (by at least three orders of magnitude faster) over exhaustive profiling. We also demonstrate that SERF supports other scheduling objectives and can be extended to any general machine learning serving system with the similar parallelism properties as above. Feng Yan 0001, Yuxiong He, Olatunji Ruwase, Evgenia Smirni |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2017 | How to Supercharge the Amazon T2: Observations and SuggestionsabstractCloud service providers adopt a credit system to allow users to obtain periods of performance bursts without additional cost. For example, the Amazon EC2 T2 instance offers low baseline performance and the capability to achieve short periods of high performance using CPU credits. Once a T2 instance is created and assigned some initial credits, while its CPU utilization is above the baseline threshold, there is a transient period where performance is boosted and the assigned CPU credits are used. After all credits are used, the maximum achievable performance drops to baseline. Credits accrue periodically, when the instance utilization is below the baseline threshold. This paper proposes a methodology to increase the performance benefits of T2 by seamlessly extending the duration of the transient period while maintaining high performance. This extension of the high performance transient period is combined with proactive migration to further take advantage of the initially assigned credits. We conduct experiments to demonstrate the benefits of this methodology for both single-tier and multi-tier applications. Feng Yan 0001, Lihua Ren, Daniel J. Dubois, Giuliano Casale, Jiawei Wen, Evgenia Smirni |
CLOUD | 6 |
| 2017 | Fill-in the gaps: Spatial-temporal models for missing dataabstractEffective workload characterization and prediction are instrumental for efficiently and proactively managing large systems. System management primarily relies on the workload information provided by underlying system tracing mechanisms that record system-related events in log files. However, such tracing mechanisms may temporarily fail due to various reasons, yielding “holes” in data traces. This missing data phenomenon significantly impedes the effectiveness of data analysis. In this paper, we study real-world data traces collected from over 80K virtual machines (VMs) hosted on 6K physical boxes in the data centers of a service provider. We discover that the usage series of VMs co-located on the same physical box exhibit strong correlation with one another, and that most VM usage series show temporal patterns. By taking advantage of the observed spatial and temporal dependencies, we propose a data-filling method to predict the missing data in the VM usage series. Detailed evaluation using trace data in the wild shows that the proposed method is sufficiently accurate as it achieves an average of 20% absolute percentage errors. We also illustrate its usefulness via a use case. Ji Xue, Bin Nie, Evgenia Smirni |
CNSM | 3 |
| 2017 | Characterizing Temperature, Power, and Soft-Error Behaviors in Data Center Systems: Insights, Challenges, and OpportunitiesabstractGPUs have become part of the mainstream high performance computing facilities that increasingly require more computational power to simulate physical phenomena quickly and accurately. However, GPU nodes also consume significantly more power than traditional CPU nodes, and high power consumption introduces new system operation challenges, including increased temperature, power/cooling cost, and lower system reliability. This paper explores how power consumption and temperature characteristics affect reliability, provides insights into what are the implications of such understanding, and how to exploit these insights toward predicting GPU errors using neural networks. Bin Nie, Ji Xue, Saurabh Gupta 0002, Christian Engelmann, Evgenia Smirni, Devesh Tiwari |
MASCOTS | 5 |
| 2017 | DyScale: A MapReduce Job Scheduler for Heterogeneous Multicore ProcessorsabstractThe functionality of modern multi-core processors is often driven by a given power budget that requires designers to evaluate different decision trade-offs, e.g., to choose between many slow, power-efficient cores, or fewer faster, power-hungry cores, or a combination of them. Here, we prototype and evaluate a new Hadoop scheduler, called DyScale, that exploits capabilities offered by heterogeneous cores within a single multi-core processor for achieving a variety of performance objectives. A typical MapReduce workload contains jobs with different performance goals: large, batch jobs that are throughput oriented, and smaller interactive jobs that are response time sensitive. Heterogeneous multi-core processors enable creating virtual resource pools based on "slow" and "fast" cores for multi-class priority scheduling. Since the same data can be accessed with either "slow" or "fast" slots, spare resources (slots) can be shared between different resource pools. Using measurements on an actual experimental setting and via simulation, we argue in favor of heterogeneous multi-core processors as they achieve "faster" (up to 40 percent) processing of small, interactive MapReduce jobs, while offering improved throughput (up to 40 percent) for large, batch jobs. We evaluate the performance benefits of DyScale versus the FIFO and Capacity job schedulers that are broadly used in the Hadoop community. Feng Yan 0001, Ludmila Cherkasova, Zhuoyao Zhang, Evgenia Smirni |
IEEE Trans. Cloud Comput. | 4 |
| 2016 | Managing Data Center Tickets: Prediction and Active SizingabstractPerformance ticket handling is an expensive operation in highly virtualized cloud data centers where physical boxes host multiple virtual machines (VMs). A large body of tickets arise from the resource usage warnings, e.g., CPU and RAM usages that exceed predefined thresholds. The transient nature of CPU and RAM usage as well as their strong correlation across time among co-located VMs drastically increase the complexity in ticket management. Based on a large resource usage data collected from production data centers, amount to 6K physical machines and more than 80K VMs, we first discover patterns of spatial dependency among co-located virtual resources. Leveraging our key findings, we develop an Active Ticket Managing(ATM) system that consists of (i) a novel time series prediction methodology and (ii) a proactive VM resizing policy for CPU and RAM resources for co-located VMs on a physical box that aims to drastically reduce usage tickets. ATM exploits the spatial dependency across multiple resources of co-located VMs for usage prediction and proactive VM resizing. Evaluation results on traces of 6K physical boxes and a prototype of a MediaWiki system show that ATM is able to achieve excellent prediction accuracy of a large number of VM time series and significant usage ticket reduction, i.e., up to 60%, at low computational overhead. Ji Xue, Robert Birke, Lydia Y. Chen, Evgenia Smirni |
DSN | 4 |
| 2016 | A large-scale study of soft-errors on GPUs in the fieldabstractParallelism provided by the GPU architecture has enabled domain scientists to simulate physical phenomena at a much faster rate and finer granularity than what was previously possible by CPU-based large-scale clusters. Architecture researchers have been investigating reliability characteristics of GPUs and innovating techniques to increase the reliability of these emerging computing devices. Such efforts are often guided by technology projections and simplistic scientific kernels, and performed using architectural simulators and modeling tools. Lack of large-scale field data impedes the effectiveness of such efforts. This study attempts to bridge this gap by presenting a large-scale field data analysis of GPU reliability. We characterize and quantify different kinds of soft-errors on the Titan supercomputer's GPU nodes. Our study uncovers several interesting and previously unknown insights about the characteristics and impact of soft-errors. Bin Nie, Devesh Tiwari, Saurabh Gupta 0002, Evgenia Smirni, James H. Rogers |
HPCA | 4 |
| 2016 | Workload interleaving with performance guarantees in data centersabstractIn the era of global, large scale data centers residing in clouds, many applications and users share the same pool of resources for the purpose of reducing costs while maintaining high performance. When multiple workloads access the same resources concurrently, their requests are interleaved, possibly causing delays. Providing performance isolation to individual workloads such that they meet their own performance objectives is important and challenging. The challenge lies in finding accurate, robust, compact metrics and models that drive algorithms which can meet different performance objectives while achieving efficient utilization of resources. This dissertation proposes a set of methodologies and tools aiming at solving the challenging performance isolation problem of workload interleaving in data centers, focusing on both storage components and computing components. At the storage node level, we consider methodologies for better interleaving user traffic with background workloads, such as tasks for improving reliability, availability, and power savings. At the storage cluster level, we propose methodologies on how to efficiently conduct work consolidation and schedule asynchronous updates without violating user performance targets. At the computing node level, we present priority scheduling middleware that employs different policies to schedule background tasks. Finally, at the computing cluster level, we develop a new Hadoop scheduler called DyScale to exploit capabilities offered by heterogeneous cores in order to achieve a variety of performance objectives. All works have been evaluated through extensive simulation using enterprise traces or real testbed implementation, and have been accepted for publications in leading performance conferences. Feng Yan 0001, Evgenia Smirni |
NOMS | 2 |
| 2016 | SERF: efficient scheduling for fast deep neural network serving via judicious parallelismabstractDeep neural networks (DNNs) has enabled a variety of artificial intelligence applications. These applications are backed by large DNN models running in serving mode on a cloud computing infrastructure. Given the compute-intensive nature of large DNN models, a key challenge for DNN serving systems is to minimize the request response latencies. This paper characterizes the behavior of different parallelism techniques for supporting scalable and responsive serving systems for large DNNs. We identify and model two important properties of DNN workloads: homogeneous request service demand, and interference among requests running concurrently due to cache/memory contention. These properties motivate the design of SERF, a dynamic scheduling framework that is powered by an interference-aware queueing-based analytical model. We evaluate SERF in the context of an image classification service using several well known benchmarks. The results demonstrate its accurate latency prediction and its ability to adapt to changing load conditions. Feng Yan 0001, Yuxiong He, Olatunji Ruwase, Evgenia Smirni |
SC | 4 |
| 2016 | Tale of Tails: Anomaly Avoidance in Data CentersabstractIt is a common practice that today's cloud data centers guard the performance by monitoring the resource usage, e.g., CPU and RAM, and issuing anomaly tickets whenever detecting usages exceeding predefined target values. Ensuring free of such usage anomaly can be extremely challenging, while catering to a large amount of virtual machines (VMs) showing bursty workloads on a limited amount of physical resource. Using resource usage data from production data centers that consist of more than 6K physical machines hosting more than 80K VMs, we identify statistic properties of anomaly instances (AIs) on physical servers, highlighting their burst duration and potential root causes. To strike a tradeoff between a strong performance guarantee and resource provisions, we propose a tail-driven anomaly avoidance policy for boxes, TailGuard, which allows a small fraction of AIs, e.g., 5% of usages can be above the target value, and still avoid severe performance degradation, typically caused by a burst of continuous AI. Specifically, TailGuard first introduces a novel usage tail prediction that explores the similarity patterns across a great number of boxes within a very recent history, and then redistributes the server load in an online fashion by proactive VM cloning and reactive load balancing. Evaluation results show that TailGuard can not only achieve an accuracy comparable with prediction methodology that relies on long history of usage data but also dramatically reduce the number of CPU AIs by 60%, with a tenfold reduction of their duration, from more than 25 time windows to only 2. Ji Xue, Robert Birke, Lydia Y. Chen, Evgenia Smirni |
SRDS | 4 |
| 2016 | PROST: Predicting Resource Usages with Spatial and Temporal DependenciesabstractWe present a tool, PROST, which can achieve scalable and accurate prediction of server workload time series in data centers. As several virtual machines are typically co-located on physical servers, the CPU and RAM show strong temporal and spatial dependencies. PROST is able to leverage the spatial dependency among co-located VMs to improve the scalability of prediction models solely based on temporal features, such as neural network. We show the benefits of PROST in obtaining accurate prediction of resource usage series and designing effective VM sizing strategies for the private data centers. Ji Xue, Evgenia Smirni, Thomas Scherer, Robert Birke, Lydia Y. Chen |
ICPE | 2 |
| 2016 | Virtualization in the Private Cloud: State of the PracticeabstractVirtualization has become a mainstream technology that allows efficient and safe resource sharing in data centers. In this paper, we present a large scale workload characterization study of 90K virtual machines hosted on 8K physical servers, across several geographically distributed corporate data centers of a major service provider. The study focuses on 19 days of operation and focuses on the state of the practice, i.e., how virtual machines are deployed across different physical resources with an emphasis on processors and memory, focusing on resource sharing and usage of physical resources, virtual machine life cycles, and migration patterns and their frequencies. This paper illustrates that indeed there is a huge tendency in over-provisioning CPU and memory resources while certain virtualization features (e.g., migration and collocation) are used rather conservatively, showing that there is significant room for the development of policies that aim to reduce operational costs in data centers. Robert Birke, Andrej Podzimek, Lydia Y. Chen, Evgenia Smirni |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2015 | Less Can Be More: Micro-managing VMs in Amazon EC2abstractMicro instances (t1. micro) are the class of Amazon EC2 virtual machines (VMs) offering the lowest operational costs for applications with short bursts in their CPU requirements. As processing proceeds, EC2 throttles CPU capacity of micro instances in a complex, unpredictable, manner. This paper aims at making micro instances more predictable and efficient to use. First, we present a characterization of EC2 micro instances that evaluates the complex interactions between cost, performance, idleness and CPU throttling. Next, we define adaptive algorithms to manage CPU consumption by learning the workload characteristics at runtime and by injecting idleness to diminish host-level throttling. We show that a gradient-hill strategy leads to favorable results. For CPU bound workloads, we observe that a significant portion of jobs (up to 65%) can have end-to-end times that are even four times shorter than those of the more expensive m1. small class. Our algorithms drastically reduce the long tails of job execution times on the micro instances, resulting to favorable comparisons against even small instances. Jiawei Wen, Giuliano Casale, Evgenia Smirni |
CLOUD | 4 |
| 2015 | PRACTISE: Robust prediction of data center time seriesabstractWe analyze workload traces from production data centers and focus on their VM usage patterns of CPU, memory, disk, and network bandwidth. Burstiness is a clear characteristic of many of these time series: there exist peak loads within clear periodic patterns but also within patterns that do not have clear periodicity. We present PRACTISE, a neural network based framework that can efficiently and accurately predict future loads, peak loads, and their timing. Extensive experimentation using traces from IBM data centers illustrates PRACTISE's superiority when compared to ARIMA and baseline neural network models, with average prediction errors that are significantly smaller. Its robustness is also illustrated with respect to the prediction window that can be short-term (i.e., hours) or long-term (i.e., a week). Ji Xue, Feng Yan 0001, Robert Birke, Lydia Y. Chen, Thomas Scherer, Evgenia Smirni |
CNSM | 6 |
| 2015 | On the DNS Deployment of Modern Web ServicesabstractAccessing Internet services relies on the Domain Name System (DNS) for translating human-readable names to routable network addresses. At the bottom level of the DNS hierarchy, the authoritative DNS (ADNS) servers maintain the actual mapping records and answer the DNS queries. Today, the increasing use of upstream ADNS services (i.e., third-party ADNS-hosting services) and Infrastructure-as-a-Service (IaaS) clouds facilitates the establishment of web services, and has been fostering the evolution of the deployment of ADNS servers. To shed light on this trend, in this paper we present a large-scale measurement to study the ADNS deployment patterns of modern web services and examine the characteristics of different deployment styles, such as performance, life-cycle of servers, and availability. Furthermore, we focus specifically on the DNS deployment for subdomains hosted in IaaS clouds. Shuai Hao 0001, Haining Wang 0001, Angelos Stavrou, Evgenia Smirni |
ICNP | 4 |
| 2014 | Optimizing Power and Performance Trade-offs of MapReduce Job Processing with Heterogeneous Multi-core ProcessorsabstractModern processors are often constrained by a given power budget that forces designers to consider different trade-offs, e.g., to choose between either many slow, power-efficient cores, or fewer faster, power-hungry cores, or to select a combination of them. In this work, we design and evaluate a new Hadoop scheduler, called DyScale, that exploits capabilities offered by heterogeneous cores within a single multi-core processor for achieving a variety of performance objectives. A typical MapReduce workload contains jobs with different performance goals: large, batch jobs that are throughput oriented, and smaller interactive jobs that are response-time sensitive. Heterogeneous multi-core processors enable creating virtual resource pools based on the different core types for multi-class priority scheduling. These virtual Hadoop clusters, based on "slow" cores versus "fast" cores can effectively support different performance objectives that cannot be achieved in a Hadoop cluster with homogeneous processors. Using detailed measurements and extensive simulation study we argue in favor of heterogeneous multi-core processors as they provide performance means for "faster" processing of the small, interactive MapReduce jobs (up to 40% faster), while at the same time offer an improved throughput (up to 40% higher) for large, batch job processing. Feng Yan 0001, Ludmila Cherkasova, Zhuoyao Zhang, Evgenia Smirni |
IEEE CLOUD | 4 |
| 2014 | (Big)data in a virtualized world: volume, velocity, and variety in cloud datacenters
Robert Birke, Mathias Björkqvist, Lydia Y. Chen, Evgenia Smirni, Antonius P. J. Engbersen |
FAST | 4 |
| 2014 | Multi-resource characterization and their (in)dependencies in production datacentersabstractPerformance studies on datacenter usage are usually limited (and perhaps even biased) due to the scale of the respective studies. In this paper, we present an extensive field study of several in-production datacenters that have been observed during a two-year time period. We present an extensive workload characterization study that sheds lights on the primary features of typical datacenter workloads, by focusing on how basic resources such as CPU, memory, and disk are used and their inter-dependencies across time. The outcome of this field study are baseline workloads that can be used to model, calibrate, manage, and plan resource management of future datacenters. Robert Birke, Lydia Y. Chen, Evgenia Smirni |
NOMS | 3 |
| 2014 | Effective resource and workload management in data centersabstractThe increasing demand for storage, computation, and business continuity has driven the growth of large data centers. Managing data centers efficiently is a difficult task because of the wide variety of datacenter applications, their ever-changing intensities, and the fact that application performance targets may differ widely. Server virtualization has been a game-changing technology for IT, providing the possibility to support multiple virtual machines (VMs) simultaneously. This dissertation focused on how virtualization technologies can be utilized to develop new tools for maintaining high resource utilization, for achieving high application performance, and for reducing the cost of data center management. This dissertation first focused on application workload management to improve web service performance especially under bursty conditions. Secondly, it concentrated on a resource measurement problem which serves as the basis of many autonomic computing solutions such as system optimization, adaptation, and troubleshooting. Thirdly, it presented AppRM, a resource allocation system that autonomically adapts to dynamic workload changes in a shared virtualized infrastructure to achieve application service level objectives (SLOs). Last, this dissertation presented PREMATCH, a tool that best co-locate different virtual machines (VMs) such that the performance of co-located VMs is maximized. All works have been implemented and tested over real enterprise applications and were accepted for publication at leading autonomic management conferences. Evgenia Smirni |
NOMS | 2 |
| 2014 | Application-driven dynamic vertical scaling of virtual machines in resource poolsabstractMost modern hypervisors offer powerful resource control primitives such as reservations, limits, and shares for individual virtual machines (VMs). These primitives provide a means to dynamic vertical scaling of VMs in order for the virtual applications to meet their respective service level objectives (SLOs). VMware DRS offers an additional resource abstraction of a resource pool (RP) as a logical container representing an aggregate resource allocation for a collection of VMs. In spite of the abundant research on translating application performance goals to resource requirements, the implementation of VM vertical scaling techniques in commercial products remains limited. In addition, no prior research has studied automatic adjustment of resource control settings at the resource pool level. In this paper, we present AppRM, a tool that automatically sets resource controls for both virtual machines and resource pools to meet application SLOs. AppRM contains a hierarchy of virtual application managers and resource pool managers. At the application level, AppRM translates performance objectives into the appropriate resource control settings for the individual VMs running that application. At the resource pool level, AppRM ensures that all important applications within the resource pool can meet their performance targets by adjusting controls at the resource pool level. Experimental results under a variety of dynamically changing workloads composed by multi-tiered applications demonstrate the effectiveness of AppRM. In all cases, AppRM is able to deliver application performance satisfaction without manual intervention. Xiaoyun Zhu, Rean Griffith, Pradeep Padala, Aashish Parikh, Evgenia Smirni |
NOMS | 7 |
| 2014 | Heterogeneous cores for MapReduce processing: Opportunity or challenge?abstractTo offer diverse computing capabilities, the emergent modern system on a chip (SoC) might include heterogeneous multi-core processors. The current SoC design is often constrained by a given power budget that forces designers to consider different decision trade-offs, e.g., to choose between many slow cores, fewer faster cores, or to select a combination of them. In this work, we design a new Hadoop scheduler, called DyScale, that exploits capabilities offered by heterogeneous cores for achieving a variety of performance objectives. Our preliminary performance evaluation results confirm potential benefits of heterogeneous multi-core processors for “faster” processing of the small, interactive MapReduce jobs, while at the same time offering an improved throughput and performance for large, batch job processing. Feng Yan 0001, Ludmila Cherkasova, Zhuoyao Zhang, Evgenia Smirni |
NOMS | 4 |
| 2014 | Agile middleware for scheduling: meeting competing performance requirements of diverse tasksabstractAs the need for scaled-out systems increases, it is paramount to architect them as large distributed systems consisting of off-the-shelf basic computing components known as compute or data nodes. These nodes are expected to handle their work independently, and often utilize off-the-shelf management tools, like those offered by Linux, to differentiate priorities of tasks. While prioritization of background tasks in server nodes takes center stage in scaled-out systems, with many tasks associated with salient features such as eventual consistency, data analytics, and garbage collection, the standard Linux tools such as nice and ionice fail to adapt to the dynamic behavior of high priority tasks in order to achieve the best trade-off between protecting the performance of high priority workload and completing as much low priority work as possible. In this paper, we provide a solution by proposing a priority scheduling middleware that employs different policies to schedule background tasks based on the instantaneous resource requirements of the high priority applications running on the server node. The selection of policies is based on off-line and on-line learning of the high priority workload characteristics and the imposed performance impact due to low priority work. In effect, this middleware uses a {\em hybrid} approach to scheduling rather than a monolithic policy. We prototype and evaluate it via measurements on a test-bed and show that this scheduling middleware is robust as it effectively and autonomically changes the relative priorities between high and low priority tasks, consistently meeting their competing performance targets. Feng Yan 0001, Shannon Hughes, Alma Riska, Evgenia Smirni |
ICPE | 4 |
| 2013 | State-of-the-practice in data center virtualization: Toward a better understanding of VM usageabstractHardware virtualization is the prevalent way to share data centers among different tenants. In this paper we present a large scale workload characterization study that aims to a better understanding of the state-of-the-practice, i.e., how data centers in the private cloud are used by their customers, how physical resources are shared among different tenants using virtualization, and how virtualization technologies are actually employed. Our study focuses on all corporate data centers of a major infrastructure provider that are geographically dispersed across the entire globe and reports on their observed usage across a 19-day period. We especially focus on how virtual machines are deployed across different physical resources with an emphasis on processors and memory, focusing on resource sharing and usage of physical resources, virtual machine life cycles, and migration patterns and frequencies. Our study illustrates that there is a huge tendency in over provisioning resources while being conservative to the several possibilities opened up by virtualization (e.g., migration and co-location), showing tremendous potential for the development of policies aiming to reduce data center operational costs. Robert Birke, Andrej Podzimek, Lydia Y. Chen, Evgenia Smirni |
DSN | 4 |
| 2013 | Predictive VM consolidation on multiple resources: Beyond load balancingabstractEffective consolidation of different applications on common resources is often akin to black art as application performance interference may result in unpredictable system and workload delays. In this paper we consider the problem of fair load balancing on multiple servers within a virtualized data center setting. We especially focus on multi-tiered applications with different resource demands per tier and address the problem on how to best match each application tier on each resource, such that performance interference is minimized. To address this problem, we propose a two-step approach. First, a fair load balancing scheme assigns different virtual machines (VMs) across different servers; this process is formulated as a multi-dimensional vector scheduling problem that uses a new polynomial-time approximation scheme (PTAS) to minimize the maximum utilization across all server resources and results in multiple load balancing solutions. Second, a queueing network analytic model is applied on the proposed min-max solutions in order to select the optimal one. We experimentally evaluate the proposed two-stage mechanism using a Xen virtualization testbed that hosts multiple RUBiS multi-tier applications. Experimental results show that the proposed mechanism is robust as it always predicts the optimal consolidation strategy. Hui Zhang 0002, Evgenia Smirni, Guofei Jiang, Kenji Yoshihira |
IWQoS | 3 |
| 2013 | Overcoming Limitations of Off-the-Shelf Priority Schedulers in Dynamic EnvironmentsabstractIt is common nowadays to architect and design scaled-out systems with off-the-shelf computing components operated and managed by off-the-shelf open-source tools. While web services represent the critical set of services offered at scale, big data analytics is emerging as a preferred service to be colocated with cloud web services at a lower priority raising the need for off-the-shelf priority scheduling. In this paper we report on the perils of Linux priority scheduling tools when used to differentiate between such complex services. We demonstrate that simple priority scheduling utilities such as nice and ionice can result in dramatically erratic behavior. We provide a remedy by proposing an autonomic priority scheduling algorithm that adjusts its execution parameters based on on-line measurements of the current resource usage of critical applications. Detailed experimentation with a user-space prototype of the algorithm on a Linux system using popular benchmarks such as SPEC and TPC-W illustrate the robustness and versatility of the proposed technique, as it provides consistency to the expected performance of a high-priority application when running simultaneously with multiple low priority jobs. Feng Yan 0001, Shannon Hughes, Alma Riska, Evgenia Smirni |
MASCOTS | 4 |
| 2013 | On load balancing: a mix-aware algorithm for heterogeneous systemsabstractToday's web services are commonly hosted on clusters of servers that are often located within computing clouds, whose computational and storage resources can be highly heterogeneous. The workload served typically exhibits disparate computation patterns (e.g., CPU-intensive or IO-intensive), that fluctuate both in terms of volume and mix. The system heterogeneity together with workload diversity further exacerbates the challenge of effective distribution of load within a computing cloud. This paper presents a novel, mix-aware load-balancing algorithm, which aims to distribute requests sent by multiple applications in heterogeneous servers such that the application response times are minimized and system resources (e.g., CPU and IO) are equally utilized. To this end, the presented algorithm tries to not only balance the total number of requests seen by each server, but also to shape the requests received by each server into a certain "mix", that is analytically shown to be optimal for response time minimization. Our experimental results---based both on simulation and on a prototype implementation---show that the mix-aware algorithm achieves robust performance in most workload mixes as well as a consistent performance improvement in comparison with one of the most robust load-balancing schemes of the Apache server. Sebastiano Spicuglia, Mathias Björkqvist, Lydia Y. Chen, Giuseppe Serazzi, Walter Binder, Evgenia Smirni |
ICPE | 6 |
| 2012 | Data Centers in the Cloud: A Large Scale Performance StudyabstractWith the advancement of virtualization technologies and the benefit of economies of scale, industries are seeking scalable IT solutions, such as data centers hosted either in-house or by a third party. Data center availability, often via a cloud setting, is ubiquitous. Nonetheless, little is known about the in-production performance of data centers, and especially the interaction of workload demands and resource availability. This study fills this gap by conducting a large scale survey of in-production data center servers within a time period that spans two years. We provide in-depth analysis on the time evolution of existing data center demands by providing a holistic characterization of typical data center server workloads, by focusing on their basic resource components, including CPU, memory, and storage systems. We especially focus on seasonality of resource demands and how this is affected by different geographical locations. This survey provides a glimpse on the evolution of data center workloads and provides a basis for an economics analysis that can be used for effective capacity planning of future data centers. Robert Birke, Lydia Y. Chen, Evgenia Smirni |
IEEE CLOUD | 3 |
| 2012 | Model-driven consolidation of Java workloads on multicoresabstractOptimal resource allocation and application consolidation on modern multicore systems that host multiple applications is not easy. Striking a balance among conflicting targets such as maximizing system throughput and system utilization while minimizing application response times is a quandary for system administrators. The purpose of this work is to offer a methodology that can automate the difficult process of identifying how to best consolidate workloads in a multicore environment. We develop a simple approach that treats the hardware and the operating system as a black box and uses measurements to profile the application resource demands. The demands become input to a queueing network model that successfully predicts application scalability and that captures the performance impact of consolidated applications on shared on-chip and off-chip resources. Extensive analysis with the widely used DaCapo Java benchmarks on an IBM Power 7 system illustrates the model's ability to accurately predict the system's optimal application mix. Danilo Ansaloni, Lydia Y. Chen, Evgenia Smirni, Walter Binder |
DSN | 3 |
| 2012 | Achieving application-centric performance targets via consolidation on multicores: myth or reality?abstractConsolidation of multiple applications with diverse and changing resource requirements is common in multicore systems as hardware resources are abundant and opportunities for better system usage are plenty. Can we maximize resource usage in such a system while respecting individual application performance targets or is it an oxymoron to simultaneously meet such conflicting measures? In this work we provide a solution to the above difficult problem by constructing a queueing-theory based tool that we use to accurately predict application scalability on multicores and that can also provide the optimal consolidation suggestions to maximize system resource usage while meeting simultaneously application performance targets. The proposed methodology is light-weight and relies on capturing application resource demands using standard tools, via nonintrusive low-level measurements. We evaluate our approach on an IBM Power7 system using the DaCapo and SPECjvm benchmark suites where each benchmark exhibits different patterns of parallelism. From 900 different consolidations of application instances, our tool accurately predicts the average iteration time of allocated applications with an average error below 10%. Lydia Y. Chen, Danilo Ansaloni, Evgenia Smirni, Akira Yokokawa, Walter Binder |
HPDC | 3 |
| 2012 | Find your best match: predicting performance of consolidated workloadsabstractModern multicore platforms allow system administrators to reduce the costs of the IT infrastructure by consolidating heterogeneous workloads on the same physical machine. To this end, it is important to develop efficient profiling techniques and accurate performance predictions to avoid violating service level objectives. In this work we present Tresa, a novel tool to automatically characterize workloads and accurately estimate the execution time of different consolidations. These results can be used to optimize consolidations depending on service-level objectives. Danilo Ansaloni, Lydia Y. Chen, Evgenia Smirni, Akira Yokokawa, Walter Binder |
ICPE | 3 |
| 2012 | Busy bee: how to use traffic information for better scheduling of background tasksabstractComputer systems, in general, and storage systems, in particular, rely on meeting their performance, reliability, and availability targets via scheduling of management and maintenance activities as background tasks.Such tasks may cause significant delays to user workload if scheduled extemporaneously. Here, we propose a scheduling policy for background tasks that is based on the statistical characteristics of the system's busy periods and that aims at completing background work expediently.Extensive trace-driven simulations show that the scheduling policy is robust and that it succeeds in completing background work faster than common practices while impacting user performance minimally. Feng Yan 0001, Alma Riska, Evgenia Smirni |
ICPE | 3 |
| 2012 | ASIdE: Using Autocorrelation-Based Size Estimation for Scheduling Bursty WorkloadsabstractTemporal dependence in workloads creates peak congestion that can make service unavailable and reduce system performance. To improve system performability under conditions of temporal dependence, a server should quickly process bursts of requests that may need large service demands. In this paper, we propose and evaluateASIdE, an Autocorrelation-based SIze Estimation, that selectively delays requests which contribute to the workload temporal dependence. ASIdE implicitly approximates the shortest job first (SJF) scheduling policy but without any prior knowledge of job service times. Extensive experiments show that (1) ASIdE achieves good service time estimates from the temporal dependence structure of the workload to implicitly approximate the behavior of SJF; and (2) ASIdE successfully counteracts peak congestion in the workload and improves system performability under a wide variety of settings. Specifically, we show that system capacity under ASIdE is largely increased compared to the first-come first-served (FCFS) scheduling policy and is highly-competitive with SJF. Ningfang Mi, Giuliano Casale, Evgenia Smirni |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2012 | Dealing with Burstiness in Multi-Tier Applications: Models and Their ParameterizationabstractWorkloads and resource usage patterns in enterprise applications often show burstiness resulting in large degradation of the perceived user performance. In this paper, we propose a methodology for detecting burstiness symptoms in multi-tier applications but, rather than identifying the root cause of burstiness, we incorporate this information into models for performance prediction. The modeling methodology is based on the index of dispersion of the service process at a server, which is inferred by observing the number of completions within the concatenated busy times of that server. The index of dispersion is used to derive a Markov-modulated process that captures burstiness and variability of the service process at each resource well and that allows us to define queueing network models for performance prediction. Experimental results and performance model predictions are in excellent agreement and argue for the effectiveness of the proposed methodology under both bursty and nonbursty workloads. Furthermore, we show that the methodology extends to modeling flash crowds that create burstiness in the stream of requests incoming to the application. Giuliano Casale, Ningfang Mi, Ludmila Cherkasova, Evgenia Smirni |
IEEE Trans. Software Eng. | 4 |
| 2011 | Approximate analysis of blocking queueing networks with temporal dependenceabstractIn this paper we extend the class of MAP queueing networks to include blocking models, which are useful to describe the performance of service instances which have a limited concurrency level. We consider two different blocking mechanisms: Repetitive Service-Random Destination (RS-RD) and Blocking After Service (BAS). We propose a methodology to evaluate MAP queueing networks with blocking based on the recently proposed Quadratic Reduction (QR), a state space transformation that decreases the number of states in the Markov chain underlying the queueing network model. From this reduced state space, we obtain boundable approximations on average performance indexes such as throughput, response time, utilizations. The two approximations that dramatically enhance the QR bounds are based on maximum entropy and on a novel minimum mutual information principle, respectively. Stress cases of increasing complexity illustrate the excellent accuracy of the proposed approximations on several models of practical interest. Vittoria de Nitto Persone, Giuliano Casale, Evgenia Smirni |
DSN | 3 |
| 2011 | Toward Automating Work Consolidation with Performance Guarantees in Storage ClustersabstractWith most of today's systems being highly distributed, from data centers to cloud and storage clusters, there is a prevalent need for robust methodologies for work consolidation to improve load balancing but also to optimize non-traditional performance measures. Such alternative measures may include power savings, e.g., it may be desirable to shut down a lowly utilized node by moving some or all of its work to another node. In this paper, we present a methodology for distributed work consolidation that keeps track of the workload in the various nodes of the cluster and makes intelligent decisions on how much work to move from a sender node to a receiver node in order to minimally "affect" the performance of the receiver node or alternatively limit any performance degradation due to consolidation in a controlled way. The proposed methodology is based on continuously monitoring the workload on sender and receiver nodes, collecting lightweight statistics in the form of histograms of coarse granularity, and deciding when and how to initiate the work transfer. Extensive experimentation using trace-driven simulation confirms the robustness of the methodology. Feng Yan 0001, Xenia Mountrouidou, Alma Riska, Evgenia Smirni |
MASCOTS | 4 |
| 2011 | Adaptive workload shaping for power savings on disk drivesabstractIn order to reduce the amount of power consumption in data centers, it is becoming necessary to shut off or slow down disks that are not actively serving user requests. In addition to exploiting disk drive idleness, system features are in place that shape a disk's workload by redirecting portions of it elsewhere, with the goal to expand the periods of idleness and the potential for power savings. In this paper, we propose several workload shaping techniques that determine which part of the working set to copy elsewhere using temporal and spatial access frequencies in the workload. These workload shaping techniques, used within an analytic estimation methodology, enable a fully automated framework that determines on-line for the current workload which, if any, shaping technique to activate such that the power saving benefits are maximized without violating performance targets. Extensive trace-driven evaluation shows that the proposed workload shaping techniques complement each-other with regard to their abilities to enhance idleness in disk drives for a wide range of workload characteristics. This results to added power savings in a data center even when performance targets are stringent and workloads intensive. Xenia Mountrouidou, Alma Riska, Evgenia Smirni |
ICPE | 3 |
| 2010 | AWAIT: Efficient Overload Management for Busy Multi-tier Web Services under Bursty Workloads
Ludmila Cherkasova, Vittoria de Nitto Persone, Ningfang Mi, Evgenia Smirni |
ICWE | 5 |
| 2010 | CWS: a model-driven scheduling policy for correlated workloadsabstractWe define CWS, a non-preemptive scheduling policy for workloads with correlated job sizes. CWS tackles the scheduling problem by inferring the expected sizes of upcoming jobs based on the structure of correlations and on the outcome of past scheduling decisions. Size prediction is achieved using a class of Hidden Markov Models (HMM) with continuous observation densities that describe job sizes. We show how the forward-backward algorithm of HMMs applies effectively in scheduling applications and how it can be used to derive closed-form expressions for size prediction. This is particularly simple to implement in the case of observation densities that are phase-type (PH-type) distributed, where existing fitting methods for Markovian point processes may also simplify the parameterization of the HMM workload model.Based on the job size predictions, CWS emulates size-based policies which favor short jobs, with accuracy depending mainly on the HMM used to parametrize the scheduling algorithm. Extensive simulation and analysis illustrate that CWS is competitive with policies that assume exact information about the workload. Giuliano Casale, Ningfang Mi, Evgenia Smirni |
SIGMETRICS | 3 |
| 2010 | Efficient resource allocation and power saving in multi-tiered systemsabstractIn this paper, we present Fastrack, a parameter-free algorithm for dynamic resource provisioning that uses simple statistics to promptly distill information about changes in workload burstiness. This information, coupled with the application's end-to-end response times and system bottleneck characteristics, guide resource allocation that shows to be very effective under a broad variety of burstiness profiles and bottleneck scenarios. Andrew Caniff, Ningfang Mi, Ludmila Cherkasova, Evgenia Smirni |
WWW | 5 |
| 2010 | Trace data characterization and fitting for Markov modeling
Giuliano Casale, Eddy Z. Zhang, Evgenia Smirni |
Perform. Evaluation | 3 |
| 2010 | KPC-Toolbox: Best recipes for automatic trace fitting using Markovian Arrival Processes
Giuliano Casale, Eddy Z. Zhang, Evgenia Smirni |
Perform. Evaluation | 3 |
| 2010 | Model-Driven System Capacity Planning under Workload BurstinessabstractIn this paper, we define and study a new class of capacity planning models called MAP queueing networks. MAP queueing networks provide the first analytical methodology to describe and predict accurately the performance of complex systems operating under bursty workloads, such as multitier architectures or storage arrays. Burstiness is a feature that significantly degrades system performance and that cannot be captured explicitly by existing capacity planning models. MAP queueing networks address this limitation by describing computer systems as closed networks of servers whose service times are Markovian Arrival Processes (MAPs), a class of Markov-modulated point processes that can model general distributions and burstiness. In this paper, we show that MAP queueing networks provide reliable performance predictions even if the service processes are bursty. We propose a methodology to solve MAP queueing networks by two state space transformations, which we call Linear Reduction (LR) and Quadratic Reduction (QR). These transformations dramatically decrease the number of states in the underlying Markov chain of the queueing network model. From these reduced state spaces, we obtain two classes of bounds on arbitrary performance indexes, e.g., throughput, response time, and utilizations. Numerical experiments show that LR and QR bounds achieve good accuracy. We also illustrate the high effectiveness of the LR and QR bounds in the performance analysis of a real multitier architecture subject to TPC-W workloads that are characterized as bursty. These results promote MAP queueing networks as a new class of robust capacity planning models. Giuliano Casale, Ningfang Mi, Evgenia Smirni |
IEEE Trans. Computers | 3 |
| 2009 | MAP-AMVA: Approximate mean value analysis of bursty systemsabstractMAP queueing networks are recently proposed models for performance assessment of enterprise systems, such as multi-tier applications, where workloads are significantly affected by burstiness. Although MAP networks do not admit a simple product-form solution, performance metrics can be estimated accurately by linear programming bounds, yet these are expensive to compute under large populations. In this paper, we introduce an approximate mean value analysis (AMVA) approach to MAP network solution that significantly reduces the computational cost of model evaluation. We define a number of balance equations that relate mean performance indices such as utilizations and response times. We show that the quality of a MAP-AMVA solution is competitive with much more complex bounds which evaluate the state space of the underlying Markov chain. Numerical results on stress cases indicate that the MVA approach is much more scalable than existing evaluation methods for MAP networks. Giuliano Casale, Evgenia Smirni |
DSN | 2 |
| 2009 | Autocorrelation-driven load control in distributed systemsabstractIn this paper, we propose a new approach for the development of load control policies in autonomic multitier systems. We control system load in a completely new way compared to existing policies: we leverage on the autocorrelation of service times and show that autocorrelation can be used to forecast future service requirements of requests and adaptively control system load. To the best of our knowledge, this is the first direct application of autocorrelation of service times to autonomic load control. We propose ALoC and D ALoC, two autocorrelation-driven policies that drop a percentage of the load in order to meet pre-defined quality-of-service levels in a distributed system. Both policies are easy to implement and rely on minimal assumptions. In particular, D ALoC is a fully no-knowledge measurement-based policy that self-adjusts its load control parameters based only on policy targets and on statistical information of requests served in the past. We illustrate the effectiveness of these new policies in a distributed multi-server setting via detailed trace driven simulations. We show that if these policies are employed in the server with a temporal dependent service process, then end-to-end response time, across all servers, reduces up to 80% by only dropping at most 13% of the incoming requests. Using real traces, we also show that, in the constrained case of being able to drop only from a portion of the incoming workload, our policy still improves request response time by up to 30%. Ningfang Mi, Giuliano Casale, Qi Zhang 0012, Alma Riska, Evgenia Smirni |
MASCOTS | 5 |
| 2009 | Automated anomaly detection and performance modeling of enterprise applicationsabstractAutomated tools for understanding application behavior and its changes during the application lifecycle are essential for many performance analysis and debugging tasks. Application performance issues have an immediate impact on customer experience and satisfaction. A sudden slowdown of enterprise-wide application can effect a large population of customers, lead to delayed projects, and ultimately can result in company financial loss. Significantly shortened time between new software releases further exacerbates the problem of thoroughly evaluating the performance of an updated application. Our thesis is that online performance modeling should be a part of routine application monitoring. Early, informative warnings on significant changes in application performance should help service providers to timely identify and prevent performance problems and their negative impact on the service. We propose a novel framework for automated anomaly detection and application change analysis. It is based on integration of two complementary techniques: (i) a regression-based transaction model that reflects a resource consumption model of the application, and (ii) an application performance signature that provides a compact model of runtime behavior of the application. The proposed integrated framework provides a simple and powerful solution for anomaly detection and analysis of essential performance changes in application behavior. An additional benefit of the proposed approach is its simplicity: It is not intrusive and is based on monitoring data that is typically available in enterprise production environments. The introduced solution further enables the automation of capacity planning and resource provisioning tasks of multitier applications in rapidly evolving IT environments. Ludmila Cherkasova, Kivanc M. Ozonat, Ningfang Mi, Julie Symons, Evgenia Smirni |
ACM Trans. Comput. Syst. | 5 |
| 2009 | Efficient management of idleness in storage systemsabstractVarious activities that intend to enhance performance, reliability, and availability of storage systems are scheduled with low priority and served during idle times. Under such conditions, idleness becomes a valuable “resource” that needs to be efficiently managed. A common approach in system design is to be nonwork conserving by “idle waiting”, that is, delay the scheduling of background jobs to avoid slowing down upcoming foreground tasks. In this article, we complement “idle waiting” with the “estimation” of background work to be served in every idle interval to effectively manage the trade-off between the performance of foreground and background tasks. As a result, the storage system is better utilized without compromising foreground performance. Our analysis shows that if idle times have low variability, then idle waiting is not necessary. Only if idle times are highly variable does idle waiting become necessary to minimize the impact of background activity on foreground performance. We further show that if there is burstiness in idle intervals, then it is possible to predict accurately the length of incoming idle intervals and use this information to serve more background jobs without affecting foreground performance. Ningfang Mi, Alma Riska, Qi Zhang 0012, Evgenia Smirni, Erik Riedel |
ACM Trans. Storage | 4 |
| 2008 | Anomaly? application change? or workload change? towards automated detection of application performance anomaly and changeabstractAutomated tools for understanding application behavior and its changes during the application life-cycle are essential for many performance analysis and debugging tasks. Application performance issues have an immediate impact on customer experience and satisfaction. A sudden slowdown of enterprise-wide application can effect a large population of customers, lead to delayed projects and ultimately can result in company financial loss. We believe that online performance modeling should be a part of routine application monitoring. Early, informative warnings on significant changes in application performance should help service providers to timely identify and prevent performance problems and their negative impact on the service. We propose a novel framework for automated anomaly detection and application change analysis. It is based on integration of two complementary techniques: i) a regression-based transaction model that reflects a resource consumption model of the application, and ii) an application performance signature that provides a compact model of run-time behavior of the application. The proposed integrated framework provides a simple and powerful solution for anomaly detection and analysis of essential performance changes in application behavior. An additional benefit of the proposed approach is its simplicity: it is not intrusive and is based on monitoring data that is typically available in enterprise production environments. Ludmila Cherkasova, Kivanc M. Ozonat, Ningfang Mi, Julie Symons, Evgenia Smirni |
DSN | 5 |
| 2008 | Scheduling for performance and availability in systems with temporal dependent workloadsabstractTemporal locality in workloads creates conditions in which a server, in order to remain available, should quickly process bursts of requests with large service requirements. In this paper, we show how to counteract the resulting peak congestions and maintain high availability by delaying selected requests that contribute to the temporal locality. We propose and evaluate SWAP, a measurement-based scheduling policy that approximates the shortest job first (SJF) scheduling without requiring any knowledge of job service times. We show that good service time estimates can be obtained from the temporal dependence structure of the workload and allow to closely approximate the behavior of SJF. Experimental results indicate that SWAP significantly improves system performability. In particular, we show that system capacity under SWAP is largely increased compared to first-come first-served (FCFS) scheduling and is highly-competitive with SJF, but without requiring a priori information of job service times. Ningfang Mi, Giuliano Casale, Evgenia Smirni |
DSN | 3 |
| 2008 | Enhancing data availability in disk drives through background activitiesabstractLatent sector errors in disk drives affect only a few data sectors. They occur silently and are detected only when the affected area is accessed again. If a latent error is detected while the storage system is operating under reduced redundancy, i.e., during a RAID rebuild, then data loss may occur. Various features such as scrubbing and intra-disk data redundancy are proposed to detect and/or recover from latent errors and avoid data loss. While such features enhance data availability in the storage system, their execution may cause performance degradation. In this paper, we evaluate the effectiveness of scrubbing and intra-disk data redundancy in improving data availability while the overall goal is to maintain user performance within predefined bounds. We show that by treating them as low priority background activities and scheduling them efficiently during idle times, these features remain performance-wise transparent to the storage system user while still improving data reliability. Detailed trace-driven simulations show that the mean time to data loss (MTTDL) improves by up to 5 orders of magnitude if these features are implemented independently. By scheduling concurrently both scrubbing and intra-disk parity updates during idle times in disk drives, MTTDL improves by as much as 8 orders of magnitude. Ningfang Mi, Alma Riska, Evgenia Smirni, Erik Riedel |
DSN | 3 |
| 2008 | Versatile models of systems using map queueing networksabstractAnalyzing the performance impact of temporal dependent workloads on hardware and software systems is a challenging task that yet must be addressed to enhance performance of real applications. For instance, existing matrix-analytic queueing models can capture temporal dependence only in systems that can be described by one or two queues, but the capacity planning of real multi-tier architectures requires larger models with arbitrary topology. To address the lack of a proper modeling technique for systems subject to temporal dependent workloads, we introduce a class of closed queueing networks where service times can have non-exponential distribution and accurately approximate temporal dependent features such as short or long range dependence. We describe these service processes using Markovian arrival processes (MAPs), which include the popular Markov-modulated Poisson processes (MMPPs) as special cases. Using a linear programming approach, we obtain for MAP closed networks tight upper and lower bounds for arbitrary performance indexes (e.g., throughput, response time, utilization). Numerical experiments indicate that our bounds achieve a mean accuracy error of 2% and promote our modeling approach for the accurate performance analysis of real multi-tier architectures. Giuliano Casale, Ningfang Mi, Evgenia Smirni |
IPDPS | 3 |
| 2008 | Burstiness in Multi-tier Applications: Symptoms, Causes, and New Models
Ningfang Mi, Giuliano Casale, Ludmila Cherkasova, Evgenia Smirni |
Middleware | 4 |
| 2008 | Analysis of application performance and its change via representative application signaturesabstractApplication servers are a core component of a multitier architecture that has become the industry standard for building scalable client-server applications. A client communicates with a service deployed as a multi-tier application via request-reply transactions. A typical server reply consists of the web page dynamically generated by the application server. The application server may issue multiple database calls while preparing the reply. Understanding the cascading effects of the various tasks that are sprung by a single request-reply transaction is a challenging task. Furthermore, significantly shortened time between new software releases further exacerbates the problem of thoroughly evaluating the performance of an updated application. We address the problem of efficiently diagnosing essential performance changes in application behavior in order to provide timely feedback to application designers and service providers. In this work, we propose a new approach based on an application signature that enables a quick performance comparison of the new application signature against the old one, while the application continues its execution in the production environment. The application signature is built based on new concepts that are introduced here, namely the transaction latency profiles and transaction signatures. These become instrumental for creating an application signature that accurately reflects important performance characteristics. We show that such an application signature is representative and stable under different workload characteristics. We also show that application signatures are robust as they effectively capture changes in transaction times that result from software updates. Application signatures provide a simple and powerful solution that can further be used for efficient capacity planning, anomaly detection, and provisioning of multi-tier applications in rapidly evolving IT environments. Ningfang Mi, Ludmila Cherkasova, Kivanc M. Ozonat, Julie Symons, Evgenia Smirni |
NOMS | 5 |
| 2008 | Bound analysis of closed queueing networks with workload burstinessabstractBurstiness and temporal dependence in service processes are often found in multi-tier architectures and storage devices and must be captured accurately in capacity planning models as these features are responsible of significant performance degradations. However, existing models and approximations for networks of first-come first-served (FCFS) queues with general independent (GI) service are unable to predict performance of systems with temporal dependence in workloads. Giuliano Casale, Ningfang Mi, Evgenia Smirni |
SIGMETRICS | 3 |
| 2008 | Performance-Guided Load (Un)balancing under Autocorrelated FlowsabstractSize-based policies have been shown in the literature to effectively balance the load and improve performance in cluster environments. Size-based policies assign jobs to servers based on the job size and their performance improvements are an outcome of separating ";short"; from ";long"; jobs, by avoiding having short jobs waiting behind long jobs for service. In this paper, we present evidence that performance improvements due to this separation quickly vanish if the arrival process to the cluster is autocorrelated. Based on our observations, we devise a new size-based policy called D_EQAL that still strives to separate jobs to servers according to job size but this separation is now biased by an effort to reduce performance loss due to autocorrelation in the arrival flows to each server. As a result of this bias, all servers may not be equally utilized (i.e., the load in the system may be ";unbalanced";), but performance benefits become significant. D_EQAL can be used on-line as it does not assume any a priori knowledge of the incoming workload. Extensive simulations show the effectiveness of D_EQAL under autocorrelated and uncorrelated arrival streams and illustrate that the policy successfully self- adjusts the degree of load unbalancing based on monitored performance measures. Qi Zhang 0012, Ningfang Mi, Alma Riska, Evgenia Smirni |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2007 | A Capacity Planning Framework for Multi-tier Enterprise Services with Real WorkloadsabstractWith complexity of systems increasing and customer requirements for QoS growing, new methods and modeling techniques that explain large-systems' behavior and help predict their future performance are required to effectively tackle the emerging performance issues. To accurately answer capacity planning questions for an existing production system with a real workload mix, we propose a new capacity planning framework that is based on the following three components: i) a Workload Profiler that dynamically builds the workload profile; ii) a Regression-based Solver that is used for deriving the CPU demand of client transactions on a given hardware; and iii) an Analytical model that is based on a network of queues representing the different tiers. To validate our approach, we conduct a detailed case study using the access logs from two heterogeneous production servers that represent customized client accesses to a popular and actively used HP Open View Service Desk application. Qi Zhang 0012, Ludmila Cherkasova, Guy Mathews, Wayne Greene, Evgenia Smirni |
Integrated Network Management | 5 |
| 2007 | New Results on the Performance Effects of Autocorrelated Flows in SystemsabstractTemporal dependence within the workload of any computing or networking system has been widely recognized as a significant factor affecting performance. More specifically, burstiness, as a form of temporal dependency, is catastrophic for performance. We use the autocorrelation function in a workload flow to formalize burstiness and also to characterize temporal dependence within a flow. We present results from two application areas: load balancing in a homogeneous cluster environment and capacity planning in a multi-tiered e-commerce system. For the load balancing problem, we show that if autocorrelation exists in the arrival stream to the cluster, classic load balancing policies become ineffective and solutions that focus on "unbalancing" the load offer superior performance. For the case of multi-tiered systems, we show that if there is autocorrelation in the flows, we observe the surprising result that in spite of the fact that the bottleneck resource in the system is far from saturation and that the measured throughput and utilizations of other resources are also modest, user response times are very high. For multi-tired systems, this underutilization of resources falsely indicates that the system can sustain higher capacities. We present analysis of the above phenomena that aims at the development better scheduling policies under auto correlated flows. Evgenia Smirni, Qi Zhang 0012, Ningfang Mi, Alma Riska, Giuliano Casale |
IPDPS | 1 |
| 2007 | R-Capriccio: A Capacity Planning and Anomaly Detection Tool for Enterprise Services with Live Workloads
Qi Zhang 0012, Ludmila Cherkasova, Guy Mathews, Wayne Greene, Evgenia Smirni |
Middleware | 5 |
| 2007 | Efficient management of idleness in systemsabstractNo abstract available. Ningfang Mi, Alma Riska, Qi Zhang 0012, Evgenia Smirni, Erik Riedel |
SIGMETRICS | 4 |
| 2007 | Future directions in performance evaluation researchabstractNo abstract available. Evgenia Smirni, Frederica Darema, Albert G. Greenberg, Adolfy Hoisie, Don Towsley |
SIGMETRICS | 1 |
| 2007 | ETAQA Solutions for Infinite Markov Processes with Repetitive StructureabstractWe describe the ETAQA (efficient technique for the solution of quasi birth-death processes) approach for the exact analysis of M/G/1 and GI/M/1-type processes, and their intersection, i.e., quasi birth-death processes. ETAQA exploits the repetitive structure of the infinite portion of the chain and derives a finite system of linear equations. In contrast to the classic techniques for solution of such systems, the solution of this finite linear system does not provide the entire probability distribution of the state space, but simply allows calculation of the aggregate probability of a finite set of classes of states from the state space, appropriately defined. Nonetheless, these aggregate probabilities allow for computation of a rich set of measures of interest such as the system queue length or any of its higher moments. The proposed solution approach is exact and, for the case of M/G/1-type processes, compares favorably to the classic methods as shown by detailed time and space complexity analysis. Detailed experimentation further corroborates that ETAQA provides significantly less expensive solutions when compared to the classic methods. Alma Riska, Evgenia Smirni |
INFORMS J. Comput. | 2 |
| 2007 | Performance impacts of autocorrelated flows in multi-tiered systems
Ningfang Mi, Qi Zhang 0012, Alma Riska, Evgenia Smirni, Erik Riedel |
Perform. Evaluation | 4 |
| 2006 | Evaluating the Performability of Systems with Background JobsabstractAs most computer systems are expected to remain operational 24 hours a day, 7 days a week, they must complete maintenance work while in operation. This work is in addition to the regular tasks of the system and its purpose is to improve system reliability and availability. Nonetheless, additional work in the system, although labeled as best effort or low priority, still affects the performance of foreground tasks, especially if background/foreground work is non-preemptive. In this paper, we propose an analytic model to evaluate the performance trade-offs of the amount of background work that a storage system can sustain. The proposed model results in a quasi-birth-death (QBD) process that is analytically tractable. Detailed experimentation using a variety of workloads shows that under dependent arrivals both foreground and background performance strongly depends on system load. In contrast, if arrivals of foreground jobs are independent, performance sensitivity to load is reduced. The model identifies dependence in the arrivals of foreground jobs as an important characteristic that controls the decision of how much background load the system can accept to maintain high availability and performance gains Qi Zhang 0012, Ningfang Mi, Evgenia Smirni, Alma Riska, Erik Riedel |
DSN | 3 |
| 2006 | Load Unbalancing to Improve Performance under Autocorrelated TrafficabstractSize-based policies have been shown to successfully balance load and improve performance in homogeneous cluster environments where a dispatcher assigns a job to a server strictly based on the job size. While the success of size-based policies is based on separating jobs to different servers according to their sizes by avoiding the unfavorable performance effects of having short jobs been stuck behind long jobs, we show that their effectiveness quickly deteriorates in the presence of job arrivals that are characterized by correlation in their dependence structure. We propose a new policy that still strives to separate jobs according to their sizes, but this separation is biased by the effort to reduce the performance loss due to autocorrelation. As a result, not all servers are equally utilized (i.e., the load in the system becomes unbalanced) but the performance benefits of this load unbalancing are significant. The proposed policy can be used on-line, i.e., it does not assume any knowledge neither of the correlation structure of the arrival stream, nor of the job size distribution in the system. Via detailed trace-driven simulation we quantify the performance benefits of the proposed policy and we show that it can effectively self adjust its configuration parameters to improve performance under continuously changing workload conditions. Qi Zhang 0012, Ningfang Mi, Alma Riska, Evgenia Smirni |
ICDCS | 4 |
| 2005 | Power-aware resource allocation in high-end systems via online simulationabstractTraditionally, scheduling in high-end parallel systems focuses on how to minimize the average job waiting time and on how to maximize the overall system utilization. Despite the development of scheduling strategies that aim at maximizing system utilization, parallel supercomputing traces that span long time periods indicate that such systems are mostly underutilized. Much of the time there is simply not enough load to keep the system fully utilized, although time periods do exist where system utilization levels peak at nearly 95%. In this paper, we propose a new family of scheduling policies that aims at minimizing power consumption and cooling costs by selectively choosing to power down (or put in "sleep" mode) parts of the system during periods of low load. Our goal is the development of a scheduling mechanism that adaptively adjusts the number of processors to the offered load while meeting predefined service-level agreements (SLAs). This scheduling mechanism uses online simulation, i.e., lightweight simulation modules that can execute while the system and its scheduler are in operation, and can guide resource provisioning in parallel systems. Detailed experimentation using traces from the Parallel Workloads Archive indicates that the proposed online mechanism is a viable alternative to conserve energy while meeting performance-based SLAs. Barry Lawson 0001, Evgenia Smirni |
ICS | 2 |
| 2005 | Bridging ETAQA and Ramaswami's formula for the solution of M/G/1-type processes
Andreas Stathopoulos, Alma Riska, Zhili Hua, Evgenia Smirni |
Perform. Evaluation | 4 |
| 2005 | Workload-Aware Load Balancing for Clustered Web ServersabstractWe focus on load balancing policies for homogeneous clustered Web servers that tune their parameters on-the-fly to adapt to changes in the arrival rates and service times of incoming requests. The proposed scheduling policy, ADAPTLOAD, monitors the incoming workload and self-adjusts its balancing parameters according to changes in the operational environment such as rapid fluctuations in the arrival rates or document popularity. Using actual traces from the 1998 World Cup Web site, we conduct a detailed characterization of the workload demands and demonstrate how online workload monitoring can play a significant part in meeting the performance challenges of robust policy design. We show that the proposed load, balancing policy based on statistical information derived from recent workload history provides similar performance benefits as locality-aware allocation schemes, without requiring locality data. Extensive experimentation indicates that ADAPTLOAD results in an effective scheme, even when servers must support both static and dynamic Web pages. Qi Zhang 0012, Alma Riska, Evgenia Smirni, Gianfranco Ciardo |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2004 | ETAQA-MG1: an efficient technique for the analysis of a class of M/G/1-type processes by aggregation
Gianfranco Ciardo, Weizhen Mao, Alma Riska, Evgenia Smirni |
Perform. Evaluation | 4 |
| 2004 | An EM-based technique for approximating long-tailed data sets with PH distributions
Alma Riska, Vesselin Diev, Evgenia Smirni |
Perform. Evaluation | 3 |
| 2004 | Exact analysis of a class of GI/G/1-type performability modelsabstractWe present an exact decomposition algorithm for the analysis of Markov chains with a GI/G/1-type repetitive structure. Such processes exhibit both M/G/1-type & GI/M/1-type patterns, and cannot be solved using existing techniques. Markov chains with a GI/G/1 pattern result when modeling open systems which accept jobs from multiple exogenous sources, and are subject to failures & repairs; a single failure can empty the system of jobs, while a single batch arrival can add many jobs to the system. Our method provides exact computation of the stationary probabilities, which can then be used to obtain performance measures such as the average queue length or any of its higher moments, as well as the probability of the system being in various failure states, thus performability measures. We formulate the conditions under which our approach is applicable, and illustrate it via the performability analysis of a parallel computer system. Alma Riska, Evgenia Smirni, Gianfranco Ciardo |
IEEE Trans. Reliab. | 2 |
| 2002 | Efficient fitting of long-tailed data sets into hyperexponential distributionsabstractWe propose a new technique for fitting long-tailed data sets into hyperexponential distributions. The approach partitions the data set in a divide and conquer fashion and uses the expectation-maximization (EM) algorithm to fit the data of each partition into a hyperexponential distribution. The fitting results of all partitions are combined to generate the fitting for the entire data set. The new method is accurate and efficient and allows one to apply existing analytic tools to analyze the behavior of queueing systems that operate under workloads that exhibit long-tail behavior, such as queues in Internet-related systems. Alma Riska, Vesselin Diev, Evgenia Smirni |
GLOBECOM | 3 |
| 2002 | ADAPTLOAD: Effective Balancing in Custered Web Servers Under Transient Load ConditionsabstractWe focus on adaptive policies for load balancing in clustered web servers, based on the size distribution of the requested documents. The proposed scheduling policy, ADAPTLOAD, adapts its balancing parameters on-the-fly, according to changes in the behavior of the customer population such as fluctuations in the intensity of arrivals or document popularity. Detailed performance comparisons via simulation using traces from the 1998 World Cup show that ADAPTLOAD is robust as it consistently outperforms traditional load balancing policies, especially under conditions of transient overload. Alma Riska, Evgenia Smirni, Gianfranco Ciardo |
ICDCS | 3 |
| 2002 | Self-Adapting Backfilling Scheduling for Parallel SystemsabstractWe focus on non-FCFS job scheduling policies for parallel systems that allow jobs to backfill, i.e., to move ahead in the queue, given that they do not delay certain previously submitted jobs. Consistent with commercial schedulers that maintain multiple queues where jobs are assigned according to the user-estimated duration, we propose a self-adapting backfilling policy that maintains multiple job queues to separate short from long jobs. The proposed policy adjusts its configuration parameters by continuously monitoring the system and quickly reacting to sudden fluctuations in the workload arrival pattern and/or severe changes in resource demands. Detailed performance comparisons via simulation using actual supercomputing, traces from the parallel workload archive indicate that the proposed policy consistently outperforms traditional backfilling. Barry Lawson 0001, Evgenia Smirni, Daniela Puiu |
ICPP | 2 |
| 2002 | Multiple-Queue Backfilling Scheduling with Priorities and Reservations for Parallel Systems
Barry Lawson 0001, Evgenia Smirni |
JSSPP | 2 |
| 2002 | Exact aggregate solutions for M/G/1-type Markov processesabstractWe introduce a new methodology for the exact analysis of M/G/1-type Markov processes. The methodology uses basic, well-known results for Markov chains by exploiting the structure of the repetitive portion of the chain and recasting the overall problem into the computation of the solution of a finite linear system. The methodology allows for the calculation of the aggregate probability of a finite set of classes of states from the state space, appropriately defined. Further, it allows for the computation of a set of measures of interest such as the system queue length or any of its higher moments. The proposed methodology is exact. Detailed experiments illustrate that the methodology is also numerically stable, and in many cases can yield significantly less expensive solutions when compared with other methods, as shown by detailed time and space complexity analysis. Alma Riska, Evgenia Smirni |
SIGMETRICS | 2 |
| 2002 | Models of Parallel Applications with Large Computation and I/O RequirementsabstractA fundamental understanding of the interplay between computation and I/O activities in parallel applications that manipulate huge amounts of data is critical to achieving good application performance, as well as correctly characterizing the workloads of large-scale high-performance parallel systems. We present a formal model of the behavior of CPU and I/O interactions in scientific applications, from which we derive various formulas that characterize application performance. Our model captures the I/O and CPU activity at different levels of granularity, where results from the model are shown to be in excellent agreement with measurement data from a set of I/O-intensive applications. Using the formulas from our model, which explicitly take I/O activity into account, we also present examples of possible applications of the model. Emilia Rosti, Giuseppe Serazzi, Evgenia Smirni, Mark S. Squillante |
IEEE Trans. Software Eng. | 3 |
| 2001 | Algorithmic modifications to the Jacobi-Davidson parallel eigensolver to dynamically balance external CPU and memory loadabstractClusters of workstations (COWs) and SMPs have become popular and cost effective means of solving scientific problems. Because such environments may be heterogenous and/or time shared, dynamic load balancing is central to achieving high performance. Our thesis is that new levels of sophistication are required in parallel algorithm design and in the interaction of the algorithms with the runtime system. To support this thesis, we illustrate a novel approach for application-level balancing of external CPU and memory load on parallel iterative methods that employ some form of local preconditioning on each node. There are two key ideas. First, because all nodes need not perform their portion of the preconditioning phase to the same accuracy, the code can achieve perfect loadbalance, dynamically adapting to external CPU load, if we stop the preconditioning phase on all processors after a fixed amount of time. Second, if the program detects memory thrashing on a node, it recedes its preconditioning phase from that node, hopefully speeding the completion of competing jobs hence the relinquishing of their resources. We have implemented our load balancing approach in a state-of-the-art, coarse grain parallel Jacobi-Davidson eigensolver. Experimental results show that the new method adapts its algorithm based on runtime system information, without compromising the overall convergence behavior. We demonstrate the effectiveness of the new algorithm in a COW environment under (a) variable CPU load and (b) variable memory availability caused by competing applications. Richard Tran Mills, Andreas Stathopoulos, Evgenia Smirni |
ICS | 3 |
| 2001 | EQUILOAD: a load balancing policy for clustered web servers
Gianfranco Ciardo, Alma Riska, Evgenia Smirni |
Perform. Evaluation | 3 |
| 1999 | Adaptive CPU Scheduling Policies for Mixed Multimedia and Best-Effort WorkloadsabstractAs multimedia applications with real-time constraints rapidly invade today's desktops, it becomes increasingly important for the operating system to provide robust resource allocation mechanisms for both multimedia and traditional best-effort workloads. We present a flexible CPU scheduling policy that adjusts the CPU proportion allocated to each application class using recent history as a feedback mechanism. The algorithm quickly adapts to varying workload conditions and compares favorably with static proportional scheduling schemes for mixed workloads. Melissa A. Rau, Evgenia Smirni |
MASCOTS | 2 |
| 1999 | Performance evaluation of parallel systems
Paolo Cremonesi, Emilia Rosti, Giuseppe Serazzi, Evgenia Smirni |
Parallel Comput. | 4 |
| 1999 | ETAQA: An Efficient Technique for the Analysis of QBD-Processes by Aggregation
Gianfranco Ciardo, Evgenia Smirni |
Perform. Evaluation | 2 |
| 1998 | The Impact of I/O on Program Behavior and Parallel SchedulingabstractIn this paper we systematically examine various performance issues involved in the coordinated allocation of processor and disk resources in large-scale parallel computer systems. Models are formulated to investigate the I/O and computation behavior of parallel programs and workloads, and to analyze parallel scheduling policies under such workloads. These models are parameterized by measurements of parallel programs, and they are solved via analytic methods and simulation. Our results provide important insights into the performance of parallel applications and resource management strategies when I/O demands are not negligible. Emilia Rosti, Giuseppe Serazzi, Evgenia Smirni, Mark S. Squillante |
SIGMETRICS | 3 |
| 1998 | A methodology for the evaluation of multiprocessor non-preemptive allocation policies
Evgenia Smirni, Emilia Rosti, Lawrence W. Dowdy, Giuseppe Serazzi |
J. Syst. Archit. | 1 |
| 1998 | Lessons from Characterizing the Input/Output Behavior of Parallel Scientific Applications
Evgenia Smirni, Daniel A. Reed |
Perform. Evaluation | 1 |
| 1998 | Processor Saving Scheduling Policies for Multiprocessor SystemsabstractIn this paper, processor scheduling policies that "save" processors are introduced and studied. In a multiprogrammed parallel system, a "processor saving" scheduling policy purposefully keeps some of the available processors idle in the presence of work to be done. The conditions under which processor saving policies can be more effective than their greedy counterparts, i.e., policies that never leave processors idle in the presence of work to be done, are examined. Sensitivity analysis is performed with respect to application speedup, system size, coefficient of variation of the applications' execution time, variability in the arrival process, and multiclass workloads. Analytical, simulation, and experimental results show that processor saving policies outperform their greedy counterparts under a variety of system and workload characteristics. Emilia Rosti, Evgenia Smirni, Lawrence W. Dowdy, Giuseppe Serazzi, Kenneth C. Sevcik |
IEEE Trans. Computers | 2 |
| 1997 | Algorithmic Influences on I/O Access Patterns and Parallel File System PerformanceabstractFor many scalable parallel applications, the input/output (I/O) barrier rivals or exceeds that of computation and interprocessor communication. Consequently, scalable parallel secondary and tertiary storage systems are necessary to satisfy the resource demands of many national challenge problems. At present, one major challenge facing the designers of such storage systems is the wide range of I/O access patterns and the lack of general purpose file system policies that achieve high performance for variable I/O requirements. We analyze the I/O behavior of two scientific applications on the Intel Paragon XP/S. Although the two applications solve the same scientific problem and their I/O access patterns are qualitatively similar, their interactions with the file system are decidedly different. Our results show that appropriate tuning of file system policy parameters to I/O demands can significantly increase I/O throughput. Evgenia Smirni, Christopher L. Elford, A. J. Lavery, Andrew A. Chien |
ICPADS | 1 |
| 1995 | Performance Gains from Leaving Idle Processors in Multiprocessor Systems
Evgenia Smirni, Emilia Rosti, Giuseppe Serazzi, Lawrence W. Dowdy, Kenneth C. Sevcik |
ICPP (3) | 1 |
| 1995 | Analysis of Non-Work-Conserving Processor Partitioning Policies
Emilia Rosti, Evgenia Smirni, Giuseppe Serazzi, Lawrence W. Dowdy |
JSSPP | 2 |
| 1994 | Robust Partitioning Policies of Multiprocessor Systems
Emilia Rosti, Evgenia Smirni, Lawrence W. Dowdy, Giuseppe Serazzi, Brian M. Carlson |
Perform. Evaluation | 2 |
| 1993 | The KSR1: Experimentation and Modeling of PoststoreabstractKendall Square Research introduced the KSR1 system in 1991. The architecture is based on a ring of rings of 64-bit microprocessora. It is a distributed, shared memory system and is scalable. The memory structure is unique and is the key to understanding the system. Different levels of caching eliminates physical memory addressing and leads to the ALLCACHE™ scheme. Since requested data may be found in any of several caches, the initial access time is variable. Once pulled into the local (sub) cache, subsequent access times are fixed and minimal. Thus, the KSR1 is a Cache-Only Memory Architecture (COMA) system.This paper describes experimentation and an analytic model of the KSR1. The focus is on the poststore programmer option. With the poststore option, the programm er can elect to broadcast the updated value of a variable to all processors that might have a copy. This may save time for threads on other processors, but delays the broadcasting thread and places additional traffic on the ring. The specific issue addressed is to determine under what conditions poststore is beneficial. The analytic model and the experimental observations are in good agreement. They indicate that the decision to use poststore depends both on the application and the current system load. Emilia Rosti, Evgenia Smirni, Thomas D. Wagner, Amy W. Apon, Lawrence W. Dowdy |
SIGMETRICS | 2 |