Alma Riska

dblp:88/4059 · DBLP profile ↗
← Back
33ranked-venue papers
9as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 6 first-authorSoftware engineering, systems software and programming languages · 7 · 1 first-authorSecurity and privacy · 2Artificial intelligence and machine learning · 1Computer networks · 1 · 1 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Storage systems · 38% Performance modeling and evaluation · 26% Cloud and datacenter computing · 19%
Software engineering, system software, and programming languages
2 papers
Operating systems · 86% Runtime systems and virtual machines · 14%

Topics — the 18 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
queueing analysis
0.122008
Performance-Guided Load (Un)balancing under Autocorrelated Flows · IEEE Trans. Parallel Distributed Syst. 2008
Exact aggregate solutions for M/G/1-type Markov processes · SIGMETRICS 2002
Storage systems › i/o optimization
write optimization
0.112008
Idle Read After Write - IRAW · USENIX ATC 2008
Performance modeling and evaluation
workload characterization
0.122006
Disk Drive Level Workload Characterization · USENIX ATC, General Track 2006
Workload-Aware Load Balancing for Clustered Web Servers · IEEE Trans. Parallel Distributed Syst. 2005
Operating systems
resource management
0.112007
Efficient management of idleness in systems · SIGMETRICS 2007
Storage systems › i/o scheduling
storage scheduling
0.112006
Storage performance virtualization via throughput and latency control · ACM Trans. Storage 2006
Parallel and multicore computing › task scheduling › process scheduling
two-level scheduling
0.112006
Storage performance virtualization via throughput and latency control · ACM Trans. Storage 2006
Cloud and datacenter computing
cluster resource management and scheduling
0.112005
Workload-Aware Load Balancing for Clustered Web Servers · IEEE Trans. Parallel Distributed Syst. 2005
Storage systems
i/o scheduling
0.112005
An interposed 2-Level I/O scheduling framework for performance virtualization · SIGMETRICS 2005
Parallel and multicore computing
load balancing
0.112005
Workload-Aware Load Balancing for Clustered Web Servers · IEEE Trans. Parallel Distributed Syst. 2005
Hardware reliability and fault tolerance
soft errors
0.012004
Susceptibility of Commodity Systems and Software to Memory Soft Errors · IEEE Trans. Computers 2004
Storage systems › magnetic storage
hard disk drive
0.022008
Idle Read After Write - IRAW · USENIX ATC 2008
Disk Drive Level Workload Characterization · USENIX ATC, General Track 2006
Performance modeling and evaluation › markov models
markov process
0.012002
Exact aggregate solutions for M/G/1-type Markov processes · SIGMETRICS 2002
Cloud and datacenter computing › cluster resource management and scheduling
cluster scheduling
0.012008
Performance-Guided Load (Un)balancing under Autocorrelated Flows · IEEE Trans. Parallel Distributed Syst. 2008
Storage systems › i/o architecture › i/o subsystem
i/o path
0.012007
Evaluating Block-level Optimization Through the IO Path · USENIX ATC 2007
Energy-efficient computing
power management
0.012007
Efficient management of idleness in systems · SIGMETRICS 2007
Cloud and datacenter computing › datacenter storage
i/o consolidation
0.012006
Storage performance virtualization via throughput and latency control · ACM Trans. Storage 2006
Cloud and datacenter computing
virtualization
0.012005
An interposed 2-Level I/O scheduling framework for performance virtualization · SIGMETRICS 2005
Performance modeling and evaluation › workload characterization › commercial workloads
web server workload
0.012005
Workload-Aware Load Balancing for Clustered Web Servers · IEEE Trans. Parallel Distributed Syst. 2005

Methods — techniques the papers use, named apart from their topics

workload characterization · 0.2write scheduling · 0.1size-based scheduling · 0.1simulation · 0.1performance-guided load unbalancing · 0.1block-level optimization · 0.1work-conserving scheduling · 0.1rate control · 0.1feedback-driven scheduling · 0.1i/o scheduling · 0.1fault injection · 0.0
YearPublicationVenuePosition
2016 Adding data analytics capabilities to scaled-out object store
Cengiz Karakoyunlu, John A. Chandy, Alma Riska
J. Syst. Softw.3
2014 Agile middleware for scheduling: meeting competing performance requirements of diverse tasks
abstract
As the need for scaled-out systems increases, it is paramount to architect them as large distributed systems consisting of off-the-shelf basic computing components known as compute or data nodes. These nodes are expected to handle their work independently, and often utilize off-the-shelf management tools, like those offered by Linux, to differentiate priorities of tasks. While prioritization of background tasks in server nodes takes center stage in scaled-out systems, with many tasks associated with salient features such as eventual consistency, data analytics, and garbage collection, the standard Linux tools such as nice and ionice fail to adapt to the dynamic behavior of high priority tasks in order to achieve the best trade-off between protecting the performance of high priority workload and completing as much low priority work as possible. In this paper, we provide a solution by proposing a priority scheduling middleware that employs different policies to schedule background tasks based on the instantaneous resource requirements of the high priority applications running on the server node. The selection of policies is based on off-line and on-line learning of the high priority workload characteristics and the imposed performance impact due to low priority work. In effect, this middleware uses a {\em hybrid} approach to scheduling rather than a monolithic policy. We prototype and evaluate it via measurements on a test-bed and show that this scheduling middleware is robust as it effectively and autonomically changes the relative priorities between high and low priority tasks, consistently meeting their competing performance targets.
Feng Yan 0001, Shannon Hughes, Alma Riska, Evgenia Smirni
ICPE3
2013 Overcoming Limitations of Off-the-Shelf Priority Schedulers in Dynamic Environments
abstract
It is common nowadays to architect and design scaled-out systems with off-the-shelf computing components operated and managed by off-the-shelf open-source tools. While web services represent the critical set of services offered at scale, big data analytics is emerging as a preferred service to be colocated with cloud web services at a lower priority raising the need for off-the-shelf priority scheduling. In this paper we report on the perils of Linux priority scheduling tools when used to differentiate between such complex services. We demonstrate that simple priority scheduling utilities such as nice and ionice can result in dramatically erratic behavior. We provide a remedy by proposing an autonomic priority scheduling algorithm that adjusts its execution parameters based on on-line measurements of the current resource usage of critical applications. Detailed experimentation with a user-space prototype of the algorithm on a Linux system using popular benchmarks such as SPEC and TPC-W illustrate the robustness and versatility of the proposed technique, as it provides consistency to the expected performance of a high-priority application when running simultaneously with multiple low priority jobs.
Feng Yan 0001, Shannon Hughes, Alma Riska, Evgenia Smirni
MASCOTS3
2012 Automated Storage Tiering Using Markov Chain Correlation Based Clustering
abstract
In this paper, we develop an automated and adaptive framework that aims to move active data to high performance storage tiers and inactive data to low cost/high capacity storage tiers by learning patterns of the storage workloads. The framework proposed is designed using efficient Markov chain correlation based clustering method (MCC), which can quickly predict or detect any changes in the current workload based on what the system has experienced before. The workload data is first normalized and Markov chains are constructed from the dynamics of the IO loads of the data storage units. Based on the correlation of one-step Markov chain transition probabilities k-means method is employed to group the storage units that have similar behavior at each point. Such framework can then easily be incorporated in various resource management policies that aim at enhancing performance, reliability, availability. The predictive nature of the model, particularly makes a storage system both faster and lower-cost at the same time, because it only uses high performance tiers when needed, and uses low cost/high capacity tiers when possible.
Malak Alshawabkeh, Alma Riska, Adnan Sahin, Motasem Awwad
ICMLA (1)2
2012 Busy bee: how to use traffic information for better scheduling of background tasks
abstract
Computer systems, in general, and storage systems, in particular, rely on meeting their performance, reliability, and availability targets via scheduling of management and maintenance activities as background tasks.Such tasks may cause significant delays to user workload if scheduled extemporaneously. Here, we propose a scheduling policy for background tasks that is based on the statistical characteristics of the system's busy periods and that aims at completing background work expediently.Extensive trace-driven simulations show that the scheduling policy is robust and that it succeeds in completing background work faster than common practices while impacting user performance minimally.
Feng Yan 0001, Alma Riska, Evgenia Smirni
ICPE2
2011 Toward Automating Work Consolidation with Performance Guarantees in Storage Clusters
abstract
With most of today's systems being highly distributed, from data centers to cloud and storage clusters, there is a prevalent need for robust methodologies for work consolidation to improve load balancing but also to optimize non-traditional performance measures. Such alternative measures may include power savings, e.g., it may be desirable to shut down a lowly utilized node by moving some or all of its work to another node. In this paper, we present a methodology for distributed work consolidation that keeps track of the workload in the various nodes of the cluster and makes intelligent decisions on how much work to move from a sender node to a receiver node in order to minimally "affect" the performance of the receiver node or alternatively limit any performance degradation due to consolidation in a controlled way. The proposed methodology is based on continuously monitoring the workload on sender and receiver nodes, collecting lightweight statistics in the form of histograms of coarse granularity, and deciding when and how to initiate the work transfer. Extensive experimentation using trace-driven simulation confirms the robustness of the methodology.
Feng Yan 0001, Xenia Mountrouidou, Alma Riska, Evgenia Smirni
MASCOTS3
2011 Adaptive workload shaping for power savings on disk drives
abstract
In order to reduce the amount of power consumption in data centers, it is becoming necessary to shut off or slow down disks that are not actively serving user requests. In addition to exploiting disk drive idleness, system features are in place that shape a disk's workload by redirecting portions of it elsewhere, with the goal to expand the periods of idleness and the potential for power savings. In this paper, we propose several workload shaping techniques that determine which part of the working set to copy elsewhere using temporal and spatial access frequencies in the workload. These workload shaping techniques, used within an analytic estimation methodology, enable a fully automated framework that determines on-line for the current workload which, if any, shaping technique to activate such that the power saving benefits are maximized without violating performance targets. Extensive trace-driven evaluation shows that the proposed workload shaping techniques complement each-other with regard to their abilities to enhance idleness in disk drives for a wide range of workload characteristics. This results to added power savings in a data center even when performance targets are stringent and workloads intensive.
Xenia Mountrouidou, Alma Riska, Evgenia Smirni
ICPE2
2009 Autocorrelation-driven load control in distributed systems
abstract
In this paper, we propose a new approach for the development of load control policies in autonomic multitier systems. We control system load in a completely new way compared to existing policies: we leverage on the autocorrelation of service times and show that autocorrelation can be used to forecast future service requirements of requests and adaptively control system load. To the best of our knowledge, this is the first direct application of autocorrelation of service times to autonomic load control. We propose ALoC and D ALoC, two autocorrelation-driven policies that drop a percentage of the load in order to meet pre-defined quality-of-service levels in a distributed system. Both policies are easy to implement and rely on minimal assumptions. In particular, D ALoC is a fully no-knowledge measurement-based policy that self-adjusts its load control parameters based only on policy targets and on statistical information of requests served in the past. We illustrate the effectiveness of these new policies in a distributed multi-server setting via detailed trace driven simulations. We show that if these policies are employed in the server with a temporal dependent service process, then end-to-end response time, across all servers, reduces up to 80% by only dropping at most 13% of the incoming requests. Using real traces, we also show that, in the constrained case of being able to drop only from a portion of the incoming workload, our policy still improves request response time by up to 30%.
Ningfang Mi, Giuliano Casale, Qi Zhang 0012, Alma Riska, Evgenia Smirni
MASCOTS4
2009 Efficient management of idleness in storage systems
abstract
Various activities that intend to enhance performance, reliability, and availability of storage systems are scheduled with low priority and served during idle times. Under such conditions, idleness becomes a valuable “resource” that needs to be efficiently managed. A common approach in system design is to be nonwork conserving by “idle waiting”, that is, delay the scheduling of background jobs to avoid slowing down upcoming foreground tasks. In this article, we complement “idle waiting” with the “estimation” of background work to be served in every idle interval to effectively manage the trade-off between the performance of foreground and background tasks. As a result, the storage system is better utilized without compromising foreground performance. Our analysis shows that if idle times have low variability, then idle waiting is not necessary. Only if idle times are highly variable does idle waiting become necessary to minimize the impact of background activity on foreground performance. We further show that if there is burstiness in idle intervals, then it is possible to predict accurately the length of incoming idle intervals and use this information to serve more background jobs without affecting foreground performance.
Ningfang Mi, Alma Riska, Qi Zhang 0012, Evgenia Smirni, Erik Riedel
ACM Trans. Storage2
2008 Enhancing data availability in disk drives through background activities
abstract
Latent sector errors in disk drives affect only a few data sectors. They occur silently and are detected only when the affected area is accessed again. If a latent error is detected while the storage system is operating under reduced redundancy, i.e., during a RAID rebuild, then data loss may occur. Various features such as scrubbing and intra-disk data redundancy are proposed to detect and/or recover from latent errors and avoid data loss. While such features enhance data availability in the storage system, their execution may cause performance degradation. In this paper, we evaluate the effectiveness of scrubbing and intra-disk data redundancy in improving data availability while the overall goal is to maintain user performance within predefined bounds. We show that by treating them as low priority background activities and scheduling them efficiently during idle times, these features remain performance-wise transparent to the storage system user while still improving data reliability. Detailed trace-driven simulations show that the mean time to data loss (MTTDL) improves by up to 5 orders of magnitude if these features are implemented independently. By scheduling concurrently both scrubbing and intra-disk parity updates during idle times in disk drives, MTTDL improves by as much as 8 orders of magnitude.
Ningfang Mi, Alma Riska, Evgenia Smirni, Erik Riedel
DSN2
2008 Idle Read After Write - IRAW
Alma Riska, Erik Riedel
USENIX ATC1
2008 Performance-Guided Load (Un)balancing under Autocorrelated Flows
abstract
Size-based policies have been shown in the literature to effectively balance the load and improve performance in cluster environments. Size-based policies assign jobs to servers based on the job size and their performance improvements are an outcome of separating ";short"; from ";long"; jobs, by avoiding having short jobs waiting behind long jobs for service. In this paper, we present evidence that performance improvements due to this separation quickly vanish if the arrival process to the cluster is autocorrelated. Based on our observations, we devise a new size-based policy called D_EQAL that still strives to separate jobs to servers according to job size but this separation is now biased by an effort to reduce performance loss due to autocorrelation in the arrival flows to each server. As a result of this bias, all servers may not be equally utilized (i.e., the load in the system may be ";unbalanced";), but performance benefits become significant. D_EQAL can be used on-line as it does not assume any a priori knowledge of the incoming workload. Extensive simulations show the effectiveness of D_EQAL under autocorrelated and uncorrelated arrival streams and illustrate that the policy successfully self- adjusts the degree of load unbalancing based on monitored performance measures.
Qi Zhang 0012, Ningfang Mi, Alma Riska, Evgenia Smirni
IEEE Trans. Parallel Distributed Syst.3
2007 New Results on the Performance Effects of Autocorrelated Flows in Systems
abstract
Temporal dependence within the workload of any computing or networking system has been widely recognized as a significant factor affecting performance. More specifically, burstiness, as a form of temporal dependency, is catastrophic for performance. We use the autocorrelation function in a workload flow to formalize burstiness and also to characterize temporal dependence within a flow. We present results from two application areas: load balancing in a homogeneous cluster environment and capacity planning in a multi-tiered e-commerce system. For the load balancing problem, we show that if autocorrelation exists in the arrival stream to the cluster, classic load balancing policies become ineffective and solutions that focus on "unbalancing" the load offer superior performance. For the case of multi-tiered systems, we show that if there is autocorrelation in the flows, we observe the surprising result that in spite of the fact that the bottleneck resource in the system is far from saturation and that the measured throughput and utilizations of other resources are also modest, user response times are very high. For multi-tired systems, this underutilization of resources falsely indicates that the system can sustain higher capacities. We present analysis of the above phenomena that aims at the development better scheduling policies under auto correlated flows.
Evgenia Smirni, Qi Zhang 0012, Ningfang Mi, Alma Riska, Giuliano Casale
IPDPS4
2007 Efficient management of idleness in systems
abstract
No abstract available.
Ningfang Mi, Alma Riska, Qi Zhang 0012, Evgenia Smirni, Erik Riedel
SIGMETRICS2
2007 Evaluating Block-level Optimization Through the IO Path
Alma Riska, James Larkby-Lahet, Erik Riedel
USENIX ATC1
2007 ETAQA Solutions for Infinite Markov Processes with Repetitive Structure
abstract
We describe the ETAQA (efficient technique for the solution of quasi birth-death processes) approach for the exact analysis of M/G/1 and GI/M/1-type processes, and their intersection, i.e., quasi birth-death processes. ETAQA exploits the repetitive structure of the infinite portion of the chain and derives a finite system of linear equations. In contrast to the classic techniques for solution of such systems, the solution of this finite linear system does not provide the entire probability distribution of the state space, but simply allows calculation of the aggregate probability of a finite set of classes of states from the state space, appropriately defined. Nonetheless, these aggregate probabilities allow for computation of a rich set of measures of interest such as the system queue length or any of its higher moments. The proposed solution approach is exact and, for the case of M/G/1-type processes, compares favorably to the classic methods as shown by detailed time and space complexity analysis. Detailed experimentation further corroborates that ETAQA provides significantly less expensive solutions when compared to the classic methods.
Alma Riska, Evgenia Smirni
INFORMS J. Comput.1
2007 Performance impacts of autocorrelated flows in multi-tiered systems
Ningfang Mi, Qi Zhang 0012, Alma Riska, Evgenia Smirni, Erik Riedel
Perform. Evaluation3
2006 Evaluating the Performability of Systems with Background Jobs
abstract
As most computer systems are expected to remain operational 24 hours a day, 7 days a week, they must complete maintenance work while in operation. This work is in addition to the regular tasks of the system and its purpose is to improve system reliability and availability. Nonetheless, additional work in the system, although labeled as best effort or low priority, still affects the performance of foreground tasks, especially if background/foreground work is non-preemptive. In this paper, we propose an analytic model to evaluate the performance trade-offs of the amount of background work that a storage system can sustain. The proposed model results in a quasi-birth-death (QBD) process that is analytically tractable. Detailed experimentation using a variety of workloads shows that under dependent arrivals both foreground and background performance strongly depends on system load. In contrast, if arrivals of foreground jobs are independent, performance sensitivity to load is reduced. The model identifies dependence in the arrivals of foreground jobs as an important characteristic that controls the decision of how much background load the system can accept to maintain high availability and performance gains
Qi Zhang 0012, Ningfang Mi, Evgenia Smirni, Alma Riska, Erik Riedel
DSN4
2006 Load Unbalancing to Improve Performance under Autocorrelated Traffic
abstract
Size-based policies have been shown to successfully balance load and improve performance in homogeneous cluster environments where a dispatcher assigns a job to a server strictly based on the job size. While the success of size-based policies is based on separating jobs to different servers according to their sizes by avoiding the unfavorable performance effects of having short jobs been stuck behind long jobs, we show that their effectiveness quickly deteriorates in the presence of job arrivals that are characterized by correlation in their dependence structure. We propose a new policy that still strives to separate jobs according to their sizes, but this separation is biased by the effort to reduce the performance loss due to autocorrelation. As a result, not all servers are equally utilized (i.e., the load in the system becomes unbalanced) but the performance benefits of this load unbalancing are significant. The proposed policy can be used on-line, i.e., it does not assume any knowledge neither of the correlation structure of the arrival stream, nor of the job size distribution in the system. Via detailed trace-driven simulation we quantify the performance benefits of the proposed policy and we show that it can effectively self adjust its configuration parameters to improve performance under continuously changing workload conditions.
Qi Zhang 0012, Ningfang Mi, Alma Riska, Evgenia Smirni
ICDCS3
2006 Disk Drive Level Workload Characterization
Alma Riska, Erik Riedel
USENIX ATC, General Track1
2006 Storage performance virtualization via throughput and latency control
abstract
I/O consolidation is a growing trend in production environments due to increasing complexity in tuning and managing storage systems. A consequence of this trend is the need to serve multiple users and/or workloads simultaneously. It is imperative to ensure that these users are insulated from each other by virtualization in order to meet any service-level objective (SLO). Previous proposals for performance virtualization suffer from one or more of the following drawbacks: (1) They rely on a fairly detailed performance model of the underlying storage system; (2) couple rate and latency allocation in a single scheduler, making them less flexible; or (3) may not always exploit the full bandwidth offered by the storage system.This article presents a two-level scheduling framework that can be built on top of an existing storage utility. This framework uses a low-level feedback-driven request scheduler, called AVATAR, that is intended to meet the latency bounds determined by the SLO. The load imposed on AVATAR is regulated by a high-level rate controller, called SARC, to insulate the users from each other. In addition, SARC is work-conserving and tries to fairly distribute any spare bandwidth in the storage system to the different users. This framework naturally decouples rate and latency allocation. Using extensive I/O traces and a detailed storage simulator, we demonstrate that this two-level framework can simultaneously meet the latency and throughput requirements imposed by an SLO, without requiring extensive knowledge of the underlying storage system.
Jianyong Zhang, Anand Sivasubramaniam, Qian Wang 0029, Alma Riska, Erik Riedel
ACM Trans. Storage4
2005 Storage Performance Virtualization via Throughput and Latency Control
abstract
I/O consolidation is a growing trend in production environments due to the increasing complexity in tuning and managing storage systems. A consequence of this trend is the need to serve multiple users/workloads simultaneously. It is imperative to make sure that these users are insulated from each other by visualization in order to meet any service level objective (SLO). This paper presents a 2-level scheduling framework that can be built on top of an existing storage utility. This framework uses a low-level feedback-driven request scheduler, called AVATAR, that is intended to meet the latency bounds determined by the SLO. The load imposed on AVATAR is regulated by a high-level rate controller, called SARC, to insulate the users from each other. In addition, SARC is work-conserving and tries to fairly distribute any spare bandwidth in the storage system to the different users. This framework naturally decouples rate and latency allocation. Using extensive I/O traces and a detailed storage simulator, we demonstrate that this 2-level framework can simultaneously meet the latency and throughput requirements imposed by an SLO, without requiring extensive knowledge of the underlying storage system.
Jianyong Zhang, Anand Sivasubramaniam, Qian Wang 0029, Alma Riska, Erik Riedel
MASCOTS4
2005 An interposed 2-Level I/O scheduling framework for performance virtualization
abstract
No abstract available.
Jianyong Zhang, Anand Sivasubramaniam, Alma Riska, Qian Wang 0029, Erik Riedel
SIGMETRICS3
2005 Bridging ETAQA and Ramaswami's formula for the solution of M/G/1-type processes
Andreas Stathopoulos, Alma Riska, Zhili Hua, Evgenia Smirni
Perform. Evaluation2
2005 Workload-Aware Load Balancing for Clustered Web Servers
abstract
We focus on load balancing policies for homogeneous clustered Web servers that tune their parameters on-the-fly to adapt to changes in the arrival rates and service times of incoming requests. The proposed scheduling policy, ADAPTLOAD, monitors the incoming workload and self-adjusts its balancing parameters according to changes in the operational environment such as rapid fluctuations in the arrival rates or document popularity. Using actual traces from the 1998 World Cup Web site, we conduct a detailed characterization of the workload demands and demonstrate how online workload monitoring can play a significant part in meeting the performance challenges of robust policy design. We show that the proposed load, balancing policy based on statistical information derived from recent workload history provides similar performance benefits as locality-aware allocation schemes, without requiring locality data. Extensive experimentation indicates that ADAPTLOAD results in an effective scheme, even when servers must support both static and dynamic Web pages.
Qi Zhang 0012, Alma Riska, Evgenia Smirni, Gianfranco Ciardo
IEEE Trans. Parallel Distributed Syst.2
2004 ETAQA-MG1: an efficient technique for the analysis of a class of M/G/1-type processes by aggregation
Gianfranco Ciardo, Weizhen Mao, Alma Riska, Evgenia Smirni
Perform. Evaluation3
2004 An EM-based technique for approximating long-tailed data sets with PH distributions
Alma Riska, Vesselin Diev, Evgenia Smirni
Perform. Evaluation1
2004 Susceptibility of Commodity Systems and Software to Memory Soft Errors
abstract
It is widely understood that most system downtime is accounted for by programming errors and administration time. However, a growing body of work has indicated an increasing cause of downtime may stem from transient errors in computer system hardware due to external factors, such as cosmic rays. This work indicates that moving to denser semiconductor technologies at lower voltages has the potential to increase these transient errors. In this paper, we investigate the susceptibility of commodity operating systems and applications on commodity PC processors to these soft-errors and we introduce ideas regarding the improved recovery from these transient errors in software. Our results indicate that, for the Linux kernel and a Java virtual machine running sample workloads, many errors are not activated, mostly due to overwriting. In addition, given current and upcoming microprocessor support, our results indicate that those errors activated, which would normally lead to system reboot, need not be fatal to the system if software knowledge is used for simple software recovery. Together, they indicate the benefits of simple memory soft error recovery handling in commodity processors and software.
Alan Messer, Philippe Bernadat, Guangrui Fu, DeQing Chen, Zoran Dimitrijevic, David Jeun Fung Lie, Durga Mannaru, Alma Riska, Dejan S. Milojicic
IEEE Trans. Computers8
2004 Exact analysis of a class of GI/G/1-type performability models
abstract
We present an exact decomposition algorithm for the analysis of Markov chains with a GI/G/1-type repetitive structure. Such processes exhibit both M/G/1-type & GI/M/1-type patterns, and cannot be solved using existing techniques. Markov chains with a GI/G/1 pattern result when modeling open systems which accept jobs from multiple exogenous sources, and are subject to failures & repairs; a single failure can empty the system of jobs, while a single batch arrival can add many jobs to the system. Our method provides exact computation of the stationary probabilities, which can then be used to obtain performance measures such as the average queue length or any of its higher moments, as well as the probability of the system being in various failure states, thus performability measures. We formulate the conditions under which our approach is applicable, and illustrate it via the performability analysis of a parallel computer system.
Alma Riska, Evgenia Smirni, Gianfranco Ciardo
IEEE Trans. Reliab.1
2002 Efficient fitting of long-tailed data sets into hyperexponential distributions
abstract
We propose a new technique for fitting long-tailed data sets into hyperexponential distributions. The approach partitions the data set in a divide and conquer fashion and uses the expectation-maximization (EM) algorithm to fit the data of each partition into a hyperexponential distribution. The fitting results of all partitions are combined to generate the fitting for the entire data set. The new method is accurate and efficient and allows one to apply existing analytic tools to analyze the behavior of queueing systems that operate under workloads that exhibit long-tail behavior, such as queues in Internet-related systems.
Alma Riska, Vesselin Diev, Evgenia Smirni
GLOBECOM1
2002 ADAPTLOAD: Effective Balancing in Custered Web Servers Under Transient Load Conditions
abstract
We focus on adaptive policies for load balancing in clustered web servers, based on the size distribution of the requested documents. The proposed scheduling policy, ADAPTLOAD, adapts its balancing parameters on-the-fly, according to changes in the behavior of the customer population such as fluctuations in the intensity of arrivals or document popularity. Detailed performance comparisons via simulation using traces from the 1998 World Cup show that ADAPTLOAD is robust as it consistently outperforms traditional load balancing policies, especially under conditions of transient overload.
Alma Riska, Evgenia Smirni, Gianfranco Ciardo
ICDCS1
2002 Exact aggregate solutions for M/G/1-type Markov processes
abstract
We introduce a new methodology for the exact analysis of M/G/1-type Markov processes. The methodology uses basic, well-known results for Markov chains by exploiting the structure of the repetitive portion of the chain and recasting the overall problem into the computation of the solution of a finite linear system. The methodology allows for the calculation of the aggregate probability of a finite set of classes of states from the state space, appropriately defined. Further, it allows for the computation of a set of measures of interest such as the system queue length or any of its higher moments. The proposed methodology is exact. Detailed experiments illustrate that the methodology is also numerically stable, and in many cases can yield significantly less expensive solutions when compared with other methods, as shown by detailed time and space complexity analysis.
Alma Riska, Evgenia Smirni
SIGMETRICS1
2001 EQUILOAD: a load balancing policy for clustered web servers
Gianfranco Ciardo, Alma Riska, Evgenia Smirni
Perform. Evaluation2