EDBT 2026 Demo / reviewers in the wild / expert
Yixin Diao
dblp:48/6719
· DBLP profile ↗
34ranked-venue papers
20as first author
0since 2021 · last 2018
0000-0003-0498-7917ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 5 first-authorSystems, architecture and hardware · 3 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 61% Cloud and datacenter computing · 30% Performance modeling and evaluation · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Database system architecture and tuning · 77% Indexing and storage engines · 23% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems › self-adaptive systems
autonomic computing |
0.1 | 1 | 2005 | A control theory foundation for self-managing computing systems · IEEE J. Sel. Areas Commun. 2005 |
Cloud and datacenter computing
resource management |
0.1 | 1 | 2005 | A control theory foundation for self-managing computing systems · IEEE J. Sel. Areas Commun. 2005 |
Distributed systems › self-adaptive systems › autonomic computing
self-managing systems |
0.1 | 1 | 2005 | A control theory foundation for self-managing computing systems · IEEE J. Sel. Areas Commun. 2005 |
Indexing and storage engines
buffer management |
0.0 | 1 | 2006 | Adaptive Self-tuning Memory in DB2 · VLDB 2006 |
Methods — techniques the papers use, named apart from their topics
self-diagnosis and self-repair · 0.1feedback regulation · 0.1control theory · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Improving response accuracy for classification- based conversational IT servicesabstractConversational IT services are expected to reduce user wait times and improve overall customer satisfaction. Cloud-based solutions are readily available for enterprise subject matter experts (SMEs) to train user-question classifiers and build conversational services with little effort. However, methodologies that the SMEs can use to improve the response accuracy and conversation quality are merely stated and evaluated. In complex service scenarios such as software support, the scope of topics is typically large and the training samples are often limited. Thus, training the classifier based on labeled samples of plain user utterances is not effective in most cases. In this paper, we identify several methods for improving classification quality and evaluate them in concrete training set scenarios. Particularly, a process-based methodology is described that builds and refines on top of service domain knowledge in order to develop a scalable solution for training accurate conversation services. Enterprises and service providers are continuously seeking new ways to improve customer experience on working with IT systems, where user wait times and service resolution quality are critical business metrics. One of the latest trends is the use of conversational IT services. Customers can interact with a conversational service to express their questions in natural language and the system can automatically return relevant answers or execute back-end processes for automated actions. Various text classification techniques have been developed and applied to understand the user questions and trigger the correct responses. For instance, in the context of IT software support, customers can use conversational systems to get answers about software product errors, licenses, or upgrade processes. While the potential benefits of building conversational services are huge, it is often difficult to effectively train classification models that cover well the scope of realistically complex services. In this paper, we propose a training methodology that addresses the limitations in both the scope of topics and the scarcity of the training set. We further evaluate the proposed methodology in a real service support scenario and share the lessons learned. Yixin Diao, Daniela Rosu 0001 |
NOMS | 1 |
| 2018 | Guest Editorial: Special Section on Advances in Big Data Analytics for ManagementabstractCloud and network analytics can harness the immense stream of operational data from clouds and networks, and can perform analytics processing to improve reliability, automated configuration, performance, and optimized network management in general. In this area, we have witnessed a growing trend towards using statistical analysis and machine learning techniques to improve operations and management of IT systems and networks. Giuliano Casale, Yixin Diao, Marco Mellia, Rajiv Ranjan 0001, Nur Zincir-Heywood |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2016 | Optimizing ATM cash flow network managementabstractAutomated Teller Machine (ATM) service providers are increasingly challenged with improving the quality of customer service while reducing the cost of cash flow management. Effectively balancing the need to have enough cash in the ATMs to avoid out-of-cash incidents as well as to reduce the cash interest cost and the cash refill cost challenges the most experienced cash flow management teams. In this paper we propose an optimization framework for managing the ATM cash flow network. The interactions among various constraints and cost factors are included in the framework to allow decision-making regarding the optimal cash refill amount and schedule. We demonstrate the effectiveness of the proposed approach using sample data from a large commercial bank. Yixin Diao, Soumitra Sarkar, Ea-Ee Jan |
NOMS | 1 |
| 2016 | Guest Editors' Introduction: Special Issue on Big Data Analytics for ManagementabstractCloud and network analytics can harness the immense stream of operational data from clouds and networks, and can perform analytics processing to improve reliability, configuration, performance, and security management. In particular, we see a growing trend towards using statistical analysis and machine learning to improve operations and management of IT systems and networks. Giuliano Casale, Yixin Diao, Hanan Lutfiyya, Philippe Owezarski, Danny Raz |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2015 | Modeling service variability in complex service delivery operationsabstractOne of the key promises of IT strategic outsourcing is to deliver greater IT service management through better quality and lower cost. However, this raises a critical question on how to model highly variable services for diverse customers with heterogeneous infrastructure and service demands. In this paper we propose the use of statistical learning approaches for service operation variability modeling. Specifically, we use the partial least squares regression that projects service attributes to explain the service volume variability, and the decision tree approach to model the service effort based on categorical customer and service properties. We demonstrate the applicability of the proposed methodology using data from a large IT service delivery environment. Yixin Diao, Larisa Shwartz |
CNSM | 1 |
| 2014 | A framework for predicting service delivery efforts using IT infrastructure-to-incident correlationabstractPredicting IT infrastructure performance under varying conditions, e.g., the addition of a new server or increased transaction loads, has become a typical IT management exercise. However, within a service delivery context, enterprise clients are demanding predictive analytics that outline future “costs” associated with changing conditions. The service delivery staffing costs incurred in addressing problems and requests (arriving in the form of incident and other problem tickets) in the managed environment is especially of high importance. This paper describes an analytical study addressing such cost prediction. Specifically, a novel approach is described in which support vector regression is used to predict service delivery workloads (measured by ticket volumes) based on managed server characteristics Additionally, a proposed framework combining various analytical models is proposed to predict service delivery staffing requirements under changing IT infrastructure characteristics and conditions. Detailed descriptions of the workload prediction techniques, as well as an evaluation using data from an actual large service delivery engagement, are presented. Joel W. Branch, Yixin Diao, Larisa Shwartz |
NOMS | 2 |
| 2014 | Predicting service delivery cost for non-standard service level agreementsabstractOne of the key promises of IT strategic outsourcing is to deliver greater IT service management through lower cost. However, this raises a critical question on how to predict service delivery cost during the service engagement phase where nonstandard service level agreements (SLAs) are negotiated and detailed service modeling data are not available. In this paper we propose a modeling framework that uses queueing model based approaches to estimate the impact of SLAs on the delivery cost. We further propose a set of approximation techniques to address the complexity of service delivery and an optimization model to predict the delivery cost subject to service level constraints and service stability conditions. We demonstrate the applicability of the proposed methodology using data from a large IT service delivery environment. Yixin Diao, Linh Lam, Larisa Shwartz, David M. Northcutt |
NOMS | 1 |
| 2014 | Modeling the Impact of Service Level Agreements During Service EngagementabstractOne of the key promises of IT strategic outsourcing is to deliver greater IT service management through lower cost. However, this raises a critical question: How can one predict the service delivery cost that will deliver the promised service level agreements (SLAs)? This is particularly challenging since such prediction is mostly needed during the service engagement phase where the SLAs and the delivery cost are negotiated, and the detailed service modeling data are not available. In this paper, we propose a modeling framework that uses queueing-model-based approaches to estimate the impact of SLAs on the delivery cost. We further propose a set of approximation techniques to address the complexity of service delivery and an optimization model to predict the delivery cost subject to service-level constraints and service stability conditions. We demonstrate the applicability of the proposed methodology using data from a large IT service delivery environment. Yixin Diao, Linh Lam, Larisa Shwartz, David M. Northcutt |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2013 | SLA impact modeling for service engagementabstractDuring the customer engagement phase it is critical for the service providers to estimate the impact of service level constraints on service personnel needs. However, it is often difficult due to the implication from customer workload. In this paper we propose an SLA impact evaluation methodology that uses queueing models to quantitatively evaluate the impact of SLAs to the engagement cost model. Yixin Diao, Linh Lam, Larisa Shwartz, David M. Northcutt |
CNSM | 1 |
| 2012 | Analysis of operational data to improve performance in service delivery systems
Yixin Diao, Aliza R. Heching |
CNSM | 1 |
| 2012 | A dynamic request dispatching system for IT service management
David Loewenstern, Yixin Diao |
CNSM | 2 |
| 2012 | Closed loop performance management for service delivery systemsabstractIT service delivery becomes an increasingly challenging business as customers demand improved quality of service while providers are driven to reduce the cost of delivery. While effective service delivery requires advances in many areas including workload management and workforce optimization, in this paper we focus on service request dispatching decision-making. Specifically, we propose a closed loop performance management solution that leverages feedback controllers to dynamically adjust the priority of service requests considering both (static) contractual service attainment targets and (dynamic) attainment levels achieved. We demonstrate the applicability of the proposed approach in a simulation testbed that models a large IT service delivery environment, and compare its performance with two open loop dispatching policies. Yixin Diao, Aliza R. Heching |
NOMS | 1 |
| 2011 | Staffing optimization in complex service delivery systems
Yixin Diao, Aliza R. Heching |
CNSM | 1 |
| 2011 | The cost of service quality in IT OutsourcingabstractIn recent years, the IT services industry is under continuous pressure to improve the quality of its services, while doing so at reduced prices. In order to determine quality, the IT industry seems to be converging towards a set of commonly used metrics, including but not limited to, equipment availability, the time to resolve incidents, etc. In order to improve the quality of their services, service providers need to adopt a consistent and continuous focus on their internal processes, the skills of their people, the organizational structure and so on. On the other hand, the IT service industry has been traditionally labor intensive. Even though a wide range of management tools is available to support even the most complex management tasks, the overall service management process still relies on the existence of humans to make important decisions, perform complex administration tasks and monitor the performance and effectiveness of the overall process. This paper studies the tradeoff between the cost of a service (as manifested by the staffing level) and the corresponding quality metrics (subject to the service level agreement). In particular, we study the cost of service quality through the use of an optimization model that takes into account the constraints and cost factors typically encountered in a service provider environment. Nikos Anerousis, Yixin Diao, Aliza R. Heching |
Integrated Network Management | 2 |
| 2010 | Dispatch tooling for global service deliveryabstractTool development to support service management is a problem in optimizing service quality and reducing redundant work subject to competing constraints, both technical requirements and political realities. However, many of the constraints only become clear over time, and often only after a series of iterations as tooling exposes new or previously discounted constraints. We present a case study: the evolution of tooling for global dispatch of work orders supporting incident tickets. We discuss different approaches, how they highlighted different constraints, and even how they competed with each other, and conclude with a discussion of approaches to balancing constraints, many of which are not initially apparent and some of which shift over time. David Loewenstern, Melissa J. Buco, Yixin Diao, Heiko Ludwig, Christopher Ward |
CNSM | 3 |
| 2010 | Elements of system design optimization in service quality managementabstractServices Quality is an area of opportunity for IT service providers to innovate and deliver outstanding results to their customers. The objective is to minimize the variance of key quality indicators and deliver predictable capabilities. In this paper we adopt the Lean Sigma methodology from the manufacturing domain as a quality control framework and propose an optimized system model for managing predictability and reducing cost in the IT incident management process. The model establishes guidelines for receiving, classifying and distributing work in a service delivery organization. It defines metrics, controls and management objectives. Using simulation, we conduct an extensive numerical study to show how the model behaves under different operational scenarios that reflect a diverse skill base, the presence of service level objectives, and incoming work with varying levels of complexity. In addition, we provide managerial insight into what drives performance and illustrate the trade-offs between different service delivery designs. Nikos Anerousis, Yixin Diao, Aliza R. Heching |
NOMS | 2 |
| 2009 | Rule-Based Problem Classification in IT Service ManagementabstractProblem management is a critical and expensive element for delivering IT service management and touches various levels of managed IT infrastructure. While problem management has been mostly reactive, recent work is studying how to leverage large problem ticket information from similar IT infrastructures to probatively predict the onset of problems. Because of the sheer size and complexity of problem tickets, supervised learning algorithms have been the method of choice for problem ticket classification, relying on labeled (or pre-classified) tickets from one managed infrastructure to automatically create signatures for similar infrastructures. However, where there are insufficient preclassified data, leveraging human expertise to develop classification rules can be more efficient. In this paper, we describe a rule-based crowdsourcing approach, where experts can author classification rules and a social networking-based platform (called xPad) is used to socialize and execute these rules by large practitioner communities. Using real data sets from several large IT delivery centers, we demonstrate that this approach balances between two key criteria: accuracy and cost effectiveness. Yixin Diao, Hani Jamjoom, David Loewenstern |
IEEE CLOUD | 1 |
| 2009 | Multi-tenant solution for IT service management: A quantitative study of benefitsabstractThe very competitive business climate dictates efficient and cost effective delivery and support of IT services. Compounded by complexity of IT environments and criticality of IT to business success, IT service providers seek multi-tenant solutions to reduce operational cost and improve service quality. In this paper we consider a multi-tenant solution for IT service management and examine its critical aspects for realizing business benefits. We conduct a quantitative study of its benefits using a complexity based value assessment methodology. By regarding complexity as a substitute for potential labor cost, we estimate the business value of the multi-tenant solution before its actual deployment. Larisa Shwartz, Yixin Diao, Genady Grabarnik |
Integrated Network Management | 2 |
| 2008 | Estimating business value of IT services through process complexity analysisabstractIn this paper we propose a methodology to estimate business value of IT services. We investigate how to apply quantitative studies for IT service management processes that allows us to link measurable performance improvements with concrete business value. Specifically, we follow an Information Technology Infrastructure Library (ITIL) defined strategy combined with process complexity analysis techniques. This helps to reason about inefficiencies in IT service processes due to problems with coordination of different roles, lack of support for task execution, and complexities for getting the source of information. Our approach consists in (1) identifying the process context using ITIL as a reference framework, (2) quantifying process baseline with typical task execution time and underlying complexity, (3) estimating performance improvement achieved by tooling deployment or process transformation, and (4) estimating business value derived for various business cases. We illustrate our methodology using the change management process as defined in IBM Tivoli Unified Process (ITUP) and estimating the business value of implementing an application discovery tool and a change management tool. Yixin Diao, Kamal Bhattacharya |
NOMS | 1 |
| 2008 | Using mixed integer programming to schedule it change requestsabstractChange management is viewed as one of the most important processes in IT management with the main challenge of scheduling the change requests. That is, given a set of change requests for different parts of an IT system, how one should schedule changes to minimize the costs and risks of service delivery. In this paper, we approach the change request scheduling problem using a simple optimization model and consider both deterministic and stochastic cases. We show that these models can be converted to the Mixed Integer Programming (MIP) framework for which promising solution techniques exist. We conclude the work with an illustrative example as an instance of our model. Leila Zia, Yixin Diao, Christopher Ward, Kamal Bhattacharya |
NOMS | 2 |
| 2007 | Predicting Labor Cost through IT Management Complexity MetricsabstractWe propose a model for relating IT management complexity metrics to key business-level performance metrics like time and labor cost. In particular, we address the problem of quantifying and predicting the value that automation and IT service management process transformation will yield before their actual deployment. Our approach looks at this problem from a different, new perspective by regarding complexity as a surrogate for potential labor cost and human-error-induced problems: It consists in (1) assessing and evaluating the complexity of IT management processes and procedures, (2) separately measuring business-level performance metrics (3) relating the collected complexity metrics to business-level performance metrics by means of a quantitative model, and (4) validating the model through a field study. Algorithms are presented for selecting a subset of the complexity metrics to use as explanatory variables in a quantitative model and for constructing the quantitative model itself. Besides improving decision making for deploying automation technologies and transforming IT service management processes, our quantitative IT management complexity model can help service providers and outsourcers predict the amount of human effort and skills that will be needed to provide a given service, thus allowing them to more effectively evaluate costs and benefits of automation technologies and IT management process transformations. Yixin Diao, Alexander Keller 0002, Sujay S. Parekh, Vladislav V. Marinov |
Integrated Network Management | 1 |
| 2006 | Controlling Quality of Service in Multi-Tier Web ApplicationsabstractThe need for service differentiation in Internet services has motivated interest in controlling multi-tier web applications. This paper describes a tier-to-tier (T2T) management architecture that supports decentralized actuator management in multi-tier systems, and a testbed implementation of this architecture using commercial software products. Based on testbed experiments and analytic models, we gain insight into the value of coordinated exploitation of actuators on multiple tiers, especially considerations for control efficiency and control granularity. For control efficiency, we show that more effective utilization of tiers can be achieved by using actuators on the bottleneck tier rather than only using actuators on the entry tier. For granularity of control (the ability to achieve a wide range of service level objectives) we show that a fine granularity of control can be achieved through a coordinated, cross-tier exploitation of coarse grained actuators (e.g., multiprogramming level), an approach that can greatly reduce controllerinduced variability. Yixin Diao, Joseph L. Hellerstein, Sujay S. Parekh, Hidayatullah Shaikh, Maheswaran Surendra |
ICDCS | 1 |
| 2006 | Dynamic Adaptation of Temporal Event Correlation for QoS Management in Distributed SystemsabstractTemporal event correlation is essential to managing quality of service in distributed systems, especially correlating events from multiple components to detect problems with availability, performance, and denial of service attacks. Two challenges in temporal event correlation are: (1) handling lost events and (2) dealing with inaccurate clocks. We show that both challenges are related to event propagation delays that result from contention for network and server resources. We develop an approach to adjusting the timer values of event correlation rules based on propagation delays in order to reduce missed alarms and false alarms. Our approach has three parts: an infrastructure for real-time measurement of propagation delay, a statistical approach to estimating propagation delays, and a controller that uses estimates of propagation delays to update timer values in temporal rules. Our approach eliminates the need for manual adjustments of timer values. Further, studies of a prototype implementation suggest that our approach produces results that are at least as good as an optimal fixed adjustment in timer values Rean Griffith, Joseph L. Hellerstein, Gail E. Kaiser, Yixin Diao |
IWQoS | 4 |
| 2006 | Modeling Differentiated Services of Multi-Tier Web ApplicationsabstractIn this paper we present a hybrid performance model for modeling differentiated service of multi-tier web applications with per-tier concurrency limits, cross-tier interactions, as well as a work-conserving resource allocation model. The service dependencies between multiple tiers are captured first using a layered queueing model. We then show how to model per-tier concurrency limits and service differentiation between multiple classes while maintaining work conservation at each tier. We use a function approximation approach combined with a coupled processor model. Our model is calibrated from an actual multitier J2EE testbed, and we show the ability of the model to accurately model common performance metrics. Our proposed (layered) model shows 78% improvement in root mean square error over a single-tier machine repair model as well as a tandem queue model. We also demonstrate one application of the model for model-based resource allocation. Yixin Diao, Joseph L. Hellerstein, Sujay S. Parekh, Hidayatullah Shaikh, Maheswaran Surendra, Asser N. Tantawi |
MASCOTS | 1 |
| 2006 | Adaptive Self-tuning Memory in DB2
Adam J. Storm, Christian Garcia-Arellano, Sam Lightstone, Yixin Diao, Maheswaran Surendra |
VLDB | 4 |
| 2005 | Control of fair queueing: modeling, implementation, and experiencesabstractFeedback control of QoS-aware servers has recently gained much popularity due to its robustness in the face of uncertainty and modeling errors. Performance of servers is characterized by the behavior of queues, which constitute the main elements of the control loop. The central role of queues in the loop motivates understanding their behavior in the context of feedback control schemes. A popular queueing policy in servers where different traffic classes must be allocated a different share of a common resource is fair queueing (FQ). This paper investigates the interactions between a FQ element and a feedback controller. It is shown that the FQ element introduces challenges that render simple feedback control ineffective and potentially unstable. These challenges are systematically exposed, explained, and resolved. An extended feedback control scheme for the FQ element is subsequently developed. The scheme is tested on an experimental prototype demonstrating higher predictability and an order of magnitude improvement in responsiveness over the initial design. The results of the paper apply in general to most systems that use a dynamic processor sharing approach for service differentiation. Yixin Diao, Shashi Parekh, Tarek F. Abdelzaher |
Integrated Network Management | 2 |
| 2005 | A control theory foundation for self-managing computing systemsabstractThe high cost of operating large computing installations has motivated a broad interest in reducing the need for human intervention by making systems self-managing. This paper explores the extent to which control theory can provide an architectural and analytic foundation for building self-managing systems. Control theory provides a rich set of methodologies for building automated self-diagnosis and self-repairing systems with properties such as stability, short settling times, and accurate regulation. However, there are challenges in applying control theory to computing systems, such as developing effective resource models, handling sensor delays, and addressing lead times in effector actions. We propose a deployable testbed for autonomic computing (DTAC) that we believe will reduce the barriers to addressing research problems in applying control theory to computing systems. The initial DTAC architecture is described along with several problems that it can be used to investigate. Yixin Diao, Joseph L. Hellerstein, Sujay S. Parekh, Rean Griffith, Gail E. Kaiser, Dan B. Phung |
IEEE J. Sel. Areas Commun. | 1 |
| 2004 | Incorporating Cost of Control into the Design of a Load Balancing ControllerabstractLoad balancing is widely used in computing systems as a way to optimize performance by reducing bottleneck utilizations, such as adjusting the size of buffer pools to balance resource demands in a database management system. Load balancing is generally approached as a constrained optimization problem in which only the benefits of load balancing are considered. However, the costs of control are important as well. Herein, we study the value of including in controller design the trade-off between the cost of transient imbalances in resource utilizations and the cost of changing resource allocations. An example of the latter are actions such as resizing buffer pools that can reduce throughputs. This is because requests for data in pools whose memory is reduced immediately have longer access times whereas requests for data in pools whose memory is increased must fill this memory with data from disk before accessed times are reduced. We frame our study of control costs in terms of the widely used linear quadratic regulator (LQR). We develop a cost model that allows us to specify the LQR Q and R matrices based on the impact on system performance of changing resource allocations and transient load imbalances. Our studies of a DB2 universal database server using benchmarks for online transaction processing and decision support workloads show that incorporating our cost model into the MIMO LQR controller results in a 14% improvement in performance beyond that achieved by dynamically allocating the size of buffers without properly considering the cost of control. Yixin Diao, Joseph L. Hellerstein, Adam J. Storm, Maheswaran Surendra, Sam Lightstone, Sujay S. Parekh, Christian Garcia-Arellano |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2004 | Service level management: A dynamic discovery and optimization approachabstractOptimizing configuration parameters for achieving service level objectives is time-consuming and skills-intensive. This paper proposes a generic approach to automating this task. By generic, we mean that the approach is relatively independent of the target system for which the optimization is done. Our approach uses online adjustment of configuration parameters to discover the system's performance characteristics. Doing so creates two challenges: (1) handling interdependencies between configuration parameters and (2) minimizing the deleterious effects on production workload while the optimization is underway. Our approach addresses (1) by including in the architecture a rule-based component that handles interdependencies between configuration parameters. For (2), we use a feedback mechanism for online optimization that searches the parameter space in a way that generally avoids poor performance at intermediate steps. Our studies of a DB2 Universal Database Server under an e-commerce workload indicate that our approach is effective in practice. Yixin Diao, Frank Eskesen, Steve Froehlich, Joseph L. Hellerstein, Alexander Keller 0002, Lisa Spainhower, Maheswaran Surendra |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2004 | Generic On-Line Discovery of Quantitative ModelsabstractQuantitative models are needed for a variety of management tasks, including identification of critical variables to use for health monitoring, anticipating service-level violations by using predictive models, and ongoing optimization of configurations. Unfortunately, constructing quantitative models requires specialized skills that are in short supply. Even worse, rapid changes in provider configurations and the evolution of business demands mean that quantitative models must be updated on an ongoing basis. This paper describes an architecture and algorithms for online discovery of quantitative models without prior knowledge of the managed elements. The architecture makes use of an element schema that describes managed elements using the Common Information Model (CIM). Algorithms are presented for selecting a subset of the element metrics to use as explanatory variables in a quantitative model and for constructing the quantitative model itself. We further describe a prototype system based onthis architecture that incorporates these algorithms. We apply the prototype to online estimation of response times for DB2 Universal Database under a TPC-W workload. Of the approximately 500 metrics available from the DB2 performance monitor, our system chooses three to construct a model that explains 72 percent of the variability of response time. Alexander Keller 0002, Yixin Diao, Frank Eskesen, Steve Froehlich, Joseph L. Hellerstein, Maheswaran Surendra, Lisa Spainhower |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2003 | Generic On-Line Discovery of Quantitative Models for Service Level Management
Yixin Diao, Frank Eskesen, Steve Froehlich, Joseph L. Hellerstein, Alexander Keller 0002, Lisa Spainhower, Maheswaran Surendra |
Integrated Network Management | 1 |
| 2003 | Online Response Time Optimization of Apache Web Server
Xue (Steve) Liu, Lui Sha, Yixin Diao, Steve Froehlich, Joseph L. Hellerstein, Sujay S. Parekh |
IWQoS | 3 |
| 2002 | Using MIMO feedback control to enforce policies for interrelated metrics with application to the Apache Web serverabstractPolicy-based management provides a means for IT systems to operate according to business needs. Unfortunately, there is often an "impedance mismatch" between the policies administrators want and the controls they are given. Consider the Apache Web server. Administrators want to control CPU and memory utilizations, but this must be done indirectly by manipulating tuning parameters such as MaxClients and KeepAlive. There has been much interest in using feedback control to bridge the impedance mismatch. However, these efforts have focused on a single metric that is manipulated by a single control and hence have not considered interactions between controls such as those that are common in computing systems. This paper shows how multiple-input, multiple-output (MIMO) control theory can be used to enforce policies for interrelated metrics. MIMO is used both to model the target system, Apache in our case, and to design feedback controllers. The MIMO model captures the interactions between KA and MC, and can be used to identify infeasible metric policies. In addition, MIMO control techniques can provide considerable benefit in handling trade-offs between speed of metric convergence and sensitivity to random fluctuations while enforcing the desired policies. Yixin Diao, Neha Gandhi, Joseph L. Hellerstein, Sujay S. Parekh, Dawn M. Tilbury |
NOMS | 1 |
| 2002 | Adaptive neural/fuzzy control for interpolated nonlinear systemsabstractAdaptive control for nonlinear time-varying systems is of both theoretical and practical importance. We propose an adaptive control methodology for a class of nonlinear systems with a time-varying structure. This class of systems is composed of interpolations of nonlinear subsystems which are input-output feedback linearizable. Both indirect and direct adaptive control methods are developed, where the spatially localized models (in the form of Takagi-Sugeno fuzzy systems or radial basis function neural networks) are used as online approximators to learn the unknown dynamics of the system. Without assumptions on rate of change of system dynamics, the proposed adaptive control methods guarantee that all internal signals of the system are bounded and the tracking error is asymptotically stable. The performance of the adaptive controller is demonstrated using a jet engine control problem. Yixin Diao, Kevin M. Passino |
IEEE Trans. Fuzzy Syst. | 1 |