Genady Grabarnik

dblp:19/5056 · also Genady Ya. Grabarnik · DBLP profile ↗
← Back
45ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0001-8068-0920ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 8 · 1 first-authorDatabases, data management, data science and information retrieval · 8 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 2 · 2 first-author
YearPublicationVenuePosition
2025 Chaos Engineering Based Kubernetes Pod Rescheduling Through Deep Sets and Reinforcement Learning
abstract
Kubernetes (K8S) is a widely used orchestration solution that helps manage complex IT applications by providing mechanisms for autoscaling, health checking, cluster formation, and replication, which are essential to deploy and manage the multitude of connected microservices. However, they may suffer in case of unexpected faults which can severely change the underlying computing infrastructure and lead to service outages, highlighting the need for resilient solutions capable of mitigating the adverse effects of faults. To address this, the TELKA sched-uler integrates Chaos Engineering (CE), Reinforcement Learning (RL), and Digital Twin (DT) to reallocate K8S pods evicted due to unexpected faults. While TELKA showed promising results in reallocating evicted pods, its preliminary implementations suffered from scalability issues, as the RL agent could only effectively operate on scenarios with the same number of nodes seen during training. To overcome this limitation, this paper improves TELKA by incorporating a neural network architecture called Deep Sets (DS), which can generalize the operation of TELKA on different numbers of nodes. Experimental results not only demonstrate the validity of the improved TELKA but also show how it can be used to identify good operating conditions.
Mattia Zaccarini, Filippo Poltronieri, Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi
NOMS6
2024 TELKA: Twin-Enhanced Learning for Kubernetes Applications
abstract
Chaos engineering is the discipline of injecting computing and network faults, such as increased network latency and unavailability of computing nodes, into an IT system to help developers in identifying problems that could arise in a production environment and tackle them. Several tools have emerged to ease the application of chaos engineering to complex IT systems, leveraging microservice and container-based applications deployed on Kubernetes. However, applying of such tools requires several phases to be put into practice, from defining a steady state to establishing an effective response plan if something goes wrong. To ease the application of chaos engineering in improving the resilience of Kubernetes applications, this work presents a smart scheduler for Kubernetes called TELKA: a Twin-Enhanced Learning for Kubernetes Applications, which combines chaos engineering, Digital Twin (DT), and Reinforcement Learning (RL) methodologies to mitigate the effects of computing and network faults. Instead of interacting directly with the physical Kubernetes application, TELKA learns by interacting with a digital twin, thus reducing the learning time and the operation costs related to the application of chaos engineering. Experiment results compare TELKA with other approaches to show its effectiveness in mitigating the adverse effects of injected faults.
Mattia Zaccarini, Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Lorenzo Manca, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi
ISCC5
2024 KubeTwin: A Digital Twin Framework for Kubernetes Deployments at Scale
abstract
Kubernetes is a well-known orchestration and management solution for complex and large-scale service architectures in the Cloud Continuum. While it provides very valuable functions from the operation perspective, the high number of control loops it implements significantly enlarges the already wide space of configuration parameters and policies to consider for management purposes. We argue that optimizing complex Kubernetes deployments considering a multi-cloud and edge computing environment would significantly benefit from a Digital Twin approach, enabling an accurate virtual representation of a Kubernetes application to optimize its deployment and management policies. Towards that goal, this work illustrates the design of KubeTwin, a framework to implement Digital Twins of Kubernetes deployments. Furthermore, we present a validation of KubeTwin in a Multi-access Edge Computing (MEC) scenario, which shows its soundness in reenacting realistic Digital Twins of complex and highly distributed Kubernetes deployments. We believe that KubeTwin can provide useful guidance to the research community working in this field.
Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Lorenzo Manca, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Mattia Zaccarini
IEEE Trans. Netw. Serv. Manag.4
2023 Characterization of Microservice Response Time in Kubernetes: A Mixture Density Network Approach
abstract
The use of microservice-based applications is becoming more prominent also in the telecommunication field. The current 5G core network, for instance, is already built around the concept of a “Service Based Architecture”, and it is foreseeable that 6G will push even further this concept to enable more flexible and pervasive deployments. However, the increasing complexity of future networks calls for sophisticated platforms that could help network providers with their deployments design. In this framework, a central research trend is the development of digital twins of the physical infrastructures. These digital representations should closely mimic the behavior of the managed system, allowing the operators to test new configurations, analyze what-if scenarios, or train their reinforcement learning algorithms in safe environments. Considering that Kubernetes is becoming the de-facto standard platform for container orchestration and microservice-based application lifecycle management, the implementation of a Kubernetes digital twin requires an accurate characterization of the microservice response time, possibly leveraging suitable Machine Learning techniques trained with measurement data collected in the field. In this paper we introduce a new methodology, based on Mixture Density Networks, to accurately estimate the statistical distribution of the response time of microservice-based applications. We show the improvement in performance with respect to simulation-based inference procedures proposed in literature.
Lorenzo Manca, Davide Borsatti, Filippo Poltronieri, Mattia Zaccarini, Domenico Scotece, Gianluca Davoli, Luca Foschini 0001, Genady Grabarnik, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Walter Cerroni
CNSM8
2023 Modeling Digital Twins of Kubernetes-Based Applications
abstract
Kubernetes provides several functions that can help service providers to deal with the management of complex container-based applications. However, most of these functions need a time-consuming and costly customization process to address service-specific requirements. The adoption of Digital Twin (DT) solutions can ease the configuration process by enabling the evaluation of multiple configurations and custom policies by means of simulation-based what-if scenario analysis. To facilitate this process, this paper proposes KubeTwin, a framework to enable the definition and evaluation of DTs of Kubernetes applications. Specifically, this work presents an innovative simulation-based inference approach to define accurate DT models for a Kubernetes environment. We experimentally validate the proposed solution by implementing a DT model of an image recognition application that we tested under different conditions to verify the accuracy of the DT model. The soundness of these results demonstrates the validity of the KubeTwin approach and calls for further investigation.
Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Mattia Zaccarini
ISCC4
2022 BDMaaS+: Business-Driven and Simulation-Based Optimization of IT Services in the Hybrid Cloud
abstract
The maturity of heterogeneous and hybrid public Cloud environments enables service providers to deploy there their complex IT services trusting these large and complex infrastructures. At the same time, evaluating the impact of changes at service configuration before and at the runtime is still a very challenging and difficult task. Moreover, a comprehensive performance evaluation of IT service configurations should not be limited just to costs for IT resource acquisition, but also include risk related elements such as Service Level Agreement (SLA) violation penalties and other intangibles. To support IT service providers in this difficult task, we developed Business-Driven Management as a Service Plus (BDMaaS+), a novel decision support tool that can evaluate IT service configuration through simulation with realistic service and network models. By allowing service providers to define expanded operational parameters, BDMaaS+ also enables what-if scenario analysis, thereby opening interesting possibilities at the planning level. Experimental results, collected from our thorough evaluations, demonstrate how a service provider can leverage BDMaaS+ to explore the potential of high-level business SLA changes and data center additions.
Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Filippo Poltronieri, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi
IEEE Trans. Netw. Serv. Manag.3
2021 Detecting Causal Structure on Cloud Application Microservices Using Granger Causality Models
abstract
The loosely-coupled microservices architecture has become increasingly popular due to the advantage of its modularity and elasticity in cloud applications. However, it also seriously complicates cloud management and degrades the performance of IT operations. Today, AI has been the locus of commerce and transactions, and transforming traditional IT operations for speed and growth. Inferring the dependencies among an application's microservices can greatly help SREs diagnose possible root causes of performance issues, which is a hard task due to the complex topology of microservices is often unknown in practice. Prior literature on detecting causal structure for cloud services requires significant application instrumentation, which rarely holds in reality. In this work, we leverage Granger causality models on just monitored log data of a microservice-based application to infer the impact of dependencies between microservices. We first describe the approach of modeling discrete log data as time series, and then formally define the Granger causality problem using both linear and nonlinear autoregressive models. Finally, we conduct an extensive comparative study to show the performance of the state-of-the-art linear and nonlinear (i.e., neural) Granger causality methods on both synthetic data and real-world log data from a publicly available benchmark microservice system. Our preliminary results indicate that neural Granger causality models outperform traditional Granger causality methods on both linear and nonlinear time series data, while for large linear time series, linear Granger causal models are more efficient with high accuracy. Using the real-world log data, we also demonstrate our interesting findings on inferred dependency graph of microservices by linear and neural Granger causality models.
Qing Wang 0016, Larisa Shwartz, Genady Grabarnik, Vijay Arya, Karthikeyan Shanmugam 0001
CLOUD3
2019 Automated Partial Credit for STEM classes
abstract
This paper introduces a system that allows the award of partial credit for exams and tests consisting of multiple choice and true/false problems. Traditional multiple-choice exams have a number of issues. For one, they result in overall lower grades than exams using the same problems in an open-ended format. It is well known, that students see multiple-choice and true/false problems as not completely fair because they do not take into account partially correct answers. Non-major courses such that Math and Statistics for Engineering and Computer Science majors create additional challenges because they require simultaneous content knowledge in more than one subject. A reliable automated method to assign partial credit for multiple-choice and true/false problems can greatly improve both quality and perception of fairness of the assessment. There are two major approaches to address the issue with multiple-choice problems, see e.g. [1], [2]. In this work we suggest a combined and generalized approach for assigning partial credit for non-open-ended problems by using Concept Inventory among other testing approaches. One of the main benefits of our suggested approach is content independence and applicability to both major and non-major STEM subjects.
Genady Grabarnik, Serge Yaskolko
FIE1
2019 Leveraging AI in Service Automation Modeling: From Classical AI Through Deep Learning to Combination Models
Qing Wang 0016, Larisa Shwartz, Genady Grabarnik, Michael Nidd, Jinho Hwang
ICSOC3
2019 What-if Scenario Analysis for IT Services in Hybrid Cloud Environments with BDMaaS+
Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Larisa Shwartz, Mauro Tortonesi
IM3
2019 Online Interactive Collaborative Filtering Using Multi-Armed Bandit with Dependent Arms
abstract
Online interactive recommender systems strive to promptly suggest users appropriate items (e.g., movies and news articles) according to the current context including both user and item content information. Such contextual information is often unavailable in practice, where only the users' interaction data on items can be utilized by recommender systems. The lack of interaction records, especially for new users and items, inflames the performance of recommendation further. To address these issues, both collaborative filtering, one of the most popular recommendation techniques relying on the interaction data only, and bandit mechanisms, capable of achieving the balance between exploitation and exploration, are adopted into an online interactive recommendation setting assuming independent items (i.e., arms). This assumption rarely holds in reality, since the real-world items tend to be correlated with each other. In this paper, we study online interactive collaborative filtering problems by considering the dependencies among items. We explicitly formulate item dependencies as the clusters of arms in the bandit setting, where the arms within a single cluster share the similar latent topics. In light of topic modeling techniques, we come up with a novel generative model to generate the items from their underlying topics. Furthermore, an efficient particle-learning based online algorithm is developed for inferring both latent parameters and states of our model by taking advantage of the fully adaptive inference strategy of particle learning techniques. Additionally, our inferred model can be naturally integrated with existing multi-armed selection strategies in an interactive collaborative filtering setting. Empirical studies on two real-world applications, online recommendations on movies and news, demonstrate both the effectiveness and efficiency of our proposed approach.
Qing Wang 0016, Chunqiu Zeng, Wubai Zhou, Tao Li 0001, S. Sitharama Iyengar, Larisa Shwartz, Genady Grabarnik
IEEE Trans. Knowl. Data Eng.7
2019 An Integrated Framework for Mining Temporal Logs from Fluctuating Events
abstract
The importance of mining time lags of hidden temporal dependencies from sequential data is highlighted in many domains including system management, stock market analysis, climate monitoring, and more. Mining time lags of temporal dependencies provides useful insights into the understanding of sequential data and predicting its evolving trend. Traditional methods mainly utilize the predefined time window to analyze the sequential items, or employ statistical techniques to identify the temporal dependencies from a sequential data. However, it is a challenging task for existing methods to find the time lag of temporal dependencies in the real world, where time lags are fluctuating, noisy, and interleaved with each other. In order to identify temporal dependencies with time lags in this setting, this paper comes up with an integrated framework from both system and algorithm perspectives. Specifically, a novel parametric model is introduced to model the noisy time lags for temporal dependencies discovery between events. Based on the parametric model, an efficient expectation maximization approach is proposed for time lag discovery with maximum likelihood. Furthermore, this paper also contributes an approximation method for learning time lag to improve the scalability in terms of the number of events, without incurring significant loss of accuracy.
Chunqiu Zeng, Wubai Zhou, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
IEEE Trans. Serv. Comput.6
2018 AISTAR: An Intelligent System for Online IT Ticket Automation Recommendation
abstract
An efficient delivery of IT services for increasingly complex IT environments demands an intelligent automated solution for resolving existing and potential issues. An automation recommender system, promptly suggesting the most proper scripted resolution to an arriving IT incident ticket, would play a significant role in IT automation services. Hence, developing a comprehensive framework supporting becomes imperative for continuous improvement of automation recommendation.In this paper, we first identify the challenges of IT services followed by a discussion on AISTAR (an intelligent system for online IT ticket automation recommendation) designed and developed to provide them. Specifically, we define and formalize automation recommendation procedure as a multi-armed bandit problem with dependent arms, which is capable of achieving the optimal tradeoff between exploitation of the system for the best automation recommendation and exploration of automation execution information for future recommendation. Two novel multi-armed bandit models are proposed and integrated to handle the aforementioned challenges in IT automation services. Empirical studies on a large ticket dataset from IBM Global Services demonstrate both the effectiveness and efficiency of our intelligent integrated system. AISTAR is earmarked for Cognitive Event Automation for IBM Service delivery.
Qing Wang 0016, Chunqiu Zeng, S. Sitharama Iyengar, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
IEEE BigData6
2018 Service Placement for Hybrid Clouds Environments based on Realistic Network Measurements
Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Larisa Shwartz, Mauro Tortonesi
CNSM3
2018 Improving STEM's Calculus education using cross-countries best teaching practices
abstract
A recently completed study by the Mathematical Association of America outlined the main points for improving Calculus education for STEM students based on the experience of teaching math in the USA. In this paper, we expand this study with the best calculus teaching practices from multiple countries. Our methodology initially followed the Program for International Student Assessment (PISA) study and then was significantly modified due to the difference between school and University education. We started by comparing the practices of two countries with traditionally different education systems. We described and analyzed factual differences and similarities of content, pedagogy and socio economics, based on an adjusted to higher education international comparison system developed by Organization for Economic Cooperation and Development for high school education. We outlined our work on culturally independent comparison methods with a goal of improving or adjusting Calculus education.
Genady Grabarnik, Luiza Kim-Tyan, Serge Yaskolko
FIE1
2018 Multi-view feature selection for labeling noisy ticket data
abstract
Service providers are facing an increasingly intense competition and growing industry requirements, which dictates efficient and cost-effective service delivery. This is largely achieved by maximizing automation of IT maintenance procedures. Automation it largely depend on classification of the tickets. The automation of ticket classification requires labeled data, which is usually triaged by manually generated rules on word features. In this paper, we propose an unsupervised approach for facilitation of rule generation for labeling tickets created by event management. We first identify and remove noisy tickets with generic resolutions, and then use sparse classic canonical analysis (CCA) for feature selection to enable an efficient rule generation. Furthermore we discuss results of an extensive empirical study of ticket data that was conducted to validate the effectiveness and efficiency of our method.
Wubai Zhou, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
NOMS5
2018 Online IT Ticket Automation Recommendation Using Hierarchical Multi-armed Bandit Algorithms
abstract
The increasing complexity of IT environments urgently requires the use of analytical approaches and automated problem resolution for more efficient delivery of IT services. In this paper, we model the automation recommendation procedure of IT automation services as a contextual bandit problem with dependent arms, where the arms are in the form of hierarchies. Intuitively, different automations in IT automation services, designed to automatically solve the corresponding ticket problems, can be organized into a hierarchy by domain experts according to the types of ticket problems. We introduce a novel hierarchical multi-armed bandit algorithms leveraging the hierarchies, which can match the coarse-to-fine feature space of arms. Empirical experiments on a real large-scale ticket dataset have demonstrated substantial improvements over the conventional bandit algorithms. In addition, a case study of dealing with the cold-start problem is conducted to clearly show the merits of our proposed algorithms.
Qing Wang 0016, Tao Li 0001, S. Sitharama Iyengar, Larisa Shwartz, Genady Grabarnik
SDM5
2017 STAR: A System for Ticket Analysis and Resolution
abstract
In large scale and complex IT service environments, a problematic incident is logged as a ticket and contains the ticket summary (system status and problem description). The system administrators log the step-wise resolution description when such tickets are resolved. The repeating service events are most likely resolved by inferring similar historical tickets. With the availability of reasonably large ticket datasets, we can have an automated system to recommend the best matching resolution for a given ticket summary. In this paper, we first identify the challenges in real-world ticket analysis and develop an integrated framework to efficiently handle those challenges. The framework first quantifies the quality of ticket resolutions using a regression model built on carefully designed features. The tickets, along with their quality scores obtained from the resolution quality quantification, are then used to train a deep neural network ranking model that outputs the matching scores of ticket summary and resolution pairs. This ranking model allows us to leverage the resolution quality in historical tickets when recommending resolutions for an incoming incident ticket. In addition, the feature vectors derived from the deep neural ranking model can be effectively used in other ticket analysis tasks, such as ticket classification and clustering. The proposed framework is extensively evaluated with a large real-world dataset.
Wubai Zhou, Ramesh Baral, Qing Wang 0016, Chunqiu Zeng, Tao Li 0001, Jian Xu 0009, Zheng Liu 0001, Larisa Shwartz, Genady Grabarnik
KDD10
2017 Knowledge Guided Hierarchical Multi-Label Classification Over Ticket Data
abstract
Maximal automation of routine IT maintenance procedures is an ultimate goal of IT service management. System monitoring, an effective and reliable means for IT problem detection, generates monitoring ticket. In light of the ticket description, the underlying categories of the IT problem are determined, and the ticket is assigned to the corresponding processing teams for problem resolving. Automatic IT problem category determination acts as a critical part during the routine IT maintenance procedures. In practice, IT problem categories are naturally organized in a hierarchy by specialization. Utilizing the category hierarchy, this paper comes up with a hierarchical multi-label classification method to classify the monitoring tickets. In order to find the most effective classification, a novel contextual hierarchy (CH) loss is introduced in accordance with the problem hierarchy. Consequently, an arising optimization problem is solved by a new greedy algorithm named GLabel. Furthermore, as well as the ticket instance itself, the knowledge from the domain experts, which partially indicates some categories the given ticket may or may not belong to, can also be leveraged to guide the hierarchical multi-label classification. Accordingly, a multi-label inference with the domain expert knowledge is conducted on the basis of the given label hierarchy. The experiment demonstrates the great performance improvement by incorporating the domain knowledge during the hierarchical multi-label classification over the ticket data.
Chunqiu Zeng, Wubai Zhou, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
IEEE Trans. Netw. Serv. Manag.5
2016 Data-driven cloud-based IT services performance forecasting
abstract
Modern Cloud computing environments are rapidly evolving, leading to a growing adoption of dynamic pricing for virtual resources and of speedier deployment tools and to the emergence of hybrid Cloud scenarios. These trends suggest the opportunity to investigate a new generation of Cloud-based IT services, capable of adapting to changes in their operating conditions and deployment environment by dynamically realigning their configuration. This calls for new and more sophisticated management tools, that are capable of atomatically evaluating the performance of alternative configurations for Cloud-based IT services and of identifying the one that aligns better to the objectives defined by the business management.
Genady Grabarnik, Mauro Tortonesi, Larisa Shwartz
IEEE BigData1
2016 Resolution Recommendation for Event Tickets in Service Management
abstract
In recent years, IT service providers have rapidly achieved an automated service delivery model. Software monitoring systems are designed to actively collect and signal event occurrences and, when necessary, automatically generate incident tickets. Repeating events generate similar tickets, which in turn have a vast number of repeated problem resolutions likely to be found in earlier tickets. In this paper, we develop techniques to recommend appropriate resolution for incoming events by making use of similarities between the events and historical resolutions of similar events. Built on the traditional k-nearest neighbor algorithm (KNN), our proposed algorithms take into account false positives often generated by monitoring systems. An additional penalty is incorporated into the algorithms to control the number of misleading resolutions in the recommendation results. Moreover, as the effectiveness of the KNN heavily relies on the underlying similarity measurement, we proposed two other approaches to significantly improve our recommendation with respect to resolution relevance. One approach uses topic-level features to incorporate resolution information into the similarity measurement; the other uses metric learning to learn a more effective similarity measure. Extensive empirical evaluations on three ticket data sets demonstrate the effectiveness and efficiency of our proposed methods.
Wubai Zhou, Chunqiu Zeng, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
IEEE Trans. Netw. Serv. Manag.6
2015 Recommending ticket resolution using feature adaptation
abstract
In recent years, IT Service Providers have been rapidly introducing automation to their service delivery model. Driven by market pressure to reduce cost and maintain quality of services, they are looking for technologies that will allow rapid progress towards attainment of truly automated service delivery. Software monitoring systems are designed to actively collect and signal event occurrences and, when necessary, automatically generate incident tickets. Repeating events generate similar tickets, which in turn have a vast number of repeated problem resolutions likely to be found in earlier tickets. In our work, we develop techniques to recommend an appropriate resolution for incoming events by making use of similarities between the events and historical resolutions of similar events. The traditional KNN (K Nearest Neighbor) algorithm has been first applied to recommend resolutions for incoming tickets. Massive heterogeneous applications as well as various monitoring software are running on clients' servers to accomplish required tasks and to monitor system health via different metrics. It leads to generation of correlated tickets that have different symptom descriptions but similar resolutions. Furthermore, change of servers' environments can also induce similar situations in which ticket descriptions differ before and after change but could have similar resolutions. These correlated tickets cause performance degradation in ticket resolution recommendation. Therefore, we propose using SCL (structural corresponding learning) based feature adaptation to uncover feature mapping in different time intervals. Moreover, to put more insights into the periodic regularities existing in our ticket datasets, we apply our algorithm on tickets grouped by different time interval granularities. Extensive empirical evaluations on real-world ticket data sets demonstrate the effectiveness and efficiency of our proposed methods.
Wubai Zhou, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
CNSM4
2015 Business-driven configuration of IT services in public and hybrid clouds based on performance forecasting
abstract
Modern Cloud computing environments are rapidly evolving, leading to a growing adoption of dynamic pricing for virtual resources and of speedier deployment tools, and to the emergence of hybrid Cloud scenarios. These trends suggest the opportunity to investigate a new generation of Cloud-based IT services, capable of adapting to changes in their operating conditions and deployment environment by dynamically realigning their configuration. This calls for new and more sophisticated management tools, that are capable of evaluating the performance of alternative configurations for Cloud-based IT services and of identifying the one that aligns better to the objectives defined by the business management. This paper presents an optimization tool for Cloud-based IT services, based on queuing theoretic analysis of service workflows and ILP optimization.
Genady Grabarnik, Mauro Tortonesi, Larisa Shwartz
IM1
2015 Resolution recommendation for event tickets in service management
abstract
In recent years, IT Service Providers have been rapidly transforming to an automated service delivery model. This is due to advances in technology and driven by the unrelenting market pressure to reduce cost and maintain quality. Tremendous progress has been made to date towards attainment of truly automated service delivery; that is, the ability to deliver the same service automatically using the same process with the same quality. However, automating Incident and Problem Management continuous to be a difficult problem, particularly due to the growing complexity of IT environments. Software monitoring systems are designed to actively collect and signal event occurrances and, when necessary, automatically generate incident tickets. Repeating events generate similar tickets, which in turn have a vast number of repeated problem resolutions likely to be found in earlier tickets. In this paper we find an appropriate resolution by making use of similarities between the events and previous resolutions of similar events. Traditional KNN (K Nearest Neighbor) algorithm has been used to recommend resolutions for incoming tickets. However, the effectiveness of recommendation heavily relies on the underlying similarity measure in KNN. In this paper, we significantly improve the similarity measure used in KNN by utilizing both the event and resolution information in historical tickets via a topic-level feature extraction using the LDA (Latent Dirichlet Allocation) model. In addition, when resolution categories are available, we propose to learn a more effective similarity measure using metric learning. Extensive empirical evaluations on three ticket data sets demonstrate the effectiveness and efficiency of our proposed methods.
Wubai Zhou, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
IM5
2014 Testing performance of IT service associates under variable load
abstract
One of the challenges in IT Service Management is an assessment of an innovation benefit introduced into service delivery on performance of IT System Administrator (SA). The question is especially complex when variability of the service requests load is taken into account. We used the sequential probability ratio test to assess effects of innovation for a case when a load of the service teams is constant over time. One of the drawbacks of this test is its requirement for a constant SAs team size while the arrival patterns of service requests might require otherwise. To address the volatility of a service request load we consider the possibility of changing number of service associates. We design a performance test based on the concepts of Repeated Significance Test (RST), which allowed us to take variable number of SAs into account. In this work we show the superiority of the suggested test in comparison to the fixed sample size test (FSST) in a number of experiments required to reach the definitive conclusion. We also compare RST with SPRT when the size of the service team is constant.
Genady Grabarnik, Yefim Haim Michlin, Larisa Shwartz
NOMS1
2014 Business-driven optimization of component placement for complex services in federated Clouds
abstract
With the advent of connected services ecosystems, new generations of services and systems are being conceived, responding to the ever growing demands of the market place. In parallel, the effective adoption of the Cloud computing paradigm is becoming an essential enabler for business enterprises. With such importance placed on the services ecosystem, the design and management of services becomes a key issue both for the providers and the users. One of the main challenges for service providers lies in the complexity of the services, comprising of a multiplicity of technologies and competing and cooperating providers, which is difficult to address through current technology-centric service design approaches, in particular for the deployment infrastructure. The work described in this paper lays a foundation for business driven service design by proposing a business goals driven model of resource allocation in the Cloud. We define a goal/loss/processing cost function for resource allocation that we optimize, while taking into account the dynamic and varying nature of requests load.
Genady Grabarnik, Larisa Shwartz, Mauro Tortonesi
NOMS1
2014 Hierarchical multi-label classification over ticket data using contextual loss
abstract
Maximal automation of routine IT maintenance procedures is an ultimate goal of IT service management. System monitoring, an effective and reliable means for IT problem detection, generates monitoring tickets to be processed by system administrators. IT problems are naturally organized in a hierarchy by specialization. The problem hierarchy is used to help triage tickets to the processing team for problem resolving. In this paper, a hierarchical multi-label classification method is proposed to classify the monitoring tickets by utilizing the problem hierarchy. In order to find the most effective classification, a novel contextual hierarchy (CH) loss is introduced in accordance with the problem hierarchy. Consequently, an arising optimization problem is solved by a new greedy algorithm. An extensive empirical study over ticket data was conducted to validate the effectiveness and efficiency of our method.
Chunqiu Zeng, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
NOMS4
2013 Identifying missed monitoring alerts based on unstructured incident tickets
abstract
Automatic system monitoring is an efficient and reliable mean for problem detection in enterprise IT infrastructures. The performance of monitoring systems depends on their configurations specified by the system administrators. In dynamic and large IT environments, the IT infrastructures are frequently changed to meet various business requirements, so the configurations may not be always consistent with the updated status. Misconfigurations can lead to false positive (false alarms) and false negative (missing alerts) for the system administrators. The false negatives can cause serious system faults. This paper presents an automatic approach for discovering the false negatives from incident tickets that are created by humans. The discovered results help the system administrators correct the misconfigurations and minimize the false negatives in future. This approach applies a text classification model for analyzing the descriptions of incident tickets and identifying the corresponding system issues. The domain knowledge for describing those issues can be incorporated to assist with this model. Experiments are conducted on real system incident tickets from a large enterprise IT infrastructure. The experimental results demonstrate the effectiveness of the proposed approach.
Tao Li 0001, Larisa Shwartz, Genady Grabarnik
CNSM4
2013 Robustness of comparison sequential test for the piloting in Service Delivery
Yefim Haim Michlin, Genady Grabarnik, Larisa Shwartz, Ofer Shaham
IM2
2013 Recommending resolutions for problems identified by monitoring
Tao Li 0001, Larisa Shwartz, Genady Grabarnik
IM4
2013 An integrated framework for optimizing automatic monitoring systems in large IT infrastructures
abstract
The competitive business climate and the complexity of IT environments dictate efficient and cost-effective service delivery and support of IT services. These are largely achieved by automating routine maintenance procedures, including problem detection, determination and resolution. System monitoring provides an effective and reliable means for problem detection. Coupled with automated ticket creation, it ensures that a degradation of the vital signs, defined by acceptable thresholds or monitoring conditions, is flagged as a problem candidate and sent to supporting personnel as an incident ticket. This paper describes an integrated framework for minimizing false positive tickets and maximizing the monitoring coverage for system faults.
Tao Li 0001, Larisa Shwartz, Florian Pinel, Genady Grabarnik
KDD5
2012 SUITS: How to make a global IT service provider sustainable?
abstract
IT Service Providers are faced with the challenge to reconcile their goals in providing services efficiently and with the quality of services required while sustaining the environment. With the introduction of carbon taxation and regulations to control green gas emissions, there is a need to expand service design to incorporate heterogeneity in carbon taxes due to regional differences. The service providers have flexibility of selecting best locations for delivery of a service by taking into consideration operational cost associated with green gas emissions. They can also buy or sell carbon credits in a carbon trading market to maintain sustainable operational performance. In this paper, we introduce the Sustainable IT Services (SUITS) framework to develop models for operational performance of an IT Service Provider in a global environment with carbon taxation and carbon credit trading markets. The models are used to formulate a revenue optimization problem for the IT Service Provider in this environment. The model solutions can provide guidance for designing operational level of different IT components and for creating effective strategies for trading in carbon markets.
Parijat Dube, Genady Grabarnik, Larisa Shwartz
NOMS2
2012 Designing pilot for operational innovation in IT service delivery
abstract
Dramatic changes are taking place in the world of IT services: in who is competing with whom, in what determines competitive success, in the technologies of product and production, and in the very ways of business' approach to innovation. Operational Innovation has become a vital necessity for IT Service Providers. It affects what their employees do every day and how they do that. Because it impacts the very core of service delivery, there is an obvious risk associated with it. The benefits of operational changes are not always easy to assess and it often creates confusion. The confusion is real and the stakes for businesses are high. We argue that direct experimentation, or piloting, is both necessary and possible for introduction of operational innovations into Service Delivery. The novelty of this work is twofold: first, we propose to use design of experiments framework for IT service delivery, addressing the essential differences between delivery of IT services and manufacturing; second, we propose robust sequential design of experiments for service processes as part of the framework and demonstrate it on a sample process.
Genady Grabarnik, Yefim Haim Michlin, Larisa Shwartz
NOMS1
2012 Optimizing system monitoring configurations for non-actionable alerts
abstract
Today's competitive business climate and the complexity of IT environments dictate efficient and cost effective service delivery and support of IT services. This is largely achieved through automating of routine maintenance procedures including problem detection, determination and resolution. System monitoring provides effective and reliable means for problem detection. Coupled with automated ticket creation, it ensures that a degradation of the vital signs, defined by acceptable thresholds or monitoring conditions, is flagged as a problem candidate and sent to supporting personnel as an incident ticket. This paper describes a novel methodology and a system for minimizing non-actionable tickets while preserving all tickets which require corrective action. Our proposed method defines monitoring conditions and the optimal corresponding delay times based on an off-line analysis of historical alerts and the matching incident tickets. Potential monitoring conditions are built on a set of predictive rules which are automatically generated by a rule-based learning algorithm with coverage, confidence and rule complexity criteria. These conditions and delay times are propagated as configurations into run-time monitoring systems.
Tao Li 0001, Florian Pinel, Larisa Shwartz, Genady Grabarnik
NOMS5
2010 Automating the delivery of IT Service Continuity Management through cloud service orchestration
abstract
IT Service Continuity Management (ITSCM) delivers the recovery of IT services in the event of a disaster. ITSCM is widely perceived as an expensive challenge for enterprise-class IT operations. Cloud computing offers a model for dynamic, scalable infrastructure resource allocation on a pay-per-use basis. These attributes promise to bring cost-efficiency to ITSCM invocation and operation processes that only in the rare event of a rehearsal or an actual disaster need to allocate infrastructure resources. We propose to use the Web Service Business Process Execution Language (BPEL) in combination with Virtual Appliances to implement standardized, testable and executable ITSCM processes. The suggested solution is described and evaluated against collected data from manual recovery processes.
Markus Klems, Stefan Tai, Larisa Shwartz, Genady Grabarnik
NOMS4
2009 Towards an optimized model of incident ticket correlation
abstract
In recent years, IT service management (ITSM) has become one of the most researched areas of IT. Incident and problem management are two of the service operation processes in the IT infrastructure library (ITIL). These two processes aim to recognize, log, isolate and correct errors which occur in the environment and disrupt the delivery of services. Incident management and problem management form the basis of the tooling provided by an incident ticket systems (ITS). In an ITS system, seemingly unrelated tickets created by end users and monitoring systems can coexist and have the same root cause. The connection between failed resource and malfunctioning services is not realized automatically, but often established manually by means of human intervention. This need for human involvement reduces productivity. The introduction of automation would increase productivity and therefore reduce the cost of incident resolution. In this paper, we propose a model to correlate incident tickets based on three criteria. First, we employ a category-based correlation that relies on matching service identifiers with associated resource identifiers, using similarity rules. Secondly, we correlate the configuration items which are critical to the failed service with the earlier identified resource tickets in order to optimize the topological comparison. Finally, we augment scheduled resource data collection with constraint adaptive probing to minimize the correlation interval for temporally correlated tickets. We present experimental data in support of our proposed correlation model.
Patricia Marcu, Genady Grabarnik, Laura Z. Luan, Daniela Rosu 0001, Larisa Shwartz, Christopher Ward
Integrated Network Management2
2009 Multi-tenant solution for IT service management: A quantitative study of benefits
abstract
The very competitive business climate dictates efficient and cost effective delivery and support of IT services. Compounded by complexity of IT environments and criticality of IT to business success, IT service providers seek multi-tenant solutions to reduce operational cost and improve service quality. In this paper we consider a multi-tenant solution for IT service management and examine its critical aspects for realizing business benefits. We conduct a quantitative study of its benefits using a complexity based value assessment methodology. By regarding complexity as a substitute for potential labor cost, we estimate the business value of the multi-tenant solution before its actual deployment.
Larisa Shwartz, Yixin Diao, Genady Grabarnik
Integrated Network Management3
2009 Comparison of the Mean Time Between Failures for Two Systems Under Short Tests
abstract
A sequential probability ratio test (SPRT) is discussed, for comparison of two systems, one "basic" (b), and the other "new" (n), with exponentially distributed times between failures (TBF). The hypothesis that the mean TBFn/MTBFbges 1 is checked, versus one that it is <1. The paper deals with tests with a low Average Sample Number (ASN), having the advantage of economy in time requirement, and cost; and it is shown that the points of possible solutions in them are sparse. Criteria are proposed for assessment of the test quality, with a view to optimization of its parameters. We present a search algorithm for the truncation apex (TA), with dependences for the search domain, and for the position of the oblique test boundaries, serving jointly as the basis for our development of the test planning algorithm.
Yefim Haim Michlin, Genady Grabarnik, E. Leshchenko
IEEE Trans. Reliab.2
2008 Closed-form supervised dimensionality reduction with generalized linear models
abstract
We propose a family of supervised dimensionality reduction (SDR) algorithms that combine feature extraction (dimensionality reduction) with learning a predictive model in a unified optimization framework, using data- and class-appropriate generalized linear models (GLMs), and handling both classification and regression problems. Our approach uses simple closed-form update rules and is provably convergent. Promising empirical results are demonstrated on a variety of high-dimensional datasets.
Irina Rish, Genady Grabarnik, Guillermo A. Cecchi, Francisco Pereira 0001, Geoffrey J. Gordon
ICML2
2008 Decomposition of IT service processes and alternative service identification using ontologies
abstract
Providers of IT services are under constant pressure to reduce cost and improve the quality of the services they provide. The ability to decompose such IT services (defined by themselves or third party suppliers) into elemental service processes characterized by core function can benefit the providers by allowing them to identify alternative service processes of potentially lower cost and/or improved quality. This paper proposes an ontology-based hierarchical service decomposition and identification approach to support service providers in managing their operational service processes by the characterization and exploitation of such elemental service processes. We further demonstrate the feasibility and sample outcomes of the proposed approach an example.
Christian Bartsch 0002, Larisa Shwartz, Christopher Ward, Genady Grabarnik, Melissa J. Buco
NOMS4
2007 Sequential Testing for Comparison of the Mean Time Between Failures for Two Systems
abstract
This study deals with simultaneous testing of two systems, one "basic" (subscript b), and the other "new" (n), both with an exponential distribution describing the times between failures. We test whether the mean TBFn/MTBFbratio equals a given value, versus whether it is smaller than the given value. These tests yield a binomial pattern. A recursive algorithm calculates the probability of a given combination of failure numbers in the systems, permitting rapid, accurate determination of the test characteristics. The influence of truncation of Wald's Sequential Probability Ratio Test (SPRT) on its characteristics is analysed, and relationships are derived for calculating the coordinates of truncation apex (TA). A test planning methodology is presented for the most common cases
Yefim Haim Michlin, Genady Grabarnik
IEEE Trans. Reliab.2
2005 Adaptive diagnosis in distributed systems
abstract
Real-time problem diagnosis in large distributed computer systems and networks is a challenging task that requires fast and accurate inferences from potentially huge data volumes. In this paper, we propose a cost-efficient, adaptive diagnostic technique called active probing. Probes are end-to-end test transactions that collect information about the performance of a distributed system. Active probing uses probabilistic reasoning techniques combined with information-theoretic approach, and allows a fast online inference about the current system state via active selection of only a small number of most-informative tests. We demonstrate empirically that the active probing scheme greatly reduces both the number of probes (from 60% to 75% in most of our real-life applications), and the time needed for localizing the problem when compared with nonadaptive (preplanned) probing schemes. We also provide some theoretical results on the complexity of probe selection, and the effect of "noisy" probes on the accuracy of diagnosis. Finally, we discuss how to model the system's dynamics using dynamic Bayesian networks (DBNs), and an efficient approximate approach called sequential multifault; empirical results demonstrate clear advantage of such approaches over "static" techniques that do not handle system's changes.
Irina Rish, Mark Brodie, Sheng Ma, Natalia Odintsova, Alina Beygelzimer, Genady Grabarnik, Karina Hernandez
IEEE Trans. Neural Networks6
2004 Real-time problem determination in distributed systems using active probing
abstract
We describe algorithms and an architecture for a real-time problem determination system that uses online selection of most-informative measurements - the approach called herein active probing. Probes are end-to-end test transactions which gather information about system components. Active probing allows probes to be selected and sent on-demand, in response to one's belief about the state of the system. At each step the most informative next probe is computed and sent. As probe results are received, belief about the system state is updated using probabilistic inference. This process continues until the problem is diagnosed. We demonstrate through both analysis and simulation that the active probing scheme greatly reduces both the number of probes and the time needed for localizing the problem when compared with non-active probing schemes.
Irina Rish, Mark Brodie, Natalia Odintsova, Sheng Ma, Genady Grabarnik
NOMS (1)5
2003 Data-driven validation, completion and construction of event relationship networks
abstract
Event management is a focal point in building and maintaining high quality information infrastructures. We have witnessed the shift of the paradigm of event management in practice from root cause analysis (RCA) to action-oriented analysis (AOA). IBM has developed a pioneer event management methodology (EMD) based on the AOA paradigm and applied it to more than two hundred production sites with success. Foreseeably, more and more event management professionals will apply AOA in different incarnations in building proactive management facilities. By that, building correct and effective Event Relationship Networks (ERNs) becomes the dominating activity in AOA service design process. Currently, the quality of ERNs and the cost of building them largely depend on the knowledge of domain experts. We believe that we can utilize historical event logs in shortening the ERNs design process and perfecting the quality of ERNs. In this paper, we describe in detail how to apply this data-driven approach in ERN validation, completion and construction.
Chang-Shing Perng, David Thoenen, Genady Grabarnik, Sheng Ma, Joseph L. Hellerstein
KDD3
2002 Progressive and Interactive Analysis of Event Data Using Event Miner
abstract
Exploring large data sets typically involves activities that iterate between data selection and data analysis, in which insights obtained from analysis result in new data selection. Further, data analysis needs to use a combination of analysis techniques: data summarization, mining algorithms and visualization. This interweaving of functions arises both from the semantics of what the analyst hopes to achieve and from scalability requirements for dealing with large data volumes. We refer to such a process as a progressive analysis. Herein is described a tool, Event Miner, that integrates data selection, mining and visualization for progressive analysis of temporal, categorical data. We discuss a data model and architecture. We illustrate how our tool can be used for complex mining tasks such as finding patterns not occurring on Monday. Further, we discuss the novel visualization employed, such as visualizing categorical data and the results of data mining. Also, we discuss the extension of the existing mining framework needed to mine temporal events with multiple attributes. Throughout, we illustrate the capabilities of Event Miner by applying it to event data from large computer networks.
Sheng Ma, Joseph L. Hellerstein, Chang-Shing Perng, Genady Grabarnik
ICDM4