Larisa Shwartz

dblp:32/6257 · also Laura Shwartz · DBLP profile ↗
← Back
65ranked-venue papers
2as first author
18since 2021 · last 2025
0000-0001-5878-0765ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 18 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 4 since 2021Databases, data management, data science and information retrieval · 10 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 since 2021Software engineering, systems software and programming languages · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CPU-Limits kill Performance: Time to rethink Resource Control
Chirag C. Shetty, Sarthak Chakraborty, Hubertus Franke, Larisa Shwartz, Chandrasekhar Narayanaswami 0001, Indranil Gupta, Saurabh Jha
SoCC4
2025 ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks
abstract
Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a framework that offers a systematic methodology for benchmarking AI agents to address real-world IT automation tasks. Our initial release targets three key areas: Site Reliability Engineering (SRE), Compliance and Security Operations (CISO), and Financial Operations (FinOps). The design enables AI researchers to understand the challenges and opportunities of AI agents for IT automation with push-button workflows and interpretable metrics. IT-Bench includes an initial set of 102 real-world scenarios, which can be easily extended by community contributions. Our results show that agents powered by state-of-the-art models resolve only 11.4% of SRE scenarios, 25.2% of CISO scenarios, and 25.8% of FinOps scenarios (excluding anomaly detection). For FinOps-specific anomaly detection (AD) scenarios, AI agents achieve an F1 score of 0.35. We expect ITBench to be a key enabler of AI-driven IT automation that is correct, safe, and fast. IT-Bench, along with a leaderboard and sample agent implementations, is available at https://github.com/ibm/itbench.
Saurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa, Yinfang Chen, Jackson Clark, Bhavya, Mudit Verma, Hirokuni Kitahara, Noah Zheutlin, Saki Takano, Divya Pathak, Felix George, Xinbo Wu, Bekir O. Turkkan, Gerard Vanloo, Michael Nidd, Oishik Chatterjee, Pranjal Gupta, Suranjana Samanta, Pooja Aggarwal, Rong Lee, Jae-wook Ahn, Debanjana Kar, Amit M. Paradkar, Yu Deng 0004, Pratibha Moogi, Prateeti Mohapatra, Naoki Abe, Chandrasekhar Narayanaswami 0001, Tianyin Xu, Lav R. Varshney, Ruchi Mahindru, Anca Sailer, Larisa Shwartz, Daby M. Sow, Nicholas C. Fuller, Ruchir Puri
ICML37
2025 Chaos Engineering Based Kubernetes Pod Rescheduling Through Deep Sets and Reinforcement Learning
abstract
Kubernetes (K8S) is a widely used orchestration solution that helps manage complex IT applications by providing mechanisms for autoscaling, health checking, cluster formation, and replication, which are essential to deploy and manage the multitude of connected microservices. However, they may suffer in case of unexpected faults which can severely change the underlying computing infrastructure and lead to service outages, highlighting the need for resilient solutions capable of mitigating the adverse effects of faults. To address this, the TELKA sched-uler integrates Chaos Engineering (CE), Reinforcement Learning (RL), and Digital Twin (DT) to reallocate K8S pods evicted due to unexpected faults. While TELKA showed promising results in reallocating evicted pods, its preliminary implementations suffered from scalability issues, as the RL agent could only effectively operate on scenarios with the same number of nodes seen during training. To overcome this limitation, this paper improves TELKA by incorporating a neural network architecture called Deep Sets (DS), which can generalize the operation of TELKA on different numbers of nodes. Experimental results not only demonstrate the validity of the improved TELKA but also show how it can be used to identify good operating conditions.
Mattia Zaccarini, Filippo Poltronieri, Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi
NOMS8
2024 SAM: Subseries Augmentation-Based Meta-Learning for Generalizing AIOps Models in Multi-Cloud Migration
abstract
In the context of cloud computing, enterprises are increasingly adopting multi-cloud strategies to enhance performance, ensure cost efficiency, and avoid vendor lock-in. This trend presents a significant challenge for the migration of AI for IT operations (AIOps) models across different cloud providers due to variations in architecture, performance, and data distribution. Traditional methods of re-training AIOps models for new cloud environments are labor-intensive and delay deployment. To address this issue, we introduce a novel framework called SAM (Subseries Augmentation-based Meta-learning), which facilitates seamless model migration between clouds without the need for re-training from scratch. SAM leverages data augmentation and meta-learning to efficiently adapt AIOps models to new cloud environments. It has proven effective in adapting anomaly detectors across various config-urations over both public and simulated datasets. We believe that SAM can also be adapted to other AI models used for automating IT tasks such as alerting and resource scaling.
Paulito Palmes, Saurabh Jha, Bekir O. Turkkan, Gerard Vanloo, Frank Bagehorn, Chandrasekhar Narayanaswami 0001, Larisa Shwartz, Naoki Abe, Yu Deng 0004, Daby M. Sow
CLOUD8
2024 Seed-Guided Fine-Grained Entity Typing in Science and Engineering Domains
abstract
Accurately typing entity mentions from text segments is a fundamental task for various natural language processing applications. Many previous approaches rely on massive human-annotated data to perform entity typing. Nevertheless, collecting such data in highly specialized science and engineering domains (e.g., software engineering and security) can be time-consuming and costly, without mentioning the domain gaps between training and inference data if the model needs to be applied to confidential datasets. In this paper, we study the task of seed-guided fine-grained entity typing in science and engineering domains, which takes the name and a few seed entities for each entity type as the only supervision and aims to classify new entity mentions into both seen and unseen types (i.e., those without seed entities). To solve this problem, we propose SEType which first enriches the weak supervision by finding more entities for each seen type from an unlabeled corpus using the contextualized representations of pre-trained language models. It then matches the enriched entities to unlabeled text to get pseudo-labeled samples and trains a textual entailment model that can make inferences for both seen and unseen types. Extensive experiments on two datasets covering four domains demonstrate the effectiveness of SEType in comparison with various baselines. Code and data are available at: https://github.com/yuzhimanhua/SEType.
Yu Zhang 0044, Yunyi Zhang 0001, Yanzhen Shen, Yu Deng 0004, Lucian Popa 0001, Larisa Shwartz, ChengXiang Zhai, Jiawei Han 0001
AAAI6
2024 TELKA: Twin-Enhanced Learning for Kubernetes Applications
abstract
Chaos engineering is the discipline of injecting computing and network faults, such as increased network latency and unavailability of computing nodes, into an IT system to help developers in identifying problems that could arise in a production environment and tackle them. Several tools have emerged to ease the application of chaos engineering to complex IT systems, leveraging microservice and container-based applications deployed on Kubernetes. However, applying of such tools requires several phases to be put into practice, from defining a steady state to establishing an effective response plan if something goes wrong. To ease the application of chaos engineering in improving the resilience of Kubernetes applications, this work presents a smart scheduler for Kubernetes called TELKA: a Twin-Enhanced Learning for Kubernetes Applications, which combines chaos engineering, Digital Twin (DT), and Reinforcement Learning (RL) methodologies to mitigate the effects of computing and network faults. Instead of interacting directly with the physical Kubernetes application, TELKA learns by interacting with a digital twin, thus reducing the learning time and the operation costs related to the application of chaos engineering. Experiment results compare TELKA with other approaches to show its effectiveness in mitigating the adverse effects of injected faults.
Mattia Zaccarini, Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Lorenzo Manca, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi
ISCC9
2024 TraceWeaver: Distributed Request Tracing for Microservices Without Application Modification
abstract
Monitoring and debugging modern cloud-based applications is challenging since even a single API call can involve many interdependent distributed microservices. To provide observability for such complex systems, distributed tracing frameworks track request flow across the microservice call tree. However, such solutions require instrumenting every component of the distributed application to add and propagate tracing headers, which has slowed adoption. This paper explores whether we can trace requests without any application instrumentation, which we refer to as request trace reconstruction. To that end, we develop TraceWeaver, a system that incorporates readily available information from production settings (e.g., timestamps) and test environments (e.g., call graphs) to reconstruct request traces with usefully high accuracy. At the heart of TraceWeaver is a reconstruction algorithm that uses request-response timestamps to effectively prune the search space for mapping requests and applies statistical timing analysis techniques to reconstruct traces. Evaluation with (1) benchmark microservice applications and (2) a production microservice dataset demonstrates that TraceWeaver can achieve a high accuracy of ~90% and can be meaningfully applied towards multiple use cases (e.g., finding slow services and A/B testing).
Sachin Ashok, Vipul Harsh, Brighten Godfrey, Radhika Mittal, Srinivasan Parthasarathy 0002, Larisa Shwartz
SIGCOMM6
2024 KubeTwin: A Digital Twin Framework for Kubernetes Deployments at Scale
abstract
Kubernetes is a well-known orchestration and management solution for complex and large-scale service architectures in the Cloud Continuum. While it provides very valuable functions from the operation perspective, the high number of control loops it implements significantly enlarges the already wide space of configuration parameters and policies to consider for management purposes. We argue that optimizing complex Kubernetes deployments considering a multi-cloud and edge computing environment would significantly benefit from a Digital Twin approach, enabling an accurate virtual representation of a Kubernetes application to optimize its deployment and management policies. Towards that goal, this work illustrates the design of KubeTwin, a framework to implement Digital Twins of Kubernetes deployments. Furthermore, we present a validation of KubeTwin in a Multi-access Edge Computing (MEC) scenario, which shows its soundness in reenacting realistic Digital Twins of complex and highly distributed Kubernetes deployments. We believe that KubeTwin can provide useful guidance to the research community working in this field.
Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Lorenzo Manca, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Mattia Zaccarini
IEEE Trans. Netw. Serv. Manag.8
2023 Fault Injection Based Interventional Causal Learning for Distributed Applications
abstract
We apply the machinery of interventional causal learning with programmable interventions to the domain of applications management. Modern applications are modularized into interdependent components or services (e.g. microservices) for ease of development and management. The communication graph among such components is a function of application code and is not always known to the platform provider. In our solution we learn this unknown communication graph solely using application logs observed during the execution of the application by using fault injections in a staging environment. Specifically, we have developed an active (or interventional) causal learning algorithm that uses the observations obtained during fault injections to learn a model of error propagation in the communication among the components. The “power of intervention” additionally allows us to address the presence of confounders in unobserved user interactions. We demonstrate the effectiveness of our solution in learning the communication graph of well-known microservice application benchmarks. We also show the efficacy of the solution on a downstream task of fault localization in which the learned graph indeed helps to localize faults at runtime in a production environment (in which the location of the fault is unknown). Additionally, we briefly discuss the implementation and deployment status of a fault injection framework which incorporates the developed technology.
Qing Wang 0016, Jesus Rios, Saurabh Jha, Karthikeyan Shanmugam 0001, Frank Bagehorn, Robert Filepp, Naoki Abe, Larisa Shwartz
AAAI9
2023 Characterization of Microservice Response Time in Kubernetes: A Mixture Density Network Approach
abstract
The use of microservice-based applications is becoming more prominent also in the telecommunication field. The current 5G core network, for instance, is already built around the concept of a “Service Based Architecture”, and it is foreseeable that 6G will push even further this concept to enable more flexible and pervasive deployments. However, the increasing complexity of future networks calls for sophisticated platforms that could help network providers with their deployments design. In this framework, a central research trend is the development of digital twins of the physical infrastructures. These digital representations should closely mimic the behavior of the managed system, allowing the operators to test new configurations, analyze what-if scenarios, or train their reinforcement learning algorithms in safe environments. Considering that Kubernetes is becoming the de-facto standard platform for container orchestration and microservice-based application lifecycle management, the implementation of a Kubernetes digital twin requires an accurate characterization of the microservice response time, possibly leveraging suitable Machine Learning techniques trained with measurement data collected in the field. In this paper we introduce a new methodology, based on Mixture Density Networks, to accurately estimate the statistical distribution of the response time of microservice-based applications. We show the improvement in performance with respect to simulation-based inference procedures proposed in literature.
Lorenzo Manca, Davide Borsatti, Filippo Poltronieri, Mattia Zaccarini, Domenico Scotece, Gianluca Davoli, Luca Foschini 0001, Genady Grabarnik, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Walter Cerroni
CNSM9
2023 Modeling Digital Twins of Kubernetes-Based Applications
abstract
Kubernetes provides several functions that can help service providers to deal with the management of complex container-based applications. However, most of these functions need a time-consuming and costly customization process to address service-specific requirements. The adoption of Digital Twin (DT) solutions can ease the configuration process by enabling the evaluation of multiple configurations and custom policies by means of simulation-based what-if scenario analysis. To facilitate this process, this paper proposes KubeTwin, a framework to enable the definition and evaluation of DTs of Kubernetes applications. Specifically, this work presents an innovative simulation-based inference approach to define accurate DT models for a Kubernetes environment. We experimentally validate the proposed solution by implementing a DT model of an image recognition application that we tested under different conditions to verify the accuracy of the DT model. The soundness of these results demonstrates the validity of the KubeTwin approach and calls for further investigation.
Davide Borsatti, Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Filippo Poltronieri, Domenico Scotece, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi, Mattia Zaccarini
ISCC7
2022 Localizing and Explaining Faults in Microservices Using Distributed Tracing
abstract
Finding the exact location of a fault in a large distributed microservices application running in containerized cloud environments can be very difficult and time-consuming. We present a novel approach that uses distributed tracing to automatically detect, localize and aid in explaining application-level faults. We demonstrate the effectiveness of our proposed approach by injecting faults into a well-known microservice-based benchmark application. Our experiments demonstrated that the proposed fault localization algorithm correctly detects and localize the microservice with the injected fault. We also compare our approach with other fault localization methods. In particular, we empirically show that our method outperforms methods in which a graph model of error propagation is used for inferring fault locations using error logs. Our work illustrates the value added by distributed tracing for localizing and explaining faults in microservices.
Jesus Rios, Saurabh Jha, Larisa Shwartz
CLOUD3
2022 Improving Model Performance Using Metric-Guided Data Selection Framework
abstract
The noisiness and low quality of IT operations management data is a major challenge in using machine learning to assist IT operations management. Our system mitigates this challenge by automatically measuring data quality, and then using the results to select data subsets that generate improved model performance. Based on a set of metrics that quantify the quality of a corpus with both structured and unstructured data, we are proposing a framework to automatically identify "well behaved" subsets in the corpus. By streaming input data to separate models for these subsets, we can achieve better performance when compared with a model trained on the full dataset. We present a motivating example that inspired our approach as well as a deployment case study of our system based on engagements with two clients which demonstrate that the proposed methodology is effective for detecting such subsets to improve model performance.
Paulina Toro Isaza, Yu Deng 0004, Michael Nidd, Amar Prakash Azad, Larisa Shwartz
IEEE Big Data5
2022 A fault injection platform for learning AIOps models
abstract
In today’s IT environment with a growing number of costly outages, increasing complexity of the systems, and availability of massive operational data, there is a strengthening demand to effectively leverage Artificial Intelligence and Machine Learning (AI/ML) towards enhanced resiliency. In this paper, we present an automatic fault injection platform to enable and optimize the generation of data needed for building AI/ML models to support modern IT operations. The merits of our platform include the ease of use, the possibility to orchestrate complex fault scenarios and to optimize the data generation for the modeling task at hand. Specifically, we designed a fault injection service that (i) combines fault injection with data collection in a unified framework, (ii) supports hybrid and multi-cloud environments, and (iii) does not require programming skills for its use. Our current implementation covers the most common fault types both at the application and infrastructure levels. The platform also includes some AI capabilities. In particular, we demonstrate the interventional causal learning capability currently available in our platform. We show how our system is able to learn a model of error propagation in a micro-service application in a cloud environment (when the communication graph among micro-services is unknown and only logs are available) for use in subsequent applications such as fault localization.
Frank Bagehorn, Jesus Rios, Saurabh Jha, Robert Filepp, Larisa Shwartz, Naoki Abe
ASE5
2022 BDMaaS+: Business-Driven and Simulation-Based Optimization of IT Services in the Hybrid Cloud
abstract
The maturity of heterogeneous and hybrid public Cloud environments enables service providers to deploy there their complex IT services trusting these large and complex infrastructures. At the same time, evaluating the impact of changes at service configuration before and at the runtime is still a very challenging and difficult task. Moreover, a comprehensive performance evaluation of IT service configurations should not be limited just to costs for IT resource acquisition, but also include risk related elements such as Service Level Agreement (SLA) violation penalties and other intangibles. To support IT service providers in this difficult task, we developed Business-Driven Management as a Service Plus (BDMaaS+), a novel decision support tool that can evaluate IT service configuration through simulation with realistic service and network models. By allowing service providers to define expanded operational parameters, BDMaaS+ also enables what-if scenario analysis, thereby opening interesting possibilities at the planning level. Experimental results, collected from our thorough evaluations, demonstrate how a service provider can leverage BDMaaS+ to explore the potential of high-level business SLA changes and data center additions.
Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Filippo Poltronieri, Larisa Shwartz, Cesare Stefanelli, Mauro Tortonesi
IEEE Trans. Netw. Serv. Manag.5
2021 Causal Modeling based Fault Localization in Cloud Systems using Golden Signals
abstract
In cloud-native applications, a large fraction of operational failures, known as outages, result in violations of Service Level Objectives (SLOs). SLOs are defined around specific measurable characteristics: availability, throughput, frequency, response time, and quality. Four metrics, latency, traffic, errors, and saturation, ensure coverage for most outages of an application. These are often called golden signals. The dynamicity and complexity of cloud-native applications complicate Site Reliability Engineers’ (SREs) efforts in problem determination, in particular in its fault localization. The fault localization is often a try-and-error process in which SREs rely on their domain knowledge and experience. It is laborious and frequently results in long Mean Time To Resolution (MTTR) for outages. This paper describes a lightweight fault localization system, that establishes causal relationships among the golden signal service errors and error logs, and further leverages PageRank centrality of the derived causal graph for generating a ranked list of faulty microservices.
Pooja Aggarwal, Seema Nagar, Larisa Shwartz, Prateeti Mohapatra, Qing Wang 0016, Amit M. Paradkar, Atri Mandal
CLOUD4
2021 A system for proactive risk assessment of application changes in cloud operations
abstract
Change is one of the biggest contributors to service outages. With more enterprises migrating their applications to cloud and using automated build and deployment the volume and rate of changes has significantly increased. Furthermore, microservice-based architectures have reduced the turnaround time for changes and increased the dependency between services. All of the above make it impossible for the Site Reliability Engineers (SREs) to use the traditional methods of manual risk assessment for changes. In order to mitigate change-induced service failures and ensure continuous improvement for cloud native services, it is critical to have an automated system for assessing the risk of change deployments. In this paper, we present an AI-based system for proactively assessing the risk associated with deployment of application changes in cloud operations. The risk assessment is accompanied with actionable risk explainability. We discuss the usage of this system in two primary scenarios of automated and manual deployment. In automated deployment scenario, our approach is able to alert SREs on 70 % of problematic changes by blocking only 1.5 % of total changes and recommending human intervention. In manual deployment scenario, our approach recommends the SREs to perform extra due diligence for 2.8 % of total changes to capture 84 % of problematic changes.
Raghav Batta, Larisa Shwartz, Michael Nidd, Amar Prakash Azad
CLOUD2
2021 Detecting Causal Structure on Cloud Application Microservices Using Granger Causality Models
abstract
The loosely-coupled microservices architecture has become increasingly popular due to the advantage of its modularity and elasticity in cloud applications. However, it also seriously complicates cloud management and degrades the performance of IT operations. Today, AI has been the locus of commerce and transactions, and transforming traditional IT operations for speed and growth. Inferring the dependencies among an application's microservices can greatly help SREs diagnose possible root causes of performance issues, which is a hard task due to the complex topology of microservices is often unknown in practice. Prior literature on detecting causal structure for cloud services requires significant application instrumentation, which rarely holds in reality. In this work, we leverage Granger causality models on just monitored log data of a microservice-based application to infer the impact of dependencies between microservices. We first describe the approach of modeling discrete log data as time series, and then formally define the Granger causality problem using both linear and nonlinear autoregressive models. Finally, we conduct an extensive comparative study to show the performance of the state-of-the-art linear and nonlinear (i.e., neural) Granger causality methods on both synthetic data and real-world log data from a publicly available benchmark microservice system. Our preliminary results indicate that neural Granger causality models outperform traditional Granger causality methods on both linear and nonlinear time series data, while for large linear time series, linear Granger causal models are more efficient with high accuracy. Using the real-world log data, we also demonstrate our interesting findings on inferred dependency graph of microservices by linear and neural Granger causality models.
Qing Wang 0016, Larisa Shwartz, Genady Grabarnik, Vijay Arya, Karthikeyan Shanmugam 0001
CLOUD2
2019 Towards Automated Planning for Enterprise Services: Opportunities and Challenges
Maja Vukovic, Scott N. Gerard, Richard Hull 0001, Michael Katz 0001, Larisa Shwartz, Shirin Sohrabi, Christian J. Muise, John J. Rofrano, Anup K. Kalia, Jinho Hwang, Yabin Dang, Zhuoxuan Jiang
ICSOC5
2019 Leveraging AI in Service Automation Modeling: From Classical AI Through Deep Learning to Combination Models
Qing Wang 0016, Larisa Shwartz, Genady Grabarnik, Michael Nidd, Jinho Hwang
ICSOC2
2019 What-if Scenario Analysis for IT Services in Hybrid Cloud Environments with BDMaaS+
Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Larisa Shwartz, Mauro Tortonesi
IM4
2019 Online Interactive Collaborative Filtering Using Multi-Armed Bandit with Dependent Arms
abstract
Online interactive recommender systems strive to promptly suggest users appropriate items (e.g., movies and news articles) according to the current context including both user and item content information. Such contextual information is often unavailable in practice, where only the users' interaction data on items can be utilized by recommender systems. The lack of interaction records, especially for new users and items, inflames the performance of recommendation further. To address these issues, both collaborative filtering, one of the most popular recommendation techniques relying on the interaction data only, and bandit mechanisms, capable of achieving the balance between exploitation and exploration, are adopted into an online interactive recommendation setting assuming independent items (i.e., arms). This assumption rarely holds in reality, since the real-world items tend to be correlated with each other. In this paper, we study online interactive collaborative filtering problems by considering the dependencies among items. We explicitly formulate item dependencies as the clusters of arms in the bandit setting, where the arms within a single cluster share the similar latent topics. In light of topic modeling techniques, we come up with a novel generative model to generate the items from their underlying topics. Furthermore, an efficient particle-learning based online algorithm is developed for inferring both latent parameters and states of our model by taking advantage of the fully adaptive inference strategy of particle learning techniques. Additionally, our inferred model can be naturally integrated with existing multi-armed selection strategies in an interactive collaborative filtering setting. Empirical studies on two real-world applications, online recommendations on movies and news, demonstrate both the effectiveness and efficiency of our proposed approach.
Qing Wang 0016, Chunqiu Zeng, Wubai Zhou, Tao Li 0001, S. Sitharama Iyengar, Larisa Shwartz, Genady Grabarnik
IEEE Trans. Knowl. Data Eng.6
2019 An Integrated Framework for Mining Temporal Logs from Fluctuating Events
abstract
The importance of mining time lags of hidden temporal dependencies from sequential data is highlighted in many domains including system management, stock market analysis, climate monitoring, and more. Mining time lags of temporal dependencies provides useful insights into the understanding of sequential data and predicting its evolving trend. Traditional methods mainly utilize the predefined time window to analyze the sequential items, or employ statistical techniques to identify the temporal dependencies from a sequential data. However, it is a challenging task for existing methods to find the time lag of temporal dependencies in the real world, where time lags are fluctuating, noisy, and interleaved with each other. In order to identify temporal dependencies with time lags in this setting, this paper comes up with an integrated framework from both system and algorithm perspectives. Specifically, a novel parametric model is introduced to model the noisy time lags for temporal dependencies discovery between events. Based on the parametric model, an efficient expectation maximization approach is proposed for time lag discovery with maximum likelihood. Furthermore, this paper also contributes an approximation method for learning time lag to improve the scalability in terms of the number of events, without incurring significant loss of accuracy.
Chunqiu Zeng, Wubai Zhou, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
IEEE Trans. Serv. Comput.5
2018 AISTAR: An Intelligent System for Online IT Ticket Automation Recommendation
abstract
An efficient delivery of IT services for increasingly complex IT environments demands an intelligent automated solution for resolving existing and potential issues. An automation recommender system, promptly suggesting the most proper scripted resolution to an arriving IT incident ticket, would play a significant role in IT automation services. Hence, developing a comprehensive framework supporting becomes imperative for continuous improvement of automation recommendation.In this paper, we first identify the challenges of IT services followed by a discussion on AISTAR (an intelligent system for online IT ticket automation recommendation) designed and developed to provide them. Specifically, we define and formalize automation recommendation procedure as a multi-armed bandit problem with dependent arms, which is capable of achieving the optimal tradeoff between exploitation of the system for the best automation recommendation and exploration of automation execution information for future recommendation. Two novel multi-armed bandit models are proposed and integrated to handle the aforementioned challenges in IT automation services. Empirical studies on a large ticket dataset from IBM Global Services demonstrate both the effectiveness and efficiency of our intelligent integrated system. AISTAR is earmarked for Cognitive Event Automation for IBM Service delivery.
Qing Wang 0016, Chunqiu Zeng, S. Sitharama Iyengar, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
IEEE BigData5
2018 Service Placement for Hybrid Clouds Environments based on Realistic Network Measurements
Walter Cerroni, Luca Foschini 0001, Genady Grabarnik, Larisa Shwartz, Mauro Tortonesi
CNSM4
2018 Multi-view feature selection for labeling noisy ticket data
abstract
Service providers are facing an increasingly intense competition and growing industry requirements, which dictates efficient and cost-effective service delivery. This is largely achieved by maximizing automation of IT maintenance procedures. Automation it largely depend on classification of the tickets. The automation of ticket classification requires labeled data, which is usually triaged by manually generated rules on word features. In this paper, we propose an unsupervised approach for facilitation of rule generation for labeling tickets created by event management. We first identify and remove noisy tickets with generic resolutions, and then use sparse classic canonical analysis (CCA) for feature selection to enable an efficient rule generation. Furthermore we discuss results of an extensive empirical study of ticket data that was conducted to validate the effectiveness and efficiency of our method.
Wubai Zhou, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
NOMS4
2018 Online IT Ticket Automation Recommendation Using Hierarchical Multi-armed Bandit Algorithms
abstract
The increasing complexity of IT environments urgently requires the use of analytical approaches and automated problem resolution for more efficient delivery of IT services. In this paper, we model the automation recommendation procedure of IT automation services as a contextual bandit problem with dependent arms, where the arms are in the form of hierarchies. Intuitively, different automations in IT automation services, designed to automatically solve the corresponding ticket problems, can be organized into a hierarchy by domain experts according to the types of ticket problems. We introduce a novel hierarchical multi-armed bandit algorithms leveraging the hierarchies, which can match the coarse-to-fine feature space of arms. Empirical experiments on a real large-scale ticket dataset have demonstrated substantial improvements over the conventional bandit algorithms. In addition, a case study of dealing with the cold-start problem is conducted to clearly show the merits of our proposed algorithms.
Qing Wang 0016, Tao Li 0001, S. Sitharama Iyengar, Larisa Shwartz, Genady Grabarnik
SDM4
2017 STAR: A System for Ticket Analysis and Resolution
abstract
In large scale and complex IT service environments, a problematic incident is logged as a ticket and contains the ticket summary (system status and problem description). The system administrators log the step-wise resolution description when such tickets are resolved. The repeating service events are most likely resolved by inferring similar historical tickets. With the availability of reasonably large ticket datasets, we can have an automated system to recommend the best matching resolution for a given ticket summary. In this paper, we first identify the challenges in real-world ticket analysis and develop an integrated framework to efficiently handle those challenges. The framework first quantifies the quality of ticket resolutions using a regression model built on carefully designed features. The tickets, along with their quality scores obtained from the resolution quality quantification, are then used to train a deep neural network ranking model that outputs the matching scores of ticket summary and resolution pairs. This ranking model allows us to leverage the resolution quality in historical tickets when recommending resolutions for an incoming incident ticket. In addition, the feature vectors derived from the deep neural ranking model can be effectively used in other ticket analysis tasks, such as ticket classification and clustering. The proposed framework is extensively evaluated with a large real-world dataset.
Wubai Zhou, Ramesh Baral, Qing Wang 0016, Chunqiu Zeng, Tao Li 0001, Jian Xu 0009, Zheng Liu 0001, Larisa Shwartz, Genady Grabarnik
KDD9
2017 Knowledge Guided Hierarchical Multi-Label Classification Over Ticket Data
abstract
Maximal automation of routine IT maintenance procedures is an ultimate goal of IT service management. System monitoring, an effective and reliable means for IT problem detection, generates monitoring ticket. In light of the ticket description, the underlying categories of the IT problem are determined, and the ticket is assigned to the corresponding processing teams for problem resolving. Automatic IT problem category determination acts as a critical part during the routine IT maintenance procedures. In practice, IT problem categories are naturally organized in a hierarchy by specialization. Utilizing the category hierarchy, this paper comes up with a hierarchical multi-label classification method to classify the monitoring tickets. In order to find the most effective classification, a novel contextual hierarchy (CH) loss is introduced in accordance with the problem hierarchy. Consequently, an arising optimization problem is solved by a new greedy algorithm named GLabel. Furthermore, as well as the ticket instance itself, the knowledge from the domain experts, which partially indicates some categories the given ticket may or may not belong to, can also be leveraged to guide the hierarchical multi-label classification. Accordingly, a multi-label inference with the domain expert knowledge is conducted on the basis of the given label hierarchy. The experiment demonstrates the great performance improvement by incorporating the domain knowledge during the hierarchical multi-label classification over the ticket data.
Chunqiu Zeng, Wubai Zhou, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
IEEE Trans. Netw. Serv. Manag.4
2016 Data-driven cloud-based IT services performance forecasting
abstract
Modern Cloud computing environments are rapidly evolving, leading to a growing adoption of dynamic pricing for virtual resources and of speedier deployment tools and to the emergence of hybrid Cloud scenarios. These trends suggest the opportunity to investigate a new generation of Cloud-based IT services, capable of adapting to changes in their operating conditions and deployment environment by dynamically realigning their configuration. This calls for new and more sophisticated management tools, that are capable of atomatically evaluating the performance of alternative configurations for Cloud-based IT services and of identifying the one that aligns better to the objectives defined by the business management.
Genady Grabarnik, Mauro Tortonesi, Larisa Shwartz
IEEE BigData3
2016 Online inference for time-varying temporal dependency discovery from time series
abstract
Large-scale time series data are prevalent across diverse application domains including system management, biomedical informatics, social networks, finance, etc. Temporal dependency discovery performs an essential part to identify the hidden interactions among the observed time series and helps to gain more insight into the behavior of the applications. However, the time-varying sparsity of the interactions among time series often poses a big challenge to temporal dependency discovery in practice. This paper formulates the temporal dependency problem with a novel Bayesian model allowing for both the sparsity and evolution of the hidden interactions among the observed time series. Taking advantage of the Bayesian modeling, an online inference method is proposed for time-varying temporal dependency discovery. Extensive empirical studies on both the synthetic and real application time series data are conducted to demonstrate the effectiveness and the efficiency of the proposed method.
Chunqiu Zeng, Qing Wang 0016, Wentao Wang 0006, Tao Li 0001, Larisa Shwartz
IEEE BigData5
2016 Towards establishing causality between change and incident
abstract
It is common knowledge in the IT service domain that changes to the system configuration are responsible for a major portion of incidents that result in client outages. However, it is typically very difficult to establish a relationship between changes and incidents as proper documentation takes lower priority at change creation time, as well as during incident management, in order to deal with the tremendous time pressure to quickly implement changes and resolve incidents. As a result, it is often not possible to leverage historical data to perform retrospective analysis to identify any emerging trends linking changes to incidents, or to build predictive models for proactive incident prevention at change creation time. In this paper, we present an approach for establishing causality between changes and incidents through an ensemble of statistics, data classification, and natural language processing techniques. We demonstrate our approach with a real world example.
Sinem Güven, Karin Murthy, Larisa Shwartz, Amit M. Paradkar
NOMS3
2016 Resolution Recommendation for Event Tickets in Service Management
abstract
In recent years, IT service providers have rapidly achieved an automated service delivery model. Software monitoring systems are designed to actively collect and signal event occurrences and, when necessary, automatically generate incident tickets. Repeating events generate similar tickets, which in turn have a vast number of repeated problem resolutions likely to be found in earlier tickets. In this paper, we develop techniques to recommend appropriate resolution for incoming events by making use of similarities between the events and historical resolutions of similar events. Built on the traditional k-nearest neighbor algorithm (KNN), our proposed algorithms take into account false positives often generated by monitoring systems. An additional penalty is incorporated into the algorithms to control the number of misleading resolutions in the recommendation results. Moreover, as the effectiveness of the KNN heavily relies on the underlying similarity measurement, we proposed two other approaches to significantly improve our recommendation with respect to resolution relevance. One approach uses topic-level features to incorporate resolution information into the similarity measurement; the other uses metric learning to learn a more effective similarity measure. Extensive empirical evaluations on three ticket data sets demonstrate the effectiveness and efficiency of our proposed methods.
Wubai Zhou, Chunqiu Zeng, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
IEEE Trans. Netw. Serv. Manag.5
2015 Modeling service variability in complex service delivery operations
abstract
One of the key promises of IT strategic outsourcing is to deliver greater IT service management through better quality and lower cost. However, this raises a critical question on how to model highly variable services for diverse customers with heterogeneous infrastructure and service demands. In this paper we propose the use of statistical learning approaches for service operation variability modeling. Specifically, we use the partial least squares regression that projects service attributes to explain the service volume variability, and the decision tree approach to model the service effort based on categorical customer and service properties. We demonstrate the applicability of the proposed methodology using data from a large IT service delivery environment.
Yixin Diao, Larisa Shwartz
CNSM2
2015 Recommending ticket resolution using feature adaptation
abstract
In recent years, IT Service Providers have been rapidly introducing automation to their service delivery model. Driven by market pressure to reduce cost and maintain quality of services, they are looking for technologies that will allow rapid progress towards attainment of truly automated service delivery. Software monitoring systems are designed to actively collect and signal event occurrences and, when necessary, automatically generate incident tickets. Repeating events generate similar tickets, which in turn have a vast number of repeated problem resolutions likely to be found in earlier tickets. In our work, we develop techniques to recommend an appropriate resolution for incoming events by making use of similarities between the events and historical resolutions of similar events. The traditional KNN (K Nearest Neighbor) algorithm has been first applied to recommend resolutions for incoming tickets. Massive heterogeneous applications as well as various monitoring software are running on clients' servers to accomplish required tasks and to monitor system health via different metrics. It leads to generation of correlated tickets that have different symptom descriptions but similar resolutions. Furthermore, change of servers' environments can also induce similar situations in which ticket descriptions differ before and after change but could have similar resolutions. These correlated tickets cause performance degradation in ticket resolution recommendation. Therefore, we propose using SCL (structural corresponding learning) based feature adaptation to uncover feature mapping in different time intervals. Moreover, to put more insights into the periodic regularities existing in our ticket datasets, we apply our algorithm on tickets grouped by different time interval granularities. Extensive empirical evaluations on real-world ticket data sets demonstrate the effectiveness and efficiency of our proposed methods.
Wubai Zhou, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
CNSM3
2015 BizMap: A framework for mapping business applications to IT infrastructure
abstract
Large enterprises are confronted with an increased need to fully understand the IT systems that support their business applications. Unfortunately, conventional communication trace-based approaches to discovering IT infrastructure dependencies are insufficient to gain a comprehensive mapping of business applications to IT infrastructure. This paper describes BizMap, a novel, information-discovery-based approach for deriving accurate business application topologies. BizMap uses a combination of information management, data analysis, and interactive visualization techniques to rapidly develop accurate topology models based on discovered information and configuration experts' insight. BizMap's evaluation is based on a comparison of the topologies created by BizMap's automated topology discovery with manually derived topologies with respect to performance and accuracy. It also includes a discussion of the user validation of the interactive BizMap steps.
Joel W. Branch, Karin Murthy, Larisa Shwartz, Emi Olsson, Robert A. Larsen
IM3
2015 Business-driven configuration of IT services in public and hybrid clouds based on performance forecasting
abstract
Modern Cloud computing environments are rapidly evolving, leading to a growing adoption of dynamic pricing for virtual resources and of speedier deployment tools, and to the emergence of hybrid Cloud scenarios. These trends suggest the opportunity to investigate a new generation of Cloud-based IT services, capable of adapting to changes in their operating conditions and deployment environment by dynamically realigning their configuration. This calls for new and more sophisticated management tools, that are capable of evaluating the performance of alternative configurations for Cloud-based IT services and of identifying the one that aligns better to the objectives defined by the business management. This paper presents an optimization tool for Cloud-based IT services, based on queuing theoretic analysis of service workflows and ILP optimization.
Genady Grabarnik, Mauro Tortonesi, Larisa Shwartz
IM3
2015 Resolution recommendation for event tickets in service management
abstract
In recent years, IT Service Providers have been rapidly transforming to an automated service delivery model. This is due to advances in technology and driven by the unrelenting market pressure to reduce cost and maintain quality. Tremendous progress has been made to date towards attainment of truly automated service delivery; that is, the ability to deliver the same service automatically using the same process with the same quality. However, automating Incident and Problem Management continuous to be a difficult problem, particularly due to the growing complexity of IT environments. Software monitoring systems are designed to actively collect and signal event occurrances and, when necessary, automatically generate incident tickets. Repeating events generate similar tickets, which in turn have a vast number of repeated problem resolutions likely to be found in earlier tickets. In this paper we find an appropriate resolution by making use of similarities between the events and previous resolutions of similar events. Traditional KNN (K Nearest Neighbor) algorithm has been used to recommend resolutions for incoming tickets. However, the effectiveness of recommendation heavily relies on the underlying similarity measure in KNN. In this paper, we significantly improve the similarity measure used in KNN by utilizing both the event and resolution information in historical tickets via a topic-level feature extraction using the LDA (Latent Dirichlet Allocation) model. In addition, when resolution categories are available, we propose to learn a more effective similarity measure using metric learning. Extensive empirical evaluations on three ticket data sets demonstrate the effectiveness and efficiency of our proposed methods.
Wubai Zhou, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
IM4
2014 A framework for predicting service delivery efforts using IT infrastructure-to-incident correlation
abstract
Predicting IT infrastructure performance under varying conditions, e.g., the addition of a new server or increased transaction loads, has become a typical IT management exercise. However, within a service delivery context, enterprise clients are demanding predictive analytics that outline future “costs” associated with changing conditions. The service delivery staffing costs incurred in addressing problems and requests (arriving in the form of incident and other problem tickets) in the managed environment is especially of high importance. This paper describes an analytical study addressing such cost prediction. Specifically, a novel approach is described in which support vector regression is used to predict service delivery workloads (measured by ticket volumes) based on managed server characteristics Additionally, a proposed framework combining various analytical models is proposed to predict service delivery staffing requirements under changing IT infrastructure characteristics and conditions. Detailed descriptions of the workload prediction techniques, as well as an evaluation using data from an actual large service delivery engagement, are presented.
Joel W. Branch, Yixin Diao, Larisa Shwartz
NOMS3
2014 Predicting service delivery cost for non-standard service level agreements
abstract
One of the key promises of IT strategic outsourcing is to deliver greater IT service management through lower cost. However, this raises a critical question on how to predict service delivery cost during the service engagement phase where nonstandard service level agreements (SLAs) are negotiated and detailed service modeling data are not available. In this paper we propose a modeling framework that uses queueing model based approaches to estimate the impact of SLAs on the delivery cost. We further propose a set of approximation techniques to address the complexity of service delivery and an optimization model to predict the delivery cost subject to service level constraints and service stability conditions. We demonstrate the applicability of the proposed methodology using data from a large IT service delivery environment.
Yixin Diao, Linh Lam, Larisa Shwartz, David M. Northcutt
NOMS3
2014 Testing performance of IT service associates under variable load
abstract
One of the challenges in IT Service Management is an assessment of an innovation benefit introduced into service delivery on performance of IT System Administrator (SA). The question is especially complex when variability of the service requests load is taken into account. We used the sequential probability ratio test to assess effects of innovation for a case when a load of the service teams is constant over time. One of the drawbacks of this test is its requirement for a constant SAs team size while the arrival patterns of service requests might require otherwise. To address the volatility of a service request load we consider the possibility of changing number of service associates. We design a performance test based on the concepts of Repeated Significance Test (RST), which allowed us to take variable number of SAs into account. In this work we show the superiority of the suggested test in comparison to the fixed sample size test (FSST) in a number of experiments required to reach the definitive conclusion. We also compare RST with SPRT when the size of the service team is constant.
Genady Grabarnik, Yefim Haim Michlin, Larisa Shwartz
NOMS3
2014 Business-driven optimization of component placement for complex services in federated Clouds
abstract
With the advent of connected services ecosystems, new generations of services and systems are being conceived, responding to the ever growing demands of the market place. In parallel, the effective adoption of the Cloud computing paradigm is becoming an essential enabler for business enterprises. With such importance placed on the services ecosystem, the design and management of services becomes a key issue both for the providers and the users. One of the main challenges for service providers lies in the complexity of the services, comprising of a multiplicity of technologies and competing and cooperating providers, which is difficult to address through current technology-centric service design approaches, in particular for the deployment infrastructure. The work described in this paper lays a foundation for business driven service design by proposing a business goals driven model of resource allocation in the Cloud. We define a goal/loss/processing cost function for resource allocation that we optimize, while taking into account the dynamic and varying nature of requests load.
Genady Grabarnik, Larisa Shwartz, Mauro Tortonesi
NOMS2
2014 Hierarchical multi-label classification over ticket data using contextual loss
abstract
Maximal automation of routine IT maintenance procedures is an ultimate goal of IT service management. System monitoring, an effective and reliable means for IT problem detection, generates monitoring tickets to be processed by system administrators. IT problems are naturally organized in a hierarchy by specialization. The problem hierarchy is used to help triage tickets to the processing team for problem resolving. In this paper, a hierarchical multi-label classification method is proposed to classify the monitoring tickets by utilizing the problem hierarchy. In order to find the most effective classification, a novel contextual hierarchy (CH) loss is introduced in accordance with the problem hierarchy. Consequently, an arising optimization problem is solved by a new greedy algorithm. An extensive empirical study over ticket data was conducted to validate the effectiveness and efficiency of our method.
Chunqiu Zeng, Tao Li 0001, Larisa Shwartz, Genady Grabarnik
NOMS3
2014 Modeling the Impact of Service Level Agreements During Service Engagement
abstract
One of the key promises of IT strategic outsourcing is to deliver greater IT service management through lower cost. However, this raises a critical question: How can one predict the service delivery cost that will deliver the promised service level agreements (SLAs)? This is particularly challenging since such prediction is mostly needed during the service engagement phase where the SLAs and the delivery cost are negotiated, and the detailed service modeling data are not available. In this paper, we propose a modeling framework that uses queueing-model-based approaches to estimate the impact of SLAs on the delivery cost. We further propose a set of approximation techniques to address the complexity of service delivery and an optimization model to predict the delivery cost subject to service-level constraints and service stability conditions. We demonstrate the applicability of the proposed methodology using data from a large IT service delivery environment.
Yixin Diao, Linh Lam, Larisa Shwartz, David M. Northcutt
IEEE Trans. Netw. Serv. Manag.3
2013 SLA impact modeling for service engagement
abstract
During the customer engagement phase it is critical for the service providers to estimate the impact of service level constraints on service personnel needs. However, it is often difficult due to the implication from customer workload. In this paper we propose an SLA impact evaluation methodology that uses queueing models to quantitatively evaluate the impact of SLAs to the engagement cost model.
Yixin Diao, Linh Lam, Larisa Shwartz, David M. Northcutt
CNSM3
2013 Identifying missed monitoring alerts based on unstructured incident tickets
abstract
Automatic system monitoring is an efficient and reliable mean for problem detection in enterprise IT infrastructures. The performance of monitoring systems depends on their configurations specified by the system administrators. In dynamic and large IT environments, the IT infrastructures are frequently changed to meet various business requirements, so the configurations may not be always consistent with the updated status. Misconfigurations can lead to false positive (false alarms) and false negative (missing alerts) for the system administrators. The false negatives can cause serious system faults. This paper presents an automatic approach for discovering the false negatives from incident tickets that are created by humans. The discovered results help the system administrators correct the misconfigurations and minimize the false negatives in future. This approach applies a text classification model for analyzing the descriptions of incident tickets and identifying the corresponding system issues. The domain knowledge for describing those issues can be incorporated to assist with this model. Experiments are conducted on real system incident tickets from a large enterprise IT infrastructure. The experimental results demonstrate the effectiveness of the proposed approach.
Tao Li 0001, Larisa Shwartz, Genady Grabarnik
CNSM3
2013 Robustness of comparison sequential test for the piloting in Service Delivery
Yefim Haim Michlin, Genady Grabarnik, Larisa Shwartz, Ofer Shaham
IM3
2013 Quality improvement and quantitative modeling - Using mashups for human error prevention
Carlos Raniery Paula dos Santos, Lisandro Z. Granville, Larisa Shwartz, Nikos Anerousis, David Loewenstern
IM3
2013 Recommending resolutions for problems identified by monitoring
Tao Li 0001, Larisa Shwartz, Genady Grabarnik
IM3
2013 An integrated framework for optimizing automatic monitoring systems in large IT infrastructures
abstract
The competitive business climate and the complexity of IT environments dictate efficient and cost-effective service delivery and support of IT services. These are largely achieved by automating routine maintenance procedures, including problem detection, determination and resolution. System monitoring provides an effective and reliable means for problem detection. Coupled with automated ticket creation, it ensures that a degradation of the vital signs, defined by acceptable thresholds or monitoring conditions, is flagged as a problem candidate and sent to supporting personnel as an incident ticket. This paper describes an integrated framework for minimizing false positive tickets and maximizing the monitoring coverage for system faults.
Tao Li 0001, Larisa Shwartz, Florian Pinel, Genady Grabarnik
KDD3
2012 A Learning Method for Improving Quality of Service Infrastructure Management in New Technical Support Groups
David Loewenstern, Florian Pinel, Larisa Shwartz, Maíra Gatti de Bayser, Ricardo Herrmann
ICSOC3
2012 Domain-Independent Data Validation and Content Assistance as a Service
abstract
In this paper we describe a scalable service for customized data validation and content assistance by means of domain-independent, user-provided sets of complex data constraints. We present an integrated architecture and a particular implementation that combines the use of existing grammar- and rule-based schema languages that allows providers to specify rules in a declarative manner. The integrated architecture provides a way to semi-automatically fill in form fields by calling the proposed service, which enumerates domains from the previously stored data validation constraints. Additionally, better error reporting can be achieved by leveraging structure from the rules' definitions. The proposed architecture has been shown to be practical and is in production use by a large organization, successfully fulfilling its role.
Maíra Gatti de Bayser, Ricardo Herrmann, David Loewenstern, Florian Pinel, Larisa Shwartz
ICWS5
2012 Discovering lag intervals for temporal dependencies
abstract
Time lag is a key feature of hidden temporal dependencies within sequential data. In many real-world applications, time lag plays an essential role in interpreting the cause of discovered temporal dependencies. Traditional temporal mining methods either use a predefined time window to analyze the item sequence, or employ statistical techniques to simply derive the time dependencies among items. Such paradigms cannot effectively handle varied data with special properties, e.g., the interleaved temporal dependencies.
Tao Li 0001, Larisa Shwartz
KDD3
2012 SUITS: How to make a global IT service provider sustainable?
abstract
IT Service Providers are faced with the challenge to reconcile their goals in providing services efficiently and with the quality of services required while sustaining the environment. With the introduction of carbon taxation and regulations to control green gas emissions, there is a need to expand service design to incorporate heterogeneity in carbon taxes due to regional differences. The service providers have flexibility of selecting best locations for delivery of a service by taking into consideration operational cost associated with green gas emissions. They can also buy or sell carbon credits in a carbon trading market to maintain sustainable operational performance. In this paper, we introduce the Sustainable IT Services (SUITS) framework to develop models for operational performance of an IT Service Provider in a global environment with carbon taxation and carbon credit trading markets. The models are used to formulate a revenue optimization problem for the IT Service Provider in this environment. The model solutions can provide guidance for designing operational level of different IT components and for creating effective strategies for trading in carbon markets.
Parijat Dube, Genady Grabarnik, Larisa Shwartz
NOMS3
2012 Designing pilot for operational innovation in IT service delivery
abstract
Dramatic changes are taking place in the world of IT services: in who is competing with whom, in what determines competitive success, in the technologies of product and production, and in the very ways of business' approach to innovation. Operational Innovation has become a vital necessity for IT Service Providers. It affects what their employees do every day and how they do that. Because it impacts the very core of service delivery, there is an obvious risk associated with it. The benefits of operational changes are not always easy to assess and it often creates confusion. The confusion is real and the stakes for businesses are high. We argue that direct experimentation, or piloting, is both necessary and possible for introduction of operational innovations into Service Delivery. The novelty of this work is twofold: first, we propose to use design of experiments framework for IT service delivery, addressing the essential differences between delivery of IT services and manufacturing; second, we propose robust sequential design of experiments for service processes as part of the framework and demonstrate it on a sample process.
Genady Grabarnik, Yefim Haim Michlin, Larisa Shwartz
NOMS3
2012 A learning feature engineering method for task assignment
abstract
Multi-domain IT services are delivered by technicians with a variety of expert knowledge in different areas. Their skills and availability are an important property of the service. However, most organizations do not have a consistent view of this information because creation and maintenance of a skill model is a difficult task, especially in light of privacy regulations, changing service catalogs and worker turnover. We propose a method for ranking technicians on their expected performance according to their suitability for receiving the assignment of a service request without maintaining an explicit skill model describing which skills are possessed by each technician. We find appropriate assignees by making use of similarities between the assignees and previous tasks performed by them.
David Loewenstern, Florian Pinel, Larisa Shwartz, Maíra Gatti de Bayser, Ricardo Herrmann, Victor F. Cavalcante
NOMS3
2012 Optimizing system monitoring configurations for non-actionable alerts
abstract
Today's competitive business climate and the complexity of IT environments dictate efficient and cost effective service delivery and support of IT services. This is largely achieved through automating of routine maintenance procedures including problem detection, determination and resolution. System monitoring provides effective and reliable means for problem detection. Coupled with automated ticket creation, it ensures that a degradation of the vital signs, defined by acceptable thresholds or monitoring conditions, is flagged as a problem candidate and sent to supporting personnel as an incident ticket. This paper describes a novel methodology and a system for minimizing non-actionable tickets while preserving all tickets which require corrective action. Our proposed method defines monitoring conditions and the optimal corresponding delay times based on an off-line analysis of historical alerts and the matching incident tickets. Potential monitoring conditions are built on a set of predictive rules which are automatically generated by a rule-based learning algorithm with coverage, confidence and rule complexity criteria. These conditions and delay times are propagated as configurations into run-time monitoring systems.
Tao Li 0001, Florian Pinel, Larisa Shwartz, Genady Grabarnik
NOMS4
2011 Performance management and quantitative modeling of IT service processes using mashup patterns
Carlos Raniery Paula dos Santos, Lisandro Z. Granville, Winnie Cheng, David Loewenstern, Larisa Shwartz, Nikos Anerousis
CNSM5
2010 Workload Migration into Clouds Challenges, Experiences, Opportunities
abstract
The steady drumbeat of Cloud as a disruptive influence for Infrastructure Service Providers (ISP's) and the enablement vehicle for Software As A Service (SAAS)providers can be heard loud and clear in the industry today. In fact, Cloud is probably at the peak of the hype curve, and already there are identified challenges associated with effective deployment for business critical applications (so called Production Applications) in mature enterprises. One of these challenges is the smooth migration of workload from the previous environment to the new cloud enabled environment in a cost effective way, with minimal disruption and risk. In this paper we introduce extensions to an integrated automation capability called the Darwin framework that enables workload migration for this scenario and discuss the impact that automated migration has on the cost and risks normally associated with migration to clouds.
Christopher Ward, N. Aravamudan, Kamal Bhattacharya, Karen Cheng, Robert Filepp, Robert D. Kearney, Brian Peterson, Larisa Shwartz, Christopher C. Young
IEEE CLOUD8
2010 Automating the delivery of IT Service Continuity Management through cloud service orchestration
abstract
IT Service Continuity Management (ITSCM) delivers the recovery of IT services in the event of a disaster. ITSCM is widely perceived as an expensive challenge for enterprise-class IT operations. Cloud computing offers a model for dynamic, scalable infrastructure resource allocation on a pay-per-use basis. These attributes promise to bring cost-efficiency to ITSCM invocation and operation processes that only in the rare event of a rehearsal or an actual disaster need to allocate infrastructure resources. We propose to use the Web Service Business Process Execution Language (BPEL) in combination with Virtual Appliances to implement standardized, testable and executable ITSCM processes. The suggested solution is described and evaluated against collected data from manual recovery processes.
Markus Klems, Stefan Tai, Larisa Shwartz, Genady Grabarnik
NOMS3
2009 Towards an optimized model of incident ticket correlation
abstract
In recent years, IT service management (ITSM) has become one of the most researched areas of IT. Incident and problem management are two of the service operation processes in the IT infrastructure library (ITIL). These two processes aim to recognize, log, isolate and correct errors which occur in the environment and disrupt the delivery of services. Incident management and problem management form the basis of the tooling provided by an incident ticket systems (ITS). In an ITS system, seemingly unrelated tickets created by end users and monitoring systems can coexist and have the same root cause. The connection between failed resource and malfunctioning services is not realized automatically, but often established manually by means of human intervention. This need for human involvement reduces productivity. The introduction of automation would increase productivity and therefore reduce the cost of incident resolution. In this paper, we propose a model to correlate incident tickets based on three criteria. First, we employ a category-based correlation that relies on matching service identifiers with associated resource identifiers, using similarity rules. Secondly, we correlate the configuration items which are critical to the failed service with the earlier identified resource tickets in order to optimize the topological comparison. Finally, we augment scheduled resource data collection with constraint adaptive probing to minimize the correlation interval for temporally correlated tickets. We present experimental data in support of our proposed correlation model.
Patricia Marcu, Genady Grabarnik, Laura Z. Luan, Daniela Rosu 0001, Larisa Shwartz, Christopher Ward
Integrated Network Management5
2009 Multi-tenant solution for IT service management: A quantitative study of benefits
abstract
The very competitive business climate dictates efficient and cost effective delivery and support of IT services. Compounded by complexity of IT environments and criticality of IT to business success, IT service providers seek multi-tenant solutions to reduce operational cost and improve service quality. In this paper we consider a multi-tenant solution for IT service management and examine its critical aspects for realizing business benefits. We conduct a quantitative study of its benefits using a complexity based value assessment methodology. By regarding complexity as a substitute for potential labor cost, we estimate the business value of the multi-tenant solution before its actual deployment.
Larisa Shwartz, Yixin Diao, Genady Grabarnik
Integrated Network Management1
2008 Decomposition of IT service processes and alternative service identification using ontologies
abstract
Providers of IT services are under constant pressure to reduce cost and improve the quality of the services they provide. The ability to decompose such IT services (defined by themselves or third party suppliers) into elemental service processes characterized by core function can benefit the providers by allowing them to identify alternative service processes of potentially lower cost and/or improved quality. This paper proposes an ontology-based hierarchical service decomposition and identification approach to support service providers in managing their operational service processes by the characterization and exploitation of such elemental service processes. We further demonstrate the feasibility and sample outcomes of the proposed approach an example.
Christian Bartsch 0002, Larisa Shwartz, Christopher Ward, Genady Grabarnik, Melissa J. Buco
NOMS2
2007 Service Provider Considerations for IT Service Management
abstract
The IT infrastructure library (ITIL) provides clarity to IT service management and its processes. ICO 20000 further describes the responsibilities of service providers operating in ITIL conformant environment. However, neither one specifically addresses external service providers' challenges that arise in a cost conserving operational model when service provider is supporting multiple customers using shared resources. This paper provides a summary of a required extension to the service provider domain for an IT service management (ITSM) solution with ITIL conformant core.
Larisa Shwartz, Naga Ayachitula, Melissa J. Buco, Maheswaran Surendra, Christopher Ward, Steve Weinberger
Integrated Network Management1
2005 Collaborative End-Point Service Modulation System (COSMOS)
Naga Ayachitula, Shu-Ping Chang, Larisa Shwartz, Maheswaran Surendra
WISE3