Rolf Stadler

dblp:s/RolfStadler · DBLP profile ↗
← Back
87ranked-venue papers
1as first author
19since 2021 · last 2025
0000-0001-6039-8493ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 47 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 1 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Learning Optimal Defender Strategies for CAGE-2 using a POMDP Model
abstract
CAGE-2 is an accepted benchmark for learning and evaluating defender strategies against cyberattacks. It reflects a scenario where a defender agent protects an IT infrastructure against various attacks. Many defender methods for CAGE-2 have been proposed in the literature. In this paper, we construct a formal model for CAGE-2 using the framework of Partially Observable Markov Decision Process (POMDP). Based on this model, we define an optimal defender strategy for CAGE-2 and introduce a method to efficiently learn this strategy. Our method, called BF-PPO, is based on PPO, and it uses particle filter to mitigate the computational complexity due to the large state space of the CAGE-2 model. We evaluate our method in the CAGE-2 CybORG environment and compare its performance with that of CARDIFF, the highest ranked method on the CAGE-2 leaderboard. We find that our method outperforms CARDIFF regarding the learned defender strategy and the required training time.
Duc Huy Le, Rolf Stadler
CNSM2
2025 Adaptive Security Response Strategies Through Conjectural Online Learning
abstract
We study the problem of learning adaptive security response strategies for an it infrastructure. We formulate the interaction between an attacker and a defender as a partially observed, non-stationary game. We relax the standard assumption that the game model is correctly specified and consider that each player has a probabilistic conjecture about the model, which may be misspecified in the sense that the true model has probability 0. This formulation allows us to capture uncertainty and misconception about the infrastructure and the intents of the players. To learn effective game strategies online, we design Conjectural Online Learning (col), a novel method where a player iteratively adapts its conjecture using Bayesian learning and updates its strategy through rollout. We prove that the conjectures converge to best fits, and we provide a bound on the performance improvement that rollout enables with a conjectured model. To characterize the steady state of the game, we propose a variant of the Berk-Nash equilibrium. We present col through an intrusion response use case. Testbed evaluations show that col produces effective security strategies that adapt to a changing environment. We also find that col enables faster convergence than current reinforcement learning techniques.
Kim Hammar, Tao Li 0046, Rolf Stadler, Quanyan Zhu
IEEE Trans. Inf. Forensics Secur.3
2024 Intrusion Tolerance for Networked Systems through Two-Level Feedback Control
abstract
We formulate intrusion tolerance for a system with service replicas as a two-level optimal control problem. On the local level node controllers perform intrusion recovery, and on the global level a system controller manages the replication factor. The local and global control problems can be formulated as classical problems in operations research, namely, the machine replacement problem and the inventory replenishment problem. Based on this formulation, we design TOLERANCE, a novel control architecture for intrusion-tolerant systems. We prove that the optimal control strategies on both levels have threshold structure and design efficient algorithms for computing them. We implement and evaluate TOLERANCE in an emulation environment where we run 10 types of network intrusions. The results show that TOLERANCE can improve service availability and reduce operational cost compared with state-of-the-art intrusion-tolerant systems.
Kim Hammar, Rolf Stadler
DSN2
2024 Online Policy Adaptation for Networked Systems using Rollout
abstract
Dynamic resource allocation in networked systems is needed to continuously achieve end-to-end management objectives. Recent research has shown that reinforcement learning can achieve near-optimal resource allocation policies for realistic system configurations. However, most current solutions require expensive retraining when changes in the system occur. We address this problem and introduce an efficient method to adapt a given base policy to system changes, e.g., to a change in the service offering. In our approach, we adapt a base control policy using a rollout mechanism, which transforms the base policy into an improved rollout policy. We perform extensive evaluations on a testbed where we run applications on a service mesh based on the Istio and Kubernetes platforms. The experiments provide insights into the performance of different rollout algorithms. We find that our approach produces policies that are equally effective as those obtained by offline retraining. On our testbed, effective policy adaptation takes seconds when using rollout, compared to minutes or hours when using retraining. Our work demonstrates that rollout, which has been applied successfully in other domains, is an effective approach for policy adaptation in networked systems.
Forough Shahab Samani, Kim Hammar, Rolf Stadler
NOMS3
2024 Comparing Transfer Learning and Rollout for Policy Adaptation in a Changing Network Environment
abstract
Dynamic resource allocation for network services is pivotal for achieving end-to-end management objectives. Previous research has demonstrated that Reinforcement Learning (RL) is a promising approach to resource allocation in networks, allowing to obtain near-optimal control policies for non-trivial system configurations. Current RL approaches however have the drawback that a change in the system or the management objective necessitates expensive retraining of the RL agent. To tackle this challenge, practical solutions including offline retraining, transfer learning, and model-based rollout have been proposed. In this work, we study these methods and present comparative results that shed light on their respective performance and benefits. Our study finds that rollout achieves faster adaptation than transfer learning, yet its effectiveness highly depends on the accuracy of the system model.
Forough Shahab Samani, Hannes Larsson, Simon Damberg, Andreas Johnsson, Rolf Stadler
NOMS5
2024 IT Intrusion Detection Using Statistical Learning and Testbed Measurements
abstract
We study automated intrusion detection in an IT infrastructure, specifically the problem of identifying the start of an attack, the type of attack, and the sequence of actions an attacker takes, based on continuous measurements from the infrastructure. We apply statistical learning methods, including Hidden Markov Model (HMM), Long Short-Term Memory (LSTM), and Random Forest Classifier (RFC) to map sequences of observations to sequences of predicted attack actions. In contrast to most related research, we have abundant data to train the models and evaluate their predictive power. The data comes from traces we generate on an in-house testbed where we run attacks against an emulated IT infrastructure. Central to our work is a machine-learning pipeline that maps measurements from a high-dimensional observation space to a space of low dimensionality or to a small set of observation symbols. Investigating intrusions in offline as well as online scenarios, we find that both HMM and LSTM can be effective in predicting attack start time, attack type, and attack actions. If sufficient training data is available, LSTM achieves higher prediction accuracy than HMM. HMM, on the other hand, requires less computational resources and less training data for effective prediction. Also, we find that the methods we study benefit from data produced by traditional intrusion detection systems like SNORT.
Rolf Stadler
NOMS2
2024 Learning Near-Optimal Intrusion Responses Against Dynamic Attackers
abstract
We study automated intrusion response and formulate the interaction between an attacker and a defender as an optimal stopping game where attack and defense strategies evolve through reinforcement learning and self-play. The gametheoretic modeling enables us to find defender strategies that are effective against a dynamic attacker, i.e. an attacker that adapts its strategy in response to the defender strategy. Further, the optimal stopping formulation allows us to prove that best response strategies have threshold properties. To obtain nearoptimal defender strategies, we develop Threshold Fictitious Self-Play (T-FP), a fictitious self-play algorithm that learns Nash equilibria through stochastic approximation. We show that T-FP outperforms a state-of-the-art algorithm for our use case. The experimental part of this investigation includes two systems: a simulation system where defender strategies are incrementally learned and an emulation system where statistics are collected that drive simulation runs and where learned strategies are evaluated. We argue that this approach can produce effective defender strategies for a practical IT infrastructure.
Kim Hammar, Rolf Stadler
IEEE Trans. Netw. Serv. Manag.2
2024 A Framework for Dynamically Meeting Performance Objectives on a Service Mesh
abstract
We present a framework for achieving end-to-end management objectives for multiple services that concurrently execute on a service mesh. We apply reinforcement learning (RL) techniques to train an agent that periodically performs control actions to reallocate resources. We develop and evaluate the framework using a laboratory testbed where we run information and computing services on a service mesh, supported by the Istio and Kubernetes platforms. We investigate different management objectives that include end-to-end delay bounds on service requests, throughput objectives, cost-related objectives, and service differentiation. Our framework supports the design of a control agent for a given management objective. The management objective is defined first and then mapped onto available control actions. Several types of control actions can be executed simultaneously, which allows for efficient resource utilization. Second, the framework separates the learning of the system model and the operating region from the learning of the control policy. By first learning the system model and the operating region from testbed traces, we can instantiate a simulator and train the agent for different management objectives. Third, the use of a simulator shortens the training time by orders of magnitude compared with training the agent on the testbed. We evaluate the learned policies on the testbed and show the effectiveness of our approach in several scenarios. In one scenario, we design a controller that achieves the management objectives with 50% less system resources than Kubernetes HPA autoscaling.
Forough Shahab Samani, Rolf Stadler
IEEE Trans. Netw. Serv. Manag.2
2023 Digital Twins for Security Automation
abstract
We present a novel emulation system for creating high-fidelity digital twins of IT infrastructures. The digital twins replicate key functionality of the corresponding infrastructures and allow to play out security scenarios in a safe environment. We show that this capability can be used to automate the process of finding effective security policies for a target infrastructure. In our approach, a digital twin of the target infrastructure is used to run security scenarios and collect data. The collected data is then used to instantiate simulations of Markov decision processes and learn effective policies through reinforcement learning, whose performances are validated in the digital twin. This closed-loop learning process executes iteratively and provides continuously evolving and improving security policies. We apply our approach to an intrusion response scenario. Our results show that the digital twin provides the necessary evaluative feedback to learn near-optimal intrusion response policies.
Kim Hammar, Rolf Stadler
NOMS2
2023 Demonstrating a System for Dynamically Meeting Management Objectives on a Service Mesh
abstract
We demonstrate a management system that lets a service provider achieve end-to-end management objectives under varying load for applications on a service mesh based on the Istio and Kubernetes platforms. The management objectives for the demonstration include end-to-end delay bounds on service requests, throughput objectives, and service differentiation. Our method for finding effective control policies includes a simulator and a control module. The simulator is instantiated with traces from a testbed, and the control module trains a reinforcement learning (RL) agent to efficiently learn effective control policies on the simulator. The learned policies are then transfered to the testbed to perform dynamic control actions based on monitored system metrics. We show that the learned policies dynamically meet management objectives on the testbed and can be changed on the fly.
Forough Shahab Samani, Kim Hammar, Rolf Stadler
NOMS3
2022 An Online Framework for Adapting Security Policies in Dynamic IT Environments
abstract
We present an online framework for learning and updating security policies in dynamic IT environments. It includes three components: a digital twin of the target system, which continuously collects data and evaluates learned policies; a system identification process, which periodically estimates system models based on the collected data; and a policy learning process that is based on reinforcement learning. To evaluate our framework, we apply it to an intrusion prevention use case that involves a dynamic IT infrastructure. Our results demonstrate that the framework automatically adapts security policies to changes in the IT infrastructure and that it outperforms a state-of-the-art method.
Kim Hammar, Rolf Stadler
CNSM2
2022 Dynamically meeting performance objectives for multiple services on a service mesh
abstract
We present a framework that lets a service provider achieve end-to-end management objectives under varying load. Dynamic control actions are performed by a reinforcement learning (RL) agent. Our work includes experimentation and evaluation on a laboratory testbed where we have implemented basic information services on a service mesh supported by the Istio and Kubernetes platforms. We investigate different management objectives that include end-to-end delay bounds on service requests, throughput objectives, and service differentiation. These objectives are mapped onto reward functions that an RL agent learns to optimize, by executing control actions, namely, request routing and request blocking. We compute the control policies not on the testbed, but in a simulator, which speeds up the learning process by orders of magnitude. In our approach, the system model is learned on the testbed; it is then used to instantiate the simulator, which produces near-optimal control policies for various management objectives. The learned policies are then evaluated on the testbed using unseen load patterns.
Forough Shahab Samani, Rolf Stadler
CNSM2
2022 A System for Interactive Examination of Learned Security Policies
abstract
We present a system for interactive examination of learned security policies. It allows a user to traverse episodes of Markov decision processes in a controlled manner and to track the actions triggered by security policies. Similar to a software debugger, a user can continue or or halt an episode at any time step and inspect parameters and probability distributions of interest. The system enables insight into the structure of a given policy and in the behavior of a policy in edge cases. We demonstrate the system with a network intrusion use case. We examine the evolution of an IT infrastructure’s state and the actions prescribed by security policies while an attack occurs. The policies for the demonstration have been obtained through a reinforcement learning approach that includes a simulation system where policies are incrementally learned and an emulation system that produces statistics that drive the simulation runs.
Kim Hammar, Rolf Stadler
NOMS2
2022 Intrusion Prevention Through Optimal Stopping
abstract
We study automated intrusion prevention using reinforcement learning. Following a novel approach, we formulate the problem of intrusion prevention as an (optimal) multiple stopping problem. This formulation gives us insight into the structure of optimal policies, which we show to have threshold properties. For most practical cases, it is not feasible to obtain an optimal defender policy using dynamic programming. We therefore develop a reinforcement learning approach to approximate an optimal threshold policy. We introduce T- SPSA, an efficient reinforcement learning algorithm that learns threshold policies through stochastic approximation. We show that T- SPSA outperforms state-of-the-art algorithms for our use case. Our overall method for learning and validating policies includes two systems: a simulation system where defender policies are incrementally learned and an emulation system where statistics are produced that drive simulation runs and where learned policies are evaluated. We show that this approach can produce effective defender policies for a practical IT infrastructure.
Kim Hammar, Rolf Stadler
IEEE Trans. Netw. Serv. Manag.2
2022 Online Feature Selection for Efficient Learning in Networked Systems
abstract
Current AI/ML methods for data-driven engineering use models that are mostly trained offline. Such models can be expensive to build in terms of communication and computing costs, and they rely on data that is collected over extended periods of time. Further, they become out-of-date when changes in the system occur. To address these challenges, we investigate online learning techniques that automatically reduce the number of available data sources for model training. We present an online algorithm called Online Stable Feature Set Algorithm (OSFS), which selects a small feature set from a large number of available data sources after receiving a small number of measurements. The algorithm is initialized with a feature ranking algorithm, a feature set stability metric, and a search policy. We perform an extensive experimental evaluation of this algorithm using traces from an in-house testbed and from two external datasets. We find that OSFS achieves a massive reduction in the size of the feature set by 1–3 orders of magnitude on all investigated datasets. Most importantly, we find that the accuracy of a predictor trained on a OSFS-produced feature set is somewhat better than when the predictor is trained on a feature set obtained through offline feature selection. OSFS is thus shown to be effective as an online feature selection algorithm and robust regarding the sample interval used for feature selection. We also find that, when concept drift in the data underlying the model occurs, its effect can be mitigated by recomputing the feature set and retraining the prediction model.
Rolf Stadler
IEEE Trans. Netw. Serv. Manag.2
2021 Learning Intrusion Prevention Policies through Optimal Stopping
abstract
We study automated intrusion prevention using reinforcement learning. In a novel approach, we formulate the problem of intrusion prevention as an optimal stopping problem. This formulation allows us insight into the structure of the optimal policies, which turn out to be threshold based. Since the computation of the optimal defender policy using dynamic programming is not feasible for practical cases, we approximate the optimal policy through reinforcement learning in a simulation environment. To define the dynamics of the simulation, we emulate the target infrastructure and collect measurements. Our evaluations show that the learned policies are close to optimal and that they indeed can be expressed using thresholds.
Kim Hammar, Rolf Stadler
CNSM2
2021 Online Feature Selection for Low-overhead Learning in Networked Systems
abstract
Data-driven functions for operation and management require measurements and readings from distributed data sources for model training and prediction. While the number of candidate data sources can be very large, research has shown that it is often possible to reduce the number of data sources significantly while still allowing for accurate prediction. Consequently, there is potential to lower communication and computing resources needed to continuously extract, collect, and process this data. We demonstrate the operation of a novel online algorithm called OSFS, which sequentially processes the collected data and reduces the number of data sources for training prediction models. OSFS builds on two main ideas, namely (1) ranking the available data sources using (unsupervised) feature selection algorithms and (2) identifying stable feature sets that include only the top features. The demonstration shows the search space exploration, the iterative selection of feature sets, and the evaluation of the stability of these sets. The demonstration uses measurements collected from a KTH testbed, and the predictions relate to end-to-end KPIs for network services.
Forough Shahab Samani, Andreas Johnsson, Rolf Stadler
CNSM4
2021 Online Learning under Resource Constraints
Rodolfo da Silva Villaça, Rolf Stadler
IM2
2021 Conditional Density Estimation of Service Metrics for Networked Services
abstract
We predict the conditional distributions of service metrics, such as response time or frame rate, from infrastructure measurements in a networked environment. From such distributions, key statistics of the service metrics, including mean, variance, or quantiles can be computed, which are essential for predicting SLA conformance and enabling service assurance. We present and assess two methods for prediction: (1) mixture models with Gaussian or Lognormal kernels, whose parameters are estimated using mixture density networks, a class of neural networks, and (2) histogram models, which require the target space to be discretized. We apply these methods to a VoD service and a KV store service running on our lab testbed. A comparative evaluation shows the relative effectiveness of the methods when applied to operational data. We find that both methods allow for accurate prediction. While mixture models provide a general and elegant solution, they incur a very high overhead related to hyper-parameter search and neural network training. Histogram models, on the other hand, allow for efficient training, but require adjustment to the specific use case.
Forough Shahab Samani, Rolf Stadler, Christofer Flinta, Andreas Johnsson
IEEE Trans. Netw. Serv. Manag.2
2020 Finding Effective Security Strategies through Reinforcement Learning and Self-Play
abstract
We present a method to automatically find security strategies for the use case of intrusion prevention. Following this method, we model the interaction between an attacker and a defender as a Markov game and let attack and defense strategies evolve through reinforcement learning and self-play without human intervention. Using a simple infrastructure configuration, we demonstrate that effective security strategies can emerge from self-play. This shows that self-play, which has been applied in other domains with great success, can be effective in the context of network security. Inspection of the converged policies show that the emerged policies reflect common-sense knowledge and are similar to strategies of humans. Moreover, we address known challenges of reinforcement learning in this domain and present an approach that uses function approximation, an opponent pool, and an autoregressive policy representation. Through evaluations we show that our method is superior to two baseline methods but that policy convergence in self-play remains a challenge.
Kim Hammar, Rolf Stadler
CNSM2
2020 Online feature selection for rapid, low-overhead learning in networked systems
abstract
Data-driven functions for operation and management often require measurements collected through monitoring for model training and prediction. The number of data sources can be very large, which requires a significant communication and computing overhead to continuously extract and collect this data, as well as to train and update the machine-learning models. We present an online algorithm, called OSFS, that selects a small feature set from a large number of available data sources, which allows for rapid, low-overhead, and effective learning and prediction. OSFS is instantiated with a feature ranking algorithm and applies the concept of a stable feature set, which we introduce in the paper. We perform extensive, experimental evaluation of our method on data from an in-house testbed. We find that OSFS requires several hundreds measurements to reduce the number of data sources by two orders of magnitude, from which models are trained with acceptable prediction accuracy. While our method is heuristic and can be improved in many ways, the results clearly suggests that many learning tasks do not require a lengthy monitoring phase and expensive offline training.
Forough Shahab Samani, Rolf Stadler
CNSM3
2020 Guest Editorial Special Issue on Advances in Artificial Intelligence and Machine Learning for Networking
abstract
https://www.youtube.com/watch?v=SQmgSOi5oos
Prosper Chemouil, Pan Hui 0001, Wolfgang Kellerer, Noura Limam, Rolf Stadler, Yonggang Wen 0001
IEEE J. Sel. Areas Commun.5
2019 Efficient Learning on High-dimensional Operational Data
abstract
In networked systems engineering, operational data gathered from sensors or logs can be used to build data-driven functions for performance prediction, anomaly detection, and other operational tasks. The number of data sources used for this purpose determines the dimensionality of the feature space for learning and can reach millions for medium-sized systems. Learning on a space with high dimensionality generally incurs high communication and computational costs for the learning process. In this work, we apply and compare a range of methods, including, feature selection, Principle Component Analysis (PCA), and autoencoders with the objective to reduce the dimensionality of the feature space while maintaining the prediction accuracy when compared with learning on the full space. We conduct the study using traces gathered from a testbed at KTH that runs a video-on-demand service and a key-value store under dynamic load. Our results suggest the feasibility of reducing the dimensionality of the feature space of operational data significantly, by one to two orders of magnitude in our scenarios, while maintaining prediction accuracy. The findings confirm the Manifold Hypothesis in machine learning, which states that real-world data sets tend to occupy a small subspace of the full feature space. In addition, we investigate the tradeoff between prediction accuracy and prediction overhead, which is crucial for applying the results to operational systems.
Forough Shahab Samani, Rolf Stadler
CNSM3
2019 Performance Prediction in Dynamic Clouds using Transfer Learning
Andreas Johnsson, Farnaz Moradi 0001, Rolf Stadler
IM3
2019 Demonstration: Predicting Distributions of Service Metrics
Forough Shahab Samani, Rolf Stadler, Andreas Johnsson, Christofer Flinta
IM2
2019 Special Issue on Artificial Intelligence and Machine Learning for Networking and Communications
abstract
Research in large-scale networking systems has been shaped and will continue to be guided by specific characteristics of applications and the underlying platforms and infrastructures. On the one hand, applications are growing at an accelerated pace, which is fundamentally unpredictable in both breadth and depth. On the other hand, the underlying networking has been the focus of a huge transformation enabled by new models resulting from virtualization and cloud computing. This has led to a number of novel architectures supported by emerging technologies such as Software-Defined Networking (SDN), Network Function Virtualization (NFV), and more recently, edge cloud and fog networking, or network slicing[1],[2]. This evolution towards enhanced design opportunities along with increasing complexity in networking and its applications has fueled the need for improved network automation in agile infrastructures. At the same time, their complexity has dramatically increased. The networking dynamics have had the effect of making it even more important and challenging to design scalable network measurement and analysis techniques and associated tools. Critical applications such as resource allocation, network monitoring, security enforcement, or dynamic network management require real-time mechanisms for online analysis as well as efficient techniques for offline deep analysis of massive historical data.
Prosper Chemouil, Pan Hui 0001, Wolfgang Kellerer, Yong Li 0008, Rolf Stadler, Dacheng Tao, Yonggang Wen 0001, Ying Zhang 0022
IEEE J. Sel. Areas Commun.5
2018 Predicting Distributions of Service Metrics using Neural Networks
Forough Shahab Samani, Rolf Stadler
CNSM2
2018 Automated diagnostic of virtualized service performance degradation
abstract
Service assurance for cloud applications is a challenging task and is an active area of research for academia and industry. One promising approach is to utilize machine learning for service quality prediction and fault detection so that suitable mitigation actions can be executed. In our previous work, we have shown how to predict service-level metrics in real-time just from operational data gathered at the server side. This gives the service provider early indications on whether the platform can support the current load demand. This paper provides the logical next step where we extend our work by proposing an automated detection and diagnostic capability for the performance faults manifesting themselves in cloud and datacenter environments. This is a crucial task to maintain the smooth operation of running services and minimizing downtime. We demonstrate the effectiveness of our approach which exploits the interpretative capabilities of Self- Organizing Maps (SOMs) to automatically detect and localize different performance faults for cloud services.
Jawwad Ahmed, Tim Josefsson, Andreas Johnsson, Christofer Flinta, Farnaz Moradi 0001, Rafael Pasquini, Rolf Stadler
NOMS7
2017 Online approach to performance fault localization for cloud and datacenter services
abstract
Automated detection and diagnosis of the performance faults in cloud and datacenter environments is a crucial task to maintain smooth operation of different services and minimize downtime. We demonstrate an effective machine learning approach based on detecting metric correlation stability violations (CSV) for automated localization of performance faults for datacenter services running under dynamic load conditions.
Jawwad Ahmed, Andreas Johnsson, Farnaz Moradi 0001, Rafael Pasquini, Christofer Flinta, Rolf Stadler
IM6
2017 Real-time resource prediction engine for cloud management
abstract
Predicting resource requirements for cloud services is critical for dimensioning, anomaly detection and service assurance. We demonstrate a system for real-time estimation of the needed amount of infrastructure resources, such as CPU and memory, for a given service. Statistical learning methods on server statistics and load parameters of the service are used for learning a resource prediction model. The model can be used as a guideline for service deployment and for real-time identification of resource bottlenecks.
Christofer Flinta, Andreas Johnsson, Jawwad Ahmed, Farnaz Moradi 0001, Rafael Pasquini, Rolf Stadler
IM6
2017 Learning end-to-end application QoS from openflow switch statistics
abstract
We use statistical learning to estimate end-to-end QoS metrics from device statistics, collected from a server cluster and an OpenFlow network. The results from our testbed, which runs a video-on-demand service and a key-value store, demonstrate that the learned models can estimate QoS metrics like frame rate or response time with errors bellow 10% for a given client. Interestingly, we find that service-level QoS metrics seem "encoded" in network statistics and it suffices to collect OpenFlow per port statistics to achieve accurate estimation at small overhead for data collection and model computation.
Rafael Pasquini, Rolf Stadler
NetSoft2
2017 2016 Reviewers for IEEE Transactions on Network and Service Management
abstract
The success and quality of this journal critically depends on the dedication and expertise of a large number of reviewers. On behalf of the Editorial Board, I would like to thank them for their excellent work.
Rolf Stadler
IEEE Trans. Netw. Serv. Manag.1
2016 Predicting SLA conformance for cluster-based services using distributed analytics
abstract
Service assurance for the telecom cloud is a challenging task and is continuously being addressed by academics and industry. One promising approach is to utilize machine learning to predict service quality in order to take early mitigation actions. In previous work we have shown how to predict service-level metrics, such as frame rate for a video application on the client side, from operational data gathered at the server side. This gives the service provider early indications on whether the platform can support the current load demand. This paper extends previous work by addressing scalability issues for cluster-based services. Operational data being generated in large volumes, from several sources, and at high velocity puts strain on computational and communication resources. We propose and evaluate a distributed machine learning system based on the Winnow algorithm to tackle scalability issues, and then compare the new distributed solution with the previously proposed centralized solution. We show that network overhead and computational execution time is substantially reduced while maintaining high prediction accuracy making it possible to achieve real-time service quality predictions in large systems.
Jawwad Ahmed, Andreas Johnsson, Rerngvit Yanggratoke, John Ardelius, Christofer Flinta, Rolf Stadler
NOMS6
2016 A bottom-up approach to real-time search in large networks and clouds
abstract
Networked systems, such as telecom networks and cloud infrastructures, generate and hold vast amounts of configuration and operational data. The goal of this work is to make all this data available through a real-time search process named network search, which will enable new real-time management solutions. The thesis contains several contributions towards engineering a network search system. Key elements of our design are a weakly structured information model that includes spatial properties, a query language that supports location- and schema-oblivious search queries, a peer-to-peer architecture, an echo protocols for scalable query processing, and an indexing protocol for efficient routing for spatial queries. The data against which network search is performed is maintained in local real-time databases close to the data sources. The design follows a bottom-up approach in the sense that the topology for query routing is constructed from the underlying network topology. We have built a prototype of the system on a cloud testbed and developed applications that use network search functionality. Testbed measurements suggest that it is feasible to engineer a network search system that processes queries at low latency and low overhead and that can scale to 100'000 nodes. Simulation results for spatial queries show that query processing achieves response times and incurs overhead close to an optimal protocol, and that it remains accurate under significant churn.
Misbah Uddin, Rolf Stadler
NOMS2
2015 Universal fault detection for NFV using SOM-based clustering
abstract
Network function virtualization (NFV) introduces additional complexity to network management, since the placement and behavior of virtualized network functions (VNFs) can be independent from the underlying hardware, and virtualization technology increases the number of monitoring points and the amount of statistical data. In our previous work, we proposed a framework for detecting anomalous behavior of VNFs using a SOM-based technique. The solution relies upon manually configuring the SOM clustering parameters and selecting the statistics for each failure type in advance, which results in a high maintenance load. In this paper, we provide a solution that is universal in the sense that a range of different faults can be detected using a single set of local statistics and SOM clustering parameters. Experimental results from a testbed show that faults, including memory leak, packet congestion, and session congestion, can be detected with high accuracy using only four types of performance statistics.
Tomonobu Niwa, Masanori Miyazawa, Michiaki Hayashi, Rolf Stadler
APNOMS4
2015 Spatial search in networked systems
abstract
Information in networked systems often has spatial properties: routers, sensors, or virtual machines have coordinates in a geographical or virtual space, for instance. In this paper, we propose a peer-to-peer design for a spatial search system that processes queries, such as range or nearest-neighbor queries, on spatial information cached on nodes inside a networked system. Key to our design is a protocol that creates a distributed index of object locations and adapts to object and node churn. The index builds upon the concept of the minimum bounding rectangle, to efficiently encode a large set of locations. We present a search protocol, which is based on an echo protocol and performs query routing. Simulations show the efficiency of the protocol in pruning the search space, thereby reducing the protocol overhead. For many queries, the protocol efficiency increases with the network size and approaches that of an optimal protocol for large systems. The protocol overhead depends on the network topology and is lower if neighboring nodes are spatially close. As a key difference to works in spatial databases, our design is bottom-up, which makes query routing network-aware and thus efficient in networked systems.
Misbah Uddin, Rolf Stadler, Alexander Clemm
CNSM2
2015 Predicting service metrics for cluster-based services using real-time analytics
abstract
Predicting the performance of cloud services is intrinsically hard. In this work, we pursue an approach based upon statistical learning, whereby the behaviour of a system is learned from observations. Specifically, our testbed implementation collects device statistics from a server cluster and uses a regression method that accurately predicts, in real-time, client-side service metrics for a video streaming service running on the cluster. The method is service-agnostic in the sense that it takes as input operating-systems statistics instead of service-level metrics. We show that feature set reduction significantly improves prediction accuracy in our case, while simultaneously reducing model computation time. We also discuss design and implementation of a real-time analytics engine, which processes streams of device statistics and service metrics from testbed sensors and produces model predictions through online learning.
Rerngvit Yanggratoke, Jawwad Ahmed, John Ardelius, Christofer Flinta, Andreas Johnsson, Daniel Gillblad, Rolf Stadler
CNSM7
2015 vNMF: Distributed fault detection using clustering approach for network function virtualization
abstract
Network function virtualization introduces additional complexity for network management through the use of virtualization environments. The amount of managed data and the operational complexity increases, which makes service assurance and failure recovery harder to realize. In response to this challenge, the paper proposes a distributed management function, called virtualized network management function (vNMF), to detect failures related to virtualized services. vNMF detects the failures by monitoring physical-layer statistics that are processed with a self-organizing map algorithm. Experimental results show that memory leaks and network congestion failures can be successfully detected and that and the accuracy of failure detection can be significantly improved compared to common k-means clustering.
Masanori Miyazawa, Michiaki Hayashi, Rolf Stadler
IM3
2015 Predicting real-time service-level metrics from device statistics
abstract
While real-time service assurance is critical for emerging telecom cloud services, understanding and predicting performance metrics for such services is hard. In this paper, we pursue an approach based upon statistical learning whereby the behavior of the target system is learned from observations. We use methods that learn from device statistics and predict metrics for services running on these devices. Specifically, we collect statistics from a Linux kernel of a server machine and predict client-side metrics for a video-streaming service (VLC). The fact that we collect thousands of kernel variables, while omitting service instrumentation, makes our approach service-independent and unique. While our current lab configuration is simple, our results, gained through extensive experimentation, prove the feasibility of accurately predicting client-side metrics, such as video frame rates and RTP packet rates, often within 10-15% error (NMAE), also under high computational load and across traces from different scenarios.
Rerngvit Yanggratoke, Jawwad Ahmed, John Ardelius, Christofer Flinta, Andreas Johnsson, Daniel Gillblad, Rolf Stadler
IM7
2015 A platform for predicting real-time service-level metrics from device statistics
abstract
Predicting performance metrics for cloud services is critical for real-time service assurance. We demonstrate a platform for estimating real-time service-level metrics. Statistical learning methods on device statistics are used to predict metrics for services running on these devices.
Rerngvit Yanggratoke, Jawwad Ahmed, John Ardelius, Christofer Flinta, Andreas Johnsson, Daniel Gillblad, Rolf Stadler
IM7
2014 Graph search for cloud network management
abstract
A large part of operational and configuration information in networks and clouds has graph structure, e.g., virtual network topologies, IP flows, communication links of distributed cloud applications. The final objective of this work is to develop a scalable management system that allows real-time management applications, such as network analytics and anomaly detection applications, to search for graph-structured operational information. The paper contains a step towards this objective. The key challenge is to devise an efficient and scalable search process on data that is volatile and distributed across the network infrastructure. Approaches that are currently pursued for distributed graph databases are not applicable in this context. This paper presents design options and possible solutions for engineering a scalable graph search system that supports management applications. It includes a simple information model based on property graphs and a search query language based on keyword search. The architecture of the system centers around a distributed search plane that performs query processing using a network of search nodes. Finally, the paper outlines the design of a search node, which contains a local database that maintains graph partitions.
Misbah Uddin, Rolf Stadler, Masanori Miyazawa, Michiaki Hayashi
NOMS2
2013 Scalable matching and ranking for network search
abstract
Network search makes operational data available in real-time to management applications. In contrast to traditional monitoring, neither the data location nor the data format needs to be known to the invoking process, which simplifies application development, but requires an efficient search plane inside the managed system. The search plane is realized as a network of search nodes that process search queries in a distributed fashion. This paper introduces matching and ranking for network search queries. We are proposing a semantic for matching and ranking, which is configurable to support different types of management applications-from exact matching for database-style queries to loose, approximate matching, which is appropriate for exploratory purposes. We describe an echo protocol for efficient distributed query processing that supports matching and ranking. Further, we present the design of a search node, which maintains a real-time database of operational information and allows for parallel processing of search queries. A prototype implementation on a cloud testbed shows that the network search system, on a 9-node cluster with 24 core servers, executes 200 global search queries/sec with the 75th percetile latency below 100 milliseconds and with a CPU utilization below 5%. The performance measurements, together with our design, suggest that a system of 100,000 servers processing the same load would exhibit the same overhead per server and a query latency of below 1 sec.
Misbah Uddin, Rolf Stadler, Alexander Clemm
CNSM2
2013 A query language for network search
Misbah Uddin, Rolf Stadler, Alexander Clemm
IM2
2013 Real-time search in clouds
Misbah Uddin, Amy Skinner, Rolf Stadler, Alexander Clemm
IM3
2012 Dynamic resource allocation with management objectives - Implementation for an OpenStack cloud
Fetahi Zebenigus Wuhib, Rolf Stadler, Hans Lindgren
CNSM2
2012 Predicting response times for the Spotify backend
Rerngvit Yanggratoke, Gunnar Kreitz, Mikael Goldmann, Rolf Stadler
CNSM4
2012 Management by network search
abstract
While networked systems hold and generate vast amounts of configuration and operational data, this data is not accessible through a simple, uniform mechanism. Rather, it must be gathered using a range of different protocols and interfaces. Our vision is to make all this data available in a simple format through a realtime search process which runs within the network and aggregates the data into a form needed by applications - a concept we call network search. We believe that such an approach, though challenging, is technically feasible and will enable rapid development of new management applications and advanced network functions. This paper motivates and formulates the concept of network search, compares it to related concepts like web search, outlines a search architecture, describes the design space and research challenges, and reports on a testbed implementation with management applications built for exploratory purposes of this new paradigm.
Misbah Uddin, Rolf Stadler, Alexander Clemm
NOMS2
2012 A Gossip Protocol for Dynamic Resource Management in Large Cloud Environments
abstract
We address the problem of dynamic resource management for a large-scale cloud environment. Our contribution includes outlining a distributed middleware architecture and presenting one of its key elements: a gossip protocol that (1) ensures fair resource allocation among sites/applications, (2) dynamically adapts the allocation to load changes and (3) scales both in the number of physical machines and sites/applications. We formalize the resource allocation problem as that of dynamically maximizing the cloud utility under CPU and memory constraints. We first present a protocol that computes an optimal solution without considering memory constraints and prove correctness and convergence properties. Then, we extend that protocol to provide an efficient heuristic solution for the complete problem, which includes minimizing the cost for adapting an allocation. The protocol continuously executes on dynamic, local input and does not require global synchronization, as other proposed gossip protocols do. We evaluate the heuristic protocol through simulation and find its performance to be well-aligned with our design goals.
Fetahi Zebenigus Wuhib, Rolf Stadler, Mike Spreitzer
IEEE Trans. Netw. Serv. Manag.2
2011 Gossip-based resource allocation for green computing in large clouds
Rerngvit Yanggratoke, Fetahi Zebenigus Wuhib, Rolf Stadler
CNSM3
2011 Distributed monitoring and resource management for large cloud environments
abstract
Over the last decade, the number, size and complexity of large-scale networked systems has been growing fast, and this trend is expected to accelerate. The best known example of a large-scale networked system is probably the Internet, while large datacenters for cloud services are the most recent ones. In such environments, a key challenge is to develop scalable and adaptive technologies for management functions. This thesis addresses the challenge by engineering several protocols for distributed monitoring and resource management that are suitable for large-scale networked systems. The protocols are evaluated through theoretical analysis, simulation studies and testbed experimentation. The evaluation results show that the protocols achieve their respective design objectives with respect to quality, efficiency, scalability, controllability and adaptability.
Fetahi Zebenigus Wuhib, Rolf Stadler
Integrated Network Management2
2010 Gossip-based resource management for cloud environments
abstract
We address the problem of resource management for a large-scale cloud environment that hosts sites. Our contribution centers around outlining a distributed middleware architecture and presenting one of its key elements, a gossip protocol that meets our design goals: fairness of resource allocation with respect to hosted sites, efficient adaptation to load changes and scalability in terms of both the number of machines and sites. We formalize the resource allocation problem as that of dynamically maximizing the cloud utility under CPU and memory constraints. While we can show that an optimal solution without considering memory constraints is straightforward (but not useful), we provide an efficient heuristic solution for the complete problem instead. We evaluate the protocol through simulation and find its performance to be well-aligned with our design goals.
Fetahi Zebenigus Wuhib, Rolf Stadler, Mike Spreitzer
CNSM2
2010 Brief announcement: the accuracy of tree-based counting in dynamic networks
abstract
We study a simple Bellman-Ford-like protocol which performs network size estimation over a tree-shaped overlay. A continuous time Markov model is constructed which allows key protocol characteristics to be estimated under churn, including the expected number of nodes at a given (perceived) distance to the root and, for each such node, the expected (perceived) size of the subnetwork rooted at that node. We validate the model by simulations, using a range of network sizes, node degrees, and churn-to-protocol rates, with convincing results.
Supriya Krishnamurthy, John Ardelius, Erik Aurell, Mads Dam, Rolf Stadler, Fetahi Zebenigus Wuhib
PODC5
2010 Distributed auto-configuration of neighboring cell graphs in radio access networks
abstract
In order to execute a handover processes in a GSM or UMTS Radio Access Network, each cell has a list of neighbors to which such handovers may be made. Today, these lists are statically configured during network planning, which does not allow for dynamic adaptation of the network to changes and unexpected events such as a cell failure. This paper advocates an autonomic, decentralized approach to dynamically configure neighboring cell lists. The main contribution of this work is a novel protocol, called DOC, which detects and continuously tracks the coverage overlaps among cells. The protocol executes on a spanning tree where the nodes are radio base stations and the links represent communication channels. Over this tree, nodes periodically exchange information about terminals that are in their respective coverage area. Bloom filters are used for efficient representations of terminal sets and efficient set operations. The protocol aggregates Bloom filters to reduce the communication overhead and also for routing messages along the tree. Using simulation, we study the system in steady state, when a base station is added or a base station fails, and also during the initialization phase where the system self-configures.
Javier Baliosian, Rolf Stadler
IEEE Trans. Netw. Serv. Manag.2
2010 H-GAP: estimating histograms of local variables with accuracy objectives for distributed real-time monitoring
abstract
We present H-GAP, a protocol for continuous monitoring, which provides a management station with the value distribution of local variables across the network. The protocol estimates the histogram of local state variables for a given accuracy and with minimal overhead. H-GAP is decentralized and asynchronous to achieve robustness and scalability, and it executes on an overlay interconnecting management processes in network devices. On this overlay, the protocol maintains a spanning tree and updates the histogram through incremental aggregation. The protocol is tunable in the sense that it allows controlling, at runtime, the trade-off between protocol overhead and an accuracy objective. This functionality is realized through dynamic configuration of local filters that control the flow of updates towards the management station. The paper includes an analysis of the problem of histogram aggregation over aggregation trees, a formulation of the global optimization problem, and a distributed solution containing heuristic, tree-based algorithms. Using SUM as an example, we show how general aggregation functions over local variables can be efficiently computed with H-GAP. We evaluate our protocol through simulation using real traces. The results demonstrate the controllability of H-GAP in a selection of scenarios and its efficiency in large-scale networks.
Dan Jurca, Rolf Stadler
IEEE Trans. Netw. Serv. Manag.2
2010 A gossiping protocol for detecting global threshold crossings
abstract
We investigate the use of gossip protocols for the detection of network-wide threshold crossings. Our design goals are low protocol overhead, small detection delay, low probability of false positives and negatives, scalability, robustness to node failures and controllability of the trade-off between overhead and detection delay. Based on push-synopses, a gossip protocol introduced by Kempe et al., we present a protocol that indicates whether a global aggregate of static local values is above or below a given threshold. For this protocol, we prove correctness and show that it converges to a state with no overhead when the aggregate is sufficiently far from the threshold. Then, we introduce an extension we call TG-GAP, a protocol that (1) executes in a dynamic network environment where local values change and (2) implements hysteresis behavior with upper and lower thresholds. Key elements of its design are the construction of snapshots of the global aggregate for threshold detection and a mechanism for synchronizing local states, both of which are realized through the underlying gossip protocol. Simulation studies suggest that TG-GAP is efficient in that the protocol overhead is minimal when the aggregate is sufficiently far from the threshold, that its overhead and the detection delay are largely independent on the system size, and that the tradeoff between overhead and detection quality can be effectively controlled. Lastly, we perform a comparative evaluation of TG-GAP against a tree-based protocol. We conclude that, for detecting global threshold crossings in the type of scenarios investigated, the tree-based protocol incurs a significantly lower overhead and a smaller detection delay than a gossip protocol such as TG-GAP.
Fetahi Zebenigus Wuhib, Mads Dam, Rolf Stadler
IEEE Trans. Netw. Serv. Manag.3
2009 Computing histograms of local variables for real-time monitoring using aggregation trees
abstract
In this paper we present a protocol for the continuous monitoring of a local network state variable. Our aim is to provide a management station with the value distribution of the local variables across the network, by means of partial histogram aggregation, with minimum protocol overhead. Our protocol is decentralized and asynchronous to achieve robustness and scalability, and it executes on an overlay interconnecting management processes in network devices. On this overlay, the protocol maintains a spanning tree and updates the histogram of the network state variables through incremental aggregation. The protocol allows to control the trade-off between protocol overhead and a global accuracy objective. This functionality is implemented by a dynamic configuration of local error filters that control whether an update is sent towards the management station or not. We evaluate our protocol by means of simulations. Our results demonstrate the controllability of our method in a wide selection of scenarios, and the scalability of our protocol for large-scale networks.
Dan Jurca, Rolf Stadler
Integrated Network Management2
2009 Controlling performance trade-offs in adaptive network monitoring
abstract
A key requirement for autonomic (i.e., self-*) management systems is a short adaptation time to changes in the networking conditions. In this paper, we show that the adaptation time of a distributed monitoring protocol can be controlled. We show this for A-GAP, a protocol for continuous monitoring of global metrics with controllable accuracy. We demonstrate through simulations that, for the case of A-GAP, the choice of the topology of the aggregation tree controls the trade-off between adaptation time and protocol overhead in steady-state. Generally, allowing a larger adaptation time permits reducing the protocol overhead. Our results suggest that the adaptation time primarily depends on the height of the aggregation tree and that the protocol overhead is strongly influenced by the number of internal nodes. We outline how A-GAP can be extended to dynamically self-configure and to continuously adapt its configuration to changing conditions, in order to meet a set of performance objectives, including adaptation time, protocol overhead, and estimation accuracy.
Alberto Gonzalez Prieto, Rolf Stadler
Integrated Network Management2
2009 Adaptive real-time monitoring for large-scale networked systems
abstract
The focus of this thesis is continuous real-time monitoring, which is essential for the realization of adaptive management systems in large-scale dynamic environments. Real-time monitoring provides the necessary input to the decision-making process of network management. We have developed, implemented, and evaluated a design for real-time continuous monitoring of global metrics with performance objectives, such as monitoring overhead and estimation accuracy. Global metrics describe the state of the system as a whole, in contrast to local metrics, such as device counters or local protocol states, which capture the state of a local entity. Global metrics are computed from local metrics using aggregation functions, such as SUM, AVERAGE and MAX. A key part in the design is a model for the distributed monitoring process that relates performance metrics to parameters that tune the behavior of a monitoring protocol. The model has been instrumental in designing a monitoring protocol that is controllable and achieves given performance objectives. Our design has proved to be effective in meeting performance objectives, efficient, adaptive to changes in the networking conditions, controllable along different performance dimensions, and scalable. We have implemented a prototype on a testbed of commercial routers, which proves the feasibility of the design, and, more generally, the feasibility of effective and efficient real-time monitoring in large network environments.
Alberto Gonzalez Prieto, Rolf Stadler
Integrated Network Management2
2009 Gossiping for threshold detection
abstract
We investigate the use of gossip protocols to detect threshold crossings of network-wide aggregates. Aggregates are computed from local device variables using functions such as SUM, AVERAGE, COUNT, MAX and MIN. The process of aggregation and detection is performed using a standard gossiping scheme. A key design element is to let nodes dynamically adjust their neighbor interaction rates according to the distance between the nodes' local estimate of the global aggregate and the threshold itself. We show that this allows considerable savings in communication overhead. In particular, the overhead becomes negligible when the aggregate is sufficiently far above or far below the threshold. We present evaluation results from simulation studies regarding protocol efficiency, quality of threshold detection, scalability, and controllability.
Fetahi Zebenigus Wuhib, Rolf Stadler, Mads Dam
Integrated Network Management2
2009 Robust monitoring of network-wide aggregates through gossiping
abstract
We investigate the use of gossip protocols for continuous monitoring of network-wide aggregates under crash failures. Aggregates are computed from local management variables using functions such as SUM, MAX, or AVERAGE. For this type of aggregation, crash failures offer a particular challenge due to the problem of mass loss, namely, how to correctly account for contributions from nodes that have failed. In this paper we give a partial solution. We present G-GAP, a gossip protocol for continuous monitoring of aggregates, which is robust against failures that are discontiguous in the sense that neighboring nodes do not fail within a short period of each other. We give formal proofs of correctness and convergence, and we evaluate the protocol through simulation using real traces. The simulation results suggest that the design goals for this protocol have been met. For instance, the tradeoff between estimation accuracy and protocol overhead can be controlled, and a high estimation accuracy (below some 5% error in our measurements) is achieved by the protocol, even for large networks and frequent node failures. Further, we perform a comparative assessment of GGAP against a tree-based aggregation protocol using simulation. Surprisingly, we find that the tree-based aggregation protocol consistently outperforms the gossip protocol for comparative overhead, both in terms of accuracy and robustness.
Fetahi Zebenigus Wuhib, Mads Dam, Rolf Stadler, Alexander Clemm
IEEE Trans. Netw. Serv. Manag.3
2008 Policy-based self-healing for radio access networks
abstract
Various centralized, distributed or cooperative management systems have been proposed to address the demands of wireless telecommunication networks. However, considering the size, complexity and heterogeneity that those networks will have in the future, current solutions either do not scale properly, or have no support for automation, or lack of the flexibility and simple control that operators will need for managing future networks in a cost-effective way. To address this problem, we designed Omega, a distributed and policy-based network management system that uses rich knowledge-modeling techniques to develop self-configuration capabilities. Omega also implements a novel conflict-resolution method that uses high-level goals and machine learning techniques to optimize its policy-based decisions. Using simulations, in this paper we show how Omega reduces the impact of a node crash on the overall availability of a radio access network by optimizing the lists of neighboring cells of the nodes in the vicinity.
Javier Baliosian, Katarína Matusíková, Karl Quinn, Rolf Stadler
NOMS4
2008 Decentralized detection of global threshold crossings using aggregation trees
Fetahi Zebenigus Wuhib, Mads Dam, Rolf Stadler
Comput. Networks3
2007 A Service Middleware that Scales in System Size and Applications
abstract
We present a peer-to-peer service management middleware that dynamically allocates system resources to a large set of applications. The system achieves scalability in number of nodes (1000s or more) through three decentralized mechanisms that run on different time scales. First, overlay construction interconnects all nodes in the system for exchanging control and state information. Second, request routing directs requests to nodes that offer the corresponding applications. Third, application placement controls the set of offered applications on each node, in order to achieve efficient operation and service differentiation. The design supports a large number of applications (100s or more) through selective propagation of configuration information needed for request routing. The control load on a node increases linearly with the number of applications in the system. Service differentiation is achieved through assigning a utility to each application, which influences the application placement process. Simulation studies show that the system operates efficiently for different sizes, adapts fast to load changes and failures and effectively differentiates between different applications under overload.
Constantin Adam, Rolf Stadler, Chunqiang Tang, Malgorzata Steinder, Mike Spreitzer
Integrated Network Management2
2007 Robust Monitoring of Network-wide Aggregates through Gossiping
abstract
We examine the use of gossip protocols for continuous monitoring of network-wide aggregates. Aggregates are computed from local management variables using functions such as AVERAGE, MIN, MAX, or SUM. A particular challenge is to develop a gossip-based aggregation protocol that is robust against node failures. In this paper, we present G-GAP, a gossip protocol for continuous monitoring of aggregates, which is robust against discontiguous failures (i.e., under the constraint that neighboring nodes do not fail within a short period of each other). We formally prove this property, and we evaluate the protocol through simulation using real traces. The simulation results suggest that the design goals for this protocol have been met. For instance, the tradeoff between estimation accuracy and protocol overhead can be controlled, and a high estimation accuracy (below some 5% error in our measurements) is achieved by the protocol, even for large networks and frequent node failures. Further, we perform a comparative assessment of G-GAP against a tree-based aggregation protocol using simulation. Surprisingly, we find that the tree-based aggregation protocol consistently outperforms the gossip protocol for comparative overhead, both in terms of accuracy and robustness.
Fetahi Zebenigus Wuhib, Mads Dam, Rolf Stadler, Alexander Clemm
Integrated Network Management3
2007 Network Patterns in Cfengine and Scalable Data Aggregation
Mark Burgess, Matthew Disney, Rolf Stadler
LISA3
2007 Decentralized Configuration of Neighboring Cells for Radio Access Networks
abstract
In order to execute a handover processes in a Radio Access Network, each cell has a configured list of neighbors to which such handovers are made. Rapid re-configuration ofthe neighborhood list in response to network failures and other events is currently not possible. To address this problem this paper suggests an autonomic approach for dynamically configuring neighboring cell lists and introduces a decentralized, three-layered framework. As a key element of this framework, a novel probabilistic protocol that detects and continuously tracks the coverage overlaps among cells is presented and evaluated. The protocol called DOC maintains a distributed graph of over-lapping cells. Dae to asing Bloom filters and aggregation techniques it exhibits a low traffic and computational overhead. A first series of simulation studies suggests that DOC is scalable with respect to network size and the namber of terminals.
Javier Baliosian, Rolf Stadler
WOWMOM2
2007 Service Middleware for Self-Managing Large-Scale Systems
abstract
Resource management poses particular challenges in large-scale systems, such as server clusters that simultaneously process requests from a large number of clients. A resource management scheme for such systems must scale both in the in the number of cluster nodes and the number of applications the cluster supports. Current solutions do not exhibit both of these properties at the same time. Many are centralized, which limits their scalability in terms of the number of nodes, or they are decentralized but rely on replicated directories, which also reduces their ability to scale. In this paper, we propose novel solutions to request routing and application placement- two key mechanisms in a scalable resource management scheme. Our solution to request routing is based on selective update propagation, which ensures that the control load on a cluster node is independent of the system size. Application placement is approached in a decentralized manner, by using a distributed algorithm that maximizes resource utilization and allows for service differentiation under overload. The paper demonstrates how the above solutions can be integrated into an overall design for a peer-to-peer management middleware that exhibits properties of self-organization. Through complexity analysis and simulation, we show to which extent the system design is scalable. We have built a prototype using accepted technologies and have evaluated it using a standard benchmark. The testbed measurements show that the implementation, within the parameter range tested, operates efficiently, quickly adapts to a changing environment and allows for effective service differentiation by a system administrator.
Constantin Adam, Rolf Stadler
IEEE Trans. Netw. Serv. Manag.2
2007 A-GAP: An Adaptive Protocol for Continuous Network Monitoring with Accuracy Objectives
abstract
We present A-GAP, a novel protocol for continuous monitoring of network state variables, which aims at achieving a given monitoring accuracy with minimal overhead. Network state variables are computed from device counters using aggregation functions, such as SUM, AVERAGE and MAX. The accuracy objective is expressed as the average estimation error. A-GAP is decentralized and asynchronous to achieve robustness and scalability. It executes on an overlay that interconnects management processes on the devices. On this overlay, the protocol maintains a spanning tree and updates the network state variables through incremental aggregation. Based on a stochastic model, it dynamically configures local filters that control whether an update is sent towards the root of the tree. We evaluate A-GAP through simulation using real traces and two different types of topologies of up to 650 nodes. The results show that we can effectively control the trade-off between accuracy and protocol overhead, and that the overhead can be reduced by almost two orders of magnitude when allowing for small errors. The protocol quickly adapts to a node failure and exhibits short spikes in the estimation error. Lastly, it can provide an accurate estimate of the error distribution in real-time.
Alberto Gonzalez Prieto, Rolf Stadler
IEEE Trans. Netw. Serv. Manag.2
2006 Distributed Real-Time Monitoring with Accuracy Objectives
Alberto Gonzalez Prieto, Rolf Stadler
Networking2
2006 A middleware design for large-scale clusters offering multiple services
abstract
We present a decentralized design that dynamically allocates resources to multiple services inside a global server cluster. The design supports QoS objectives (maximum response time and maximum loss rate) for each service. A system administrator can modify policies that assign relative importance to services and, in this way, control the resource allocation process. Distinctive features of our design are the use of an epidemic protocol to disseminate state and control information, as well as the decentralized evaluation of utility functions to control resource partitioning among services. Simulation results show that the system operates both effectively and efficiently; it meets the QoS objectives and dynamically adapts to load changes and to failures. In case of overload, the service quality degrades gracefully, controlled by the cluster policies.
Constantin Adam, Rolf Stadler
IEEE Trans. Netw. Serv. Manag.2
2005 Adaptable server clusters with QoS objectives
abstract
We present a decentralized design for a server cluster that supports a single service with response time guarantees. Three distributed mechanisms represent the key elements of our design. Topology construction maintains a dynamic overlay of cluster nodes. Request routing directs service requests towards available servers. Membership control allocates/releases servers to/from the cluster, in response to changes in the external load. We advocate a decentralized approach, because it is scalable, fault-tolerant, and has a lower configuration complexity than a centralized solution. We demonstrate through simulations that our system operates efficiently by comparing it to an ideal centralized system. In addition, we show that our system rapidly adapts to changing load. We found that the interaction of the various mechanisms in the system leads to desirable global properties. More precisely, for a fixed connectivity c (i.e., the number of neighbors of a node in the overlay), the average experienced delay in the cluster is independent of the external load. In addition, increasing c increases the average delay but decreases the system size for a given load. Consequently, the cluster administrator can use c as a management parameter that permits control of the tradeoff between a small system size and a small experienced delay for the service.
Constantin Adam, Rolf Stadler
Integrated Network Management2
2005 Real-time views of network traffic using decentralized management
abstract
The ability to create views of a network on a fast time scale becomes increasingly important as the complexity and diversity of networks increase. These views, which combine information from many distributed points in the network, can provide an administrator with a better understanding of the interdependencies and interactions between network elements and traffic conditions. Applications that could benefit from being able to compute such "near" real-time views of the network range from performance monitoring to fault management. In this paper, we present the architecture of a distributed management infrastructure that enables such views to be computed. Based on our earlier work on decentralized management, our architecture takes a novel database approach that combines the expressive power of SQL with distributed algorithms. We describe the implementation of the system on platform of embedded Linux devices attached to a network of routers. We provide specific examples of how the system can be used as a powerful distributed real-time monitoring platform. Finally, we derive a performance model of the system and validate it with a set of experiments.
Koon-Seng Lim, Rolf Stadler
Integrated Network Management2
2004 Patterns for routing and self-stabilization
abstract
The paper contributes towards engineering self-stabilizing networks and services. We propose the use of navigation patterns, which define how information for state updates is disseminated in the system, as fundamental building blocks for self-stabilizing systems. We present two navigation patterns for self-stabilization: the progressive wave pattern and the stationary wave pattern. The progressive wave pattern defines the update dissemination in Internet routing systems running the DUAL and OSPF (open shortest path first) protocols. Similarly, the stationary wave pattern defines the interactions of peer nodes in structured peer-to-peer systems, including Chord, Pastry, Tapestry, and CAN. It turns out that the two patterns are related. They both disseminate information in the form of waves, i.e, sets of messages that originate from single events. Patterns can be instrumented to obtain wave statistics, which enables monitoring the process of self-stabilization in a system. We focus on Internet routing and peer-to-peer systems, since we believe that studying these (existing) systems can lead to engineering principles for self-stabilizing systems in various application areas.
Constantin Adam, Rolf Stadler
NOMS (1)2
2003 Weaver: Realizing a Scalable Management Paradigm on Commodity Routers
Koon-Seng Lim, Rolf Stadler
Integrated Network Management2
2002 Service deployment on high performance active network nodes
abstract
In the context of realizing service deployment on high performance active nodes, we address the problem of installing and configuring software components in complex, heterogeneous node environments. Service deployment is difficult on high-performance nodes due to their complex architectures. They are often based on multiprocessors and run service components in multiple concurrent execution environments. Many types of execution environments have been developed each of which is usually optimized for a certain type of tasks. Advanced active nodes provide more than one to support the whole spectrum of services. Also, different types of active nodes usually support different sets of execution environments. All this motivates us to develop a service deployment scheme that can cope with heterogeneous active nodes. This paper presents our approach to this problem, called Chameleon. It has two important aspects. First, the service model we propose is based on components with two types of interfaces - a data flow interface for programming packet flows and a control interface for controlling and managing the service components. A service is structured as an arbitrary tree of such components. Second, the service specification is independent of any particular node architecture. During the service deployment phase, the service specification is resolved recursively on each node offering the service and is driven by node-specific parameters. The result of this resolution is a tree of service components, which can differ among different types of nodes. Our solution allows a service to take full advantage of specific node features, such as those related to performance or security. This paper advocates a service model that is specialized for active networking. We base the model on two basic abstractions, namely composable containers and connectors.
Matthias Bossardt, Lukas Ruf, Bernhard Plattner, Rolf Stadler
NOMS4
2001 A Navigation Pattern for Scalable Internet Management
abstract
Performing global management operations on the Internet in an efficient way is difficult, because of the continuous changes to the Internet topology, its large number of nodes and the lack of an up-to-date global database. In practice, these difficulties appear in the management of large private IP networks and large autonomous systems, which form the sub-topologies of the Internet and are under independent administration. This paper introduces the echo pattern, a scheme for distributing management operations, which addresses these difficulties. Management operations based on this pattern do not need knowledge of the network topology, they can dynamically adapt to changes in the topology, and they scale well in very large networks. A management operation based on the echo pattern has two phases. In a first phase, the network is being flooded with management commands to be run on the network elements. In the second phase, the results of the local management operations are aggregated inside the network. We analyze the echo pattern with respect to time and traffic complexity and compare its performance to that of a centralized management scheme. Our results show that "typical" echo-based management operations could be executed within some 18 seconds on the entire Internet. This short time is due to (1) the high degree of parallelism and distributed control in this pattern and (2) some specific properties of the Internet topology.
Koon-Seng Lim, Rolf Stadler
Integrated Network Management2
2001 Selected Topics in Network and Systems Management
Emil C. Lupu, Subrata Mazumdar, Rolf Stadler
Comput. Networks3
2000 A middleware architecture for active distributed management of IP networks
abstract
We argue that a management platform for the future Internet has to be inherently distributed and programmable. This motivates us to introduce a new management architecture, named the active distributed management (ADM) architecture, which exploits the active network and mobile agent paradigms and provides the properties of distributed control and programmability inside the network. We realize the ADM architecture as a management middleware composed of several layers. In order to facilitate the development of efficient and correct programs, these layers include patterns for distributed algorithms that are typical for management applications and a set of building blocks for constructing management programs. First results of an ADM prototype system are presented.
Ryutaro Kawamura, Rolf Stadler
NOMS2
1999 The Impact of Active Networking Technology on Service Management in a Telecom Environment
abstract
Active networking, where network nodes perform customized processing of packets, is a rapidly expanding field of research. This paper is based on the assumption that active networking technology will mature to a point where it can be commercially deployed on a larger scale. We investigate the realization of service provisioning and service management in a telecommunication environment that is based on active networking technology, primarily with respect to customer-provider interactions. Compared to conventional networking technology, active networking concepts enable additional flexibility in supporting management tasks. We outline a framework that allows customers, on the one hand, to access and manage a service in a provider's domain, and, on the other hand, to outsource a service and its management to a service provider. Our framework has the properties of supporting: (1) generic, i.e., service-independent, interfaces for service provisioning and management; and (2) customized service abstractions and control functions, according to a customer's requirements. Further, we describe how some of the key concepts of this framework can be realized in an active networking testbed that we are in the process of building.
Marcus Brunner, Rolf Stadler
Integrated Network Management2
1999 Service enabling platforms for networked multimedia systems
David Hutchison 0001, Giovanni Pacifici, Bernhard Plattner, Rolf Stadler, Joseph S. Sventek
IEEE J. Sel. Areas Commun.4
1998 Building open programmable multimedia networks
abstract
Recent advances in distributed systems and transportable software and increasing demand for better quality-of-service (QOS) control in multiservice networks are driving a re-examination of network software architectures. We established the COMET Group (Control Management and Telemedia) at Columbia University's Center for Telecommunications Research to provide a comprehensive understanding of network software architecture of the 1990s and beyond. In this paper, we present an overview of our activities, focusing on new research initiatives, international forum participation and on-going research projects. Collectively, these activities shape our vision of a new era driven by open programmable networking.
Andrew T. Campbell, Aurel A. Lazar, Henning Schulzrinne, Rolf Stadler
Comput. Commun.4
1997 Customer Management and Control of Broadband VPN Services
Mun Choon Chan, Aurel A. Lazar, Rolf Stadler
Integrated Network Management3
1996 Prototyping Network Architectures on a Supercomputer
abstract
Outlines a methodology for developing network control systems which allows for an evaluation of the dynamic behavior and overall performance at an early stage of the development process. Our approach is to build a software prototype which is designed according to the architecture under consideration and runs the intended control algorithms. The functional and dynamic properties of this prototype are tested and evaluated on an emulation platform that we built for this purpose. By providing support for real-time visualization and interactive emulation, this platform can be used to study multimedia networks in various scenarios, such as different load patterns, network sizes and management operations. The current implementation runs on a KSR-1 and an SP2 parallel processor, which are connected to a graphics workstation via ATM links. We use the platform in several projects, one of which aims at developing an architecture for managing multimedia network services.
Mun Choon Chan, Giovanni Pacifici, Rolf Stadler
HPDC3
1996 An architecture for broadband virtual networks under customer control
abstract
Emerging ATM-based virtual private network (VPN) services offer customers a flexible way to interconnect customer premises networks (CPNs) via high-speed links. Compared with traditional leased lines, these services allow for rapid provisioning of VPN bandwidth through cooperative control between customer and provider. Customers can dynamically renegotiate the VPN bandwidth according to their current needs, paying only for the resources they actually use. In order to meet the various requirements and demands of different classes of VPN customers, a VPN provider must provide customers with the flexibility to choose their own control schemes and objectives. The focus of this paper is on enhancing the customer's capability of controlling a VPN. First, we propose a new scheme for a broadband VPN service, which is based on the virtual path group (VPG) concept. In our scheme, the customer performs VP control operations without interacting with the VPN provider, thus enabling the following merits: (1) the customer can share bandwidth among VPs that traverse the same physical network link in the provider's domain, thus using the VPN bandwidth more efficiently; (2) customers can perform VP control operations according to their own requirements and control objectives. Second, we outline an architecture for a customer-operated control system, which utilizes a VPG-based VPN service. The system is structured into three layers of control, which execute on different time scales. The functionalities of these layers are call processing, VP control, and VPN control, respectively. Finally, we evaluate the effectiveness of the control system, with respect to VP control.
Mun Choon Chan, Hisaya Hadama, Rolf Stadler
NOMS3
1995 An architecture for performance management of multimedia networks
Giovanni Pacifici, Rolf Stadler
Integrated Network Management2
1995 Managing Real-Time Services in Multimedia Networks Using Dynamic Visualization and High-Level Controls
abstract
No abstract available.
Mun Choon Chan, Giovanni Pacifici, Rolf Stadler
ACM Multimedia3
1995 Real-Time Emulation and Visualization of Large Multimedia Networks
Mun Choon Chan, Giovanni Pacifici, Rolf Stadler
ACM Multimedia3