Alessandro Vittorio Papadopoulos

dblp:92/9596 · also Alessandro Papadopoulos, Alessandro V. Papadopoulos · DBLP profile ↗
← Back
86ranked-venue papers
6as first author
49since 2021 · last 2026
0000-0002-1364-8127ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 49 · 3 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 8 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Security and privacy · 2 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Introduction to the Special Issue on Control of Computing Systems
abstract
This work is licensed under Creative Commons Attribution-NonCommercial-NoDerivatives International.
Sophie Cerf, Alessandro Vittorio Papadopoulos, Éric Rutten
ACM Trans. Auton. Adapt. Syst.2
2025 Stochastic Scheduling for Human-Robot Collaboration in Dynamic Manufacturing Environments
abstract
Collaborative human-robot teams enhance efficiency and adaptability in manufacturing, but task scheduling in mixed-agent systems remains challenging due to the uncertainty of task execution times and the need for synchronization of agent actions. Existing task allocation models often rely on deterministic assumptions, limiting their effectiveness in dynamic environments. We propose a stochastic scheduling framework that models uncertainty through probabilistic makespan estimates, using convolutions and stochastic max operators for realistic performance evaluation. Our approach employs metaheuristic optimization to generate executable schedules aligned with human preferences and system constraints. It features a novel deadlock detection and repair mechanism to manage cross-schedule dependencies and prevent execution failures. This framework offers a robust, scalable solution for real-world human-robot scheduling in uncertain, interdependent task environments.
Anders Lager, Branko Miloradovic, Giacomo Spampinato, Thomas Nolte, Alessandro Vittorio Papadopoulos
RO-MAN5
2025 Nip it in the Bud: Job Acceptance Multi-Server
abstract
Computationally demanding tasks with highly variable execution times may require parallel processing. Scheduling such tasks with low deadline miss rates but without significant overprovisioning is challenging. This issue arises in applications like nonlinear optimization for Model Predictive Control (MPC). The Constant Bandwidth Server (CBS) provides timing isolation, supporting both hard and soft real-time tasks. However, scheduling parallel, time-varying jobs across multiple CBS instances requires static job-to-server assignments, which can lead to resource underutilization due to queued jobs awaiting specific servers. This paper introduces the Job Acceptance Multi-Server (JAMS), a mechanism in which multiple CBS instances share a common job queue, enabling flexible job dispatching for parallel workloads. JAMS incorporates a job dismissal mechanism to address overloads, ensuring that only jobs with guaranteed resource availability are accepted. Each CBS instance checks if it can complete a job by its deadline, given probabilistic knowledge on its execution times, dismissing unfeasible jobs to avoid excessive tardiness across queued tasks. Implemented in Linux, JAMS is evaluated with computation times drawn from an MPC task and synthetic datasets. The extensive experimental results we provide demonstrate that JAMS effectively controls the deadline miss rate, maintaining it below a specified design threshold.
Anna Friebe, Tommaso Cucinotta, Filip Markovic 0001, Alessandro Vittorio Papadopoulos, Thomas Nolte
RTAS4
2025 Resource Management for Stochastic Parallel Synchronous Tasks: Bandits to the Rescue
abstract
Abstract In scheduling real-time tasks, we face the challenge of meeting hard deadlines while optimizing for some other objective, such as minimizing energy consumption. Formulating the optimization as a Multi-Armed Bandit (MAB) problem allows us to use MAB strategies to balance the exploitation of good choices based on observed data with the exploration of potentially better options. In this paper, we integrate hard real-time constraints with MAB strategies for resource management of a Stochastic Parallel Synchronous Task. On a platform with $$M$$ M cores available for the task, $$m\le M$$ m ≤ M cores are initially assigned. Prior work has shown how to compute a virtual deadline such that assigning all $$M$$ M cores to the task if it has not completed by this virtual deadline guarantees that the deadline will be met. An MAB strategy is used to select the value of $$m$$ m . A Dynamic Power Management (DPM) energy model considering CPU sockets and sleep states is described. Experimental evaluation shows that MAB strategies learn consistently suitable $$m$$ m , and perform well compared to binary exponential search and greedy methods.
Anna Friebe, Alberto Marchetti-Spaccamela, Tommaso Cucinotta, Alessandro Vittorio Papadopoulos, Thomas Nolte, Sanjoy Baruah
Real Time Syst.4
2025 Rule-Based Predictive Control for Battery Scheduling in Microgrids Under Power Generation and Load Uncertainties
abstract
This paper addresses the control of the state of charge (SoC) of a Battery Energy Storage System (BESS) in a microgrid, considering uncertainties in load and Renewable Energy Sources (RES) generated power estimations. To achieve this objective, we propose RubPC, a novel rule-based Model Predictive Control (MPC). We partition the feasible operation space of the microgrid into two subzones, referred to as the white and yellow zones. The yellow zone represents the boundary space between the feasible and unfeasible operation spaces. In RubPC, we initially implement MPC on a predefined optimization window to determine the optimal SoC of the BESS, aiming to keep the microgrid within the white zone. Noting that mismatches between estimated and actual load and generated power may lead to constraint violations, we introduce a rule-based controller as a supervisory control. This controller monitors the microgrid’s state, and if the microgrid enters the yellow zone, it adjusts the control to maintain the microgrid within the white zone. We validate our proposed method by simulating it using data from an electrified quarry site in Sweden. Note to Practitioners—Optimizing the charge and discharge schedule of BESSs in microgrids offers a promising avenue for substantial economic and technical benefits. However, the successful realization of these benefits hinges on accurately estimating and aligning the values of load and RES-generated power. In industries, the consequences of mismatches between these estimates and actual values can translate into unexpected costs that may outweigh the anticipated economic benefits of BESS’s optimal charge and discharge schedule. This paper underscores the critical importance of addressing this concern to ensure the viability of BESS applications in various industries. To tackle this challenge, we present RubPC, an innovative Rule-Based MPC framework. Unlike conventional approaches, RubPC is specifically designed to effectively handle discrepancies between estimated and actual values, thereby preventing potential constraint violations. Our aim is to offer practitioners a robust solution that not only brings economic benefits but also ensures their safe and reliable operation.
Mojtaba Kaheni, Jiali Fu, Alessandro Vittorio Papadopoulos
IEEE Trans Autom. Sci. Eng.3
2025 dcGuard: A Holistic Approach for Detecting and Isolating Malicious Nodes in Cloud Data Centers
abstract
This paper presentsdcGuard, a unified security approach for detecting and isolating misbehaving computing and forwarding nodes in multi-tenant virtualized cloud data centers.dcGuardemploys technological advancements in Virtual Machine Introspection (VMI), Software-Defined Networking (SDN), and secure probabilistic sketching to detect and isolate parts of the Virtual Machines (VMs) and network switches experiencing malicious behavior dynamically. The main contribution lies in designing a divide-and-conquer strategy that utilizes VMI and network programmability to apply focused distributed task and packet probing mechanisms on portions of the data center network rather than focusing the security functions on the entire physical network. The processing VMs and network switches are recursively partitioned into independent logical groups inspected individually to localize abnormal/malicious computing and switching nodes incrementally. This remarkably enhances the efficiency of the detection mechanisms, which opportunistically approaches a logarithmic time complexity in the number of protocol steps towards convergence (compared to a linear time complexity in traditional intrusion detection systems) when a relatively low number of hostile VMs and switches are present. Real experiments are evaluated, and a test-bed blueprint of the proposed design is emulated in a virtualized cloud environment using the Mininet emulator. The performance, convergence, and accuracy benchmarks corroborate the analytical advantage of the proposed security approach.
Wassim Itani, Maha Shamseddine, Auday Aldulaimy, Thomas Nolte, Alessandro Vittorio Papadopoulos
IEEE Trans. Dependable Secur. Comput.5
2024 A Data-Driven Predictive Control Driver for Racing Car Simulation
abstract
The capability to accurately simulate the behavior of a racing car is paramount in modern-day racing competitions to quickly find a good base setup to kick-start the work on track. Typically, a professional driver is employed to drive the simulated race car and provide feedback. However, this operation is expensive and time-consuming, as capable human drivers quickly become a bottleneck. In conjunction with highly accurate simulations of the physical car’s behavior, a capable virtual driver could thus accelerate the car setup and development to a great extent. In this paper, we propose to apply a data-driven predictive control approach called Data-enabled Predictive Control to model a racing driver by tracking a pre-defined trajectory. We compare our proposed approach with an industrial first-choice Proportional-Integral-Derivative controller and state-of-the-art Model Predictive Control controller, finding that the approach is feasible, and it can provide significant improvements over the state-of-the-art, especially for trajectories whose feasibility is at the edge of the car’s capabilities.
Ruslan Shaiakhmetov, Danilo Pianini, Valter Venusti, Alessandro Vittorio Papadopoulos
DS-RT4
2024 OPC UA PubSub and Industrial Controller Redundancy
abstract
Industrial controllers constitute the core of numerous automation solutions. Continuous control system operation is crucial in certain sectors, where hardware duplication serves as a strategy to mitigate the risk of unexpected operational halts due to hardware failures. Standby controller redundancy is a commonly adopted strategy for process automation. This approach involves an active primary controller managing the process while a passive backup is on standby, ready to resume control should the primary fail. Typically, redundant controllers are paired with redundant networks and devices to eliminate any single points of failure. The process automation domain is on the brink of a paradigm shift towards greater interconnectivity and interoperability. OPC UA is emerging as the standard that will facilitate this shift, with OPC UA PubSub as the communication standard for cyclic real-time data exchange. Our work investigates standby redundancy using OPC UA PubSub, analyzing a system with redundant controllers and devices in publisher-subscriber roles. The analysis reveals that failovers are not subscriber-transparent without synchronized publisher states. We discuss solutions and experimentally validate an internal stack state synchronization alternative.
Bjarne Johansson, Olof Holmgren, Martin Dahl, Håkan Forsberg, Thomas Nolte, Alessandro Vittorio Papadopoulos
ETFA6
2024 Towards High-Integrity Redundancy Role Leasing
abstract
Control systems are often an integral part of automation solutions where high reliability is crucial due to the high cost of downtime. The risk of unplanned downtime is typically reduced with redundant solutions. Additionally, safety-critical automation functions require high-integrity controllers. Today, the prevalent redundancy solution is a standby scheme, where one active primary controller drives the process while a standby backup controller is ready to take over in case of primary failure. This redundant controller pair can consist of high - integrity controllers. The automation industry is trending towards Ethernet as the sole communication medium. Our work presents an initial study of a high-integrity realization of a redundancy failure detection mechanism that guarantees only one primary controller, even in the case of network partitioning between the redundant controller pair. The failure detection is a lease-based function that leases the primary role from a central lease broker. This work discusses a high-integrity realization of the primary redundancy role leasing. We deduce and present the high-integrity-related requirements and a high-level design as an initial step towards a high-integrity realization of the redundancy role leasing.
Bjarne Johansson, Olof Holmgren, Håkan Forsberg, Thomas Nolte, Alessandro Vittorio Papadopoulos
ETFA5
2024 Robust Online Epistemic Replanning of Multi-Robot Missions
abstract
As Multi-Robot Systems (MRS) become more affordable and computing capabilities grow, they provide significant advantages for complex applications such as environmental monitoring, underwater inspections, or space exploration. However, accounting for potential communication loss or the unavailability of communication infrastructures in these application domains remains an open problem. Much of the applicable MRS research assumes that the system can sustain communication through proximity regulations and formation control or by devising a framework for separating and adhering to a predetermined plan for extended periods of disconnection. The latter technique enables an MRS to be more efficient, but breakdowns and environmental uncertainties can have a domino effect throughout the system, particularly when the mission goal is intricate or time-sensitive. To deal with this problem, our proposed framework has two main phases: i) a centralized planner to allocate mission tasks by rewarding intermittent rendezvous between robots to mitigate the effects of the unforeseen events during mission execution, and ii) a decentralized replanning scheme leveraging epistemic planning to formalize belief propagation and a Monte Carlo tree search for policy optimization given distributed rational belief updates. The proposed framework outperforms a baseline heuristic and is validated using simulations and experiments with aerial vehicles.
Lauren Bramblett, Branko Miloradovic, Patrick Sherman, Alessandro Vittorio Papadopoulos, Nicola Bezzo
IROS4
2024 Risk-Aware Planning of Collaborative Mobile Robot Applications with Uncertain Task Durations
abstract
The efficiency of collaborative mobile robot applications is influenced by the inherent uncertainty introduced by humans’ presence and active participation. This uncertainty stems from the dynamic nature of the working environment, various external factors, and human performance variability. The observed makespan of an executed plan will deviate from any deterministic estimate. This raises questions about whether a calculated plan is optimal given uncertainties, potentially risking failure to complete the plan within the estimated timeframe. This research addresses a collaborative task planning problem for a mobile robot serving multiple humans through tasks such as providing parts and fetching assemblies. To account for uncertainties in the durations needed for a single robot and multiple humans to perform different tasks, a probabilistic modeling approach is employed, treating task durations as random variables. The developed task planning algorithm considers the modeled uncertainties while searching for the most efficient plans. The outcome is a set of the best plans, where no plan is better than the other in terms of stochastic dominance. Our proposed methodology offers a systematic framework for making informed decisions regarding selecting a plan from this set, considering the desired risk level specific to the given operational context.
Anders Lager, Branko Miloradovic, Giacomo Spampinato, Thomas Nolte, Alessandro Vittorio Papadopoulos
RO-MAN5
2024 An Improved Worst-Case Response Time Analysis for AVB Traffic in Time-Sensitive Networks
abstract
Time-Sensitive Networking (TSN) has become one of the most relevant communication networks in many application areas. Among several traffic classes supported by TSN networks, Audio-Video Bridging (AVB) traffic requires a Worst-Case Response Time Analysis (WCRTA) to ensure that AVB frames meet their time requirements. In this paper, we evaluate the existing WCRTAs that cover various features of TSN, including Scheduled Traffic (ST) interference and preemption. When considering the effect of the ST interference, we detect optimism problems in two of the existing WCRTAs, namely (i) the analysis based on the busy period calculation and (ii) the analysis based on the eligible interval. Therefore, we propose a new analysis including a new ST interference calculation that can extend the analysis based on the eligible interval approach. The new analysis covers the effect of the ST interference, the preemption by the ST traffic, and the multi-hop architecture. The resulting WCRTA, while safe, shows a significant improvement in terms of pessimism level compared to the existing analysis approaches relying on either the concept of busy period or the Network Calculus model.
Daniel Bujosa, Julián Proenza, Alessandro Vittorio Papadopoulos, Thomas Nolte, Mohammad Ashjaei
RTSS3
2024 Efficiently bounding deadline miss probabilities of Markov chain real-time tasks
abstract
Abstract In real-time systems analysis, probabilistic models, particularly Markov chains, have proven effective for tasks with dependent executions. This paper improves upon an approach utilizing Gaussian emission distributions within a Markov task execution model that analyzes bounds on deadline miss probabilities for tasks in a reservation-based server. Our method distinctly addresses the issue of runtime complexity, prevalent in existing methods, by employing a state merging technique. This not only maintains computational efficiency but also retains the accuracy of the deadline-miss probability estimations to a significant degree. The efficacy of this approach is demonstrated through the timing behavior analysis of a Kalman filter controlling a Furuta pendulum, comparing the derived deadline miss probability bounds against various benchmarks, including real-time Linux server metrics. Our results confirm that the proposed method effectively upper-bounds the actual deadline miss probabilities, showcasing a significant improvement in computational efficiency without significantly sacrificing accuracy.
Anna Friebe, Filip Markovic 0001, Alessandro Vittorio Papadopoulos, Thomas Nolte
Real Time Syst.3
2024 Hierarchical Resource Orchestration Framework for Real-time Containers
abstract
Container-based virtualization is a promising deployment model in fog and edge computing applications, because it allows a seamless co-existence of virtualized applications in a heterogeneous environment without introducing significant overhead. Certain application domains (e.g., industrial automation, automotive, or aerospace) mandate that applications exhibit a certain degree of temporal predictability. Container-based virtualization cannot be easily used for such applications, since the technology is not designed to support real-time properties and handle temporal disturbances. This article proposes a framework consisting of a static offline and a dynamic online phase for resource allocation and adaptive re-dimensioning of real-time containers. In the offline phase, the optimal initial deployment and dimensioning of containers are decided based on ideal system models. Additionally, to adapt to dynamic variations caused by changing workloads or interferences, the online phase adapts the CPU usage and limits of real-time containers at runtime to improve the real-time behavior of the real-time containerized applications while optimizing resource usage. We implement the framework in a real Linux-based system and show through a series of experiments that the proposed framework is able to adjust and re-distribute computing resources between containers to improve the real-time behavior of containerized applications in the presence of temporal disturbances while optimizing resource usage.
Václav Struhár, Silviu S. Craciunas, Mohammad Ashjaei, Moris Behnam, Alessandro Vittorio Papadopoulos
ACM Trans. Embed. Comput. Syst.5
2023 The SPEC-RG Reference Architecture for The Compute Continuum
abstract
As the next generation of diverse workloads like autonomous driving and augmented/virtual reality evolves, computation is shifting from cloud-based services to the edge, leading to the emergence of a cloud-edge compute continuum. This continuum promises a wide spectrum of deployment opportunities for workloads that can leverage the strengths of cloud (scalable infrastructure, high reliability) and edge (energy efficient, low latencies). Despite its promises, the continuum has only been studied in silos of various computing models, thus lacking strong end-to-end theoretical and engineering foundations for computing and resource management across the continuum. Consequently, devel-opers resort to ad hoc approaches to reason about performance and resource utilization of workloads in the continuum. In this work, we conduct a first-of-its-kind systematic study of various computing models, identify salient properties, and make a case to unify them under a compute continuum reference architecture. This architecture provides an end-to-end analysis framework for developers to reason about resource management, workload distribution, and performance analysis. We demonstrate the utility of the reference architecture by analyzing two popular continuum workloads, deep learning and industrial IoT. We have developed an accompanying deployment and benchmarking framework and first-order analytical model for quantitative reasoning of continuum workloads. The framework is open-sourced and available at https://github.com/atlarge-research/continuum.
Matthijs Jansen, Auday Aldulaimy, Alessandro Vittorio Papadopoulos, Animesh Trivedi, Alexandru Iosup
CCGrid3
2023 Change-Point and Model Estimation with Heteroskedastic Noise and Unknown Model Structure
abstract
In this paper, we investigate the problem of modeling time-series as a process generated through (i) switching between several independent sub-models; (ii) where each sub-model has heteroskedastic noise, and (iii) a polynomial bias, describing nonlinear dependency on system input. First, we propose a generic nonlinear and heteroskedastic statistical model for the process. Then, we design Maximum Likelihood (ML) parameters estimation method capable of handling heteroscedasticity and exploiting constraints on model structure. We investigate solving the intractable ML optimization using population-based stochastic numerical methods. We then find possible model change-points that maximize the likelihood without over-fitting measurement noise. Finally, we verify the usefulness of the proposed technique in a practically relevant case study, the execution-time of odometry estimation for a robot operating radar sensor, and evaluate the different proposed procedures using both simulations and field data.
Anas W. Alhashimi, Thomas Nolte, Alessandro Vittorio Papadopoulos
CoDIT3
2023 Introducing Guard Frames to Ensure Schedulability of All TSN Traffic Classes
abstract
Offline scheduling of Scheduled Traffic (ST) in Time-Sensitive Networks (TSN) without taking into account the quality of service of non-ST traffic, e.g., time-sensitive traffic such as Audio-Video Bridging (AVB) traffic, can potentially cause deadline misses for non-ST traffic. In this paper, we report our ongoing work to propose a solution that, regardless of the ST scheduling algorithm being used, can ensure meeting timing requirements for non-ST traffic. To do this, we define a frame called Guard Frame (GF) that will be scheduled together with all ST frames. We show that a proper design for the GFs will leave necessary porosity in the ST schedules to ensure that all non-ST traffic will meet their timing requirements.
Daniel Bujosa, Julián Proenza, Alessandro Vittorio Papadopoulos, Thomas Nolte, Mohammad Ashjaei
ETFA3
2023 Towards a holistic approach to security validation of construction machinery through HIL systems
abstract
The construction industry is increasingly equipping its machinery with sophisticated embedded systems and modern connectivity. Technology advancements in connected safety-critical systems are complex, with cyber-security becoming a more critical factor. Due to interdependencies and network connectivity, attack surfaces and vulnerabilities have increased significantly. Consequently, it is imperative to perform a risk assessment and implement robust security testing methods in order to prevent cyber-attacks on machinery segments. This paper presents a method for identifying potential security threats that also affect machine functional safety, facilitated by identifying threats in the threat modeling process and analyzing safety-security synergies. By identifying such risks, attack scenarios are created to simulate cyber-attacks and create test cases for validation. This approach integrates security testing into the current testing process by using penetration testing tools and utilizing a Hardware-in-the-Loop(HIL) test setup and it is verified with a simulated Denial of Service attack over a CAN network.
Sheela Hariharan, Andreas Erséus, Thomas Nolte, Alessandro Vittorio Papadopoulos
ETFA4
2023 Consistency Before Availability: Network Reference Point based Failure Detection for Controller Redundancy
abstract
Distributed control systems constitute the automation solution backbone in domains where downtime is costly. Redundancy reduces the risk of faults leading to unplanned downtime. The Industry 4.0 appetite to utilize the device-to-cloud continuum increases the interest in network-based hardware-agnostic controller software. Functionality, such as controller redundancy, must adhere to the new ground rules of pure network dependency. In a standby controller redundancy, only one controller is the active primary. When the primary fails, the backup takes over. A typical network-based failure detection uses a cyclic message with a known interval, a.k.a. a heartbeat. Such a failure detection interprets heartbeat absences as a failure of the supervisee; consequently, a network partitioning could be indistinguishable from a node failure. Hence, in a network partitioning situation, a conventional heartbeat-based failure detection causes more than one active controller in the redundancy set, resulting in inconsistent outputs. We present a failure detection algorithm that uses network reference points to prevent network partitioning from leading to dual primary controllers. In other words, a failure detection that prioritizes consistency before availability.
Bjarne Johansson, Mats Rågberger, Alessandro Vittorio Papadopoulos, Thomas Nolte
ETFA3
2023 Dependability and Security Aspects of Network-Centric Control
abstract
Industrial automation and control systems are responsible for running our most important infrastructures, providing electricity and clean water, producing medicine and food, along with many other services and products we take for granted. The safe and secure operation of these systems is therefore of great importance.One of the emerging trends in industrial automation systems is the transition from static hierarchical controller-centric systems to flexible network-centric systems. This transition has a great impact on the characteristics of industrial automation systems. In this article we describe the network-centric design strategy for industrial automation systems and describe the impact on dependability and security aspects that this strategy brings, looking at both challenges and possibilities.
Björn Leander, Bjarne Johansson, Tomas Lindström, Olof Holmgren, Thomas Nolte, Alessandro Vittorio Papadopoulos
ETFA6
2023 Challenges in the Automated Disassembly Process of Electric Vehicle Battery Packs
abstract
The surge in the development and adoption of Electric Vehicles (EVs) globally is a trend many countries are paying close attention to. This inevitably means that a significant number of EV batteries will soon reach their End-of-Life (EoL). This looming issue reveals a notable challenge: there’s currently a lack of sustainable strategies for managing Lithium-ion Batteries (LiBs) when they reach their EoL stage. The process of disassembling these battery packs is challenging due to their intricate design, involving several different materials and components integrated tightly for performance and safety. Consequently, effective disassembly and subsequent recycling procedures require highly specialized methods and equipment, and involve significant safety and health risks. Moreover, existing recycling technologies often fail to recover all valuable and potentially hazardous materials, leading to both economic and environmental loss. This paper provides an overview and analysis of possible challenges arising in the domain of automated battery disassembly and recycling of EV batteries that reached their EoL. We provide insight into the disassembly process as well as optimization of the disassembly sequence with the goal of minimizing the overall cost and environmental footprint.
Branko Miloradovic, Eduard Marti Bigorra, Thomas Nolte, Alessandro Vittorio Papadopoulos
ETFA4
2023 Dispatching Deadline Constrained Jobs in Edge Computing Systems
abstract
The edge computing paradigm extends the architectural space of real-time systems by bringing the capabilities of the cloud to the edge. Unlike cloud-native systems designed for mean response times, real-time industrial embedded systems are designed to control a single physical system, such as a manipulator arm or a mobile robot, that requires temporal predictability. We consider the problem of dispatching and scheduling of jobs with deadlines that can be offloaded to the edge and propose DAL, a deadline-aware load balancing and scheduling framework that leverages the availability of on-demand computing resources along with an on-arrival dispatching scheme to manage temporal requirements of such offloaded applications. The evaluation indicates that DAL can achieve reasonably good performance even when execution times, arrival times, and deadlines vary.
Shaik Mohammed Salman, Alessandro Vittorio Papadopoulos, Saad Mubeen, Thomas Nolte
ETFA2
2023 Resource Adaptation for Real-Time Containers Considering Quality of Control
abstract
Container-based virtualization has become a promising deployment model for industrial applications mainly due to its benefits, such as providing support for co-located applications in heterogeneous environments. However, such facilitation brings challenges, including full temporal isolation among real-time applications and support for time-critical applications. In this paper, we tackle such challenges, in particular when the applications are time-sensitive Control Applications. The literature suggests that flexible timing constraints for Control Applications are beneficial in responding to disturbances and minimizing response deviation. Therefore, we propose a mechanism to support such a runtime adaptation in container-based virtualization. To show the performance of the proposed mechanism, we implement our approach on a Linux-based hierarchical scheduling platform, and we evaluate it for a Control application.
Václav Struhár, Mohammad Ashjaei, Moris Behnam, Alessandro Vittorio Papadopoulos, Silviu S. Craciunas
ETFA4
2023 Evaluating Dispatching and Scheduling Strategies for Firm Real-Time Jobs in Edge Computing
abstract
We consider the problem of on-arrival dispatching and scheduling jobs with stochastic execution times, inter-arrival times, and deadlines in multi-server fog and edge computing platforms. In terms of mean response times, it has been shown that size-based scheduling policies, when combined with dispatching policies such as join-shortest-queue, provide better performance over policies such as first-in-first-out. Since job sizes may not always be known apriori, prediction-based policies have been shown to perform reasonably well. However, little is known about the performance of prediction-based policies for jobs with firm deadlines. In this paper, we address this issue by considering the number of jobs that complete within their deadlines as a performance metric and investigate, using simulations, the performance of a prediction-based shortest-job-first scheduling policy for the considered metric and compare it against scheduling policies that prioritize based on deadlines (EDF) and arrival times (FIFO). The evaluation indicates that in under-loaded conditions, the prediction-based policy is outperformed by both FIFO and EDF policies. However, in overloaded scenarios, the prediction-based policy offers slightly better performance.
Shaik Mohammed Salman, Alessandro Vittorio Papadopoulos, Saad Mubeen, Thomas Nolte
IECON2
2023 Scheduling Firm Real-time Applications on the Edge with Single-bit Execution Time Prediction
abstract
The edge computing paradigm brings the capabilities of the cloud such as on-demand resource availability to the edge for applications with low-latency and real-time requirements. While cloud-native load balancing and scheduling algorithms strive to improve performance metrics like mean response times, real-time systems, that govern physical systems, must satisfy deadline requirements. This paper explores the potential of an edge computing architecture that utilizes the on-demand availability of computational resources to satisfy firm real-time requirements for applications with stochastic execution and inter-arrival times. As it might be difficult to know precise execution times of individual jobs prior to completion, we consider an admission policy that relies on single-bit execution time predictions for dispatching. We evaluate its performance in terms of the number of jobs that complete by their deadlines via simulations. The results indicate that the prediction-based admission policy can achieve reasonable performance for the considered settings.
Shaik Mohammed Salman, Van-Lan Dao, Alessandro Vittorio Papadopoulos, Saad Mubeen, Thomas Nolte
ISORC3
2023 Continuous-Emission Markov Models for Real-Time Applications: Bounding Deadline Miss Probabilities
abstract
Probabilistic approaches have gained attention over the past decade, providing a modeling framework that enables less pessimistic analysis of real-time systems. Among the different proposed approaches, Markov chains have been shown effective for analyzing real-time systems, particularly in estimating the pending workload distribution and deadline miss probability. However, the state-of-the-art mainly considered discrete emission distributions without investigating the benefits of continuous ones. In this paper, we propose a method for analyzing the workload probability distribution and bounding the deadline miss probability for a task executing in a reservation-based server, where execution times are described by a Markov model with Gaussian emission distributions. The evaluation is performed for the timing behavior of a Kalman filter for Furuta pendulum control. Deadline miss probability bounds are derived with a workload accumulation scheme. The bounds are compared to 1) measured deadline miss ratios of tasks running under the Linux Constant Bandwidth Server with SCHED-DEADLINE, 2) estimates derived from a Markov Model with discrete-emission distributions (PROSIT), 3) simulation-based estimates, and 4) an estimate assuming independent execution times. The results suggest that the proposed method successfully upper bounds actual deadline miss probabilities. Compared to the discrete-emission counterpart, the computation time is independent of the range of the execution times under analysis, and resampling is not required.
Anna Friebe, Filip Markovic 0001, Alessandro Vittorio Papadopoulos, Thomas Nolte
RTAS3
2023 CTA: A Correlation-Tolerant Analysis of the Deadline-Failure Probability of Dependent Tasks
abstract
Estimating the worst-case deadline failure probability (WCDFP) of a real-time task is notoriously difficult, primarily because a task's execution time typically depends on prior activations (i.e., history dependence) and the execution of other tasks (e.g., via shared inputs). Previous analyses have either assumed that execution times are probabilistically independent (which is unrealistic and unsafe), or relied on complex upper-bounding abstractions such as probabilistic worst-case execution time (pWCET), which mask dependencies with pessimism. Exploring an analytically novel direction, this paper proposes the first closed-form upper bound on WCDFP that accounts for dependent execution times. The proposed correlation-tolerant analysis (CTA), based on Cantelli's inequality, targets fixed-priority scheduling and requires only two basic summary statistics of each task's ground- truth execution time distribution: upper bounds on the mean and standard deviation (for any possible job-arrival sequence). Notably, CTA does not use pWCET, nor does it require the full execution-time distribution to be known. Core parts of the analysis have been verified with the Coq proof assistant. Empirical comparison with state-of-the-art WCDFP analyses reveals that CTA can yield significantly improved bounds (e.g., a lower WCDFP than any pWCET-based method for ~70% of the workloads tested at 90% pWCET utilization and 60% average utilization). Beyond accuracy gains, the favorable results highlight the potential of the previously unexplored analytical direction underlying CTA.
Filip Markovic 0001, Pierre Roux 0001, Sergey Bozhko, Alessandro Vittorio Papadopoulos, Björn B. Brandenburg
RTSS4
2023 Interplay of Human and AI Solvers on a Planning Problem
abstract
With the rapidly growing use of Multi-Agent Systems (MASs), which can exponentially increase the system complexity, the problem of planning a mission for MASs became more intricate. In some MASs, human operators are still involved in various decision-making processes, including manual mission planning, which can be an ineffective approach for any non-trivial problem. Mission planning and re-planning can be represented as a combinatorial optimization problem. Computing a solution to these types of problems is notoriously difficult and not scalable, posing a challenge even to cutting-edge solvers. As time is usually considered an essential resource in MASs, automated solvers have a limited time to provide a solution. The downside of this approach is that it can take a substantial amount of time for the automated solver to provide a sub-optimal solution. In this work, we are interested in the interplay between a human operator and an automated solver and whether it is more efficient to let a human or an automated solver handle the planning and re-planning problems, or if the combination of the two is a better approach. We thus propose an experimental setup to evaluate the effect of having a human operator included in the mission planning and re-planning process. Our tests are performed on a series of instances with gradually increasing complexity and involve a group of human operators and a metaheuristic solver based on a genetic algorithm. We measure the effect of the interplay on both the quality and structure of the output solutions. Our results show that the best setup is to let the operator come up with a few solutions, before letting the solver improve them.
Afshin Ameri, Branko Miloradovic, Baran Çürüklü, Alessandro Vittorio Papadopoulos, Mikael Ekström, Johann Dréo
SMC4
2023 Feedback-based resource management for multi-threaded applications
abstract
Abstract Reconciling the constraint of guaranteeing to always meet deadlines with the optimization objective of reducing waste of computing capacity lies at the heart of a large body of research on real-time systems. Most approaches to doing so require the application designer to specify a deeper characterization of the workload (and perhaps extensive profiling of its run-time behavior), which then enables shaping the resource assignment to the application. In practice, such approaches are weak as they load the designer with the heavy duty of a detailed workload characterization. We seek approaches for reducing the waste of computing resources for recurrent real-time workloads in the absence of such additional characterization, by monitoring the minimal information that needs to be observable about the run-time behavior of a real-time system: its response time. We propose two resource control strategies to assign resources: one based on binary-exponential search and the other, on principles of control. Both approaches are compared against the clairvoyant scenario in which the average/typical behavior is known. Via an extensive simulation, we show that both techniques are useful approaches to reducing resource computation while meeting hard deadlines.
Alessandro Vittorio Papadopoulos, Kunal Agrawal 0001, Enrico Bini, Sanjoy Baruah
Real Time Syst.1
2022 Nodeguard: A Virtualized Introspection Security Approach for the Modern Cloud Data Center
abstract
This paper presents Nodeguard, a security approach for detecting and isolating misbehaving Virtual Machines (VMs) in multi-tenant virtualized cloud data centers, based on the Virtual Machine Introspection (VMI) monitoring primitives. Nodeguard employs a divide-and-conquer strategy that checks logical groups of VMs to ensure the efficiency of the detection mechanisms which opportunistically approaches a complexity of$\mathcal{O}(\log_{2}(n))$when there is a relatively low number of hostile VMs. This greatly enhances the algorithmic time complexity of the pro-posed security system compared to the$\mathcal{O}(n)$complexity achieved by the traditional VMI inspection strategy that checks each VM separately. The approach has been evaluated in a virtualized cloud environment using the Mininet network emulator.
Maha Shamseddine, Auday Aldulaimy, Wassim Itani, Thomas Nolte, Alessandro Vittorio Papadopoulos
CCGRID5
2022 TOLERANCER: A Fault Tolerance Approach for Cloud Manufacturing Environments
abstract
The paper presents an approach to solve the software and hardware related failures in edge-cloud environments, more precisely, in cloud manufacturing environments. The proposed approach, called TOLERANCER, is composed of distributed components that continuously interact in a peer to peer fashion. Such interaction aims to detect stress situations or node failures, and accordingly, TOLERANCER makes decisions to avoid or solve any potential system failures. The efficacy of the proposed approach is validated through a set of experiments, and the performance evaluation shows that it responds effectively to different faults scenarios.
Auday Aldulaimy, Christian Sicari, Alessandro Vittorio Papadopoulos, Antonino Galletta, Massimo Villari, Mohammad Ashjaei
ETFA3
2022 The Effects of Clock Synchronization in TSN Networks with Legacy End-Stations
abstract
In this paper, we present our ongoing work on proposing solutions to integrate legacy end-stations into Time-Sensitive Network (TSN) communication systems where the legacy end-stations are synchronized via their legacy clock synchronization protocol. To this end, we experimentally identify the effects of lacking synchronization or partial synchronization in TSN networks. In the experiments we show the effects of clock synchronization in different scenarios on jitter and clock drifts. Based on the experiments, we propose preliminary solutions to overcome the identified effects.
Daniel Bujosa, Andreas Johansson, Mohammad Ashjaei, Alessandro Vittorio Papadopoulos, Julián Proenza, Thomas Nolte
ETFA4
2022 On In-Vehicle Network Security Testing Methodologies in Construction Machinery
abstract
In construction machinery, connectivity delivers higher advantages in terms of higher productivity, lower costs, and most importantly safer work environment. As the machinery grows more dependent on internet-connected technologies, data security and product cybersecurity become more critical than ever. These machines have more cyber risks compared to other automotive segments since there are more complexities in software, larger after-market options, use more standardized SAE J1939 protocol, and connectivity through long-distance wireless communication channels (LTE interfaces for fleet management systems). Construction machinery also operates throughout the day, which means connected and monitored endlessly. Till today, construction machinery manufacturers are investigating the product cybersecurity challenges in threat monitoring, security testing, and establishing security governance and policies. There are limited security testing methodologies on SAE J1939 CAN protocols. There are several testing frameworks proposed for fuzz testing CAN networks according to [1]. This paper proposes security testing methods (Fuzzing, Pen testing) for in-vehicle communication protocols in construction machinery.
Sheela Hariharan, Alessandro Vittorio Papadopoulos, Thomas Nolte
ETFA2
2022 Priority Based Ethernet Handling in Real-Time End System with Ethernet Controller Filtering
abstract
This work addresses the impact of best-effort traffic on network-dependent real-time functions in distributed control systems. Motivated by the increased Ethernet use in real-time dependent domains, such as the automation industry, a growth driven by Industry 4.0, interconnectivity desires, and data thirst. Ethernet allows different network-based functions to converge on one physical network infrastructure. In the automation domain, converged networks imply that functions with different criticality and real-time requirements coexist and share the same physical resources. The IEEE 60802 Time-Sensitive Networking profile for Industrial Automation targets the automation industry and addresses Ethernet network determinism on converged networks. However, the profile is still in the draft stage at the time of writing this paper. Meanwhile, Ethernet already provides attributes utilized by network equipment to prioritize time-critical communication. This paper shows that Ethernet Controller filtering with prioritized processing is a prominent solution for preserving real-time guarantees while supporting best-effort traffic. A solution capable of eliminating all best-effort traffic interference in the real-time application is exemplified and evaluated on a VxWorks system.
Bjarne Johansson, Mats Rågberger, Thomas Nolte, Alessandro Vittorio Papadopoulos
IECON4
2022 Analytical Approximations in Probabilistic Analysis of Real-Time Systems
abstract
Probabilistic timing and schedulability analysis of real-time systems is constrained by the problem of often intractable exact computations. The intractability problem is present whenever there is a large number of entities to be analysed, e.g., jobs, tasks, etc. In the last few years, the analytical approximations for deadline-miss probability emerged as an important solution in the above problem domain. In this paper, we explore analytical solutions for two major problems that are present in the probabilistic analysis of real-time systems. First, for a safe approximation of the entire probability distributions (e.g., of the accumulated execution workloads) we show how the Berry-Esseen theorem can be used. Second, we propose an approximation built on the Berry-Esseen theorem for efficient computation of the quantile functions of probability execution distributions. We also show the asymptotic bounds on the execution distribution of the fixed-priority preemptive tasks. In the evaluation, we investigate the complexity and accuracy of the proposed methods as the number of analysed jobs and tasks increases. The methods are compared with the circular convolution approach. We also investigate the memory footprint comparison between the proposed Berry-Esseen-based solutions and the circular convolution.. The contributions and results presented in this paper complement the state-of-the-art in accurate and efficient probabilistic analysis of real-time systems.
Filip Markovic 0001, Thomas Nolte, Alessandro Vittorio Papadopoulos
RTSS3
2022 Multi-processor scheduling of elastic applications in compositional real-time systems
abstract
Scheduling of real-time applications modelled according to the periodic and the sporadic task model under hierarchical and compositional real-time systems has been widely studied to provide temporal isolation among independent applications running on shared resources. However, for some real-time applications which are amenable to variation in their timing behaviour, usage of these tasks models can result in pessimistic solutions. The elastic task model addresses this pessimism by allowing the timing requirements of an application’s tasks to be specified as a range of values instead of a single value. Although the scheduling of elastic applications on dedicated resources has received considerable attention, there is limited work on scheduling of such applications in hierarchical and compositional settings. In this paper, we evaluate different earliest deadline first scheduling algorithms to schedule elastic applications in a minimum parallelism supply form reservation on a multiprocessor system. Our evaluation indicates that the proposed approach provides performance comparable to the current state-of-art algorithms for scheduling elastic applications on dedicated processors in terms of schedulability.
Shaik Mohammed Salman, Alessandro Vittorio Papadopoulos, Saad Mubeen, Thomas Nolte
J. Syst. Archit.2
2022 GMP: A Genetic Mission Planner for Heterogeneous Multirobot System Applications
abstract
The use of multiagent systems (MASs) in real-world applications keeps increasing, and diffuses into new domains, thanks to technological advances, increased acceptance, and demanding productivity requirements. Being able to automate the generation of mission plans for MASs is critical for managing complex missions in realistic settings. In addition, finding the right level of abstraction to represent any generic MAS mission is important for being able to provide general solution to the automated planning problem. In this article, we show how a mission for heterogeneous MASs can be cast as an extension of the traveling salesperson problem (TSP), and we propose a mixed-integer linear programming formulation. In order to solve this problem, a genetic mission planner (GMP), with a local plan refinement algorithm, is proposed. In addition, the comparative evaluation of CPLEX and GMP is presented in terms of timing and optimality of the obtained solutions. The algorithms are benchmarked on a proposed set of different problem instances. The results show that, in the presence of timing constraints, GMP outperforms CPLEX in the majority of test instances.
Branko Miloradovic, Baran Çürüklü, Mikael Ekström, Alessandro Vittorio Papadopoulos
IEEE Trans. Cybern.4
2022 STRETCH: Virtual Shared-Nothing Parallelism for Scalable and Elastic Stream Processing
abstract
Stream processing applications extract value from raw data through Directed Acyclic Graphs of data analysis tasks. Shared-nothing (SN) parallelism is the de-facto standard to scale stream processing applications. Given an application, SN parallelism ins9tantiates several copies of each analysis task, making each instance responsible for a dedicated portion of the overall analysis, and relies on dedicated queues to exchange data among connected instances. On the one hand, SN parallelism can scale the execution of applications both up and out since threads can run task instances within and across processes/nodes. On the other hand, its lack of sharing can cause unnecessary overheads and hinder the scaling up when threads operate on data that could be jointly accessed in shared memory. This trade-off motivated us in studying a way for stream processing applications to leverage shared memory and boost the scale up (before the scale out) while adhering to the widely-adopted and SN-based APIs for stream processing applications. We introduceSTRETCH, a framework that maximizes the scale up and offers instantaneous elastic reconfigurations (without state transfer) for stream processing applications. We propose the concept of Virtual Shared-Nothing (VSN) parallelism and elasticity and provide formal definitions and correctness proofs for the semantics of the analysis tasks supported bySTRETCH, showing they extend the ones found in common Stream Processing Engines. We also provide a fully implemented prototype and show thatSTRETCH's performance exceeds that of state-of-the-art frameworks such as Apache Flink and offers, to the best of our knowledge, unprecedented ultra-fast reconfigurations, taking less than 40 ms even when provisioning tens of new task instances.
Vincenzo Gulisano, Hannaneh Najdataei, Yiannis Nikolakopoulos, Alessandro Vittorio Papadopoulos, Marina Papatriantafilou, Philippas Tsigas
IEEE Trans. Parallel Distributed Syst.4
2021 On the Convolution Efficiency for Probabilistic Analysis of Real-Time Systems
abstract
This paper addresses two major problems in probabilistic analysis of real-time systems: space and time complexity of convolution of discrete random variables. For years, these two problems have limited the applicability of many methods for the probabilistic analysis of real-time systems, that rely on convolution as the main operation. Convolution in probabilistic analysis leads to a substantial space explosion and therefore space reductions may be necessary to make the problem tractable. However, the reductions lead to pessimism in the obtained probabilistic distributions, affecting the accuracy of the timing analysis. In this paper, we propose an optimal algorithm for down-sampling, which minimises the probabilistic expectation (i.e., the pessimism) in polynomial time. The second problem relates to the time complexity of the convolution between discrete random variables. It has been shown that quadratic time complexity of a single linear convolution, together with the space explosion of probabilistic analysis, limits its applicability for systems with a large number of tasks, jobs, and other analysed entities. In this paper, we show that the problem can be solved with a complexity of 𝒪(n log(n)), by proposing an algorithm that utilises circular convolution and vector space reductions. Evaluation results show several important improvements with respect to other state-of-the-art techniques.
Filip Markovic 0001, Alessandro Vittorio Papadopoulos, Thomas Nolte
ECRTS2
2021 Self-adapting Industrial Augmented Reality Applications with Proactive Dynamic Software Product Lines
abstract
Industrial Augmented Reality (IAR) is a key enabling technology for Industry 4.0. However, its adoption poses several challenges because it requires the execution of computing-intensive tasks in devices with poor computational resources, which contributes to a faster draining of the device batteries. Proactive self-adaptation techniques could overcome these problems that affect the quality of experience by optimizing computational resources and minimizing user disturbance. In this work, we propose to apply ProDSPL, a proactive Dynamic Software Product Line, for the self-adaptation of IAR applications to satisfy the quality requirements. PRODSPL is compared against MODAGAME, a multi-objective DSPL approach that uses a genetic algorithm to generate quasi-optimal feature model configurations at runtime. The evaluation with randomly generated feature models running on mobile devices shows that PRODSPL gives results closer to the Pareto optimal than MODAGAME.
Inmaculada Ayala, Mercedes Amor, Lidia Fuentes, Alessandro Vittorio Papadopoulos
ETFA4
2021 LETRA: Mapping Legacy Ethernet-Based Traffic into TSN Traffic Classes
abstract
This paper proposes a method to efficiently map the legacy Ethernet-based traffic into Time Sensitive Networking (TSN) traffic classes considering different traffic characteristics. Traffic mapping is one of the essential steps for industries to gradually move towards TSN, which in turn significantly mitigates the management complexity of industrial communication systems. In this paper, we first identify the legacy Ethernet traffic characteristics and properties. Based on the legacy traffic characteristics we presented a mapping methodology to map them into different TSN traffic classes. We implemented the mapping method as a tool, named Legacy Ethernet-based Traffic Mapping Tool or LETRA, together with a TSN traffic scheduling and performed a set of evaluations on different synthetic networks. The results show that the proposed mapping method obtains up to 90% improvement in the schedulability ratio of the traffic compared to an intuitive mapping method on a multi-switch network architecture.
Daniel Bujosa, Mohammad Ashjaei, Alessandro Vittorio Papadopoulos, Julián Proenza, Thomas Nolte
ETFA3
2021 Installation Order in Automatic Fabrication of Reinforcement Rebar Cages
abstract
Despite the significant development of automation in the manufacturing industry, the construction industry has not yet comparably gained much from automated processes. Fabrication of reinforcement rebar cages is one good example where automation has a limited application. Several challenges have to be tackled to introduce and take advantage of the automatic fabrication of reinforcement rebar cages. One important challenge is how and in what order the rebars should be installed one after another so that the fabrication of a reinforcement rebar cage is feasible. In this paper, we present our ongoing work towards proposing a method that gives a solution to finding the installation order of rebars.
Johan Relefors, Mahdi Momeni, Lars Pettersson, Alessandro Vittorio Papadopoulos, Thomas Nolte
ETFA4
2021 Scheduling Elastic Applications in Compositional Real-Time Systems
abstract
Many real-time applications have functional behaviour that requires variability in timing properties at runtime. The elastic task model provides a convenient mechanism to specify and encapsulate such variability and enables the modification of an application's periods during run-time to keep the application schedulable. Additionally, reservation-based scheduling techniques were proposed for the same purpose of taming unpredictability of timing variations, but with a different solution, i.e., by providing the spatial and temporal isolation for executing independent applications on the same hardware. In this paper, we combine the two approaches by proposing a two-level adaptive scheduling framework which is based on the elastic task model and the compositional framework based on the periodic resource model. The proposed framework minimises the number of requests for bandwidth adaption at the reservation (system) level and primarily enables schedulability by accounting for the application's elasticity by adjusting the periods. The motivation for this design choice is to rather localise the effect of the modifications within the application, without necessarily affecting all the applications at the system level compared to the changes made at the application level. The evaluation results show that the local application changes may often be enough to solve the problem of variability, significantly reducing the number of bandwidth adjustments, and therefore reducing the potential negative impact on all the applications of a system.
Shaik Mohammed Salman, Saad Mubeen, Filip Markovic 0001, Alessandro Vittorio Papadopoulos, Thomas Nolte
ETFA4
2021 REACT: Enabling Real-Time Container Orchestration
abstract
Fog and edge computing offer the flexibility and decentralized architecture benefits of cloud computing without suffering from the latency issues inherent in the cloud. This makes fog computing very attractive in real-time and safety-critical applications, especially if combined with container-based technologies. Whereas different orchestration systems are available to manage the container placement based on their resource demand, no orchestration system is considering real-time requirements for containerized applications. In this paper, we present the architecture and design of a real-time container orchestrator based on Kubernetes. Moreover, this paper defines metrics for the performance evaluation of real-time containers, and describes an initial model for allocating a mixture of real-time and non-real-time containers. We present an initial implementation of our real-time container extension and evaluate its feasibility on Linux-based systems.
Václav Struhár, Silviu S. Craciunas, Mohammad Ashjaei, Moris Behnam, Alessandro Vittorio Papadopoulos
ETFA5
2021 Adaptive Runtime Estimate of Task Execution Times using Bayesian Modeling
abstract
In the recent works that analyzed execution-time variation of real-time tasks, it was shown that such variation may conform to regular behavior. This regularity may arise from multiple sources, e.g., due to periodic changes in hardware or program state, program structure, inter-task dependence or inter-task interference. Such complexity can be better captured by a Markov Model, compared to the common approach of assuming independent and identically distributed random variables. However, despite the regularity that may be described with a Markov model, over time, the execution times may change, due to irregular changes in input, hardware state, or program state. In this paper, we propose a Bayesian approach to adapt the emission distributions of the Markov Model at runtime, in order to account for such irregular variation. A preprocessing step determines the number of states and the transition matrix of the Markov Model from a portion of the execution time sequence. In the preprocessing step, segments of the execution time trace with similar properties are identified and combined into clusters. At runtime, the proposed method switches between these clusters based on a Generalized Likelihood Ratio (GLR). Using a Bayesian approach, clusters are updated and emission distributions estimated. New clusters can be identified and clusters can be merged at runtime. The time complexity of the online step is $O(N^{2}+ NC)$ where N is the number of states in the Hidden Markov Model (HMM) that is fixed after the preprocessing step, and C is the number of clusters.
Anna Friebe, Filip Markovic 0001, Alessandro Vittorio Papadopoulos, Thomas Nolte
RTCSA3
2021 LOOPS: A Holistic Control Approach for Resource Management in Cloud Computing
abstract
In cloud computing model, resource sharing introduces major benefits for improving resource utilization and total cost of ownership, but it can create technical challenges on the running performance. In practice, orchestrators are required to allocate sufficient physical resources to each Virtual Machine (VM) to meet a set of predefined performance goals. To ensure a specific service level objective, the orchestrator needs to be equipped with a dynamic tool for assigning computing resources to each VM, based on the run-time state of the target environment. To this end, we present LOOPS, a multi-loop control approach, to allocate resources to VMs based on the service level agreement (SLA) requirements and the run-time conditions. LOOPS is mainly composed of one essential unit to monitor VMs, and three control levels to allocate resources to VMs based on requests from the essential node. A tailor-made controller is proposed with each level to regulate contention among collocated VMs, to reallocate resources if required, and to migrate VMs from one host to another. The three levels work together to meet the required SLA. The experimental results have shown that the proposed approach can meet applications' performance goals by assigning the resources required by cloud-based applications.
Auday Aldulaimy, Javid Taheri, Alessandro Vittorio Papadopoulos, Thomas Nolte
ICPE3
2021 A systematic methodology to migrate complex real-time software systems to multi-core platforms
abstract
This paper proposes a systematic three-stage methodology for migrating complex real-time industrial software systems from single-core to multi-core computing platforms. Single-core platforms have limited computational capabilities that prevent integration of computationally demanding applications such as image processing within the existing system. Modern multi-core processors offer a promising solution to address these limitations by providing increased computational power and allowing parallel execution of different applications within the system. However, the transition from traditional single-core to contemporary multi-core computing platforms is non-trivial and requires a systematic and well-defined migration process. This paper reviews some of the existing migration methods and provides a systematic multi-phase migration process with emphasis on software architecture recovery and transformation to explicitly address the timing and dependability attributes expected of industrial software systems. The methodology was evaluated using a survey-based approach and the results indicate that the presented methodology is feasible, useable and useful for real-time industrial software systems.
Shaik Mohammed Salman, Alessandro Vittorio Papadopoulos, Saad Mubeen, Thomas Nolte
J. Syst. Archit.2
2021 ProDSPL: Proactive self-adaptation based on Dynamic Software Product Lines
Inmaculada Ayala, Alessandro Vittorio Papadopoulos, Mercedes Amor, Lidia Fuentes
J. Syst. Softw.2
2021 Methodological Principles for Reproducible Performance Evaluation in Cloud Computing
abstract
The rapid adoption and the diversification of cloud computing technology exacerbate the importance of a sound experimental methodology for this domain. This work investigates how to measure and report performance in the cloud, and how well the cloud research community is already doing it. We propose a set of eight important methodological principles that combine best-practices from nearby fields with concepts applicable only to clouds, and with new ideas about the time-accuracy trade-off. We show how these principles are applicable using a practical use-case experiment. To this end, we analyze the ability of the newly released SPEC Cloud IaaS benchmark to follow the principles, and showcase real-world experimental studies in common cloud environments that meet the principles. Last, we report on a systematic literature review including top conferences and journals in the field, from 2012 to 2017, analyzing if the practice of reporting cloud performance measurements follows the proposed eight principles. Worryingly, this systematic survey and the subsequent two-round human reviews, reveal that few of the published studies follow the eight experimental principles. We conclude that, although these important principles are simple and basic, the cloud community is yet to adopt them broadly to deliver sound measurement of cloud environments.
Alessandro Vittorio Papadopoulos, Laurens Versluis, André Bauer 0001, Nikolas Herbst, Jóakim von Kistowski, Ahmed Ali-Eldin, Cristina L. Abad, José Nelson Amaral, Petr Tuma 0001, Alexandru Iosup
IEEE Trans. Software Eng.1
2020 IoT and Fog Analytics for Industrial Robot Applications
abstract
The rapid development of IoT, cloud and fog computing has increased the potential for developing smart services for IoT devices. Such services require not only connectivity and high computing capacity, but also fast response time and throughput of inferencing results. In this paper we present our ongoing work, investigating the potential for implementing smart services in the context of industrial robot applications with focus on analytic inferencing on fog and cloud computing platforms. We review different use cases that we have found in the literature and we divide them into two suggested categories, "distributed deep models" and "distributed interconnected models". We analyze the characteristics of IoT data in industrial robot applications and present two concrete use cases of smart services where inferencing in a fog and a cloud architecture, respectively, is needed. We also reason about important considerations and design decisions for the development process of analytic services.
Anders Lager, Alessandro Vittorio Papadopoulos, Thomas Nolte
ETFA2
2020 Clock Synchronization in Integrated TSN-EtherCAT Networks
abstract
Moving towards new technologies, such as Time Sensitive Networking (TSN), in industries should be gradual with a proper integration process instead of replacing the existing ones to make it beneficial in terms of cost and performance. Within this context, this paper identifies the challenges of integrating a legacy EtherCAT network, as a commonly used technology in the automation domain, into a TSN network. We show that clock synchronization plays an essential role when it comes to EtherCAT-TSN network integration with important requirements. We propose a clock synchronization mechanism based on the TSN standards to obtain a precise synchronization among EtherCAT nodes, resulting to an efficient data transmission. Based on a formal verification framework using UPPAAL tool we show that the integrated EtherCAT-TSN network with the proposed clock synchronization mechanism achieves at least 3 times higher synchronization precision compared to not using any synchronization.
Daniel Bujosa, Daniel Hallmans, Mohammad Ashjaei, Alessandro Vittorio Papadopoulos, Julián Proenza, Thomas Nolte
ETFA4
2020 Enabling Fog-based Industrial Robotics Systems
abstract
Low latency and on demand resource availability enable fog computing to host industrial applications in a cloud like manner. One industrial domain which stands to benefit from the advantages of fog computing is robotics. However, the challenges in developing and implementing a fog-based robotic system are manifold. To illustrate this, in this paper we discuss a system involving robots and robot cells at a factory level, and then highlight the main building blocks necessary for achieving such functionality in a fog-based system. Further, we elaborate on the challenges in implementing such an architecture, with emphasis on resource virtualization, memory interference management, real-time communication and the system scalability, dependability and safety. We then discuss the challenges from a system perspective where all these aspects are interrelated.
Shaik Mohammed Salman, Václav Struhár, Zeinab Bakhshi, Van-Lan Dao, Nitin Desai, Alessandro Vittorio Papadopoulos, Thomas Nolte, Vasileios Karagiannis, Stefan Schulte 0002, Alexandre Venito, Gerhard Fohler
ETFA6
2020 Modeling the Willingness to Interact in Cooperative Multi-robot Systems
abstract
When multiple robots are required to collaborate in order to accomplish a specific task, they need to be coordinated in order to operate efficiently. To allow for scalability and robustness, we pro ...
Mirgita Frasheri, Lukas Esterle, Alessandro Vittorio Papadopoulos
ICAART (1)3
2020 Heartbeat Bully: Failure Detection and Redundancy Role Selection for Network-Centric Controller
abstract
High availability and reliability are fundamental for distributed control systems in the automation industry. Redundancy solutions, with duplicated hardware, is the common way to increase availability. With the advent of Industry 4.0, the automation industry is undergoing a paradigm shift; a peer-to-peer mesh oriented architecture is replacing the traditional hierarchical automation pyramid. With generic computational power provided anywhere in the cloud - device continuum, the conventional control centric solutions are becoming obsolete. The paradigm shift imposes new challenges and possibilities on the redundancy solutions used. We present and evaluate a hardware-agnostic algorithm suitable for failure detection and redundancy role selection in the new automation paradigm. The algorithm is modeled, evaluated and validated with the model checking tool UPPAAL.
Bjarne Johansson, Mats Rågberger, Alessandro Vittorio Papadopoulos, Thomas Nolte
IECON3
2020 A Systematic Migration Methodology for Complex Real-time Software Systems
abstract
This paper provides a systematic three-stage methodology for migrating complex real-time industrial software systems from single-core to multi-core computing platforms. Single-core platforms have limited computational capabilities that prevent integration of computationally demanding applications such as image processing within the existing system. Modern multi-core processors provide increased computing capacity and allow the parallel execution of different applications within the system. However, this transition is non-trivial and requires a systematic and well-defined migration process. This paper reviews some of the existing migration methods and provides a systematic multi-phase migration process with emphasis on software architecture recovery and transformation to explicitly address the timing and dependability attributes expected of industrial software systems.
Shaik Mohammed Salman, Alessandro Vittorio Papadopoulos, Saad Mubeen, Thomas Nolte
ISORC2
2020 Identification and Validation of Markov Models with Continuous Emission Distributions for Execution Times
abstract
It has been shown that in some robotic applications, where the execution times cannot be assumed to be independent and identically distributed, a Markov Chain with discrete emission distributions can be an appropriate model. In this paper we investigate whether execution times can be modeled as a Markov Chain with continuous Gaussian emission distributions. The main advantage of this approach is that the concept of distance is naturally incorporated. We propose a framework based on Hidden Markov Model (HMM) methods that 1) identifies the number of states in the Markov Model from observations and fits the Markov Model to observations, and 2) validates the proposed model with respect to observations. Specifically, we apply a tree-based cross-validation approach to automatically find a suitable number of states in the Markov model. The estimated models are validated against observations, using a data consistency approach based on log likelihood distributions under the proposed model. The framework is evaluated using two test cases executed on a Raspberry Pi Model 3B+ single-board computer running Arch Linux ARM patched with PREEMPT_RT. The first is a simple test program where execution times intentionally vary according to a Markov model, and the second is a video decompression using the ffmpeg program. The results show that in these cases the framework identifies Markov Chains with Gaussian emission distributions that are valid models with respect to the observations.
Anna Friebe, Alessandro Vittorio Papadopoulos, Thomas Nolte
RTCSA2
2020 Packet priority assignment for wireless control systems of multiple physical systems
Wenchen Wang, Daniel Mossé, Alessandro Vittorio Papadopoulos
J. Syst. Archit.3
2019 Towards a Framework for Safe and Secure Adaptive Collaborative Systems
abstract
Real-time adaptive systems are complex systems capable to adapt their behavior to changing conditions in the environment, and/or internal state changes. Highly dynamic and possibly unpredictable environments, and uncertain operating conditions call for new paradigms of software design, and run-time adaptation mechanisms, to overcome the lack of knowledge at design time. Main application areas include vehicles or robots that need to collaborate to achieve a common task, e.g., minimize fuel consumption, moving objects at a construction site, or performing a set of operations in a factory. Moreover, these vehicles or robots need to interact and possibly collaborate with humans in a safe way, e.g., avoiding accidents or collisions, and prevent hazardous situations that may harm humans and/or machines. % This paper proposes a framework for developing safe and secure adaptive collaborative systems, with run-time guarantees. To enable this, our focus is on requirement engineering and safety assurance techniques to capture the specific safety and security properties for the collaborative system, and to provide an assurance case guaranteeing that the system is sufficiently safe. Moreover, the paper proposes an architecture and behavioral models to analyze the requirements at run-time. Finally, we design a suitable deployment platform to perform the run-time analysis and planning while guaranteeing the real-time constraints.
Aida Causevic, Alessandro Vittorio Papadopoulos, Marjan Sirjani
COMPSAC (2)2
2019 Classification of PROFINET I/O Configurations utilizing Neural Networks
abstract
In process automation installations, the I/O system connect the field devices to the process controller over a fieldbus, a reliable, real-time capable communication link with signal values cyclical being exchanged with a 10-100 millisecond rate. If a deviation from intended behaviour occurs, analyzing the potentially vast data recordings from the field can be a time consuming and cumbersome task for an engineer. For the engineer to be able to get a full understanding of the problem, knowledge of the used I/O configuration is required. In the problem report, the configuration description is sometimes missing. In such cases it is difficult to use the recorded data for analysis of the problem.In this paper we present our ongoing work towards using neural network models as assistance in the interpretation of an industrial fieldbus communication recording. To show the potential of such an approach we present an example using an industrial setup where fieldbus data is collected and classified. In this context we present an evaluation of the suitability of different neural net configurations and sizes for the problem at hand.
Bjarne Johansson, Björn Leander, Aida Causevic, Alessandro Vittorio Papadopoulos, Thomas Nolte
ETFA4
2019 Towards Reactive Robot Applications in Dynamic Environments
abstract
Traditionally, industrial robots have been deployed in fairly static environments, to perform highly dedicated tasks. These robots perform with very high precision and throughput. However, nowadays there is an increasing demand for utilizing robots in more dynamic environments, also performing more flexible and less specialized operations - high mix/low volume. Both traditional industrial robots and force-limited robots are used in collaborative, dynamic environments. Such robot applications introduce new challenges when it comes to efficiency and robustness. In this paper, we propose an architecture for reactive multi-robot applications in the context of dynamic environments, and we analyze the main research challenges that must be tackled for its realization. A logistics use case, with robots picking customer orders from the shelves of a warehouse, is used as a running example to support the description of the key challenges.
Anders Lager, Giacomo Spampinato, Alessandro Vittorio Papadopoulos, Thomas Nolte
ETFA3
2019 Towards Automated Installation of Reinforcement Using Industrial Robots
abstract
The construction industry is today among the least automated industries with a long tradition of utilizing manual labour. Despite the potential benefits of automation, only a few examples of using robots to automate (parts of) construction have been presented over the past years. In this paper we present our ongoing work towards automated installation of reinforcement, a traditionally very heavy and labour intensive work. We use industrial robots and we discuss the potential benefits and challenges of such robotic automation in construction. Our overall goal is to achieve a fully automated robotic solution for flexible serial production of custom made non-identical reinforcement cages. In the paper we highlight and analyse the main challenges that must be addressed in order to reach a functioning and efficient solution.
Johan Relefors, Mahdi Momeni, Lars Pettersson, Erik Hellström, Anders Thunell, Alessandro Vittorio Papadopoulos, Thomas Nolte
ETFA6
2019 Probabilistic Timing Analysis of a Periodic Task on a Microcontroller
abstract
In this paper we present our ongoing work towards a realistic probabilistic timing analysis of embedded software systems subject to timing requirements. In order to provide such an analysis that captures necessary and important behavioural features of the software system under analysis, including the underlying platform, we have implemented a real-time system running on a Rasberry Pi microcontroller on which we have performed a series of experiments and measurements. The results so far suggest a new model for analysis that captures more detailed behaviour and consequently provides a more accurate and correct probabilistic analysis.
Jonathan Thörn, Najda Vidimlic, Anna Friebe, Alessandro Vittorio Papadopoulos, Thomas Nolte
ETFA4
2019 Extended Colored Traveling Salesperson for Modeling Multi-Agent Mission Planning Problems
abstract
In recent years, multi-agent systems have been widely used in different missions, ranging from underwater to airborne. A mission typically involves a large number of agents and tasks, making it very hard for the human operator to create a good plan. A search for an optimal plan may take too long, and it is hard to make a time estimate of when the planner will finish. A genetic algorithm based planner is proposed in order to overcome this issue. The contribution of this paper is threefold. First, an Integer Linear Programming (ILP) formulation of a novel Extensive Colored Traveling Salesperson Problem (ECTSP) is given. Second, a new objective function suitable for multi-agent mission planning problems is proposed. Finally, a reparation algorithm to allow usage of common variation operators for ECTSP has been developed.
Branko Miloradovic, Baran Çürüklü, Mikael Ekström, Alessandro Vittorio Papadopoulos
ICORES4
2019 DART: Dynamic Bandwidth Distribution Framework for Virtualized Software Defined Networks
abstract
In this paper we address a network architecture that uses a combination of network virtualization and software defined networking in order to reduce complexity of network management and at the same time support high quality of service. Within this network architecture, we propose a framework to be able to dynamically distribute the network bandwidth to various services such that the network resources are utilized efficiently. In many industrial domains, multiple services may use the same hardware platform for the sake of a better resource utilization. Therefore, bandwidth distribution among the services should be done in an efficient way during runtime. We also develop an admission control in this framework which dynamically coordinates the bandwidth distributions based on requested quality of services. We show the applicability of the proposed framework by implementing it on a common SDN controller. Moreover, we conduct a set of experiments to show the performance of the proposed framework.
Václav Struhár, Mohammad Ashjaei, Moris Behnam, Silviu S. Craciunas, Alessandro Vittorio Papadopoulos
IECON5
2019 Packet Priority Assignment for Wireless Control Systems of multiple Physical Systems
abstract
Wireless control systems (WCSs) have gained much attention lately, due to their easy deployment and flexibility compared to wired control systems. However, this comes at the cost of possibly increased network delay and packet losses, that can significantly impact the control system performance, and possibly its stability. Such problems become even more relevant if the network is shared among different control systems, and thus becomes a scarce resource, like in Industrial Internet of Things applications. In this paper, we describe how to assign packet priorities dynamically when there are many physical systems sharing a given network, aiming at minimizing the performance degradation of the WCS. Towards that, we present a network model including both delay and packet losses, both of which are very important for the control system performance. Our solution is evaluated over two different use cases to show the generality of the approach: the WCS for a set of inverted pendula, and the WCS for small modular reactors in a nuclear power plant. The results show that the proposed approach allows for a more stable performance even in presence of highly nonlinear systems, sensitive to time-varying delays, as well as in presence of high network interference.
Wenchen Wang, Daniel Mossé, Alessandro Vittorio Papadopoulos
ISORC3
2019 TAMER: Task Allocation in Multi-robot Systems Through an Entity-Relationship Model
Branko Miloradovic, Mirgita Frasheri, Baran Çürüklü, Mikael Ekström, Alessandro Vittorio Papadopoulos
PRIMA5
2019 Work-in-Progress: Validation of Probabilistic Timing Models of a Periodic Task with Interference - A Case Study
abstract
Probabilistic timing analysis techniques have been proposed for real-time systems to remedy the problems that deterministic estimates of the task's Worst-Case Execution Time and Worst-Case Response-Time can be both intractable and overly pessimistic. Often, assumptions are made that a task's response time and execution time probability distributions are independent of the other tasks. This assumption may not hold in real systems. In this paper, we analyze the timing behavior of a simple periodic task on a Raspberry Pi model 3 running Arch Linux ARM. In particular, we observe and analyze the distributions of wake-up latencies and execution times for the sequential jobs released by a simple periodic task. We observe that the timing behavior of jobs is affected by release events during the job's execution time, and of other processes running in between subsequent jobs of the periodic task. Using a data consistency approach we investigate whether it is reasonable to model the timing distribution of jobs affected by release events and intermediate processes as translations of the empirical timing distribution of non-affected jobs. According to the analysis, this paper shows that a translated distribution model of non-affected jobs is invalid for the execution time distribution of jobs affected by intermediate processes. Regarding the wake-up latency distribution with intermediate processes, a translated distribution model is improbable, but cannot be completely ruled out.
Anna Friebe, Alessandro Vittorio Papadopoulos, Thomas Nolte
RTSS2
2018 AdaptMC: A Control-Theoretic Approach for Achieving Resilience in Mixed-Criticality Systems
abstract
A system is said to be resilient if slight deviations from expected behavior during run-time does not lead to catastrophic degradation of performance: minor deviations should result in no more than minor performance degradation. In mixed-criticality systems, such degradation should additionally be criticality-cognizant. The applicability of control theory is explored for the design of resilient run-time scheduling algorithms for mixed-criticality systems. Recent results in control theory have shown how appropriately designed controllers can provide guaranteed service to hard-real-time servers; this prior work is extended to allow for such guarantees to be made concurrently to multiple criticality-cognizant servers. The applicability of this approach is explored via several experimental simulations in a dual-criticality setting. These experiments demonstrate that our control-based run-time schedulers can be synthesized in such a manner that bounded deviations from expected behavior result in the high-criticality server suffering no performance degradation and the lower-criticality one, bounded performance degradation.
Alessandro Vittorio Papadopoulos, Enrico Bini, Sanjoy Baruah, Alan Burns 0001
ECRTS1
2018 Fog computing for adaptive human-robot collaboration: work-in-progress
abstract
Fog computing is an emerging technology that enables the design of novel time sensitive industrial applications. This new computing paradigm also opens several new research challenges in different scientific domains, ranging from computer architectures to networks, from robotics to real-time systems. In this paper, we present a use case in the human-robot collaboration domain, and we identify some of the most relevant research challenges.
Václav Struhár, Alessandro Vittorio Papadopoulos, Moris Behnam
EMSOFT2
2018 Hybrid Adaptive Checkpointing for Virtual Machine Fault Tolerance
abstract
Active Virtual Machine (VM) replication is an application independent and cost-efficient mechanism for high availability and fault tolerance, with several recently proposed implementations based on checkpointing. However, these methods may suffer from large impacts on application latency, excessive resource usage overheads, and/or unpredictable behavior for varying workloads. To address these problems, we propose a hybrid approach through a Proportional-Integral (PI) controller to dynamically switch between periodic and on-demand check-pointing. Our mechanism automatically selects the method that minimizes application downtime by adapting itself to changes in workload characteristics. The implementation is based on modifications to QEMU, LibVirt, and OpenStack, to seamlessly provide fault tolerant VM provisioning and to enable the controller to dynamically select the best checkpointing mode. Our evaluation is based on experiments with a video streaming application, an e-commerce benchmark, and a software development tool. The experiments demonstrate that our adaptive hybrid approach improves both application availability and resource usage compared to static selection of a checkpointing method, with application performance gains and neglectable overheads.
Abel Souza, Alessandro Vittorio Papadopoulos, Luis Tomás, Johan Tordsson
IC2E2
2018 Engineering Self-Adaptive Software Systems: From Requirements to Model Predictive Control
abstract
Self-adaptive software systems monitor their operation and adapt when their requirements fail due to unexpected phenomena in their environment. This article examines the case where the environment changes dynamically over time and the chosen adaptation has to take into account such changes. In control theory, this type of adaptation is known as Model Predictive Control and comes with a well-developed theory and myriad successful applications. The article focuses on modeling the dynamic relationship between requirements and possible adaptations. It then proposes a controller that exploits this relationship to optimize the satisfaction of requirements relative to a cost function. This is accomplished through a model-based framework for designing self-adaptive software systems that can guarantee a certain level of requirements satisfaction over time by dynamically composing adaptation strategies when necessary. The proposed framework is illustrated and evaluated through two simulated systems, namely, the Meeting-Scheduling exemplar and an E-Shop.
Konstantinos Angelopoulos, Alessandro Vittorio Papadopoulos, Vítor E. Silva Souza, John Mylopoulos
ACM Trans. Auton. Adapt. Syst.2
2017 KPI-agnostic Control for Fine-Grained Vertical Elasticity
abstract
Applications hosted in the cloud have become indispensable in several contexts, with their performance often being key to business operation and their running costs needing to be minimized. To minimize running costs, most modern virtualization technologies such as Linux Containers, Xen, and KVM offer powerful resource control primitives for individual provisioning - that enable adding or removing of fraction of cores and/or megabytes of memory for as short as few seconds. Despite the technology being ready, there is a lack of proper techniques for fine-grained resource allocation, because there is an inherent challenge in determining the correct composition of resources an application needs, with varying workload, to ensure deterministic performance. This paper presents a control-based approach for the management of multiple resources, accounting for the resource consumption, together with the application performance, enabling fine-grained vertical elasticity. The control strategy ensures that the application meets the target performance indicators, consuming as less resources as possible. We carried out an extensive set of experiments using different applications - interactive with response-time requirements, as well as noninteractive with throughput desires - by varying the workload mixes of each application over time. The results demonstrate that our solution precisely provides guaranteed performance while at the same time avoiding both resource over-and underprovisioning.
Ewnetu Bayuh Lakew, Alessandro Vittorio Papadopoulos, Martina Maggio, Cristian Klein, Erik Elmroth
CCGrid2
2017 Distributed Approach to the Holistic Resource Management of a Mobile Cloud Network
abstract
The Mobile Cloud Network is an emerging cost and capacity heterogeneous distributed cloud topological paradigm that aims to remedy the application performance constraints imposed by centralised cloud infrastructures. A centralised cloud infrastructure and the adjoining Telecom network will struggle to accommodate the exploding amount of traffic generated by forthcoming highly interactive applications. Cost effectively managing a Mobile Cloud Network computing infrastructure while meeting individual application's performance goals is non-trivial and is at the core of our contribution. Due to the scale of a Mobile Cloud Network, a centralised approach is infeasible. Therefore, in this paper a distributed algorithm that addresses these challenges is presented. The presented approach works towards meeting individual application's performance objectives, constricting system-wide operational cost, and mitigating resource usage skewness. The presented distributed algorithm does so by iteratively and independently acting on the objectives of each component with a common heuristic objective function. Systematic evaluations reveal that the presented algorithm quickly converges and performs near optimal in terms of system-wide operational cost and application performance, and significantly outperforms similar naïve and random methods.
William Tärneberg, Alessandro Vittorio Papadopoulos, Amardeep Mehta, Johan Tordsson, Maria Kihl
ICFEC2
2017 Automated control of multiple software goals using multiple actuators
abstract
Modern software should satisfy multiple goals simultaneously: it should provide predictable performance, be robust to failures, handle peak loads and deal seamlessly with unexpected conditions and changes in the execution environment. For this to happen, software designs should account for the possibility of runtime changes and provide formal guarantees of the software's behavior. Control theory is one of the possible design drivers for runtime adaptation, but adopting control theoretic principles often requires additional, specialized knowledge. To overcome this limitation, automated methodologies have been proposed to extract the necessary information from experimental data and design a control system for runtime adaptation. These proposals, however, only process one goal at a time, creating a chain of controllers. In this paper, we propose and evaluate the first automated strategy that takes into account multiple goals without separating them into multiple control strategies. Avoiding the separation allows us to tackle a larger class of problems and provide stronger guarantees. We test our methodology's generality with three case studies that demonstrate its broad applicability in meeting performance, reliability, quality, security, and energy goals despite environmental or requirements changes.
Martina Maggio, Alessandro Vittorio Papadopoulos, Antonio Filieri, Henry Hoffmann
ESEC/SIGSOFT FSE2
2017 An Experimental Performance Evaluation of Autoscaling Policies for Complex Workflows
abstract
Simplifying the task of resource management and scheduling for customers, while still delivering complex Quality-of-Service (QoS), is key to cloud computing. Many autoscaling policies have been proposed in the past decade to decide on behalf of cloud customers when and how to provision resources to a cloud application utilizing cloud elasticity features. However, in prior work, when a new policy is proposed, it is seldom compared to the state-of-the-art, and is often compared only to static provisioning using a predefined QoS target. This reduces the ability of cloud customers and of cloud operators to choose and deploy an autoscaling policy. In our work, we conduct an experimental performance evaluation of autoscaling policies, using as application model workflows, a commonly used formalism for automating resource management for applications with well-defined yet complex structure. We present a detailed comparative study of general state-of-the-art autoscaling policies, along with two new workflow-specific policies. To understand the performance differences between the 7 policies, we conduct various forms of pairwise and group comparisons. We report both individual and aggregated metrics. Our results highlight the trade-offs between the suggested policies, and thus enable a better understanding of the current state-of-the-art.
Alexey Ilyushkin, Ahmed Ali-Eldin, Nikolas Herbst, Alessandro Vittorio Papadopoulos, Bogdan Ghit, Dick H. J. Epema, Alexandru Iosup
ICPE4
2017 FLOPSYNC-QACS: Quantization-aware clock synchronization for wireless sensor networks
Federico Terraneo, Alessandro Vittorio Papadopoulos, Alberto Leva, Maria Prandini
J. Syst. Archit.2
2017 Control Strategies for Self-Adaptive Software Systems
abstract
The pervasiveness and growing complexity of software systems are challenging software engineering to design systems that can adapt their behavior to withstand unpredictable, uncertain, and continuously changing execution environments. Control theoretical adaptation mechanisms have received growing interest from the software engineering community in the last few years for their mathematical grounding, allowing formal guarantees on the behavior of the controlled systems. However, most of these mechanisms are tailored to specific applications and can hardly be generalized into broadly applicable software design and development processes. This article discusses a reference control design process, from goal identification to the verification and validation of the controlled system. A taxonomy of the main control strategies is introduced, analyzing their applicability to software adaptation for both functional and nonfunctional goals. A brief extract on how to deal with uncertainty complements the discussion. Finally, the article highlights a set of open challenges, both for the software engineering and the control theory research communities.
Antonio Filieri, Martina Maggio, Konstantinos Angelopoulos, Nicolás D'Ippolito, Ilias Gerostathopoulos, Andreas B. Hempel, Henry Hoffmann, Pooyan Jamshidi, Evangelia Kalyvianaki, Cristian Klein, Filip Krikava, Sasa Misailovic, Alessandro Vittorio Papadopoulos, Suprio Ray, Amir Molzam Sharifloo, Stepan Shevtsov, Mateusz Ujma, Thomas Vogel 0001
ACM Trans. Auton. Adapt. Syst.13
2016 Which Cloud Auto-Scaler Should I Use for my Application?: Benchmarking Auto-Scaling Algorithms
abstract
Rapid elasticity is one of the essential characteristics of cloud computing identified by NIST [17]. Elasticity allows resources to be provisioned and released to scale rapidly out ward and in ward according to demand. Tens -- if not hundreds -- of algorithms have been proposed in the literature to automatically achieve elastic provisioning [15, 23, 14, 21, 13, 20, 6, 12, 16, 10]. These algorithms are typically referred to as elasticity algorithms, dynamic provisioning techniques or autoscalers. While trying to solve the same problem, sometimes with differing assumption, many of these algorithms are either compared to static provisioning or to a predefined QoS target, e.g., predefined response time target, with very little -- or no -- comparison to previously published work. This reduces the ability of an application owner or a cloud operator to choose and deploy a suitable algorithm from the literature. Many of these algorithms have been tested with one single -- real or synthetic -- workload in a specific use-case [13, 14, 10]. While all published algorithms are shown to work in the specific use-case they were designed for with the, typically short, workloads tested with, it is seldom the case that the real scenarios will be any thing close to the test cases for which the algorithms are shown to work. Bursts occur in workloads occasionally. Workload dynamics change over time and the load-mix of an application significantly affects how provisioning should be done [21].
Ahmed Ali-Eldin, Alexey Ilyushkin, Bogdan Ghit, Nikolas Herbst, Alessandro Vittorio Papadopoulos, Alexandru Iosup
ICPE5
2015 Reverse Flooding: Exploiting Radio Interference for Efficient Propagation Delay Compensation in WSN Clock Synchronization
abstract
Clock synchronization is a necessary component in modern distributed systems, especially Wireless Sensor Networks (WSNs). Despite the great effort and the numerous improvements, the existing synchronization schemes do not yet address the cancellation of propagation delays. Up to a few years ago, this was not perceived as a problem, because the time-stamping precision was a more limiting factor for the accuracy achievable with a synchronization scheme. However, the recent introduction of efficient flooding schemes based on constructive interference has greatly improved the achievable accuracy, to the point where propagation delays can effectively become the main source of error. In this paper, we propose a method to estimate and compensate for the network propagation delays. Our proposal does not require to maintain a spanning tree of the network, and exploits constructive interference even to transmit packets whose content are slightly different. To show the validity of the approach, we implemented the propagation delay estimator on top of the FLOPSYNC-2 synchronization scheme. Experimental results prove the feasibility of measuring propagation delays using off-the-shelf microcontrollers and radio transceivers, and show how the proposed solution allows to achieve sub-microsecond clock synchronization even for networks where propagation delays are significant.
Federico Terraneo, Alberto Leva, Silvano Seva, Martina Maggio, Alessandro Vittorio Papadopoulos
RTSS5
2015 Hard real-time guarantees in feedback-based resource reservations
Alessandro Vittorio Papadopoulos, Martina Maggio, Alberto Leva, Enrico Bini
Real Time Syst.1
2014 Model reduction of switched affine systems: a method based on balanced truncation and randomized optimization
abstract
This paper proposes an approach to build a reduced order model for a Switched Affine (SA) system. The main idea is to transform the SA system into an equivalent Switched Linear (SL) system with state reset, and then apply balanced truncation to each mode and redefine the reset maps so as to best reproduce the free evolution of the system output. A randomized method is proposed for order selection in the case when the input is stochastic and one is interested in reproducing the output of the original SA system over a finite time-horizon. The performance of the approach is shown on a benchmark example.
Alessandro Vittorio Papadopoulos, Maria Prandini
HSCC1
2014 FLOPSYNC-2: Efficient Monotonic Clock Synchronisation
abstract
Time synchronisation is crucial for distributed systems, and particularly for Wireless Sensor Networks (WSNs), where each node is executing concurrent operations to achieve a real-time objective. However, synchronisation is quite difficult to achieve in WSNs, due to the unpredictable deployment conditions and to physical effects like thermal stress, that cause drifts in the local node clocks. As a result, state-of-the-art synchronisation schemes do not guarantee monotonicity of the nodes clock, or are relying on external hardware assistance. In this paper we present FLOPSYNC-2, a scheme to synchronise the clocks of multiple nodes in a WSN, requiring no additional hardware, and based on the application of control-theoretical principles. The scheme guarantees low overhead, low power consumption and synchronisation with clock monotonicity. We propose an implementation of FLOPSYNC-2 on top of the microcontroller operating system Miosix, and prove the validity of our claims with several-days-long experiments on an eight-hop network. The experimental results show that the average clock difference among nodes is limited to a hundred of ns, with a sub-microsecond standard deviation. By introducing a suitable power model, we also prove that synchronisation is achieved with a sub-μA consumption overhead.
Federico Terraneo, Luigi Rinaldi, Martina Maggio, Alessandro Vittorio Papadopoulos, Alberto Leva
RTSS4
2014 Improving Cloud Service Resilience Using Brownout-Aware Load-Balancing
abstract
We focus on improving resilience of cloud services (e.g., e-commerce website), when correlated or cascading failures lead to computing capacity shortage. We study how to extend the classical cloud service architecture composed of a load-balancer and replicas with a recently proposed self-adaptive paradigm called brownout. Such services are able to reduce their capacity requirements by degrading user experience (e.g., disabling recommendations). Combining resilience with the brownout paradigm is to date an open practical problem. The issue is to ensure that replica self-adaptivity would not confuse the load-balancing algorithm, overloading replicas that are already struggling with capacity shortage. For example, load-balancing strategies based on response times are not able to decide which replicas should be selected, since the response times are already controlled by the brownout paradigm. In this paper we propose two novel brownout-aware load-balancing algorithms. To test their practical applicability, we extended the popular lighttpd web server and load-balancer, thus obtaining a production-ready implementation. Experimental evaluation shows that the approach enables cloud services to remain responsive despite cascading failures. Moreover, when compared to Shortest Queue First (SQF), believed to be near-optimal in the non-adaptive case, our algorithms improve user experience by 5%, with high statistical significance, while preserving response time predictability.
Cristian Klein, Alessandro Vittorio Papadopoulos, Manfred Dellkrantz, Jonas Durango, Martina Maggio, Karl-Erik Årzén, Francisco Hernández-Rodriguez, Erik Elmroth
SRDS2
2013 Generation of human walking paths
abstract
This work investigates the way humans plan their paths in a goal-directed motion. The person can be viewed as an optimal controller that plans the path minimizing a certain (unknown) cost function. Taking this viewpoint, the problem can be formulated as an inverse optimal control one, i.e., starting from control and state trajectories we want to figure out the cost function used by a person while planning the path. To test the envisaged ideas, a set of walking paths of different volunteers were recorded using a motion capture facility. The collected data have been used to compare a solution to the inverse control problem coming from the literature to a novel one. The obtained results, ranked using the discrete Fréchet distance, show the effectiveness of the proposed approach.
Alessandro Vittorio Papadopoulos, Luca Bascetta, Gianni Ferretti
IROS1
2012 Comparison of Decision-Making Strategies for Self-Optimization in Autonomic Computing Systems
abstract
Autonomic computing systems are capable of adapting their behavior and resources thousands of times a second to automatically decide the best way to accomplish a given goal despite changing environmental conditions and demands. Different decision mechanisms are considered in the literature, but in the vast majority of the cases a single technique is applied to a given instance of the problem. This article proposes a comparison of some state of the art approaches for decision making, applied to a self-optimizing autonomic system that allocates resources to a software application. A variety of decision mechanisms, from heuristics to control-theory and machine learning, are investigated. The results obtained with these solutions are compared by means of case studies using standard benchmarks. Our results indicate that the most suitable decision mechanism can vary depending on the specific test case but adaptive and model predictive control systems tend to produce good performance and may work best in a priori unknown situations.
Martina Maggio, Henry Hoffmann, Alessandro Vittorio Papadopoulos, Jacopo Panerati, Marco D. Santambrogio, Anant Agarwal, Alberto Leva
ACM Trans. Auton. Adapt. Syst.3
2011 Looking for "Good" Recommendations: A Comparative Evaluation of Recommender Systems
Paolo Cremonesi, Franca Garzotto, Sara Negro, Alessandro Vittorio Papadopoulos, Roberto Turrin
INTERACT (3)4