Stefan Nastic

dblp:123/3328 · DBLP profile ↗
← Back
34ranked-venue papers
3as first author
25since 2021 · last 2026
0000-0003-0410-6315ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Compass: Optimizing Compound AI Workflows for Dynamic Adaptation
abstract
Compound AI is a distributed intelligence approach that represents a unified system orchestrating specialized AI/ML models with engineered software components into AI workflows. Compound AI production deployments must satisfy accuracy, latency, and cost objectives under varying loads. However, many deployments operate on fixed infrastructure where horizontal scaling is not viable. Existing approaches optimize solely for accuracy and do not consider changes in workload conditions. We observe that compound AI systems can switch between configurations to fit infrastructure capacity, trading accuracy for latency based on current load. This requires discovering multiple Pareto-optimal configurations from a combinatorial search space and determining when to switch between them at runtime. We present Compass, a novel framework that enables dynamic configuration switching through offline optimization and online adaptation. Compass consists of three components: COMPASS-V algorithm for configuration discovery, Planner for switching policy derivation, and Elastico Controller for runtime adaptation. COMPASS-V discovers accuracy-feasible configurations using finite-difference guided search and a combination of hill-climbing and lateral expansion. Planner profiles these configurations on target hardware and derives switching policies using a queuing theory based model. Elastico monitors queue depth and switches configurations based on derived thresholds. Across two compound AI workflows, COMPASS-V achieves 100% recall while reducing configuration evaluations by 57.5% on average compared to exhaustive search, with efficiency gains reaching 95.3% at tight accuracy thresholds. Runtime adaptation achieves 90-98% SLO compliance under dynamic load patterns, improving SLO compliance by 71.6% over static high-accuracy baselines, while simultaneously improving accuracy by 3-5% over static fast baselines.
Milos Gravara, Juan Luis Herrera 0001, Stefan Nastic
CCGrid3
2026 Constella: A Novel Framework for Cost-Efficient Distributed AI Inference in LEO Space Data Centers
Andrija Stanisic, Milos Gravara, Juan Luis Herrera 0001, Stefan Nastic
Euro-Par (2)4
2026 ClusterLess: Deadline-Aware Serverless Workflow Orchestration on Federated Edge Clusters
Reza Farahani, Mario Colosi, Ilir Murturi, Stefan Nastic, Massimo Villari, Schahram Dustdar, Radu Prodan
ICDCS4
2025 Gaia: Hybrid Hardware Acceleration for Serverless AI in the 3D Compute Continuum
abstract
Serverless computing offers elastic scaling and pay-per-use execution, making it well-suited for AI workloads. As these workloads run in heterogeneous environments such as the Edge-Cloud-Space 3D Continuum, they often require intensive parallel computation, which GPUs can perform far more efficiently than CPUs. However, current platforms struggle to manage hardware acceleration effectively, as static user-device assignments fail to ensure SLO compliance under varying loads or placements, and one-time dynamic selections often lead to suboptimal or cost-inefficient configurations.
Maximilian Reisecker, Cynthia Marcelino, Thomas W. Pusztai, Stefan Nastic
BDCAT4
2025 EdgeCloudForge: Simulation-Driven Synthetic Dataset Generation For Proactive Serverless Edge Function Autoscaling
abstract
Serverless edge computing promises autonomous function management across the heterogeneous edge-cloud continuum. Specifically, autoscaling of functions increases resource efficiency by creating and destroying instances on demand. However, reactive function instance creation can cause high end-to-end latency due to requests waiting for processing. Researchers and practitioners use proactive autoscaling mechanisms to overcome this by predicting the required number of function instances in advance. Proactive approaches pose challenges because prediction models must be fine-tuned for new environments and updated over time. Therefore, we propose a simulation-driven framework, EdgeCloudForge, that generates datasets to train prediction models for autoscaling in edge-cloud environments ahead of their deployment, thus preparing them in advance for new environments. We introduce the concept of Edge-Cloud Domain Space to describe different infrastructure attributes. EdgeCloudForge uses a two-step method based on initial sampling and Bayesian optimization to navigate this domain space. Simulation results show that our approach can reduce SLO violations by up to 79% across different request patterns compared to the reactive approach and is able to outperform approaches from related work by up to 91%. EdgeCloudForge also shows the capability of training models that can be used across different infrastructures.
Philipp Raith, Alireza Furutanpey, Nikola Lukic, Vijay Thurimella, Stefan Nastic
IC2E5
2025 FedCCL: Federated Clustered Continual Learning Framework for Privacy-Focused Energy Forecasting
abstract
Privacy-preserving distributed model training is crucial for modern machine learning applications, yet existing Federated Learning approaches struggle with heterogeneous data distributions and varying computational capabilities. Traditional solutions either treat all participants uniformly or require costly dynamic clustering during training, leading to reduced efficiency and delayed model specialization. We present FedCCL (Federated Clustered Continual Learning), a framework specifically designed for environments with static organizational characteristics but dynamic client availability. By combining static pretraining clustering with an adapted asynchronous FedAvg algorithm, Fed-CCL enables new clients to immediately profit from specialized models without prior exposure to their data distribution, while maintaining reduced coordination overhead and resilience to client disconnections. Our approach implements an asynchronous Federated Learning protocol with a three-tier model topology - global, cluster-specific, and local models - that efficiently manages knowledge sharing across heterogeneous participants. Evaluation using photovoltaic installations across central Europe demonstrates that FedCCL's location-based clustering achieves an energy prediction error of 3.93% (±0.21%), while maintaining data privacy and showing that the framework maintains stability for population-independent deployments, with 0.14 percentage point degradation in performance for new installations. The results demonstrate that FedCCL offers an effective framework for privacy-preserving distributed learning, maintaining high accuracy and adaptability even with dynamic participant populations.
Michael A. Helcig, Stefan Nastic
ICFEC2
2025 Roadrunner: Accelerating Data Delivery to WebAssembly-Based Serverless Functions
abstract
Serverless computing provides infrastructure management and elastic auto-scaling, therefore reducing operational overhead. By design serverless functions are stateless, which means they typically leverage external remote services to store and exchange data. Transferring data over a network typically involves serialization and deserialization. These operations usually require multiple data copies and transitions between user and kernel space, resulting in overhead from context switching and memory allocation, contributing significantly to increased latency and resource consumption.
Cynthia Marcelino, Thomas W. Pusztai, Stefan Nastic
Middleware3
2025 Cosmos: A Cost Model for Serverless Workflows in the 3D Compute Continuum
abstract
Due to the high scalability, infrastructure management, and pay-per-use pricing model, serverless computing has been adopted in a wide range of applications such as real-time data processing, IoT, and AI-related workflows. However, deploying serverless functions across dynamic and heterogeneous environments such as the 3D (Edge-Cloud-Space) Continuum introduces additional complexity. Each layer of the 3D Continuum shows different performance capabilities and costs according to workload characteristics. Cloud services alone often show significant differences in performance and pricing for similar functions, further complicating cost management. Additionally, serverless workflows consist of functions with diverse character-istics, requiring a granular understanding of performance and cost trade-offs across different infrastructure layers to be able to address them individually. In this paper, we present Cosmos, a cost- and a performance-cost-tradeoff model for serverless workflows that identifies key factors that affect cost changes across different workloads and cloud providers. We present a case study analyzing the main drivers that influence the costs of serverless workflows. We demonstrate how to classify the costs of serverless workflows in leading cloud providers AWS and GCP. Our results show that for data-intensive functions, data transfer and state management costs contribute to up to 75% of the costs in AWS and 52% in GCP. For compute-intensive functions such as AI inference, the cost results show that BaaS services are the largest cost driver, reaching up to 83 % in AWS and 97 % in GCP.
Cynthia Marcelino, Sebastian Gollhofer-Berger, Thomas W. Pusztai, Stefan Nastic
SMARTCOMP4
2025 Databelt: A continuous data path for serverless workflows in the 3D compute continuum
abstract
Serverless computing allows for dynamic and flexible execution of FaaS functions while simplifying infrastructure management. Typically, serverless functions rely on remote storage services for managing state, which can result in increased latency and network communication overhead. In a dynamic environment such as the 3D (Edge-Cloud-Space) Compute Continuum, serverless functions face additional challenges due to frequent changes in network topology. As satellites move in and out of the range of ground stations, functions must make multiple hops to access cloud services, leading to high-latency state access and unnecessary data transfers. In this paper, we present Databelt, a state management framework for serverless workflows designed for the dynamic environment of the 3D Compute Continuum. Databelt introduces an SLO-aware state propagation mechanism that enables the function state to move continuously in orbit. Databelt proactively offloads function states to the most suitable node, such that when functions execute, the data is already present on the execution node or nearby, thus minimizing state access latency and reducing the number of network hops. Additionally, Databelt introduces a function state fusion mechanism that abstracts state management for functions sharing the same serverless runtime. When functions are fused, Databelt seamlessly retrieves their state as a group, reducing redundant network and storage operations and improving overall workflow efficiency. Our experimental results show that Databelt reduces workflow execution time by up to 66% and increases throughput by 50% compared to the baselines. Furthermore, our results show that Databelt function state fusion reduces storage operations latency by up to 20%, by reducing repetitive storage requests for functions within the same runtime, ensuring efficient execution of serverless workflows in highly dynamic network environments such as the 3D Continuum.
Cynthia Marcelino, Leonard Guelmino, Thomas W. Pusztai, Stefan Nastic
J. Syst. Archit.4
2025 ChunkFunc: Dynamic SLO-Aware Configuration of Serverless Functions
abstract
Serverless computing promises to be a cost effective form of on demand computing. To fully utilize its cost saving potential, workflows must be configured with the appropriate amount of resources to meet their response time Service Level Objective (SLO), while keeping costs at a minimum. Since determining and updating these configuration models manually is a nontrivial and error prone task, researchers have developed solutions for automatically finding configurations that meet the aforementioned requirements. However, our initial experiments show that even when following best practices and using state-ofthe- art configuration tools, resources may still be considerably over- or underprovisioned, depending on the size of functions' input payload. In this paper we present ChunkFunc, an SLOand input data-aware framework for tuning serverless workflows. Our main contributions include: i) an SLO- and input sizeaware function performance model for optimized configurations in serverless workflows, ii) ChunkFunc Profiler, an auto-tuned, Bayesian Optimization-guided profiling mechanism for profiling serverless functions with typical input data sizes to build a performance model, and iii) ChunkFunc Workflow Optimizer, which uses these models to determine an input size dependent configuration for each serverless function in a workflow to meet the SLO, while keeping costs to a minimum. We evaluate ChunkFunc on real-life serverless workflows and compare it to two state-of-the-art solutions, showing that it increases SLO adherence by a factor of 1.04 to 2.78, depending on the workflow, and reduces costs by up to 61%.
Thomas W. Pusztai, Stefan Nastic
IEEE Trans. Parallel Distributed Syst.2
2025 Performance Isolation for Serverless Functions
abstract
Serverless computing has emerged as a flexible model for deploying applications in multi-tenant environments, where small, isolated functions often share the same host and compete for local resources. This co-location can lead to resource contention, making performance isolation a fundamental challenge, particularly given the variability of workloads, infrastructure, and fine-grained resource sharing. Although existing surveys address performance isolation in cloud systems, they do not account for the unique characteristics of serverless computing, such as cold starts, fine-grained scaling, and function-level isolation. Therefore, in this paper, we provide insights into state-of-the-art methods that deal with the challenges of performance isolation for serverless functions. The selected approaches are evaluated based on multiple criteria, including the technique used to achieve isolation, the virtualization level at which isolation is enforced, the decision-making approach, and the primary isolation technique. We analyze and classify existing performance isolation techniques, organizing them into runtime, provisioning, and hybrid approaches. Building on this classification, we outline performance and reliability engineering mechanisms applicable to serverless computing that isolate functions while addressing serverless-specific challenges. Our findings show that i) Isolation is often treated as a secondary goal, primarily to reduce latency or SLO violations, rather than a primary objective. ii) Existing solutions frequently focus on CPU or memory contention while overlooking other critical shared components, such as the network, and iii) offer limited ways to tune the trade-off between isolation and performance, iv) unpredictability of serverless functions is the main challenge to performance isolation, v) novel metrics are required to monitor and quantify performance interference. Addressing these gaps will require more comprehensive and adaptive hybrids that unify multiple aspects of performance isolation. By consolidating and structuring these techniques in the context of serverless computing, this paper lays the foundation for developing resilient, efficient, and interference-tolerant serverless platforms.
Rastko Gajanin, Cynthia Marcelino, Stefan Nastic
IEEE Trans. Serv. Comput.3
2024 SimuScale: Optimizing Parameters for Autoscaling of Serverless Edge Functions Through Co-Simulation
abstract
Serverless Edge Computing is growing in popularity, and while commercial providers are starting to offer edge-oriented products, much research is still being done on orches-trating functions (e.g., autoscaling). These approaches range from threshold- to AI-based strategies and support various Service Level Objectives (SLOs), such as Round-Trip-Time (RTT) and re-source usage. However, the Quality of Service (Qo$S$) continuously deteriorates due to the dynamic edge-cloud continuum and static parameterization of orchestration strategy parameters. Platforms must adapt the orchestration parameters during runtime to counteract this drift that causes SLO violations. To this end, we introduce the Orchestration Parameter Optimization Prob-lem (OPOP), which aims to find parameters for orchestration strategies to minimize SLO violations. We propose a novel self-adaptive Simulation-based Scaling (SimuScale) approach that uses co-simulation to solve OPOP for autoscalers during runtime. SimuScale uses live monitoring data to feed the simulation and perform parameter optimization. Our Proof of Concept is inte-grated with Kubernetes and evaluated on a real-world edge-cloud testbed. While this work focuses on a threshold-based autoscaler, it can be extended to optimize other orchestration components (e.g., schedulers). Our experimental results show that SimuScale finds parameters that decrease RTT SLO violations between 15% and 40%. SimuScale also can reduce resource usage by 29.87% while maintaining the target 95th RTT percentile. Moreover, it can reduce variance caused by different request patterns, making orchestration strategies more resilient in realistic scenarios.
Philipp Raith, Stefan Nastic, Schahram Dustdar
CLOUD2
2024 Opportunistic Energy-Aware Scheduling for Container Orchestration Platforms Using Graph Neural Networks
abstract
Reducing the energy consumption of data centers is critical to meeting international climate goals and lowering operation costs. Container orchestration platforms can help counteract this trend by optimally placing applications across the infrastructure to increase resource utilization and reduce energy consumption. But platforms in use today are still energy-agnostic and do not offer any insights into energy consumption. In this paper, we present a monitoring framework and a new modeling approach for resource usage in data centers. The model captures heterogeneous hardware and software and acts as input for a Graph Neural Network (GNN) to predict power consumption. Based on this model, we derive a set of container scheduling algorithms that opportunistically schedule applications based on the estimated energy impact of incoming containers. Our results show that the GNN-based prediction model is very accurate and achieves an average RMSE (Root Mean Square Error) of 7.5%. We have implemented a custom scheduler to demonstrate the benefits of using our prediction, and our scheduler can decrease energy consumption on average by 6.2% without any code changes for the application and without increasing workload completion time compared to the default Kubernetes scheduler.
Philipp Raith, Gourav Rattihalli, Aditya Dhakal, Sai Rahul Chalamalasetti, Dejan S. Milojicic, Eitan Frachtenberg, Stefan Nastic, Schahram Dustdar
CCGrid7
2024 MISO: A CRDT-based Middleware for Stateful Objects in the Serverless Edge-Cloud Continuum
abstract
Serverless functions typically depend on external services to manage the application state, which can be difficult at the Edge due to high latency and network costs. Current solutions for stateful serverless functions at the Edge either have limited support for data locality or require mutual consensus for write operations which is hard to achieve at the Edge. This paper introduces MISO, a novel middleware for serverless computing that enables stateful serverless functions across the Edge-Cloud continuum. The middleware provides MISO Objects offering data locality. It is interoperable with existing serverless platforms and allows concurrent state modifications in a decentralized manner. The main contributions of our work include: i) A novel conceptual model to maintain application state in serverless functions called MISO Objects, ii) MISO middleware and an SDK for serverless functions, and iii) The asynchronous state replication mechanism of MISO Objects using an overlay network to optimize data transfer and resource consumption. Our evaluation demonstrated that MISO outperforms the state-of-theart by up to $243 \%$ in terms of total execution time for AllReducetype operations. Furthermore, the state replication exhibits $O(n)$ scalability regarding time, throughput, memory usage, and data volume. We further demonstrate that our work can seamlessly be integrated into an open-source serverless platform and that our SDK requires up to $150 \%$ fewer lines of code and exhibits up to $75 \%$ less cognitive complexity than the state-of-the-art.
Valentin Goronjic, Stefan Nastic
IC2E2
2024 VATE: Edge-Cloud System for Object Detection in Real-Time Video Streams
abstract
In the realm of edge intelligence, emerging video analytics applications are often based on resource constrained edge devices. These applications need systems which are able to provide both low-latency and high-accuracy video stream processing, such as for object detection in real-time video streams. State-of-the-art systems tackle this challenge by leveraging edge computing and cloud computing. Such edge-cloud approaches typically combine low-latency results from the edge and high accuracy results from the cloud when processing a frame of the video stream. However, the accuracy achieved so far leaves much room for improvement. Furthermore, using more accurate object detection often requires having more capable hardware. This limits the edge devices which can be used. Applications related to autonomous drones, with the drone being the edge device, give one example. A wide variety of objects needs to be detected reliably for drones to operate safely. Drones with more computing capabilities are often more expensive and suffer from short battery life, as they consume more energy. In this paper, we introduce VATE, a novel edge-cloud system for object detection in real-time video streams. An enhanced approach for edge-cloud fusion is presented, leading to improved object detection accuracy. A novel multi-object tracker is introduced, allowing VATE to run on less capable edge devices. The architecture of VATE enables it to be used when edge devices are capable of running on-device object detection frequently and when edge devices need to minimise on-device object detection to preserve battery life. Its performance is evaluated on a challenging, drone-based video dataset. The experimental results show that VATE improves accuracy by up to 27.5% compared to the state-of-the-art system, while running on less capable and cheaper hardware.
Maximilian Maresch, Stefan Nastic
ICFEC2
2024 StoreLess: Serverless Workflow Scheduling with Federated Storage in Sky Computing
Sashko Ristov, Mika Hautz, Philipp Gritsch, Stefan Nastic, Radu Prodan, Michael Felderer
ICSOC (2)4
2024 Deployment Architectures of MQTT Brokers in Event-Driven Industrial Internet of Things
abstract
The Industrial Internet of Things (IIoT) involves various standards, protocols, and tools, requiring extensive expertise for system design. The traditional automation pyramid limits scalability due to tightly-coupled components. Integrating IIoT data in the cloud through event-driven communication, like MQTT brokers, provides loose coupling. This integration also aids in resource monitoring and planning, enhancing IIoT application performance. Despite many IIoT architecture studies, there is a lack of empirical comparisons of MQTT broker deployment strategies. This paper examines an edge-cloud aggregation scenario, where IIoT device requests are aggregated at the edge and fog services before being transmitted to the cloud. We study four MQTT deployment architectures and compare the results empirically. We use an Arduino Opta as an IIoT device. Also, we use software-based load generation to support replicability of our experiment and reproducibility of our results. Findings indicate that a central broker on a dedicated virtual machine gives 13.8% improvements of the mean response time compared to the shared deployment of MQTT broker and device gateways. Our results offer insights into effective deployment strategies.
Amirali Amiri, Valentin Philipp Just, Gernot Steindl, Stefan Nastic, Wolfgang Kastner, Ian Gorton
IECON4
2024 HyperDrive: Scheduling Serverless Functions in the Edge-Cloud-Space 3D Continuum
abstract
The number of Low Earth Orbit (LEO) satellites has grown enormously in the past years. Their abundance and low orbits allow for low latency communication with a satellite almost anywhere on Earth, and high-speed inter-satellite laser links (ISLs) enable a quick exchange of large amounts of data among satellites. As the computational capabilities of LEO satellites grow, they are becoming eligible as general-purpose compute nodes. In the 3D continuum, which combines Cloud and Edge nodes on Earth and satellites in space into a seamless computing fabric, workloads can be executed on any of the aforementioned compute nodes, depending on where it is most beneficial. However, scheduling on LEO satellites moving at approx. 27,000 km/h requires picking the satellite with the lowest latency to all data sources (ground and, possibly, earth observation satellites). Dissipating heat from onboard hardware is challenging when facing the sun and workloads must not drain the satellite's batteries. These factors make meeting SLOs more challenging than in the Edge-Cloud continuum, i.e., on Earth alone. We present HyperDrive, an SLOaware scheduler for serverless functions specifically designed for the 3D continuum. It places functions on Cloud, Edge, or Space compute nodes, based on their availability and ability to meet the SLO requirements of the workflow. We evaluate HyperDrive using a wildfire disaster response use case with high Earth Observation data processing requirements and stringent SLOs, showing that it enables the design and execution of such next-generation 3D scenarios with 71% lower network latency than the best baseline scheduler.
Thomas W. Pusztai, Cynthia Marcelino, Stefan Nastic
SEC3
2024 BaaSLess: Backend-as-a-Service (BaaS)-Enabled Workflows in Federated Serverless Infrastructures
abstract
Serverless is a popular paradigm for expressing compute-intensive applications as serverless workflows. In practice, a significant portion of the computing is typically offloaded to various Backend-as-a-Service (BaaS) cloud services. There cent rise of federated serverless and Sky computing offers cost and performance advantages for these BaaS-enabled serverless workflows. However, due to vendor lock-in and lack of service interoperability, many challenges remain that impact the development, deployment, and scheduling of BaaS-enabled serverless workflows in federated serverless infrastructures. This paper introduces BAASLESS – a novel platform that delivers global and dynamic federated BaaS to serverless work flows. BAASLESS provides: (i) a novel SDK for uniform and dynamic access to federated BaaS services, reducing the complexity associated with the development of BaaS-enabled serverless workflows, (ii) a novel globally-federated serverless BaaS framework that delivers a suite of BaaS -less ML services, including text-to-speech, speech-to-text, translation, and OCR, together with a globally federated storage infrastructure, comprising AWS and Google cloud providers, and (iii) a novel model and an algorithm for scheduling BaaS-enabled serverless workflows to improve their performance. Experimental results using three complementary BaaS-enabled serverless workflows show that BAASLESS improves workflow execution time by up to 2.95× compared to the state-of-the-art serverless schedulers, often at a lower cost.
Thomas Larcher, Philipp Gritsch, Stefan Nastic, Sashko Ristov
IEEE Trans. Cloud Comput.3
2023 Demystifying deep learning in predictive monitoring for cloud-native SLOs
abstract
The complexity inherent in managing cloud computing systems calls for novel solutions that can effectively enforce high-level Service Level Objectives (SLOs) promptly. Unfortunately, most of the current SLO management solutions rely on reactive approaches, i.e., correcting SLO violations only after they have occurred. Further, the few methods that explore predictive techniques to prevent SLO violations focus solely on forecasting low-level system metrics, such as CPU and Memory utilization. Although valid in some cases, these metrics do not necessarily provide clear and actionable insights into application behavior. This paper presents a novel approach that directly predicts high-level SLOs using low-level system metrics. We target this goal by training and optimizing two state-of-the-art neural network models, a Short-Term Long Memory - LSTM, and a Transformer-based model. Our models provide actionable insights into application behavior by establishing proper connections between the evolution of low-level workload-related metrics and the high-level SLOs. We demonstrate our approach to selecting and preparing the data. We show in practice how to optimize LSTM and Transformer by targeting efficiency as a high-level SLO metric and performing a comparative analysis. We show how these models behave when the input workloads come from different distributions. Consequently, we demonstrate their ability to generalize in heterogeneous systems. Finally, we operationalize our two models by integrating them into the Polaris framework we have been developing to enable a performance-driven SLO-native approach to Cloud computing.
Andrea Morichetta 0002, Thomas W. Pusztai, Deepak Vij, Víctor Casamayor-Pujol, Philipp Raith, Stefan Nastic, Schahram Dustdar, Zhaobo Zhang
CLOUD7
2023 Vela: A 3-Phase Distributed Scheduler for the Edge-Cloud Continuum
abstract
The amalgamation of multiple Edge and Cloud clusters into an Edge-Cloud continuum requires efficient scheduling techniques to cope with high numbers of infrastructure nodes and computing jobs. Since monolithic schedulers typically do not scale well beyond a certain cluster size, distributed scheduling approaches are usually employed to address such scalability issues. Distributed schedulers are often designed for Cloud environments and lack support for the Edge. Conversely, many Edge schedulers focus on single clusters and provide limited support to deal with the scale of the Edge-Cloud continuum. In this paper, we present the Vela Distributed Scheduler, a globally distributed scheduler, which is specifically tailored for the Edge-Cloud continuum. The main contributions of our work include: i) A novel, globally distributed and orchestrator-independent scheduler with a 3-phase scheduling workflow; ii) A two-level, informed sampling mechanism, which reduces latency for globally distributed sampling and leverages job requirements to produce high quality node samples; And iii) a MultiBind mechanism that significantly reduces job evictions and rescheduling due to scheduling conflicts. We implement Vela on top of Kubernetes and evaluate it in a realistic large-scale setup using multiple interconnected, globally distributed, and production-ready MicroK8s clusters with up to 20,000 total simulated nodes. Our results show that Vela’s performance scales linearly with infrastructure size and that it reduces scheduling conflicts by a factor of 10.
Thomas W. Pusztai, Stefan Nastic, Philipp Raith, Schahram Dustdar, Deepak Vij
IC2E2
2023 CWASI: A WebAssembly Runtime Shim for Inter-Function Communication in the Serverless Edge-Cloud Continuum
abstract
Serverless Computing brings advantages to the Edge-Cloud continuum, like simplified programming and infrastructure management. In composed workflows, where serverless functions need to exchange data constantly, serverless platforms rely on remote services such as object storage and key-value stores as a common approach to exchange data. In WebAssembly, functions leverage WebAssembly System Interface to connect to the network and exchange data via remote services. As a consequence, co-located serverless functions need remote services to exchange data, increasing latency and adding network overhead. To mitigate this problem, in this paper, we introduce CWASI: a WebAssembly OCI-compliant runtime shim that determines the best inter-function data exchange approach based on the serverless function locality. CWASI introduces a three-mode communication model for the Serverless Edge-Cloud continuum. This communication model enables CWASI Shim to optimize inter-function communication for co-located functions by leveraging the function host mechanisms. Experimental results show that CWASI reduces the communication latency between the co-located serverless functions by up to 95% and increases the communication throughput by up to 30x.
Cynthia Marcelino, Stefan Nastic
SEC2
2021 Polaris Scheduler: Edge Sensitive and SLO Aware Workload Scheduling in Cloud-Edge-IoT Clusters
abstract
Application workload scheduling in hybrid Cloud-Edge-IoT infrastructures has been extensively researched over the last years. The recent trend of containerizing application workloads, both in the cloud and on the edge, has further fueled the need for more advanced scheduling solutions in these hybrid infrastructures. Unfortunately, most of the current approaches are not fully sensitive to the edge properties and also lack adequate support for Service Level Objective (SLO) awareness. Previously, we introduced software defined gateways (SDGs), which enable managing novel edge resources at scale. At the same time Kubernetes was initially released. In spite of not being specifically developed for the edge, Kubernetes implements many of the design principles introduced by our SDGs, making it suitable for building SDG extensions on top of it. In this paper we present Polaris Scheduler - a novel scheduling framework, which enables edge sensitive and SLO aware scheduling in the Cloud-Edge-IoT Continuum. Polaris Scheduler is being developed as a part of Linux Foundation's Centaurus project. We discuss the main research challenges, the approach, and the vision of SLO aware edge sensitive scheduling.
Stefan Nastic, Thomas W. Pusztai, Andrea Morichetta 0002, Víctor Casamayor-Pujol, Schahram Dustdar, Deepak Vij
CLOUD1
2021 A Novel Middleware for Efficiently Implementing Complex Cloud-Native SLOs
abstract
Service Level Objectives (SLOs) guide the elasticity of cloud applications, e.g., by deciding when and how much the resources provisioned to an application should be changed. Evaluating SLOs requires metrics, which can be directly measured on the application or system, or, more elaborately, be composed from multiple low-level metrics. The implementation of such metrics and SLOs, the triggering of elasticity strategies, and allowing configurability by the user deploying an application, requires a flexible middleware. In this paper, we present a middleware that provides an orchestrator-independent SLO controller for periodically evaluating SLOs and triggering elasticity strategies, while decoupling SLOs from the elasticity strategies to increase flexibility, and provider-independent services for obtaining low-level metrics and composing them into higher-level metrics. We evaluate our middleware by implementing a motivating use case, featuring a cost efficiency SLO for an application deployed on Kubernetes.
Thomas W. Pusztai, Andrea Morichetta 0002, Víctor Casamayor-Pujol, Schahram Dustdar, Stefan Nastic, Xiaoning Ding, Deepak Vij
CLOUD5
2021 SLO Script: A Novel Language for Implementing Complex Cloud-Native Elasticity-Driven SLOs
abstract
Service Level Objectives (SLOs) allow defining expected performance of cloud services, such that cloud service providers know what they guarantee and service consumers know what to expect. Most approaches focus on low-level SLOs, closely related to resources, e.g., average CPU or memory usage, and are usually bound to specific elasticity controllers. We present SLO Script, a language and accompanying framework, motivated by real-world, industrial needs to allow service providers to define complex, high-level SLOs in an orchestrator-independent manner. The main features of SLO Script include: i) novel abstractions (StronglyTypedSLO) with type safety features, ensuring compatibility between SLOs and elasticity strategies, ii) abstractions that enable decoupling of SLOs from elasticity strategies, iii) a strongly typed metrics API, and iv) an orchestrator-independent object model that enables language extensibility. We present a case study about a real-world, cloud-native application and evaluate our language while implementing a realistic Cost Efficiency SLO.
Thomas W. Pusztai, Andrea Morichetta 0002, Víctor Casamayor-Pujol, Schahram Dustdar, Stefan Nastic, Xiaoning Ding, Deepak Vij
ICWS5
2019 Towards Resilient Internet of Things: Vision, Challenges, and Research Roadmap
abstract
Internet of Things (IoT) systems open up massive versatility and opportunity to our world. Providing solutions for smart cities, healthcare, energy, and mobility, such systems increasingly permeate critical aspects of human activity. In a flourish of growth, these complex systems run software, are dynamic, without stable spatial and temporal boundaries, and involve mostly independent software components with different lifespans and evolution models. IoT systems provide data-centric, device-centric and service-centric functionalities that are subject to continuous disruption, under limitations such as resource-constrained devices, platforms heterogeneity, deployment in adverse environments and administrative domains. As these systems evolve and gain complexity, resilience becomes a crucial system property. Bolstering resilience entails understanding and systematically managing dynamic behavior and decentralizing operations. We advocate that to systematically engineer resilience in IoT systems, a complete rethink is necessary regarding their design and operation. In this paradigm shift, systems demand conceptual frameworks, techniques, and mathematically-backed formalisms to treat change and achieve decentralization. We outline a vision for addressing fundamental challenges that software engineering and distributed systems research encounters when building resilient IoT systems. Within a roadmap, we identify techniques and methods that can be leveraged to maintain resilience in the face of disruption, especially in the absence of central control and persistently at the system's runtime.
Christos Tsigkanos, Stefan Nastic, Schahram Dustdar
ICDCS2
2018 EMMA: Distributed QoS-Aware MQTT Middleware for Edge Computing Applications
abstract
Publish-subscribe middleware is a popular technology for facilitating device-to-device communication in large-scale distributed Internet of Things (IoT) scenarios. However, the stringent quality of service (QoS) requirements imposed by many applications cannot be met by cloud-based solutions alone. Edge computing is considered a key enabler for such applications. Client mobility and dynamic resource availability are prominent challenges in edge computing architectures. In this paper, we present EMMA, an edge-enabled publish-subscribe middleware that addresses these challenges. EMMA continuously monitors network QoS and orchestrates a network of MQTT protocol brokers. It transparently migrates MQTT clients to brokers in close proximity to optimize QoS. Experiments in a real-world testbed show that EMMA can significantly reduce end-to-end latencies that incur from network link usage, even in the face of client mobility and unpredictable resource availability.
Thomas Rausch, Stefan Nastic, Schahram Dustdar
IC2E2
2017 Data and control points: A programming model for resource-constrained iot cloud edge devices
abstract
Recent emergence of IoT Cloud systems has fostered proliferation of various applications mainly driven by urgent need to respond to volume, velocity and variety of data generated by IoT Cloud, but also to enable timely propagation of actuation decisions, crucial for business operation, to the Edge of the infrastructure. In such systems, utilizing currently untapped Edge resources such as sensory gateways, and enabling the IoT devices as first-class execution environments plays a crucial role. However, enabling virtually exclusive access to the underlying devices, e.g., field bus sensors and supporting flexible, application-specific customizations for such devices still remain a challenge. In this paper, we introduce Data- and Control Points - a novel programming model and framework for developing applications specifically tailored for resource-constrained Edge devices. Our framework offers programming constructs that enable applications to define custom configurations for and their own view of the underlying devices. By providing an illusion of an exclusive access to the underlying sensors and actuators, our framework supports execution of multiple applications within a single Edge device.
Stefan Nastic, Hong Linh Truong 0001, Schahram Dustdar
SMC1
2017 Deviceless edge computing: extending serverless computing to the edge of the network
abstract
The serverless paradigm has been rapidly adopted by developers of cloud-native applications, mainly because it relieves them from the burden of provisioning, scaling and operating the underlying infrastructure. In this paper, we propose a novel computing paradigm - Deviceless Edge Computing that extends the serverless paradigm to the edge of the network, enabling IoT and Edge devices to be seamlessly integrated as application execution infrastructure. We also discuss open challenges to realize Deviceless Edge Computing, based on our experience in prototyping a deviceless platform.
Alex Glikson, Stefan Nastic, Schahram Dustdar
SYSTOR2
2016 On Engineering Analytics for Elastic IoT Cloud Platforms
Hong Linh Truong 0001, Georgiana Copil, Schahram Dustdar, Duc-Hung Le, Daniel Moldovan, Stefan Nastic
ICSOC6
2015 Governing Elastic IoT Cloud Systems under Uncertainty
abstract
Emerging IoT cloud systems create unified IoT cloud infrastructures that offer large pools of elastic resources, which need to be governed through their entire lifecycle. However, numerous uncertainties are inherently present in such infrastructures, mainly due to the novel interactions of IoT elements, network elements, cloud resources and humans. They pose a plethora of challenges for the governance of such IoT cloud systems. In this paper we introduce U-GovOps -- a novel framework for dynamic, on-demand governance of elastic IoT cloud systems under uncertainty. We introduce a declarative policy language to simplifythe development of uncertainty-and elasticity-aware governance strategies. Based on that we develop runtime mechanisms, which enable mitigating the uncertainties by monitoring and governing the IoT cloud systems through specified strategies. We evaluate our approach using a real-life case study in the domain of predictive maintenance.
Stefan Nastic, Georgiana Copil, Hong Linh Truong 0001, Schahram Dustdar
CloudCom1
2015 iCOMOT - A Toolset for Managing IoT Cloud Systems
abstract
Developing and operating IoT cloud systems require novel features for deploying, controlling, monitoring and testing both IoT units and cloud services in an integrated environment spanning different infrastructures. In this paper, we demonstrate iCOMOT -- a novel toolset offering these features. Using iCOMOT we can perform various activities, such as dynamically reconfiguration of sensors, communication protocols, and cloud services in an elastic manner, suitable for testing and assuring quality of IoT cloud systems configurations. We will demonstrate our iCOMOT with a real-world predictive maintenance case study.
Hong Linh Truong 0001, Georgiana Copil, Schahram Dustdar, Duc-Hung Le, Daniel Moldovan, Stefan Nastic
MDM (1)6
2014 SALSA: A Framework for Dynamic Configuration of Cloud Services
abstract
Contemporary cloud services are constructed from different types of software and deployed on multiple cloud infrastructures, which offer various configuration options, and can change dynamically at runtime. Due to this complexity, such cloud services require substantial configuration efforts. Currently we lack techniques for automating the complex tasks and providing fine-grained configuration features for multi-cloud services. In this paper, we present a novel multi-level configuration approach for complex cloud services on multi-cloud environments. We develop techniques for automating configuration orchestration activities. Our solution enables the fine-grained configuration at different application abstraction levels and supports the dynamic change of cloud services at runtime. We provide the SALSA framework to implement our approach and demonstrate its usefulness with several real-world services.
Duc-Hung Le, Hong Linh Truong 0001, Georgiana Copil, Stefan Nastic, Schahram Dustdar
CloudCom4
2012 A programming model for context-aware applications in large-scale pervasive systems
abstract
In recent years, new business and research opportunities have been increasingly emerging in the field of large-scale context-aware pervasive systems (e.g. pervasive health-care, city traffic monitoring, environmental monitoring, smart grids). These large-scale pervasive systems are characterized by the need to employ large number of context sources, process massive amounts of real-time context data, provide services to numerous context-aware applications, and cope with higher volatility of the environment. This paper proposes the Origins Model — a programming model for context-aware applications in large-scale pervasive systems. In the Origins Model, an origin is an abstraction of any source of context information. Origins are universal, discoverable, composable, migratable, and replicable components that are associated with type and meta-information. They create an adequate foundation for the development of context-aware applications. Based on them, four processing operations are defined in the Origins Model: filter, infer, aggregate, and compose. As such, these operations provide a powerful mechanism to express a rich set of processing schemes in context-aware applications. Based on the Origins Model, we present the Origins Toolkit — a proof-of-concept implementation developed using the Scala programming language and the Akka toolkit to provide a distributed, scalable, and fault-tolerant solution.
Sanjin Sehic, Fei Li 0002, Stefan Nastic, Schahram Dustdar
WiMob3