Stefan Pedratscher

dblp:276/9928 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-6164-880XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MARLTC: A Multi-Agent Reinforcement Learning-Based Interference-Aware Transmission Control for LoRaWAN IoT Devices
abstract
LoRaWAN has become a foundational technology in the Internet of Things (IoT) landscape due to its long-range communication and energy efficiency. However, its default Adaptive Data Rate (ADR) mechanism struggles to adapt to dynamic environments and dense network deployments, where co-channel interference and the coupling between transmission parameters limit its ability to ensure reliable and energy-efficient communication. To address these challenges, this paper proposes MARLTC, an ADR mechanism based on Multi-Agent Rein-forcement Learning (MARL) that jointly optimizes spreading factor and transmission power using realistic observable metrics, including recent transmission history, observed signal-to-noise ratios, and distribution of spreading factors in the network. The transmission configuration problem is modeled as a cooperative Markov Game and solved using the Centralized Training and Decentralized Execution (CTDE) paradigm, where pre-trained policies are deployed at end devices to infer suitable transmission parameters with reduced convergence time. The results using a realistic LoRaWAN simulator show that MARLTC achieves up to 67.9% faster convergence, 62.1% higher energy efficiency, and a 5.0% better packet delivery ratio compared to state-of-the-art approaches, highlighting its scalability and responsiveness in dense deployments. The practical feasibility of MARLTC is further validated using physical LoRaWAN hardware, proving that the resulting policies meet the memory and timing constraints of resource-constrained IoT devices.
Juan Aznar-Poveda, Laura Acosta-Garcia, Fabian Margreiter, Marlon Etheredge, Abolfazl Younesi, Stefan Pedratscher, Joan García-Haro, Thomas Fahringer, Antonio-Javier García-Sánchez
IEEE Internet Things J.6
2026 Pulse: Multi-objective scheduling of service-based applications in multi-cluster cloud-edge-IoT infrastructures
abstract
The rapid growth of cloud computing and the expansion of edge and IoT technologies are becoming essential for meeting the performance, scalability, and latency requirements of modern distributed service-based applications. While these applications facilitate the accommodation of real-world workloads, their placement across a computing continuum spanning the cloud, edge, and IoT remains challenging. Existing service scheduling works often rely on simulations or orchestration systems limited to a single cluster. While simulations fail to capture real-world constraints, single-cluster orchestration systems introduce significant overhead and cannot capture the heterogeneity, network latency, and cross-cluster dependencies across cloud, edge, and IoT layers. In this paper, we introduce Pulse, a fully distributed scheduling system designed to optimize the deployment of distributed service-based applications across multi-cluster environments. Pulse leverages a two-phase distributed multi-objective optimization approach: locally optimizing for monetary cost and fairness within clusters, and globally optimizing for monetary cost and latency across multiple clusters. Pulse is built atop a lightweight orchestration framework, which enables service coordination across cloud, edge, and IoT layers, ensuring adaptability to heterogeneous infrastructures and the latency among geo-distributed clusters. To validate our approach, we conduct a comprehensive real-world evaluation on the Grid’5000 infrastructure, demonstrating that Pulse outperforms state-of-the-art scheduling methods by improving total resource utilization by 34.5%, reducing monetary cost by 82.6%, and lowering the average end-to-end network latency among services by 75.0%. These results highlight Pulse’s effectiveness in managing large-scale, service-based applications in realistic, multi-cluster environments.
Marlon Etheredge, Juan Aznar-Poveda, Stefan Pedratscher, Abolfazl Younesi, Thomas Fahringer
J. Netw. Comput. Appl.3
2025 SmartKV: A cost-effective and low-latency geo-distributed key-value store for the computing continuum
abstract
Many data-intensive and distributed applications rely on low-latency and scalable key–value storage systems across the Computing Continuum. Key–value storage systems typically use consistent hashing or hash slot-sharding mechanisms to distribute data across storage nodes, which ensures load balancing but often leads to sub-optimal response times and monetary costs, particularly in geo-distributed systems where nodes might have different unit prices and be widely dispersed. In this paper, we propose SmartKV , a cost-efficient geo-distributed key–value store that optimizes data placement dynamically, abstracting the intricacies of data organization, transfer, access, and processing. SmartKV integrates a decentralized data placement algorithm that optimizes the replication factor and selects suitable locations for key–value pairs and replicas, balancing cost and access latency while keeping optimization overhead low. We employ a realistic cost model based on public and private Cloud and Edge providers that consider data transfer, request, and storage costs. In addition to conventional key–value pairs, SmartKV supports active key–value pairs, which enable the definition of custom data types and the execution of user-defined functions directly on the storage side. This contributes to reducing data transfer costs and round-trip times. We thoroughly evaluate SmartKV across different regions of the Chameleon testbed using several realistic workloads. Results show that the utilized decentralized data placement strategy allows SmartKV to reduce round trip times between 9 and 84% while reducing costs up to 4.84 × under different client workloads and consistency models compared to state-of-the-art data placement strategies. • Novel geo-distributed KV store with custom data placement strategies. • Decentralized data placement algorithm to optimize costs and round trip times. • Active KV pairs support remote execution to reduce costs and round trip times.
Juan Aznar-Poveda, Maximilian Franz Ebner, Thomas Fahringer, Zahra Najafabadi Samani, Marlon Etheredge, Stefan Pedratscher, Nishant Saurabh
Future Gener. Comput. Syst.6
2024 Dynamic Workflow Scheduling in the Edge-Cloud Continuum: Optimizing Runtimes Under Budget Constraints
abstract
Scientific workflows are increasingly adopting hy-brid Edge-Cloud infrastructures to benefit from the computational and storage capacity of the Cloud and the cost savings and data locality of the Edge. Workflow scheduling is one of the most challenging problems for the Edge-Cloud continuum. State-of-the-art workflow schedulers often rely on a centralized runtime system and are based on static algorithms that either focus on Cloud or Edge systems (but not both). In this paper, we introduce a novel, open-source, and dynamic scheduler for scientific workflows that targets the Edge-Cloud continuum by design using fully decentralized runtime system instances. This not only reduces data transfer times but also leverages the benefits of the continuum. The proposed scheduler optimizes for runtime while adhering to a given cost limit by dynamically mapping tasks to resources and orchestrating groups of workflow tasks on runtime system instances. Furthermore, the scheduler adapts to real-time updates in task durations, accommodating for variations in resource performance, to efficiently use the cost limit and to reduce the total runtime of the workflow. We compare our approach against a state-of-the-art dynamic scheduler (JIT-C) for four well-known scientific workflows. Experiments demonstrate that by using the smallest cost derived by JIT-C as a cost limit for our scheduler, we achieve an average runtime improvement of 56% and an average cost reduction of 34% compared to JIT-C.
Stefan Pedratscher, Thomas Fahringer, Juan Aznar-Poveda
CLOUD1
2023 $xAFCL$xAFCL: Run Scalable Function Choreographies Across Multiple FaaS Systems
abstract
Most well-known cloud providers offer advanced support for serverless applications that goes beyond single function invocation by enabling developers to build entire workflows, which are known as serverless function choreographies (FCs). Current support for FCs by many FaaS systems uncovered important problems including maximum number of parallel function executions, unexpected considerable delays, and provider lock-in. These limitations can result in longer execution times or even failure to execute individual functions or entire FCs. To overcome some of these limitations, we introduce a scalable middleware service xAFCL that can schedule and execute different functions of the same FC across multiple FaaS systems (currently supporting all top five providers). In order to support scheduling under xAFCL, we introduce a novel FaaS model which estimates the completion time of functions by considering FaaS system limitations, submission delays, and overheads for executing functions. Experimental results demonstrate that xAFCLs FaaS model shows very low inaccuracy of up to 2.9% for AWS and 20% for IBM for real-life BWA data-bound FC that uses S3. Moreover, xAFCL outperforms an earliest start time (EST) scheduler by up to 43% for makespan and 2.7x for throughput.
Sashko Ristov, Stefan Pedratscher, Thomas Fahringer
IEEE Trans. Serv. Comput.2
2022 xAFCL: Run Scalable Function Choreographies Across Multiple FaaS Systems
abstract
[J1C2 Presentation Abstract at IEEE SERVICES 2021 for IEEE Transactions on Services Computing DOI 10.1109/TSC.2021.3128137]
Sashko Ristov, Stefan Pedratscher, Thomas Fahringer
SERVICES2
2022 M2FaaS: Transparent and fault tolerant FaaSification of Node.js monolith code blocks
abstract
Porting existing monoliths to the Function-as-a-Service (FaaS) (FaaSification) can be very challenging for software developers due to different architectural styles. For a successful porting, developers need to resolve various dependencies, such as method invocations of external packages or user-defined codes, as well as global and local variables used in and after the code block that should be faasified. To bridge the gap and automatize FaaSification, this paper introduces M2FaaS, a FaaSifier that automatically converts a Node.js monolith into a hybrid by faasifying annotated code blocks as serverless functions on multiple FaaS providers. M2FaaS is a novel FaaSifier that resolves many challenges for the resulting monolith to work properly after the FaaSification. Developers may annotate all dependencies that need to be resolved for the generated functions to run properly and specify variables that should be returned by the function to the monolith because they are used later in the monolith. Moreover, M2FaaS is the first FaaSifier that faasifies arbitrary code blocks. The current M2FaaS prototype supports FaaSification of individual functions on two FaaS providers, AWS Lambda and IBM Cloud Functions. Finally, M2FaaS introduces an optional annotation for alternative functions to be invoked in case the primary faasified function fails. The resulting hybrid application invokes the automatically deployed serverless functions, while the original code remains executable. Experiments with four complementary monoliths demonstrate that M2FaaS outperforms state-of-the-art FaaSifiers in terms of development effort by up to 73.3%. Moreover, with the fault tolerance support, M2FaaS finishes all submitted functions, thereby achieving by 18.5% higher throughput than the other FaaSifiers.
Stefan Pedratscher, Sashko Ristov, Thomas Fahringer
Future Gener. Comput. Syst.1
2021 AFCL: An Abstract Function Choreography Language for serverless workflow specification
abstract
Serverless workflow applications or function choreographies (FCs), which connect serverless functions by data- and control-flow, have gained considerable momentum recently to create more sophisticated applications as part of Function-as-a-Service (FaaS) platforms. Initial experimental analysis of the current support for FCs uncovered important weaknesses, including provider lock-in, and limited support for important data-flow and control-flow constructs. To overcome some of these weaknesses, we introduce the Abstract Function Choreography Language (AFCL) for describing FCs at a high-level of abstraction, which abstracts the function implementations from the developer. AFCL is a YAML-based language that supports a rich set of constructs to express advanced control-flow (e.g. parallelFor loops, parallel sections, dynamic loop iterations counts) and data-flow (e.g multiple input and output parameters of functions, DAG-based data-flow). We introduce data collections which can be distributed to loop iterations and parallel sections that may substantially reduce the delays for function invocations due to reduced data transfers between functions. We also support asynchronous functions to avoid delays due to blocking functions. AFCL supports properties (e.g. expected size of function input data) and constraints (e.g. minimize execution time) for the user to optionally provide hints about the behavior of functions and FCs and to control the optimization by the underlying execution environment. We implemented a prototype AFCL environment that supports AFCL as input language with multiple backends (AWS Lambda and IBM Cloud Functions) thus avoiding provider lock-in which is a common problem in serverless computing. We created two realistic FCs from two different domains and encoded them with AWS Step Functions, IBM Composer and AFCL. Experimental results demonstrate that our current implementation of the AFCL environment substantially outperforms AWS Step Functions and IBM Composer in terms of development effort, economic costs, and makespan.
Sashko Ristov, Stefan Pedratscher, Thomas Fahringer
Future Gener. Comput. Syst.2