Dimitrios Spatharakis

dblp:237/5155 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-0157-7173ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dynamic Task Scheduling and Function Orchestration for Serverless LLM
Myrsini Kellari, Christina Diamanti, Dimitrios Spatharakis, Aris Leivadeas, Symeon Papavassiliou
HPSR3
2026 A platform perspective for the computing continuum: Synergetic orchestration of compute and network resources for hyper-distributed applications
abstract
The rapid advancements in technologies across the Computing Continuum have reinforced the need for the interplay of various network and compute orchestration mechanisms within distributed infrastructure architectures to support the hyper-distributed application (HDA) deployments. A unified approach to managing heterogeneous components is crucial for reconciling conflicting objectives and creating a synergetic framework. To undertake these challenges, we present NEPHELE, a platform that realizes a hierarchical multi-layered orchestration architecture that incorporates infrastructure and application orchestration workflows across diverse resource management layers. The proposed platform integrates well-defined components spanning network and multi-cluster compute domains to enable intent-driven, dynamic orchestration. At its core, the Synergetic Meta Orchestrator (SMO) integrates diverse application requirements, generating deployment plans by interfacing with underlying orchestrators over distributed compute and network infrastructure. In the current work, we present the NEPHELE architecture, enumerate its interaction workflows, and evaluate key components of the overall architecture based on the instantiation and usage of the NEPHELE platform. The platform is evaluated in a multi-domain infrastructure setup to assess the operational overhead of the introduced orchestration functionality, considering also the assessment of different topology configurations on resource instantiation times and allocation dynamics, and network latency. Finally, we demonstrate the platform’s effectiveness in orchestrating distributed application graphs under varying placement intents, performance constraints, and workload stress conditions. The evaluation results outline the effectiveness of NEPHELE in orchestrating various infrastructure layers and application lifecycle scenarios through a unified interface.
Nikos Filinis, Ioannis Dimolitsas, Dimitrios Spatharakis, Paolo Bono, Anastasios Zafeiropoulos, Cristina Emilia Costa, Roberto Bruschi, Symeon Papavassiliou
Comput. Networks3
2026 A scalable and modular open-source stack for computing continuum digital twins
abstract
The exponential rise of intelligent Internet of Things (IoT) devices and the development of Cyber-Physical Systems (CPS) pose new challenges and requirements for modern applications. These include the need for seamless interconnectivity and interoperable interaction between various physical and virtual elements. The enrichment and transformation of IoT technologies to support such interactions is undergoing, considering the need for convergence with edge and cloud computing technologies and the management of IoT applications across resources in the computing continuum. This broader sense of connectivity is tightly connected with the development of Digital Twins (DT), which take advantage of the development of virtual counterparts of IoT devices and CPS. Novel architectural approaches are required to manage complex DTs’ topologies, collectively forming a Digital Twin Network (DTN) that acts as a middleware to provide advanced communication, efficient orchestration, and autonomous decision-making capabilities. This manuscript presents an architectural approach and a relevant open-source software stack implementation -called VOStack- for developing DTs. VOStack is open and modular by design, while it tackles IoT interoperability and convergence challenges with edge and cloud computing technologies. VOStack is thoroughly evaluated under various deployment schemas, virtualization techniques, and based on the provision of an IoT application in the context of a smart city scenario, demonstrating efficient utilization of resources and high efficiency of Machine Learning (ML)-driven orchestration mechanisms.
Nikos Filinis, Dimitrios Spatharakis, Ioannis Dimolitsas, Eleni Fotopoulou, Constantinos Vassilakis, Anastasios Zafeiropoulos, Symeon Papavassiliou
Future Gener. Comput. Syst.2
2026 A task offloading and batch scheduling framework for Edge-assisted inference
abstract
Real-time processing of inference tasks generated by resource-constrained devices in an Edge Computing environment demands carefully designed solutions that guarantee the performance and achieve a high level of accuracy. To reduce transmission time, inference tasks are often compressed to reduce the transmission time and offloaded to the edge infrastructure for parallel batch processing in GPUs. In this setting, an interesting tradeoff arises, characteristic of Approximate Computing, where the quality of inference and the system’s end-to-end latency are competing objectives. In this paper, we formulate a joint optimization problem to maximize the quality of inference while minimizing the overall latency for the GPU-enabled batch processing of inference applications. The optimization problem is NP-hard, and we split it into two subproblems to obtain optimal values for the compression of offloaded tasks and select the offloading strategy that minimizes the total latency. By carefully examining the results of the compression problem, we identify that compressing the tasks in such a way to arrive simultaneously for remote processing significantly increases the performance of batch processing. To compute an offloading strategy, we employ a semidefinite relaxation (SDR)-based approach and a randomized mapping to obtain feasible solutions. Therefore, we design an iterative alternating algorithm to solve both problems and obtain a near-optimal solution in polynomial complexity. Simulation results indicate that the proposed framework outperforms all compared solutions by reducing the total cost by 50%.
Dimitrios Spatharakis, Christos Pelekis, Dimitrios Dechouniotis, Symeon Papavasileiou
Future Gener. Comput. Syst.1
2025 Multi-Partner Project: Orchestrating Deployment and Real-Time Monitoring - NEPHELE Multi-Cloud Ecosystem
Manolis Katsaragakis, Orfeas Filippopoulos, Christos Sad, Dimosthenis Masouros, Dimitrios Spatharakis, Ioannis Dimolitsas, Nikos Filinis, Anastasios Zafeiropoulos, Kostas Siozios, Dimitrios Soudris, Symeon Papavassiliou
DATE5
2025 SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
abstract
The growing gap between the increasing complexity of large language models (LLMs) and the limited computational budgets of edge devices poses a key challenge for efficient on-device inference, despite gradual improvements in hardware capabilities. Existing strategies, such as aggressive quantization, pruning, or remote inference, trade accuracy for efficiency or lead to substantial cost burdens. This position paper introduces a new framework that leverages speculative decoding, previously viewed primarily as a decoding acceleration technique for autoregressive generation of LLMs, as a promising approach specifically adapted for edge computing by orchestrating computation across heterogeneous devices. We propose SLED, a framework that allows lightweight edge devices to draft multiple candidate tokens locally using diverse draft models, while a single, shared edge server verifies the tokens utilizing a more precise target model. To further increase the efficiency of verification, the edge server batches the diverse verification requests from devices. This approach supports heterogeneous devices and reduces server-side memory footprint by sharing a single upstream target model across devices. Our initial experiments with Jetson Orin Nano, Raspberry Pi 4B/5, and an edge server equipped with 4 Nvidia A100 GPUs indicate substantial benefits: ×2.2 higher system throughput, ×2.8 higher system capacity, and better cost efficiency, all without sacrificing model accuracy.
Xiangchen Li, Dimitrios Spatharakis, Saeid Ghafouri, Jiakun Fan, Hans Vandierendonck, Chacko John Deepu, Bo Ji 0001, Dimitrios S. Nikolopoulos
SEC2
2024 Virtual Objects for Robots and Sensor Nodes in Distributed Applications over the Cloud Continuum
abstract
The cloud-to-edge-to-IoT continuum represents a seamless flow of data processing and management, spanning from centralized cloud services to distributed edge computing and interconnected IoT devices. This paradigm can become very challenging in real implementations, especially in the presence of multiple stakeholders using proprietary and heterogeneous software and hardware. The Horizon Europe NEPHELE project proposes the virtualization of IoT devices through a specific software stack called Virtual Object Stack (VOStack) that promotes openness and interoperability. In this paper, we present an implementation of VOStack using W3C Web of Things (WoT) standard in a post-disaster domain for two different types of IoT devices: a ground robot (Turtlebot2) for navigation and mapping in an unknown environment, and a Raspberry Pi 3 acting as a wireless sensor network gateway. We propose an application graph for the resulting hyper-distributed application (HDA) and present our first implementation to validate the proposed solution.
Adriana Arteaga Arce, Nikos Filinis, Carol Habib, Leonardo Militano, Dimitrios Spatharakis, Anastasios Zafeiropoulos, Thomas Michael Bohnert, Nathalie Mitton, Symeon Papavassiliou
ISCC6
2024 Intent-driven orchestration of serverless applications in the computing continuum
Nikos Filinis, Ioannis Tzanettis, Dimitrios Spatharakis, Eleni Fotopoulou, Ioannis Dimolitsas, Anastasios Zafeiropoulos, Constantinos Vassilakis, Symeon Papavassiliou
Future Gener. Comput. Syst.3
2023 Multi-Application Hierarchical Autoscaling for Kubernetes Edge Clusters
abstract
The dynamic workload demands of smart city applications hosted on edge infrastructures require the development of advanced scaling mechanisms. Recent studies proposed single-application autoscaling solutions based on various technical approaches. However, for edge infrastructures with limited resource availability, it is essential to simultaneously manage heterogeneous application requirements, aiming at optimal resource allocation and minimal operational costs. This study introduces a multi-application hierarchical autoscaling framework for Kubernetes Edge Clusters. An application-based mechanism nominates the best applications’ deployments based on workload prediction and several criteria that guarantee the application’s performance while minimizing the infrastructure provider’s cost. For the joint application orchestration, an aggregation mechanism composes the candidate scaling solutions for the cluster. Then, a cluster autoscaling mechanism, based on the Analytic Hierarchy Process, undertakes the cluster’s scaling decision to optimize the resource allocation and energy consumption of the cluster. The evaluation illustrates the benefits of the proposed scaling strategy, achieving significant improvement in the average allocated resources and energy consumption compared to single-application approaches.
Ioannis Dimolitsas, Dimitrios Spatharakis, Dimitrios Dechouniotis, Anastasios Zafeiropoulos, Symeon Papavassiliou
SMARTCOMP2
2023 Resource-Aware Estimation and Control for Edge Robotics: A Set-Based Approach
abstract
The evolution of the Industrial Internet of Things (IIoT) and edge computing enables resource-constrained mobile robots to offload the computationally intensive localization algorithms. Naturally, utilizing the remote resources of an edge server to offload these tasks encounters the challenge of a joint co-design in communication, control, estimation, and computing infrastructure. We introduce a set-based estimation offloading framework, for the specific case of the navigation of a unicycle robot toward a target position. The robot is subject to modeling and measurement uncertainties, and the estimation set is calculated using overapproximation techniques that alleviate additional computations. A switching set-based control mechanism provides accurate navigation and triggers more precise estimation algorithms when needed. To guarantee the convergence of the system and optimize the utilization of remote resources, a utility-based offloading mechanism is designed, which takes into account both the dynamic network conditions and the available computing resources at the network edge. The performance of the proposed framework is demonstrated through simulations and comparison with alternative offloading schemes.
Dimitrios Spatharakis, Marios Avgeris, Nikolaos Athanasopoulos, Dimitrios Dechouniotis, Symeon Papavassiliou
IEEE Internet Things J.1
2022 Distributed Resource Autoscaling in Kubernetes Edge Clusters
abstract
Maximizing the performance of modern applications requires timely resource management of the virtualized resources. However, proactively deploying resources for meeting specific application requirements subject to a dynamic workload profile of incoming requests is extremely challenging. To this end, the fundamental problems of task scheduling and resource autoscaling must be jointly addressed. This paper presents a scalable architecture compatible with the decentralized nature of Kubernetes [1], to solve both. Exploiting the stability guarantees of a novel AIMD-like task scheduling solution, we dynamically redirect the incoming requests towards the containerized application. To cope with dynamic workloads, a prediction mechanism allows us to estimate the number of incoming requests. Additionally, a Machine Learning-based (ML) Application Profiling Modeling is introduced to address the scaling, by co-designing the theoretically-computed service rates obtained from the AIMD algorithm with the current performance metrics. The proposed solution is compared with the state-of-the-art autoscaling techniques under a realistic dataset in a small edge infrastructure and the trade-off between resource utilization and QoS violations are analyzed. Our solution provides better resource utilization by reducing CPU cores by 8% with only an acceptable increase in QoS violations.
Dimitrios Spatharakis, Ioannis Dimolitsas, Eleftherios E. Vlahakis, Dimitrios Dechouniotis, Nikolaos Athanasopoulos, Symeon Papavassiliou
CNSM1
2022 AHP4HPA: An AHP-based Autoscaling Framework for Kubernetes Clusters at the Network Edge
abstract
Autoscaling resources in a power-efficient way is essential to enable Green Computing resource management solutions. The development of dynamic resource provisioning techniques could lead to the minimization of power consumption and simultaneously guarantee high quality of service (QoS) inline with the workload demand. In this work, we introduce AHP4HPA, an autoscaling framework for Kubernetes Clusters, which is aligned with the Kubernetes architecture and state-of-the-art practices. We define resource profiles, namely a mapping between the QoS and the computing resources, to maximize the performance. Furthermore, Analytic Hierarchy Process (AHP) is exploited to dictate the scaling decision of the resources under various Key Performance Indicators (KPIs) toward power optimization of the allocated resources. To guarantee maximum performance of the deployed image classification application, an ARIMA model is dedicated to providing predictions regarding the incoming workload traffic. The framework is evaluated against a realistic dataset in a small-scale testbed. Numerical results indicate at least a 9% reduction of the average energy consumption when compared to other state of the art techniques.
Ioannis Dimolitsas, Dimitrios Spatharakis, Dimitrios Dechouniotis, Symeon Papavassiliou
GLOBECOM2
2022 Edge Robotics Experimentation over Next Generation IIoT Testbeds
abstract
The emergence of Industrial Internet of Things (IIoT) requires the interconnection between robots, sensors, and the underlying network and computing infrastructure. Edge Robotics has emerged as a flexible paradigm that enables resource-constrained mobile robots to offload computationally intensive tasks of time/mission-critical applications. In this context, Edge Computing is essential for providing additional resources towards confronting the stringent performance specifications. This article presents the architectural concepts and capabilities of the NETMODE testbed, member of the Fed4FIRE+ federation, for the state-of-the-art experimentation with robotic applications. An evaluation of the proposed architecture is conducted using a SLAM algorithm which is a compute-intensive application.
Dimitrios Dechouniotis, Dimitrios Spatharakis, Symeon Papavassiliou
NOMS2
2021 Task offloading in Edge and Cloud Computing: A survey on mathematical, artificial intelligence and control theory solutions
Firdose Saeik, Marios Avgeris, Dimitrios Spatharakis, Nina Santi, Dimitrios Dechouniotis, John Violos, Aris Leivadeas, Nikolaos Athanasopoulos, Nathalie Mitton, Symeon Papavassiliou
Comput. Networks3
2020 COSMOS: An Orchestration Framework for Smart Computation Offloading in Edge Clouds
abstract
The evolution of Internet of Things (IoT) has sparked significant research interest in edge computing. Within this scope and given the ever-increasing number of IoT and mobile devices, computation offloading is emerging as a cutting-edge and significant research area with enormous potential and practical applications.In this respect, we present the architecture design and experimental evaluation of an orchestration framework for smart computation offloading from IoT or mobile devices to edge cloud servers. The proposed orchestration platform, namely COSMOS, includes control-plane components for workload prediction, load balancing, and admission control. COSMOS is particularly tailored to the needs of an object identification service that receives images from a multitude of Points of Interest (PoIs), performs object identification using a trained model (based on Tensorflow), calculates the prediction accuracy, and finally returns to the end-users the identification outcome and accuracy along with useful information about the identified object. COSMOS has been deployed and evaluated in a large-scale experimental facility that employs OpenStack and OpenSourceMANO (OSM) for Network Function Virtualization (NFV) orchestration. Our experimental results indicate the feasibility of computation offloading for this object identification service and further uncover useful insights in terms of performance and scalability.
George Papathanail, Ioakeim Fotoglou, Christos Demertzis, Angelos Pentelas, Kyriakos Sgouromitis, Panagiotis Papadimitriou 0001, Dimitrios Spatharakis, Ioannis Dimolitsas, Dimitrios Dechouniotis, Symeon Papavassiliou
NOMS7
2020 A scalable Edge Computing architecture enabling smart offloading for Location Based Services
Dimitrios Spatharakis, Ioannis Dimolitsas, Dimitrios Dechouniotis, George Papathanail, Ioakeim Fotoglou, Panagiotis Papadimitriou 0001, Symeon Papavassiliou
Pervasive Mob. Comput.1