Achilleas Tzenetopoulos

dblp:276/6308 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2027
0000-0001-6084-4297ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 CECAIServe: Facilitating ML inference serving across the cloud-edge-continuum
abstract
Machine learning (ML) inference-serving has become a core operational component of MLOps, particularly as ML services are increasingly deployed across the cloud–edge continuum. However, existing inference-serving tools typically provide limited support for heterogeneous hardware, weak energy observability, and cumbersome integration of device-specific accelerated inference frameworks and model-specific pre-/post-processing. In this work, we present CECAIServe, an open-source, unified, and vendor-neutral framework that automatically generates deployment-ready Accelerated Inference Serving Containers (AISCs) from high-level TensorFlow and PyTorch models. Through its modular design, CECAIServe abstracts device- and framework-specific complexity while supporting CPUs, GPUs, edge accelerators, and FPGA-based systems. The framework further incorporates MLOps-oriented functionality, including fine-grained latency instrumentation, integrated power monitoring, and an interface for pre-/post-processing. Our comprehensive evaluation on 12 models demonstrates that CECAIServe can automatically generate AISCs across 7 diverse devices in less than 10 min. Additionally, we demonstrate that CECAIServe facilitates effective benchmarking and design space exploration on HW-accelerated inference-serving on devices across the cloud–edge-continuum. Finally, we validate the extensibility of the framework by adapting it to support Large Language Model inference-serving.
Aimilios Leftheriotis, Achilleas Tzenetopoulos, George Lentaris, Dimitrios Soudris, George Theodoridis
Future Gener. Comput. Syst.2
2026 $\Omega$Ωkypous: Harnessing Timing Slacks and Coordinated DVFS for Power-Efficient Serverless Workflows
abstract
Serverless workflows have emerged in Function-as-a-Service (FaaS) platforms to represent the operational structure of traditional applications. With latency propagation effects becoming increasingly prominent, step-wise resource tuning is required to address Service-Level-Objectives (SLOs). Modern processors’ allowance for fine-grained Dynamic Voltage and Frequency Scaling (DVFS), coupled with serverless workflows’ intermittent nature, presents a unique opportunity to reduce power while meeting SLOs. We introduce Ωkypous, an SLOdriven DVFS framework for serverless workflows. Ωkypous employs a grey-box model that predicts functions’ execution latency and power under different Core and Uncore frequency combinations. Based on these predictions and the timing slacks between workflow functions, Ωkypous uses a closed-loop control mechanism to dynamically adjust Core and Uncore frequencies, reducing power consumption without compromising predefined end-to-end latency constraints. Our evaluation on real-world traces from Azure demonstrates an average power consumption reduction of 16% compared to state-of-the-art power management frameworks, while consistently maintaining low SLO violation rates (1.8%), even when operating under power caps.
Achilleas Tzenetopoulos, Dimosthenis Masouros, Sotirios Xydis, Dimitrios Soudris
IEEE Trans. Computers1
2024 Seamless HW-accelerated AI serving in heterogeneous MEC Systems with AI@EDGE
abstract
The advancement towards B5G/6G relies on the synthesis of connect-compute platforms and their use in highly heterogeneous clusters featuring hardware accelerators. While these accelerators offer improved computational efficiency, sill, they make development, deployment, and orchestration of services more complex, with limited flexibility, and necessitate domain-specific knowledge. In AI@EDGE we are targeting seamless integration of such diverse platforms for executing AI-related tasks. This paper focuses on acceleration aspects and presents a MEC system that facilitates AI servicing over a cluster of FPGA, GPU, and CPU nodes. To this end, we develop our custom tools for generating multi-variant AI models, informative function descriptors, flexible MEC orchestrators, and runtime resource managers. The results show successful interoperability, with generic Python models getting deployed/migrated across distinct platforms for performance gains in the area of 10x.
Achilleas Tzenetopoulos, George Lentaris, Aimilios Leftheriotis, Panos Chrysomeris, Javier Palomares, Estefanía Coronado, Raman Kazhamiakin, Dimitrios Soudris
HPDC1
2024 Disaggregated RDDs: Extending and Analyzing Apache Spark for Memory Disaggregated Infrastructures
abstract
Apache Spark has become essential in large-scale data processing as the demand for scalable data analytics grows. With memory costs constituting a significant portion of server expenses, the under-utilization and fragmentation of resources pose a substantial challenge for data center operators reliant on economies of scale. Memory disaggregation emerges as a solution to this challenge, by leveraging remote memory pools to reduce resource fragmentation and under-utilization. Yet, these advantages are not without cost. Disaggregated memory systems introduce increased latency and reduced bandwidth, significantly impacting job execution latency. This necessitates careful optimization and management strategies to effectively balance the trade-offs between accessibility and performance. This paper introduces cache-remote, a custom Apache Spark configuration balancing memory disaggregation benefits with execution efficiency. Cache-remote uses remote memory for RDD caching (Disaggregated RDDs) and local memory for latency-sensitive computations. Our work includes a comprehensive evaluation of different memory allocation policies and Spark configurations on a hardware setup that supports memory disaggregation. We expand upon prior work by exploring a range of solutions that cater to varying tolerances for job completion latency, introducing new points to the latency-memory usage Pareto. Notably, our cache-remote approach enhances the efficiency of current disaggregated memory allocation strategies. It achieves a substantial reduction in local memory utilization—up to ${2 4. 8 \%}$—while incurring minimal execution time overhead of merely $7 \%$, compared to local-only policies.
Achilleas Tzenetopoulos, Michele Gazzetti, Dimosthenis Masouros, Christian Pinto, Sotirios Xydis, Dimitrios Soudris
IC2E1
2023 IRIS: Interference and Resource Aware Predictive Orchestration for ML Inference Serving
abstract
Over the last years, the ever-growing number of Machine Learning(ML) and Artificial Intelligence(AI) applications deployed in the Cloud has led to high demands on the computing resources required for efficient processing. Multiple users deploy multiple applications on the same server node to maximize Quality of Service(QoS); however, this leads to increased interference. In addition, Cloud providers aim to minimize their operating costs by efficiently utilizing the available resources. These conflicting optimization goals form a complex paradigm where efficient scheduling is required. In this work, we present IRIS, an interference- and resource-aware predictive inference scheduling framework for ML inference serving in the cloud. We target the multi-objective problem of QoS maximization with effective CPU utilization based on Queries per Second(QPS) predictions by proposing a modelless ML-based solution and integrating it into the Kubernetes platform. Our approach is evaluated over real hardware infrastructure and a set of ML applications. Our experimental analysis shows that under various QoS constraints, the model specific interference-aware scheduler violates QoS constraints less frequently by achieving 1.8x fewer violations, on average, compared to over-provisioning and 3.1 x fewer violations compared to under-provisioning, through efficient exploitation of available CPU resources. The model-less feature is able to cause, on average, 1.5x fewer violations compared to the model-specific scheduler, while further reducing the average CPU utilization by$\approx 30{\%}$.
Aggelos Ferikoglou, Panos Chrysomeris, Achilleas Tzenetopoulos, Manolis Katsaragakis, Dimosthenis Masouros, Dimitrios Soudris
CLOUD3
2023 Darly: Deep Reinforcement Learning for QoS-aware scheduling under resource heterogeneity Optimizing serverless video analytics
abstract
Today, video analytics are becoming extremely popular due to the increasing need for extracting valuable information from videos available in public sharing services through camera-driven streams. Typically, video analytics are organized as a set of separate tasks, each of which has different resource requirements (e.g., computational- vs. memory-intensive tasks). The serverless computing paradigm forms a very promising approach for mapping such types of applications, as it enables fine-grained deployment and management in a per-function manner. However, modern serverless frameworks suffer from performance variability issues, due to i) the interference introduced due to co-location of third-party workloads with the serverless funcations and ii) the increasing hardware heterogeneity introduced in public clouds. To this end, this work introduces Darly, a QoS- and heterogeneity-aware Deep Reinforcement Learning-based Scheduler for serverless video analytics deployments. The proposed framework incorporates a DRL agent which exploits low-level performance counters to identify the levels of interference and the degree of heterogeneity in the underlying infrastructure and combines this information along with user-defined QoS requirements to dynamically optimize resource allocations by deciding the placement, migration, or horizontal scaling of serverless functions. Promising results are produced withing our experiments, which are accompanied with the intent to further build upon this groundwork.
Dimitrios Giagkos, Achilleas Tzenetopoulos, Dimosthenis Masouros, Dimitrios Soudris, Sotirios Xydis
CLOUD2
2022 Sequence Clock: A Dynamic Resource Orchestrator for Serverless Architectures
abstract
Function-as-a-service (FaaS) represents the next frontier in the evolution of cloud computing being an emerging paradigm that removes the burden of configuration and management issues from users. This is achieved by replacing the well-established monolithic approach with graphs of standalone, small, stateless, event-driven components called functions. At the same time, from the cloud providers’ perspective, problems such as availability, load balancing and scalability need to be resolved without being aware of the functionality, behavior or resource requirements of their tenants’ code. However, in this context, functions’ containers coexist with others inside a host of finite resources, where a passive resource allocation technique does not guarantee a well-defined quality of service (QoS) in regards to time latency. In this paper, we present Sequence Clock, an expandable latency targeting tool that actively monitors serverless invocations in a cluster and offers execution of a sequential chain of functions, also known as pipelines or sequences, while achieving the targeted time latency. Two regulation methods were utilized, with one of them achieving up to 82% decrease in the severity of time violations and in some cases even eliminating them completely.
Ioannis Fakinos, Achilleas Tzenetopoulos, Dimosthenis Masouros, Sotirios Xydis, Dimitrios Soudris
CLOUD2
2022 EVOLVE: Towards Converging Big-Data, High-Performance and Cloud-Computing Worlds
abstract
EVOLVE is a pan European Innovation Action that aims to fully-integrate High-Performance-Computing (HPC) hardware with state-of-the-art software technologies under a unique testbed, that enables the convergence of HPC, Cloud and Big-Data worlds and increases our ability to extract value from massive and demanding datasets. EVOLVE's advanced compute platform combines HPC-enabled capabilities, with transparent deployment in high abstraction level, and a versatile Big-Data processing stack for end-to-end workflows. Hence, domain experts have the potential to improve substantially the efficiency of existing services or introduce new models in the respective domains, e.g., automotive services, bus transportation, maritime surveillance and others. In this paper, we describe EVOLVE's testbed, and evaluate the performance of the integrated pilots from different domains.
Achilleas Tzenetopoulos, Dimosthenis Masouros, Konstantina Koliogeorgi, Sotirios Xydis, Dimitrios Soudris, Antony Chazapis, Christos Kozanitis, Angelos Bilas, Christian Pinto, Huy-Nam Nguyen, Stelios Louloudakis, Georgios Gardikis, George Vamvakas, Michelle Aubrun, Christi Symeonidou, Vassilis Spitadakis, Konstantinos F. Xylogiannopoulos, Bernhard Peischl, Tahir Emre Kalayci, Alexander Stocker, Jean-Thomas Acquaviva
DATE1
2021 FADE: FaaS-inspired application decomposition and Energy-aware function placement on the Edge
abstract
Lately, more and more applications are deployed on heterogeneous, power-constrained edge-computing devices. Bringing computation closer to the data, contributes both to latency and energy consumption reduction due to the elimination of excessive data transfers. However, while the main concern in such environments is the minimization of energy consumption, the heterogeneity in compute resources found at the edge may lead to Quality of Service (QoS) violations. At the same time, Serverless computing, the next frontier of Cloud computing has emerged to offer unprecedented elasticity by utilizing fine-grained, stateless functions. The reduction in the execution time and the modest memory footprint of such decomposed applications, allow for fine-grained resource multiplexing. In this work, we propose a methodology for application decomposition into fine-grained functions and energy-aware function placement on a cluster of edge devices subject to user-specified QoS guarantees.
Achilleas Tzenetopoulos, Charalampos Marantos, Giannos Gavrielides, Sotirios Xydis, Dimitrios Soudris
SCOPES1