László Toka

dblp:96/5813 · DBLP profile ↗
← Back
39ranked-venue papers
15as first author
15since 2021 · last 2026
0000-0003-1045-9205ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 20 · 10 first-author · 3 since 2021Systems, architecture and hardware · 5 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
5 papers
Software-defined and programmable networks · 42% Routing and switching · 27% Network optimization and economics · 13%
Computer architecture, parallel and distributed computing, and storage systems
7 papers
Cloud and datacenter computing · 65% Storage systems · 19% Distributed systems · 10%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%
Computer graphics and multimedia
1 paper
Virtual and augmented reality · 100%

Topics — the 21 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software-defined and programmable networks
network function virtualization
0.822020
5G Applications From Vision to Reality: Multi-Operator Orchestration · IEEE J. Sel. Areas Commun. 2020
FERO: Fast and Efficient Resource Orchestrator for a Data Plane Built on Docker and DPDK · INFOCOM 2018
Software-defined and programmable networks
SDN deployment
0.412020
Transition to SDN is HARMLESS: Hybrid Architecture for Migrating Legacy Ethernet Switches to SDN · IEEE/ACM Trans. Netw. 2020
Virtual and augmented reality › virtual reality
VR platform
0.412019
Towards Human-Robot Collaboration: An Industry 4.0 VR Platform with Clouds Under the Hood · ICNP 2019
Human-robot interaction › robot navigation
collision avoidance
0.412019
Towards Human-Robot Collaboration: An Industry 4.0 VR Platform with Clouds Under the Hood · ICNP 2019
Human-robot interaction
human-robot collaboration
0.412019
Towards Human-Robot Collaboration: An Industry 4.0 VR Platform with Clouds Under the Hood · ICNP 2019
Routing and switching
data plane
0.312018
FERO: Fast and Efficient Resource Orchestrator for a Data Plane Built on Docker and DPDK · INFOCOM 2018
Cloud and datacenter computing › virtualization › containerization
container-based virtualization
0.312018
FERO: Fast and Efficient Resource Orchestrator for a Data Plane Built on Docker and DPDK · INFOCOM 2018
Cloud and datacenter computing
virtualization
0.312018
FERO: Fast and Efficient Resource Orchestrator for a Data Plane Built on Docker and DPDK · INFOCOM 2018
Network optimization and economics
resource allocation
0.312017
A resource-aware and time-critical IoT framework · INFOCOM 2017
Distributed systems
peer-to-peer systems
0.222012
Redundancy management for P2P backup · INFOCOM 2012
Managing a Peer-to-Peer Data Storage System in a Selfish Society · IEEE J. Sel. Areas Commun. 2008
Cloud and datacenter computing › computation offloading
cloud offloading
0.222019
Towards Human-Robot Collaboration: An Industry 4.0 VR Platform with Clouds Under the Hood · ICNP 2019
A resource-aware and time-critical IoT framework · INFOCOM 2017
Storage systems › storage reliability
durability
0.112012
Redundancy management for P2P backup · INFOCOM 2012
Hardware reliability and fault tolerance › redundancy
redundancy management
0.112012
Redundancy management for P2P backup · INFOCOM 2012
Storage systems
storage reliability
0.112012
Redundancy management for P2P backup · INFOCOM 2012
Cloud and datacenter computing › datacenter network
datacenter network management
0.112020
Transition to SDN is HARMLESS: Hybrid Architecture for Migrating Legacy Ethernet Switches to SDN · IEEE/ACM Trans. Netw. 2020
Cellular and mobile networks
resource orchestration
0.112018
FERO: Fast and Efficient Resource Orchestrator for a Data Plane Built on Docker and DPDK · INFOCOM 2018
Network optimization and economics › mechanism design
incentive mechanism
0.112008
Managing a Peer-to-Peer Data Storage System in a Selfish Society · IEEE J. Sel. Areas Commun. 2008
Storage systems › distributed storage
peer-to-peer storage
0.112008
Managing a Peer-to-Peer Data Storage System in a Selfish Society · IEEE J. Sel. Areas Commun. 2008
Storage systems › storage reliability
data recovery
0.012012
Redundancy management for P2P backup · INFOCOM 2012
Algorithmic game theory and mechanism design
non-cooperative game
0.012008
Managing a Peer-to-Peer Data Storage System in a Selfish Society · IEEE J. Sel. Areas Commun. 2008
Algorithmic game theory and mechanism design
social welfare
0.012008
Managing a Peer-to-Peer Data Storage System in a Selfish Society · IEEE J. Sel. Areas Commun. 2008

Methods — techniques the papers use, named apart from their topics

digital twin · 1.1software switch emulation · 0.9openflow · 0.9information modeling · 0.9graph-based embedding algorithm · 0.9service graph embedding · 0.7open vswitch · 0.7docker · 0.7DPDK · 0.7optimization · 0.6prediction · 0.3mechanism design · 0.1game theory · 0.1
YearPublicationVenuePosition
2026 Temporal Graph Network Framework for Quantifying Pass Reception Probabilities Against Defensive Structures
abstract
Abstract Passing decisions in soccer are heavily influenced by the opposing team’s defensive organization. Existing approaches often decompose the problem into pass selection and success probabilities, sometimes incorporating pressure-related features to account for defensive constraints. In this study, we propose a framework for evaluating passing decisions against defensive structures using temporal graph networks (TGNs). Rather than separately modeling selection and success, we estimate the probability of a pass being received by each teammate or intercepted by an opponent, leveraging temporal, spatial, and relational data to capture dynamic interactions. We focus on forward passes originating in the middle third of the pitch, where teams frequently encounter structured defensive shapes. Specifically, we analyze passes that (1) bypass defensive lines, (2) reach teammates inside defensive structure, or (3) penetrate, exit, and reach teammates outside defensive structure. Our evaluation compares pass reception prediction accuracy against baselines, examines receiver availability against defensive structure, and assesses the situational value of passing options. The results suggest that TGNs improve pass probability estimates while offering practical insights for decision-making against organized defenses.
Pegah Rahimian, Jesse Davis, László Toka
Mach. Learn.3
2025 Collaborative HD Map Creation: A Stackelberg Evolutionary Game Approach
abstract
High-definition (HD) maps are becoming essential for Advanced Driver Assistance Systems and Fully Self-Driving technology. Given the vast scale and dynamic nature of road networks, crowdsensing appears to be the only viable solution for maintaining these maps. However, sustaining user contributions requires effective incentive mechanisms. In this work, we analyze such a system using game theory, nonlinear programming, and deep learning. Specifically, we employ a Stackelberg evolutionary game framework and optimize incentive strategies using a Deep Deterministic Policy Gradient method enhanced with Evolutionary Algorithms. Our recent theoretical results highlight the necessity of incentives to ensure sustained participation, while also revealing challenges in rural areas due to participation fluctuations, which may impact long-term data reliability.
Marcell Szabó, László Toka
VTC2025-Fall2
2025 Routing LEO satellite traffic in adverse weather
abstract
This research investigates the planning and optimization of sixth-generation (6G) non-terrestrial networks (NTNs) by leveraging Free-Space Optical (FSO) communication to enhance speed, coverage, and resilience in global communication. Building on the theoretical foundation of NTN technologies, this study develops a dynamic simulation model to represent interactions among satellites, ground stations, and environmental variables. Our model allows detailed analysis of weather-induced disruptions, revealing that conditions such as rain and fog can significantly impact radio and FSO-based communication by reducing network speed and reliability. Simulation results show that integrating FSO links alongside traditional microwave connections enhances network adaptability, reducing latency and hop counts, particularly in high-traffic scenarios. In stress tests simulating ground station outages, the network demonstrated resilience by dynamically rerouting traffic over inter-orbit inter-satellite links, maintaining continuity even under significant regional disruptions. These findings underscore the importance of hybrid connectivity models and adaptive routing strategies in 6G NTN planning, providing critical insights for designing robust, high-performance networks capable of meeting future connectivity demands.
László Toka, Zoltan Illes, Endre Angelus Papp, László Hévizi, István Gódor
Comput. Networks1
2024 In-game soccer outcome prediction with offline reinforcement learning
abstract
Abstract Predicting outcomes in soccer is crucial for various stakeholders, including teams, leagues, bettors, the betting industry, media, and fans. With advancements in computer vision, player tracking data has become abundant, leading to the development of sophisticated soccer analytics models. However, existing models often rely solely on spatiotemporal features derived from player tracking data, which may not fully capture the complexities of in-game dynamics. In this paper, we present an end-to-end system that leverages raw event and tracking data to predict both offensive and defensive actions, along with the optimal decision for each game scenario, based solely on historical game data. Our model incorporates the effectiveness of these actions to accurately predict win probabilities at every minute of the game. Experimental results demonstrate the effectiveness of our approach, achieving an accuracy of 87% in predicting offensive and defensive actions. Furthermore, our in-game outcome prediction model exhibits an error rate of 0.1, outperforming counterpart models and bookmakers’ odds.
Pegah Rahimian, Balazs Mark Mihalyi, László Toka
Mach. Learn.3
2023 A Comprehensive Performance Analysis of Stream Processing with Kafka in Cloud Native Deployments for IoT Use-cases
abstract
The constant growth of the number of Internet of Things devices drives a huge increase in data that needs to be analyzed, at times in real time. Multiple platforms are available for delivering such data to analytics engines that can perform various operations on the data with low processing latency. These platforms can find their home in cloud native environments where high availability and scaling to the actual workload can be easily achieved. While the deployment environment is elastic, clusters still need to be adequately dimensioned to accommodate the components of the platforms even under high load.In this paper, we provide an analysis in this regard: we discuss key performance indicators of the popular Kafka message bus and the related Kafka Streams processing engine. Namely, we analyze latency, throughput, CPU and memory resource footprint aspects of these services under varying load and processing tasks that appear in Internet of Things applications. We find subsecond processing latency and linear but heavily task-dependent scaling behavior in the other performance indicators’ case.
István Pelle, Bence Szoke, Abdulhalim Fayad, Tibor Cinkler, László Toka
NOMS5
2023 5G on the roads: optimizing the latency of federated analysis in vehicular edge networks
abstract
At the dawn of autonomous driving, vehicular communications and coordination become more vital than ever. Fast information gathering, processing and sharing creates the basis of safety and efficiency, the main promises of conceding the control of vehicles from humans to machines. In this paper we propose to deploy an information gathering and distributing system that aims exactly at minimizing the latency of delivering the essential information to the end clients. We specifically tackle the crowd-sourced maintenance of high definition maps, i.e., road maps with extremely high accuracy and environmental fidelity containing dynamic information about the traffic as well, via a federated analysis scheme, and by broadcasting those maps through a 5G network. The system is designed for minimizing the latency of information delivery: analytical models based on queuing theory and optimization are proposed, and a wide range of system parameters are evaluated in numerical simulations. We find that the latency of delivering timely high quality information to end clients can be reduced with careful dimensioning of the system. According to our measurements, high-speed 5G data connection is a must, as we reach the optimal latency by building map segments with 1km in diameter via Gb/s uplink speeds in densely populated central metropolitan settings.
László Toka, Márk Konrád, István Pelle, Balázs Sonkoly, Marcell Szabó, Bhavishya Sharma, Shashwat Kumar, Madhuri Annavazzala, Sree Teja Deekshitula, A. Antony Franklin
NOMS1
2023 Federated learning for vehicular coordination use cases
abstract
Vehicular coordination and communication tasks are crucial aspects of enabling autonomous driving, guaranteeing safety and efficiency. In our present work, we explore methods for collecting and distributing information among participants by employing collaboratively-built high-definition maps that contain fine-grained contextual data. We leverage a hierarchical federated learning structure and anticipatory onboarding of the maps through a mobility-aware content caching scheme and minimize the delay of data delivery in both subsystems. We provide analytical models built on queuing theory and integer linear programming and evaluate essential system parameters in an emulation testbed. Based on our results, we conclude that we can significantly reduce the delay in delivering timely information to vehicular clients by introducing intermediary layers in the federated learning structure and by pre-loading current map tiles corresponding to vehicle paths.
László Toka, Márk Konrád, István Pelle, Balázs Sonkoly, Marcell Szabó, Bhavishya Sharma, Shashwat Kumar, Madhuri Annavazzala, Sree Teja Deekshitula, A. Antony Franklin
NOMS1
2023 Real-Time FaaS: Towards a Latency Bounded Serverless Cloud
abstract
Today, Function-as-a-Service is the most promising concept of serverless cloud computing. It makes possible for developers to focus on application development without any system management effort: FaaS ensures resource allocation, fast response time, schedulability, scalability, resiliency, and upgradability. Applications of 5G, IoT, and Industry 4.0 raise the idea to open cloud-edge computing infrastructures for time-critical applications too, i.e., there is a strong desire to pose real-time requirements for computing systems like FaaS. However, multi-node systems make real-time scheduling significantly complex since guaranteeing real-time task execution and communication is challenging even on one computing node with multi-core processors. In this paper, we present an analytical model and a heuristic partitioning scheduling algorithm suitable for real-time FaaS platforms of multi-node clusters. We show that our task scheduling heuristics could outperform existing algorithms by 55%. Furthermore, we propose three conceptual designs to enable the necessary real-time communications. We present the architecture of the envisioned real-time FaaS platform, emphasize its benefits and the requirements for the underlying network and nodes, and survey the related work that could meet these demands.
Mark Szalay, Péter Mátray, László Toka
IEEE Trans. Cloud Comput.3
2022 Optimizing Performance and Resource Consumption of Cloud-Native Logging Application Stacks
abstract
Nowadays cloud-based applications and Internet of Things use-cases are becoming more and more common in the field of IT benefiting from the virtually limitless resources and microservice-based deployment options available in the cloud. Observability in such environments is key for tracing application execution to detect possible malfunctions and anomalies. Collecting logs can greatly help in this regard, however, a high volume of logging data can add huge costs for the maintenance of the infrastructure gathering monitoring data. In order to increase the profitability of the application, monitoring-related infrastructure needs to have the lowest cost possible while still being able to fully serve the application’s monitoring needs. In this work, we investigate this aspect and provide an evaluation of the resource footprint of one of the most prominent log collection services, Elastic Stack, from the perspective of its write path.
Gergo Csáti, István Pelle, László Toka
NOMS3
2022 Optimal Resource Provisioning for Data-intensive Microservices
abstract
With the continuous progress of cloud computing, many microservices and complex multi-component applications arise for which resource planning is a great challenge. For example, when it comes to data-intensive cloud-native applications, the tenant might be eager to provision cloud resources in an economical manner while ensuring that the application performance meets the requirements in terms of data throughput. However, due to the complexity of the interplay between the building blocks, adequately setting resource limits of the components separately for various data rates is nearly impossible. In this paper, we propose a comprehensive approach that consists of measuring the resource footprint and data throughput performance of such a microservices-based application, analyzing the measurement results by data mining techniques, and finally formulating an optimization problem that aims to minimize the allocated resources given the performance constraints. We illustrate the benefits of the proposed approach on Cortex, an extension to Prometheus for storing monitored metrics data. The data-intensive nature of this illustrative example stems from real-time monitoring of metrics exposed by a multitude of applications running in a data center and the continuous analysis performed on the collected data that can be fetched from Cortex. We present Cortex’s performance vs resource footprint trade-off, and then we build regression models to predict the microservices’ resource consumption and draw a mathematical programming formulation to optimize the most important configuration parameters. Our most important finding is the linear relationship between resource consumption and application performance, which allows for applying linear regression and linear programming models. After the optimization, we compare our results to Cortex’s recommendation, leading to a CPU reservation reduced by 50-80%.
Roland Mark Erdei, László Toka
NOMS2
2021 Real-time task scheduling in a FaaS cloud
abstract
Today, Function-as-a-Service is the most promising concept of serverless cloud computing. It makes possible for developers to focus on application development without any system management effort: FaaS ensures resource allocation, fast response time, schedulability, scalability, resiliency, and upgrad-ability. Applications of 5G, IoT, and Industry 4.0 raise the idea to open cloud-edge computing infrastructures for time-critical applications too, i.e., there is a strong desire to pose real-time requirements for computing systems like FaaS. However, multinode systems make real-time scheduling significantly complex since guaranteeing real-time task execution is challenging even on one computing node with multi-core processors. In this paper, we present an analytical model and a heuristic partitioning scheduling algorithm for a partitioned scheduling system suitable for real-time FaaS platforms of multi-node clusters. We present the architecture of the envisioned real-time FaaS platform, emphasize its benefits and the requirements for the underlying network and nodes, and survey the related work that could meet these demands.
Mark Szalay, Péter Mátray, László Toka
CLOUD3
2021 Towards optimized actions in critical situations of soccer games with deep reinforcement learning
abstract
Soccer is a sparse rewarding game: any smart or careless action in critical situations can change the result of the match. Therefore players, coaches, and scouts are all curious about the best action to be performed in critical situations, such as the times with a high probability of losing ball possession or scoring a goal. This work proposes a new state representation for the soccer game and a batch reinforcement learning to train a smart policy network. This network gets the contextual information of the situation and proposes the optimal action to maximize the expected goal for the team. We performed extensive numerical experiments on the soccer logs made by InStat for 104 European soccer matches. The results show that in all 104 games, the optimized policy obtains higher rewards than its counterpart in the behavior policy. Besides, our framework learns policies that are close to the expected behavior in the real world. For instance, in the optimized policy, we observe that some actions such as foul, or ball out can be sometimes more rewarding than a shot in specific situations.
Pegah Rahimian, Afshin Oroojlooyjadid, László Toka
DSAA3
2021 Predicting cloud-native application failures based on monitoring data of cloud infrastructure
László Toka, Gergely Dobreff, Dávid Haja, Mark Szalay
IM1
2021 Ultra-Reliable and Low-Latency Computing in the Edge with Kubernetes
abstract
Abstract Novel applications will require extending traditional cloud computing infrastructure with compute resources deployed close to the end user. Edge and fog computing tightly integrated with carrier networks can fulfill this demand. The emphasis is on integration: the rigorous delay constraints, ensuring reliability on the distributed, remote compute nodes, and the sheer scale of the system altogether call for a powerful resource provisioning platform that offers the applications the best of the underlying infrastructure. We therefore propose Kubernetes-edge-scheduler that provides high reliability for applications in the edge, while provisioning less than 10% of resources for this purpose, and at the same time, it guarantees compliance with the latency requirements that end users expect. We present a novel topology clustering method that considers application latency requirements, and enables scheduling applications even on a worldwide scale of edge clusters. We demonstrate that in a potential use case, a distributed stream analytics application, our orchestration system can reduce the job completion time to 40% of the baseline provided by the default Kubernetes scheduler.
László Toka
J. Grid Comput.1
2021 Machine Learning-Based Scaling Management for Kubernetes Edge Clusters
abstract
Kubernetes, the container orchestrator for cloud-deployed applications, offers automatic scaling for the application provider in order to meet the ever-changing intensity of processing demand. This auto-scaling feature can be customized with a parameter set, but those management parameters are static while incoming Web request dynamics often change, not to mention the fact that scaling decisions are inherently reactive, instead of being proactive. We set the ultimate goal of making cloud-based applications' management easier and more effective. We propose a Kubernetes scaling engine that makes the auto-scaling decisions apt for handling the actual variability of incoming requests. In this engine various machine learning forecast methods compete with each other via a short-term evaluation loop in order to always give the lead to the method that suits best the actual request dynamics. We also introduce a compact management parameter for the cloud-tenant application provider to easily set their sweet spot in the resource over-provisioning vs. SLA violation trade-off. We motivate our scaling solution with analytical modeling and evaluation of the current Kubernetes behavior. The multi-forecast scaling engine and the proposed management parameter are evaluated both in simulations and with measurements on our collected Web traces to show the improved quality of fitting provisioned resources to service demand. We find that with just a few, but fundamentally different, and competing forecast methods, our auto-scaler engine, implemented in Kubernetes, results in significantly fewer lost requests with just slightly more provisioned resources compared to the default baseline.
László Toka, Gergely Dobreff, Balázs Fodor, Balázs Sonkoly
IEEE Trans. Netw. Serv. Manag.1
2020 Adaptive AI-based auto-scaling for Kubernetes
abstract
Kubernetes, the prevalent container orchestrator for cloud-deployed web applications, offers an automatic scaling feature for the application provider in order to meet the ever-changing amount of demand from its clients. This auto-scaling service, however, requires a seemingly difficult parameter set to be customized by the application provider, and those management parameters are static while incoming web request dynamics often change, not to mention the fact that scaling decisions are inherently reactive, instead of being proactive. Therefore we set the ultimate goal of making cloud-based web applications' management easier and more effective. We propose a Kubernetes scaling engine that makes the auto-scaling decisions apt for handling the actual variability of incoming requests. In this engine various AI-based forecast methods compete with each other via a short-term evaluation loop in order to always give the lead to the method that suits best the actual request dynamics, as soon as possible. We also introduce a compact management parameter for the cloud-tenant application provider in order to easily set their sweet spot in the resource over-provisioning vs. SLA violation trade-off. The multi-forecast scaling engine and the proposed management parameter are evaluated both in simulations and with measurements on our collected web traces to show the improved quality of fitting provisioned resources to service demand. We find that with just a few competing forecast methods, our auto-scaling engine, implemented in Kubernetes, results in significantly less lost requests with slightly more provisioned resources compared to the default baseline.
László Toka, Gergely Dobreff, Balázs Fodor, Balázs Sonkoly
CCGRID1
2020 AnnaBellaDB: Key-Value Store Made Cloud Native
abstract
The cloud-native paradigm has become a well-known approach to ensure the elasticity and reliability of applications running in the cloud. One recurrent motif is the stateless design of applications, which aims to decouple the life-cycle of application states from the life-cycle of individual application instances. Application data is written to and read from cloud databases, deployed close to the application code to ensure low latency bounds on state access. When applying a stateless design, the performance of the cloud service is often limited by the cloud database. In order not to become a bottleneck, database instances are distributed on multiple hosts, and strive to ensure data locality for all application functions. However, the shared nature of certain states, and the inevitable dynamics of the application workload necessarily lead to inter-host data access. If the service is geographically distributed, this is even across data centers and edge servers resulting in a significant delay. To minimize the service performance loss due to the stateless design of applications, we propose a latency and access pattern aware state storage method, called state-layer, that can be easily applied in any kind of key-value store with the ability of deciding where to store replicas in the cluster and measure networking/computing delay. By adapting our solution to Anna, a key-value store from academia, we show the proposed state-layer is ideal to use as a cloud database for storing application states. To foster further research in this area, we make our proof-of-concept solution open-source.
Mark Szalay, Péter Mátray, László Toka
CNSM3
2020 To boost or not to boost: a stochastic game in wireless access networks
abstract
Resource allocation in wireless access networks has been an intensively researched topic recently: many proposed solutions tackle radio channel access and dynamic spectrum allocation, but traditional issues of queuing, bandwidth sharing and packet processing at wireless access points have been targeted as well. In most of the related work the competition for high quality of service is usually solved by central coordination among users via optimizing a specific target aspect of the overall communication. In this paper we take a turn and provide users with the possibility of resource allocation suggestions. We propose a wireless access sharing framework, in which users have a say in optimizing their quality of service on the long term, and we tackle its analysis with the tool set of stochastic game theory. Our findings show that greedy users become polite against their counterparts when the load is relatively low with the goal of preparing for situations with high load.
László Toka, Mark Szalay, Dávid Haja, Géza Szabó, Sándor Rácz, Miklós Telek
ICC1
2020 Scalable edge cloud platforms for IoT services
abstract
Nowadays, online applications are moving to the cloud, and for delay-sensitive ones, the cloud is being extended with edge/fog domains. Emerging cloud platforms that tightly integrate compute and network resources enable novel services, such as versatile IoT (Internet of Things), augmented reality or Tactile Internet applications. Virtual infrastructure managers (VIMs), network controllers and upper-level orchestrators are in charge of managing these distributed resources. A key and challenging task of these orchestrators is to find the proper placement for software components of the services. As the basic variant of the related theoretical problem (Virtual Network Embedding) is known to be NP-hard, heuristic solutions and approximations can be addressed. In this paper, we propose two architecture options together with proof-of-concept prototypes and corresponding embedding algorithms, which enable the provisioning of delay-sensitive IoT applications. On the one hand, we extend the VIM itself with network-awareness, typically not available in today's VIMs. On the other hand, we propose a multi-layer orchestration system where an orchestrator is added on top of VIMs and network controllers to integrate different resource domains. We argue that the large-scale performance and feasibility of the proposals can only be evaluated with complete prototypes, including all relevant components. Therefore, we implemented fully-fledged solutions and conducted large-scale experiments to reveal the scalability characteristics of both approaches. We found that our VIM extension can be a valid option for single-provider setups encompassing even 100 edge domains (Points of Presence equipped with multiple servers) and serving a few hundreds of customers. Whereas, our multi-layer orchestration system showed better scaling characteristics in a wider range of scenarios at the cost of a more complex control plane including additional entities and novel APIs (Application Programming Interfaces).
Balázs Sonkoly, Dávid Haja, Balázs Németh 0001, Mark Szalay, János Czentye, Róbert Szabó, Rehmat Ullah 0001, Byung-Seo Kim, László Toka
J. Netw. Comput. Appl.9
2020 5G Applications From Vision to Reality: Multi-Operator Orchestration
abstract
Envisioned 5G applications and services, such as Tactile Internet, Industry 4.0 use-cases, remote control of drone swarms, pose serious challenges to the underlying networks and cloud platforms. On the one hand, evolved cloud infrastructures provide the IT basis for future applications. On the other hand, networking is in the middle of a momentous revolution and important changes are mainly driven by Network Function Virtualization (NFV) and Software Defined Networking (SDN). A diverse set of cloud and network resources, controlled by different technologies and owned by cooperating or competing providers, should be coordinated and orchestrated in a novel way in order to enable future applications and fulfill application level requirements. In this paper, we propose a novel cross domain orchestration system which provides wholesale XaaS (Anything as a Service) services over multiple administrative and technology domains. Our goal is threefold. First, we design a novel orchestration system exploiting a powerful information model and propose a versatile embedding algorithm with advanced capabilities as a key enabler. The main features of the architecture include i) efficient and multi-purpose service embedding algorithms which can be implemented based on graph models, ii) inherent multidomain support, iii) programmable aggregation of different resources, iv) information hiding together with flexible delegation of certain requirements enabling multi-operator use-cases, and v) support for legacy technologies. Second, we present our proof-of-concept prototype implementing the proposed system. Third, we establish a dedicated test environment spanning across multiple European sites encompassing sandbox environments from both operators and the academia in order to evaluate the operation of the system. Dedicated experiments confirm the feasibility and good scalability of the whole framework.
Balázs Sonkoly, Róbert Szabó, Balázs Németh 0001, János Czentye, Dávid Haja, Mark Szalay, Janos Doka, Balázs Péter Gero, Dávid Jocha, László Toka
IEEE J. Sel. Areas Commun.10
2020 Transition to SDN is HARMLESS: Hybrid Architecture for Migrating Legacy Ethernet Switches to SDN
abstract
Software-Defined Networking (SDN) offers a new way to operate, manage, and deploy communication networks and to overcome many long-standing problems of legacy networking. However, widespread SDN adoption has not occurred yet due to the lack of a viable incremental deployment path and the relatively immature present state of SDN-capable devices on the market. While continuously evolving software switches may alleviate the operational issues of commercial hardware-based SDN offerings, namely lagging standards-compliance, performance regressions, and poor scaling, they fail to match the cost-efficiency and port density. In this paper, we propose HARMLESS, a new SDN switch design that seamlessly adds SDN capability to legacy network gear, by emulating the OpenFlow switch OS in a separate software switch component. This way, HARMLESS enables a quick and easy leap into SDN, combining the rapid innovation and upgrade cycles of software switches with the port density and cost-efficiency of hardware-based appliances into a fully dataplane-transparent and vendor-neutral solution. HARMLESS incurs an order of magnitude smaller initial expenditure for an SDN deployment than existing turnkey vendor SDN solutions while, at the same time, yields matching, or even better, data plane performance for smaller enterprises.
Levente Csikor, Mark Szalay, Gábor Rétvári, Gergely Pongrácz, Dimitrios P. Pezaros, László Toka
IEEE/ACM Trans. Netw.6
2019 Improving Big Data Application Performance in Edge-Cloud Systems
abstract
Data analysis is widely used in all domains of the economy. While the amount of data to process grows, the time criteria and the resource consumption constraints get stricter. These phenomena call for advanced resource orchestration for the big data applications. The challenge is actually even greater at the advent of edge computing: orchestration of big data resources in a hybrid edge-cloud infrastructure is challenging. The difficulty stems from the fact that wide-area networking and all its well-known issues come into play and affect the performance of the application. In this paper we present the steps we made towards network-aware big data application design over such distributed systems. We propose a HDFS block placement algorithm for the network reliability problem we identify in geographically distributed topologies. The heuristic algorithm we propose provides better big data application performance compared to the default block placement method. We implement our solution in our simulation environment and show the improved quality of big data applications.
Dávid Haja, Balázs Vass, László Toka
CLOUD3
2019 Industrial-Scale Stateless Network Functions
abstract
While the industry is still struggling to embrace the network function virtualization paradigm, recently a novel approach has appeared with the promise of improving the state-of-the-art: stateless virtualized network functions. Rooted in cloud-native computing, this design outsources the state embedded in virtual network functions to a dedicated "state storage" layer, facilitating elastic scaling and resiliency. While related work mostly focuses on performance, we in this paper pinpoint all other factors that weigh in when it comes to deploying the stateless design in a carrier-grade operator network. Among those we argue that reliability and flexibility are key, and we propose a system design that can be adapted to any telco use case without the need for complex coordination among the network control, the stateless network functions, and the state storage backend. Then, in extensive evaluations on synthetic use cases we show that the additional flexibility provided by our design does not come at a performance penalty; in fact, in certain cases our design outperforms the state-of-the-art significantly. Finally, we present what to our knowledge is the first product-phase realization of the stateless paradigm, an operational virtualized IP Multimedia Subsystem that can restore the live call records of thousands of mobile subscribers under a couple of seconds with half the resources required by a traditional "stateful" design.
Mark Szalay, Máté Nagy 0002, Daniel Gehberger, Zoltán Lajos Kis, Péter Mátray, Felician Németh, Gergely Pongrácz, Gábor Rétvári, László Toka
CLOUD9
2019 Optimizing Latency Sensitive Applications for Amazon's Public Cloud Platform
abstract
Recent cloud technologies enable a diverse set of novel applications with capabilities never seen before. Cloud native programming, microservices, serverless architectures are novel paradigms reducing the burden on both software developers and operators while enabling cloud-grade service deployments. Several types of applications fit in well with the new concepts, however, latency sensitive applications with strict delay constraints pose additional challenges on the platforms. Can we run these applications on today's public cloud platforms making use of the brand new tools and techniques? In this paper, we try to answer this question by addressing one of the most widely used and versatile public cloud platforms, namely Amazon's AWS, and we propose a novel mechanism to optimize the software "layout" based on dynamic performance measurements. Our contribution is threefold. First, we define a combined performance and cost model on CaaS/FaaS (Container/Function as a Service) platforms, specifically for AWS, based on a comprehensive performance analysis, and we also provide an application model capturing the performance requirements. Second, we formulate an optimization problem which minimizes the deployment costs on AWS while meeting the latency constraints. A polynomial algorithm finding the optimal solution is also given. Third, we evaluate the model and the algorithm for different scenarios and investigate the performance on today's system.
János Czentye, István Pelle, András Kern, Balázs Péter Gero, László Toka, Balázs Sonkoly
GLOBECOM5
2019 Towards Human-Robot Collaboration: An Industry 4.0 VR Platform with Clouds Under the Hood
abstract
Safe and efficient Human-Robot Collaboration (HRC) is an essential feature of future Industry 4.0 production systems which requires sophisticated collision avoidance mechanisms with intense computation need. Digital twins provide a novel way to test the impact of different control decisions in a simulated virtual environment even in parallel. In addition, Virtual/Augmented Reality (VR/AR) applications can revolutionize future industry environments. Each component requires extreme computational power which can be provided by cloud platforms but at the cost of higher delay and jitter. Moreover, clouds bring a versatile set of novel techniques easing the life of both developers and operators. Can these applications be realized and operated on today's systems? In this demonstration, we give answers to this question via real experiments.
Bálint György Nagy, Janos Doka, Sándor Rácz, Géza Szabó, István Pelle, János Czentye, László Toka, Balázs Sonkoly
ICNP7
2018 FERO: Fast and Efficient Resource Orchestrator for a Data Plane Built on Docker and DPDK
abstract
Future services and applications, such as Tactile Internet, coordinated remote driving or wireless controlled exoskeletons, pose serious challenges on the underlying networks and IT platforms in terms of reliability, latency, or capacity, just to mention a few. Towards those services, virtualization is a key enabler from both technological and economic aspects which significantly reshaped the IT and networking ecosystem. On the one hand, cloud computing and the services based on that are evident results of last years' efforts; on the other hand, networking is in the middle of a momentous revolution and important changes mainly driven by Network Function Virtualization (NFV) and Software Defined Networking (SDN). In order to enable carrier grade network services with strict QoS requirements, we need a novel data plane supporting high performance and flexible, fine granular programmability and control. As the network functions (implemented by virtual machines or containers) use the same hardware resources (cpu, memory) as the components responsible for networking, we need a low-level resource orchestrator which is capable of jointly controlling these resources. In this paper, we propose a novel resource orchestrator (RO) for a data plane making use of open source components such as, Docker, DPDK and OVS. Our goal is threefold. First, we propose a novel data plane resource model which is capable of abstracting several hardware architectures. Second, we provide an adapter module which can automatically discover the underlying hardware and build the model on-the-fly. Third, we design and implement a novel RO building on the aforementioned components and a publicly available Service Graph embedding engine. As a proof of the concept, two software switches (OVS, ERFS) are adapted and different hardware platforms are evaluated
Balázs Sonkoly, Marton Szabo, Balázs Németh 0001, András Majdán, Gergely Pongrácz, László Toka
INFOCOM6
2018 Realizing services and slices across multiple operator domains
abstract
Supporting end-to-end network slices and services across operators has become an important use case of study for 5G networks as can be seen by 5G use cases published in 3GPP, ETSI as well as NGMN. This paper presents the in- depth architecture, implementation and experiment on a multi-domain orchestration framework that is ab le to deploy such multi-operator service as well as monitor the service for SLA compliance. Our implemented architecture allows operators to abstract their sensitive details while exposing the relevant amount of information to support inter-operator slice creation. Our experiment shows that the implemented framework is capable of creating services across operators while fulfilling the respective service requirements.
Ishan Vaishnavi, János Czentye, Molka Gharbaoui, Giovanni Giuliani, Dávid Haja, János Harmatos, Dávid Jocha, Yoonhee Kim, Barbara Martini, Javier Melian, Paolo Monti 0001, Balázs Németh 0001, Wint Yi Poe, Aurora Ramos, Andrea Sgambelluri, Balázs Sonkoly, László Toka, Francesco Tusa, Carlos J. Bernardos, Róbert Szabó
NOMS17
2017 On Pricing of 5G Services
abstract
IT and telco providers are preparing for the era of 5G; in terms of technology, the driving force is virtualization, both for computing and networking. The 5G services will be superior than today's online services not only in technological aspects, but also from an economic and business perspective: fast service creation, effective utilization of resources, dynamic adaption to actual demand are all direct benefits of the virtualized infrastructure. In this paper we study the economic interactions between 5G resource providers and customers: we formalize how resources should be priced and selected for being booked. In particular we show that usage-based pricing is an income-maximizing scheme for providers, and we derive the problem the customers need to solve for cost-optimizing service deployment.
László Toka, János Tapolcai, George Darzanos, Balázs Sonkoly
GLOBECOM1
2017 A resource-aware and time-critical IoT framework
abstract
Internet of Things (IoT) systems produce great amount of data, but usually have insufficient resources to process them in the edge. Several time-critical IoT scenarios have emerged and created a challenge of supporting low latency applications. At the same time cloud computing became a success in delivering computing as a service at affordable price with great scalability and high reliability. We propose an intelligent resource allocation system that optimally selects the important IoT data streams to transfer to the cloud for processing. The optimization runs on utility functions computed by predictor algorithms that forecast future events with some probabilistic confidence based on a dynamically recalculated data model. We investigate ways of reducing specifically the upload bandwidth of IoT video streams and propose techniques to compute the corresponding utility functions. We built a prototype for a smart squash court and simulated multiple courts to measure the efficiency of dynamic allocation of network and cloud resources for event detection during squash games. By continuously adapting to the observed system state and maximizing the expected quality of detection within the resource constraints our system can save up to 70% of the resources compared to the naive solution.
László Toka, Balázs Lajtha, Éva Hosszu, Bence Formanek, Daniel Gehberger, János Tapolcai
INFOCOM1
2016 Sharing is Power: Incentives for Information Exchange in Multi-Operator Service Delivery
abstract
A majority of 5G verticals have the potential to generate large revenues, but are expected to have strict Quality of Service (QoS) guarantees, and are projected to be delivered as a service chain of multiple, independent operators. Such multi-operator service delivery requires a set of interdependent Service Level Agreements (SLAs) between operators. The amount and aggregation-level of information shared between stakeholders inside such SLAs will determine how efficient the coordinated traffic engineering between the operators will be. Sharing more details on one's network is uncommon in today's interactions due to the fear of losing competitive advantage and regulations with regard to national security. In this paper, we analyze the economic incentives for information exchange in the context of multi-operator service delivery. We show that the current practice of exchanging only highly aggregated information can lead to both significant under- and overestimation of the risk of not meeting user-facing Quality of Service guarantees. We also show that economic incentives for mutually sharing an optimal amount of information do exist, and optimal information exchange between operators is viable in the long run. Moreover, through a simple numerical example, we demonstrate how the mutually shared information and the resulting risk estimation affect the revenues of the operators from the end-user market. We believe this work opens up a new line of research connecting the economics of multi-operator service delivery and network performability.
Poul E. Heegaard, Gergely Biczók, László Toka
GLOBECOM3
2015 Adaptive redundancy management for durable P2P backup
abstract
We design and analyze the performance of a redundancy management mechanism for peer-to-peer backup applications . Armed with the realization that a backup system has peculiar requirements – namely, data is read over the network only during restore processes caused by data loss – redundancy management targets data durability , i.e. guaranteeing that data is not lost, rather than attempting to make each piece of information available at any time. In our approach each peer determines, in an on-line manner, an amount of redundancy sufficient to counter the effects of peer deaths, while preserving acceptable data restore times. Our experiments, based on trace-driven simulations, indicate that our mechanism can reduce the redundancy by a factor between two and three with respect to redundancy policies aiming for data availability. These results imply an according increase in storage capacity and decrease in time to complete backups, at the expense of longer times required to restore data. We believe this is a very reasonable price to pay, given the nature of the application. We complete our work with a discussion on practical issues, and their solutions, related to which encoding technique is more suitable to support our scheme.
Matteo Dell'Amico, Pietro Michiardi, László Toka, Pasquale Cataldi
Comput. Networks3
2012 Redundancy management for P2P backup
abstract
We propose a redundancy management mechanism for peer-to-peer backup applications. Since, in a backup system, data is read over the network only during restore processes caused by data loss, redundancy management targets data durability rather than attempting to make each piece of information availabile at any time. Each peer determines, in an on-line manner, an amount of redundancy sufficient to counter the effects of peer deaths, while preserving acceptable data restore times. Our experiments, based on trace-driven simulations, indicate that our mechanism can reduce the redundancy by a factor between two and three with respect to redundancy policies aiming for data availability. These results imply an according increase in storage capacity and decrease in time to complete backups, at the expense of longer times required to restore data.We believe this is a very reasonable price to pay, given the nature of the application.
László Toka, Pasquale Cataldi, Matteo Dell'Amico, Pietro Michiardi
INFOCOM1
2011 Data transfer scheduling for P2P storage
abstract
In Peer-to-Peer storage and backup applications, large amounts of data have to be transferred between nodes. In general, recipient of data transfers are not chosen randomly from the whole set of nodes in the Peer-to-Peer networks, but they are chosen according to peer selection rules imposing several criteria, such as resource contributions, position in DHTs, or trust between nodes. Imposing too stringent restrictions on the choice of nodes that are eligible to receive data can have a negative impact on the amount of time needed to complete data transfer, and scheduling choices influence this result as well. We formalize the problem of data transfer scheduling, and devise means for calculating (knowing a posteriori the availability patterns of nodes) optimal scheduling choices; we then propose and evaluate realistic scheduling policies, and evaluate their overheads in transfer times with respect to the optimal. We show that allowing even a small flexibility in choosing nodes after the peer selection step results in large improvements on time to complete transfers, and that even simple informed scheduling policies can significantly reduce transfer time overhead.
László Toka, Matteo Dell'Amico, Pietro Michiardi
Peer-to-Peer Computing1
2011 Incentivizing the global wireless village
Gergely Biczók, László Toka, András Gulyás, Tuan Anh Trinh, Attila Vidács
Comput. Networks2
2010 Online Data Backup: A Peer-Assisted Approach
abstract
In this work we study the benefits of a peer- assisted approach to online backup applications, in which spare bandwidth and storage space of end- hosts complement that of an online storage service. Via simulations, we analyze the interplay between two key aspects of such applications: data placement and bandwidth allocation. Our analysis focuses on metrics such as the time required to complete a backup and a restore operation, as well as the storage costs. We show that, by using adequate bandwidth allocation policies in which storage space at a cloud provider can be used temporarily, hybrid systems can achieve performance comparable to traditional client-server architectures at a fraction of the costs. Moreover, we explore the impact of mechanisms to impose fairness and conclude that a peer-assisted approach does not discriminate peers in terms of performance, but associates a storage cost to peers contributing with little resources.
László Toka, Matteo Dell'Amico, Pietro Michiardi
Peer-to-Peer Computing1
2009 Selfish Neighbor Selection in Peer-to-Peer Backup and Storage Applications
Pietro Michiardi, László Toka
Euro-Par2
2009 General distributed economic framework for dynamic spectrum allocation
László Toka, Attila Vidács
Comput. Commun.1
2008 A dynamic exchange game
abstract
Our work aims to study a game based on an extended variant of the stable fixtures problem where multiple matches can be established between pairs of players, moreover preference orders are subject to alteration due to player strategies.
László Toka, Pietro Michiardi
PODC1
2008 Managing a Peer-to-Peer Data Storage System in a Selfish Society
abstract
We compare two possible mechanisms to manage a peer-to-peer storage system, where participants can store data online on the disks of peers in order to increase data availability and accessibility. Due to the lack of incentives for peers to contribute to the service, we suggest that either each peer's use of the service be limited to her contribution level (symmetric schemes), or that storage space be bought from and sold to peers by a system operator that seeks to maximize profit. Using a noncooperative game model to take into account user selfishness, we study those mechanisms with respect to the social welfare performance measure, and give necessary and sufficient conditions for one scheme to socially outperform the other.
Patrick Maillé, László Toka
IEEE J. Sel. Areas Commun.2