Enes Bajrovic

dblp:69/7757 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Hybrid Reactive Autoscaling for Task-Based Pipelines on Kubernetes
abstract
We present Python-to-Kubernetes (PTK), a hybrid autoscaling framework for pipeline-oriented, task-based Python applications on Kubernetes. PTK coordinates queue-length-driven horizontal scaling for CPU, memory, and GPU, together with reactive in-place vertical scaling of CPU and memory. The framework introduces source-code annotations, enabling users to define task-specific scaling constraints and automatically generate Kubernetes manifests. A periodic controller uses utilization and queue metrics to coordinate horizontal and vertical scaling, improving resource efficiency while maintaining pipeline performance. In a streaming machine learning (ML) inference pipeline, PTK sustains the target throughput while reducing hourly cost by 40.6%, CPU by 32.1 %, and memory by 22.4%, and lowering the GPU count from 4 to 3, compared with an uncoordinated baseline that combines the Horizontal Pod Autoscaler (HPA) and the Vertical Pod Autoscaler (VPA). It also cuts peak cost by 23.6% compared with a queue-driven HPA baseline.
Andrey Nagiyev, Enes Bajrovic, Siegfried Benkner
CloudCom2
2025 Deadline-Aware Resource Allocation and Scheduling of Serverless Workloads on Heterogeneous Clusters
abstract
Serverless computing has become widely adopted as a cloud deployment model due to its ease of use and finegrained pay-as-you-go pricing. By hiding infrastructure complexity, it simplifies access to cloud resources and lets developers focus on application code. However, most serverless platforms operate on a best-effort basis and provide minimal control over performance tuning. Combined with limited visibility into underlying hardware, this makes it difficult to reliably meet Service Level Objectives (SLOs). To address this, we introduce DHRT, a deadline- and heterogeneity-aware scheduling and resource allocation framework for performance-critical serverless workloads. DHRT applies heuristic-driven online optimisation to iteratively refine resource estimates by leveraging real-time metrics and historical data from live executions. To fulfil SLOs, it accounts for both workload characteristics and node heterogeneity. We evaluate DHRT on synthetic workloads by comparing it against baseline scheduling and resource allocation policies commonly used in FaaS platforms. Results show that DHRT accurately estimates resource demands within a few live executions, eliminating the need for manual resource tuning. By exploiting node heterogeneity and dynamically scaling vCPU allocations as workloads near their deadlines, DHRT improves resource efficiency and significantly reduces deadline violations.
Matthias Fritz, Siegfried Benkner, Enes Bajrovic
CLUSTER3
2025 Provisioning of Kubernetes Clusters for Task-Based Python Applications
abstract
We present Python-to-Kubernetes (PTK), a framework that automates the provisioning and deployment of taskbased Python applications on Kubernetes. PTK introduces compact source-code annotations for tasks, resource needs, grouping, and data-size hints. From these annotations, it provisions an application-specific cluster, builds container images, generates manifests, and selects the data-transfer mechanism based on placement. A scoring-based mapping co-locates bandwidth-heavy neighbors to reduce cross-node traffic and right-sizes nodes after placement. On a six-task machine learning (ML) ResNet50 imageclassification pipeline (ImageNet$\mathbf{5 \%} \boldsymbol{/} \mathbf{1 0 \%}$subsets), PTK achieves up to$6.43 \times$faster runtime and$7.26 \times$lower cost per run than the Kubernetes Default Scheduler, with higher CPU/memory utilization and fewer/smaller nodes. These results indicate that lightweight annotations plus application-aware provisioning can substantially improve price-performance for Kubernetes-based ML pipelines.
Andrey Nagiyev, Enes Bajrovic, Siegfried Benkner
ICPADS2
2024 Python to Kubernetes: A Programming and Resource Management Framework for Compute-and Data-intensive Applications
abstract
In this paper, we introduce the Python to Kubernetes (PTK) framework, a high-level Python-based programming framework for deploying Python applications on top of Kubernetes clusters. PTK supports a task-based programming approach with extensions for specifying resource requirements and performance constraints. A major goal of PTK is to provide users with high-level control for deploying compute- and data-intensive applications on different types and configurations of heterogeneous clusters, while ensuring performance and/or cost constraints.
Andrey Nagiyev, Enes Bajrovic, Siegfried Benkner
ICPADS2
2018 Pipeline Patterns on Top of Task-Based Runtimes
Enes Bajrovic, Siegfried Benkner, Jirí Dokulil
PDCAT1
2018 A multi-aspect online tuning framework for HPC applications
Michael Gerndt, Siegfried Benkner, Eduardo César, Carmen B. Navarrete, Enes Bajrovic, Jirí Dokulil, Carla Guillén, Robert Mijakovic, Anna Sikora
Softw. Qual. J.5
2013 HyPHI - Task Based Hybrid Execution C++ Library for the Intel Xeon Phi Coprocessor
abstract
The Intel Threading Building Blocks (TBB) C++ library introduced task parallelism to a wide audience of application developers. The library is easy to use and powerful, but it is limited to shared-memory machines. In this paper we present HyPHI, a novel library for the Intel Xeon Phi coprocessor for building applications which execute using a hybrid parallel model that exploits parallelism across host CPUs and Xeon Phi coprocessors simultaneously. Our library currently provides hybrid for-each and map-reduce. It hides the details of parallelization, work distribution and computation offloading from users while using internally TBB as its foundation. Despite the higher level of abstraction provided by our library we show that for certain types of applications we outperform codes that rely on the built-in offload support currently provided by the Intel compiler. We have performed a set of experiments with the library and created guidelines that help the developers decide in which situations they should use the HyPHI library.
Jirí Dokulil, Enes Bajrovic, Siegfried Benkner, Martin Sandrieser, Beverly Bachmayer
ICPP2
2012 High-Level Support for Pipeline Parallelism on Many-Core Architectures
Siegfried Benkner, Enes Bajrovic, Erich Marth, Martin Sandrieser, Raymond Namyst, Samuel Thibault
Euro-Par2
2011 Using MPI Derived Datatypes in Numerical Libraries
Enes Bajrovic, Jesper Larsson Träff
EuroMPI1
2009 Experimental Study of Multithreading to Improve Memory Hierarchy Performance of Multi-core Processors for Scientific Applications
abstract
In this paper we study performance characteristics and parallelization strategies for recently shipped, powerful multi-core processors - IBM Power6 and Sun T2 Plus - for high-end scientific computing. Central aspect is data locality. First, we investigate the impacts of good and bad data locality by modifying data accesses. Next, we study the impact of multithreading with respect to data locality based on the data-parallel programming approach. The level of parallelism is increased by assigning multiple threads onto one core in order to hide processor stalls caused by bad data locality. We measure the impacts of data locality and multithreading in terms of execution times and bandwidth for synthetic micro-benchmarks, a matrix multiplication kernel, and an application from Bioinformatics. The results indicate that substantial performance improvements can be obtained with minor effort by utilizing multithreading.
Enes Bajrovic, Eduard Mehofer
CISIS1