VLDB 2026 Research / reviewers in the wild / expert
Pawel Zuk
dblp:272/5219
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-4904-7171ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AeroResQ: Edge-accelerated UAV framework for scalable, resilient and collaborative escape route planning in wildfire scenarios
Suman Raj, Radhika Mittal, Rajiv Mayani, Pawel Zuk, Anirban Mandal, Michael Zink, Yogesh L. Simmhan, Ewa Deelman |
Future Gener. Comput. Syst. | 4 |
| 2025 | A Greedy Consensus-Based Approach to Distributed Job Selection: Toward Fully-Decentralized Workload Management SystemabstractCurrent approaches to resilience for highly distributed, heterogeneous, large-scale scientific workflows are limited. Most existing workflow and resource management systems have a single point of failure and resilience strategies are often static, depend on a centralized control, and require considerable design effort from experts. The increasing scale and complexity of workflows coupled with limited resilience capabilities in centralized systems necessitates a fully decentralized, adaptive resource management approach. This paper addresses a very important slice of the overall problem by leveraging the advances in multi-agent systems (MAS). In particular, we explore the suitability of a MAS consisting of globally distributed agents to perform distributed job selection from a dynamic job pool in a truly decentralized, performant, and resilient manner. We present a novel consensus formulation of the distributed job selection problem. By introducing a cost function encapsulating the requirements and constraints of the job and resource loads, we design a novel, greedy consensus algorithm leveraging the Practical Byzantine Fault Tolerance (PBFT)-based consensus method, allowing agents to collectively select jobs in a resilient manner. We compared our algorithms with other state of the art approaches by deploying them in a network testbed infrastructure to emulate distributed job selection. Our evaluation results demonstrated that our greedy consensus algorithm employing the cost-function and PBFT-based consensus method outperforms the ones using the vanilla PBFT-based consensus method - improving scheduling latency by as much as 63.5 % and reducing resource idle time by as much as 63.8 %, with benefits increasing with higher numbers of agents emulated. Komal Thareja, Raghavan Krishnan, Anirban Mandal, Pawel Zuk, Imtiaz Mahmud, Mariam Kiran, Ewa Deelman |
CCGrid | 4 |
| 2024 | DISTRI: Development and Integration of Simulation Tools for Resilient InfrastructureabstractIn contemporary scientific research, data acquisition and analysis platforms have grown increasingly complex, often spanning multiple facilities with diverse internal structures. Efficiently managing the interactions between job scheduling, resource allocation, and networking across these distributed systems requires a robust simulation framework. However, existing simulators fall short in capturing the detailed interactions necessary for comprehensive analysis of large-scale distributed environments. To address this gap, we introduce DISTRI, a versatile framework specifically designed for the development and testing of distributed multi-facility workflows. DISTRI allows for customizable facility configurations and includes built-in support for distributed, resilient scheduling and resource management, alongside detailed network simulation for data communication. Key features of DISTRI encompass inter- and intra-facility resource management, agent-based distributed scheduling, and extensive performance metrics logging for both resource and network management. By providing these essential tools, DISTRI enables thorough analysis and optimization, thereby advancing research in the resilience and efficiency of multi-facility systems. Imtiaz Mahmud, Pawel Zuk, Cong Wang 0014, Mariam Kiran, Kesheng Wu, Komal Thareja, Raghavan Krishnan, Anirban Mandal, Ewa Deelman |
IEEE Big Data | 2 |
| 2024 | sAirflow: Adopting Serverless in a Legacy Workflow Scheduler
Filip Mikina, Pawel Zuk, Krzysztof Rzadca |
Euro-Par (1) | 2 |
| 2024 | Large Language Models for Anomaly Detection in Computational Workflows: From Supervised Fine-Tuning to In-Context LearningabstractAnomaly detection in computational workflows is critical for ensuring system reliability and security. However, traditional rule-based methods struggle to detect novel anomalies. This paper leverages large language models (LLMs) for workflow anomaly detection by exploiting their ability to learn complex data patterns. Two approaches are investigated: (1) supervised fine-tuning (SFT), where pretrained LLMs are fine-tuned on labeled data for sentence classification to identify anomalies, and (2) in-context learning (ICL), where prompts containing task descriptions and examples guide LLMs in few-shot anomaly detection without fine-tuning. The paper evaluates the performance, efficiency, and generalization of SFT models and explores zeroshot and few-shot ICL prompts and interpretability enhancement via chain-of-thought prompting. Experiments across multiple workflow datasets demonstrate the promising potential of LLMs for effective anomaly detection in complex executions. George Papadimitriou 0002, Raghavan Krishnan, Pawel Zuk, Prasanna Balaprakash, Cong Wang 0014, Anirban Mandal, Ewa Deelman |
SC | 4 |
| 2022 | Divide (CPU Load) and Conquer: Semi-Flexible Cloud Resource AllocationabstractCloud resource management is often modeled by two-dimensional bin packing with a set of items that correspond to tasks having fixed CPU and memory requirements. However, applications running in clouds are much more flexible: modern frameworks allow to (horizontally) scale a single application to dozens, even hundreds of instances; and then the load balancer can precisely divide the workload between them. We analyze a model that captures this (semi)-flexibility of cloud resource management. Each cloud application is characterized by its memory footprint and its momentary CPU load. Combining the scheduler and the load balancer, the resource manager decides how many instances of each application will be created and how the CPU load will be balanced between them. In contrast to the divisible load model, each instance of the application requires a certain amount of memory, independent of the number of instances. Thus, the resource manager effectively trades additional memory for more evenly balanced load. We study two objectives: the bin-packing-like minimization of the number of machines used; and the makespan-like minimization of the maximum load among all the machines. We prove NP-hardness of the general problems, but also propose polynomial-time exact algorithms for boundary special cases. Notably, we show that (semi)-flexibility may result in reducing the required number of machines by a tight factor of 2 - ε. For the general case, we propose heuristics that we validate by simulation on instances derived from the Azure trace. Bartlomiej Przybylski, Pawel Zuk, Krzysztof Rzadca |
CCGRID | 2 |
| 2022 | Call Scheduling to Reduce Response Time of a FaaS SystemabstractIn an overloaded FaaS cluster, individual worker nodes strain under lengthening queues of requests. Although the cluster might be eventually horizontally-scaled, adding a new node takes dozens of seconds. As serving applications are tuned for tail serving latencies, and these greatly increase under heavier loads, the current workaround is resource over-provisioning. In fact, even though a service can withstand a steady load of, e.g., 70% CPU utilization, the autoscaler is triggered at, e.g., 30–40% (thus the service uses twice as many nodes as it would be needed). We propose an alternative: a worker-level method handling heavy load without increasing the number of nodes. FaaS executions are not interactive, compared to, e.g., text editors: end-users do not benefit from the CPU allocated to processes often, yet for short periods. Inspired by scheduling methods for High Performance Computing, we take a radical step of replacing the classic OS preemption by (1) queuing requests based on their historical characteristics; (2) once a request is being processed, setting its CPU limit to exactly one core (with no CPU oversubscription). We extend OpenWhisk and measure the efficiency of the proposed solutions using the SeBS benchmark. In a loaded system, our method decreases the average response time by a factor of 4. The improvement is even higher for shorter requests, as the average stretch is decreased by a factor of 18. This leads us to show that we can provide better response-time statistics with 3 machines compared to a 4-machine baseline. Pawel Zuk, Bartlomiej Przybylski, Krzysztof Rzadca |
CLUSTER | 1 |
| 2022 | Using Unused: Non-Invasive Dynamic FaaS Infrastructure with HPC-WhiskabstractModern HPC workload managers and their careful tuning contribute to the high utilization of HPC clusters. However, due to inevitable uncertainty it is impossible to completely avoid node idleness. Although such idle slots are usually too short for any HPC job, they are too long to ignore them. Function-as-a-Service (FaaS) paradigm promisingly fills this gap, and can be a good match, as typical FaaS functions last seconds, not hours. Here we show how to build a FaaS infrastructure on idle nodes in an HPC cluster in such a way that it does not affect the performance of the HPC jobs significantly. We dynamically adapt to a changing set of idle physical machines, by integrating open-source software Slurm and OpenWhisk. We designed and implemented a prototype solution that allowed us to cover up to 90% of the idle time slots on a 50k-core cluster that runs production workloads. Bartlomiej Przybylski, Maciej Pawlik, Pawel Zuk, Bartlomiej Lagosz, Maciej Malawski, Krzysztof Rzadca |
SC | 3 |
| 2022 | Reducing response latency of composite functions-as-a-service through scheduling
Pawel Zuk, Krzysztof Rzadca |
J. Parallel Distributed Comput. | 1 |
| 2021 | Data-driven scheduling in serverless computing to reduce response timeabstractIn Function as a Service (FaaS), a serverless computing variant, customers deploy functions instead of complete virtual machines or Linux containers. It is the cloud provider who maintains the runtime environment for these functions. FaaS products are offered by all major cloud providers (e.g. Amazon Lambda, Google Cloud Functions, Azure Functions); as well as standalone open-source software (e.g. Apache OpenWhisk) with their commercial variants (e.g. Adobe I/O Runtime or IBM Cloud Functions). We take the bottom-up perspective of a single node in a FaaS cluster. We assume that all the execution environments for a set of functions assigned to this node have been already installed. Our goal is to schedule individual invocations of functions, passed by a load balancer, to minimize performance metrics related to response time. Deployed functions are usually executed repeatedly in response to multiple invocations made by end-users. Thus, our scheduling decisions are based on the information gathered locally: the recorded call frequencies and execution times. We propose a number of heuristics, and we also adapt some theoretically-grounded ones like SEPT or SERPT. Our simulations use a recently-published Azure Functions Trace. We show that, compared to the baseline FIFO or round-robin, our data-driven scheduling decisions significantly improve the performance. Bartlomiej Przybylski, Pawel Zuk, Krzysztof Rzadca |
CCGRID | 2 |
| 2020 | Scheduling Methods to Reduce Response Latency of Function as a ServiceabstractFunction as a Service (FaaS) permits cloud customers to deploy to cloud individual functions, in contrast to complete virtual machines or Linux containers. All major cloud providers offer FaaS products (Amazon Lambda, Google Cloud Functions, Azure Serverless); there are also popular open-source implementations (Apache OpenWhisk) with commercial offerings (Adobe I/O Runtime, IBM Cloud Functions). A new feature of FaaS is function composition: a function may (sequentially) call another function, which, in turn, may call yet another function - forming a chain of invocations. From the perspective of the infrastructure, a composed FaaS is less opaque than a virtual machine or a container. We show that this additional information enables the infrastructure to reduce the response latency. In particular, knowing the sequence of future invocations, the infrastructure can schedule these invocations along with environment preparation. We model resource management in FaaS as a scheduling problem combining (1) sequencing of invocations, (2) deploying execution environments on machines, and (3) allocating invocations to deployed environments. For each aspect, we propose heuristics. We explore their performance by simulation on a range of synthetic workloads. Our results show that if the setup times are long compared to invocation times, algorithms that use information about the composition of functions consistently outperform greedy, myopic algorithms, leading to significant decrease in response latency. Pawel Zuk, Krzysztof Rzadca |
SBAC-PAD | 1 |