EDBT 2026 Demo / reviewers in the wild / expert
Diana Arroyo
dblp:24/11531
· DBLP profile ↗
8ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via FaroabstractThis paper tackles the challenge of running multiple ML inference jobs (models) under time-varying workloads, on a constrained on-premises production cluster. Our system Faro takes in latency Service Level Objectives (SLOs) for each job, auto-distills them into utility functions, "sloppifies" these utility functions to make them amenable to mathematical optimization, automatically predicts workload via probabilistic prediction, and dynamically makes implicit cross-job resource allocations, in order to satisfy cluster-wide objectives, e.g., total utility, fairness, and other hybrid variants. A major challenge Faro tackles is that using precise utilities and high-fidelity predictors, can be too slow (and in a sense too precise!) for the fast adaptation we require. Faro's solution is to "sloppify" (relax) its multiple design components to achieve fast adaptation without overly degrading solution quality. Faro is implemented in a stack consisting of Ray Serve running atop a Kubernetes cluster. Trace-driven cluster deployments show that Faro achieves 2.3×-23× lower SLO violations compared to state-of-the-art systems. Beomyeol Jeon, Chen Wang 0039, Diana Arroyo, Alaa Youssef, Indranil Gupta |
EuroSys | 3 |
| 2025 | The Cloud, Like Building and Running Go BinariesabstractWe document the UX challenges of targeting distributed batch-processing code to Cloud resources. The challenges largely stem from the need for every team member to know everything about the mechanics of achieving scale. Thus, both running and writing code become overwhelming tasks. We present Lunchpail, a tool designed to address these challenges.We present two case studies, one of a team running AI/ML workloads and one of the code they wrote to make it happen. We quantify the challenges with two novel UX metrics: multiplicity and divergence. We show that the code base manifests a multitude of concerns, including distribution, packaging, and automation; 64–98% of the team’s code diverges from the main goal of the application. The story is paralleled when running workloads. Users switch between 3–7 types of tasks on a daily basis (high multiplicity). The nature of these tasks differ greatly from the users’ core competencies (high divergence). In particular, we show that all users assume the daily burdens of cluster operators.We demonstrate that four angles of attack, combined, can yield significant reductions in complexity: 1) Adopt a Serverless approach, allowing code to focus on that core "2%". 2) Treat application packaging like building a Golang binary via go build. This binary embeds source, configuration, deployment logic, and a lightweight runtime that channels data to workers with fan-out and queuing. 3) Treat running distributed applications pipelines against Cloud resources like launching said binaries, with simple bash "|" syntax; 4) When possible, avoid multi-tenancy, and instead target Cloud virtual machines directly.We present a large experimental study to quantify the viability of obtaining a dedicated "burst" of cloud resources for every job run. We show VMs can be ready in well under a minute, which is 10-20x faster than scaling a Kubernetes cluster.We embody this approach in Lunchpail. Lunchpail itself is small, weighing in at 12k lines of code (10% of the size of Kubeflow, 2.5% of Ray, 1% of Kueue). We validate Lunchpail against AI/ML code, legacy chip design workloads, and show that it adds little overhead on top of acquiring Cloud VMs. Diana Arroyo, Paul Castro, Thuan Doan, Nick Mitchell, Sara Kokkila Schumacher, Ed Seabolt, Aleksander Slominski, Ansu Varghese, Lionel Villard, Cora Coleman |
ICDCS | 1 |
| 2024 | Optimizing Simultaneous Autoscaling for Serverless Cloud ComputingabstractThis paper explores resource allocation in server-less cloud computing platforms and proposes an optimization approach for autoscaling systems. Serverless computing relieves users from resource management tasks, enabling focus on application functions. However, dynamic resource allocation and function replication based on changing loads remain crucial. Typically, autoscalers in these platforms utilize threshold-based mechanisms to adjust function replicas independently. We model applications as interconnected graphs of functions, where requests probabilistically traverse the graph, triggering associated function execution. Our objective is to develop a control policy that optimally allocates resources on servers, minimizing failed requests and response time in reaction to load changes. Using a fluid approximation model and Separated Continuous Linear Programming (SCLP), we derive an optimal control policy that determines the number of resources per replica and the required number of replicas over time. We evaluate our approach using a simulation framework built with Python and simpy. Comparing against threshold-based autoscaling, our approach demonstrates significant improvements in average response times and failed requests, ranging from 15% to over 300% in most cases. We also explore the impact of system and workload parameters on performance, providing insights into the behavior of our optimization approach under different conditions. Overall, our study contributes to advancing resource allocation strategies, enhancing efficiency and reliability in serverless cloud computing platforms. Harold J. Ship, Evgeny Shindin, Chen Wang 0039, Diana Arroyo, Asser N. Tantawi |
CLOUD | 4 |
| 2024 | Cloud-native Workflow Scheduling using a Hybrid Priority Rule, Dynamic Resource Allocation, and Dynamic Task PartitionabstractAs cloud-native workflow orchestration tools become increasingly important for complex data science workloads, there is a growing need for more efficient scheduling. Existing cloud schedulers rely on basic heuristics and user choice for task partitioning for parallel computing, leading to under-utilization of cluster resources and prolonged job completion times. To address this, we propose a novel workflow scheduling algorithm that leverages workflow characteristics to enhance resource utilization and reduce weighted job completion time. The algorithm combines three sub-algorithms, each reflecting a distinct aspect of the scheduling strategy: 1) Hybrid Maximum Children (MC) -Weighted Shortest Critical Path Time (WSCPT) rule alternates between two heuristics, MC and WSCPT, which prioritize jobs based on workflow structure and critical path, respectively. The choice between these heuristics is dynamically adjusted according to the cluster queue size. 2) Dynamic Resource Allocation (DRA), which dynamically adjusts the number of executors assigned to each workflow, and 3) Dynamic Task Partition (DTP), which autonomously determines the task parallelism level. We tested our algorithm with extensive experiments on various workflow types using Spark-imitated simulation. Our algorithm outperformed other schedulers, including learning-based models, by reducing 21-47% of the combined performance of average job completion time and makespan for unweighted workflows and reducing at least 50% of weighted job completion time for weighted workflows. Jungeun Shin, Diana Arroyo, Asser N. Tantawi, Chen Wang 0039, Alaa Youssef, Rakesh Nagi |
SoCC | 2 |
| 2022 | Cloud-native workflow scheduling using a hybrid priority rule and dynamic task parallelismabstractDemand for efficient cloud-native workflow scheduling is growing as many data science workloads are composed of several tasks with dependencies. As container technology becomes more prevalent in cloud communities, containerized workflow orchestration tools are introduced and become standard for scheduling workflows. However, current schedulers use simple heuristics and rely on the user's choice on priority and parallelism level of tasks without accounting for workflow-specific information. Jungeun Shin, Diana Arroyo, Asser N. Tantawi, Chen Wang 0039, Alaa Youssef, Rakesh Nagi |
SoCC | 2 |
| 2019 | FfDL: A Flexible Multi-tenant Deep Learning PlatformabstractDeep learning (DL) is becoming increasingly popular in several application domains and has made several new application features involving computer vision, speech recognition and synthesis, self-driving automobiles, drug design, etc. feasible and accurate. As a result, large scale "on-premise" and "cloud-hosted" deep learning platforms have become essential infrastructure in many organizations. These systems accept, schedule, manage and execute DL training jobs at scale. K. R. Jayaram, Vinod Muthusamy, Parijat Dube, Vatche Isahagian, Chen Wang 0039, Benjamin Herta, Scott Boag, Diana Arroyo, Asser N. Tantawi, Archit Verma, Falk Pollok, Rania Khalaf |
Middleware | 8 |
| 2014 | COLD: Cloud Optimized Workload DeployerabstractWe demonstrate a prototype system called COLD that we are developing at IBM Research which provides optimized deployment of workload in the cloud. A workload refers to an application, consisting of virtual entities (e.g. VM, volume), to be deployed in a cloud infrastructure, consisting of physical entities (e.g. PM, storage). The resource requirements of the virtual entities, as well as metadata describing relations among virtual entities (e.g. location proximity requirement), are described using a declarative workload definition language, namely using a HOT template. COLD provides a clean separation between the underlying mechanisms offered by a given cloud environment for lifecycle management of virtual resources and policies that influence optimal placement. The current implementation of COLD is tied to Open Stack, though we have plans in the future to make it cloud agnostic. It is designed to be fault-tolerant. In a software defined environment of an Open Stack cloud with various hyper visors, we show the operation of COLD to place various Heat stacks that contain specific policies such as rack-level antic location. Diana Arroyo, Iqbal Mohomed |
MASCOTS | 1 |
| 2012 | Cost-aware replication for dataflowsabstractIn this work we are concerned with the cost associated with replicating intermediate data for dataflows in Cloud environments. This cost is attributed to the extra resources required to create and maintain the additional replicas for a given data set. Existing data-analytic platforms such as Hadoop provide for fault-tolerance guarantee by relying on aggressive replication of intermediate data. We argue that the decision to replicate along with the number of replicas should be a function of the resource usage and utility of the data in order to minimize the cost of reliability. Furthermore, the utility of the data is determined by the structure of the dataflow and the reliability of the system. We propose a replication technique, which takes into account resource usage, system reliability and the characteristic of the dataflow to decide what data to replicate and when to replicate. The replication decision is obtained by solving a constrained integer programming problem given information about the dataflow up to a decision point. In addition, we built a working prototype, CARDIO of our technique which shows through experimental evaluation using a real testbed that finds an optimal solution. Claris Castillo, Asser N. Tantawi, Diana Arroyo, Malgorzata Steinder |
NOMS | 3 |