EDBT 2026 Demo / reviewers in the wild / expert
Alaa Youssef
dblp:69/4290
· DBLP profile ↗
21ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 5 since 2021Computer networks · 6 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via FaroabstractThis paper tackles the challenge of running multiple ML inference jobs (models) under time-varying workloads, on a constrained on-premises production cluster. Our system Faro takes in latency Service Level Objectives (SLOs) for each job, auto-distills them into utility functions, "sloppifies" these utility functions to make them amenable to mathematical optimization, automatically predicts workload via probabilistic prediction, and dynamically makes implicit cross-job resource allocations, in order to satisfy cluster-wide objectives, e.g., total utility, fairness, and other hybrid variants. A major challenge Faro tackles is that using precise utilities and high-fidelity predictors, can be too slow (and in a sense too precise!) for the fast adaptation we require. Faro's solution is to "sloppify" (relax) its multiple design components to achieve fast adaptation without overly degrading solution quality. Faro is implemented in a stack consisting of Ray Serve running atop a Kubernetes cluster. Trace-driven cluster deployments show that Faro achieves 2.3×-23× lower SLO violations compared to state-of-the-art systems. Beomyeol Jeon, Chen Wang 0039, Diana Arroyo, Alaa Youssef, Indranil Gupta |
EuroSys | 4 |
| 2025 | The bottlenecks of AI: challenges for embedded and real-time research in a data-centric ageabstractAbstract Recent advances in AI culminate a shift in science and engineering away from strong reliance on algorithmic and symbolic knowledge towards new data-driven approaches. How does the emerging intelligent data-centric world impact research on real-time and embedded computing? We argue for two effects: (1) new challenges in embedded system contexts, and (2) new opportunities for community expansion beyond the embedded domain. First, on the embedded system side, the shifting nature of computing towards data-centricity affects the types of bottlenecks that arise. At training time, the bottlenecks are generally data-related. Embedded computing relies on scarce sensor data modalities, unlike those commonly addressed in mainstream AI, necessitating solutions for efficient learning from scarce sensor data. At inference time, the bottlenecks are resource-related, calling for improved resource economy and novel scheduling policies. Further ahead, the convergence of AI around large language models (LLMs) introduces additional model-related challenges in embedded contexts. Second, on the domain expansion side, we argue that community expertise in handling resource bottlenecks is becoming increasingly relevant to a new domain: the cloud environment, driven by AI needs. The paper discusses the novel research directions that arise in the data-centric world of AI, covering data-, resource-, and model-related challenges in embedded systems as well as new opportunities in the cloud domain. Tarek F. Abdelzaher, Yigong Hu, Denizhan Kara, Tomoyoshi Kimura, Ashitabh Misra, Vishakha Ramani, Olivier Tardieu, Tianshi Wang 0002, Maggie B. Wigness, Alaa Youssef |
Real Time Syst. | 10 |
| 2024 | Cloud-native Workflow Scheduling using a Hybrid Priority Rule, Dynamic Resource Allocation, and Dynamic Task PartitionabstractAs cloud-native workflow orchestration tools become increasingly important for complex data science workloads, there is a growing need for more efficient scheduling. Existing cloud schedulers rely on basic heuristics and user choice for task partitioning for parallel computing, leading to under-utilization of cluster resources and prolonged job completion times. To address this, we propose a novel workflow scheduling algorithm that leverages workflow characteristics to enhance resource utilization and reduce weighted job completion time. The algorithm combines three sub-algorithms, each reflecting a distinct aspect of the scheduling strategy: 1) Hybrid Maximum Children (MC) -Weighted Shortest Critical Path Time (WSCPT) rule alternates between two heuristics, MC and WSCPT, which prioritize jobs based on workflow structure and critical path, respectively. The choice between these heuristics is dynamically adjusted according to the cluster queue size. 2) Dynamic Resource Allocation (DRA), which dynamically adjusts the number of executors assigned to each workflow, and 3) Dynamic Task Partition (DTP), which autonomously determines the task parallelism level. We tested our algorithm with extensive experiments on various workflow types using Spark-imitated simulation. Our algorithm outperformed other schedulers, including learning-based models, by reducing 21-47% of the combined performance of average job completion time and makespan for unweighted workflows and reducing at least 50% of weighted job completion time for weighted workflows. Jungeun Shin, Diana Arroyo, Asser N. Tantawi, Chen Wang 0039, Alaa Youssef, Rakesh Nagi |
SoCC | 5 |
| 2024 | Dexter: A Performance-Cost Efficient Resource Allocation Manager for Serverless Data AnalyticsabstractLeveraging serverless platforms for the efficient execution of distributed data analytics frameworks, such as Apache Spark [3], has gained substantial interest since early 2022. The elasticity, free-of-management, and on-demand scalability of serverless have motivated the effort in deploying distributed data analytics applications to serverless platforms. However, effectively auto-scaling resources for such complex workloads so that we can fully benefit from the resource elasticity of serverless remains challenging. Mis-configuration can result in severe performance and cost issues arising from resource under- and over-provisioning. Anna Maria Nestorov, Diego Marron, Alberto Gutierrez-Torre, Chen Wang 0039, Claudia Misale, Alaa Youssef, David Carrera 0001, Josep Lluís Berral |
Middleware | 6 |
| 2023 | A Carbon-aware Workload Dispatcher in Cloud Computing SystemsabstractThe amount of carbon emission associated with the computational energy consumption in data centers depends, in a significant way, on the schedule of the workloads. Due to the inconsistent availability of renewable energy over time, in addition to the existence of various sources of power in grid regions, the carbon intensity of data centers changes over time and location. Thus, the placement and scheduling of flexible workloads, based on the carbon intensity of power sources in data centers, can remarkably decrease the carbon emission. In this paper, we address the problem of placement and scheduling of workloads over geographically distributed data centers. We propose two algorithms that take the variability of carbon intensity of the power sources of the data centers, as well as their computational resource availability, into account when deciding about the placement and scheduling of the workloads. The first is a randomized rounding approximation algorithm that provides solutions that are guaranteed to be within a given distance from the optimal solution. The second is a sample-based algorithm that improves the solutions obtained by the randomized rounding approximation algorithm. The experimental results show that the proposed algorithms can solve the problem efficiently. Tayebeh Bahreini, Asser N. Tantawi, Alaa Youssef |
CLOUD | 3 |
| 2023 | AWARE: Automate Workload Autoscaling with Reinforcement Learning in Production Cloud Systems
Haoran Qiu, Weichao Mao, Chen Wang 0039, Hubertus Franke, Alaa Youssef, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer |
USENIX ATC | 5 |
| 2022 | An Approximation Algorithm for Minimizing the Cloud Carbon Footprint through Workload SchedulingabstractIn this paper, we address the problem of workload scheduling in data centers, while considering the greenness of the power sources. We prove that finding a feasible solution for the problem is NP-hard. Therefore, we develop an LP-based approximation algorithm to solve the problem in polynomial time. The proposed algorithm provides strong approximation bounds on the constraints and the objective of the problem. We conduct an extensive experimental analysis to evaluate the performance of the proposed algorithm using real world data. Tayebeh Bahreini, Asser N. Tantawi, Alaa Youssef |
CLOUD | 3 |
| 2022 | Cloud-native workflow scheduling using a hybrid priority rule and dynamic task parallelismabstractDemand for efficient cloud-native workflow scheduling is growing as many data science workloads are composed of several tasks with dependencies. As container technology becomes more prevalent in cloud communities, containerized workflow orchestration tools are introduced and become standard for scheduling workflows. However, current schedulers use simple heuristics and rely on the user's choice on priority and parallelism level of tasks without accounting for workflow-specific information. Jungeun Shin, Diana Arroyo, Asser N. Tantawi, Chen Wang 0039, Alaa Youssef, Rakesh Nagi |
SoCC | 5 |
| 2021 | Theta-Scan: Leveraging Behavior-Driven Forecasting for Vertical Auto-Scaling in Container CloudabstractDetection of behavior patterns on resource usage in containerized Cloud applications is necessary for proper resource provisioning. Applications can use CPU/Memory with repetitive patterns, following a trend over time independently. By identifying such patterns, resource forecasting models can be fit better, reducing over/under-provisioning via fewer resizing operations. Here we present ThetaScan, a time-series analysis method for vertical auto-scaling of containers in the Cloud, based on the detection of stationarity/trending and periodicity on resource consumption. Our method leverages the Theta Forecaster algorithm with deseasonalization that, in our provisioning scenario, only requires the estimated periodicity for resource consumption as principal hyper-parameter. Commonly used behavior detection methods require manual hyper-parameter tuning, making them infeasible for automation. Besides, it can be used at multi-scales (minute/hour/day), detecting hourly and daily patterns to improve resource usage prediction. Experiments show that we can detect behaviors in resource consumption that common methods miss, without requiring extensive manual tuning. We can reduce the resizing triggers compared to fixed-size scheduling around ~ 10% – 15%, reduce over-provisioning of CPU and Memory through periodic-based provisioning. Also a ~ 60% on multiscale resource forecasting for traces showing periodicity at different levels in respect to single-scale. Josep Lluís Berral, David Buchaca Prats, Claudia Herron, Chen Wang 0039, Alaa Youssef |
CLOUD | 5 |
| 2021 | Performance Evaluation of Data-Centric Workloads in Serverless EnvironmentsabstractServerless computing is a cloud-based execution paradigm that allows provisioning resources on-demand, freeing developers from infrastructure management and operational concerns. It typically involves deploying workloads as stateless functions that take no resources when not in use, and is meant to scale transparently. To make serverless effective, providers impose limits on a per-function level, such as maximum duration, fixed amount of memory, and no persistent local storage. These constraints make it challenging for data-intensive workloads to take advantage of serverless because they lead to sharing significant amounts of data through remote storage. In this paper, we build a performance model for serverless workloads that considers how data is shared between functions, including the amount of data and the underlying technology that is being used. The model's accuracy is assessed by running a real workload in a cluster using Knative, a state-of-the-art serverless environment, showing a relative error of 5.52%. With the proposed model, we evaluate the performance of data-intensive workloads in serverless, analyzing parallelism, scalability, resource requirements, and scheduling policies. We also explore possible solutions for the data-sharing problem, like using local memory and storage. Our results show that the performance of data-intensive workloads in serverless can be up to 4.32= faster depending on how these are deployed. Anna Maria Nestorov, Jorda Polo, Claudia Misale, David Carrera 0001, Alaa Youssef |
CLOUD | 5 |
| 2020 | Proactive Container Auto-scaling for Cloud Native Machine Learning ServicesabstractUnderstanding the resource usage behaviors of the ever-increasing machine learning workloads are critical to cloud providers offering Machine Learning (ML) services. Capable of auto-scaling resources for customer workloads can significantly improve resource utilization, thus greatly reducing the cost. Here we leverage the AI4DL framework [1] to characterize workload and discover resource consumption phases. We advance the existing technology to an incremental phase discovery method that applies to more general types of ML workload for both training and inference. We use a time-window MultiLayer Perceptron (MLP) to predict phases in containers with different types of workload. Then, we propose a predictive vertical auto-scaling policy to resize the container dynamically according to phase predictions. We evaluate our predictive auto-scaling policies on 561 long-running containers with multiple types of ML workloads. The predictive policy can reduce up to 38% of allocated CPU compared to the default resource provisioning policies by developers. By comparing our predictive policies with commonly used reactive auto-scaling policies, we find that they can accurately predict sudden phase transitions (with an F1-score of 0.92) and significantly reduce the number of out-of-memory errors (350 vs. 20). Besides, we show that the predictive auto-scaling policy maintains the number of resizing operations close to the best reactive policies. David Buchaca Prats, Josep Lluís Berral, Chen Wang 0039, Alaa Youssef |
CLOUD | 4 |
| 2018 | Resource Profile Advisor for Containers in Cognitive PlatformabstractContainers have transformed the cluster management into an application oriented endeavor, thus being widely used as the deployment units (i.e., micro-services) of large scale cloud services. As opposed to VMs, containers allow for resource provisioning with fine granularity and their resource usage directly reflects the micro-service behaviors. Container management systems like Kubernetes and Mesos provision resources to containers according to the capacity requested by the developers. Resource usages estimated by the developers are grossly inaccurate. They tend to be risk-averse and over provision resources, as under-provisioning would cause poor runtime performance or failures. Mehmet Fatih Aktas, Chen Wang 0039, Alaa Youssef, Malgorzata Steinder |
SoCC | 3 |
| 2017 | Reducing tail latencies in micro-batch streaming workloadsabstractSpark Streaming discretizes streams of data into micro-batches, each of which is further sub-divided into tasks and processed in parallel to improve job throughput. Previous work [2, 3] has lowered end-to-end latency in Spark Streaming. However, two causes of high tail latencies remain unaddressed: 1) data is not load-balanced across tasks, and 2) straggler tasks can increase end-to-end latency by 8 times more than the median task on a production cluster [1]. We propose a feedback-control mechanism that allows frameworks to adaptively load-balance workloads across tasks according to their processing speeds. The task runtimes are thus equalized, lowering end-to-end tail latency. Further, this reduces load on machines that have transient resource bottlenecks, thus resolving the bottlenecks and preventing them from having an enduring impact on task runtimes. Faria Kalim, Asser N. Tantawi, Stefania Costache 0002, Alaa Youssef |
SoCC | 4 |
| 2005 | Performance management for cluster-based web servicesabstractWe present an architecture and prototype implementation of a performance management system for cluster-based web services. The system supports multiple classes of web services traffic and allocates server resources dynamically so to maximize the expected value of a given cluster utility function in the face of fluctuating loads. The cluster utility is a function of the performance delivered to the various classes, and this leads to differentiated service. In this paper, we will use the average response time as the performance metric. The management system is transparent: it requires no changes in the client code, the server code, or the network interface between them. The system performs three performance management tasks: resource allocation, load balancing, and server overload protection. We use two nested levels of management. The inner level centers on queuing and scheduling of request messages. The outer level is a feedback control loop that periodically adjusts the scheduling weights and server allocations of the inner level. The feedback controller is based on an approximate first-principles model of the system, with parameters derived from continuous monitoring. We focus on SOAP-based web services. We report experimental results that show the dynamic behavior of the system. Giovanni Pacifici, Mike Spreitzer, Asser N. Tantawi, Alaa Youssef |
IEEE J. Sel. Areas Commun. | 4 |
| 2003 | Performance Management for Cluster Based Web Services
Ronald M. Levy, Jay Nagarajarao, Giovanni Pacifici, Mike Spreitzer, Asser N. Tantawi, Alaa Youssef |
Integrated Network Management | 6 |
| 1999 | General and Scalable State Feedback for Multimedia SystemsabstractObtaining feedback information regarding the state of receivers in a multicast session is a fundamental problem that often arises in collaborative multimedia systems. In this paper we present a generalized abstraction of the state feedback problem. Then, we present a feedback protocol that addresses some of the special cases that commonly arise. The presented feedback protocol is suitable for application in best-effort unreliable networks such as the Internet. It allows for obtaining the desired state feedback about a group of receivers, where each receiver may be in one of a set of finite states. The efficiency of the proposed protocol in eliminating the reply implosion problem is illustrated by simulation experiments. Alaa Youssef, Hussein M. Abdel-Wahab, Kurt Maly, Mohamed G. Gouda |
ISCC | 1 |
| 1998 | The Software Architecture of a Distributed Quality of Session Control LayerabstractCollaborative multimedia systems demand overall session quality control beyond the level of quality of service (QoS) pertaining to individual streams in isolation of others. To this end, the authors have recently introduced the concept of quality of session (QoSess) control. At every instant in time, the quality of the session depends on the actual QoS offered by the system to each of the application streams, as well as on the relative priorities of these streams according to the application semantics. The authors present a framework for achieving QoSess control, and describe the architecture of a distributed QoSess control layer. In addition, they describe a new inter-stream bandwidth adaptation mechanism, which is used by the QoSess control layer to dynamically control the bandwidth shares of the streams belonging to a session. Alaa Youssef, Hussein M. Abdel-Wahab, Kurt Maly |
HPDC | 1 |
| 1998 | Controlling Quality of Session in Adaptive Multimedia Multicast SystemsabstractControlling the quality of collaborative multimedia sessions, that deploy multiple media streams, is a challenging problem. In this paper we present a framework for achieving quality of session (QoSess) control focusing on two main components of the QoSess control layer. The first component is a scalable and robust feedback mechanism which allows for determining the worst case state among a group of receivers of a stream. This mechanism is used for controlling the transmission rate of multimedia sources in the cases of layered and single-rate streams. The second component is the inter-stream bandwidth adaptation mechanism that dynamically controls the bandwidth shares of the streams belonging to a session. We compare the performance of several adaptation algorithms. Additionally, in order to ensure stability and responsiveness in the inter-stream adaptation process, several measures are taken, including devising a domain rate control protocol. The performance of our mechanisms is analyzed and their advantages are demonstrated by simulation and experimental results. Alaa Youssef, Hussein M. Abdel-Wahab, Kurt Maly |
ICNP | 1 |
| 1998 | Distributed management of exclusive resources in collaborative multimedia systemsabstractCollaborative multimedia systems encompass many Internet applications such as desktop conferencing and interactive distance learning. These applications often contain resources, such as audio, video and shared applications, that must be accessed exclusively by one participant at a time. We present a distributed algorithm that manages the access to these exclusive resources. The algorithm is based on the assumption that the transport layer provides reliable multicasting. Resources are classified into two main classes: primitive and composite. Composite resources consist of a set of two or more primitive resources. A token is associated with each resource unit, and a participant must obtain the resource's token before using the resource. To use a resource, certain permissions may be needed from certain entities such as the session coordinator, the current resource holder and, in some cases, the resource itself. The algorithm guarantees that at any given time, the resource is held by exactly one participant and the token of any resource will never be lost under all possible failure conditions. Hussein M. Abdel-Wahab, Alaa Youssef, Kurt Maly |
ISCC | 2 |
| 1997 | Inter-stream adaptation for collaborative multimedia applicationsabstractIn new collaborative multimedia applications, there is a need for overall control, beyond the level of quality of service (QoS) as pertaining to individual streams in isolation of others. At every instant in time, the quality of the session, as perceived by the end user, depends on the priorities of the on-going streams, according to the application semantics, as well as on the actual QoS offered by the system to each of these streams. We introduce the concept of "Quality of Session" control. This is achieved by employing a monitoring mechanism for measuring the perceived QoS of each stream. In addition, in order to react to existing or potential bottlenecks in the network or end-systems, or skewness in the synchronization of views, an inter-stream adaptation mechanism is applied. Alaa Youssef, Hussein M. Abdel-Wahab, Kurt Maly, Mohamed G. Gouda |
ISCC | 1 |
| 1996 | Interactive remote instruction: initial experiences
Kurt Maly, J. Christian Wild, C. Michael Overstreet, Hussein M. Abdel-Wahab, Ajay Gupta 0003, Alaa Youssef, Emilia Stoica, R. Talla, A. Prabhu |
ITiCSE | 6 |