EDBT 2026 Demo / reviewers in the wild / expert
Abel Souza
dblp:215/2394
· DBLP profile ↗
15ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0001-6952-1195ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Untangling the Carbon-Cost Tradeoffs and Stampede Effect Challenges in Cloud ComputingabstractAs society explores new computing applications, data centers have seen a sharp rise in capacity and energy use, driving up data centers’ carbon footprint and raising concerns about their sustainability. In response, efforts to reduce emissions now complement the traditional goals of cutting energy costs and boosting performance. Given the spatiotemporal variability of the grid carbon intensity, researchers are increasingly adopting workload shifting strategies to minimize the operational carbon footprint of cloud workloads. However, shifting from cost- and performance-centric operations to carbon-aware operations introduces performance penalties and cost overheads. Moreover, large-scale spatial and temporal shifts can cause stampede effects, with synchronized migration to the same low-carbon regions or times, thereby straining resources and creating imbalances. In this paper, we quantify the carbon-cost-performance trade-offs of carbon-aware scheduling and evaluate the risk of such stampede effects. We provide real-world examples illustrating these tradeoffs for application and cloud providers. Lastly, we offer practical insights to help mitigate these challenges and guide sustainable scheduling decisions. Walid A. Hanafy, Thanathorn Sukprasert, Abel Souza, David Irwin 0001, Prashant J. Shenoy |
IEEE Trans. Computers | 3 |
| 2025 | CarbonEdge: Leveraging Mesoscale Spatial Carbon-Intensity Variations for Low Carbon Edge ComputingabstractThe proliferation of latency-critical and compute-intensive edge applications is driving increases in computing demand and carbon emissions at the edge. To better understand carbon emissions at the edge, we analyze granular carbon intensity traces at intermediate "mesoscales," such as within a single US state or among neighboring countries in Europe, and observe significant variations in carbon intensity at these spatial scales. Importantly, our analysis shows that carbon intensity variations, which are known to occur at large continental scales (e.g., cloud regions), also occur at much finer spatial scales, making it feasible to exploit geographic workload shifting in the edge computing context. Motivated by these findings, we propose CarbonEdge, a carbon-aware framework for edge computing that optimizes the placement of edge workloads across mesoscale edge data centers to reduce carbon emissions while meeting latency SLOs. We implement CarbonEdge and evaluate it on a real edge computing testbed and through large-scale simulations for multiple edge workloads and settings. Our experimental results on a real testbed demonstrate that CarbonEdge can reduce emissions by up to 78.7% for a regional edge deployment in central Europe. Moreover, our CDN-scale experiments show potential savings of 49.5% and 67.8% in the US and Europe, respectively, while limiting the one-way latency increase to less than 5.5 ms. Walid A. Hanafy, Abel Souza, Jan Harkes, David Irwin 0001, Mahadev Satyanarayanan, Prashant J. Shenoy |
HPDC | 3 |
| 2025 | A Decentralized Microservice Scheduling Approach Using Service Mesh in Cloud-Edge SystemsabstractAs microservice-based systems scale across the cloud-edge continuum, traditional centralized scheduling mechanisms increasingly struggle with latency, coordination overhead, and fault tolerance. This paper presents a new architectural direction: leveraging service mesh sidecar proxies as decentralized, in-situ schedulers to enable scalable, low-latency coordination in large-scale, cloud-native environments. We propose embedding lightweight, autonomous scheduling logic into each sidecar, allowing scheduling decisions to be made locally without centralized control. This approach leverages the growing maturity of service mesh infrastructures, which support programmable distributed traffic management. We describe the design of such an architecture and present initial results demonstrating its scalability potential in terms of response time and latency under varying request rates. Rather than delivering a finalized scheduling algorithm, this paper presents a system-level architectural direction and preliminary evidence to support its scalability potential. Yangyang Wen, Paul Townend, Per-Olov Östberg, Abel Souza, Clément Courageux-Sudan |
JCC | 4 |
| 2024 | Going Green for Less Green: Optimizing the Cost of Reducing Cloud Carbon EmissionsabstractThe continued exponential growth of cloud datacenter capacity has increased awareness of the carbon emissions when executing large compute-intensive workloads. To reduce carbon emissions, cloud users often temporally shift their batch workloads to periods with low carbon intensity. While such time shifting can increase job completion times due to their delayed execution, the cost savings from cloud purchase options, such as reserved instances, also decrease when users operate in a carbon-aware manner. This happens because carbon-aware adjustments change the demand pattern by periodically leaving resources idle, which creates a trade-off between carbon emissions and cost. In this paper, we present GAIA, a carbon-aware scheduler that enables users to address the three-way trade-off between carbon, performance, and cost in cloud-based batch schedulers. Our results quantify the carbon-performance-cost trade-off in cloud platforms and show that compared to existing carbon-aware scheduling policies, our proposed policies can double the amount of carbon savings per percentage increase in cost, while decreasing the performance overhead by 26%. Walid A. Hanafy, Qianlin Liang, Noman Bashir, Abel Souza, David Irwin 0001, Prashant J. Shenoy |
ASPLOS (3) | 4 |
| 2024 | SLO-Power: SLO and Power-aware Elastic Scaling for Web ServicesabstractManaging the performance of online web services in cloud data centers while optimizing resource allocation and power consumption is a multifaceted challenge. Often, resource and power management techniques, such as elastic scaling and power capping, are handled independently, leading to conflicts and sub-optimal power-performance trade-offs. To tackle this issue, we introduce SLO-Power, a system that coordinates the resource and power scaling techniques to achieve power savings while adhering to service level objectives (SLOs), such as tail latency constraints. Our approach employs a combination of analytic queuing models and feedback-driven techniques to jointly allocate resources and power to cloud applications in an SLO and power-aware manner. We implement a prototype of our system and evaluate it using realistic workloads to demonstrate its ability to harmonize elastic and power scaling, enabling enhanced resource utilization and reduced power consumption while ensuring the application performance. Our findings indicate that SLO-Power achieves exceptional power and resource efficiency, approaching near-optimal power-efficiency levels at 90%, all while preventing SLO violations. Furthermore, compared to state-of-the-art solutions, SLO-Power demonstrates lower P95 latency, accompanied by a 12% reduction in resource usage. Mehmet Savasci, Abel Souza, David Irwin 0001, Ahmed Ali-Eldin, Prashant J. Shenoy |
CCGrid | 2 |
| 2024 | TailClipper: Reducing Tail Response Time of Distributed Services Through System-Wide SchedulingabstractReducing tail latency has become a crucial issue for optimizing the performance of online cloud services and distributed applications. In distributed applications, there are many causes of high end-to-end tail latency, including operating system delays, request re-ordering due to fan-out/fanin, and network congestion. Although recent research has focused on reducing tail latency for individual application components, such as by replicating requests and scheduling, in this paper, we argue for a holistic approach for reducing the end-to-end tail latency across application components. We propose TailClipper, a distributed scheduler that tags each arriving request with an arrival timestamp, and propagates it across the microservices' call chain. TailClipper then uses arrival timestamps to implement an oldest request first scheduler that combines global first-come first serve with a limited form of processor sharing to reduce end-to-end tail latency. In doing so, TailClipper can counter the performance degradation caused by request reordering in multi-tiered and microservices-based applications. We implement TailClipper as a userspace Linux scheduler and evaluate it using cloud workload traces and a real-world microservices application. Compared to state-of-the-art schedulers, our experiments reveal that TailClipper improves the 99th percentile response time by up to 81%, while also improving the mean response time and the system throughput by up to 54% and 29% respectively under high loads. Nathan Ng 0002, Abel Souza, Ahmed Ali-Eldin, David Irwin 0001, Don Towsley, Prashant J. Shenoy |
SoCC | 2 |
| 2024 | On the Limitations of Carbon-Aware Temporal and Spatial Workload Shifting in the CloudabstractCloud platforms have been focusing on reducing their carbon emissions by shifting workloads across time and locations to when and where low-carbon energy is available. Despite the prominence of this idea, prior work has only quantified the potential of spatiotemporal workload shifting in narrow settings, i.e., for specific workloads in select regions. In particular, there has been limited work on quantifying an upper bound on the ideal and practical benefits of carbon-aware spatiotemporal workload shifting for a wide range of cloud workloads. To address the problem, we conduct a detailed data-driven analysis to understand the benefits and limitations of carbon-aware spatiotemporal scheduling for cloud workloads. We utilize carbon intensity data from 123 regions, encompassing most major cloud sites, to analyze two broad classes of workloads---batch and interactive---and their various characteristics, e.g., job duration, deadlines, and SLOs. Our findings show that while spatiotemporal workload shifting can reduce workloads' carbon emissions, the practical upper bounds of these carbon reductions are currently limited and far from ideal. We also show that simple scheduling policies often yield most of these reductions, with more sophisticated techniques yielding little additional benefit. Notably, we also find that the benefit of carbon-aware workload scheduling relative to carbon-agnostic scheduling will decrease as the energy supply becomes "greener." Thanathorn Sukprasert, Abel Souza, Noman Bashir, David Irwin 0001, Prashant J. Shenoy |
EuroSys | 2 |
| 2024 | Acies-OS: A Content-Centric Platform for Edge AI Twinning and OrchestrationabstractThis paper describes Acies-OS, a content-centric platform for edge AI twinning and orchestration that allows easy deployment, re-configuration, and control of edge AI services, augmented by a digital twin. The work is motivated by the proliferation of edge AI in a plethora of IoT applications, ranging from home automation to military defense, and the emergence of digital twins that go beyond monitoring and emulation into configuration management and optimization of edge capabilities. While past work focused on either the edge capabilities themselves or the digital twin, this work focuses on their seamless interactions, offering abstractions that enable the digital twin to manage and optimize an increasingly diverse edge AI system. Acies-OS features a structured namespace, a thin client library with flexible pub/sub-based communication, health monitoring support, and a control plane for twin-based value-added analysis and optimization. To illustrate the use of Acies-OS, we implemented a multi-node multi-modality vehicle classification application and used Acies-OS to interface it to a digital twin. We then deployed the system in the field to showcase run-time twin-based optimizations of inference latency, classification accuracy, and robustness to failures in noisy and challenging conditions. Jinyang Li 0004, Yizhuo Chen, Tomoyoshi Kimura, Tianshi Wang 0002, Ruijie Wang 0004, Denizhan Kara, Yigong Hu, Walid A. Hanafy, Abel Souza, Prashant J. Shenoy, Maggie B. Wigness, Joydeep Bhattacharyya, Jae Kim, Guijun Wang, Greg Kimberly, Josh D. Eckhardt, Denis Osipychev, Tarek F. Abdelzaher |
ICCCN | 10 |
| 2023 | Ecovisor: A Virtual Energy System for Carbon-Efficient ApplicationsabstractCloud platforms' rapid growth is raising significant concerns about their carbon emissions. To reduce carbon emissions, future cloud platforms will need to increase their reliance on renewable energy sources, such as solar and wind, which have zero emissions but are highly unreliable. Unfortunately, today's energy systems effectively mask this unreliability in hardware, which prevents applications from optimizing their carbon-efficiency, or work done per kilogram of carbon emitted. To address the problem, we design an "ecovisor", which virtualizes the energy system and exposes software-defined control of it to applications. An ecovisor enables each application to handle clean energy's unreliability in software based on its own specific requirements. We implement a small-scale ecovisor prototype that virtualizes a physical energy system to enable software-based application-level i) visibility into variable grid carbon-intensity and local renewable generation and ii) control of server power usage and battery charging and discharging. We evaluate the ecovisor approach by showing how multiple applications can concurrently exercise their virtual energy system in different ways to better optimize carbon-efficiency based on their specific requirements compared to general system-wide policies. Abel Souza, Noman Bashir, Jorge Murillo, Walid A. Hanafy, Qianlin Liang, David Irwin 0001, Prashant J. Shenoy |
ASPLOS (2) | 1 |
| 2021 | Enabling Sustainable Clouds: The Case for Virtualizing the Energy SystemabstractCloud platforms' growing energy demand and carbon emissions are raising concern about their environmental sustainability. The current approach to enabling sustainable clouds focuses on improving energy-efficiency and purchasing carbon offsets. These approaches have limits: many cloud data centers already operate near peak efficiency, and carbon offsets cannot scale to near zero carbon where there is little carbon left to offset. Instead, enabling sustainable clouds will require applications to adapt to when and where unreliable low-carbon energy is available. Applications cannot do this today because their energy use and carbon emissions are not visible to them, as the energy system provides the rigid abstraction of a continuous, reliable energy supply. This vision paper instead advocates for a "carbon first" approach to cloud design that elevates carbon-efficiency to a firs--class metric. To do so, we argue that cloud platforms should virtualize the energy system by exposing visibility into, and software-defined control of, it to applications, enabling them to define their own abstractions for managing energy and carbon emissions based on their own requirements. Noman Bashir, Tian Guo 0001, Mohammad Hajiesmaili, David Irwin 0001, Prashant J. Shenoy, Ramesh K. Sitaraman, Abel Souza, Adam Wierman |
SoCC | 7 |
| 2021 | AdCom: Adaptive Combiner for Streaming AggregationsabstractContinuous applications such as device monitoring and anomaly detection often require real-time aggregated statistics over unbounded data streams. While existing stream processing systems such as Flink, Spark, and Storm support processing of streaming aggregations, their optimizations are limited with respect to the dynamic nature of the data, and therefore are suboptimal when the workload changes and/or when there is data skew. In this paper we present AdCom, which is an adaptive combiner for stream processing engines. The use of AdCom in aggregation queries enables pre-aggregating tuples upstream (i.e., before data shuffling) followed by global aggregation downstream. In contrast to existing approaches, AdCom can automatically adjust the number of tuples to pre-aggregate depending on the data rate and available network. Our experimental study using real-world streaming workloads shows that using AdCom leads to 2.5-9× higher sustainable throughput without compromising latency. Felipe Oliveira Gutierrez, Kaustubh Beedkar, Abel Souza, Volker Markl |
EDBT | 3 |
| 2021 | A HPC Co-scheduler with Reinforcement Learning
Abel Souza, Kristiaan Pelckmans, Johan Tordsson |
JSSPP | 1 |
| 2020 | ASA - The Adaptive Scheduling ArchitectureabstractIn High Performance Computing (HPC), resources are controlled by batch systems and may not be available due to long queue waiting times, negatively impacting application deadlines. This is noticeable in low latency scientific workflows where resource planning and timely allocation are key for efficient processing. On the one hand, peak allocations guarantee the fastest possible workflows execution time, at the cost of extended queue waiting times and costly resource usage. On the other hand, dynamic allocations following specific workflow stage requirements optimizes resource usage, though it increases the total workflow makespan. To enable new scheduling strategies and features in workflows, we propose ASA: the Adaptive Scheduling Architecture, a novel scheduling method to reduce perceived queue waiting times as well as to optimize workflows resource usage. Reinforcement learning is used to estimate queue waiting times, and based on these estimates ASA pro-actively submit resource change requests, minimizing total workflow inter-stage waiting times, idle resources, and makespan. Experiments with three scientific workflows at two HPC centers show that ASA combines the best of the two aforementioned approaches, with average queue waiting time and makespan reductions of up to 10% and 2% respectively, with up to 100% prediction accuracy, while obtaining near optimal resource utilization. Abel Souza, Kristiaan Pelckmans, Devarshi Ghoshal, Lavanya Ramakrishnan, Johan Tordsson |
HPDC | 1 |
| 2019 | Hybrid Resource Management for HPC and Data Intensive WorkloadsabstractHigh Performance Computing (HPC) and Data Intensive (DI) workloads have been executed on separate clusters using different tools for resource and application management. With increasing convergence, where modern applications are composed of both types of jobs in complex workflows, this separation becomes a growing overhead and the need for a common platform increases. Executing both workload classes on the same clusters not only enables hybrid workflows, but can also increase system efficiency, as available hardware often is not fully utilized by applications. While HPC systems are typically managed in a coarse grained fashion, with exclusive resource allocations, DI systems employ a finer grained regime, enabling dynamic allocation and control based on application needs. On the path to full convergence, a useful and less intrusive step is a hybrid resource management system allowing the execution of DI applications on top of standard HPC scheduling systems. In this paper we present the architecture of a hybrid system enabling dual-level scheduling for DI jobs in HPC infrastructures. Our system takes advantage of real-time resource profiling to efficiently co-schedule HPC and DI applications. The architecture is easily extensible to current and new types of distributed applications, allowing efficient combination of hybrid workloads on HPC resources with increased job throughput and higher overall resource utilization. The implementation is based on the Slurm and Mesos resource managers for HPC and DI jobs. Experimental evaluations in a real cluster based on a set of representative HPC and DI applications demonstrate that our hybrid architecture improves resource utilization by 20%, with 12% decrease on queue makespan while still meeting all deadlines for HPC jobs. Abel Souza, Mohamad Rezaei, Erwin Laure, Johan Tordsson |
CCGRID | 1 |
| 2018 | Hybrid Adaptive Checkpointing for Virtual Machine Fault ToleranceabstractActive Virtual Machine (VM) replication is an application independent and cost-efficient mechanism for high availability and fault tolerance, with several recently proposed implementations based on checkpointing. However, these methods may suffer from large impacts on application latency, excessive resource usage overheads, and/or unpredictable behavior for varying workloads. To address these problems, we propose a hybrid approach through a Proportional-Integral (PI) controller to dynamically switch between periodic and on-demand check-pointing. Our mechanism automatically selects the method that minimizes application downtime by adapting itself to changes in workload characteristics. The implementation is based on modifications to QEMU, LibVirt, and OpenStack, to seamlessly provide fault tolerant VM provisioning and to enable the controller to dynamically select the best checkpointing mode. Our evaluation is based on experiments with a video streaming application, an e-commerce benchmark, and a software development tool. The experiments demonstrate that our adaptive hybrid approach improves both application availability and resource usage compared to static selection of a checkpointing method, with application performance gains and neglectable overheads. Abel Souza, Alessandro Vittorio Papadopoulos, Luis Tomás, Johan Tordsson |
IC2E | 1 |