Walid A. Hanafy

dblp:238/6458 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0001-5765-8194ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 10 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FM-CAC: Carbon-Aware Control for Battery-Buffered Edge AI via Time-Series Foundation Models
Kang Yang 0005, Walid A. Hanafy, Prashant J. Shenoy, Mani Srivastava 0001
ISLPED2
2026 Untangling the Carbon-Cost Tradeoffs and Stampede Effect Challenges in Cloud Computing
abstract
As society explores new computing applications, data centers have seen a sharp rise in capacity and energy use, driving up data centers’ carbon footprint and raising concerns about their sustainability. In response, efforts to reduce emissions now complement the traditional goals of cutting energy costs and boosting performance. Given the spatiotemporal variability of the grid carbon intensity, researchers are increasingly adopting workload shifting strategies to minimize the operational carbon footprint of cloud workloads. However, shifting from cost- and performance-centric operations to carbon-aware operations introduces performance penalties and cost overheads. Moreover, large-scale spatial and temporal shifts can cause stampede effects, with synchronized migration to the same low-carbon regions or times, thereby straining resources and creating imbalances. In this paper, we quantify the carbon-cost-performance trade-offs of carbon-aware scheduling and evaluate the risk of such stampede effects. We provide real-world examples illustrating these tradeoffs for application and cloud providers. Lastly, we offer practical insights to help mitigate these challenges and guide sustainable scheduling decisions.
Walid A. Hanafy, Thanathorn Sukprasert, Abel Souza, David Irwin 0001, Prashant J. Shenoy
IEEE Trans. Computers1
2025 FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
abstract
Model serving systems have become popular for deploying deep learning models for various latency-sensitive inference tasks. While traditional replication-based methods have been used for failure-resilient model serving in the cloud, such methods are often infeasible in edge environments due to significant resource constraints that preclude full replication. To address this problem, this paper presents FailLite, a failure-resilient model serving system that employs (i) a heterogeneous replication where the failover model is a smaller variant of the original one, (ii) an intelligent approach that uses warm replicas to ensure quick failover for critical applications while using cold replicas, and (iii) progressive failover to provide low mean time to recovery (MTTR) for the remaining applications. We implement a full prototype of our system and demonstrate its efficacy on an experimental edge testbed and large-scale simulations. Our results using 27 models show that FailLite can recover all failed applications with 2× lower MTTR and only a 0.6% reduction in accuracy. Under extreme failure scenarios, where 50% of edge sites fail simultaneously, FailLite improves recovery rate by at least 39.3% compared to the baseline methods.
Walid A. Hanafy, Tarek F. Abdelzaher, David Irwin 0001, Jesse Milzman, Prashant J. Shenoy
SoCC2
2025 CarbonEdge: Leveraging Mesoscale Spatial Carbon-Intensity Variations for Low Carbon Edge Computing
abstract
The proliferation of latency-critical and compute-intensive edge applications is driving increases in computing demand and carbon emissions at the edge. To better understand carbon emissions at the edge, we analyze granular carbon intensity traces at intermediate "mesoscales," such as within a single US state or among neighboring countries in Europe, and observe significant variations in carbon intensity at these spatial scales. Importantly, our analysis shows that carbon intensity variations, which are known to occur at large continental scales (e.g., cloud regions), also occur at much finer spatial scales, making it feasible to exploit geographic workload shifting in the edge computing context. Motivated by these findings, we propose CarbonEdge, a carbon-aware framework for edge computing that optimizes the placement of edge workloads across mesoscale edge data centers to reduce carbon emissions while meeting latency SLOs. We implement CarbonEdge and evaluate it on a real edge computing testbed and through large-scale simulations for multiple edge workloads and settings. Our experimental results on a real testbed demonstrate that CarbonEdge can reduce emissions by up to 78.7% for a regional edge deployment in central Europe. Moreover, our CDN-scale experiments show potential savings of 49.5% and 67.8% in the US and Europe, respectively, while limiting the one-way latency increase to less than 5.5 ms.
Walid A. Hanafy, Abel Souza, Jan Harkes, David Irwin 0001, Mahadev Satyanarayanan, Prashant J. Shenoy
HPDC2
2025 LLM-Driven Auto Configuration for Transient IoT Device Collaboration
abstract
Today's Internet of Things (IoT) has evolved from simple sensing and actuation devices to those with embedded processing and intelligent services, enabling rich collaborations between users and their devices. However, enabling such collaboration becomes challenging when transient devices need to interact with host devices in temporarily visited environments. In such cases, fine-grained access control policies are necessary to ensure secure interactions; however, manually implementing them is often impractical for non-expert users. Moreover, at run-time, the system must automatically configure the devices and enforce such fine-grained access control rules. Additionally, the system must address the heterogeneity of devices.
Hetvi Shastri, Walid A. Hanafy, David Irwin 0001, Mani Srivastava 0001, Prashant J. Shenoy
SEC2
2025 Poster Abstract: Rethinking Collaboration Among Mobile Devices in IoT Environments
abstract
Many emerging IoT devices are mobile, enabling them to visit new environments and networks beyond their home networks. Mobile devices often have to interact and collaborate with users and their devices, which belong to the different administrative environments they are temporarily visiting. In this paper, we envision a system for seamless collaboration among transient devices in IoT environments. The system is based on zero-conf collaboration and allows for fine-grained access control. Our proposed design supports hardware-independent interfaces and supports a large number of devices.
Hetvi Shastri, Walid A. Hanafy, David Irwin 0001, Mani Srivastava 0001, Prashant J. Shenoy
SenSys2
2024 Going Green for Less Green: Optimizing the Cost of Reducing Cloud Carbon Emissions
abstract
The continued exponential growth of cloud datacenter capacity has increased awareness of the carbon emissions when executing large compute-intensive workloads. To reduce carbon emissions, cloud users often temporally shift their batch workloads to periods with low carbon intensity. While such time shifting can increase job completion times due to their delayed execution, the cost savings from cloud purchase options, such as reserved instances, also decrease when users operate in a carbon-aware manner. This happens because carbon-aware adjustments change the demand pattern by periodically leaving resources idle, which creates a trade-off between carbon emissions and cost. In this paper, we present GAIA, a carbon-aware scheduler that enables users to address the three-way trade-off between carbon, performance, and cost in cloud-based batch schedulers. Our results quantify the carbon-performance-cost trade-off in cloud platforms and show that compared to existing carbon-aware scheduling policies, our proposed policies can double the amount of carbon savings per percentage increase in cost, while decreasing the performance overhead by 26%.
Walid A. Hanafy, Qianlin Liang, Noman Bashir, Abel Souza, David Irwin 0001, Prashant J. Shenoy
ASPLOS (3)1
2024 CDN-Shifter: Leveraging Spatial Workload Shifting to Decarbonize Content Delivery Networks
abstract
Content Delivery Networks (CDNs) are Internet-scale systems that deliver streaming and web content to users from many geographically distributed edge data centers. Since large CDNs can comprise hundreds of thousands of servers deployed in thousands of global data centers, they can consume a large amount of energy for their operations and thus are responsible for large amounts of Green House Gas (GHG) emissions. As these networks scale to cope with increased demand for bandwidth-intensive content, their emissions are expected to rise further, making sustainable design and operation an important goal for the future. Since different geographic regions vary in the carbon intensity and cost of their electricity supply, in this paper, we consider spatial shifting as a key technique to jointly optimize the carbon emissions and energy costs of a CDN. We present two forms of shifting: spatial load shifting, which operates within the time scale of minutes, and VM capacity shifting, which operates at a coarse time scale of days or weeks. The proposed techniques jointly reduce carbon and electricity costs while considering the performance impact of increased request latency from such optimizations. Using real-world traces from a large CDN and carbon intensity and energy prices data from electric grids in different regions, we show that increasing the latency by 60ms can reduce carbon emissions by up to 35.5%, 78.6%, and 61.7% across the US, Europe, and worldwide, respectively. In addition, we show that capacity shifting can increase carbon savings by up to 61.2%. Finally, we analyze the benefits of spatial shifting and show that it increases carbon savings from added solar energy by 68% and 130% in the US and Europe, respectively.
Jorge Murillo, Walid A. Hanafy, David Irwin 0001, Ramesh K. Sitaraman, Prashant J. Shenoy
SoCC2
2024 Acies-OS: A Content-Centric Platform for Edge AI Twinning and Orchestration
abstract
This paper describes Acies-OS, a content-centric platform for edge AI twinning and orchestration that allows easy deployment, re-configuration, and control of edge AI services, augmented by a digital twin. The work is motivated by the proliferation of edge AI in a plethora of IoT applications, ranging from home automation to military defense, and the emergence of digital twins that go beyond monitoring and emulation into configuration management and optimization of edge capabilities. While past work focused on either the edge capabilities themselves or the digital twin, this work focuses on their seamless interactions, offering abstractions that enable the digital twin to manage and optimize an increasingly diverse edge AI system. Acies-OS features a structured namespace, a thin client library with flexible pub/sub-based communication, health monitoring support, and a control plane for twin-based value-added analysis and optimization. To illustrate the use of Acies-OS, we implemented a multi-node multi-modality vehicle classification application and used Acies-OS to interface it to a digital twin. We then deployed the system in the field to showcase run-time twin-based optimizations of inference latency, classification accuracy, and robustness to failures in noisy and challenging conditions.
Jinyang Li 0004, Yizhuo Chen, Tomoyoshi Kimura, Tianshi Wang 0002, Ruijie Wang 0004, Denizhan Kara, Yigong Hu, Walid A. Hanafy, Abel Souza, Prashant J. Shenoy, Maggie B. Wigness, Joydeep Bhattacharyya, Jae Kim, Guijun Wang, Greg Kimberly, Josh D. Eckhardt, Denis Osipychev, Tarek F. Abdelzaher
ICCCN9
2024 INSERT: In-Network Stateful End-to-End RDMA Telemetry
abstract
Remote Direct Memory Access (RDMA) has been widely adopted in modern data centers thanks to its high-throughput, low-latency data transfer capability and reduced CPU overhead. However, traditional network-flow-based monitoring is poor at interpreting RDMA communication and hence inadequate for gaining insights. In this paper, we present INSERT, an end-to-end RDMA telemetry system that enables seamless visibility into RDMA communication from the network layer all the way to the application layer. To this end, INSERT combines (i) eBPF-based transparent RDMA tracing on end-hosts and (ii) stateful RDMA network telemetry on programmable data plane. We implement RDMA network telemetry on programmable SmartNICs, where we address practical challenges for maintaining fine-grained state on massively-parallel packet processing pipelines. We demonstrate that INSERT can perform reasonably accurate telemetry at line-rate for different types of RDMA traffic even in the presence of out-of-order packets, and finally showcase two practical use cases that can benefit from it.
Hyunseok Chang, Walid A. Hanafy, Sarit Mukherjee, Limin Wang 0010
INFOCOM2
2023 Ecovisor: A Virtual Energy System for Carbon-Efficient Applications
abstract
Cloud platforms' rapid growth is raising significant concerns about their carbon emissions. To reduce carbon emissions, future cloud platforms will need to increase their reliance on renewable energy sources, such as solar and wind, which have zero emissions but are highly unreliable. Unfortunately, today's energy systems effectively mask this unreliability in hardware, which prevents applications from optimizing their carbon-efficiency, or work done per kilogram of carbon emitted. To address the problem, we design an "ecovisor", which virtualizes the energy system and exposes software-defined control of it to applications. An ecovisor enables each application to handle clean energy's unreliability in software based on its own specific requirements. We implement a small-scale ecovisor prototype that virtualizes a physical energy system to enable software-based application-level i) visibility into variable grid carbon-intensity and local renewable generation and ii) control of server power usage and battery charging and discharging. We evaluate the ecovisor approach by showing how multiple applications can concurrently exercise their virtual energy system in different ways to better optimize carbon-efficiency based on their specific requirements compared to general system-wide policies.
Abel Souza, Noman Bashir, Jorge Murillo, Walid A. Hanafy, Qianlin Liang, David Irwin 0001, Prashant J. Shenoy
ASPLOS (2)4
2023 Energy Time Fairness: Balancing Fair Allocation of Energy and Time for GPU Workloads
abstract
Traditionally, multi-tenant cloud and edge platforms use fair-share schedulers to fairly multiplex resources across applications. These schedulers ensure applications receive processing time proportional to a configurable share of the total time. Unfortunately, enforcing time-fairness across applications often violates energy-fairness, such that some applications consume more than their fair share of energy. This occurs because applications either do not fully utilize their resources or operate at a reduced frequency/voltage during their time-slice. The problem is particularly acute for machine learning (ML) applications using GPUs, where model size largely dictates utilization and energy usage. Enforcing energy-fairness is also important since energy is a costly and limited resource. For example, in cloud platforms, energy dominates operating costs and is limited by the power delivery infrastructure, while in edge platforms, energy is often scarce and limited by energy harvesting and battery constraints.
Qianlin Liang, Walid A. Hanafy, Noman Bashir, David Irwin 0001, Prashant J. Shenoy
SEC2
2023 Understanding the Benefits of Hardware-Accelerated Communication in Model-Serving Applications
abstract
It is commonly assumed that the end-to-end networking performance of edge offloading is purely dictated by that of the network connectivity between end devices and edge computing facilities, where ongoing innovation in 5G/6G networking can help. However, with the growing complexity of edge-offloaded computation and dynamic load balancing requirements, an offloaded task often goes through a multi-stage pipeline that spans across multiple compute nodes and proxies interconnected via a dedicated network fabric within a given edge computing facility. As the latest hardware-accelerated transport technologies such as RDMA and GPUDirect RDMA are adopted to build such network fabric, there is a need for good understanding of the full potential of these technologies in the context of computation offload and the effect of different factors such as GPU scheduling and characteristics of computation on the net performance gain achievable by these technologies. This paper unveils detailed insights into the latency overhead in typical machine learning (ML)-based computation pipelines and analyzes the potential benefits of adopting hardware-accelerated communication. To this end, we build a model-serving framework that supports various communication mechanisms. Using the framework, we identify performance bottlenecks in state-of-the-art model-serving pipelines and show how hardware-accelerated communication can alleviate them. For example, we show that GPUDirect RDMA can save 15-50% of model-serving latency, which amounts to 70–160 ms.
Walid A. Hanafy, Limin Wang 0010, Hyunseok Chang, Sarit Mukherjee, T. V. Lakshman, Prashant J. Shenoy
IWQoS1
2023 Model-driven Cluster Resource Management for AI Workloads in Edge Clouds
abstract
Since emerging edge applications such as Internet of Things (IoT) analytics and augmented reality have tight latency constraints, hardware AI accelerators have been recently proposed to speed up deep neural network (DNN) inference run by these applications. Resource-constrained edge servers and accelerators tend to be multiplexed across multiple IoT applications, introducing the potential for performance interference between latency-sensitive workloads. In this article, we design analytic models to capture the performance of DNN inference workloads on shared edge accelerators, such as GPU and edgeTPU, under different multiplexing and concurrency behaviors. After validating our models using extensive experiments, we use them to design various cluster resource management algorithms to intelligently manage multiple applications on edge accelerators while respecting their latency constraints. We implement a prototype of our system in Kubernetes and show that our system can host 2.3× more DNN applications in heterogeneous multi-tenant edge clusters with no latency violations when compared to traditional knapsack hosting algorithms.
Qianlin Liang, Walid A. Hanafy, Ahmed Ali-Eldin, Prashant J. Shenoy
ACM Trans. Auton. Adapt. Syst.2