Khaled Diab 0001

dblp:137/0579 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0002-1939-0673ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 AdaGen: Workload-Adaptive Cluster Scheduler for Latency-Optimal LLM Inference Serving
abstract
The inference workloads of Large Language Models (LLMs) pose significant latency and cost challenges due to increasing model sizes and demand for real-time responses. Existing cluster schedulers for multi-instance LLM serving primarily focus on load balancing to optimize memory usage, which is insufficient for workloads with diverse request characteristics. In such cases, the compute layout — the arrangement of tokens across iterations within each instance—plays a crucial role in determining latency. We propose AdaGen, a workload-adaptive cluster scheduler that minimizes latency and thus maximizes SLO attainment by optimizing compute layouts across instances. AdaGen employs a multi-step scheduling strategy: it first classifies requests based on prefill and decode lengths, then balances load, and finally performs selective distributed execution across instances. Each step incrementally refines the scheduling based on the compute layouts derived from the decision of the previous step. To avoid the overhead of actual execution to generate the layouts, AdaGen introduces a novel simulation-based estimator. Extensive experiments using production workloads show that AdaGen achieves up to 3.6× higher SLO attainment and 2× better cost-efficiency compared to the existing systems, while ensuring scalability.
Sudipta Saha Shubha, Ayush Goel, Diman Zad Tootaghaj, Khaled Diab 0001, Hardik Soni 0001, K. K. Ramakrishnan, Puneet Sharma 0001, Haiying Shen
EuroSys4
2026 Griffin: Coherency-Aware Task Scheduling and Memory Allocation for CXL Interconnects
abstract
CXL is an emerging interconnect that has the potential to efficiently realize memory disaggregation. This is because CXL enables the expansion of memory beyond individual hosts, and supports coherent memory sharing among multiple hosts. However, CXL introduces several performance overheads due to the cache coherency protocol for memory sharing, as well as placement constraints for shared data, which, if ignored, can lead to correctness issues. This paper presents the first analysis of the impact of CXL memory sharing and shows that the overheads of hardware-based coherency in CXL interconnects are substantial. We then propose Griffin, a new coherency-aware task and memory allocator for CXL disaggregated memory systems. Griffin introduces new abstractions and algorithms that allow it to prioritize which data is allocated remotely and to which memory node, to efficiently reduce the coherence overheads associated with both the amount of shared data and the load on CXL coherence resources. Our simulation results show that Griffin reduces the total memory time by up to 4.29 × compared to a standard baseline and 1.71 × compared to an advanced baseline.
Suyeon Lee, Khaled Diab 0001, Diman Zad Tootaghaj, Lianjie Cao, Puneet Sharma 0001, Ada Gavrilovska
ICS2
2026 CCSwitch: A Scalable Data Plane for Non-Blocking In-Network Collective Communication
abstract
Collective communication operations in AI and HPC workloads generate heavy network traffic. Offloading these operations to network switches reduces latency, but performing arithmetic and replication at line rate is difficult, especially as port counts and link speeds grow. Existing in-network approaches rely on accumulation buffers that not only limit throughput but also require complex state management to handle stragglers and congestion. We present CCSwitch, a modular switching fabric built from 4×4 non-blocking Collective Engines (CEs). Each CE combines spatial and temporal parallelism to perform reductions without accumulation buffers. CEs compose into k-ary n-tree topologies, scaling to 32- and 256-port switches while preserving non-blocking throughput. Source routing and flit-level synchronization keep per-switch state minimal. Our FPGA implementation shows that CCSwitch's quaternary-tree reduction fabric uses up to 23% fewer LUTs and 12–30% fewer flip-flops than a comparable Clos-based design at equal throughput. Enabling the full feature set—source routing, replication, and time-multiplexed VCs—uses 1.4–1.8× more LUTs than the circuit-switched baseline, well below the 3–5× overhead typical of packet-switched NoC routers, while supporting concurrent collectives on shared links.
Sumukh Pinge, Hardik Soni 0001, Bob Lantz, Khaled Diab 0001, Lianjie Cao, Tajana Rosing, Puneet Sharma 0001
SIGCOMM4
2026 DynamoServe: A Distributed Tiered Memory System for Multi-tenant LLM Serving
abstract
The rapid adoption of large language models (LLMs) has increased the need for efficient multi-tenant inference systems that maximize GPU utilization. However, existing frameworks struggle to scale due to the high memory demands of model weights and key-value (KV) caches. We present DynamoServe, a multi-tenant LLM serving framework that addresses these challenges through three key innovations: (1) leveraging stranded GPU memory to offload model weights and KV caches, (2) mitigating resource fragmentation in multi-workload environments, and (3) improving memory locality through coordinated data placement and demand-driven weight migration across GPUs. Together, these techniques enable high-throughput, low-latency inference. Experiments on state-of-the-art models show that DynamoServe significantly improves memory efficiency without sacrificing latency.
Diman Zad Tootaghaj, Khaled Diab 0001, Bob Lantz, Hanjiang Wu, K. K. Ramakrishnan, Md Ashfaqur Rahaman, Ryan Stutsman, Puneet Sharma 0001, Tushar Krishna
SIGCOMM2
2024 Horus: Granular In-Network Task Scheduler for Cloud Datacenters
Parham Yassini, Khaled Diab 0001, Saeed Mahloujifar, Mohamed Hefeeda
NSDI2
2022 Yeti: Stateless and Generalized Multicast Forwarding
Khaled Diab 0001, Mohamed Hefeeda
NSDI1
2022 Orca: Server-assisted Multicast for Datacenter Networks
Khaled Diab 0001, Parham Yassini, Mohamed Hefeeda
NSDI1
2021 DeepGame: Efficient Video Encoding for Cloud Gaming
abstract
Cloud gaming enables users to play games on virtually any device. This is achieved by offloading the game rendering and encoding to cloud datacenters. As game resolutions and frame rates increase, cloud gaming platforms face a major challenge to stream high quality games due to the high bandwidth and low latency requirements. In this paper, we propose a new video encoding pipeline, called DeepGame, for cloud gaming platforms to reduce the bandwidth requirements with limited to no impact on the player quality of experience. DeepGame learns the player's contextual interest in the game and the temporal correlation of that interest using a spatio-temporal deep neural network. Then, it encodes various areas in the video frames with different quality levels proportional to their contextual importance. DeepGame does not change the source code of the video encoder or the video game, and it does not require any additional hardware or software at the client side. We implemented DeepGame in an open-source cloud gaming platform and evaluated its performance using multiple popular games. We also conducted a subjective study with real players to demonstrate the potential gains achieved by DeepGame and its practicality. Our results show that DeepGame can reduce the bandwidth requirements by up to 36% compared to the baseline encoder, while maintaining the same level of perceived quality for players and running in real time.
Omar Mossad, Khaled Diab 0001, Ihab Amer, Mohamed Hefeeda
ACM Multimedia2
2020 Oktopus: Service Chaining for Multicast Traffic
abstract
Multicast service chaining refers to the orchestration of network services for multicast traffic. Paths of a multicast session that span the source, destinations and required services form a complex structure that we refer to as the multicast distribution graph. In this paper, we propose a new path-based algorithm, called Oktopus, that runs at the control plane of the ISP network to calculate the multicast distribution graph for a given session. Oktopus aims at minimizing the routing cost for each multicast session while satisfying all service chaining requirements. Oktopus consists of two steps. The first one generates a set of segments from the given ISP network topology, and the second step uses these segments to efficiently calculate the multicast distribution graph. Oktopus has a fine-grained control over the selection of links in the distribution graphs that leads to significant improvements. Specifically, Oktopus increases the number of allocated sessions because it can reach ISP locations that have the required services, and thus includes them in the calculated graph. Moreover, Oktopus can reduce the routing cost per session as it carefully chooses links belonging to the graph. We compared Oktopus against the optimal and closest algorithms using real ISP topologies. Our results show that Oktopus has an optimality gap of 5% on average, and it computes the distribution graphs multiple orders of magnitude faster than the optimal algorithm. Moreover, Oktopus outperforms the closest algorithm in the literature in terms of the number of allocated multicast sessions by up to 37%.
Khaled Diab 0001, Carlos Lee, Mohamed Hefeeda
ICNP1
2019 Joint Content Distribution and Traffic Engineering of Adaptive Videos in Telco-CDNs
abstract
Telco-CDNs refer to content distribution networks deployed and managed by Internet Service Providers (ISPs). They are getting popular among major ISPs because they offer new revenue streams and have the potential of providing better performance compared to traditional CDNs. Managing telco-CDNs is, however, a complex problem, because it requires jointly managing the network resources (links and switches) and the caching resources (processing and storage capacities), while supporting the adaptive nature and skewed popularity of multimedia content. To address this problem, we present a new algorithm called CAD (Cooperative Active Distribution), which strives to serve as much as possible of the requested multimedia objects within the ISP while carefully engineering the traffic paths through the network. This is achieved by enabling the cooperation among caches within the ISP not only to serve various representations of multimedia objects, but also to create them on demand using the available processing capacity of caches. We have implemented CAD and evaluated it on top of a network emulator that runs deployment code and processes real traffic. Using an actual ISP topology, our experimental results show that CAD achieves substantial performance improvements compared to the closest work in the literature, e.g., up to 64% reduction in the total inter-domain traffic.
Khaled Diab 0001, Mohamed Hefeeda
INFOCOM1
2019 Content-aware video encoding for cloud gaming
abstract
Cloud gaming allows users with thin-clients to play complex games on their end devices as the bulk of processing is offloaded to remote servers. A thin-client is only required to have basic decoding capabilities which exist on most modern devices. The result of the remote processing is an encoded video that gets streamed to the client. As modern games are complex in terms of graphics and motion, the encoded video requires high bandwidth to provide acceptable Quality of Experience (QoE) to end users. The cost incurred by the cloud gaming service provider to stream the encoded video at such high bandwidth grows rapidly with the increase in the number of users. In this paper, we present a content-aware video encoding method for cloud gaming (referred to as CAVE) to improve the perceptual quality of the streamed video frames with comparable bandwidth requirements. This is a challenging task because of the stringent requirements on latency in cloud gaming, which impose additional restrictions on frame sizes as well as processing time to limit the total latency perceived by clients. Unlike many of the previous works, the proposed method is suitable for the state-of-the-art High Efficiency Video Coding (HEVC) encoder, which by itself offers substantial bitrate savings compared to prior encoders. The proposed method leverages information from the game such as the Regions-of-Interest (ROIs), and optimizes the quality by allocating different amounts of bits to various areas in the video frames. Through actual implementation in an open-source cloud gaming platform, we show that the proposed method achieves quality gains in ROIs that can be translated to bitrate savings between 21% and 46% against the baseline HEVC encoder and between 12% and 89% against the closest work in the literature.
Mohamed Hegazy, Khaled Diab 0001, Mehdi Saeedi, Boris Ivanovic, Ihab Amer, Gabor Sines, Mohamed Hefeeda
MMSys2
2017 MASH: A rate adaptation algorithm for multiview video streaming over HTTP
abstract
Multiview videos offer unprecedented experience by allowing users to explore scenes from different angles and perspectives. Thus, such videos have been gaining substantial interest from major content providers such as Google and Facebook. Adaptive streaming of multiview videos is, however, challenging because of the Internet dynamics and the diversity of user interests and network conditions. To address this challenge, we propose a novel rate adaptation algorithm for multiview videos (called MASH). Streaming multiview videos is more user centric than single-view videos, because it heavily depends on how users interact with the different views. To efficiently support this interactivity, MASH constructs probabilistic view switching models that capture the switching behavior of the user in the current session, as well as the aggregate switching behavior across all previous sessions of the same video. MASH then utilizes these models to dynamically assign relative importance to different views. Furthermore, MASH uses a new buffer-based approach to request video segments of various views at different qualities, such that the quality of the streamed videos is maximized while the network bandwidth is not wasted. We have implemented a multiview video player and integrated MASH in it. We compare MASH versus the state-of-the-art algorithm used by YouTube for streaming multiview videos. Our experimental results show that MASH can produce much higher and smoother quality than the algorithm used by YouTube, while it is more efficient in using the network bandwidth. In addition, we conduct large-scale experiments with up to 100 concurrent multiview streaming sessions, and we show that MASH maintains fairness across competing sessions, and it does not overload the streaming server.
Khaled Diab 0001, Mohamed Hefeeda
INFOCOM1
2016 Depth Personalization and Streaming of Stereoscopic Sports Videos
abstract
Current three-dimensional displays cannot fully reproduce all depth cues used by a human observer in the real world. Instead, they create only an illusion of looking at a three-dimensional scene. This leads to a number of challenges during the content creation process. To assure correct depth reproduction and visual comfort, either the acquisition setup has to be carefully controlled or additional postprocessing techniques have to be applied. Furthermore, these manipulations need to account for a particular setup that is used to present the content, for example, viewing distance or screen size. This creates additional challenges in the context of personal use when stereoscopic content is shown on TV sets, desktop monitors, or mobile devices. We address this problem by presenting a new system for streaming stereoscopic content. Its key feature is a computationally efficient depth adjustment technique which can automatically optimize viewing experience for videos of field sports such as soccer, football, and tennis. Additionally, the method enables depth personalization to allow users to adjust the amount of depth according to their preferences. Our stereoscopic video streaming system was implemented, deployed, and tested with real users.
Kiana Calagari, Tarek Elgamal, Khaled Diab 0001, Krzysztof Templin, Piotr Didyk, Wojciech Matusik, Mohamed Hefeeda
ACM Trans. Multim. Comput. Commun. Appl.3
2014 Anahita: A System for 3D Video Streaming with Depth Customization
abstract
Producing high-quality stereoscopic 3D content requires significantly more effort than preparing regular video footage. In order to assure good depth perception and visual comfort, 3D videos need to be carefully adjusted to specific viewing conditions before they are shown to viewers. While most stereoscopic 3D content is designed for viewing in movie theaters, where viewing conditions do not vary significantly, adapting the same content for viewing on home TV-sets, desktop displays, laptops, and mobile devices requires additional adjustments. To address this challenge, we propose a new system for 3D video streaming that provides automatic depth adjustments as one of its key features. Our system takes into account both the content and the display type in order to customize 3D videos and maximize their perceived quality. We propose a novel method for depth adjustment that is well-suited for videos of field sports such as soccer, football, and tennis. Our method is computationally efficient and it does not introduce any visual artifacts. We have implemented our 3D streaming system and conducted two user studies, which show: (i) adapting stereoscopic 3D videos for different displays is beneficial, and (ii) our proposed system can achieve up to 35% improvement in the perceived quality of the stereoscopic 3D content.
Kiana Calagari, Krzysztof Templin, Tarek Elgamal, Khaled Diab 0001, Piotr Didyk, Wojciech Matusik, Mohamed Hefeeda
ACM Multimedia4
2014 Storage optimization for 3D streaming systems
abstract
Three dimensional (3D) content is becoming attractive in entertainment events such as soccer games and movies. Also, 3D displays are widespread at homes, offices, and theaters. Yet, 3D content may lack good 3D experience due to varying display technologies and sizes. In addition, 3D content providers may not be able to deliver their content to all potential subscribers, which leads to viewership reduction or dissatisfaction. In this work, we propose the design of a system for enhanced 3D content streaming. In order to support all 3D display technologies and sizes in the system, we design different 3D versions of the original videos that are optimized for various displays. Moreover, we propose a storage optimization algorithm that optimizes storage usage in our system depending on versions popularity as well as storage and processing requirements. The algorithm satisfies the limited processing resources, maximum delay, and request rate requirements. We implemented and deployed the proposed system on the cloud for live testing. We simulated the proposed algorithm to study its effect on the storage requirements in 3D streaming systems. The results of the simulations show that the algorithm can achieve storage gain up to 360x compared to storing all versions.
Khaled Diab 0001, Tarek Elgamal, Kiana Calagari, Mohamed Hefeeda
MMSys1