Ahmed M. Abdelmoniem

dblp:207/3449 · also Ahmed M. Abdel-Moniem, Ahmed Mohamed Abdelmoniem · DBLP profile ↗
← Back
58ranked-venue papers
17as first author
41since 2021 · last 2026
0000-0002-1374-1882ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 26 · 12 first-author · 15 since 2021Systems, architecture and hardware · 12 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 8 since 2021Security and privacy · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 R-CASH: Receiver-assisted Class-aware AQM via in-Switch Hysteresis Control for Data Center Networks
abstract
In data center networks (DCNs), incast traffic, primarily composed of small flows (S-Flows), faces congestion due to buffer bloating caused by large flows (L-Flows). Consequently, S-Flows experience inflated flow completion times (FCTs). This paper introduces R-CASH, a receiver-assisted class-aware active queue management (AQM) scheme that exploits in-switch hysteresis to address this issue. R-CASH relies on flow classification at the source host to differentiate S-Flows and L-Flows based on their size. While switches monitor queue lengths to determine congestion based on three hysteresis thresholds that divide the buffer into four zones, the state is encoded in two ToS bits on the packet header. Receivers decode this state and selectively clamp the receiver window (Rwnd) in the corresponding outgoing TCP ACKs to throttle L-Flows as necessary, thereby safeguarding S-Flows. R-CASH has been implemented in a small testbed and has demonstrated improved FCTs for S-Flows while maintaining sustained throughput for L-Flows compared to centralized AQMs.
Waheed G. Gadallah, Brahim Bensaou, Ahmed M. Abdelmoniem
ICC3
2026 FairMoE-FL: A Communication-Efficient and Fair Federated Mixture-of-Experts Framework
Asadullah Tariq, Mohamed Adel Serhani, Ahmed M. Abdelmoniem, Ikbal Taleb
ICC3
2026 Breaking the Boundary Barrier: Robust Model Fingerprinting via Unlearnable Examples in Model-Parameter Space
abstract
Deep learning models represent valuable intellectual property due to their high development costs. To protect model ownership, existing fingerprinting techniques have been proposed to use adversarial examples to fingerprint a model's decision boundaries. However, these fingerprints are inherently fragile, as model decision boundaries are highly sensitive to common model modifications such as fine-tuning, pruning, and adversarial training. In this paper, we propose MFUE (Model Fingerprinting via Unlearnable Examples), a novel fingerprinting methodology that leverages the stable unlearnability of unlearnable examples to fingerprint arbitrary modified models in parameter space, fundamentally circumventing the inherent vulnerability of decision boundaries. To achieve robust model fingerprinting in parameter space, we are the first to identify that unlearnable examples, owing to their persistent training resistance, can serve as stable fingerprints beyond the model's decision boundaries. To endow unlearnable examples with robustness against arbitrary model modifications, we introduce adversarial training that simulates the randomness of model modifications by jointly optimizing the unlearnable examples over models at different training stages. We evaluate the performance of MFUE against six different attack types, including both model and input tampering. Through extensive experiments, we demonstrate that MFUE outperforms four existing methods in terms of robustness and uniqueness.
Tianlong Xu, Zixiong Wang, Gaoyang Liu, Jian Chen 0046, Ahmed M. Abdelmoniem, Chen Wang 0011
KDD (1)5
2026 Federated Learning for Edge Computing Enabled Artificial Intelligence of Things: A comprehensive survey
abstract
Among contemporary AI computing paradigms, Federated Learning (FL) stands out as an innovative method and has shown great potential in conjunction with edge computing. The two techniques combined serve as a building block forthe development of the Artificial Intelligence of Things (AIoT). This paper sheds light on the synergistic integration of FL with edge computing to propel AIoT’s capabilities in decentralized environments. By executing computing tasks closer to the data, FL at the edge not only alleviates latency and bandwidth limitations inherent in cloud-centric architectures, but also presents a robust solution to privacy concerns—a crucial obstacle in traditional centralized training setups. This paper delves into how FL tackles these privacy issues, providing an intricate explanation of its operational principles, applications, and the resultant benefits for AIoT systems. Through this scrutiny, we highlight FL’s potential in bolstering the efficiency and privacy of AIoT deployments while also delineating future research directions and the expected impact across various domains. This study aims to comprehensively comprehend FL for Edge Computing-enabled AIoT and foster developments in intelligent technologies and applications in an interconnected world.
Qilei Li, Mingliang Gao 0001, Wenzhe Zhai, Wentai Wu, Chen Wang 0011, Ahmed M. Abdelmoniem
Knowl. Based Syst.6
2026 UDMP: Unified Delay-Driven Multipath Protocol for AI Clusters
abstract
Distributed AI model training generates bursty, low-entropy elephant flows that challenge existing single-path transport protocols in multi-stage Clos networks, leading to congestion and inefficiency. Multipath transport emerges as a promising solution, leveraging multiple paths to balance traffic and enhance resilience. However, current multipath RDMA solutions suffer from scalability, congestion control, and load-balancing inefficiencies. This paper introduces Unified Delay-driven Multipath Protocol (UDMP), a novel approach that co-designs congestion control and load balancing using network delay as a unified signal. UDMP employs delay-gradient-based congestion control to precisely resolve unavoidable congestion. Moreover, UDMP leverages delay-assisted load balancing to shift traffic across paths with minimal latency adaptively, maintaining throughput when encountering avoidable congestion. A novel Token Pool design integrates these components, eliminating per-path state overhead while achieving fine-grained traffic distribution. Implementations on DPDK and NS3 demonstrate that UDMP achieves up to 2x higher throughput and reduces flow completion times by up to 30% compared to state-of-the-art methods like MPRDMA and QP-Scaling. These results highlight UDMP’s effectiveness in meeting the stringent performance requirements of modern distributed AI training workloads.
Chengyuan Huang, Zhengqi Cui, Jun Xu 0037, Zhaochen Zhang, Li Wang 0110, Peirui Cao, Zhongming Ji, Jilei Chen, Shengju Zhang, Lingkun Meng, Ahmed M. Abdelmoniem, Fu Xiao 0001, Wan-Chun Dou, Guihai Chen, Keqiang He, Chen Tian 0001
IEEE Trans. Netw.12
2025 Prototype Surgery: Tailoring Neural Prototypes via Soft Labels for Efficient Machine Unlearning
abstract
The rapid advancements and widespread application of deep neural networks (DNNs), coupled with their reliance on sensitive and private data, have sparked growing concerns regarding data privacy and the ''right to be forgotten''. To address these concerns, machine unlearning has been proposed to efficiently eliminate the influence of specific training data from trained DNNs. However, existing machine unlearning methods struggle with the large number of parameters in trained DNNs, which lead to slow execution and high memory consumption, making them impractical for large-scale models. In this paper, we shift our focus to the small set of weights in the final classification layer of DNNs, which are defined as as ''prototypes'' for different classes. Our key observation is that the prototype associated with the unlearned training data undergoes a significant shift, whereas prototypes of unrelated classes exhibit only minor changes when comparing the prototypes of original and retrained models. Based on this observation, we propose a novel machine unlearning approach that efficiently achieves machine unlearning by directly adjusting the prototypes of DNNs. We first introduce Naive Prototype Surgery (Naive PS), a fast and simplified method that uses a closed-form solution to approximate unlearning effect by directly adjusting the prototype associated with the unlearned data. Next, we propose Prototype Surgery (PS), which incorporates soft label information to fine-tune the prototypes of all classes, to achieve a more effective unlearning. Both methods achieve data unlearning by only modifying the prototypes in the DNNs, thus avoiding the challenges posed by the large number of model parameters. Extensive experiments on four datasets demonstrate that our methods significantly accelerate the unlearning process while achieving comparable results to five existing methods in terms of both unlearning performance and privacy guarantee.
Gaoyang Liu, Xijie Wang, Zixiong Wang, Chen Wang 0011, Ahmed M. Abdelmoniem, Desheng Wang 0001
CCS5
2025 Discovering Latent Knowledge Prototypes for Heterogeneous Federated Learning
abstract
Federated learning (FL) is crucial for ensuring data privacy, a major concern in many applications. However, FL faces significant challenges due to data and model heterogeneity arising from diverse learning environments and the varying capabilities of participating entities. Most existing methods primarily concentrate on aggregating knowledge that is represented by models, logits, or features. which rely on specific assumptions that may not hold in real-world scenarios and thus fail to address both data and model heterogeneity simultaneously. In this work, we aim to address these challenges by tackling heterogeneity from both model and data perspectives while maintaining efficiency. To this end, we leverage locally encoded latent prototypes produced from the local knowledge memory bank to represent per-client knowledge updates, which are then aggregated on the server and transferred back to the clients for knowledge decoding and integration as global constraints for further local training. Considering the heterogeneity in model architectures, we design the knowledge encoder and decoder to be compatible with different model architectures and ensure robust prototype aggregation by aligning latent spaces to a common prior distribution, to enhance compatibility under diverse data distributions. We evaluate our method on multiple benchmarks and demonstrate its superior performance in terms of accuracy and effectiveness under various heterogeneous settings.
Qilei Li, Ahmed M. Abdelmoniem
ECAI2
2025 Split Fine-Tuning of BERT-Based Music Models in the Edge-Cloud Continuum: An Empirical Analysis
abstract
Transformer models have been instrumental in the proliferation of AI in modern life, with larger models demonstrating incredible performance in many domains and applications. However, with the rise of edge and mobile computing, there remains challenges with viably deploying these models in such settings. While these models can be split across edge and cloud nodes, the communication overhead can be substantial, creating a bottleneck and impacting the viability of real-world deployment. This work applies various compression techniques to split transformer models, focusing on two BERT models from different domains. Using a real edge device setup, we measure the communication time across the split, as well as the overhead introduced by compressing the communicated tensors. The impact on model accuracy by systematically moving the split layer choice is also systematically explored.
Bradley Aldous, Ahmed M. Abdelmoniem
HPCC2
2025 Query-based Knowledge Transfer for Heterogeneous Learning Environments
abstract
Decentralized collaborative learning under data heterogeneity and privacy constraints has rapidly advanced. However, existing solutions like federated learning, ensembles, and transfer learning, often fail to adequately serve the unique needs of clients, especially when local data representation is limited. To address this issue, we propose a novel framework called Query-based Knowledge Transfer (QKT) that enables tailored knowledge acquisition to fulfill specific client needs without direct data exchange. It employs a data-free masking strategy to facilitate the communication-efficient query-focused knowledge transformation while refining task-specific parameters to mitigate knowledge interference and forgetting. Our experiments, conducted on both standard and clinical benchmarks, show that QKT significantly outperforms existing collaborative learning methods by an average of 20.91% points in single-class query settings and an average of 14.32% points in multi-class query scenarios. Further analysis and ablation studies reveal that QKT effectively balances the learning of new and existing knowledge, showing strong potential for its application in decentralized learning.
Norah Alballa, Wenxuan Zhang 0003, Ziquan Liu, Ahmed M. Abdelmoniem, Mohamed Elhoseiny 0001, Marco Canini
ICLR4
2025 Hierarchical Knowledge Structuring for Effective Federated Learning in Heterogeneous Environments
abstract
Federated learning enables collaborative model training across distributed entities while maintaining individual data privacy. A key challenge in federated learning is balancing the personalization of models for local clients with generalization for the global model. Recent efforts leverage logit-based knowledge aggregation and distillation to overcome these issues. However, due to the non-IID nature of data across diverse clients and the imbalance in the client’s data distribution, directly aggregating the logits often produces biased knowledge that fails to apply to individual clients and obstructs the convergence of local training. To solve this issue, we propose a Hierarchical Knowledge Structuring (HKS) framework that formulates sample logits into a multi-granularity codebook to represent logits from personalized per-sample insights to globalized per-class knowledge. The unsupervised bottom-up clustering method is leveraged to enable the global server to provide multi-granularity responses to local clients. These responses allow local training to integrate supervised learning objectives with global generalization constraints, which results in more robust representations and improved knowledge sharing in subsequent training rounds. The proposed framework’s effectiveness is validated across various benchmarks and model architectures.
Wai Fong Tam, Qilei Li, Ahmed M. Abdelmoniem
IJCNN3
2025 From Expansion to Retraction: Long-tailed Machine Unlearning via Boundary Manipulation
abstract
Machine unlearning aims to remove the information of specific data from a trained machine learning model while retaining its utility for the remaining data, so as to meet the requirements of privacy regulations. Existing unlearning methods often assume a balanced data distribution, but neglect the real-world, long-tailed scenarios, where the decision boundaries of tail classes are frequently distorted due to insufficient sample representation, thereby reducing the unlearning efficacy. In this paper, we propose the first Long-Tailed Machine Unlearning (LTMU) framework from a unified decision-boundary perspective. Our framework begins with a directional boundary repair scheme designed to enrich the distorted decision boundary of the tail class, and then develop a novel boundary retraction approach tailored for long-tailed unlearning, dispersing both the augmented and original features throughout the feature space. This bidirectional manipulation not only offers a unified interpretation of the relationship between long-tailed learning and unlearning, but also enables flexible control over both repair and unlearning processes through the generation of augmented features, thereby effectively accomplishing the long-tailed unlearning task. Extensive experiments across multiple datasets and neural network architectures demonstrate the effectiveness of our framework in achieving complete unlearning of tail classes in long-tailed distributions.
Weizhuo Gao, Chen Wang 0011, Gaoyang Liu, Ahmed M. Abdelmoniem, Kai Peng 0001
KDD (2)5
2025 Unlocking the power of 4G/5G mobile networks: An empirical dive into quality and energy efficiency in YouTube Edge services
abstract
The advancements in 5G mobile networks and Edge computing offer great potential for services like augmented reality and Cloud gaming, thanks to their low latency and high bandwidth capabilities. However, the practical limitations of achieving optimal latency on real applications remain uncertain. This paper investigates the latency, bandwidth, and energy consumption of 5G Networks and leverages YouTube Edge service as the practical use case. We analyze how latency, bandwidth, and energy consumption differ between 4G LTE and 5G networks and how the location of YouTube Edge servers impacts these metrics. Surprisingly, our observations show that the 5G ecosystems have average latency hikes of up to 2 × , demonstrating that they are far from achieving their proclaimed promises. Our research study reveals over 10 significant observations and implications, indicating that the primary constraints on 4G/LTE and 5G capabilities are the ecosystem and energy efficiency of mobile devices’ downstream data. Moreover, our study demonstrates that to unlock the potential of 5G and its applications fully, it is crucial to prioritize efforts to improve the 5G ecosystem and introduce better methods and techniques to enhance energy efficiency.
Peixuan Song, Ahmed M. Abdelmoniem, Lev Mukhanov
Comput. Networks3
2025 Edge Intelligence for Intelligent Transport Systems: Approaches, challenges, and future directions
abstract
Intelligent Transportation Systems (ITS) are entering a new era with the integration of Distributed Edge Intelligence, which brings the power of artificial intelligence to the edge of the network. This survey provides a comprehensive review of the role of Distributed Edge Intelligence in ITS, emphasizing its applications, challenges, and implications. Unlike previous studies that focus on specific technologies such as communication, blockchain, cloud and fog computing, and security, this work highlights the unique integration of Edge Intelligence across various ITS components, including vehicles, infrastructure, and communication systems. The paper systematically examines these integrations, identifies key technical challenges, and offers insights into future research directions. By focusing on the transformative impact of Edge Intelligence, this study aims to complement existing surveys and guide researchers, practitioners, and policymakers in shaping the future of smart, sustainable transportation. Through this, we contribute to advancing ITS technology and fostering innovation in the transportation sector.
Arezoo Ghasemi, Amin Keshavarzi, Ahmed M. Abdelmoniem, Omid Reza Nejati, Tajedin Derikvand
Expert Syst. Appl.3
2025 A Data Poisoning Resistible and Privacy Protection Federated-Learning Mechanism for Ubiquitous IoT
abstract
As a novel distributed learning paradigm, federated learning (FL) allows clients to train global models collaboratively without exchanging private data. However, recent research not only demonstrates the vulnerability of FL against privacy attacks where adversaries try to recover private data by intercepting local gradients/models but also its inadequacy in defending against poisoning attacks launched by malicious adversaries, who modify local datasets to disrupt the global training process. Even though many solutions have been proposed to defend against these attacks, there is still a gap in mitigating the risks in more complex nonindependent and identically distributed (Non-IID) scenarios that are prevalent in Internet of Things (IoT) systems. To fill this gap, this article proposes a data poisoning resistible and privacy protection FL mechanism (DPR-PPFL) for ubiquitous IoT. Based on representational similarity analysis, DPR-PPFL allows clients to construct asymmetric local models in defending against data inversion attacks, and also the server to detect and aggregate benign local models uploaded by the clients to correctly train the global model in the face of data poisoning attacks. By comparing the performance of DPR-PPFL with state-of-the-art baselines, its merits in securing the learning process under IID and Non-IID scenes of IoT are demonstrated.
Gengxiang Chen, Linlin You, Ahmed M. Abdelmoniem, Yan Zhang 0002, Chau Yuen
IEEE Internet Things J.4
2025 AFML: An Asynchronous Federated Meta-Learning Mechanism for Charging Station Occupancy Prediction With Biased and Isolated Data
abstract
Electric vehicles (EVs) are driving green and low-carbon transport in modern cities. It makes charging station occupancy prediction (CSOP) critual for intelligent transportation systems (ITS) to achieve a balance between the supply and demand in resolving the dynamics between EVs and changing stations. Even though several Big Data-based solutions have been discussed, they are still struggling to collaboratively utilize heterogeneous data and distributed computing resources located at both physically and logicially isolated charging stations to better support context-driven CSOP. To addres this challenge, we propose an Asynchronous Federated Meta-learning Mechanism (AFML) for CSOP, which can train a meta-model with strong adaptation ability in an asynchronous and collaborative manner. In general, it incorporates an adaptive reptile algorithm (AR) and an weighted aggregation strategy (WA) to jointly ensure the training efficiency and model adaptivity. Evaluations on real-world CSOP datasets demonstrate that compared to the second best method, AFML can significantly improve forecasting accuracy by 14%, accelerate model convergence by 9% and enhance model generalizability by 10%, illustrating its merits in support CSOP to embrace a smart and sustainable city.
Linlin You, Haohao Qu, Ahmed M. Abdelmoniem, Chau Yuen
IEEE Trans. Big Data4
2025 Space-Frequency and Global-Local Attentive Networks for Sequential Deepfake Detection
abstract
The widespread misinformation generated by deepfake systems has emerged as a significant challenge in the dynamic realm of digital media. It poses threats to credibility, privacy, and security of information in daily life. Moreover, the increasing accessibility to facial editing tools further enables users to alter facial characteristics subtly through a series of intricate steps. To address the issue, we introduce a space–frequency and global–local attentive network (SFGLA-Net) for sequential deepfake detection. This method is designed to identify and analyze the sophisticated manipulated attributes of deepfake images. Specifically, we introduce a space–frequency fusion module to leverage the deep feature extracted in spatial and frequency domains, so as to exploit subtle inconsistencies and artifacts that are not perceptible in the spatial domain alone. Additionally, we design a global–local attention module to pinpoint the manipulated areas more accurately. Extensive experiments demonstrate the superior performance of the proposed method by significantly outperforming existing techniques in sequential deepfake detection. The code is available athttps://github.com/guishengzhanga/SFGLA.
Guisheng Zhang, Qilei Li, Mingliang Gao 0001, Siyou Guo, Gwanggil Jeon, Ahmed M. Abdelmoniem
IEEE Trans. Comput. Soc. Syst.6
2025 Poisoning as a Post-Protection: Mitigating Membership Privacy Leakage From Gradient and Prediction of Federated Models
abstract
Federated learning (FL) is a distributed learning paradigm that enables multiple clients to train a unified model without sharing their private data. However, recent works demonstrate that FL models are vulnerable to membership inference attacks (MIAs), which can infer whether a data sample was used to train a given FL model. Existing countermeasures either require far-reaching modifications of FL training process or enforce extra processing in prediction phase, yielding them unlikely to be applied well in practice. In this paper, we design a post-protection mechanism, dubbedP$^{2}$-Protection, which degrades the inference performance of MIAs by simultaneously poisoning the prediction and gradient of the target FL model to reduce the privacy leakage of training data while keeping the model prediction accuracy.P$^{2}$-Protectiononly involves one additional training round to embed the poisoned prediction and gradient into the target FL model, without requiring model retraining or training process modification. We evaluateP$^{2}$-Protectionand compare it with two state-of-the-art defenses against three MIAs on five realistic datasets. Experimental results show thatP$^{2}$-Protectionoutperforms the existing defenses by offering limited implement overhead and improved utility-privacy trade-off.
Gaoyang Liu, Tianlong Xu, Yang Yang 0060, Ahmed M. Abdelmoniem, Chen Wang 0011, Jiangchuan Liu
IEEE Trans. Dependable Secur. Comput.4
2025 Generic Representation Learning for Vehicle Association Guided by Foundational Models
abstract
Vehicle association is a vital yet complex task to retrieve specific vehicles across various camera angles, time frames, and geographical locations. In environments supported by autonomous driving and 6G networks, this task plays a vital role in urban surveillance and traffic management by enabling the real-time sharing of vehicle location and status information through ultra-high-speed, low-latency 6G communication. The success of a retrieval model largely depends on the quality of the extracted representations, which can be influenced by factors such as background diversity and occlusions. This study proposes a method to extract representations that remain consistent across different domains while retaining the discriminative power necessary to determine a vehicle’s spatial location, regardless of background or environmental variations. To achieve this, we introduce a framework called Generic Representation Learning (GRL). Within GRL, we leverage large-scale pre-trained foundational models to provide spatial priors of vehicles, specifically the Grounding DINO model for object detection and the SAM model for object segmentation. These modules collaborate to help the network understand the spatial context of the object, enabling the feature extractor to focus on discriminative areas while minimizing interference. Additionally, we introduce a complementary feature alignment mechanism based on a memory bank to explore globally applicable knowledge within the learned representation of the object. These constituent elements collectively form SRP, to enhance its capability for outstanding performance in vehicle retrieval. Extensive experimentation demonstrates that SRP significantly outperforms existing models on widely recognized benchmarks.
Qilei Li, Mingliang Gao 0001, Wenzhe Zhai, Gwanggil Jeon, Ahmed M. Abdelmoniem
IEEE Trans. Intell. Transp. Syst.6
2025 Troubleshooting Programmable Data Planes via Real-Time Table Information Recording
abstract
While the flexibility of programmable switches brings opportunities, it also introduces security risks. Hence, it is vital to conduct effective troubleshooting in the programmable switch to mitigate frequent network failures. However, troubleshooting programmable switch failures is challenging due to their enhanced flexibility and functionality compared to regular switches, posing increased difficulty in debugging, particularly with limited debugging tools and information. To address this problem, we propose an efficient troubleshooting method that records real-time information about packets in the data plane, including the tables involved in packet processing. Unfortunately, due to hardware limitations, it is infeasible to record all tables’ information in the data plane. Thus, the key is to find the table set reflecting the execution path a packet goes through while minimizing the resource overhead. We first represent P4 programs as a probabilistic transition directed acyclic graph (DAG) and employ information entropy to quantify the information within a set of tracked tables. Then, we adopt a two-step approach and design algorithms to find both optimal and approximately optimal table record plans. The evaluation results show the efficacy of the proposed method, including achieving the same path recovery rate as the related works with less than one-third of the resource consumption.
Chengyuan Huang, Yibo Xiao, Tianfan Zhang, Bingheng Yan, Ahmed M. Abdelmoniem, Gianni Antichi, Xiaoliang Wang 0001, Fu Xiao 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
IEEE Trans. Netw.7
2024 FLOAT: Federated Learning Optimizations with Automated Tuning
abstract
Federated Learning (FL) has emerged as a powerful approach that enables collaborative distributed model training without the need for data sharing. However, FL grapples with inherent heterogeneity challenges leading to issues such as stragglers, dropouts, and performance variations. Selection of clients to run an FL instance is crucial, but existing strategies introduce biases and participation issues and do not consider resource efficiency. Communication and training acceleration solutions proposed to increase client participation also fall short due to the dynamic nature of system resources. We address these challenges in this paper by designing FLOAT, a novel framework designed to boost FL client resource awareness. FLOAT optimizes resource utilization dynamically for meeting training deadlines, and mitigates stragglers and dropouts through various optimization techniques; leading to enhanced model convergence and improved performance. FLOAT leverages multi-objective Reinforcement Learning with Human Feedback (RLHF) to automate the selection of the optimization techniques and their configurations, tailoring them to individual client resource conditions. Moreover, FLOAT seamlessly integrates into existing FL systems, maintaining non-intrusiveness and versatility for both asynchronous and synchronous FL settings. As per our evaluations, FLOAT increases accuracy by up to 53%, reduces client dropouts by up to 78×, and improves communication, computation, and memory utilization by up to 81×, 44×, and 20× respectively.
Ahmad Khan 0001, Azal Ahmad Khan, Ahmed M. Abdelmoniem, Samuel Fountain, Ali Raza Butt, Ali Anwar 0001
EuroSys3
2024 EMPRN: Reinforcement Learning-based ECN Tuning Using Message Passing Graph Recurrent Networks for Datacenters
abstract
Congestion control (CC) based on explicit congestion notification (ECN) is a common method for reducing latency and increasing link utilization in data center networks (DCN). Proper ECN tuning significantly impacts the performance of ECN-based CC algorithms. Due to the fast buffer buildup and dynamic spatial-temporal nature of traffic in high-speed DCNs, fast and online ECN tuning can reduce latency and packet loss. Most existing approaches do not capture the spatial dependencies between egress ports of switches. In this paper, we propose EMPRN, a novel in-network CC algorithm based on multi-agent reinforcement learning (MARL). We design a graph recurrent neural network for online ECN tuning. EMPRN is implemented in a distributed manner on switches, and it can be adapted to most ECN-based CC protocols. We use a message-passing neural network (MPNN) architecture to capture the spatial dependencies between egress ports. We integrate the proposed MPNN with a gated recurrent unit (GRU) network to learn both the spatial and temporal dependencies. Our simulation results show that our proposed approach achieves up to 21% and 87.9% reductions in terms of flow completion time (FCT) and average queue length, respectively, compared to a state-of-the-art reinforcement learning-based approach for online ECN tuning.
Kasra Zaeri, Ahmed M. Abdelmoniem
ICC2
2024 Decentralised Moderation for Interoperable Social Networks: A Conversation-Based Approach for Pleroma and the Fediverse
abstract
The recent development of decentralised and interoperable social networks (such as the "fediverse") creates new challenges for content moderators. This is because millions of posts generated on one server can easily "spread" to another, even if the recipient server has very different moderation policies. An obvious solution would be to leverage moderation tools to automatically tag (and filter) posts that contravene moderation policies, e.g. related to toxic speech. Recent work has exploited the conversational context of a post to improve this automatic tagging, e.g. using the replies to a post to help classify if it contains toxic speech. This has shown particular potential in environments with large training sets that contain complete conversations. This, however, creates challenges in a decentralised context, as a single conversation may be fragmented across multiple servers. Thus, each server only has a partial view of an entire conversation because conversations are often federated across servers in a non-synchronized fashion. To address this, we propose a decentralised conversation-aware content moderation approach suitable for the fediverse. Our approach employs a graph deep learning model (GraphNLI) trained locally on each server. The model exploits local data to train a model that combines post and conversational information captured through random walks to detect toxicity. We evaluate our approach with data from Pleroma, a major decentralised and interoperable micro-blogging network containing 2 million conversations. Our model effectively detects toxicity on larger instances, exclusively trained using their local post information (0.8837 macro-F1). Yet, we show that this approach does not perform well on smaller instances that do not possess sufficient local training data. Thus, in cases where a server contains insufficient data, we strategically retrieve information (posts or model parameters) from other servers to reconstruct larger conversations and improve results. With this, we show that we can attain a macro-F1 of 0.8826. Our approach has considerable scope to improve moderation in decentralised and interoperable social networks such as Pleroma or Mastodon.
Vibhor Agarwal, Aravindh Raman, Nishanth Sastry, Ahmed M. Abdelmoniem, Gareth Tyson, Ignacio Castro
ICWSM4
2024 MpScope: Enabling multi-pipeline monitoring inside a switch
Chengyuan Huang, Tianfan Zhang, Li Wang 0110, Yibo Xiao, Chen Tian 0001, Xiaoliang Wang 0001, Bingheng Yan, Ahmed M. Abdelmoniem, Wan-Chun Dou, Guihai Chen
Comput. Networks10
2024 Towards Energy-Aware Federated Learning via Collaborative Computing Approach
abstract
This research delves into the consequences of the high complexity of on-device operations executed during the federated learning process. We investigate how the varying computational capabilities and battery levels among mobile devices can introduce performance disparities and influence training quality. Hence, in order to deal with these challenges, we propose EAFL+, a novel energy optimization technique, that focuses on managing power consumption in devices with limited battery capacity. EAFL+ is a cloud–edge-terminal collaborative approach that provides a new architectural design for achieving power-aware FL training by leveraging resource diversity and computation offloading. The innovative scheme enables the efficient selection of an approximately-optimal offloading target, from a set of Cloud-tier, Edge-tier, and Terminal-tier resources and achieves the best cost-quality tradeoff for the devices taking part in the FL system. Our evaluation shows EAFL+ can help conserve the devices’ energy participating in training, which improves the participation rates and increases the clients’ contributions, hence achieving higher accuracy and faster convergence. Through experiments on real datasets and traces in an emulated FL environment, EAFL+ reduces the drop-out of clients to zero and enhances accuracy by up to 24% and 9% compared to EAFL and Oort, respectively.
Amna Arouj, Ahmed M. Abdelmoniem
Comput. Commun.2
2024 Dimensioning the pending interest table in content-centric networks
abstract
In chunk-based interest-driven content-centric networks, each interest packet forwarded upstream by a node on a face implies the return of at most one data chunk to the node from that face shortly after. As a consequence, the congestion of the downstream transmission buffer in the data path of the node is highly correlated with the occupancy of the pending interest table (PIT). Therefore a systematic study and analysis of the PIT occupancy are of paramount importance to understanding congestion in CCN. In particular, in this paper, we propose an analytical model to estimate the PIT occupancy distribution via a continuous-time Markov chain (CTMC) model that considers the effects of interest blocking, interest timeout and retries. To validate our model and assumptions, we invoke simulation with realistic traffic streams and show how the filtering effects of caching and interest aggregation make the Markov assumption reasonable in the nodes that are the most susceptible to experiencing congestion. To solve the model numerically, we use two alternative approximations, state space truncation and state aggregation, and then give some numerical results to demonstrate the accuracy of our approximations.
Amuda James Abu, Brahim Bensaou, Ahmed M. Abdelmoniem
Future Gener. Comput. Syst.3
2024 Alleviating Congestion via Switch Design for Fair Buffer Allocation in Datacenters
abstract
In data-centers, the composite origin and bursty nature of traffic, the small bandwidth-delay product and the tiny switch buffers lead to unusual congestion patterns that are not handled well by traditional end-to-end congestion control mechanisms such as TCP. Existing works address the problem by modifying TCP to adapt it to the idiosyncrasies of data-centers. While this is feasible in private environments, it remains almost impossible to achieve practically in public multi-tenant clouds where a multitude of operating systems and thus congestion control protocols coexist. In this work, we design a simple switch-based active queue management scheme to deal with such congestion issues adequately. Our approach requires no modification to TCP which enables easy and seamless deployment in public data-centers via switch software/firmware updates. We present a simple analysis to show its stability and effectiveness, then discuss three different real implementations in software and hardware switches on the NetFPGA platform. Numerical results from NS-2 simulation and experimental results from a small testbed cluster demonstrate the effectiveness of our approach in achieving high overall throughput, good fairness, smaller flow completion times (FCT) for short-lived flows, and reduction in the tail of the FCT distribution by as much as two orders of magnitude.
Ahmed M. Abdelmoniem, Brahim Bensaou
IEEE Trans. Cloud Comput.1
2024 FLAIR: A Fast and Low-Redundancy Failure Recovery Framework for Inter Data Center Network
abstract
Due to the fast developments of 5G and IoT technologies, Inter-Datacenter (Inter-DC) networks are facing unprecedented pressure to duplicate large volumes of geographically distributed user data in a real-time manner. Meanwhile, with the expansion of Inter-DC networks scale, link/node failures also become increasingly frequent, negatively affecting the data transmission efficiency. Therefore, link failure recovery methods become of utmost importance. Many works investigated fast failure recovery, yet none of them consider the deployment overhead of such recovery schemes. While in this paper, we found that the side-effect of deploying recovery strategies and the future availability of the recovered transmissions are also crucial for fast recovery. So we propose a fast and low-redundancy failure recovery framework, FLAIR, which consists of a fast recovery strategy FRAVaR and a redundancy removal algorithm ROSE. FRAVaR takes full consideration of deployment overhead by minimizing shuffle traffic. On its base, ROSE regularly eliminates the cumulative rerouting redundancy by removing unnecessary routing updates. The experiment results on 4 realistic network topologies show that FLAIR successfully reduces up to 48.2% deployment overhead compared with the state-of-the-art solutions, and thus reduces up to 70.2% recovery speed and improves up to 36% network utilization.
Yuchao Zhang 0004, Haoqiang Huang, Ahmed M. Abdelmoniem, Gaoxiong Zeng, Chenyue Zheng, Xirong Que, Wendong Wang 0003, Ke Xu 0002
IEEE Trans. Cloud Comput.3
2024 Manipulating Pre-Trained Encoder for Targeted Poisoning Attacks in Contrastive Learning
abstract
In recent years, contrastive learning has become very powerful for representation learning using large-scale unlabeled data, by involving pre-trained encoders to fine-tune downstream classifiers. However, the latest research indicates that contrastive learning can potentially suffer from the risks of data poisoning attacks, where the attacker injects maliciously crafted poisoned samples into the unlabeled pre-training data. To step forward, in this paper, we present a more stealthy poisoning attack dubbed PA-CL to directly poison the pre-trained encoder, such that the downstream classifier’s behavior on a single target instance to the attacker-desired class can be manipulated without affecting the overall downstream classification performance. We observe that a high similarity exists between the feature representation generated by the poisoned pre-trained encoder for the target sample and samples from the attacker-desired class. This leads to the downstream classifier misclassifying the target sample with the attacker-desired class. Therefore, we formulate our attack as an optimization problem, and design two novel loss functions, namely, the target effectiveness loss to effectively poison the pre-trained encoder, and the model utility loss to maintain the downstream classification performance. Experimental results on four real-world datasets demonstrate that the attack success rate of the proposed attack is 40% higher on average than that of the three baseline attacks, and the fluctuation of the downstream classifier’s prediction accuracy is within 5%.
Jian Chen 0046, Gaoyang Liu, Ahmed M. Abdelmoniem, Chen Wang 0011
IEEE Trans. Inf. Forensics Secur.4
2024 Leveraging chaos for enhancing encryption and compression in large cloud data transfers
Shiladitya Bhattacharjee, Himanshi Sharma, Tanupriya Choudhury, Ahmed M. Abdelmoniem
J. Supercomput.4
2024 Publisher Correction to: Leveraging chaos for enhancing encryption and compression in large cloud data transfers
Shiladitya Bhattacharjee, Himanshi Sharma, Tanupriya Choudhury, Ahmed M. Abdelmoniem
J. Supercomput.4
2023 REFL: Resource-Efficient Federated Learning
abstract
Federated Learning (FL) enables distributed training by learners using local data, thereby enhancing privacy and reducing communication. However, it presents numerous challenges relating to the heterogeneity of the data distribution, device capabilities, and participant availability as deployments scale, which can impact both model convergence and bias. Existing FL schemes use random participant selection to improve the fairness of the selection process; however, this can result in inefficient use of resources and lower quality training. In this work, we systematically address the question of resource efficiency in FL, showing the benefits of intelligent participant selection, and incorporation of updates from straggling participants. We demonstrate how these factors enable resource efficiency while also improving trained model quality.
Ahmed M. Abdelmoniem, Atal Narayan Sahu, Marco Canini, Suhaib A. Fahmy
EuroSys1
2023 A2FL: Availability-Aware Selection for Machine Learning on Clients with Federated Big Data
abstract
Recent advances in Big Data Analytics are primarily driven by innovations in Artificial Intelligence and Machine Learning Methods. Due to the richness of data sources at the edge and with the increasing privacy concerns, Distributed privacy-preserving machine learning (ML) methods are increasingly becoming the norm for training ML models on federated big data. In a popular approach known as Federated learning (FL), service providers leverage end-user data to train ML models to improve services such as text auto-completion, virtual keyboards, and item recommendations. FL is expected to grow in importance with the increasing focus on big data, privacy and 5G/6G technologies. However, FL faces significant challenges such as heterogeneity, communication overheads, and privacy preservation. In practice, training models via FL is time-intensive and worse its dependent on client participation who may not always be available to join the training. Our empirical analysis shows that client availability can significantly impact the model quality which motivates the design of an availability-aware selection scheme. We propose A2FL to mitigate the quality degradation caused by the under-representation of the global client population by prioritizing the least available clients. Our results show that, compared to state-of-the-art methods, A2FL can improve the client diversity during the training and hence boost the trained model quality.
Ahmed M. Abdelmoniem, Yomna M. Abdelmoniem, Ahmed Elzanaty
ICC1
2023 Achieving Zero-copy Serialization for Datacenter RPC
abstract
Remote Procedure Call (RPC) is widely used in distributed systems and it usually needs to serialize data before transmission. Serialization accounts for a large proportion of the overhead in RPC and becomes a bottleneck for RPC communications. Because the size of the output serialized message cannot be predicted in advance, there could be multiple memory reallocations and copies in typical serialization libraries (e.g., FlatBuffers), which dominates the overhead. We propose the novel serialization library, zFlatBuffers, to eliminate these avoidable copies during the serialization process and realize zero copy during communication. Unlike the typical serialization library, FlatBuffers, the message generated by zFlatBuffers consists of multiple non-contiguous buffers due to its zero-copy nature. Moreover, we integrate zFlatBuffers with RDMA-based RPC systems. For RDMA Unreliable Datagram, we modify the message buffer of eRPC to enable it to transmit messages composed of multiple buffers. We also build the zRPC system based on RDMA Reliable Connection, which transmits the zFlatBuffers message by the scatter/gather function. Compared to the original FlatBuffers, zFlatBuffers improves the throughput of eRPC and zRPC by 11.2%-33.7% and 5.8%-53.6%, respectively.
Tianfan Zhang, Huaping Zhou, Chengyuan Huang, Chen Tian 0001, Xiaoliang Wang 0001, Ahmed M. Abdelmoniem, Matthew Tan, Wan-Chun Dou, Guihai Chen
IPCCC8
2023 A Comprehensive Empirical Study of Heterogeneity in Federated Learning
abstract
Federated learning (FL) is becoming a popular paradigm for collaborative learning over distributed, private data sets owned by nontrusting entities. FL has seen successful deployment in production environments, and it has been adopted in services, such as virtual keyboards, auto-completion, item recommendation, and several IoT applications. However, FL comes with the challenge of performing training over largely heterogeneous data sets, devices, and networks that are out of the control of the centralized FL server. Motivated by this inherent challenge, we aim to empirically characterize the impact of device and behavioral heterogeneity on the trained model. We conduct an extensive empirical study spanning nearly 1.5K unique configurations on five popular FL benchmarks. Our analysis shows that these sources of heterogeneity have a major impact on both model quality and fairness, causing up to$4.6\times $and$2.2\times $degradation in the quality and fairness, respectively, thus shedding light on the importance of considering heterogeneity in FL system design.
Ahmed M. Abdelmoniem, Chen-Yu Ho 0001, Pantelis Papageorgiou, Marco Canini
IEEE Internet Things J.1
2023 Knowledge Representation of Training Data With Adversarial Examples Supporting Decision Boundary
abstract
Deep learning (DL) has achieved tremendous success in recent years in many fields. The success of DL typically relies on a considerable amount of training data and the expensive model optimization process. Therefore, a trained DL model and its corresponding training data have become valuable assets whose intellectual property (IP) needs to be protected. Once a DL model or its training dataset is released, there is currently no mechanism for the entity that owns one part to establish a clear relationship with the other. In this paper, we aim to reveal the integrated relationship between a given DL model and the corresponding training dataset, by framing the problem of knowledge representation of a dataset with respect to DL models trained on it:how to effectively represent the knowledge transferred from a training dataset to a DL model?Our basic idea is that the knowledge transferred from a training dataset to a DL model can be uniquely represented by the model’s decision boundary. Therefore, we design a novel generation method that utilizes geometric consistency to find the samples supporting the decision boundary, which can serve as the proxy for the knowledge representation. We evaluate our method in three different cases: IP audit of training data, IP audit of DL models, and adversarial knowledge distillation. The experimental results show that our method can improve the performance of existing works in all cases, which confirm that our method can effectively represent the knowledge transferred from a training dataset to a DL model.
Zehao Tian, Zixiong Wang, Ahmed M. Abdelmoniem, Gaoyang Liu, Chen Wang 0011
IEEE Trans. Inf. Forensics Secur.3
2023 Enhancing TCP via Hysteresis Switching: Theoretical Analysis and Empirical Evaluation
abstract
In this paper we study the relationship between the TCP packet loss cycle and the performance of time-sensitive traffic in data centers. Using real traffic measurements and analysis, we find that such loss cycles are not long enough to enable most partition-aggregate time-sensitive TCP applications to recover their packet losses via the TCP 3-dup ACKs mechanism. As a result, the Timeout (RTO) mechanism is frequently triggered, leading to the expansion of the flow completion times (FCT) of such applications by orders of magnitude. Hence, we seek an alternative method that does not change the virtual machines and that can effectively expand the loss cycle duration to enable short flows to finish their transfer without incurring the cost of the RTO. To this end, we propose a novel TCP-AQM mechanism that alternates between a slow constant bitrate (CBR) mode and a fast TCP rate via hysteresis switching to expand the loss cycle. We prove the stability of the proposed TCP-AQM via a control theoretic model, then evaluate its performance gains via small and large scale NS2 simulation and by real FPGA implementation of a prototype on the NetFPGA platform. The results show considerable improvements in FCT distribution and reduction of missed deadlines in simulation and real experiments.
Ahmed M. Abdelmoniem, Brahim Bensaou
IEEE/ACM Trans. Netw.1
2021 A Two-tiered Caching Scheme for Information-Centric Networks
abstract
In information centric networking (ICN), by default, forwarder nodes along the paths from content producers to consumers, cache and reuse content chunks ubiquitously, invoking the Least Recently Used (LRU) replacement policy when needed. Due to the cache filtering effect, this ubiquitous-LRU strategy is inefficient: popular contents that are cached a few hops away from the edge are of little utility. Most alternative proposals adopt unified on-path schemes that rely on popularity statistics or other global state information to improve the performance. In the Internet, it is difficult to imagine different administrative-entities exposing such information to each other, and so we argue that such schemes are not realistic; and secondly, a unified caching scheme across the network ignores the administrative autonomy, which is one of the tenets of the Internet. In this paper we argue for a two-tiered caching scheme that maintains an on-path caching scheme (e.g., Ubiquitous-LRU), yet embraces AS-autonomy by adding within the AS an off-path cooperative-caching and redundancy elimination, to make better use of caches in the edge where they are most valuable. We describe the implementation of our scheme and highlight the design choices to make the system practical. Our evaluation results show that our scheme can reduce cache misses and upstream traffic by up to 15% compared to the state-of-the-art.
Kelvin H. T. Chiu, Jason Min Wang, Ahmed M. Abdelmoniem, Brahim Bensaou
HPSR3
2021 GRACE: A Compressed Communication Framework for Distributed Machine Learning
abstract
Powerful computer clusters are used nowadays to train complex deep neural networks (DNN) on large datasets. Distributed training increasingly becomes communication bound. For this reason, many lossy compression techniques have been proposed to reduce the volume of transferred data. Unfortunately, it is difficult to argue about the behavior of compression methods, because existing work relies on inconsistent evaluation testbeds and largely ignores the performance impact of practical system configurations. In this paper, we present a comprehensive survey of the most influential compressed communication methods for DNN training, together with an intuitive classification (i.e., quantization, sparsification, hybrid and low-rank). Next, we propose GRACE, a unified framework and API that allows for consistent and easy implementation of compressed communication on popular machine learning toolkits. We instantiate GRACE on TensorFlow and PyTorch, and implement 16 such methods. Finally, we present a thorough quantitative evaluation with a variety of DNNs (convolutional and recurrent), datasets and system configurations. We show that the DNN architecture affects the relative performance among methods. Interestingly, depending on the underlying communication library and computational cost of compression / decompression, we demonstrate that some methods may be impractical. GRACE and the entire benchmarking suite are available as open-source.
Chen-Yu Ho 0001, Ahmed M. Abdelmoniem, Aritra Dutta, El Houcine Bergou, Konstantinos Karatsenidis, Marco Canini, Panos Kalnis
ICDCS3
2021 DC2: Delay-aware Compression Control for Distributed Machine Learning
abstract
Distributed training performs data-parallel training of DNN models which is a necessity for increasingly complex models and large datasets. Recent works are identifying major communication bottlenecks in distributed training. These works seek possible opportunities to speed-up the training in systems supporting distributed ML workloads. As communication reduction, compression techniques are proposed to speed up this communication phase. However, compression comes at the cost of reduced model accuracy, especially when compression is applied arbitrarily. Instead, we advocate a more controlled use of compression and propose DC2, a delay-aware compression control mechanism. DC2 couples compression control and network delays in applying compression adaptively. DC2 not only compensates for network variations but can also strike a better trade-off between training speed and accuracy. DC2 is implemented as a drop-in module to the communication library used by the ML toolkit and can operate in a variety of network settings. We empirically evaluate DC2 in network environments exhibiting low and high delay variations. Our evaluation of different popular CNN models and datasets shows that DC2 improves training speed-ups of up to 41× and 5.3 × over baselines with no-compression and uniform compression, respectively.
Ahmed M. Abdelmoniem, Marco Canini
INFOCOM1
2021 Rethinking gradient sparsification as total error minimization
abstract
Gradient compression is a widely-established remedy to tackle the communication bottleneck in distributed training of large deep neural networks (DNNs). Under the error-feedback framework, Top-$k$ sparsification, sometimes with $k$ as little as 0.1% of the gradient size, enables training to the same model quality as the uncompressed case for a similar iteration count. From the optimization perspective, we find that Top-$k$ is the communication-optimal sparsifier given a per-iteration $k$ element budget.We argue that to further the benefits of gradient sparsification, especially for DNNs, a different perspective is necessary — one that moves from per-iteration optimality to consider optimality for the entire training.We identify that the total error — the sum of the compression errors for all iterations — encapsulates sparsification throughout training. Then, we propose a communication complexity model that minimizes the total error under a communication budget for the entire training. We find that the hard-threshold sparsifier, a variant of the Top-$k$ sparsifier with $k$ determined by a constant hard-threshold, is the optimal sparsifier for this model. Motivated by this, we provide convex and non-convex convergence analyses for the hard-threshold sparsifier with error-feedback. We show that hard-threshold has the same asymptotic convergence and linear speedup property as SGD in both the case, and unlike with Top-$k$ sparsifier, has no impact due to data-heterogeneity. Our diverse experiments on various DNNs and a logistic regression model demonstrate that the hard-threshold sparsifier is more communication-efficient than Top-$k$.
Atal Narayan Sahu, Aritra Dutta, Ahmed M. Abdelmoniem, Trambak Banerjee, Marco Canini, Panos Kalnis
NeurIPS3
2021 T-RACKs: A Faster Recovery Mechanism for TCP in Data Center Networks
abstract
Cloud interactive data-driven applications generate swarms of small TCP flows that compete for the small switch buffer space in data-center. Such applications require a small flow completion time (FCT) to be effective. Unfortunately, TCP is myopic with respect to the composite nature of application data. In addition it tends to artificially inflate the FCT of individual flows by several orders of magnitude, because of its Internet-centric design, that fixes the retransmission timeout (RTO) to be at least hundreds of milliseconds. To better understand this problem, in this paper, we use empirical measurements in a small data center testbed to study, at a microscopic level, the effects of various types of packet losses on TCP's performance. In particular, we single out packet losses that impact the tail end of small flows, as well as bursty losses that span a significant fraction of small TCP congestion windows, and show a non-negligible effect of such losses on the FCT. Based on this, we propose the so-called, timely-retransmitted ACKs (or T-RACKs), a simple loss recovery mechanism that conceals the drawbacks of the long RTO even in the presence of heavy packet losses. Interestingly enough, T-RACKS achieves this transparently to TCP itself as it does not require any change to TCP in the tenant's virtual machine (VM) or container. T-RACKs can be implemented as a software shim layer in the hypervisor between the VMs and the server's NIC or in hardware as a networking function in a SmartNIC. Simulation and real testbed results show remarkable performance improvements.
Ahmed M. Abdelmoniem, Brahim Bensaou
IEEE/ACM Trans. Netw.1
2020 On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep Learning
abstract
Compressed communication, in the form of sparsification or quantization of stochastic gradients, is employed to reduce communication costs in distributed data-parallel training of deep neural networks. However, there exists a discrepancy between theory and practice: while theoretical analysis of most existing compression methods assumes compression is applied to the gradients of the entire model, many practical implementations operate individually on the gradients of each layer of the model.In this paper, we prove that layer-wise compression is, in theory, better, because the convergence rate is upper bounded by that of entire-model compression for a wide range of biased and unbiased compression methods. However, despite the theoretical bound, our experimental study of six well-known methods shows that convergence, in practice, may or may not be better, depending on the actual trained model and compression ratio. Our findings suggest that it would be advantageous for deep learning frameworks to include support for both layer-wise and entire-model compression.
Aritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho 0001, Atal Narayan Sahu, Marco Canini, Panos Kalnis
AAAI3
2020 Reducing Latency in Multi-Tenant Data Centers via Cautious Congestion Watch
abstract
Modern data centers host a plethora of interactive data-intensive applications. These are known to often generate large numbers of short-lived parallel flows that must complete their transfer quickly to meet the stringent performance requirements of interactive applications. Network resources (e.g., switch buffer space) in the data center are scarce and are easily congested, which unfortunately results in long latency. As a first solution, following a traditional design rationale, Data Center TCP (DCTCP) was proposed to speed up the completion time of short-lived flow by maintaining a low buffer occupancy in the switch. In general, DCTCP performs well in homogeneous environments, however, its performance degrades quickly in heterogeneous environments or when it uses a large initial congestion window. To resolve this problem, we propose a Hypervisor-based congestion watching mechanism (HWatch), which measures the network load in the data center via ECN and uses the resulting statistics to determine the appropriate initial congestion window size to avoid congestion. HWatch neither needs modification to the TCP stack in the VMs nor requires any specialized network hardware features to meet its targets. In our evaluation, we demonstrate the benefits of HWatch in improving the performance of TCP flows through large-scale ns2 simulation and via testbed experiments in a small data center.
Ahmed M. Abdelmoniem, Hengky Susanto, Brahim Bensaou
ICPP1
2019 Creek: Inter Many-to-Many Coflows Scheduling for Datacenter Networks
abstract
Datacenter networked applications, often require multiple data transfer flows that semantically constitute a coflow group. A coflow is thus considered completed when all the transfers in the coflow are completed. Hence, application performance is optimized whenever the completion time of a coflow is minimized, rather than that of the flows composing it. Currently, popular coflow scheduling algorithms are mostly centralized, and they incur high overheads. The decentralized approach in the “many-to-many” scenario also incurs high communication overheads due to the communication among the local controllers. In this paper, we present a coflow scheduling mechanism that aims to minimize the coflow completion time for coflows that show a many-to-many communication pattern, and as a byproduct communication overhead cost is also minimized. Our algorithm preserves compatibility with existing commodity switches and network protocols and improves the coflow completion times on average by 1.8 times compared to the baseline as demonstrated via testbed implementation and large-scale simulation.
Hengky Susanto, Ahmed M. Abdelmoniem, Brahim Bensaou
ICC2
2019 Taming Latency in Data Centers Via Active Congestion-Probing
abstract
In cloud environments, interactive applications deployed in data centers often generate swarms of short-lived data transfers (or flows) that face dramatic competition for the scarce switch buffer space from other short-lived as well as the long-lived flows. In the presence of bloated queues, such short-lived flows often experience multiple packet losses per round-trip time which often triggers the timeout-based loss recovery mechanism. A direct consequence of this is an inflated application response time that turns out to be orders of magnitude larger than what it should be. A data center aware TCP protocol (DCTCP) was designed as a new TCP specifically to address this issue, however, it does not consider its co-existence with other transport protocol (e.g., CuBIC and NewReno of Linux). In such situations, which are abundant in multi-tenant data centers, the legacy large initial congestion window sizes (e.g., 10 segments), induce multiple packet losses at the onset of a TCP flow, which forces timeout and even binary exponential backoff. In this paper, we propose a novel Hypervisor-based, application-transparent approach for active congestion probing to enable the hypervisor to infer on-path congestion before the TCP connection is fully established for new traffic to avoid such massive packet losses and timeout. The so-called ProBoSCIS mechanism does not require any changes to TCP, works with all versions of TCP and does not need any special network hardware features other than those that exist in today's data center commodity switches. We show its effectiveness via ns2 simulation and demonstrate its practical feasibility by implementing and deploying it in a small-scale data center test-bed. We show the significant reduction in application latency by adopting ProBoSCIS in a series of real experiments.
Ahmed M. Abdelmoniem, Brahim Bensaou, Hengky Susanto
ICDCS1
2019 A Near Optimal Multi-Faced Job Scheduler for Datacenter Workloads
abstract
As data-parallel applications process more complex data, the dependencies between computation jobs in a multi-stage job also become more complicated. However, most of the existing scheduling solutions primarily rely on total bytes sent (job size) to differentiate jobs where jobs with fewer? bytes sent are prioritized over the larger ones. This approach overlooks the fact that jobs may consist of multiple computation stages, and that the completion of a computation job stage depends on the completion of other jobs' stage. In this paper, we present a coflow scheduler of multi-stage jobs that minimizes the average job completion time. Our solution prioritizes jobs based on the multi-faceted characteristics of multi-stage job structure per stage, instead of total bytes sent. Our experiments show that our approach provides twice the performance of existing solutions on average and by four times in bursty traffic scenario.
Hengky Susanto, Ahmed M. Abdelmoniem, Honggang Zhang 0003, Benyuan Liu, Don Towsley
ICDCS2
2019 Hysteresis-based Active Queue Management for TCP Traffic in Data Centers
abstract
Much of the incremental improvement to TCP over the past three decades had the ultimate goal of making it more effective in using the long-fat pipes of the global Internet. This resulted in a rigid set of mechanisms in the protocol that put TCP at a disadvantage in small-delay environments such as data centers. In particular, in the presence of the shallow buffers of commodity switches and the short round trip times in data centers, the continued use of a large TCP initial congestion window and a huge minimum retransmission timeout (both inherited from the Internet-centric design) results in a very short TCP loss cycle that affects particularly the flow completion times of short-lived incast flows. In this paper, we first investigate empirically the TCP loss cycle and discuss its impact on packet losses, recovery and delay; then we propose a switch-based congestion controller with hysteresis (HSCC) that aims to stretch the TCP loss cycle without modifying TCP itself. To protect incast flows from severe congestion, HSCC is designed to transparently induce the TCP source to alternate between its native TCP congestion control algorithm and a slower more conservative constant bit rate flow control mode that is activated when congestion is imminent. We show the stability of HSCC via analytical modelling, and demonstrate its effectiveness via simulation and implementation in a small testbed.
Ahmed M. Abdelmoniem, Brahim Bensaou
INFOCOM1
2019 Uranus: Congestion-proportionality among slices based on Weighted Virtual Congestion Control
abstract
Modern data centers are the host for multitude of large-scale distributed applications. These applications generate tremendous amount of network flows to complete their tasks. At this scale, efficient network control manages the network traffic at the level of flow aggregates (or slices ) who need to share the network with respect to operator’s proportionality policy. Existing slice scheduling mechanisms can not meet this goal in multi-path data center networks. Hence, in this paper, we aim to fulfil this goal and satisfy the congestion proportionality policy for network sharing. The policy is applied to the traffic traversing congested links in the network. We propose Uranus, a novel slice scheduler based on a combination of flow-level control mechanisms. The scheduler implements two-tier weight allocation to individual flows. Then, relying on a non-blocking big switch abstraction, slice weights are allocated at the inter-rack level by aggregating the weights of rack-to-rack flows. Finally, Uranus can dynamically divide the rack-level weight to its constituent flows. We also implement Weighted Virtual Congestion Control (WVCC), an end-host shim-layer that enforces weighted bandwidth sharing among competing flows. Trace-driven NS3 simulations demonstrate that Uranus closely approximates the congestion-proportionality and is able to improve the proportional fairness by 31.49% compared to the state-of-the-art mechanisms. The results also prove Uranus’s capability of intra-slice scheduling optimization. Moreover, Uranus’s throughput in Clos fabrics outperforms the state-of-the-art mechanisms by 10%.
Jiaqing Dong, Chen Tian 0001, Ahmed M. Abdelmoniem, Huaping Zhou, Bo Bai 0001, Gong Zhang 0001
Comput. Networks4
2018 IncastGuard: An Efficient TCP-Incast Mitigation Mechanism for Cloud Networks
abstract
TCP's internet-centric design is oblivious to the composite nature of the data center traffic, and in the presence of small switch buffers and other long-lived flows (elephants), short-lived (mice) flows often experience “Incast” congestion events that significantly affect their flow completion time (FCT). In this paper, we implement a switch- assisted Incast congestion mitigation system that takes into consideration the existence of other elastic traffic in data center networks (DCNs). The so-called IncastGuard system does not require any changes to the TCP protocol at the sender nor at the receiver and is shown to be effective in mitigating the effects of incast congestion via a prototype implementation in NetFPGA and experiments in a small-scale cluster testbed.
Ahmed M. Abdelmoniem, Brahim Bensaou, Victor Barsoum
GLOBECOM1
2018 Curbing Timeouts for TCP-Incast in Data Centers via A Cross-Layer Faster Recovery Mechanism
abstract
We first study, at a microscopic level, the effects of various types of packet losses on TCP performance in a small data center. Then based on the findings we propose a simple recovery mechanism to combat the drawbacks of the long retransmission timeout. We emphasize through our empirical study that packet losses that occur at the tail of short-lived flows$a$nd/or bursty losses that span a large fraction of the congestion window are frequent in data center networks; and, in most cases, especially for short-lived flows, they result in a loss recovery that incurs waiting for a long retransmission timeout (RTO). The negative effect of frequent RTOs on the FCT is dramatic, yet recovery via RTO is merely a symptom of the pathological design of TCP's minimum RTO mechanism (set by default to the Internet scale). We propose the so-called Timely Retransmitted ACKs (T-RACKs), a very simple recovery mechanism for data centers, implemented as a shim layer between the virtual machines layer and the end-host NIC, to bridge the gap between TCP's huge RTO and the actual round trip times experienced in the data center. Compared to alternative solutions such as DCTCP, our T-RACKS has the virtue of not requiring any modification to TCP, which makes it readily deployable in virtualized multi-tenant public data centers. Experimental results show considerable improvements in the FCT distribution.
Ahmed M. Abdelmoniem, Brahim Bensaou
INFOCOM1
2017 SICC: SDN-based incast congestion control for data centers
abstract
Due to the partition/aggregate nature of many distributed cloud-based applications, incast traffic carried by TCP abounds in data center networks. TCP, being agnostic to such applications' traffic patterns and their delay-sensitivity, cannot cope with the resulting congestion events, leading to severe performance degradation. The co-existence of such incast traffic with other throughput-demanding elastic traffic flows in the network worsens the performance degradation further. In this paper, relying on the programmability of Software Defined Networks (SDN), we address this problem in an efficient and easily deployable manner. The proposed SDN-based incast congestion control framework relies on the SDN controller and the hypervisor programmability to solve such congestion problems without altering the guest virtual machines nor the network switches. We assess the performance of the proposed scheme via real deployment in a small-scale testbed and ns2 simulation in larger environments.
Ahmed M. Abdelmoniem, Brahim Bensaou, Amuda James Abu
ICC1
2017 Enforcing Transport-Agnostic Congestion Control in SDN-Based Data Centers
abstract
To meet the deadlines of interactive applications, congestion-agnostic transport protocols like UDP are increasingly used side-by-side with congestion-responsive TCP. As bandwidth is not totally virtualized in data centers, service outage may occur (for some applications) when such diverse traffics contend for the small buffers in the switches. In this paper we present SDN-GCC, a simple and practical software-based congestion control mechanism that puts monitoring and control decisions in a centralized controller and traffic control enforcement in the hypervisors on the servers. SDN-GCC builds a congestion control loop between the SDN controller and hypervisors without assuming any cooperation from tenants applications (or transport protocol), ultimately making it deployable in existing data centers without any service disruption or hardware upgrade. SDN-GCC is implemented and evaluated via simulation in ns2 as well as real-life small-scale test-bed deployment and experiments.
Ahmed M. Abdelmoniem, Brahim Bensaou
LCN1
2017 A Resilient Auction Framework for Deadline-Aware Jobs in Cloud Spot Market
abstract
Public cloud providers, such as Amazon EC2, offer idle computing resources known as spot instances at a much cheaper rate compared to On-Demand instances. Spot instance prices are set dynamically according to market demand. Cloud users request spot instances by submitting their bid, and if user's bid price exceeds current spot price then a spot instance is assigned to that user. The problem however is that while spot instances are executing their jobs, they can be revoked whenever the spot price rises above the current bid of the user. In such scenarios and to complete jobs reliably, we propose a set of improvements for the cloud spot market which benefits both the provider and users. Typically, the new framework allows users to bid different prices depending on their perceived urgency and nature of the running job. Hence, it practically allow them to negotiate the current bid price in a way that guarantees the timely completion of their jobs. To complement our intuition, we have conducted an empirical study using real cloud spot price traces to evaluate our framework strategies which aim to achieve a resilient deadline-aware auction framework.
Abadhan Saumya Sabyasachi, Hussain Mohammed Dipu Kabir, Ahmed M. Abdelmoniem, Subrota K. Mondal
SRDS3
2016 HyGenICC: Hypervisor-based generic IP congestion control for virtualized data centers
abstract
In today's modern cloud supported applications, the traffic relies on both congestion-responsive transport protocols like TCP and congestion-oblivious transport like UDP including all their variations. As bandwidth is not totally virtualized in today's data centers, the diverse responses of these protocols to congestion lead to inefficiencies and service disruption to some applications. In this paper we present HyGenICC, a simple, distributed and practical congestion control mechanism that puts congestion control back where it belongs, in the network layer, albeit, without modifying the network layer behaviour. To this end IP ECN marking in the switches to indicate congestion, combined with additional packet processing in the hypervisor enable HyGenICC to partition bandwidth among competing virtual machines (VMs) in datacenters effectively. To enable easy deployment in existing data center, HyGenICC design is subjected to several constraints such as: i) freedom of changes to the guest VMs' congestion control mechanism, and ii) reliance only on switch capabilities that are already available in today's commodity switches. We evaluate HyGenICC via extensive simulation in the ns2 network simulator.
Ahmed M. Abdelmoniem, Brahim Bensaou, Amuda James Abu
ICC1
2016 A Markov model of CCN pending interest table occupancy with interest timeout and retries
abstract
The occupancy of the pending interest table (PIT) in content-centric networks (CCN) and other information-centric networking (ICN) architectures has an undeniable effect on the congestion level and the performance of the network. Each interest packet forwarded upstream by a node results in at most one data packet arriving back at the node and being possibly forwarded to many interfaces downstream. Clearly, to regulate the rate of arrival of new data chunks at a node, one can regulate the arrival rate of interest packets. To be able to design effective mechanisms to control the traffic in CCN, a systematic study and analysis of the PIT occupancy is thus imperative. In this paper, we derive an analytical model to estimate the PIT occupancy distribution via an approximate continuous time Markov (CTMC) model and validate the model via simulation experiments with different system parameters. Our proposed model is not only useful in the design of traffic controllers for managing the PIT occupancy but also offers an effective method for dimensioning the PIT.
Amuda James Abu, Brahim Bensaou, Ahmed M. Abdelmoniem
ICC3
2016 Inferring and Controlling Congestion in CCN via the Pending Interest Table Occupancy
abstract
Despite the improvements brought about by content-centric networking's (CCN) in-network caching and interest aggregation, congestion can still take place in such networks due to the dominance of non-reusable content, high cache churn, high delay variations, premature timeouts and interest retransmission. This becomes even more dramatic when multi-path routing is adopted. Identifying that at a given node in CCN, the Pending Interest Table (PIT) occupancy can give a good estimate of the data workload to arrive to the node in the near future, we propose in this paper a novel mechanism to control congestion in CCN based on this idea. Our mechanism uses the average occupancy of the PIT to estimate the anticipated data packet transmission queue length and sends explicit congestion notification signals to the content requesters to reduce their interest sending rates when such anticipated queue size exceeds a threshold. We demonstrate the effectiveness of our proposed mechanism via ns-3 simulation.
Amuda James Abu, Brahim Bensaou, Ahmed M. Abdelmoniem
LCN3
2015 Incast-Aware Switch-Assisted TCP Congestion Control for Data Centers
abstract
Due to the partition/aggregate nature of many cloud applications, incast traffic is preponderant in data center networks (DCNs). Because TCP is agnostic to this composite nature of the applications traffic and their quality of service requirements, a few congestion events often degrade significantly the user perceived quality of service. This is exacerbated by the co-existence of such incast traffic with other elastic traffic flows in the network. In this paper we address the congestion problems of incast traffic and its interaction with other elastic traffic in DCNs. We propose a switch- assisted TCP congestion control via some small modifications to the switch software that do not require any modification to the TCP protocol nor to the TCP sender or receiver logic. We assess the performance of the proposed scheme via ns2 simulation as well as a real deployment in a small-scale testbed1.
Ahmed M. Abdelmoniem, Brahim Bensaou
GLOBECOM1
2010 An ant colony optimization algorithm for the mobile ad hoc network routing problem based on AODV protocol
abstract
In this paper, we present a modified on-demand routing algorithm for mobile ad-hoc networks (MANETs). The proposed algorithm is based on both the standard Ad-hoc On-demand Distance Vector (AODV) protocol and ant colony based optimization. The modified routing protocol is highly adaptive, efficient and scalable. The main goal in the design of the protocol was to reduce the routing overhead, response time, end-to-end delay and increase the performance. We refer to the new modified protocol as the Multi-Route AODV Ant routing algorithm (MRAA).
Ahmed M. Abdelmoniem, Marghny H. Mohamed, Abdel-Rahman Hedar
ISDA1