Arsany Guirguis

dblp:150/6425 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0002-0898-0387ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Computer networks · 3 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2024 Accelerating Transfer Learning with Near-Data Computation on Cloud Object Stores
abstract
Storage disaggregation underlies today's cloud and is naturally complemented by pushing down some computation to storage, thus mitigating the potential network bottleneck between the storage and compute tiers. We show how ML training benefits from storage pushdowns by focusing on transfer learning (TL), the widespread technique that democratizes ML by reusing existing knowledge on related tasks. We propose HAPI, a new TL processing system centered around two complementary techniques that address challenges introduced by disaggregation. First, applications must carefully balance execution across tiers for performance. HAPI judiciously splits the TL computation during the feature extraction phase yielding pushdowns that not only improve network time but also improve total TL training time by overlapping the execution of consecutive training iterations across tiers. Second, operators want resource efficiency from the storage-side computational resources. HAPI employs storage-side batch size adaptation allowing increased storage-side pushdown concurrency without affecting training accuracy. HAPI yields up to 2.5× training speed-up while choosing in 86.8% of cases the best performing split point or one that is at most 5% off from the best.
Diana Petrescu, Arsany Guirguis, Do Le Quoc, Javier Picorel, Rachid Guerraoui, Florin Dinu
SoCC2
2022 Genuinely distributed Byzantine machine learning
abstract
Abstract Machine learning (ML) solutions are nowadays distributed, according to the so-called server/worker architecture. One server holds the model parameters while several workers train the model. Clearly, such architecture is prone to various types of component failures, which can be all encompassed within the spectrum of a Byzantine behavior. Several approaches have been proposed recently to tolerate Byzantine workers. Yet all require trusting a central parameter server. We initiate in this paper the study of the “general” Byzantine-resilient distributed machine learning problem where no individual component is trusted. In particular, we distribute the parameter server computation on several nodes. We show that this problem can be solved in an asynchronous system, despite the presence of $$\frac{1}{3}$$ 1 3 Byzantine parameter servers (i.e., $$n_{ps} > 3f_{ps}+1$$ n ps > 3 f ps + 1 ) and $$\frac{1}{3}$$ 1 3 Byzantine workers (i.e., $$n_w > 3f_w$$ n w > 3 f w ), which is asymptotically optimal. We present a new algorithm, ByzSGD, which solves the general Byzantine-resilient distributed machine learning problem by relying on three major schemes. The first, scatter/gather, is a communication scheme whose goal is to bound the maximum drift among models on correct servers. The second, distributed median contraction (DMC), leverages the geometric properties of the median in high dimensional spaces to bring parameters within the correct servers back close to each other, ensuring safe and lively learning. The third, Minimum-diameter averaging (MDA), is a statistically-robust gradient aggregation rule whose goal is to tolerate Byzantine workers. MDA requires a loose bound on the variance of non-Byzantine gradient estimates, compared to existing alternatives [e.g., Krum (Blanchard et al., in: Neural information processing systems, pp 118-128, 2017)]. Interestingly, ByzSGD ensures Byzantine resilience without adding communication rounds (on a normal path), compared to vanilla non-Byzantine alternatives. ByzSGD requires, however, a larger number of messages which, we show, can be reduced if we assume synchrony. We implemented ByzSGD on top of both TensorFlow and PyTorch, and we report on our evaluation results. In particular, we show that ByzSGD guarantees convergence with around 32% overhead compared to vanilla SGD. Furthermore, we show that ByzSGD’s throughput overhead is 24–176% in the synchronous case and 28–220% in the asynchronous case.
El Mahdi El Mhamdi, Rachid Guerraoui, Arsany Guirguis, Lê-Nguyên Hoang, Sébastien Rouault
Distributed Comput.3
2021 GARFIELD: System Support for Byzantine Machine Learning (Regular Paper)
abstract
We present GARFIELD, a library to transparently make machine learning (ML) applications, initially built with popular (but fragile) frameworks, e.g., TensorFlow and PyTorch, Byzantine-resilient. GARFIELD relies on a novel object-oriented design, reducing the coding effort, and addressing the vulnerability of the shared-graph architecture followed by classical ML frameworks. GARFIELD encompasses various communication patterns and supports computations on CPUs and GPUs, allowing addressing the general question of the practical cost of Byzantine resilience in ML applications. We report on the usage of GARFIELD on three main ML architectures: (a) a single server with multiple workers, (b) several servers and workers, and (c) peer-to-peer settings. Using GARFIELD, we highlight interesting facts about the cost of Byzantine resilience. In particular, (a) Byzantine resilience, unlike crash resilience, induces an accuracy loss, (b) the throughput overhead comes more from communication than from robust aggregation, and (c) tolerating Byzantine servers costs more than tolerating Byzantine workers.
Rachid Guerraoui, Arsany Guirguis, Jérémy Plassmann, Anton Ragot, Sébastien Rouault
DSN2
2021 Collaborative Learning in the Jungle (Decentralized, Byzantine, Heterogeneous, Asynchronous and Nonconvex Learning)
abstract
We study \emph{Byzantine collaborative learning}, where $n$ nodes seek to collectively learn from each others' local data. The data distribution may vary from one node to another. No node is trusted, and $f < n$ nodes can behave arbitrarily. We prove that collaborative learning is equivalent to a new form of agreement, which we call \emph{averaging agreement}. In this problem, nodes start each with an initial vector and seek to approximately agree on a common vector, which is close to the average of honest nodes' initial vectors. We present two asynchronous solutions to averaging agreement, each we prove optimal according to some dimension. The first, based on the minimum-diameter averaging, requires $n \geq 6f+1$, but achieves asymptotically the best-possible averaging constant up to a multiplicative constant. The second, based on reliable broadcast and coordinate-wise trimmed mean, achieves optimal Byzantine resilience, i.e., $n \geq 3f+1$. Each of these algorithms induces an optimal Byzantine collaborative learning protocol. In particular, our equivalence yields new impossibility theorems on what any collaborative learning algorithm can achieve in adversarial and heterogeneous environments.
El Mahdi El Mhamdi, Sadegh Farhadkhani, Rachid Guerraoui, Arsany Guirguis, Lê-Nguyên Hoang, Sébastien Rouault
NeurIPS4
2020 FeGAN: Scaling Distributed GANs
abstract
Existing approaches to distribute Generative Adversarial Networks (GANs) either (i) fail to scale for they typically put the two components of a GAN (the generator and the discriminator) on different machines, inducing significant communication overhead, or (ii) they face GAN training specific issues, exacerbated by distribution.
Rachid Guerraoui, Arsany Guirguis, Anne-Marie Kermarrec, Erwan Le Merrer
Middleware2
2020 Genuinely Distributed Byzantine Machine Learning
abstract
Machine Learning (ML) solutions are nowadays distributed, according to the so-called server/worker architecture. One server holds the model parameters while several workers train the model. Clearly, such architecture is prone to various types of component failures, which can be all encompassed within the spectrum of a Byzantine behavior. Several approaches have been proposed recently to tolerate Byzantine workers. Yet all require trusting a central parameter server. We initiate in this paper the study of the "general" Byzantine-resilient distributed machine learning problem where no individual component is trusted. In particular, we distribute the parameter server computation on several nodes.
El Mahdi El Mhamdi, Rachid Guerraoui, Arsany Guirguis, Lê-Nguyên Hoang, Sébastien Rouault
PODC3
2019 Primary User-Aware Optimal Discovery Routing for Cognitive Radio Networks
abstract
Routing protocols in multi-hop cognitive radio networks (CRNs) can be classified into two main categories: local and global routing. Local routing protocols aim at decreasing the overhead of the routing process while exploring the route by choosing, in a greedy manner, one of the direct neighbors. On the contrary, global routing protocols choose the optimal route by exploring the whole network to the destination paying the flooding overhead cost. In this paper, we propose a primary user-aware$k$-hop routing scheme where$k$is the discovery radius. This scheme can be plugged into any CRN routing protocol to adapt, in real time, to network dynamics like the number and activity of primary users. The aim of this scheme is to cover the gap between local and global routing protocols for CRNs. It is based on balancing the routing overhead and the route optimality, in terms of primary users avoidance, according to a user-defined utility function. We analytically derive the optimal discovery radius ($k$) that achieves this target. Evaluations on NS2 with a side-by-side comparison with traditional CRNs protocols show that our scheme can achieve the user-defined balance between the route optimality, which in turn reflected on throughput and packet delivery ratio, and the routing overhead in real time.
Arsany Guirguis, Fadel F. Digham, Karim G. Seddik, Mohamed Ibrahim Ahmed 0001, Khaled A. Harras, Moustafa Youssef 0001
IEEE Trans. Mob. Comput.1
2018 Cooperation-based multi-hop routing protocol for cognitive radio networks
Arsany Guirguis, Mohammed Karmoose, Karim Habak, Mustafa ElNainay, Moustafa Youssef 0001
J. Netw. Comput. Appl.1
2015 Primary User Aware k-Hop Routing for Cognitive Radio Networks
abstract
We propose a primary user-aware k-hop routing scheme that can be plugged into any cognitive radio network routing protocol to adapt, in real time, to the environmental changes. The main use of this scheme is to make the compromise required between the route overhead and its optimality based on a user-defined utility function. We analytically derive the optimal discovery radius (k) that achieves this target. Evaluations on NS2 show that our scheme can enhance the current routing protocols in terms of throughput with minimal overhead.
Arsany Guirguis, Mohamed Ibrahim Ahmed 0001, Karim G. Seddik, Khaled A. Harras, Fadel F. Digham, Moustafa Youssef 0001
GLOBECOM1