Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

John Stephan

dblp:185/5372 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
8since 2021 · last 2025
0009-0007-0293-1967ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 50% Trustworthy machine learning · 28% Optimization for machine learning · 22%
Network and information security
3 papers
Privacy and data protection · 89% Web and mobile security · 11%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
federated learning
1.722025
Towards Trustworthy Federated Learning with Untrusted Participants · ICML 2025
Adaptive Gradient Clipping for Robust Federated Learning · ICLR 2025
Machine learning › Trustworthy machine learning › robustness
byzantine robustness
1.732023
Robust Collaborative Learning with Linear Gradient Overhead · ICML 2023
Byzantine Machine Learning Made Easy By Resilient Averaging of Momentums · ICML 2022
Differential Privacy and Byzantine Resilience in SGD: Do They Add Up? · PODC 2021
Machine learning › Trustworthy machine learning
robustness
1.522025
Towards Trustworthy Federated Learning with Untrusted Participants · ICML 2025
On the Privacy-Robustness-Utility Trilemma in Distributed Learning · ICML 2023
Privacy and data protection
differential privacy
1.522025
Towards Trustworthy Federated Learning with Untrusted Participants · ICML 2025
On the Privacy-Robustness-Utility Trilemma in Distributed Learning · ICML 2023
Machine learning › Efficient and distributed learning
distributed training
1.432023
Robust Collaborative Learning with Linear Gradient Overhead · ICML 2023
Byzantine Machine Learning Made Easy By Resilient Averaging of Momentums · ICML 2022
Differential Privacy and Byzantine Resilience in SGD: Do They Add Up? · PODC 2021
Machine learning › Optimization for machine learning › adaptive optimization
adaptive gradient clipping
0.912025
Adaptive Gradient Clipping for Robust Federated Learning · ICLR 2025
Machine learning › Optimization for machine learning
gradient clipping
0.912025
Adaptive Gradient Clipping for Robust Federated Learning · ICLR 2025
Machine learning › Efficient and distributed learning › federated learning
robust federated learning
0.912025
Adaptive Gradient Clipping for Robust Federated Learning · ICLR 2025
Machine learning › Efficient and distributed learning › distributed training
robust gradient aggregation
0.912025
Towards Trustworthy Federated Learning with Untrusted Participants · ICML 2025
Machine learning › Efficient and distributed learning › federated learning
trustworthy federated learning
0.912025
Towards Trustworthy Federated Learning with Untrusted Participants · ICML 2025
Privacy and data protection › differential privacy › noise addition
correlated noise mechanisms
0.912025
Towards Trustworthy Federated Learning with Untrusted Participants · ICML 2025
Privacy and data protection › differential privacy
distributed differential privacy
0.712023
On the Privacy-Robustness-Utility Trilemma in Distributed Learning · ICML 2023
Distributed systems
fault tolerance
0.712023
Robust Collaborative Learning with Linear Gradient Overhead · ICML 2023
Machine learning › Optimization for machine learning
distributed optimization
0.612022
Byzantine Machine Learning Made Easy By Resilient Averaging of Momentums · ICML 2022
Machine learning › Optimization for machine learning
stochastic gradient descent
0.112021
Differential Privacy and Byzantine Resilience in SGD: Do They Add Up? · PODC 2021

Methods — techniques the papers use, named apart from their topics

distributed SGD · 1.9shared randomness · 1.7correlated noise · 1.7robust aggregation · 1.3polyak momentum · 1.3nearest-neighbor averaging · 1.3mean estimation · 1.3lower bound analysis · 1.3resilient averaging · 0.6momentum · 0.6app hiding · 0.4
YearPublicationVenuePosition
2025 Adaptive Gradient Clipping for Robust Federated Learning
abstract
Robust federated learning aims to maintain reliable performance despite the presence of adversarial or misbehaving workers. While state-of-the-art (SOTA) robust distributed gradient descent (Robust-DGD) methods were proven theoretically optimal, their empirical success has often relied on pre-aggregation gradient clipping. However, existing static clipping strategies yield inconsistent results: enhancing robustness against some attacks while being ineffective or even detrimental against others. To address this limitation, we propose a principled adaptive clipping strategy, Adaptive Robust Clipping (ARC), which dynamically adjusts clipping thresholds based on the input gradients. We prove that ARC not only preserves the theoretical robustness guarantees of SOTA Robust-DGD methods but also provably improves asymptotic convergence when the model is well-initialized. Extensive experiments on benchmark image classification tasks confirm these theoretical insights, demonstrating that ARC significantly enhances robustness, particularly in highly heterogeneous and adversarial settings.
Youssef Allouah, Rachid Guerraoui, Nirupam Gupta, Ahmed Jellouli, Geovani Rizk, John Stephan
ICLR6
2025 Towards Trustworthy Federated Learning with Untrusted Participants
abstract
Resilience against malicious participants and data privacy are essential for trustworthy federated learning, yet achieving both with good utility typically requires the strong assumption of a trusted central server. This paper shows that a significantly weaker assumption suffices: each pair of participants shares a randomness seed unknown to others. In a setting where malicious participants may collude with an untrusted server, we propose CafCor, an algorithm that integrates robust gradient aggregation with correlated noise injection, using shared randomness between participants. We prove that CafCor achieves strong privacy-utility trade-offs, significantly outperforming local differential privacy (DP) methods, which do not make any trust assumption, while approaching central DP utility, where the server is fully trusted. Empirical results on standard benchmarks validate CafCor's practicality, showing that privacy and robustness can coexist in distributed systems without sacrificing utility or trusting the server.
Youssef Allouah, Rachid Guerraoui, John Stephan
ICML3
2024 Towards Practical Homomorphic Aggregation in Byzantine-Resilient Distributed Learning
abstract
The growing availability of distributed data has led to the increased use of machine learning (ML) algorithms in distributed topologies, where multiple nodes collaborate to train models under the coordination of a central server. However, distributed learning faces two significant challenges: the risk of Byzantine nodes corrupting the learning process by sending incorrect information, and the potential for a curious server to violate the privacy of individual nodes, even reconstructing their private data. While homomorphic encryption (HE) has been a promising solution for privacy preservation in distributed settings, its high computational cost, especially for high-dimensional ML models, has made it challenging to design robust (non-linear) Byzantine-resilient algorithms using HE.
Antoine Choffrut, Rachid Guerraoui, Rafael Pinot, Renaud Sirdey, John Stephan, Martin Zuber
Middleware5
2023 Fixing by Mixing: A Recipe for Optimal Byzantine ML under Heterogeneity
abstract
Byzantine machine learning (ML) aims to ensure the resilience of distributed learning algorithms to misbehaving (or Byzantine) machines. Although this problem received significant attention, prior works often assume the data held by the machines to be homogeneous, which is seldom true in practical settings. Data heterogeneity makes Byzantine ML considerably more challenging, since a Byzantine machine can hardly be distinguished from a non-Byzantine outlier. A few solutions have been proposed to tackle this issue, but these provide suboptimal probabilistic guarantees and fare poorly in practice. This paper closes the theoretical gap, achieving optimality and inducing good empirical results. In fact, we show how to automatically adapt existing solutions for (homogeneous) Byzantine ML to the heterogeneous setting through a powerful mechanism, we call nearest neighbor mixing (NNM), which boosts any standard robust distributed gradient descent variant to yield optimal Byzantine resilience under heterogeneity. We obtain similar guarantees (in expectation) by plugging NNM in the distributed stochastic heavy ball method, a practical substitute to distributed gradient descent. We obtain empirical results that significantly outperform state-of-the-art Byzantine ML solutions.
Youssef Allouah, Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot, John Stephan
AISTATS6
2023 On the Privacy-Robustness-Utility Trilemma in Distributed Learning
abstract
The ubiquity of distributed machine learning (ML) in sensitive public domain applications calls for algorithms that protect data privacy, while being robust to faults and adversarial behaviors. Although privacy and robustness have been extensively studied independently in distributed ML, their synthesis remains poorly understood. We present the first tight analysis of the error incurred by any algorithm ensuring robustness against a fraction of adversarial machines, as well as differential privacy (DP) for honest machines' data against any other curious entity. Our analysis exhibits a fundamental trade-off between privacy, robustness, and utility. To prove our lower bound, we consider the case of mean estimation, subject to distributed DP and robustness constraints, and devise reductions to centralized estimation of one-way marginals. We prove our matching upper bound by presenting a new distributed ML algorithm using a high-dimensional robust aggregation rule. The latter amortizes the dependence on the dimension in the error (caused by adversarial workers and DP), while being agnostic to the statistical properties of the data.
Youssef Allouah, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot, John Stephan
ICML5
2023 Robust Collaborative Learning with Linear Gradient Overhead
abstract
Collaborative learning algorithms, such as distributed SGD (or D-SGD), are prone to faulty machines that may deviate from their prescribed algorithm because of software or hardware bugs, poisoned data or malicious behaviors. While many solutions have been proposed to enhance the robustness of D-SGD to such machines, previous works either resort to strong assumptions (trusted server, homogeneous data, specific noise model) or impose a gradient computational cost that is several orders of magnitude higher than that of D-SGD. We present MoNNA, a new algorithm that (a) is provably robust under standard assumptions and (b) has a gradient computation overhead that is linear in the fraction of faulty machines, which is conjectured to be tight. Essentially, MoNNA uses Polyak’s momentum of local gradients for local updates and nearest-neighbor averaging (NNA) for global mixing, respectively. While MoNNA is rather simple to implement, its analysis has been more challenging and relies on two key elements that may be of independent interest. Specifically, we introduce the mixing criterion of $(\alpha, \lambda)$-reduction to analyze the non-linear mixing of non-faulty machines, and present a way to control the tension between the momentum and the model drifts. We validate our theory by experiments on image classification and make our code available at https://github.com/LPD-EPFL/robust-collaborative-learning.
Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Lê-Nguyên Hoang, Rafael Pinot, John Stephan
ICML6
2022 Byzantine Machine Learning Made Easy By Resilient Averaging of Momentums
abstract
Byzantine resilience emerged as a prominent topic within the distributed machine learning community. Essentially, the goal is to enhance distributed optimization algorithms, such as distributed SGD, in a way that guarantees convergence despite the presence of some misbehaving (a.k.a., Byzantine) workers. Although a myriad of techniques addressing the problem have been proposed, the field arguably rests on fragile foundations. These techniques are hard to prove correct and rely on assumptions that are (a) quite unrealistic, i.e., often violated in practice, and (b) heterogeneous, i.e., making it difficult to compare approaches. We present RESAM (RESilient Averaging of Momentums), a unified framework that makes it simple to establish optimal Byzantine resilience, relying only on standard machine learning assumptions. Our framework is mainly composed of two operators: resilient averaging at the server and distributed momentum at the workers. We prove a general theorem stating the convergence of distributed SGD under RESAM. Interestingly, demonstrating and comparing the convergence of many existing techniques become direct corollaries of our theorem, without resorting to stringent assumptions. We also present an empirical evaluation of the practical relevance of RESAM.
Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot, John Stephan
ICML5
2021 Differential Privacy and Byzantine Resilience in SGD: Do They Add Up?
abstract
This paper addresses the problem of combining Byzantine resilience with privacy in machine learning (ML). Specifically, we study if a distributed implementation of the renowned Stochastic Gradient Descent (SGD) learning algorithm is feasible withboth differential privacy (DP) and (α,f)-Byzantine resilience. To the best of our knowledge, this is the first work to tackle this problem from a theoretical point of view. A key finding of our analyses is that the classical approaches to these two (seemingly) orthogonal issues are incompatible. More precisely, we show that a direct composition of these techniques makes the guarantees of the resulting SGD algorithm depend unfavourably upon the number of parameters of the ML model, making the training of large models practically infeasible. We validate our theoretical results through numerical experiments on publicly-available datasets; showing that it is impractical to ensure DP and Byzantine resilience simultaneously.
Rachid Guerraoui, Nirupam Gupta, Rafael Pinot, Sébastien Rouault, John Stephan
PODC5
2019 HideMyApp: Hiding the Presence of Sensitive Apps on Android
Anh Pham, Italo Dacosta, Eleonora Losiouk, John Stephan, Kévin Huguenin, Jean-Pierre Hubaux
USENIX Security Symposium4