VLDB 2026 Research / reviewers in the wild / expert
John Stephan
dblp:185/5372
· DBLP profile ↗
9ranked-venue papers
0as first author
8since 2021 · last 2025
0009-0007-0293-1967ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Efficient and distributed learning · 50% Trustworthy machine learning · 28% Optimization for machine learning · 22% | |
| Network and information security
3 papers |
Privacy and data protection · 89% Web and mobile security · 11% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 100% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
federated learning |
1.7 | 2 | 2025 | Towards Trustworthy Federated Learning with Untrusted Participants · ICML 2025 Adaptive Gradient Clipping for Robust Federated Learning · ICLR 2025 |
Machine learning › Trustworthy machine learning › robustness
byzantine robustness |
1.7 | 3 | 2023 | Robust Collaborative Learning with Linear Gradient Overhead · ICML 2023 Byzantine Machine Learning Made Easy By Resilient Averaging of Momentums · ICML 2022 Differential Privacy and Byzantine Resilience in SGD: Do They Add Up? · PODC 2021 |
Machine learning › Trustworthy machine learning
robustness |
1.5 | 2 | 2025 | Towards Trustworthy Federated Learning with Untrusted Participants · ICML 2025 On the Privacy-Robustness-Utility Trilemma in Distributed Learning · ICML 2023 |
Privacy and data protection
differential privacy |
1.5 | 2 | 2025 | Towards Trustworthy Federated Learning with Untrusted Participants · ICML 2025 On the Privacy-Robustness-Utility Trilemma in Distributed Learning · ICML 2023 |
Machine learning › Efficient and distributed learning
distributed training |
1.4 | 3 | 2023 | Robust Collaborative Learning with Linear Gradient Overhead · ICML 2023 Byzantine Machine Learning Made Easy By Resilient Averaging of Momentums · ICML 2022 Differential Privacy and Byzantine Resilience in SGD: Do They Add Up? · PODC 2021 |
Machine learning › Optimization for machine learning › adaptive optimization
adaptive gradient clipping |
0.9 | 1 | 2025 | Adaptive Gradient Clipping for Robust Federated Learning · ICLR 2025 |
Machine learning › Optimization for machine learning
gradient clipping |
0.9 | 1 | 2025 | Adaptive Gradient Clipping for Robust Federated Learning · ICLR 2025 |
Machine learning › Efficient and distributed learning › federated learning
robust federated learning |
0.9 | 1 | 2025 | Adaptive Gradient Clipping for Robust Federated Learning · ICLR 2025 |
Machine learning › Efficient and distributed learning › distributed training
robust gradient aggregation |
0.9 | 1 | 2025 | Towards Trustworthy Federated Learning with Untrusted Participants · ICML 2025 |
Machine learning › Efficient and distributed learning › federated learning
trustworthy federated learning |
0.9 | 1 | 2025 | Towards Trustworthy Federated Learning with Untrusted Participants · ICML 2025 |
Privacy and data protection › differential privacy › noise addition
correlated noise mechanisms |
0.9 | 1 | 2025 | Towards Trustworthy Federated Learning with Untrusted Participants · ICML 2025 |
Privacy and data protection › differential privacy
distributed differential privacy |
0.7 | 1 | 2023 | On the Privacy-Robustness-Utility Trilemma in Distributed Learning · ICML 2023 |
Distributed systems
fault tolerance |
0.7 | 1 | 2023 | Robust Collaborative Learning with Linear Gradient Overhead · ICML 2023 |
Machine learning › Optimization for machine learning
distributed optimization |
0.6 | 1 | 2022 | Byzantine Machine Learning Made Easy By Resilient Averaging of Momentums · ICML 2022 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.1 | 1 | 2021 | Differential Privacy and Byzantine Resilience in SGD: Do They Add Up? · PODC 2021 |
Methods — techniques the papers use, named apart from their topics
distributed SGD · 1.9shared randomness · 1.7correlated noise · 1.7robust aggregation · 1.3polyak momentum · 1.3nearest-neighbor averaging · 1.3mean estimation · 1.3lower bound analysis · 1.3resilient averaging · 0.6momentum · 0.6app hiding · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptive Gradient Clipping for Robust Federated LearningabstractRobust federated learning aims to maintain reliable performance despite the presence of adversarial or misbehaving workers. While state-of-the-art (SOTA) robust distributed gradient descent (Robust-DGD) methods were proven theoretically optimal, their empirical success has often relied on pre-aggregation gradient clipping.
However, existing static clipping strategies yield inconsistent results: enhancing robustness against some attacks while being ineffective or even detrimental against others.
To address this limitation, we propose a principled adaptive clipping strategy, Adaptive Robust Clipping (ARC), which dynamically adjusts clipping thresholds based on the input gradients. We prove that ARC not only preserves the theoretical robustness guarantees of SOTA Robust-DGD methods but also provably improves asymptotic convergence when the model is well-initialized. Extensive experiments on benchmark image classification tasks confirm these theoretical insights, demonstrating that ARC significantly enhances robustness, particularly in highly heterogeneous and adversarial settings. Youssef Allouah, Rachid Guerraoui, Nirupam Gupta, Ahmed Jellouli, Geovani Rizk, John Stephan |
ICLR | 6 |
| 2025 | Towards Trustworthy Federated Learning with Untrusted ParticipantsabstractResilience against malicious participants and data privacy are essential for trustworthy federated learning, yet achieving both with good utility typically requires the strong assumption of a trusted central server. This paper shows that a significantly weaker assumption suffices: each pair of participants shares a randomness seed unknown to others.
In a setting where malicious participants may collude with an untrusted server, we propose CafCor, an algorithm that integrates robust gradient aggregation with correlated noise injection, using shared randomness between participants.
We prove that CafCor achieves strong privacy-utility trade-offs, significantly outperforming local differential privacy (DP) methods, which do not make any trust assumption, while approaching central DP utility, where the server is fully trusted.
Empirical results on standard benchmarks validate CafCor's practicality, showing that privacy and robustness can coexist in distributed systems without sacrificing utility or trusting the server. Youssef Allouah, Rachid Guerraoui, John Stephan |
ICML | 3 |
| 2024 | Towards Practical Homomorphic Aggregation in Byzantine-Resilient Distributed LearningabstractThe growing availability of distributed data has led to the increased use of machine learning (ML) algorithms in distributed topologies, where multiple nodes collaborate to train models under the coordination of a central server. However, distributed learning faces two significant challenges: the risk of Byzantine nodes corrupting the learning process by sending incorrect information, and the potential for a curious server to violate the privacy of individual nodes, even reconstructing their private data. While homomorphic encryption (HE) has been a promising solution for privacy preservation in distributed settings, its high computational cost, especially for high-dimensional ML models, has made it challenging to design robust (non-linear) Byzantine-resilient algorithms using HE. Antoine Choffrut, Rachid Guerraoui, Rafael Pinot, Renaud Sirdey, John Stephan, Martin Zuber |
Middleware | 5 |
| 2023 | Fixing by Mixing: A Recipe for Optimal Byzantine ML under HeterogeneityabstractByzantine machine learning (ML) aims to ensure the resilience of distributed learning algorithms to misbehaving (or Byzantine) machines. Although this problem received significant attention, prior works often assume the data held by the machines to be homogeneous, which is seldom true in practical settings. Data heterogeneity makes Byzantine ML considerably more challenging, since a Byzantine machine can hardly be distinguished from a non-Byzantine outlier. A few solutions have been proposed to tackle this issue, but these provide suboptimal probabilistic guarantees and fare poorly in practice. This paper closes the theoretical gap, achieving optimality and inducing good empirical results. In fact, we show how to automatically adapt existing solutions for (homogeneous) Byzantine ML to the heterogeneous setting through a powerful mechanism, we call nearest neighbor mixing (NNM), which boosts any standard robust distributed gradient descent variant to yield optimal Byzantine resilience under heterogeneity. We obtain similar guarantees (in expectation) by plugging NNM in the distributed stochastic heavy ball method, a practical substitute to distributed gradient descent. We obtain empirical results that significantly outperform state-of-the-art Byzantine ML solutions. Youssef Allouah, Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot, John Stephan |
AISTATS | 6 |
| 2023 | On the Privacy-Robustness-Utility Trilemma in Distributed LearningabstractThe ubiquity of distributed machine learning (ML) in sensitive public domain applications calls for algorithms that protect data privacy, while being robust to faults and adversarial behaviors. Although privacy and robustness have been extensively studied independently in distributed ML, their synthesis remains poorly understood. We present the first tight analysis of the error incurred by any algorithm ensuring robustness against a fraction of adversarial machines, as well as differential privacy (DP) for honest machines' data against any other curious entity. Our analysis exhibits a fundamental trade-off between privacy, robustness, and utility. To prove our lower bound, we consider the case of mean estimation, subject to distributed DP and robustness constraints, and devise reductions to centralized estimation of one-way marginals. We prove our matching upper bound by presenting a new distributed ML algorithm using a high-dimensional robust aggregation rule. The latter amortizes the dependence on the dimension in the error (caused by adversarial workers and DP), while being agnostic to the statistical properties of the data. Youssef Allouah, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot, John Stephan |
ICML | 5 |
| 2023 | Robust Collaborative Learning with Linear Gradient OverheadabstractCollaborative learning algorithms, such as distributed SGD (or D-SGD), are prone to faulty machines that may deviate from their prescribed algorithm because of software or hardware bugs, poisoned data or malicious behaviors. While many solutions have been proposed to enhance the robustness of D-SGD to such machines, previous works either resort to strong assumptions (trusted server, homogeneous data, specific noise model) or impose a gradient computational cost that is several orders of magnitude higher than that of D-SGD. We present MoNNA, a new algorithm that (a) is provably robust under standard assumptions and (b) has a gradient computation overhead that is linear in the fraction of faulty machines, which is conjectured to be tight. Essentially, MoNNA uses Polyak’s momentum of local gradients for local updates and nearest-neighbor averaging (NNA) for global mixing, respectively. While MoNNA is rather simple to implement, its analysis has been more challenging and relies on two key elements that may be of independent interest. Specifically, we introduce the mixing criterion of $(\alpha, \lambda)$-reduction to analyze the non-linear mixing of non-faulty machines, and present a way to control the tension between the momentum and the model drifts. We validate our theory by experiments on image classification and make our code available at https://github.com/LPD-EPFL/robust-collaborative-learning. Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Lê-Nguyên Hoang, Rafael Pinot, John Stephan |
ICML | 6 |
| 2022 | Byzantine Machine Learning Made Easy By Resilient Averaging of MomentumsabstractByzantine resilience emerged as a prominent topic within the distributed machine learning community. Essentially, the goal is to enhance distributed optimization algorithms, such as distributed SGD, in a way that guarantees convergence despite the presence of some misbehaving (a.k.a., Byzantine) workers. Although a myriad of techniques addressing the problem have been proposed, the field arguably rests on fragile foundations. These techniques are hard to prove correct and rely on assumptions that are (a) quite unrealistic, i.e., often violated in practice, and (b) heterogeneous, i.e., making it difficult to compare approaches. We present RESAM (RESilient Averaging of Momentums), a unified framework that makes it simple to establish optimal Byzantine resilience, relying only on standard machine learning assumptions. Our framework is mainly composed of two operators: resilient averaging at the server and distributed momentum at the workers. We prove a general theorem stating the convergence of distributed SGD under RESAM. Interestingly, demonstrating and comparing the convergence of many existing techniques become direct corollaries of our theorem, without resorting to stringent assumptions. We also present an empirical evaluation of the practical relevance of RESAM. Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot, John Stephan |
ICML | 5 |
| 2021 | Differential Privacy and Byzantine Resilience in SGD: Do They Add Up?abstractThis paper addresses the problem of combining Byzantine resilience with privacy in machine learning (ML). Specifically, we study if a distributed implementation of the renowned Stochastic Gradient Descent (SGD) learning algorithm is feasible withboth differential privacy (DP) and (α,f)-Byzantine resilience. To the best of our knowledge, this is the first work to tackle this problem from a theoretical point of view. A key finding of our analyses is that the classical approaches to these two (seemingly) orthogonal issues are incompatible. More precisely, we show that a direct composition of these techniques makes the guarantees of the resulting SGD algorithm depend unfavourably upon the number of parameters of the ML model, making the training of large models practically infeasible. We validate our theoretical results through numerical experiments on publicly-available datasets; showing that it is impractical to ensure DP and Byzantine resilience simultaneously. Rachid Guerraoui, Nirupam Gupta, Rafael Pinot, Sébastien Rouault, John Stephan |
PODC | 5 |
| 2019 | HideMyApp: Hiding the Presence of Sensitive Apps on Android
Anh Pham, Italo Dacosta, Eleonora Losiouk, John Stephan, Kévin Huguenin, Jean-Pierre Hubaux |
USENIX Security Symposium | 4 |