EDBT 2026 Demo / reviewers in the wild / expert
Yi-Rui Yang
dblp:260/0404
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Optimization for machine learning · 39% Efficient and distributed learning · 37% Trustworthy machine learning · 24% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
2.7 | 4 | 2024 | Ordered Momentum for Asynchronous SGD · NeurIPS 2024 On the Effect of Batch Size in Byzantine-Robust Distributed Learning · ICLR 2024 Buffered Asynchronous SGD for Byzantine Learning · J. Mach. Learn. Res. 2023 |
Machine learning › Trustworthy machine learning
robustness |
2.0 | 3 | 2025 | On the Tension between Byzantine Robustness and No-Attack Accuracy in Distributed Learning · ICML 2025 Buffered Asynchronous SGD for Byzantine Learning · J. Mach. Learn. Res. 2023 BASGD: Buffered Asynchronous SGD for Byzantine Learning · ICML 2021 |
Machine learning › Optimization for machine learning
convergence analysis |
1.6 | 2 | 2025 | On the Tension between Byzantine Robustness and No-Attack Accuracy in Distributed Learning · ICML 2025 Ordered Momentum for Asynchronous SGD · NeurIPS 2024 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
1.5 | 2 | 2024 | Ordered Momentum for Asynchronous SGD · NeurIPS 2024 On the Effect of Batch Size in Byzantine-Robust Distributed Learning · ICLR 2024 |
Machine learning › Optimization for machine learning › stochastic gradient descent
asynchronous SGD |
1.3 | 2 | 2024 | Ordered Momentum for Asynchronous SGD · NeurIPS 2024 BASGD: Buffered Asynchronous SGD for Byzantine Learning · ICML 2021 |
Machine learning › Trustworthy machine learning › robustness
byzantine robustness |
1.2 | 2 | 2023 | Buffered Asynchronous SGD for Byzantine Learning · J. Mach. Learn. Res. 2023 BASGD: Buffered Asynchronous SGD for Byzantine Learning · ICML 2021 |
Machine learning › Optimization for machine learning › convergence guarantees
gradient descent convergence |
0.9 | 1 | 2025 | On the Tension between Byzantine Robustness and No-Attack Accuracy in Distributed Learning · ICML 2025 |
Machine learning › Efficient and distributed learning › federated learning › model aggregation
robust aggregation |
0.9 | 1 | 2025 | On the Tension between Byzantine Robustness and No-Attack Accuracy in Distributed Learning · ICML 2025 |
Distributed systems › fault tolerance
byzantine fault tolerance |
0.9 | 1 | 2025 | On the Tension between Byzantine Robustness and No-Attack Accuracy in Distributed Learning · ICML 2025 |
Distributed systems
distributed machine learning |
0.9 | 1 | 2025 | On the Tension between Byzantine Robustness and No-Attack Accuracy in Distributed Learning · ICML 2025 |
Machine learning › Efficient and distributed learning › distributed training › robust distributed learning
byzantine-robust learning |
0.8 | 1 | 2024 | On the Effect of Batch Size in Byzantine-Robust Distributed Learning · ICLR 2024 |
Machine learning › Efficient and distributed learning › distributed training › asynchronous training
asynchronous stochastic gradient descent |
0.7 | 1 | 2023 | Buffered Asynchronous SGD for Byzantine Learning · J. Mach. Learn. Res. 2023 |
Methods — techniques the papers use, named apart from their topics
convergence analysis · 3.3robust aggregators · 1.7ordered momentum · 0.8normalized momentum · 0.8delay-adaptive learning rate · 0.8momentum · 0.7buffered asynchronous SGD · 0.7aggregation rule · 0.7stochastic gradient descent · 0.5buffered asynchronous update · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Tension between Byzantine Robustness and No-Attack Accuracy in Distributed LearningabstractByzantine-robust distributed learning (BRDL), which refers to distributed learning that can work with potential faulty or malicious workers (also known as Byzantine workers), has recently attracted much research attention. Robust aggregators are widely used in existing BRDL methods to obtain robustness against Byzantine workers. However, Byzantine workers do not always exist in applications. As far as we know, there is almost no existing work theoretically investigating the effect of using robust aggregators when there are no Byzantine workers. To bridge this knowledge gap, we theoretically analyze the aggregation error for robust aggregators when there are no Byzantine workers. Specifically, we show that the worst-case aggregation error without Byzantine workers increases with the increase of the number of Byzantine workers that a robust aggregator can tolerate. The theoretical result reveals the tension between Byzantine robustness and no-attack accuracy, which refers to accuracy without faulty workers and malicious workers in this paper. Furthermore, we provide lower bounds for the convergence rate of gradient descent with robust aggregators for non-convex objective functions and objective functions that satisfy the Polyak-Lojasiewicz (PL) condition, respectively. We also prove the tightness of the lower bounds. The lower bounds for convergence rate reveal similar tension between Byzantine robustness and no-attack accuracy. Empirical results further support our theoretical findings. Yi-Rui Yang, Chang-Wei Shi, Wu-Jun Li |
ICML | 1 |
| 2024 | On the Effect of Batch Size in Byzantine-Robust Distributed LearningabstractByzantine-robust distributed learning (BRDL), in which computing devices are likely to behave abnormally due to accidental failures or malicious attacks, has recently become a hot research topic. However, even in the independent and identically distributed (i.i.d.) case, existing BRDL methods will suffer a significant drop on model accuracy due to the large variance of stochastic gradients. Increasing batch sizes is a simple yet effective way to reduce the variance. However, when the total number of gradient computation is fixed, a too-large batch size will lead to a too-small iteration number (update number), which may also degrade the model accuracy. In view of this challenge, we mainly study the effect of batch size when the total number of gradient computation is fixed in this work. In particular, we show that when the total number of gradient computation is fixed, the optimal batch size corresponding to the tightest theoretical upper bound in BRDL increases with the fraction of Byzantine workers. Therefore, compared to the case without attacks, a larger batch size is preferred when under Byzantine attacks. Motivated by the theoretical finding, we propose a novel method called Byzantine-robust stochastic gradient descent with normalized momentum (ByzSGDnm) in order to further increase model accuracy in BRDL. We theoretically prove the convergence of ByzSGDnm for general non-convex cases under Byzantine attacks. Empirical results show that when under Byzantine attacks, compared to the cases of small batch sizes, setting a relatively large batch size can significantly increase the model accuracy, which is consistent with our theoretical results. Moreover, ByzSGDnm can achieve higher model accuracy than existing BRDL methods when under deliberately crafted attacks. In addition, we empirically show that increasing batch sizes has the bonus of training acceleration. Yi-Rui Yang, Chang-Wei Shi, Wu-Jun Li |
ICLR | 1 |
| 2024 | Ordered Momentum for Asynchronous SGDabstractDistributed learning is essential for training large-scale deep models.
Asynchronous SGD (ASGD) and its variants are commonly used distributed learning methods, particularly in scenarios where the computing capabilities of workers in the cluster are heterogeneous.
Momentum has been acknowledged for its benefits in both optimization and generalization in deep model training. However, existing works have found that naively incorporating momentum into ASGD can impede the convergence.
In this paper, we propose a novel method called ordered momentum (OrMo) for ASGD. In OrMo, momentum is incorporated into ASGD by organizing the gradients in order based on their iteration indexes. We theoretically prove the convergence of OrMo with both constant and delay-adaptive learning rates for non-convex problems. To the best of our knowledge, this is the first work to establish the convergence analysis of ASGD with momentum without dependence on the maximum delay. Empirical results demonstrate that OrMo can achieve better convergence performance compared with ASGD and other asynchronous methods with momentum. Chang-Wei Shi, Yi-Rui Yang, Wu-Jun Li |
NeurIPS | 2 |
| 2023 | Buffered Asynchronous SGD for Byzantine LearningabstractDistributed learning has become a hot research topic due to its wide application in cluster-based large-scale learning, federated learning, edge computing, and so on. Most traditional distributed learning methods typically assume no failure or attack. However, many unexpected cases, such as communication failure and even malicious attack, may happen in real applications. Hence, Byzantine learning (BL), which refers to distributed learning with failure or attack, has recently attracted much attention. Most existing BL methods are synchronous, which are impractical in some applications due to heterogeneous or offline workers. In these cases, asynchronous BL (ABL) is usually preferred. In this paper, we propose a novel method, called buffered asynchronous stochastic gradient descent (BASGD), for ABL. To the best of our knowledge, BASGD is the first ABL method that can resist non-omniscient attacks without storing any instances on the server. Furthermore, we also propose an improved variant of BASGD, called BASGD with momentum (BASGDm), by introducing local momentum into BASGD. Compared with those methods which need to store instances on server, BASGD and BASGDm have a wider scope of application. Both BASGD and BASGDm are compatible with various aggregation rules. Moreover, both BASGD and BASGDm are proved to be convergent and able to resist failure or attack. Empirical results show that our methods significantly outperform existing ABL baselines when there exists failure or attack on workers. Yi-Rui Yang, Wu-Jun Li |
J. Mach. Learn. Res. | 1 |
| 2021 | BASGD: Buffered Asynchronous SGD for Byzantine LearningabstractDistributed learning has become a hot research topic due to its wide application in cluster-based large-scale learning, federated learning, edge computing and so on. Most traditional distributed learning methods typically assume no failure or attack. However, many unexpected cases, such as communication failure and even malicious attack, may happen in real applications. Hence, Byzantine learning (BL), which refers to distributed learning with failure or attack, has recently attracted much attention. Most existing BL methods are synchronous, which are impractical in some applications due to heterogeneous or offline workers. In these cases, asynchronous BL (ABL) is usually preferred. In this paper, we propose a novel method, called buffered asynchronous stochastic gradient descent (BASGD), for ABL. To the best of our knowledge, BASGD is the first ABL method that can resist malicious attack without storing any instances on server. Compared with those methods which need to store instances on server, BASGD has a wider scope of application. BASGD is proved to be convergent, and be able to resist failure or attack. Empirical results show that BASGD significantly outperforms vanilla asynchronous stochastic gradient descent (ASGD) and other ABL baselines when there exists failure or attack on workers. Yi-Rui Yang, Wu-Jun Li |
ICML | 1 |