Shi-Shang Wang

dblp:427/0216 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 56% Optimization for machine learning · 44%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › distributed training
asynchronous training
1.012026
Ordered Local Momentum for Asynchronous Distributed Learning Under Arbitrary Delays · AAAI 2026
Machine learning › Optimization for machine learning › stochastic gradient descent
stochastic gradient descent with momentum
1.012026
Ordered Local Momentum for Asynchronous Distributed Learning Under Arbitrary Delays · AAAI 2026
Machine learning › Efficient and distributed learning
local updates
0.312026
Ordered Local Momentum for Asynchronous Distributed Learning Under Arbitrary Delays · AAAI 2026

Methods — techniques the papers use, named apart from their topics

momentum SGD · 1.0local updates · 1.0convergence analysis · 1.0
YearPublicationVenuePosition
2026 Ordered Local Momentum for Asynchronous Distributed Learning Under Arbitrary Delays
abstract
Momentum SGD (MSGD) serves as a foundational optimizer in training deep models due to momentum's key role in accelerating convergence and enhancing generalization. Meanwhile, asynchronous distributed learning is crucial for training large-scale deep models, especially when the computing capabilities of the workers in the cluster are heterogeneous. To reduce communication frequency, local updates are widely adopted in distributed learning. However, how to implement asynchronous distributed MSGD with local updates remains unexplored. To solve this problem, we propose a novel method, called ordered local momentum (OrLoMo), for asynchronous distributed learning. In OrLoMo, each worker runs MSGD locally. Then the local momentum from each worker will be aggregated by the server in order based on its global iteration index. To the best of our knowledge, OrLoMo is the first method to implement asynchronous distributed MSGD with local updates. We prove the convergence of OrLoMo for non-convex problems under arbitrary delays. Experiments validate that OrLoMo can outperform its synchronous counterpart and other asynchronous methods.
Chang-Wei Shi, Shi-Shang Wang, Wu-Jun Li
AAAI2