Arto Maranjyan

dblp:332/0784 · also Artavazd Maranjyan · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0001-8409-817XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 88% Optimization for machine learning · 12%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
2.632025
Ringmaster ASGD: The First Asynchronous SGD with Optimal Time Complexity · ICML 2025
ATA: Adaptive Task Allocation for Efficient Resource Management in Distributed Machine Learning · ICML 2025
LoCoDL: Communication-Efficient Distributed Learning with Local Training and Compression · ICLR 2025
Machine learning › Efficient and distributed learning › distributed training › asynchronous training
asynchronous stochastic gradient descent
0.912025
Ringmaster ASGD: The First Asynchronous SGD with Optimal Time Complexity · ICML 2025
Machine learning › Efficient and distributed learning › distributed training
asynchronous training
0.912025
ATA: Adaptive Task Allocation for Efficient Resource Management in Distributed Machine Learning · ICML 2025
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
0.912025
LoCoDL: Communication-Efficient Distributed Learning with Local Training and Compression · ICLR 2025
Machine learning › Efficient and distributed learning
federated learning
0.912025
LoCoDL: Communication-Efficient Distributed Learning with Local Training and Compression · ICLR 2025
Machine learning › Optimization for machine learning
stochastic gradient descent
0.912025
Ringmaster ASGD: The First Asynchronous SGD with Optimal Time Complexity · ICML 2025

Methods — techniques the papers use, named apart from their topics

theoretical analysis · 1.7sparsification · 0.9quantization · 0.9local training · 0.9gradient compression · 0.9adaptive task allocation · 0.9
YearPublicationVenuePosition
2025 LoCoDL: Communication-Efficient Distributed Learning with Local Training and Compression
abstract
In $D$istributed optimization and $L$earning, and even more in the modern framework of federated learning, communication, which is slow and costly, is critical. We introduce LoCoDL, a communication-efficient algorithm that leverages the two popular and effective techniques of $Lo$cal training, which reduces the communication frequency, and $Co$mpression, in which short bitstreams are sent instead of full-dimensional vectors of floats. LoCoDL works with a large class of unbiased compressors that includes widely-used sparsification and quantization methods. LoCoDL provably benefits from local training and compression and enjoys a doubly-accelerated communication complexity, with respect to the condition number of the functions and the model dimension, in the general heterogeneous regime with strongly convex functions. This is confirmed in practice, with LoCoDL outperforming existing algorithms.
Laurent Condat, Arto Maranjyan, Peter Richtárik
ICLR2
2025 ATA: Adaptive Task Allocation for Efficient Resource Management in Distributed Machine Learning
abstract
Asynchronous methods are fundamental for parallelizing computations in distributed machine learning. They aim to accelerate training by fully utilizing all available resources. However, their greedy approach can lead to inefficiencies using more computation than required, especially when computation times vary across devices. If the computation times were known in advance, training could be fast and resource-efficient by assigning more tasks to faster workers. The challenge lies in achieving this optimal allocation without prior knowledge of the computation time distributions. In this paper, we propose ATA (Adaptive Task Allocation), a method that adapts to heterogeneous and random distributions of worker computation times. Through rigorous theoretical analysis, we show that ATA identifies the optimal task allocation and performs comparably to methods with prior knowledge of computation times. Experimental results further demonstrate that ATA is resource-efficient, significantly reducing costs compared to the greedy approach, which can be arbitrarily expensive depending on the number of workers.
Arto Maranjyan, El Mehdi Saad, Peter Richtárik, Francesco Orabona
ICML1
2025 Ringmaster ASGD: The First Asynchronous SGD with Optimal Time Complexity
abstract
Asynchronous Stochastic Gradient Descent (Asynchronous SGD) is a cornerstone method for parallelizing learning in distributed machine learning. However, its performance suffers under arbitrarily heterogeneous computation times across workers, leading to suboptimal time complexity and inefficiency as the number of workers scales. While several Asynchronous SGD variants have been proposed, recent findings by Tyurin & Richtárik (NeurIPS 2023) reveal that none achieve optimal time complexity, leaving a significant gap in the literature. In this paper, we propose Ringmaster ASGD, a novel Asynchronous SGD method designed to address these limitations and tame the inherent challenges of Asynchronous SGD. We establish, through rigorous theoretical analysis, that Ringmaster ASGD achieves optimal time complexity under arbitrarily heterogeneous and dynamically fluctuating worker computation times. This makes it the first Asynchronous SGD method to meet the theoretical lower bounds for time complexity in such scenarios.
Arto Maranjyan, Alexander Tiurin, Peter Richtárik
ICML1
2025 MindFlayer SGD: Efficient Parallel SGD in the Presence of Heterogeneous and Random Worker Compute Times
abstract
We investigate the problem of minimizing the expectation of smooth nonconvex functions in a distributed setting with multiple parallel workers that are able to compute stochastic gradients. A significant challenge in this context is the presence of arbitrarily heterogeneous and stochastic compute times among workers, which can severely degrade the performance of existing parallel stochastic gradient descent (SGD) methods. While some parallel SGD algorithms achieve optimal performance under deterministic but heterogeneous delays, their effectiveness diminishes when compute times are random-a scenario not explicitly addressed in their design. To bridge this gap, we introduce MindFlayer SGD, a novel parallel SGD method specifically designed to handle stochastic and heterogeneous compute times. Through theoretical analysis and empirical evaluation, we demonstrate that MindFlayer SGD consistently outperforms existing baselines, particularly in environments with heavy-tailed noise. Our results highlight its robustness and scalability, making it a compelling choice for large-scale distributed learning tasks.
Arto Maranjyan, Omar Shaikh Omar, Peter Richtárik
UAI1