Adel Nabli

dblp:269/9664 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0003-3180-5445ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
3 papers
Mathematical optimization · 98% Graph algorithms and graph theory · 2%
Artificial intelligence
3 papers
Efficient and distributed learning · 82% Reinforcement learning · 18%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Distributed systems · 82% High-performance computing · 10% Parallel and multicore computing · 8%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
1.522025
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training · NeurIPS 2025
A2CiD2: Accelerating Asynchronous Communication in Decentralized Deep Learning · NeurIPS 2023
Machine learning › Efficient and distributed learning
distributed training
1.522025
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training · NeurIPS 2025
A2CiD2: Accelerating Asynchronous Communication in Decentralized Deep Learning · NeurIPS 2023
Mathematical optimization
distributed optimization
1.522025
Decentralized Asynchronous Optimization with DADAO allows Decoupling and Acceleration · J. Mach. Learn. Res. 2025
DADAO: Decoupled Accelerated Decentralized Asynchronous Optimization · ICML 2023
Mathematical optimization › continuous optimization › convex optimization
first-order methods
1.522025
Decentralized Asynchronous Optimization with DADAO allows Decoupling and Acceleration · J. Mach. Learn. Res. 2025
DADAO: Decoupled Accelerated Decentralized Asynchronous Optimization · ICML 2023
Distributed systems
gossip protocols
0.922025
DADAO: Decoupled Accelerated Decentralized Asynchronous Optimization · ICML 2023
Decentralized Asynchronous Optimization with DADAO allows Decoupling and Acceleration · J. Mach. Learn. Res. 2025
Machine learning › Efficient and distributed learning › distributed training
gradient aggregation
0.912025
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training · NeurIPS 2025
Mathematical optimization
primal-dual method
0.912025
Decentralized Asynchronous Optimization with DADAO allows Decoupling and Acceleration · J. Mach. Learn. Res. 2025
Distributed systems › distributed communication
decentralized communication
0.712023
DADAO: Decoupled Accelerated Decentralized Asynchronous Optimization · ICML 2023
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
accelerated gradient methods
0.712023
DADAO: Decoupled Accelerated Decentralized Asynchronous Optimization · ICML 2023
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.412020
Curriculum learning for multilevel budgeted combinatorial problems · NeurIPS 2020
Machine learning › Reinforcement learning
value-based reinforcement learning
0.412020
Curriculum learning for multilevel budgeted combinatorial problems · NeurIPS 2020
Mathematical optimization
combinatorial optimization
0.412020
Curriculum learning for multilevel budgeted combinatorial problems · NeurIPS 2020
Distributed systems
consensus
0.312025
Decentralized Asynchronous Optimization with DADAO allows Decoupling and Acceleration · J. Mach. Learn. Res. 2025
High-performance computing
large-scale training
0.312025
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training · NeurIPS 2025
Distributed systems › operating system support › interprocess communication
asynchronous communication
0.212023
A2CiD2: Accelerating Asynchronous Communication in Decentralized Deep Learning · NeurIPS 2023
Parallel and multicore computing
parallel programming models
0.212023
A2CiD2: Accelerating Asynchronous Communication in Decentralized Deep Learning · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

poisson point process · 3.1optimizer state sharding · 1.7laplacian matrix · 1.7delayed gradient synchronization · 1.7SDP relaxation · 1.7laplacian matrix analysis · 1.3gradient compression · 1.3asynchronous SGD · 1.3multi-agent reinforcement learning · 0.9graph neural network · 0.9curriculum learning · 0.9
YearPublicationVenuePosition
2025 ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
abstract
Training LLMs relies on distributed implementations using multiple GPUs to compute gradients in parallel with sharded optimizers. However, synchronizing gradients in data parallel setups introduces communication overhead that grows with the number of workers, limiting parallelization efficiency. Local optimization algorithms reduce communications but incur high memory costs as they prevent optimizer state sharding, hindering scalability. To address this, we propose $\textbf{AC}$cumulate while $\textbf{CO}$mmunicate ($\texttt{ACCO}$), a memory-efficient optimization algorithm for distributed LLM training. By synchronizing delayed gradients while computing new ones, $\texttt{ACCO}$ reduces GPU idle time and supports heterogeneous hardware. To mitigate the convergence issues caused by delayed updates, we introduce a novel technique ensuring training dynamics align with standard distributed optimization. Compared to ZeRO-1, our approach is significantly faster and scales effectively across heterogeneous hardware.
Adel Nabli, Louis Fournier, Pierre Erbacher, Louis Serrano, Eugene Belilovsky, Edouard Oyallon
NeurIPS1
2025 Decentralized Asynchronous Optimization with DADAO allows Decoupling and Acceleration
abstract
DADAO is the first decentralized, accelerated, asynchronous, primal, first-order algorithm to minimize a sum of $L$-smooth and $\mu$-strongly convex functions distributed over a network of size $n$. Modeling the gradient updates and gossip communication procedures with separate independent Poisson Point Processes allows us to decouple the computation and communication steps, which can be run in parallel, while making the whole approach completely asynchronous. This leads to communication acceleration compared to synchronous approaches. Our method employs primal gradients and avoids using a multi-consensus inner loop and other ad-hoc mechanisms. By relating the smallest positive eigenvalue $1/\chi_1$ of the Laplacian matrix $\Lambda$ and the maximal resistance $\chi_2\leq \chi_1$ of the graph to a sufficient minimal communication rate, we show that DADAO requires $\mathcal{O}(n\sqrt{\frac{L}{\mu}}\log(\frac{1}{\epsilon}))$ local gradients and only $\mathcal{O}(\sqrt{\chi_1\chi_2}\operatorname{Tr}\Lambda\sqrt{\frac{L}{\mu}}\log(\frac{1}{\epsilon}))$ communications to reach $\epsilon$-precision, up to logarithmic terms. Thus, we simultaneously obtain an accelerated rate for computations and communications, leading to an improvement over state-of-the-art works, our simulations further validating the strength of our relatively unconstrained method. Moreover, we propose a SDP relaxation to find the gossip rate of each edge minimizing the total number of communications for a given graph, resulting in faster convergence compared to standard approaches relying on uniform communication weights.
Adel Nabli, Edouard Oyallon
J. Mach. Learn. Res.1
2023 DADAO: Decoupled Accelerated Decentralized Asynchronous Optimization
abstract
This work introduces DADAO: the first decentralized, accelerated, asynchronous, primal, first-order algorithm to minimize a sum of $L$-smooth and $\mu$-strongly convex functions distributed over a given network of size $n$. Our key insight is based on modeling the local gradient updates and gossip communication procedures with separate independent Poisson Point Processes. This allows us to decouple the computation and communication steps, which can be run in parallel, while making the whole approach completely asynchronous. This leads to communication acceleration compared to synchronous approaches. Our new method employs primal gradients and does not use a multi-consensus inner loop nor other ad-hoc mechanisms such as Error Feedback, Gradient Tracking, or a Proximal operator. By relating the inverse of the smallest positive eigenvalue of the Laplacian matrix $\chi_1$ and the maximal resistance $\chi_2\leq \chi_1$ of the graph to a sufficient minimal communication rate between the nodes of the network, we show that our algorithm requires $\mathcal{O}(n\sqrt{\frac{L}{\mu}}\log(\frac{1}{\epsilon}))$ local gradients and only $\mathcal{O}(n\sqrt{\chi_1\chi_2}\sqrt{\frac{L}{\mu}}\log(\frac{1}{\epsilon}))$ communications to reach a precision $\epsilon$, up to logarithmic terms. Thus, we simultaneously obtain an accelerated rate for both computations and communications, leading to an improvement over state-of-the-art works, our simulations further validating the strength of our relatively unconstrained method.
Adel Nabli, Edouard Oyallon
ICML1
2023 A2CiD2: Accelerating Asynchronous Communication in Decentralized Deep Learning
Adel Nabli, Eugene Belilovsky, Edouard Oyallon
NeurIPS1
2022 Speech Sequence Embeddings using Nearest Neighbors Contrastive Learning
abstract
International audience
Robin Algayres, Adel Nabli, Benoît Sagot, Emmanuel Dupoux
INTERSPEECH2
2022 Complexity of the multilevel critical node problem
Adel Nabli, Margarida Carvalho, Pierre Hosteins
J. Comput. Syst. Sci.1
2020 Curriculum learning for multilevel budgeted combinatorial problems
abstract
Learning heuristics for combinatorial optimization problems through graph neural networks have recently shown promising results on some classic NP-hard problems. These are single-level optimization problems with only one player. Multilevel combinatorial optimization problems are their generalization, encompassing situations with multiple players taking decisions sequentially. By framing them in a multi-agent reinforcement learning setting, we devise a value-based method to learn to solve multilevel budgeted combinatorial problems involving two players in a zero-sum game over a graph. Our framework is based on a simple curriculum: if an agent knows how to estimate the value of instances with budgets up to $B$, then solving instances with budget $B+1$ can be done in polynomial time regardless of the direction of the optimization by checking the value of every possible afterstate. Thus, in a bottom-up approach, we generate datasets of heuristically solved instances with increasingly larger budgets to train our agent. We report results close to optimality on graphs up to $100$ nodes and a $185 \times$ speedup on average compared to the quickest exact solver known for the Multilevel Critical Node problem, a max-min-max trilevel problem that has been shown to be at least $\Sigma_2^p$-hard.
Adel Nabli, Margarida Carvalho
NeurIPS1