Muhammed Abdullah Soyturk

dblp:408/2470 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0000-0002-2880-0857ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 50% Deep learning architectures and training · 50%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 100%

Topics — the 1 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › large-scale training
large-scale distributed training
0.312025
Balanced and Elastic End-to-end Training of Dynamic LLMs · SC 2025

Methods — techniques the papers use, named apart from their topics

gradual pruning · 1.7early exit · 1.7dynamic sparse attention · 1.7mixture-of-experts · 0.9mixture-of-depths · 0.9mixture of experts · 0.9mixture of depths · 0.9
YearPublicationVenuePosition
2025 Balanced and Elastic End-to-end Training of Dynamic LLMs
abstract
To reduce the computational and memory overhead of Large Language Models, various approaches have been proposed. These include a) Mixture of Experts (MoEs), where token routing affects compute balance; b) gradual pruning of model parameters; c) dynamically freezing layers; d) dynamic sparse attention mechanisms; e) early exit of tokens as they pass through model layers; and f) Mixture of Depths (MoDs), where tokens bypass certain blocks. While these approaches are effective in reducing overall computation, they often introduce significant workload imbalance across workers. In many cases, this imbalance is severe enough to render the techniques impractical for large-scale distributed training, limiting their applicability to toy models due to poor efficiency.
Mohamed Wahib, Muhammed Abdullah Soyturk, Didem Unat
SC2