Samar Aseeri

dblp:238/5574 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2024
0000-0001-6140-2749ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2024 SUARA: A scalable universal allreduce communication algorithm for acceleration of parallel deep learning applications
abstract
Parallel and distributed deep learning (PDNN) has become an effective strategy to reduce the long training times of large-scale deep neural networks. Mainstream PDNN software packages based on the message-passing interface (MPI) and employing synchronous stochastic gradient descent rely crucially on the performance of MPI allreduce collective communication routine. In this work, we propose a novel scalable universal allreduce meta-algorithm called SUARA. In general, SUARA consists of L serial steps, where L≥2, executed by all MPI processes involved in the allreduce operation. At each step, SUARA partitions this set of processes into subsets, which execute optimally selected library allreduce algorithms to solve sub-allreduce problems on these subsets in parallel, to accomplish the whole allreduce operation after completing all the L steps. We then design, theoretically study and implement a two-step SUARA (L=2) called SUARA2 on top of the Open MPI library. We prove that the theoretical asymptotic speedup of SUARA2 executed by P processes over the base Open MPI routine is O(P). Our experiments on Shaheen-II supercomputer employing 1024 nodes demonstrate over 2x speedup of SUARA2 over native Open MPI allreduce routine, which translates into the performance improvement of training of ResNet-50 DNN on ImageNet by 9%.
Emin Nuriyev, Ravi Reddy, Samar Aseeri, Mahendra K. Verma, Alexey L. Lastovetsky
J. Parallel Distributed Comput.3
2023 A scheduling policy to save 10% of communication time in parallel fast Fourier transform
abstract
Summary The fast Fourier transform (FFT) has applications in almost every frequency related study, for example, in image and signal processing, and radio astronomy. It is also used as a Poisson operator inversion kernel in partial differential equations in fluid flows, in density functional theory, many‐body theory, and others. The three‐dimensional FFT has large time complexity . Hence, parallelization is used to compute such FFTs. Popular libraries perform slab division or pencil decomposition of data. None of the existing libraries achieve perfect inverse scaling of time with cores because FFT requires all‐to‐all communication and clusters hitherto do not have physical all‐to‐all connections. Dragonfly, one of the popular topologies for the interconnect, supports hierarchical connections among the components. We show that if we align the all‐to‐all communication of FFT with the physical connections of Dragonfly topology we will achieve a better scaling and reduce communication time.
Samar Aseeri, Anando Gopal Chatterjee, Mahendra K. Verma, David E. Keyes
Concurr. Comput. Pract. Exp.1
2023 mpi4py.futures: MPI-Based Asynchronous Task Execution for Python
abstract
We present mpi4py.futures, a lightweight, asynchronous task execution framework targeting the Python programming language and using the Message Passing Interface (MPI) for interprocess communication. mpi4py.futures follows the interface of the concurrent.futures package from the Python standard library and can be used as its drop-in replacement, while allowing applications to scale over multiple compute nodes. We discuss the design, implementation, and feature set of mpi4py.futures and compare its performance to other solutions on both shared and distributed memory architectures. On a shared-memory system, we show mpi4py.futures to consistently outperform Python's concurrent.futures with speedup ratios between 1.4X and 3.7X in throughput (tasks per second) and between 1.9X and 2.9X in bandwidth. On a Cray XC40 system, we compare mpi4py.futures to Dask – a well-known Python parallel computing package. Although we note more varied results, we show mpi4py.futures to outperform Dask in most scenarios.
Marcin Rogowski, Samar Aseeri, David E. Keyes, Lisandro Dalcín
IEEE Trans. Parallel Distributed Syst.2