Majid Salimi Beni

dblp:285/6169 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-8634-7712ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Simulating MPI Collectives on Tofino Smart Switches in SimGrid
abstract
Programmable smart switches enable In-Network Computing, e.g., to accelerate HPC workloads by offloading collective operations from host CPUs. However, evaluating the benefits of these devices remains challenging due to the cost and complexity of deployment on real hardware. In this paper, we address this limitation by simulating Intel Tofino smart switches using the SimGrid framework. We take advantage of the components of the S4U and SMPI modules and introduce a new network component that represents smart switches that reproduce the latency and computational capabilities of the Tofino architecture. We validate this model on a physical testbed and present a performance evaluation of MPI collective operations offloading. Although we focus on simulating Tofino-class switches, our approach can be adapted to other smart switch architectures. Our preliminary results indicate that small-scale simulations achieve latency comparable to the real hardware when offloading MPI_Allreduce. This study lays the groundwork for future assessments of Tofino smart switches at scale.
Ahmad Moh'd Saleh A. Belbeisi, Majid Salimi Beni, Thomas Erbesdobler, Ehab Saleh, Matthew Tovey, Amir Raoofy, Josef Weidendorfer
CF2
2025 Phase-Based Frequency Scaling for Energy-Efficient Heterogeneous Computing
abstract
Energy efficiency has been a major challenge for exascale computing. Frequency scaling is a powerful technique to achieve energy savings in modern heterogeneous systems, and can be applied either at a coarse granularity, by application, or at a fine granularity, by setting the frequency for each computational kernel. The chosen granularity significantly impacts the performance and energy consumption of applications due to frequency-change overhead. We propose a novel phase-based method that minimizes the frequency-change overhead and improves performance and energy efficiency on heterogeneous multi-GPU systems. Our approach detects different phases through application profiling and DAG analysis, and sets an optimal frequency for each phase. Our methodology also considers MPI programs, where the overhead can be hidden by overlapping frequency-change with communication. Experimental results show up to 37 % energy saving and$1.87 \times$speedup for various benchmarks on a single GPU, and 68 % energy saving and$3.63 \times$speedup on two multiGPU applications.
Lorenzo Carpentieri, Antonio De Caro, Majid Salimi Beni, Kaijie Fan, Biagio Cosenza
IPDPS3
2024 MPI Collective Algorithm Selection in the Presence of Process Arrival Patterns
abstract
The Message Passing Interface (MPI) is a programming model for developing high-performance applications on large-scale machines. A key component of MPI is its collective communication operations. While the MPI standard defines the semantics of these operations, it leaves the algorithmic implementation to the MPI libraries. Each MPI library contains various algorithms for each collective, and selecting the best algorithm typically relies on performance metrics obtained from micro-benchmarks. In such micro-benchmarks, processes are typically synchronized using an MPI_Barrier before invoking a collective operation. However, in real-world scenarios, processes often arrive at a collective in diverse patterns, often due to resource contention. The performance of collective algorithms can vary significantly depending on the arrival pattern type. In this work, we address the challenge of selecting the most efficient algorithm for a given collective, taking into account process arrival patterns. First, we demonstrate through a simulation study that arrival patterns significantly influence the choice of the optimal collective algorithm for specific communication instances. Second, we conduct a comprehensive micro-benchmark analysis to illustrate the sensitivity of MPI collectives to these arrival patterns. Third, we show that our innovative micro-benchmarking methodology is effective in selecting the best-performing collective algorithm for real-world applications.
Majid Salimi Beni, Biagio Cosenza, Sascha Hunold
CLUSTER1
2024 Analysis and prediction of performance variability in large-scale computing systems
abstract
Abstract The development of new exascale supercomputers has dramatically increased the need for fast, high-performance networking technology. Efficient network topologies, such as Dragonfly+, have been introduced to meet the demands of data-intensive applications and to match the massive computing power of GPUs and accelerators. However, these supercomputers still face performance variability mainly caused by the network that affects system and application performance. This study comprehensively analyzes performance variability on a large-scale HPC system with Dragonfly+ network topology, focusing on factors such as communication patterns, message size, job placement locality, MPI collective algorithms, and overall system workload. The study also proposes an easy-to-measure metric for estimating network background traffic generated by other users, which can be used to estimate the performance of our job accurately. The insights gained from this study contribute to improving performance predictability, enhancing job placement policies and MPI algorithm selection, and optimizing resource management strategies in supercomputers.
Majid Salimi Beni, Sascha Hunold, Biagio Cosenza
J. Supercomput.1
2023 EMPI: Enhanced Message Passing Interface in Modern C++
abstract
Message Passing Interface (MPI) is a well-known standard for programming distributed and HPC systems. While the community has been continuously improving MPI to address the requirements of next-generation architectures and applications, its interface has not substantially evolved. In fact, MPI only provides an interface to C and Fortran and does not support recent features of modern C++. Moreover, MPI programs are error-prone and subject to different syntactic and semantic errors. This paper introduces EMPI, an Enhanced Message Passing Interface based on modern C++, which is directly mapped to the OpenMPI implementation and exploits modern C++ for safe and efficient distributed programming. EMPI proposes novel C++RAII-based semantics and constant specialization to prevent error-prone code patterns such as parameter mismatch, and reduce the overhead of handling multiple objects and perinvocation time. Consequently, EMPI programs are safer: six out of nine well-known MPI error patterns do not occur while correctly using EMPI semantics. Experimental results on five microbenchmarks and two applications on a large-scale cluster using up to 1024 processes show that EMPI's performance is very similar to native MPI and considerably faster than the MPL C++ interface.
Majid Salimi Beni, Luigi Crisci, Biagio Cosenza
CCGrid1
2022 An Analysis of Performance Variability on Dragonfly+topology
abstract
Large-scale compute clusters are highly affected by performance variability that originates from different sources. Among these sources, the network plays an essential role as a shared resource between users and their jobs in a supercomputer. In this paper, we analyze the effect of some network-related sources on the performance variability of a modern compute cluster equipped with a Dragonfly+ interconnect. Specifically, we focus on the impacts of job placement, communication patterns, routing strategy, and network background traffic on the performance variability of communication-intensive workloads. To quantify the effect of network congestion (background traffic) on the performance variability, we propose a heuristic that can successfully estimate the amount of communication on the network produced by other jobs running on the cluster simultaneously. Then, we show how this network congestion contributes to the performance variability of different communication patterns and real-world communication-intensive applications.
Majid Salimi Beni, Biagio Cosenza
CLUSTER1
2021 Ignite-GPU: a GPU-enabled in-memory computing architecture on clusters
Amir Hossein Sojoodi, Majid Salimi Beni, Farshad Khunjush
J. Supercomput.2