Amir Raoofy

dblp:189/1229 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-9664-8298ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 first-author · 10 since 2021
YearPublicationVenuePosition
2026 Simulating MPI Collectives on Tofino Smart Switches in SimGrid
abstract
Programmable smart switches enable In-Network Computing, e.g., to accelerate HPC workloads by offloading collective operations from host CPUs. However, evaluating the benefits of these devices remains challenging due to the cost and complexity of deployment on real hardware. In this paper, we address this limitation by simulating Intel Tofino smart switches using the SimGrid framework. We take advantage of the components of the S4U and SMPI modules and introduce a new network component that represents smart switches that reproduce the latency and computational capabilities of the Tofino architecture. We validate this model on a physical testbed and present a performance evaluation of MPI collective operations offloading. Although we focus on simulating Tofino-class switches, our approach can be adapted to other smart switch architectures. Our preliminary results indicate that small-scale simulations achieve latency comparable to the real hardware when offloading MPI_Allreduce. This study lays the groundwork for future assessments of Tofino smart switches at scale.
Ahmad Moh'd Saleh A. Belbeisi, Majid Salimi Beni, Thomas Erbesdobler, Ehab Saleh, Matthew Tovey, Amir Raoofy, Josef Weidendorfer
CF6
2026 A Dedicated CPU Core for MPI Progress: Towards Improved Overlap in Non-Blocking Two-Sided Communication
Ehab Saleh, Amir Raoofy, Robert Mijakovic, Ahmad Moh'd Saleh A. Belbeisi, Josef Weidendorfer
ISPDC2
2025 POSTER: Performance Comparison of GPU Programming Models Using HeCBench Benchmarks
abstract
GPUs play an important role in High-Performance Computing.The choice of GPU programming models plays a crucial role in achieving portability and performance.High-level programming models, such as SYCL and OpenMP offloading, have emerged, offering unified abstractions that enable developers to target multiple architectures with a single, maintainable codebase.However, achieving consistent performance across different models remains a significant challenge due to variations in abstraction levels, compiler optimizations, and runtime behavior.We present a profiling-based methodology for systematically comparing GPU programming models on NVIDIA and AMD GPUs.We apply our methodology to over 150 benchmarks of HeCBench, demonstrating its effectiveness in identifying performance issues in OpenMP, SYCL, HIP and CUDA implementations for AMD and NVIDIA GPUs. CCS Concepts• General and reference
Jakob Schäffeler, Bengisu Elis, Amir Raoofy, Josef Weidendorfer, Martin Schulz 0001
CF3
2025 Reproducibility Report for SC25 Paper KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPU
abstract
This reproducibility report provides details about the artifact evaluation done with regards to the Artifact Description and Evaluation appendix of SC25 paper KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPU by Xiaowen Tian et al. The work was done as part of the Reproducibility Initiative of SC25. The author is a member of the SC25 Reproducibilty Committee.
Amir Raoofy
SC1
2024 A Portable Tool to Compare Performance Profiles from GPU Offloading Programming Models
abstract
GPUs are growingly dominating the High-Performance Computing ecosystem, and therefore, the ease of their programming is getting increasingly important. Standard and high-level offloading methods, like OpenMP offloading and OpenACC, facilitate portable and efficient offloading across different GPU platforms. However, pinpointing and troubleshooting performance variations among different models, implementations, or architectures poses a challenge due to varying abstraction levels and profilers employed. Therefore, to tackle this problem and to unwind the performance issues related to various offloading abstractions and models that are entangled together in practice, in this work, we introduce a portable tool to enable the comparison of performance profiles acquired from various offloading models and GPU platforms. For this, the tool first processes the collected profiles by different profilers to extract key performance indicatory metrics. For ease of comparison, the tool utilizes plots depicting the metrics of all target variants for relative comparison. Moreover, we demonstrate the tool's capabilities by discussing specific issues discovered by using the tool when comparing OpenMP offloading and CUDA implementations of Babelstream.
Jakob Schäffeler, Bengisu Elis, Amir Raoofy, Josef Weidendorfer, Martin Schulz 0001
CF3
2023 Machine Learning Application Benchmark
abstract
This paper presents the MLAB project, a research and development activity funded by ESA General Support Technology Programme under the lead of Airbus Defence and Space GmbH, with the goal of developing a machine learning application benchmark for space applications. First, the need for a benchmark dedicated to machine learning applications in spacecraft is explained, and examples of applications are described including their design challenges. Then the benchmark design is presented, including the rules of the metrics, guidelines and scenarios for references. These scenarios include a description of the reference workloads that have been selected during the activity as representative for spacecraft applications. Lastly, the submission concept is introduced.
Michael Petry, Max Ghiglione, Amir Raoofy, Gabriel Dax, Gianluca Furano, Martin Werner 0001, Carsten Trinitis, Martin Langer
CF4
2023 Phase-aware System-Side Sampling for HPC
abstract
HPC compute centers always benefit from better insights into the application mix running on their systems. We present a low-overhead statistical sampling tool running in the background on the system side, which can capture application compute phases. Our tool leverages eBPF (extended Berkeley Packet Filter) from modern Linux kernels to extract phase information provided by developers via instrumentation or instruction pointers. We outline how this tool can be integrated into the monitoring framework DCDB, and we show resulting performance insights.
Julian Scheipl, Amir Raoofy, Michael Ott 0001, Josef Weidendorfer
CF2
2022 Benchmarking and feasibility aspects of machine learning in space systems
abstract
Compute in space, e.g., in miniaturized satellites, requires dealing with special physical and boundary constraints, including the limited energy budget. These constraints impose strict operational conditions on the on-board data processing system and its capability in dealing with sophisticated workloads suchlike Machine Learning (ML). In the meantime, the breakthroughs in ML based on Deep Neural Networks (DNNs) in the last decade promise innovative solutions to expand the functional capabilities of on-board data processing and to drive the space industry forward. Therefore, due to the aforementioned special requirements, performance- and power-efficient, and novel solutions and architectures for deploying ML via, e.g., FPGA-enabled SoC, particularly Commercial-Off-The-Shelf (COTS) solutions, are gaining significant interest in the space industry. Therefore it is essential to conduct extensive benchmarking and feasibility and efficiency analyses in different aspects: such analyses would require the investigation of options for programming and deployment as well as the investigation of various real-world models and datasets. To this end, a research and development activity is funded by the European Space Agency (ESA) General Support Technology Programme and is led by Airbus Defence and Space GmbH with the goal of developing an ML Application Benchmark (MLAB) that covers benchmarking aspects mentioned above.
Amir Raoofy, Gabriel Dax, Vittorio Serra, Max Ghiglione, Martin Werner 0001, Carsten Trinitis
CF1
2022 Always-on instrumentation for application introspection in HPC
abstract
Obtaining insights into the dynamic behavior of user code is crucial for supercomputing centers to support both better operation and co-design of future systems. To this end, always-on instrumentation is the key: enabling all running code to dynamically forward metadata such as compute phase changes to the system would provide important information for those goals. To keep the overhead low, the system must be able to deactivate instrumentation points with high trigger frequency on demand. In this poster, we present a simple always-on instrumentation method for C/C++ which can be easily used by developers, copying a single source file into their code base. Our evaluations show that the overhead in the deactivated state is low enough for the manual instrumentation to stay compiled in, all the time.
Amir Raoofy, Josef Weidendorfer, Michael Ott 0001
CF1
2022 Exploiting Reduced Precision for GPU-based Time Series Mining
abstract
The mining of multi-dimensional time series is a crucial step in gaining insights into data obtained from physical systems and from monitoring infrastructures. A widely accepted approach for this challenge is the matrix profile, which, however, is computationally very expensive. It relies on calculating large correlation matrices coupled with sort operations across all dimensions of the data, as well as on performing inclusive scans. All of these steps are inherently data parallel and can, therefore, benefit from execution on GPUs, and even more so from horizontal scaling on multiple GPUs. In addition, the nature of the matrix profile calculation allows the exploitation of reduced precision on GPUs. This offers further improvements to enable the analysis of ever growing data sets in real-world scenarios. Based on these motivations, we introduce the first parallel algorithm for multi-dimensional matrix profile on multiple GPUs exploiting reduced precision modes and provide a highly opti-mized implementation using novel optimization techniques. On one NVIDIA A100 GPU, our implementation achieves a 54x performance improvement in comparison to an optimized single-node execution on a state-of-the-art CPU-based implementation relying on double-precision computation and an additional factor of 1.4x when switching to reduced precision while maintaining sufficient accuracy. We study the accuracy and performance trade-offs for our proposed algorithm in detail and present synthetic and real-world case studies to demonstrate how the reduced precision improves the performance, while accomplishing sufficiently accurate results.
Yi Ju, Amir Raoofy, Dai Yang, Erwin Laure, Martin Schulz 0001
IPDPS2
2021 Living on the Edge: Efficient Handling of Large Scale Sensor Data
abstract
Real-time sensor monitoring is critical in many industrial applications and is, e.g., used to model and predict operating conditions to optimize operations as well as to prevent damage in machinery and systems. In many cases, this data is generated by a myriad of sensors and stored or transmitted for post-processing by data analysts. Handling this data near its origin-on the edge-imposes significant challenges for storage and compression: it is necessary to store it in a format that is suitable for large data analytics algorithms, which in most cases means columnar storage. Furthermore, to provide efficient storage and transmission of such sensor data, it must be compressed efficiently. However, existing solutions do not address these challenges sufficiently. In this work, we present a holistic approach for fast streaming of large scale sensor data directly into columnar storage and integrate it with a proven compression scheme. Our approach uses a pipelined scheme for streaming and transposing the data layout, combined with a byte-level transformation of data representation and compression, which we evaluate in comprehensive experiments. As a result, our approach enables transformation of large scale sensor data streams into an efficient, analytics-friendly format already at the sensor site, i.e., on the edge, at data ingestion time. By implementing our optimized approach in the open and widely used columnar storage format Apache Parquet, which we already partly upstreamed, we ensure its accessibility to the community.
Roman Karlstetter, Amir Raoofy, Martin Radev, Carsten Trinitis, Jakob Hermann, Martin Schulz 0001
CCGRID2
2016 A parallel arbitrary-order accurate AMR algorithm for the scalar advection-diffusion equation
abstract
We present a numerical method for solving the scalar advection-diffusion equation using adaptive mesh refinement. Our solver has three unique characteristics: (1) it supports arbitrary-order accuracy in space; (2) it allows different discretizations for the velocity and scalar advected quantity; (3) it combines the method of characteristics with an integral equation formulation; and (4) it supports shared and distributed memory architectures. In particular, our solver is based on a second-order accurate, unconditionally stable, semi-Lagrangian scheme combined with a spatially-adaptive Chebyshev octree for discretization. We study the convergence, single-node performance, strong scaling, and weak scaling of our scheme for several challenging flows that cannot be resolved efficiently without using high-order accurate discretizations. For example, we consider problems for which switching from 4th order to 14th order approximation results in two orders of magnitude speedups for a computation in which we keep the target accuracy in the solution fixed. For our largest run, we solve a problem with one billion unknowns on a tree with maximum depth equal to 10 and using 14th-order elements on 16,384 x86 cores on the “STAMPEDE“ system at the Texas Advanced Computing Center.
Arash Bakhtiari, Dhairya Malhotra, Amir Raoofy, Miriam Mehl, Hans-Joachim Bungartz, George Biros
SC3