Ahmad Tarraf

dblp:241/0969 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0002-9174-5598ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2025 EquilibrIO: Taming the I/O Tides in High-Performance Computing
abstract
In high-performance computing systems, jobs typically have exclusive compute access but share storage resources, such as the parallel file system, often becoming a point of contention. Concurrent execution of data-intensive jobs can exacerbate this phenomenon, as jobs compete for shared resources, impeding each other's progress while suffering from limited I/O bandwidth. As a result, the increasing I/O intensity of workloads places greater demands on resource management systems to optimize the scheduling of data-intensive jobs. Although scheduling decisions significantly impact shared storage systems, scheduling algorithms on production systems generally ignore the I/O intensity of individual jobs. In this work, we present EquilibrIO, a novel job scheduling algorithm that minimizes resource contention and maintains fairness by balancing computation and I/O over time, while requiring minimal information collected by tools commonly used in high-performance computing systems. We show that, depending on the desired level of fairness, our algorithm can reduce the I/O slowdown caused by contention from 64 % to 4 %. The results further demonstrate that 25 % of the jobs augmented with additional I/O information are sufficient to minimize file system congestion, cutting the effect of I/O slowdown by half.
Taylan Özden, Ahmad Tarraf, Felix Wolf 0001
CLUSTER2
2025 A Deep Look into the Temporal I/O Behavior of HPC Applications
abstract
The increasing gap between compute and I/O speeds in high-performance computing (HPC) systems imposes the need for techniques to improve applications' I/O performance. Such techniques must rely on assumptions about I/O behavior in order to efficiently allocate I/O resources such as burst buffers, to schedule accesses to the shared parallel file system or to delay certain applications at the batch scheduler level to prevent contention, for instance. In this paper, we verify these common assumptions about I/O behavior, specifically about temporal behavior, using over 440,000 traces from real HPC systems. By combining traces from diverse systems, we characterize the behaviors observed in real HPC workloads. Among other findings, we show that I/O activity tends to last for a few seconds, and that periodic jobs are the minority, but responsible for a large portion of the I/O time. Furthermore, we make projections for the expected improvement yielded by popular approaches for I/O performance improvement. Our work provides valuable insights to everyone working to alleviate the I/O bottleneck in HPC.
Francieli Zanon Boito, Luan Teylo, Mihail Popov, Théo Jolivel, Francois Tessier, Jakob Lüttgau, Julien Monniot, Ahmad Tarraf, Andre Ramos Carneiro, Carla Osthoff
IPDPS8
2024 I/O Behind the Scenes: Bandwidth Requirements of HPC Applications with Asynchronous I/O
abstract
I/O bandwidth is a critical resource in an HPC cluster. As with all shared resources, its availability is impacted significantly by the users and the applications they execute. Without proper restrictions, jobs consuming more prominent portions of the I/O bandwidth can severely affect other jobs by notably prolonging their runtime. In such a context, applications that perform asynchronous I/O bring unique properties that allow for the reduction of such effects. That is, by limiting the bandwidth to the required one to perform the I/O in the background of the compute phases, I/O bursts can be flattened without significantly prolonging the application time, if at all. Hence, the bandwidth consumption of such applications is limited to what they need, sparing much of the system bandwidth to other applications. At the same time, these applications achieve higher parallel efficiency due to the overlapping of different resources (e.g., compute and I/O). This paper shows these aspects and demonstrates our approach to finding the required bandwidth for applications that use asynchronous I/O. Moreover, we apply it automatically using an MPI implementation of a bandwidth limitation approach at the application level. We validate our approach with several experiments on a large production cluster and show the impact of our approach on the application behavior and its importance for the system throughput.
Ahmad Tarraf, Javier Fernández 0001, David E. Singh, Taylan Özden, Jesús Carretero 0001, Felix Wolf 0001
CLUSTER1
2024 Capturing Periodic I/O Using Frequency Techniques
abstract
Many HPC applications perform their I/O in bursts that follow a periodic pattern. This allows for making predictions as to when a burst occurs. System providers can take advantage of such knowledge to reduce file-system contention by actively scheduling I/O bandwidth. The effectiveness of this approach, however, depends on the ability to detect and quantify the periodicity of I/O patterns online. In this paper, we introduce FTIO, an online method to detect periodic I/O phases, which is based on discrete Fourier transform (DFT), combined with outlier detection. We provide metrics that gauge the confidence in the output and tell how far from being periodic the signal is. We validate our approach with large-scale experiments on a production system and examine its limitations extensively. Our experiments show that FTIO has a mean error below 11%. Finally, we demonstrate that FTIO allowed the I/O scheduler Set10 to boost system utilization by 26% and reduce I/O slowdown by 56%.
Ahmad Tarraf, Alexis Bandet, Francieli Zanon Boito, Guillaume Pallez, Felix Wolf 0001
IPDPS1
2024 Malleability in Modern HPC Systems: Current Experiences, Challenges, and Future Opportunities
abstract
With the increase of complex scientific simulations driven by workflows and heterogeneous workload profiles, managing system resources effectively is essential for improving performance and system throughput, especially due to trends like heterogeneous HPC and deeply integrated systems with on-chip accelerators. For optimal resource utilization, dynamic resource allocation can improve productivity across all system and application levels, by adapting the applications' configurations to the system's resources. In this context, malleable jobs, which can change resources at runtime, can increase the system throughput and resource utilization while bringing various advantages for HPC users (e.g., shorter waiting time). Malleability has received much attention recently, even though it has been an active research area for almost two decades [1]. This paper presents the state-of-the-art of malleable implementations in HPC systems, targeting mainly malleability in compute and I/O resources. Based on our experiences, we state our current concerns and list future opportunities for research.
Ahmad Tarraf, Martin Schreiber 0001, Alberto Cascajo, Jean-Baptiste Besnard, Marc-Andre Vef, Dominik Huber, Sonja Happ, André Brinkmann, David E. Singh, Hans-Christian Hoppe, Alberto Miranda, Antonio J. Peña, Marta Garcia-Gasulla, Martin Schulz 0001, Paul M. Carpenter, Simon Pickartz, Tiberiu Rotaru, Sergio Iserte, Víctor López 0003, Jorge Ejarque, Heena Sirwani, Jesús Carretero 0001, Felix Wolf 0001
IEEE Trans. Parallel Distributed Syst.1
2020 Establishing Reachset Conformance for the Formal Analysis of Analog Circuits
abstract
We present the first work on the automated generation of reachset conformant models for analog circuits. Our approach applies reachset conformant synthesis to add nondeterminism to piecewise-linear circuit models so that they enclose all recorded behaviors of the real system. To achieve this, we present a novel technique to compute the required nondeterminism for the piecewise-linear models. The effectiveness of our approach is demonstrated on a real analog circuit. Since the resulting models enclose all measurements, they can be used for formal verification.
Niklas Kochdumper, Ahmad Tarraf, Malgorzata Rechmal, Markus Olbrich, Lars Hedrich, Matthias Althoff
ASP-DAC2
2019 Behavioral Modeling of Transistor-Level Circuits using Automatic Abstraction to Hybrid Automata
abstract
Accurate abstracted behavioral modeling of analog circuits is still an open problem, especially when the abstraction process is automated. In this paper we present an automated abstraction technique of transistor level circuits with full SPICE accuracy alongside a significant simulation speed-up. The methodology computes a hybrid automaton which is transformed into a behavioral model in Verilog-A. The resulting hybrid automaton exhibits linear behavior as well as the technology dependent nonlinear e.g. limiting behavior. The accuracy and speed-up of the methodology is evaluated on several transistor level circuits ranging from simple operational amplifiers up to a complex industrial OTA-based Gm/C filter. Finally, we formally verify the equivalence between the generated model and the original circuit.
Ahmad Tarraf, Lars Hedrich
DATE1
2019 Multi-agent Learning for Energy-Aware Placement of Autonomous Vehicles
abstract
Mobility gets an increasing amount of meaning and significance in the modern society. In this paper, we introduce a multi-agent learning application for a multi-agent system in e-mobility. In particular, we propose a geospatial model for free-floating and autonomously driving e-trikes and demonstrate a calculation method of positioning e-trikes on a given area by using different methods of cluster analysis. The solution of the cluster analysis contains cluster centers which represent a positioning for the e-trikes. The solution is then evaluateded by a simulation model with more sophisticated parameters. This research field opens different opportunities of application scenarios, which are discussed in the conclusion.
Ömer Ibrahim Erduran, Mirjam Minor, Lars Hedrich, Ahmad Tarraf, Frederik Ruehl, Hans Schroth
ICMLA4