Taylan Özden

dblp:337/7547 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-4540-4717ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 EquilibrIO: Taming the I/O Tides in High-Performance Computing
abstract
In high-performance computing systems, jobs typically have exclusive compute access but share storage resources, such as the parallel file system, often becoming a point of contention. Concurrent execution of data-intensive jobs can exacerbate this phenomenon, as jobs compete for shared resources, impeding each other's progress while suffering from limited I/O bandwidth. As a result, the increasing I/O intensity of workloads places greater demands on resource management systems to optimize the scheduling of data-intensive jobs. Although scheduling decisions significantly impact shared storage systems, scheduling algorithms on production systems generally ignore the I/O intensity of individual jobs. In this work, we present EquilibrIO, a novel job scheduling algorithm that minimizes resource contention and maintains fairness by balancing computation and I/O over time, while requiring minimal information collected by tools commonly used in high-performance computing systems. We show that, depending on the desired level of fairness, our algorithm can reduce the I/O slowdown caused by contention from 64 % to 4 %. The results further demonstrate that 25 % of the jobs augmented with additional I/O information are sufficient to minimize file system congestion, cutting the effect of I/O slowdown by half.
Taylan Özden, Ahmad Tarraf, Felix Wolf 0001
CLUSTER1
2024 I/O Behind the Scenes: Bandwidth Requirements of HPC Applications with Asynchronous I/O
abstract
I/O bandwidth is a critical resource in an HPC cluster. As with all shared resources, its availability is impacted significantly by the users and the applications they execute. Without proper restrictions, jobs consuming more prominent portions of the I/O bandwidth can severely affect other jobs by notably prolonging their runtime. In such a context, applications that perform asynchronous I/O bring unique properties that allow for the reduction of such effects. That is, by limiting the bandwidth to the required one to perform the I/O in the background of the compute phases, I/O bursts can be flattened without significantly prolonging the application time, if at all. Hence, the bandwidth consumption of such applications is limited to what they need, sparing much of the system bandwidth to other applications. At the same time, these applications achieve higher parallel efficiency due to the overlapping of different resources (e.g., compute and I/O). This paper shows these aspects and demonstrates our approach to finding the required bandwidth for applications that use asynchronous I/O. Moreover, we apply it automatically using an MPI implementation of a bandwidth limitation approach at the application level. We validate our approach with several experiments on a large production cluster and show the impact of our approach on the application behavior and its importance for the system throughput.
Ahmad Tarraf, Javier Fernández 0001, David E. Singh, Taylan Özden, Jesús Carretero 0001, Felix Wolf 0001
CLUSTER4
2024 Performance-driven scheduling for malleable workloads
abstract
Abstract The development of adaptive scheduling algorithms that take advantage of malleability has become a crucial area of research in many large-scale projects. Malleable workloads can improve the system’s performance but, at the same time, provide an extra dimension to the scheduling problem. This paper proposes an adaptive, performance-based job scheduling method that emphasizes the backfilling concept with malleability. The proposed method performs the malleability operations only when the estimated execution time of the involved applications is better than or equal to the execution time on the allocated resources without reconfiguration. The reconfiguration feasibility is determined by performance models considering the application scalability and reconfiguration overheads. Different policies for implementing malleability are presented, each targeting a specific workload in terms of job size and scalability. The comprehensive evaluation shows an improvement in the slowdown up to 49% compared to the non-adaptive baseline scheduling algorithm.
Njoud O. Almaaitah, David E. Singh, Taylan Özden, Jesús Carretero 0001
J. Supercomput.3
2022 ElastiSim: A Batch-System Simulator for Malleable Workloads
abstract
As high-performance computing infrastructures move towards exascale, the role of resource and job management systems is more critical now than ever. Simulating batch systems to improve scheduling algorithms and resource management efficiency is an indispensable option, as running large-scale experiments is expensive and time-consuming. Batch-system simulators are responsible for simulating the computing infrastructure and the types of jobs that constitute the workload. In contrast to rigid jobs, malleable jobs can dynamically reconfigure their resources during runtime. Although studies indicate that malleability can improve system performance, no simulator exists to investigate malleable scheduling policies. In this work, we present ElastiSim, a batch-system simulator supporting the combined scheduling of rigid and malleable jobs. To facilitate the simulation, we propose a malleable workload model and introduce a scheduling protocol that enables the evaluation of topology-, I/O-, and progress-aware scheduling algorithms. We validate the scaling behavior of our workload model by comparing training runtimes of various deep-learning models against the results achieved by ElastiSim. We use real-world cluster trace files to generate workloads and simulate various scheduling algorithms (FCFS, SJF, DRF, SRTF) to analyze their implications on the simulated platform. The results demonstrate that real-world executions show the same scaling behavior as our proposed workload model. We further show that ElastiSim can capture the complex interplay between emerging workloads and modern platforms to support algorithm designers by providing consistently meaningful results. ElastiSim is publicly available as an open-source project on https://github.com/elastisim.
Taylan Özden, Tim Beringer, Arya Mazaheri, Hamid Mohammadi Fard, Felix Wolf 0001
ICPP1