Ozgur O. Kilic

dblp:252/7149 · also Ozgur Ozan Kilic · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0003-2129-408XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Alternative mixed integer linear programming optimization for joint job scheduling and data allocation in grid computing
abstract
This paper presents a novel approach to the joint optimization of job scheduling and data allocation in grid computing environments. We formulate this joint optimization problem as a mixed integer quadratically constrained program. To tackle the nonlinearity in the constraint, we alternatively fix a subset of decision variables and optimize the remaining ones via Mixed Integer Linear Programming (MILP). We solve the MILP problem at each iteration via an off-the-shelf MILP solver. Our experimental results show that our method significantly outperforms existing heuristic methods, employing either independent optimization or joint optimization strategies. We have also verified the generalization ability of our method over grid environments with various sizes and its high robustness to the algorithm setting.
Shengyu Feng, Jaehyung Kim 0001, Yiming Yang 0002, Joseph Boudreau, Tasnuva Chowdhury, Adolfy Hoisie, Raees Khan, Ozgur O. Kilic, Scott Klasky, Tatiana Korchuganova, Paul Nilsson, Verena Ingrid Martinez Outschoorn, David Keetae Park, Norbert Podhorszki, Yihui Ren 0001, Frédéric Suter, Sairam Sri Vatsavai, Shinjae Yoo, Tadashi Maeno, Alexei Klimentov
Future Gener. Comput. Syst.8
2026 Predicting runtime and resource utilization of jobs on integrated cloud and HPC systems
Esma Yildirim, Mohab Hussein, Mikhail Titov, Ozgur O. Kilic
Future Gener. Comput. Syst.4
2024 Workflow Mini-Apps: Portable, Scalable, Tunable & Faithful Representations of Scientific Workflows
abstract
Workflows are critical for scientific discovery. However, the sophistication, heterogeneity, and scale of workflows make building, testing, and optimizing them increasingly challenging. Furthermore, their complexity and heterogeneity make performance reproducibility hard. In this paper, we propose workflow mini-apps as a tool to address the challenges in building and testing workflows while controlling the fidelity of representing real-world workflows. Workflow mini-apps are deployed and run on various HPC systems and architectures without workflow-specific constraints. We offer insight into their design and implementation, providing an analysis of their performance and reproducibility. Workflow mini-apps thus advance the science of workflows by providing simple, portable, and managed (fidelity) representations of otherwise complex and difficult-to-control real workflows.
Ozgur O. Kilic, Tianle Wang 0001, Matteo Turilli, Mikhail Titov, André Merzky, Line C. Pouchard, Shantenu Jha
CCGrid1
2024 Enabling Performance Observability for Heterogeneous HPC Workflows with SOMA
abstract
Heterogeneous workflows represent a promising approach for overcoming traditional application performance limitations and to accelerate scientific insight on high-performance computing (HPC) platforms. As HPC platforms grow in size and complexity, managing and optimizing workflow resources while maximizing scientific output assumes vital importance. Optimal workflow resource allocation requires high-quality and timely information about the state of the hardware resources, the status of the pending tasks, the performance of the tasks that have already been executed, and the current status of the workflow itself. A robust performance observability framework that captures and delivers this information can fundamentally improve the quality of decision-making within the workflow system, setting the stage for the adaptive execution of workflow tasks. We propose the use of SOMA, a service-based performance observability framework for such HPC workflows. With the RADICAL-Pilot runtime system as a development vehicle, SOMA demonstrates that service-based architectures coupled with an appropriate data model can serve the performance monitoring needs of large-scale ensemble workflows in a low-overhead fashion. Effective observability of workflow performance requires exporting, storing, and analyzing several types of performance data from across the application and workflow software stacks. Our study finds significant benefits in integrating observability frameworks as first-class citizens within an HPC workflow software stack. In this paper, we demonstrate how SOMA can simultaneously observe the performance states of the individual tasks, system hardware, and the workflow as a whole. Such information can then be employed to calculate better resource allocation and task configuration.
Dewi Yokelson, Mikhail Titov, Srinivasan Ramesh, Ozgur O. Kilic, Matteo Turilli, Shantenu Jha, Allen D. Malony
ICPP4
2024 Radical-Cylon: A Heterogeneous Data Pipeline for Scientific Computing
Arup Kumar Sarker, Aymen Alsaadi, Niranda Perera, Mills Staylor, Gregor von Laszewski, Matteo Turilli, Ozgur O. Kilic, Mikhail Titov, André Merzky, Shantenu Jha, Geoffrey C. Fox
JSSPP7
2023 Building the I (Interoperability) of FAIR for Performance Reproducibility of Large-Scale Composable Workflows in RECUP
abstract
Scientific computing communities increasingly run their experiments using complex data- and compute-intensive workflows that utilize distributed and heterogeneous architectures targeting numerical simulations and machine learning, often executed on the Department of Energy Leadership Computing Facilities (LCFs). We argue that a principled, systematic approach to implementing FAIR principles at scale, including fine-grained metadata extraction and organization, can help with the numerous challenges to performance reproducibility posed by such workflows. We extract workflow patterns, propose a set of tools to manage the entire life cycle of performance metadata, and aggregate them in an HPC-ready framework for reproducibility (RECUP). We describe the challenges in making these tools interoperable, preliminary work, and lessons learned from this experiment.
Bogdan Nicolae, Tanzima Z. Islam, Robert B. Ross, Huub J. J. Van Dam, Kevin Assogba, Polina Shpilker, Mikhail Titov, Matteo Turilli, Tianle Wang 0001, Ozgur O. Kilic, Shantenu Jha, Line C. Pouchard
e-Science10
2023 Asynchronous Execution of Heterogeneous Tasks in ML-Driven HPC Workflows
Vincent R. Pascuzzi, Ozgur O. Kilic, Matteo Turilli, Shantenu Jha
JSSPP2
2022 MemGaze: Rapid and Effective Load-Level Memory Trace Analysis
abstract
A challenge of memory trace analysis is combining detailed analysis and low overhead measurement. Currently, hardware/software-based analysis of load-level sequences easily incurs time slowdowns of 100x. We present MemGaze, a tool for low-overhead, high-resolution memory trace analysis. MemGaze uses Intel's Processor Tracing (PT) instruction ptwrite to collect sampled and compressed memory address traces for load-level, sequence-aware analysis of data reuse. We describe multi-resolution analysis for locations vs. operations, accesses vs. spatio-temporal reuse, and reuse (distance, rate, volume) vs. access patterns. Both trace size and resolution are controllable. We use MemGaze to elucidate the memory effects of different data structures and algorithms. For sampled traces that are ≈ 1 % of a full one, analysis metrics have 1-25% MAPE for histograms of varying dynamic sequence lengths. With current suboptimal kernel support (PT runs continuously), MemGaze's time overhead is typically 10-95%; 7x at worst. However, when PT runs only during samples, overhead is 10–35 % on memory intensive regions and correlates with executed ptwrites.
Ozgur O. Kilic, Nathan R. Tallent, Yasodha Suriyakumar, Chenhao Xie 0001, Andrés Márquez 0001, Stéphane Eranian
CLUSTER1
2020 Rapid Memory Footprint Access Diagnostics
abstract
Footprint and reuse distance measure temporal locality and therefore do not capture the significance of access patterns (spacial locality). A strided access pattern has the largest possible footprint but usually has the best performance. To highlight exposed memory latency, we separate footprint into strided (prefetchable) and irregular (non-prefetchable) access components and calculate the growth rate of each. To rapidly compute these footprint access diagnostics, we present two methods, whole-program and precise. Current footprint analyses can cause 200× or more slowdown with realistic inputs and are therefore impractical. Our whole-program method reduces the overhead to 10% by computing upper bounds, but still yields inter-procedural insight through a call path profile. Our precise method uses additional static analysis and profiling to refine the upper bounds for intra-procedural loop nests. We evaluate our approaches using benchmarks that vary access patterns (strided vs. unpredictable), sparsity (all words in a cache line vs. some), and reuse (varying and repeated accesses per element). Notably, for loop nests with unpredictable accesses, the precise method's accuracy is within 10% of ideal. The whole-program method has sufficient accuracy to diagnose bottlenecks.
Ozgur O. Kilic, Nathan R. Tallent, Ryan D. Friese
ISPASS1
2019 Rapidly Measuring Loop Footprints
abstract
Knowing a loop's footprint - the unique data items it accesses - enables important locality and capacity analysis. Unfortunately, current methods for computing footprint cause integer factors of slowdown and are therefore difficult to use with realistic inputs. Current methods are slow because they are based on software tracing of load instructions and data addresses. We present an approach that combines lightweight measurements and static binary analysis. We use static analysis to reason about each load's expected reuse. We then calculate the footprint for each loop using inference rules informed by measurements from common performance counters. We validate our method on benchmarks that vary access patterns (strided vs. unpredictable), reuse (varying and repeated accesses per element), and sparsity (all words in a cache line vs. some). For strided patterns, errors are within 1%; for unpredictable ones, errors are 5-10%. Our overheads are under 10%. A tool based on software tracing has an error of less than 1%, but either introduces at least 130× overhead or hangs.
Ozgur O. Kilic, Nathan R. Tallent, Ryan D. Friese
CLUSTER1
2018 Overcoming Virtualization Overheads for Large-vCPU Virtual Machines
abstract
Virtual Machines (VM) frequently run parallel applications in cloud environments, and high performance computing platforms. It is well known that configuring a VM with too many virtual processors (vCPUs) worsens application performance due to scheduling cross-talk between the hypervisor and the guest OS. Specifically, when the number of vCPUs assigned to a VM exceeds available physical CPUs then parallel applications in the VM experience worse performance, even when number of application threads remains fixed. In this paper, we first track the root cause of this performance loss to inefficient hypervisor-level emulation of inter-vCPU synchronization events. We then present three techniques to minimize hypervisor-induced overheads on parallel workloads in large-VCPU VMs. The first technique pins application threads to dedicated vCPUs to eliminate inter-vCPU thread migrations, reducing the overhead of emulating inter-processor interrupts (IPIs). The second technique para-virtualizes inter-vCPU TLB flush operations. The third technique enables faster reactivation of idle vCPUs by prioritizing the delivery of rescheduling IPIs. Unlike existing solutions which rely on heavyweight and slow vCPU hotplug mechanisms, our techniques are lightweight and provide more flexibility in migrating large-vCPU VMs. Using several parallel benchmarks, we demonstrate the effectiveness of our prototype implementation in the Linux KVM/QEMU virtualization platform. Specifically, we demonstrate that with our techniques, parallel applications can maintain their performance even when up to 255 VCPUs are assigned to a VM running on as few as 6 physical cores.
Ozgur O. Kilic, Spoorti Doddamani, Aprameya Bhat, Hardik Bagdi, Kartik Gopalan
MASCOTS1