EDBT 2026 Demo / reviewers in the wild / expert
Ozan Tuncer
dblp:155/4386
· DBLP profile ↗
7ranked-venue papers
3as first author
0since 2021 · last 2019
0000-0003-4222-1648ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Distributed systems · 58% Performance modeling and evaluation · 29% High-performance computing · 9% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
anomaly diagnosis |
0.4 | 1 | 2019 | Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning · IEEE Trans. Parallel Distributed Syst. 2019 |
Performance modeling and evaluation
performance variability |
0.4 | 1 | 2019 | Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning · IEEE Trans. Parallel Distributed Syst. 2019 |
Distributed systems › anomaly detection
runtime anomaly detection |
0.4 | 1 | 2019 | Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning · IEEE Trans. Parallel Distributed Syst. 2019 |
High-performance computing
system resilience |
0.1 | 1 | 2019 | Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning · IEEE Trans. Parallel Distributed Syst. 2019 |
Energy-efficient computing
datacenter energy efficiency |
0.1 | 1 | 2015 | Leakage-Aware Cooling Management for Improving Server Energy Efficiency · IEEE Trans. Parallel Distributed Syst. 2015 |
Methods — techniques the papers use, named apart from their topics
time series analysis · 0.4statistical feature extraction · 0.4machine learning · 0.4fan speed optimization · 0.2empirical power modeling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Online Diagnosis of Performance Variation in HPC Systems Using Machine LearningabstractAs the size and complexity of high performance computing (HPC) systems grow in line with advancements in hardware and software technology, HPC systems increasingly suffer from performance variations due to shared resource contention as well as software- and hardware-related problems. Such performance variations can lead to failures and inefficiencies, which impact the cost and resilience of HPC systems. To minimize the impact of performance variations, one must quickly and accurately detect and diagnose the anomalies that cause the variations and take mitigating actions. However, it is difficult to identify anomalies based on the voluminous, high-dimensional, and noisy data collected by system monitoring infrastructures. This paper presents a novel machine learning based framework to automatically diagnose performance anomalies at runtime. Our framework leverages historical resource usage data to extract signatures of previously-observed anomalies. We first convert collected time series data into easy-to-compute statistical features. We then identify the features that are required to detect anomalies, and extract the signatures of these anomalies. At runtime, we use these signatures to diagnose anomalies with negligible overhead. We evaluate our framework using experiments on a real-world HPC supercomputer and demonstrate that our approach successfully identifies 98 percent of injected anomalies and consistently outperforms existing anomaly diagnosis techniques. Ozan Tuncer, Emre Ates, Yijia Zhang 0002, Ata Turk, Jim M. Brandt, Vitus J. Leung, Manuel Egele, Ayse K. Coskun |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Taxonomist: Application Detection Through Rich Monitoring Data
Emre Ates, Ozan Tuncer, Ata Turk, Vitus J. Leung, Jim M. Brandt, Manuel Egele, Ayse K. Coskun |
Euro-Par | 2 |
| 2018 | Level-Spread: A New Job Allocation Policy for Dragonfly NetworksabstractThe dragonfly network topology has attracted attention in recent years owing to its high radix and constant diameter. However, the influence of job allocation on communication time in dragonfly networks is not fully understood. Recent studies have shown that random allocation is better at balancing the network traffic, while compact allocation is better at harnessing the locality in dragonfly groups. Based on these observations, this paper introduces a novel allocation policy called Level-Spread for dragonfly networks. This policy spreads jobs within the smallest network level that a given job can fit in at the time of its allocation. In this way, it simultaneously harnesses node adjacency and balances link congestion. To evaluate the performance of Level-Spread, we run packet-level network simulations using a diverse set of application communication patterns, job sizes, and communication intensities. We also explore the impact of network properties such as the number of groups, number of routers per group, machine utilization level, and global link bandwidth. Level-Spread reduces the communication overhead by 16% on average (and up to 71%) compared to the state-of-the-art allocation policies. Yijia Zhang 0002, Ozan Tuncer, Fulya Kaplan, Katzalin Olcoz, Vitus J. Leung, Ayse K. Coskun |
IPDPS | 2 |
| 2017 | Unveiling the Interplay Between Global Link Arrangements and Network Management Algorithms on Dragonfly NetworksabstractNetwork messaging delay historically constitutes a large portion of the wall-clock time for High Performance Computing (HPC) applications, as these applications run on many nodes and involve intensive communication among their tasks. Dragonfly network topology has emerged as a promising solution for building exascale HPC systems owing to its low network diameter and large bisection bandwidth. Dragonfly includes local links that form groups and global links that connect these groups via high bandwidth optical links. Many aspects of the dragonfly network design are yet to be explored, such as the performance impact of the connectivity of the global links, i.e., global link arrangements, the bandwidth of the local and global links, or the job allocation algorithm. This paper first introduces a packet-level simulation framework to model the performance of HPC applications in detail. The proposed framework is able to simulate known MPI (message passing interface) routines as well as applications with custom-defined communication patterns for a given job placement algorithm and network topology. Using this simulation framework, we investigate the coupling between global link bandwidth and arrangements, communication pattern and intensity, job allocation and task mapping algorithms, and routing mechanisms in dragonfly topologies. We demonstrate that by choosing the right combination of system settings and workload allocation algorithms, communication overhead can be decreased by up to 44%. We also show that circulant arrangement provides up to 15% higher bisection bandwidth compared to the other arrangements, but for realistic workloads, the performance impact of link arrangements is less than 3%. Fulya Kaplan, Ozan Tuncer, Vitus J. Leung, Karl S. Hemmert, Ayse K. Coskun |
CCGrid | 2 |
| 2015 | PaCMap: Topology Mapping of Unstructured Communication Patterns onto Non-contiguous AllocationsabstractIn high performance computing (HPC), applications usually have many parallel tasks running on multiple machine nodes. As these tasks intensively communicate with each other, the communication overhead has a significant impact on an application's execution time. This overhead is determined by the application's communication pattern as well as the network distances between communicating tasks. By mapping the tasks to the available machine nodes in a communication-aware manner, the network distances and the execution times can be significantly reduced. Ozan Tuncer, Vitus J. Leung, Ayse K. Coskun |
ICS | 1 |
| 2015 | Leakage-Aware Cooling Management for Improving Server Energy EfficiencyabstractThe computational and cooling power demands of enterprise servers are increasing at an unsustainable rate. Understanding the relationship between computational power, temperature, leakage, and cooling power is crucial to enable energy-efficient operation at the server and data center levels. This paper develops empirical models to estimate the contributions of static and dynamic power consumption in enterprise servers for a wide range of workloads, and analyzes the interactions between temperature, leakage, and cooling power for various workload allocation policies. We propose a cooling management policy that minimizes the server energy consumption by setting the optimum fan speed during runtime. Our experimental results on a presently shipping enterprise server demonstrate that including leakage awareness in workload and cooling management provides additional energy savings without any impact on performance. Marina Zapater, Ozan Tuncer, José Luis Ayala, José Manuel Moya, Kalyan Vaidyanathan, Kenny C. Gross, Ayse K. Coskun |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | CoolBudget: Data center power budgeting with workload and cooling asymmetry awarenessabstractPower over-subscription challenges and emerging cost management strategies motivate designing efficient data center power capping techniques. During capping, provisioned power must be budgeted among the computational and cooling units. This work presents a data center power budgeting policy that simultaneously improves the quality-of-service (QoS) and power efficiency by considering the workload- and cooling-induced asymmetries among the servers. Proposed policy finds the most efficient data center temperature and the power distribution among servers while guaranteeing reliable temperature levels for the server internal components. Experiments based on real servers demonstrate 21% increase in throughput compared to existing techniques. Ozan Tuncer, Kalyan Vaidyanathan, Kenny C. Gross, Ayse K. Coskun |
ICCD | 1 |