Emre Ates

dblp:201/1248 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
2since 2021 · last 2022
0000-0002-2292-2626ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 61% Performance modeling and evaluation · 30% High-performance computing · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
anomaly diagnosis
0.412019
Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning · IEEE Trans. Parallel Distributed Syst. 2019
Performance modeling and evaluation
performance variability
0.412019
Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning · IEEE Trans. Parallel Distributed Syst. 2019
Distributed systems › anomaly detection
runtime anomaly detection
0.412019
Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning · IEEE Trans. Parallel Distributed Syst. 2019
High-performance computing
system resilience
0.112019
Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning · IEEE Trans. Parallel Distributed Syst. 2019

Methods — techniques the papers use, named apart from their topics

time series analysis · 0.4statistical feature extraction · 0.4machine learning · 0.4
YearPublicationVenuePosition
2022 Praxi: Cloud Software Discovery That Learns From Practice
abstract
With today’s rapidly-evolving cloud landscape embracing continuous integration and delivery, users of cloud systems must monitor software running on theircontainersandvirtual machines (VMs)to ensure compliance, security, and efficiency. Traditional solutions to this problem rely on manually-createdrulesthat identify software installations and modifications, but these require expert authors and are often unmaintainable. Recently, automated techniques for software discovery have emerged. Some techniques use examples of software to train machine learning models to predict which software has been installed on a system. Others leverage the knowledge of packaging practices to aid in discovery without requiring any pre-training, but these practice-based methods cannot provide precise-enough information to perform discovery by themselves. This article introduces Praxi, a new software discovery method that builds upon the strengths of prior approaches by combining the accuracy of learning-based methods with the efficiency of practice-based methods. In tests using samples collected on real-world cloud systems, Praxi correctly classifies installations at least 97.6 percent of the time, while running 14.8 times faster and using 87 percent less disk space than a similar learning-based method. Using a diverse software dataset, this article quantitatively compares Praxi to systematic rule-, learning-, and practice-based methods, and discusses the best uses for each.
Anthony Byrne, Emre Ates, Ata Turk, Vladimir Pchelin, Sastry S. Duri, Shripad Nadgowda, Canturk Isci, Ayse K. Coskun
IEEE Trans. Cloud Comput.2
2021 Automating instrumentation choices for performance problems in distributed applications with VAIF
abstract
Developers use logs to diagnose performance problems in distributed applications. However, it is difficult to know a priori where logs are needed and what information in them is needed to help diagnose problems that may occur in the future. We present the Variance-driven Automated Instrumentation Framework (VAIF), which runs alongside distributed applications. In response to newly-observed performance problems, VAIF automatically searches the space of possible instrumentation choices to enable the logs needed to help diagnose them. To work, VAIF combines distributed tracing (an enhanced form of logging) with insights about how response-time variance can be decomposed on the critical-path portions of requests' traces. We evaluate VAIF by using it to localize performance problems in OpenStack and HDFS. We show that VAIF can localize problems related to slow code paths, resource contention, and problematic third-party code while enabling only 3-34% of the total tracing instrumentation.
Mert Toslali, Emre Ates, Alex Ellis, Darby Huye, Samantha Puterman, Ayse K. Coskun, Raja R. Sambasivan
SoCC2
2019 An automated, cross-layer instrumentation framework for diagnosing performance problems in distributed applications
abstract
Diagnosing performance problems in distributed applications is extremely challenging. A significant reason is that it is hard to know where to place instrumentation a priori to help diagnose problems that may occur in the future. We present the vision of an automated instrumentation framework, Pythia, that runs alongside deployed distributed applications. In response to a newly-observed performance problem, Pythia searches the space of possible instrumentation choices to enable the instrumentation needed to help diagnose it. Our vision for Pythia builds on workflow-centric tracing, which records the order and timing of how requests are processed within and among a distributed application's nodes (i.e., records their workflows). It uses the key insight that localizing the sources high performance variation within the workflows of requests that are expected to perform similarly gives insight into where additional instrumentation is needed.
Emre Ates, Lily Sturmann, Mert Toslali, Orran Krieger, Richard Megginson, Ayse K. Coskun, Raja R. Sambasivan
SoCC1
2019 HPAS: An HPC Performance Anomaly Suite for Reproducing Performance Variations
abstract
Modern high performance computing (HPC) systems, including supercomputers, routinely suffer from substantial performance variations. The same application with the same input can have more than 100% performance variation, and such variations cause reduced efficiency and wasted resources. There have been recent studies on performance variability and on designing automated methods for diagnosing "anomalies" that cause performance variability. These studies either observe data collected from HPC systems, or they rely on synthetic reproduction of performance variability scenarios. However, there is no standardized way of creating performance variability inducing synthetic anomalies; so, researchers rely on designing ad-hoc methods for reproducing performance variability.
Emre Ates, Yijia Zhang 0002, Burak Aksar, Jim M. Brandt, Vitus J. Leung, Manuel Egele, Ayse K. Coskun
ICPP1
2019 Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning
abstract
As the size and complexity of high performance computing (HPC) systems grow in line with advancements in hardware and software technology, HPC systems increasingly suffer from performance variations due to shared resource contention as well as software- and hardware-related problems. Such performance variations can lead to failures and inefficiencies, which impact the cost and resilience of HPC systems. To minimize the impact of performance variations, one must quickly and accurately detect and diagnose the anomalies that cause the variations and take mitigating actions. However, it is difficult to identify anomalies based on the voluminous, high-dimensional, and noisy data collected by system monitoring infrastructures. This paper presents a novel machine learning based framework to automatically diagnose performance anomalies at runtime. Our framework leverages historical resource usage data to extract signatures of previously-observed anomalies. We first convert collected time series data into easy-to-compute statistical features. We then identify the features that are required to detect anomalies, and extract the signatures of these anomalies. At runtime, we use these signatures to diagnose anomalies with negligible overhead. We evaluate our framework using experiments on a real-world HPC supercomputer and demonstrate that our approach successfully identifies 98 percent of injected anomalies and consistently outperforms existing anomaly diagnosis techniques.
Ozan Tuncer, Emre Ates, Yijia Zhang 0002, Ata Turk, Jim M. Brandt, Vitus J. Leung, Manuel Egele, Ayse K. Coskun
IEEE Trans. Parallel Distributed Syst.2
2018 Taxonomist: Application Detection Through Rich Monitoring Data
Emre Ates, Ozan Tuncer, Ata Turk, Vitus J. Leung, Jim M. Brandt, Manuel Egele, Ayse K. Coskun
Euro-Par1