VLDB 2026 Research / reviewers in the wild / expert
Naser Ezzati-Jivan
dblp:54/10412
· DBLP profile ↗
33ranked-venue papers
3as first author
26since 2021 · last 2026
0000-0003-1435-6297ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 19 · 1 first-author · 16 since 2021Systems, architecture and hardware · 6 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | B-Perf: Black-box Performance Antipattern Detection Using System-level Execution TracingabstractPerformance antipatterns capture recurring behaviours that degrade software efficiency. Black-box approaches aim to detect such issues without modifying the application. This paper presents B-Perf, a system-level black-box method that reconstructs execution, memory, and messaging behaviour from kernel-level traces. By analysing scheduling, allocation, and communication events, B-Perf derives workload-dependent behavioural trends and reports antipattern indicators grounded in resource usage and contention. To handle large trace volumes, the approach follows a pipeline of workload generation, event gathering, trace handling, and antipattern inference. Morteza Noferesti, Mahsa Panahandeh, Naser Ezzati-Jivan |
ICPE | 3 |
| 2026 | LMAT: An adaptive tracing approach based on efficient system behavior analysis using language modelsabstractWe introduce LMAT, a Language Model-based Adaptive Tracing framework designed for host-level observability that provides granular monitoring without excessive overhead. LMAT leverages a multi-task architecture to jointly predict kernel event sequences and classify event durations, thereby capturing both control-flow and temporal dynamics. By continuously comparing live trace data against model predictions, LMAT automatically signals deviations, dynamically adjusting trace granularity only when needed. This approach significantly reduces trace volume, along with associated energy and storage costs, achieving a 70.6% reduction in our experiments. Additionally, LMAT utilizes prediction discrepancies to drive an efficient root-cause classifier, mapping detected anomalies directly to their potential fault sources and providing actionable feedback for operations teams. We evaluate LMAT on two architecturally distinct single-host environments—an Apache2 web-server stack and the Sock Shop containerized microservice benchmark—using kernel traces that include standard workloads, duration-centric noise scenarios, and controlled CPU, disk, memory, and network stress injections. On the Apache workload, LMAT demonstrates up to 97.7% accuracy in anomaly detection and root-cause identification, surpassing state-of-the-art methods relying solely on event sequences. On Sock Shop, the same design remains effective for host-local change detection, while root-cause attribution in the microservice setting remains more challenging. A deployment-oriented overhead study shows that under a stable load, asynchronous LMAT inference introduces no measurable additional tail-latency overhead beyond tracing, while maintaining consistent throughput. Our findings illustrate that LMAT is a practical approach for adaptive tracing in the evaluated single-host environments, improving detection quality while keeping deployment overhead negligible. Kasra Darvishi, Morteza Noferesti, Yuvraj Sehgal, Naser Ezzati-Jivan |
J. Syst. Softw. | 4 |
| 2026 | CARE: Context Aware Root Cause Identification Using Distributed Traces and Profiling MetricsabstractRoot cause localization in microservices is challenging due to intricate service dependencies and the high volume and heterogeneity of collected monitoring data, which add complexity to the analysis. Conventional methods often overlook nuanced propagation patterns and contextual interactions among services, and they are limited in leveraging multi-source observability data for comprehensive root cause identification. This study introduces CARE, a context-aware, spectrum-analysis-based approach that integrates multi-source observability data and employs network analysis to prioritize the contextual significance of components in propagating anomalies across individual services, service communities, and requests. CARE’s weighted spectrum analysis leverages these prioritized contexts to pinpoint underlying performance issues. Evaluations on 224 cases from the TrainTicket benchmark and a real-world Internet service provider’s production system demonstrate CARE’s substantial accuracy gains, with top-1 accuracy of 72%-89% and top-5 accuracy of 84%-99% for single root causes, outperforming baselines by 8%-41%. CARE also shows significant improvements in dual root cause identification, exceeding baseline performance by 18%-37%, all while maintaining efficient resource usage, establishing CARE as a robust and resource-effective solution for root cause localization in complex microservice environments. Mahsa Panahandeh, Naser Ezzati-Jivan, Abdelwahab Hamou-Lhadj, James Miller 0001 |
IEEE Trans. Software Eng. | 2 |
| 2025 | Execution Trace Reconstruction Using Diffusion-Based Generative ModelsabstractExecution tracing is essential for understanding system and software behaviour, yet lost trace events can significantly compromise data integrity and analysis. Existing solutions for trace reconstruction often fail to fully leverage available data, particularly in complex and high-dimensional contexts. Recent advancements in generative artificial intelligence, particularly diffusion models, have set new benchmarks in image, audio, and natural language generation. This study conducts the first comprehensive evaluation of diffusion models for reconstructing incomplete trace event sequences. Using nine distinct datasets generated from the Phoronix Test Suite, we rigorously test these models on sequences of varying lengths and missing data ratios. Our results indicate that the SSSDS4model, in particular, achieves superior performance, in terms of accuracy, perfect rate, and ROUGE-L score across diverse imputation scenarios. These findings underscore the potential of diffusion-based models to accurately reconstruct missing events, thereby maintaining data integrity and enhancing system monitoring and analysis. Madeline Janecek, Naser Ezzati-Jivan, Abdelwahab Hamou-Lhadj |
ICSE | 2 |
| 2025 | HybridRCA: Lightweight Critical-Path-Aware Hybrid Tracing for Root-Cause Analysis in Production Microservicesabstract[Context] Distributed cloud-native systems operated by our industrial partners, including Ericsson and Ciena, generate millions of trace spans daily. Capturing and analyzing this data at full granularity is infeasible due to excessive storage and computational overhead. [Objective] We aim to enable fast and accurate RCA with minimal trace volume and system overhead, quickly pinpointing the service causing a latency spike, making it practical for large-scale production environments. [Method] We present HybridRCA, a critical-path-aware RCA pipeline that (1) extracts the critical path of each request, (2) applies a PageRank-weighted spectrum analysis to identify suspicious spans, and (3) collects system metrics only for targeted spans. [Results] Across three microservice benchmarks (HotRod, TrainTicket, OnlineBoutique), HybridRCA improves recall by an average of$\text{0.45 \%}$over the best existing methods, while analyzing up to 22.6 % fewer spans and reducing kernel-level storage usage by over 99%. [Significance] HybridRCA addresses key observability challenges faced by our industry partners, enabling scalable, low-overhead RCA in real-world distributed systems. Maryam Ekhlasi, Arnaud Fiorini, Michel R. Dagenais, Naser Ezzati-Jivan, Maxime Lamothe |
ICSME | 4 |
| 2025 | Developing a Taxonomy for Advanced Log Parsing TechniquesabstractLogs are widely used in various software engineering applications, including debugging, program comprehension, failure prediction, and anomaly detection. Despite their value, the unstructured nature of logs complicates the extraction of meaningful insights. In response, various log parsing techniques leveraging methods like machine learning and pattern recognition have been developed. Nevertheless, existing parsers frequently fail to achieve consistent accuracy, especially when handling complex log formats. To address this challenge, we conduct a comprehensive study to understand the characteristics of log events that lead to parsing errors. Using 16 different log datasets and 8 log parsers, we apply open coding techniques to derive a taxonomy of log event characteristics that contribute to parsing errors. We also examine how different log parsers are impacted by each category in the taxonomy. The resulting taxonomy not only provides insights into the complexity of parsing log data but can also guide the development of advanced parsing tools capable of handling the unique characteristics of diverse log formats. Issam Sedki, Abdelwahab Hamou-Lhadj, Otmane Aït Mohamed, Naser Ezzati-Jivan |
ICPC | 4 |
| 2025 | Utilizing Graph Neural Networks for Effective Link Prediction in Microservice ArchitecturesabstractManaging microservice architectures in distributed systems is complex and resource-intensive due to the high frequency and dynamic nature of inter-service interactions. Accurate prediction of these future interactions can enhance adaptive monitoring, enabling proactive maintenance and resolution of potential performance issues before they escalate. This study introduces a Graph Neural Network (GNN)-based approach, specifically using a Graph Attention Network (GAT), for link prediction in microservice Call Graphs. Unlike social networks, where interactions tend to occur sporadically and are often less frequent, microservice Call Graphs involve highly frequent and time-sensitive interactions that are essential to operational performance. Ghazal Khodabandeh, Alireza Ezaz, Majid Babaei, Naser Ezzati-Jivan |
ICPE | 4 |
| 2025 | Optimization Strategies for Enhancing Resource Efficiency in Transformers & Large Language ModelsabstractAdvancements in Natural Language Processing are heavily reliant on Transformer architectures, whose improvements come at substantial resource costs due to ever-growing model sizes. This study explores optimization techniques, including quantization, knowledge distillation, and pruning, focusing on energy and computational efficiency while retaining performance. Among standalone methods, 4-Bit quantization significantly reduces energy use with minimal accuracy loss. Hybrid approaches, like NVIDIA's Minitron approach combining KD and structured pruning, further demonstrate promising trade-offs between size reduction and accuracy retention. A novel optimization framework is introduced, offering a flexible framework for comparing various methods. Through the investigation of these compression methods, we provide valuable insights for developing more sustainable and efficient LLMs, shining a light on the often-ignored concern of energy efficiency. Tom Wallace, Beatrice M. Ombuki-Berman, Naser Ezzati-Jivan |
ICPE | 3 |
| 2024 | Picturing Ambiguity: A Visual Twist on the Winograd Schema ChallengeabstractLarge Language Models (LLMs) have demonstrated remarkable success in tasks like the Winograd Schema Challenge (WSC), showcasing advanced textual common-sense reasoning.However, applying this reasoning to multimodal domains, where understanding text and images together is essential, remains a substantial challenge.To address this, we introduce WINOVIS, a novel dataset specifically designed to probe text-to-image models on pronoun disambiguation within multimodal contexts.Utilizing GPT-4 for prompt generation and Diffusion Attentive Attribution Maps (DAAM) for heatmap analysis, we propose a novel evaluation framework that isolates the models' ability in pronoun disambiguation from other visual processing challenges.Evaluation of successive model versions reveals that, despite incremental advancements, Stable Diffusion 2.0 achieves a precision of 56.7% on WINOVIS, showing minimal improvement from past iterations and only marginally surpassing random guessing.Further error analysis identifies important areas for future research aimed at advancing text-to-image models in their ability to interpret and interact with the complex visual world. Brendan Park, Madeline Janecek, Naser Ezzati-Jivan, Ali Emami |
ACL (1) | 3 |
| 2024 | Assessing Predictive Models for Energy Consumption Across Varied Software EnvironmentsabstractThis study contributes to a deeper understanding of energy consumption in software applications, emphasizing the critical need for energy efficiency. We focus on integrating performance counter events and system call data to build energy predictive models using advanced machine learning techniques, including linear regression, multi-layer perceptrons, and random forests. These models are carefully calibrated against empirical energy measurements obtained through the Perf framework. Our study addresses variability in model outcomes that stem from differences in feature selection and the inherent discrepancies of operating systems. Through various experimentation, we demonstrate that our models robustly predict energy consumption across diverse scenarios, with particularly promising results in unseen datasets. However, challenges persist in cross-application efficacy. Event-based models particularly stand out, offering reliable energy estimations in novel applications. This research validates the effectiveness of our methodologies and also illuminates the complex landscape of precise energy consumption modeling in contemporary software environments. Sarwat Islam Dipanzan, Leila Tahmooresnejad, Naser Ezzati-Jivan |
IEEE Big Data | 4 |
| 2024 | An Adaptive Logging System (ALS): Enhancing Software Logging with Reinforcement Learning TechniquesabstractThe efficient management of software logs is crucial in software performance evaluation, enabling detailed examination of runtime information for postmortem analysis. Recognizing the importance of logs and the challenges developers face in making informed log-placement decisions, there is a clear need for a robust log-placement framework that supports developers. Existing frameworks, however, are limited by their inability to adapt to customized logging objectives, a concern highlighted by our industrial partner, Ciena, who required a system for their specific logging goals in resource-limited environments like routers. Moreover, these frameworks often show poor cross-project consistency. This study introduces a novel performance logging objective designed to uncover potential performance-bugs, categorized into three classes-Loops, Synchronization, and API Misuses-and defines 12 source code features for their detection. We present an Adaptive Logging System (ALS), based on reinforcement learning, which adjusts to specified logging objectives, particularly for identifying performance-bugs. This framework, not restricted to specific projects, demonstrates stable cross-project performance. We trained and evaluated ALS on Python source code from 17 diverse open-source projects within the Apache and Django ecosystems. Our findings suggest that ALS has the potential to significantly enhance current logging practices by providing a more targeted, efficient, and context-aware logging approach, particularly beneficial for our industry partner who requires a flexible system that adapts to varied performance objectives and logging needs in their unique operational environments. Amirmahdi Khosravi Tabrizi, Naser Ezzati-Jivan, François Tetreault |
ICPE | 2 |
| 2024 | Enhancing empirical software performance engineering research with kernel-level events: A comprehensive system tracing approach
Morteza Noferesti, Naser Ezzati-Jivan |
J. Syst. Softw. | 2 |
| 2023 | AltOOM: A Data-driven Out of Memory Root Cause Identification StrategyabstractResource-constrained devices face significant performance challenges when encountering memory pressure situations due to limited hardware resources. Existing approaches mainly focus on reactive and instantaneous approaches, but they often fail to accurately identify the root cause of memory pressure, resulting in delayed and ineffective response strategies. In this paper, we address this limitation by proposing an alternative data-driven approach to proactively detect memory pressure and identify the responsible process in resource-limited devices. Our method enables the activation and deactivation of extended process-level profiling based on the predicted memory pressure, facilitating the identification of the root cause process. Through evaluation, we achieved an 85% accuracy in forecasting memory pressure situations and correctly identified the responsible process in 83% of use-cases. These results demonstrate the effectiveness of our strategy to stream large amount of trace data in mitigating memory pressure issues in resource-constrained systems. This approach has the potential to enhance system performance and improve overall system architecture in such devices. Pranjal Chakraborty, Naser Ezzati-Jivan, Seyed Vahid Azhari, François Tetreault |
IEEE Big Data | 2 |
| 2023 | Towards a Classification of Log Parsing ErrorsabstractLog parsing is used to extract structures from unstructured log data. It is a key enabler for many software engineering tasks including debugging, fault diagnosis, and anomaly detection. In recent years, we have seen an increase in the number of log parsing techniques and tools. The accuracy of these tools varies significantly. To improve log parsing tools, we need to understand the type of parsing errors they make, which is the purpose of this early research track paper. We achieve this by examining errors of four leading log parsing tools when applied to the parsing of four log datasets generated from various systems. Based on this analysis, we suggest a preliminary classification of log parsing errors, which contains nine categories of errors. We believe that this classification is a good starting point for improving the accuracy of log parsing tools, and also defining better logging practices. Issam Sedki, Abdelwahab Hamou-Lhadj, Otmane Aït Mohamed, Naser Ezzati-Jivan |
ICPC | 4 |
| 2023 | PASD: A Performance Analysis Approach Through the Statistical Debugging of Kernel EventsabstractDynamic performance analysis plays a crucial role in optimizing systems and identifying performance bottlenecks. Traditional software debugging methods frequently encounter difficulties when trying to pinpoint performance problems in complex software settings. This is often because performance issues remain hidden during the code execution within debugging tools or under certain run-time circumstances, making them challenging to identify and address. This paper introduces PASD (Performance Analysis through Statistical Debugging), a dynamic performance analysis approach based on statistical debugging of kernel-level trace events. Importantly, this approach requires no application code instrumentation and purely utilizes operating system kernel trace events for analysis. PASD collects kernel trace events generated during software execution and utilizes heuristics to analyze their performance issues and the root-causes. Through statistical debugging techniques, PASD identifies the most important functions correlated with performance problems. It notably does so without disrupting the software’s normal functions and ensuring that any issues are detected in the software’s typical operating conditions, thus avoiding additional complexity in the debugging process. We have conducted two empirical studies to assess the effectiveness of PASD on performance issues in the Firefox web browser as well as the ‘ls’ tool (a common utility in Unix-like systems). Our experiments demonstrate that PASD successfully identifies performance issues and their causes in software without prior knowledge of the architecture or source code instrumentation. By providing an overview of software behavior through the kernel-level, our proposed method can aid developers and testers in quickly pinpointing performance problems in the source code. This, in turn, can result in improved software quality, increased user satisfaction, and the prevention of critical system failures. Mohammed Adib Khan, Morteza Noferesti, Naser Ezzati-Jivan |
SCAM | 3 |
| 2023 | Multi-level Adaptive Execution Tracing for Efficient Performance AnalysisabstractTroubleshooting system performance issues is a challenging task that requires a deep understanding of various factors that may impact system performance. This process involves analyzing trace logs from the kernel and user space using tools such as ftrace, strace, DTrace, or LTTng. However, pre-set tracing instrumentation can lead to missing important data where not enough components of the system include observability coverage. Also, having too much coverage may result in unnecessary noise in the data, making it extremely difficult to debug. This paper proposes an adaptive instrumentation technique for execution tracing, which dynamically makes decisions not only for which components to trace but also when to trace, thus reducing the risk of missing important data related to the performance problem and increasing the accuracy of debugging by reducing unwanted noises. Our case study results show that the proposed method is capable of handling tracing instrumentation dynamically for both kernel and application levels while maintaining a low overhead. Mohammed Adib Khan, Naser Ezzati-Jivan |
SERA | 2 |
| 2023 | EMD-SCS: A Dynamic Behavioral Approach for Early Malware Detection with Sonification of System Call SequencesabstractThe privacy and security of users are increasingly threatened due to the rising frequency of malware assaults. Both Host-based Intrusion Detection Systems (HIDSes) and Antivirus software rely on signature-based or anomaly-based techniques for malware detection. However, the escalating diversity and sophistication of malware pose significant obstacles. In this research, we introduce EMD-SCS, an early malware detection methodology, employing sonification and system call sequence analysis. In our methodology, we interpret an executing program/process as a sequence of system calls, leveraging a Long Short-Term Memory network (LSTM) to hold a record of preceding system calls within this chain, consequently facilitating the prediction of future calls. After this prediction phase, the BLEU and hamming distance scores are utilized to classify the system call sequence. Importantly, these results are attained by analyzing just a small segment of the data for early prediction, which is crucial for a sonification-based approach as it enables us to notify administrators in advance of potential threats. This early warning system would allow admins to protect the host before a potential compromise. EMD-SCS uses sonification to convey the prediction outcomes using natural and animal sounds, offering a broader monitoring scope than visual observation. Evaluation results from the ADFA-LD dataset suggest that EMD-SCS surpasses prior techniques in early malware detection with an accuracy of 91.2%, a detection rate of 87.7%, and a false-positive rate of 15.3%, achieved by only processing 40% of the input system call sequences before they infiltrate the host. Raghav Bhardwaj, Morteza Noferesti, Madeline Janecek, Naser Ezzati-Jivan |
TrustCom | 4 |
| 2022 | Poster Paper: Operating System Support for Applications Performance AnalysisabstractIn order to classify common performance issues of multi-core applications, used in cloud computing and distributed systems, and offer solutions to them, performance antipatterns have been introduced by researchers. Performance antipatterns help developers refactor inefficient code, and are exceptionally useful for multi-threaded applications, where problems can be difficult to diagnose. However, existing performance antipattern detection methods do not properly examine operating system-wide resources, leading to imprecise metrics and results. In this paper, a novel system-level execution tracing method is presented for detecting the One Lane Bridge performance antipattern. This method is validated through a case study performed on an open-source multi-threaded application, where we diagnosed performance issues. Riley VanDonge, Naser Ezzati-Jivan |
IC2E | 2 |
| 2022 | Performance anomaly detection through sequence alignment of system-level tracesabstractIdentifying and diagnosing performance anomalies is essential for maintaining software quality, yet it can be a complex and time-consuming task. Low level kernel events have been used as an excellent data source to monitor performance, but raw trace data is often too large to easily conduct effective analyses. To address this shortcoming, in this paper, we propose a framework for uncovering performance problems using execution critical path data. A critical path is the longest execution sequence without wait delays, and it can provide valuable insight into a program's internal and external dependencies. Upon extracting this data, course grained anomaly detection techniques are employed to determine if a finer grained analysis is required. If this is the case, the critical paths of individual executions are grouped together with machine learning clustering to identify different execution types, and outlying anomalies are identified using performance indicators. Finally, multiple sequence alignment is used to pinpoint specific abnormalities in the identified anomalous executions, allowing for improved application performance diagnosis and overall program comprehension. Madeline Janecek, Naser Ezzati-Jivan, Abdelwahab Hamou-Lhadj |
ICPC | 2 |
| 2022 | N-Lane Bridge Performance Antipattern Analysis Using System-Level Execution TracingabstractPerformance problems caused by the improper use of multi-threading can be incredibly difficult to diagnose. There are countless resources that could introduce latency into an application when multiple cooperating threads interact improperly. As a matter of program comprehension, it is crucial to know which resources are being misused by the program causing that program to run slower. The concept of performance antipatterns has been introduced in order to classify common performance problems and bundle them with a solution. The One Lane Bridge (OLB) antipattern in particular deals with latency due to the incorrect use of multi-threading. However, existing methods to detect the OLB antipattern do not consider latency caused by active resources and use imprecise metrics. In this paper, we present a new category of OLB, the N-Lane Bridge antipattern, to cover situations of latency caused by the overuse of active resources. Moreover, a novel system-level execution tracing approach is presented to detect both categories of OLB antipatterns. As a proof-of-concept, we applied our approach to the popular Firefox web browser application and we were able to identify several OLB antipatterns, enabling us to diagnose and understand a critical performance issue. Riley VanDonge, Naser Ezzati-Jivan |
SCAM | 2 |
| 2022 | Execution trace-based model verification to analyze multicore and real-time systemsabstractAbstract As a key part of model‐driven development, modeling allows users to represent the application workflow or to automatically generate source code. This is convenient for developers, particularly to create or improve real‐time applications embedded in complex systems. Multicore systems are difficult to debug because the concurrently running processes can interfere with each other. In real‐time systems, timing constraints add to the complexity, invalidating results when a deadline is missed. Tracing is usually the most accurate and reliable tool to study the runtime behaviour of those applications. However, the interpretation of voluminous detailed execution traces requires a deep understanding of the operating system and application behaviour, and time to dig through the millions of trace events.In this paper, we present the use of model‐based constraints on top of user‐space and kernel traces to provide weighted analysis results. Our algorithms have been applied to multiple traces showing common problems for multi‐core real‐time systems. The experimental results show that our algorithms can quickly identify many different types of problems with a low runtime, even for traces with millions of events, thus helping to save time when analyzing thousands of trace events for complex systems. Raphaël Beamonte, Naser Ezzati-Jivan, Michel R. Dagenais |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Performance evaluation of complex multi-thread applications through execution path analysis
Majid Rezazadeh, Naser Ezzati-Jivan, Seyed Vahid Azhari, Michel R. Dagenais |
Perform. Evaluation | 2 |
| 2021 | Efficient heap monitoring tool for memory leak detection and root-cause analysisabstractMemory leaks are of serious concern in software programs written in languages that do not have a dedicated garbage collector. Memory leaks are the result of prolonged unnecessary usage of the system memory resources. This increases the workload, paging, and lowers the response rate leading to performance degradation of the system or software. Such leaks are difficult to detect due to a software’s development testing span and different working environments.This research paper proposes an algorithm to detect memory leaks based on the growth analysis of the memory blocks. The proposed method captures the allocation made in memory blocks via malloc, calloc, and realloc and stores the captured results onto the files. The files are then analyzed using the leak threshold determined by the developer or tester. Unlike other approaches such as ASAN/LSAN which require multiple snapshots to analyze the leaks, the proposed method works without a snapshot for the summary analysis algorithm and only requires one snapshot, collected at the end of the execution. This approach does not need to be integrated within the application to analyze the software and detect memory leaks. Our method provides more accurate results by visualizing the captured and analyzed data for the developers, making it more convenient to detect the root cause of the leakage. Seyed Vahid Azhari, Simar Bhamra, Naser Ezzati-Jivan, François Tetreault |
IEEE BigData | 3 |
| 2021 | Container Workload Characterization Through Host System TracingabstractThe use of containers within cloud environments has become increasing popular due to their lightweight nature, scalability, and efficiency. However, as containers share their host's resources, advanced resource management techniques are essential to avoid performance impacting resource contention. Coarse measures such as CPU, disk, and network usage collected from internal agents are often considered, yet these methods may be improved upon to garner a more precise view of container workloads. In this paper, we present a container workload characterization method using host system tracing. Features derived from thread execution states are taken from tracing data to reveal container runtime behaviour. A PageRank-based algorithm is then used to identify the most significant threads for further analysis. Once this data is collected and vectorized, a two stage K-Means clustering technique is used to generate groups of containers with similar workloads. This eliminates the need for manual analysis of individual containers, and instead allows administrators to view and address container behaviours collectively. Experimental results show that our methodology can identify a variety of execution behaviours. Administrators may use these results to remove idle containers to free up system resources. Moreover, they may identify clusters of containers that are at risk of resource contention, allowing for more effective resource assignment. Madeline Janecek, Naser Ezzati-Jivan, Seyed Vahid Azhari |
IC2E | 2 |
| 2021 | Integrated modeling tool for indexing and analyzing state machine traceabstractIt is important to model and understand an application or system runtime behavior to identify potential performance problems. Execution tracing, the basis of various dynamic analysis methods includes the collection of events, metrics, and statistics about the runtime behaviors of systems and applications. However, comprehensive execution tracing can result in very large trace files, most of which are irrelevant to the problem at hand. This is compounded by the inflexibility and complexity of common tools in how the user specifies what to capture, making the collection of relevant statistics difficult. While existing solutions allow for an adaptive collection of metrics and statistics, they often require users to write large and complex scripts in a domain-specific language. In this paper, we propose a state machine based modeling tool that simplifies the creation of user-defined and data-driven trace-based analyses. The proposed method combines advanced kernel-space and user-space execution trace events with powerful and adaptable modeling in order to automatically generating event-based analysis based on users’ specific requirements and problems. The difficulty and complexity of user-defined event tracing is drastically reduced. We demonstrate the efficiency, effectiveness, and simplicity of our proposed tool through real use cases of multi-level dynamic execution tracing in the Linux kernel. Simon Delisle, Naser Ezzati-Jivan, Michel R. Dagenais |
ISNCC | 2 |
| 2021 | Automated Cause Analysis of Latency Outliers Using System-Level Dependency GraphsabstractDetecting performance issues and identifying their root causes in the runtime is a challenging task. Typically, developers use methods such as logging and tracing to identify bottlenecks. These solutions are, however, not ideal as they are time-consuming and require manual effort. In this paper, we propose a method to automate the task of detecting latency outliers using system-level traces and then comparing them to identify the root cause(s). Our method makes use of dependency graphs to show internal interactions between threads and system resources. With these graphs, one can pinpoint where performance issues occur. However, a single trace can be composed of a large number of requests, each generating one graph. To automate the task of identifying outliers within the dataset, we use machine learning density-based models and statistical calculations such as$Z$-score. Our evaluation shows an accuracy greater than 97 % on outlier detection, making them appropriate for in-production servers and industry-level use cases. Sneh Patel, Brendan Park, Naser Ezzati-Jivan, Quentin Fournier |
QRS | 3 |
| 2020 | DepGraph: Localizing Performance Bottlenecks in Multi-Core Applications Using Waiting Dependency Graphs and Software TracingabstractThis paper addresses the challenge of understanding the waiting dependencies between the threads and hardware resources required to complete a task. The objective is to improve software performance by detecting the underlying bottlenecks caused by system-level blocking dependencies. In this paper, we use a system level tracing approach to extract a Waiting Dependency Graph that shows the breakdown of a task execution among all the interleaving threads and resources. The method allows developers and system administrators to quickly discover how the total execution time is divided among its interacting threads and resources. Ultimately, the method helps detecting bottlenecks and highlighting their possible causes. Our experiments show the effectiveness of the proposed approach in several industry-level use cases. Three performance anomalies are analysed and explained using the proposed approach. Evaluating the method efficiency reveals that the imposed overhead never exceeds 10.1%, therefore making it suitable for in-production environments. Naser Ezzati-Jivan, Quentin Fournier, Michel R. Dagenais, Abdelwahab Hamou-Lhadj |
SCAM | 1 |
| 2019 | Efficient large-scale heterogeneous debugging using dynamic tracingabstractHeterogeneous multi-core and many-core processors are increasingly common in personal computers and industrial systems. Efficient software development on these platforms needs suitable debugging tools, beyond traditional interactive debuggers . An alternative, to interactively follow the execution flow of a program, is tracing within the debugging environment , as long as the tracer has a minimal overhead. In this paper, the dynamic tracing infrastructure of GNU debugger (GDB) was investigated to understand its performance limitations. Thereafter, we propose an improved architecture for dynamic tracing on many-core processors within GDB, and demonstrate its scalability on highly parallel platforms. In addition, the scalability of the thread data collection and presentation component was studied and new views were proposed within the Eclipse Debugging Service Framework and the Trace Compass visualization tool. With these scalability enhancements, debuggers such as GDB can more efficiently help debugging multi-threaded programs on heterogeneous many-core processors composed of multi-core CPUs, and GPUs containing thousands of cores. Didier Nadeau, Naser Ezzati-Jivan, Michel R. Dagenais |
J. Syst. Archit. | 2 |
| 2017 | Multi-scale navigation of large trace data: A surveyabstractSummary Dynamic analysis through execution traces is frequently used to analyze the runtime behavior of software systems. However, tracing long running executions generates voluminous data, which are complicated to analyze and manage. Extracting interesting performance or correctness characteristics out of large traces of data from several processes and threads is a challenging task. Trace abstraction and visualization are potential solutions to alleviate this challenge. Several efforts have been made over the years in many subfields of computer science for trace data collection, maintenance, analysis, and visualization. Many analyses start with an inspection of an overview of the trace, before digging deeper and studying more focused and detailed data. These techniques are common and well supported in geographical information systems, automatically adjusting the level of details depending on the scale. However, most trace visualization tools operate at a single level of representation, which are not adequate to support multilevel analysis. Sophisticated techniques and heuristics are needed to address this problem. Multi‐scale (multilevel) visualization with support for zoom and focus operations is an effective way to enable this kind of analysis. Considerable research and several surveys are proposed in the literature in the field of trace visualization. However, multi‐scale visualization has yet received little attention. In this paper, we provide a survey and methodological structure for categorizing tools and techniques aiming at multi‐scale abstraction and visualization of execution trace data and discuss the requirements and challenges faced to be able to meet evolving user demands. Naser Ezzati-Jivan, Michel R. Dagenais |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Hardware-assisted software event tracingabstractSummary Event tracing is a reliable and a low‐intrusiveness method to debug and optimize systems and processes. Low overhead is particularly important in embedded systems where resources and energy consumption is critical. The most advanced tracing infrastructures achieve a very low footprint on the traced software, bringing each tracepoint overhead to less than a microsecond. To reduce this still non‐negligible impact, the use of dedicated hardware resources is promising. In this paper, we propose complementary methods for tracing that rely on hardware modules to assist software tracing. We designed solutions to take advantage of CoreSight STM, CoreSight ETM, and Intel BTS, which are present on most newer ARM‐based systems‐on‐chip and Intel x86 processors. Our results show that the time overhead for tracing can be reduced by up to 10 times when assisted by hardware, as compared to software tracing with LTTng, a high‐performance tracer for Linux. We also propose a modification to the Perf tool to speed BTS execution tracing up to 65%. Adrien Vergé, Naser Ezzati-Jivan, Michel R. Dagenais |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | A declarative framework for stateful analysis of execution traces
Florian Wininger, Naser Ezzati-Jivan, Michel R. Dagenais |
Softw. Qual. J. | 2 |
| 2015 | Cube data model for multilevel statistics computation of live execution tracesabstractSummary Execution trace logs are used to analyze system run‐time behaviour and detect problems. Trace analysis tools usually read the input logs and gather either a detailed or brief summary of them to later process and inspect in the analysis steps. However, continuous and lengthy trace streams contained in the live tracing mode make it difficult to indefinitely record all events or even a detailed summary of the whole stream. This situation is further complicated when the system aims to compare different parts of the trace and provide a multilevel and multidimensional analysis. This paper presents an architecture with corresponding data structures and algorithms to process stream events, generate an adequate summary—detailed enough for recent data and succinct enough for old data—and organize them to enable an efficient multilevel and multidimensional analysis, similar to online analytical processing analyses in the database applications. The proposed solution arranges data in a compact manner using interval forms and enables the range queries for any arbitrary time durations. Because this feature makes it possible to compare of different system parameters in different time areas, it significantly influences the system's ability to provide a comprehensive trace analysis. Although the Linux operating system trace logs are used to evaluate the solution, we propose a generic architecture that can be used to summarize various types of stream data. Copyright © 2014 John Wiley & Sons, Ltd. Naser Ezzati-Jivan, Michel R. Dagenais |
Concurr. Comput. Pract. Exp. | 1 |
| 2013 | Efficient Model to Query and Visualize the System States Extracted from Trace Data
Alexandre Montplaisir, Naser Ezzati-Jivan, Florian Wininger, Michel R. Dagenais |
RV | 2 |