EDBT 2026 Demo / reviewers in the wild / expert
Abdelwahab Hamou-Lhadj
dblp:70/2136 · also Wahab Hamou-Lhadj
· DBLP profile ↗
80ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0002-3319-5006ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 68 · 3 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Systems, architecture and hardware · 4 · 1 since 2021Security and privacy · 2Computer networks · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CARE: Context Aware Root Cause Identification Using Distributed Traces and Profiling MetricsabstractRoot cause localization in microservices is challenging due to intricate service dependencies and the high volume and heterogeneity of collected monitoring data, which add complexity to the analysis. Conventional methods often overlook nuanced propagation patterns and contextual interactions among services, and they are limited in leveraging multi-source observability data for comprehensive root cause identification. This study introduces CARE, a context-aware, spectrum-analysis-based approach that integrates multi-source observability data and employs network analysis to prioritize the contextual significance of components in propagating anomalies across individual services, service communities, and requests. CARE’s weighted spectrum analysis leverages these prioritized contexts to pinpoint underlying performance issues. Evaluations on 224 cases from the TrainTicket benchmark and a real-world Internet service provider’s production system demonstrate CARE’s substantial accuracy gains, with top-1 accuracy of 72%-89% and top-5 accuracy of 84%-99% for single root causes, outperforming baselines by 8%-41%. CARE also shows significant improvements in dual root cause identification, exceeding baseline performance by 18%-37%, all while maintaining efficient resource usage, establishing CARE as a robust and resource-effective solution for root cause localization in complex microservice environments. Mahsa Panahandeh, Naser Ezzati-Jivan, Abdelwahab Hamou-Lhadj, James Miller 0001 |
IEEE Trans. Software Eng. | 3 |
| 2025 | Mining Five Years of Actively Exploited VulnerabilitiesabstractCompanies and researchers rely heavily on the National Vulnerability Database (NVD) in order to be cognizant of the threat environment they face. However, research shows that only about 5% of reported vulnerabilities are eventually exploited. Our study compares exploited and non-exploited vulnerabilities to provide valuable insights for advancing research on effective vulnerability prioritization. To achieve this, we compile a database of over 4,000 exploited vulnerabilities, spanning 5 years, from 2019 to 2023. We further mine this dataset to uncover trends and patterns that relate to how exploited vulnerabilities differ from non-exploited ones and how exploited vulnerabilities evolved over the 5-year spans of our study. We found that exploited vulnerabilities differ from non-exploited vulnerabilities with respect to combinations of CVSS score attributes, but not with respect to the attributed CVSS score when considered in isolation. We further found that the CVSS scores of exploited vulnerabilities were largely stable over time for the duration of the study, and that widely used security resources are not concordant with the data we observed. Raphaël Khoury, Kobra Khanmohammadi, Justin Vallé, Abdelwahab Hamou-Lhadj |
COMPSAC | 4 |
| 2025 | Execution Trace Reconstruction Using Diffusion-Based Generative ModelsabstractExecution tracing is essential for understanding system and software behaviour, yet lost trace events can significantly compromise data integrity and analysis. Existing solutions for trace reconstruction often fail to fully leverage available data, particularly in complex and high-dimensional contexts. Recent advancements in generative artificial intelligence, particularly diffusion models, have set new benchmarks in image, audio, and natural language generation. This study conducts the first comprehensive evaluation of diffusion models for reconstructing incomplete trace event sequences. Using nine distinct datasets generated from the Phoronix Test Suite, we rigorously test these models on sequences of varying lengths and missing data ratios. Our results indicate that the SSSDS4model, in particular, achieves superior performance, in terms of accuracy, perfect rate, and ROUGE-L score across diverse imputation scenarios. These findings underscore the potential of diffusion-based models to accurately reconstruct missing events, thereby maintaining data integrity and enhancing system monitoring and analysis. Madeline Janecek, Naser Ezzati-Jivan, Abdelwahab Hamou-Lhadj |
ICSE | 3 |
| 2025 | Developing a Taxonomy for Advanced Log Parsing TechniquesabstractLogs are widely used in various software engineering applications, including debugging, program comprehension, failure prediction, and anomaly detection. Despite their value, the unstructured nature of logs complicates the extraction of meaningful insights. In response, various log parsing techniques leveraging methods like machine learning and pattern recognition have been developed. Nevertheless, existing parsers frequently fail to achieve consistent accuracy, especially when handling complex log formats. To address this challenge, we conduct a comprehensive study to understand the characteristics of log events that lead to parsing errors. Using 16 different log datasets and 8 log parsers, we apply open coding techniques to derive a taxonomy of log event characteristics that contribute to parsing errors. We also examine how different log parsers are impacted by each category in the taxonomy. The resulting taxonomy not only provides insights into the complexity of parsing log data but can also guide the development of advanced parsing tools capable of handling the unique characteristics of diverse log formats. Issam Sedki, Abdelwahab Hamou-Lhadj, Otmane Aït Mohamed, Naser Ezzati-Jivan |
ICPC | 2 |
| 2024 | The Effectiveness of Compact Fine-Tuned LLMs in Log ParsingabstractLog parsing is defined as the process of extracting structured information from unstructured log data. It is an important step prior to many log analytics tasks. The emergence of Large Language Models (LLMs), like Generative Pre-trained Transformers (GPTs), has driven the development of novel log parsing methods. Existing studies have examined the effectiveness of large-scale general-purpose LLMs in log parsing. In this paper, we argue that the long-term adoption of such LLMs pose challenges of data privacy, cost, and tool integration. To address these challenges, we explore the viability of supervised fine-tuning of an open-source compact LLM for log parsing as a prospective alternative. To this end, we fine-tune the Mistral-7B-Instruct LLM on a diverse set of log files and evaluate its performance, in terms of both accuracy and robustness, against OpenAI's GPT-4-Turbo using different configuration settings. We apply two evaluation approaches, namely metric-based and LLM-based. Our overall findings show that fine-tuning a compact LLM such as Mistral-7B provides similar and sometimes better results than using a large-scale LLM, in our case GPT-4-Turbo. These findings are important because they enable companies to use a smaller LLM that they can readily adapt to parsing their log data, and integrate into their log analytics tools, without the need to rely on third-party LLM providers. Maryam Mehrabi, Abdelwahab Hamou-Lhadj, Hossein Moosavi |
ICSME | 2 |
| 2024 | ServiceAnomaly: An anomaly detection approach in microservices using distributed traces and profiling metrics
Mahsa Panahandeh, Abdelwahab Hamou-Lhadj, Mohammad Hamdaqa, James Miller 0001 |
J. Syst. Softw. | 2 |
| 2024 | AML: An accuracy metric model for effective evaluation of log parsing techniquesabstractLogs are essential for the maintenance of large software systems. Software engineers often analyze logs for debugging, root cause analysis , and anomaly detection tasks. Logs, however, are partly structured, making the extraction of useful information from massive log files a challenging task. Recently, many log parsing techniques have been proposed to automatically extract log templates from unstructured log files. These parsers, however, are evaluated using different accuracy metrics. In this paper, we show that these metrics have several drawbacks, making it challenging to understand the strengths and limitations of existing parsers. To address this, we propose a novel accuracy metric, called AML (Accuracy Metric for Log Parsing). AML is a robust accuracy metric that is inspired by research in the field of remote sensing . It is based on measuring omission and commission errors. We use AML to assess the accuracy of 14 log parsing tools applied to the parsing of 16 log datasets. We also show how AML compares to existing accuracy metrics. Our findings demonstrate that AML is a promising accuracy metric for log parsing compared to alternative solutions, which enables a comprehensive evaluation of log parsing tools to help better decision-making in selecting and improving log parsing techniques. Issam Sedki, Abdelwahab Hamou-Lhadj, Otmane Aït Mohamed |
J. Syst. Softw. | 2 |
| 2024 | Commit-time defect prediction using one-class classification
Mohammed A. Shehab, Wael Khreich, Abdelwahab Hamou-Lhadj, Issam Sedki |
J. Syst. Softw. | 3 |
| 2023 | Towards a Classification of Log Parsing ErrorsabstractLog parsing is used to extract structures from unstructured log data. It is a key enabler for many software engineering tasks including debugging, fault diagnosis, and anomaly detection. In recent years, we have seen an increase in the number of log parsing techniques and tools. The accuracy of these tools varies significantly. To improve log parsing tools, we need to understand the type of parsing errors they make, which is the purpose of this early research track paper. We achieve this by examining errors of four leading log parsing tools when applied to the parsing of four log datasets generated from various systems. Based on this analysis, we suggest a preliminary classification of log parsing errors, which contains nine categories of errors. We believe that this classification is a good starting point for improving the accuracy of log parsing tools, and also defining better logging practices. Issam Sedki, Abdelwahab Hamou-Lhadj, Otmane Aït Mohamed, Naser Ezzati-Jivan |
ICPC | 2 |
| 2023 | JITBoost: Boosting Just-In-Time Defect Prediction using Boolean Combination of ClassifiersabstractJust-In-Time Software Defects Prediction (JIT-SDP) plays a critical role in software engineering by enabling the early identification of potential defects before they impact system performance. This study investigates the effectiveness of Boolean Combination of Classifiers (BCC) in building effective JIT-SDP models. We propose the JITBoost framework, which leverages three BCC algorithms, namely Brute-force Boolean Combination (BBC), Iterative Boolean Combination (IBC), and Weighted Pruning Iterative Boolean Combination (WPIBC). JITBoost combines the decisions of six traditional machine learning algorithms and one deep learning algorithm. When applied to 259K commits of 34 projects, we show that JITBoost models perform better than traditional machine learning and deep learning algorithms when used individually. Specifically, JITBoost-BBC, JITBoost-IBC, and JITBoost-WPIBC achieve mean AUCs of 0.891, 0.879, and 0.886, respectively, with cross-validation. With a time-aware data-splitting approach, they achieve mean AUCs of 0.863, 0.854, and 0.857, respectively. Overall, the findings suggest that combining machine learning models within the JITBoost framework can lead to improved performance in JIT-SDP models. Mohammed A. Shehab, Abdelwahab Hamou-Lhadj, Venkata Sai Gunda |
QRS | 2 |
| 2023 | Open Science in Software Engineering: A Study on Deep Learning-Based Vulnerability DetectionabstractOpen science is a practice that makes scientific research publicly accessible to anyone, hence is highly beneficial. Given the benefits, the software engineering (SE) community has been diligently advocating open science policies during peer reviews and publication processes. However, to this date, there has been few studies that look into the status and issues of open science in SE from a systematic perspective. In this paper, we set out to start filling this gap. Given the great breadth of SE in general, we constrained our scope to a particular topic area in SE as an example case. Recently, an increasing number of deep learning (DL) approaches have been explored in SE, includingDL-based software vulnerability detection, a popular, fast-growing topic that addresses an important problem in software security. We exhaustively searched the literature in this area and identified 55 relevant works that propose a DL-based vulnerability detection approach. This was then followed by comprehensively investigating the four integral aspects of open science:availability,executability,reproducibility, andreplicability. Among other findings, our study revealed that only a small percentage (25.5%) of the studied approaches provided publiclyavailabletools. Some of these available tools did not provide sufficient documentation and complete implementation, making them notexecutableor notreproducible. The uses of balanced or artificially generated datasets caused significantly overrated performance of the respective techniques, making most of them notreplicable. Based on our empirical results, we made actionable suggestions on improving the state of open science in each of the four aspects. We note that our results and recommendations on most of these aspects (availability,executability,reproducibility) are not tied to the nature of the chosen topic (DL-based vulnerability detection) hence are likely applicable to other SE topic areas. We also believe our results and recommendations onreplicabilityto be applicable to other DL-based topics in SE as they are not tied to (the particular application of DL in) detecting software vulnerabilities. Yu Nong, Rainy Sharma, Abdelwahab Hamou-Lhadj, Xiapu Luo, Haipeng Cai |
IEEE Trans. Software Eng. | 3 |
| 2022 | An Effective Approach for Parsing Large Log FilesabstractBecause of their contribution to the overall reliability assurance process, software logs have become important data assets for the analysis of software systems. Logs are often the only data points that can shed light on how a software system behaves once deployed. Unfortunately, logs are often unstructured data items, hindering viable analysis of their content. There are studies that aim to automatically parse large log files. The primary goal is to create templates from raw log data samples that can later be used to recognize future logs. In this paper, we propose ULP, a Unified Log Parsing tool, which is highly accurate and efficient. ULP combines string matching and local frequency analysis to parse large log files in an efficient manner. First, log events are organized into groups using a text processing method. Frequency analysis is then applied locally to instances of the same group to identify static and dynamic content of log events. When applied to 10 log datasets of the LogPai benchmark, ULP achieves an average accuracy of 89.2%, which outperforms the accuracy of four leading log parsing tools, namely Drain, Logram, SPELL and AEL. Additionally, ULP can parse up to four million log events in less than 3 minutes. ULP is available online as an open source and can be readily used by practitioners and researchers to parse effectively and efficiently large log files so as to support log analysis tasks. Issam Sedki, Abdelwahab Hamou-Lhadj, Otmane Aït Mohamed, Mohammed A. Shehab |
ICSME | 2 |
| 2022 | Performance anomaly detection through sequence alignment of system-level tracesabstractIdentifying and diagnosing performance anomalies is essential for maintaining software quality, yet it can be a complex and time-consuming task. Low level kernel events have been used as an excellent data source to monitor performance, but raw trace data is often too large to easily conduct effective analyses. To address this shortcoming, in this paper, we propose a framework for uncovering performance problems using execution critical path data. A critical path is the longest execution sequence without wait delays, and it can provide valuable insight into a program's internal and external dependencies. Upon extracting this data, course grained anomaly detection techniques are employed to determine if a finer grained analysis is required. If this is the case, the critical paths of individual executions are grouped together with machine learning clustering to identify different execution types, and outlying anomalies are identified using performance indicators. Finally, multiple sequence alignment is used to pinpoint specific abnormalities in the identified anomalous executions, allowing for improved application performance diagnosis and overall program comprehension. Madeline Janecek, Naser Ezzati-Jivan, Abdelwahab Hamou-Lhadj |
ICPC | 3 |
| 2022 | ClusterCommit: A Just-in-Time Defect Prediction Approach Using Clusters of ProjectsabstractExisting Just-in-Time (JIT) bug prediction techniques are designed to work on single projects. In this paper, we present ClusterCommit, a JIT bug prediction approach geared towards clusters of projects that share common libraries and functionalities. Unlike existing techniques, ClusterCommit trains a machine learning model by combining commits from a set of projects that are part of a larger cluster. Once this model is built, ClusterCommit can be used to detect buggy commits in each of these projects. When applying ClusterCommits to 16 projects that revolve around the Hadoop ecosystem and 10 projects of the Hive ecosystem, the results show that ClusterCommit achieves an F1-score of 73% and MCC of 0.44 for both clusters. These preliminary results are very promising and may lead to new JIT bug prediction techniques geared towards projects that are part of a large cluster. Mohammed A. Shehab, Abdelwahab Hamou-Lhadj, Luay Alawneh |
SANER | 2 |
| 2022 | HealMA: a model-driven framework for automatic generation of IoT-based Android health monitoring applications
Maryam Mehrabi, Bahman Zamani, Abdelwahab Hamou-Lhadj |
Autom. Softw. Eng. | 3 |
| 2022 | The sense of logging in the Linux kernel
Keyur Patel, João Guilherme Faccin, Abdelwahab Hamou-Lhadj, Ingrid Nunes |
Empir. Softw. Eng. | 3 |
| 2022 | Guest Editors' introduction to the special section on the 12th system analysis and modelling conference (SAM 2020)
Abdelouahed Gherbi, Abdelwahab Hamou-Lhadj |
Inf. Softw. Technol. | 2 |
| 2022 | Locating and categorizing inefficient communication patterns in HPC systems using inter-process communication traces
Luay Alawneh, Abdelwahab Hamou-Lhadj |
J. Syst. Softw. | 2 |
| 2021 | EnHMM: On the Use of Ensemble HMMs and Stack Traces to Predict the Reassignment of Bug Report FieldsabstractBug reports (BR) contain vital information that can help triaging teams prioritize and assign bugs to developers who will provide the fixes. However, studies have shown that BR fields often contain incorrect information that need to be reassigned, which delays the bug fixing process. There exist approaches for predicting whether a BR field should be reassigned or not. These studies use mainly BR descriptions and traditional machine learning algorithms (SVM, KNN, etc.). As such, they do not fully benefit from the sequential order of information in BR data, such as function call sequences in BR stack traces, which may be valuable for improving the prediction accuracy. In this paper, we propose a novel approach, called EnHMM, for predicting the reassignment of BR fields using ensemble Hidden Markov Models (HMMs), trained on stack traces. EnHMM leverages the natural ability of HMMs to represent sequential data to model the temporal order of function calls in BR stack traces. When applied to Eclipse and Gnome BR repositories, EnHMM achieves an average precision, recall, and F-measure of 54%, 76%, and 60% on Eclipse dataset and 41%, 69%, and 51% on Gnome dataset. We also found that EnHMM improves over the best single HMM by 36% for Eclipse and 76% for Gnome. Finally, when comparing EnHMM to Im.ML.KNN, a recent approach in the field, we found that the average F-measure score of EnHMM improves the average F-measure of Im.ML.KNN by 6.80% and improves the average recall of Im.ML.KNN by 36.09%. However, the average precision of EnHMM is lower than that of Im.ML.KNN (53.93% as opposed to 56.71%). Abdelwahab Hamou-Lhadj, Korosh Koochekian Sabor, Mohammad Hamdaqa, Haipeng Cai |
SANER | 2 |
| 2021 | ALBA: a model-driven framework for the automatic generation of android location-based apps
Mohammadali Gharaat, Mohammadreza Sharbaf, Bahman Zamani, Abdelwahab Hamou-Lhadj |
Autom. Softw. Eng. | 4 |
| 2021 | The evolution of IoT Malwares, from 2008 to 2019: Survey, taxonomy, process simulator and perspectives
Benjamin Vignau, Raphaël Khoury, Sylvain Hallé, Abdelwahab Hamou-Lhadj |
J. Syst. Archit. | 4 |
| 2021 | MUPPIT: a method for using proper patterns in model transformations
Mahsa Panahandeh, Mohammad Hamdaqa, Bahman Zamani, Abdelwahab Hamou-Lhadj |
Softw. Syst. Model. | 4 |
| 2020 | DepGraph: Localizing Performance Bottlenecks in Multi-Core Applications Using Waiting Dependency Graphs and Software TracingabstractThis paper addresses the challenge of understanding the waiting dependencies between the threads and hardware resources required to complete a task. The objective is to improve software performance by detecting the underlying bottlenecks caused by system-level blocking dependencies. In this paper, we use a system level tracing approach to extract a Waiting Dependency Graph that shows the breakdown of a task execution among all the interleaving threads and resources. The method allows developers and system administrators to quickly discover how the total execution time is divided among its interacting threads and resources. Ultimately, the method helps detecting bottlenecks and highlighting their possible causes. Our experiments show the effectiveness of the proposed approach in several industry-level use cases. Three performance anomalies are analysed and explained using the proposed approach. Evaluating the method efficiency reveals that the imposed overhead never exceeds 10.1%, therefore making it suitable for in-production environments. Naser Ezzati-Jivan, Quentin Fournier, Michel R. Dagenais, Abdelwahab Hamou-Lhadj |
SCAM | 4 |
| 2020 | MobiLogLeak: A Preliminary Study on Data Leakage Caused by Poor Logging PracticesabstractLogging is an essential software practice that is used by developers to debug, diagnose and audit software systems. Despite the advantages of logging, poor logging practices can potentially leak sensitive data. The problem of data leakage is more severe in applications that run on mobile devices, since these devices carry sensitive identification information ranging from physical device identifiers (e.g., IMEI MAC address) to communications network identifiers (e.g., SIM, IP, Bluetooth ID), and application-specific identifiers related to the location and the users' accounts. This preliminary study explores the impact of logging practices on data leakage of such sensitive information. Particularly, we want to investigate whether log-related statements inserted into an application code could lead to data leakage. While studying logging practices in mobile applications is an active research area, to our knowledge, this is the first study that explores the interplay between logging and security in the context of mobile applications for Android. We propose an approach called MobiLogLeak, an approach that identifies log statements in deployed apps that leak sensitive data. MobiLogLeak relies on taint flow analysis. Among 5,000 Android apps that we studied, we found that 200 apps leak sensitive data through logging. Mohammad Hamdaqa, Haipeng Cai, Abdelwahab Hamou-Lhadj |
SANER | 4 |
| 2020 | A study of run-time behavioral evolution of benign versus malicious apps in android
Haipeng Cai, Xiaoqin Fu, Abdelwahab Hamou-Lhadj |
Inf. Softw. Technol. | 3 |
| 2020 | A systematic literature review on automated log abstraction techniques
Diana El-Masri, Fábio Petrillo, Yann-Gaël Guéhéneuc, Abdelwahab Hamou-Lhadj, Anas Bouziane |
Inf. Softw. Technol. | 4 |
| 2020 | Automatic prediction of the severity of bugs using stack traces and categorical features
Korosh Koochekian Sabor, Mohammad Hamdaqa, Abdelwahab Hamou-Lhadj |
Inf. Softw. Technol. | 3 |
| 2020 | Lossless compaction of model execution traces
Fazilat Hojaji, Bahman Zamani, Abdelwahab Hamou-Lhadj, Tanja Mayerhofer, Erwan Bousse |
Softw. Syst. Model. | 3 |
| 2019 | Empirical study of android repackaged applications
Kobra Khanmohammadi, Neda Ebrahimi Koopaei, Abdelwahab Hamou-Lhadj, Raphaël Khoury |
Empir. Softw. Eng. | 3 |
| 2019 | Exploiting Parts-of-Speech for effective automated requirements traceability
Nasir Ali, Haipeng Cai, Abdelwahab Hamou-Lhadj, Jameleddine Hassine |
Inf. Softw. Technol. | 3 |
| 2019 | An HMM-based approach for automatic detection and classification of duplicate bug reports
Neda Ebrahimi Koopaei, Abdelaziz Trabelsi, Abdelwahab Hamou-Lhadj, Kobra Khanmohammadi |
Inf. Softw. Technol. | 4 |
| 2019 | Model execution tracing: a systematic mapping study
Fazilat Hojaji, Tanja Mayerhofer, Bahman Zamani, Abdelwahab Hamou-Lhadj, Erwan Bousse |
Softw. Syst. Model. | 4 |
| 2018 | CLEVER: combining code metrics with clone detection for just-in-time fault prevention and resolution in large industrial projectsabstractAutomatic prevention and resolution of faults is an important research topic in the field of software maintenance and evolution. Existing approaches leverage code and process metrics to build metric-based models that can effectively prevent defect insertion in a software project. Metrics, however, may vary from one project to another, hindering the reuse of these models. Moreover, they tend to generate high false positive rates by classifying healthy commits as risky. Finally, they do not provide sufficient insights to developers on how to fix the detected risky commits. In this paper, we propose an approach, called CLEVER (Combining Levels of Bug Prevention and Resolution techniques), which relies on a two-phase process for intercepting risky commits before they reach the central repository. When applied to 12 Ubisoft systems, the results show that CLEVER can detect risky commits with 79% precision and 65% recall, which outperforms the performance of Commit-guru, a recent approach that was proposed in the literature. In addition, CLEVER is able to recommend qualitative fixes to developers on how to fix risky commits in 66.7% of the cases. Mathieu Nayrolles, Abdelwahab Hamou-Lhadj |
MSR | 2 |
| 2018 | MASKED: A MapReduce Solution for the Kappa-Pruned Ensemble-Based Anomaly Detection SystemabstractDetecting system anomalies at run-time is critical for system reliability and security. Studies in this area focused mainly on effectiveness of the proposed approaches; that is, the ability to detect anomalies with high accuracy. However, less attention was given to efficiency. In this paper, we propose an efficient MapReduce Solution for the Kappa-pruned Ensemble based Anomaly Detection System (MASKED). It profiles the heterogeneous features from large-scale traces of system calls and processes them by heterogeneous anomaly detectors which are Sequence-Time Delay Embedding (STIDE), Hidden Markov Model (HMM), and One-class Support Vector Machine (OCSVM). We deployed MASKED on a Hadoop cluster using the MapReduce programming model. We compared their efficiency and scalability by varying the size of the cluster. We assessed the performance of the proposed approach using the CANALI-WD dataset which consists of 180 GB of execution traces, collected from 10 different machines. Experimental results show that MASKED becomes more efficient and scalable as the file size is increased (e.g., 6-node cluster is 8 times faster than the 2-node cluster). Moreover, the throughput achieved on a 6-node solution is up to 5 times better than a 2-node solution. Korosh Koochekian Sabor, Abdelaziz Trabelsi, Abdelwahab Hamou-Lhadj, Luay Alawneh |
QRS | 4 |
| 2018 | A framework for the recovery and visualization of system availability scenarios from execution traces
Jameleddine Hassine, Abdelwahab Hamou-Lhadj, Luay Alawneh |
Inf. Softw. Technol. | 2 |
| 2018 | Combining heterogeneous anomaly detectors for improved software security
Wael Khreich, Syed Shariyar Murtaza, Abdelwahab Hamou-Lhadj, Chamseddine Talhi |
J. Syst. Softw. | 3 |
| 2018 | Anomaly Detection Techniques Based on Kappa-Pruned EnsemblesabstractEnsemble-based anomaly detection systems (ADSs), using Boolean combination, have been shown to reduce the false alarm rate over that of a single detector. However, the existing Boolean combination methods rely on an exponential number of combinations making them impractical, even for a small number of detectors. In this paper, we propose weighted pruning-based Boolean combination, an efficient approach for selecting and combining accurate and diverse anomaly detectors. It works in three phases. The first phase selects a subset of the available base diverse soft detectors by pruning all the redundant soft detectors based on a weighted version of Cohen's kappa measure of agreement. The second phase selects a subset of diverse and accurate crisp detectors from the base soft detectors (selected in Phase1) based on the unweighted kappa measure. The selected complementary crisp detectors are then combined in the final phase using Boolean combinations. The results on two large scale datasets show that the proposed weighted pruning approach is able to maintain and even improve the accuracy of existing Boolean combination techniques, while significantly reducing the combination time and the number of detectors selected for combination. Wael Khreich, Abdelwahab Hamou-Lhadj |
IEEE Trans. Reliab. | 3 |
| 2017 | HyDroid: A Hybrid Approach for Generating API Call Traces from Obfuscated Android Applications for Mobile SecurityabstractThe growing popularity of Android applications makes them vulnerable to security threats. There exist several studies that focus on the analysis of the behaviour of Android applications to detect the repackaged and malicious ones. These techniques use a variety of features to model the application's behaviour, among which the calls to Android API, made by the application components, are shown to be the most reliable. To generate the APIs that an application calls is not an easy task. This is because most malicious applications are obfuscated and do not come with the source code. This makes the problem of identifying the API methods invoked by an application an interesting research issue. In this paper, we present HyDroid, a hybrid approach that combines static and dynamic analysis to generate API call traces from the execution of an application's services. We focus on services because they contain key characteristics that allure attackers to misuse them. We show that HyDroid can be used to extract API call trace signatures of several malware families. Kobra Khanmohammadi, Abdelwahab Hamou-Lhadj |
QRS | 2 |
| 2017 | DURFEX: A Feature Extraction Technique for Efficient Detection of Duplicate Bug ReportsabstractThe detection of duplicate bug reports can help reduce the processing time of handling field crashes. This is especially important for software companies with a large client base where multiple customers can submit bug reports, caused by the same faults. There exist several techniques for the detection of duplicate bug reports; many of them rely on some sort of classification techniques applied to information extracted from stack traces. They classify each report using functions invoked in the stack trace associated with the bug report. The problem is that typical bug repositories may have stack traces that contain tens of thousands of functions, which causes the curse of dimensionality problem. In this paper, we propose a feature extraction technique that reduces the feature size and yet retains the information that is most critical for the classification. The proposed feature extraction approach starts by abstracting stack traces of function calls into sequences of package names, by replacing each function with the package in which it is defined. We then segment these traces into multiple N-grams of variable length and map them to fixed-size sparse feature vectors, which are used to measure the distance between the stack trace of incoming bug report with a historical set of bug reports stack traces. The linear combination of stack trace similarity and non-textual fields such as component and severity are then used to measure the distance of a bug report with a historical set of bug reports. We show the effectiveness of our approach by applying it to the Eclipse bug repository that contains tens of thousands of bug reports. Our approach outperforms the approach that uses distinct function names, while significantly reducing the processing time. Korosh Koochekian Sabor, Abdelwahab Hamou-Lhadj, Alf Larsson |
QRS | 2 |
| 2017 | An anomaly detection system based on variable N-gram features and one-class SVM
Wael Khreich, Babak Khosravifar, Abdelwahab Hamou-Lhadj, Chamseddine Talhi |
Inf. Softw. Technol. | 3 |
| 2017 | A bug reproduction approach based on directed model checking and crash tracesabstractAbstract Reproducing a bug that caused a system to crash is an important task for uncovering the causes of the crash and providing appropriate fixes. In this paper, we propose a novel crash reproduction approach that combines directed model checking and backward slicing to identify the program statements needed to reproduce a crash. Our approach, named JCHARMING (Java CrasH Automatic Reproduction by directed Model checkING), uses information found in crash traces combined with static program slices to guide a model checking engine in an optimal way. We show that JCHARMING is efficient in reproducing bugs from 10 different open source systems. Overall, JCHARMING is able to reproduce 80% of the bugs used in this study in an average time of 19 min. Copyright © 2016 John Wiley & Sons, Ltd. Mathieu Nayrolles, Abdelwahab Hamou-Lhadj, Sofiène Tahar, Alf Larsson |
J. Softw. Evol. Process. | 2 |
| 2016 | Model driven performance simulation of cloud provisioned Hadoop mapreduce applicationsabstractHadoop is a widely adopted open source implementation of MapReduce. A Hadoop cluster can be fully provisioned by a Cloud service provider to provide elasticity in computational resource allocation. Understanding the performance characteristics of a Hadoop job can help achieve an optimal balance between resource usage (cost) and job latency on a cloud-based cluster. This paper presents a method that estimates the performance of a MapReduce application in a Cloud provisioned Hadoop cluster. We develop a model-driven approach that models a cloud provided independent Hadoop MapReduce model and customizes it for a specific Cloud deployment. These models are further transformed into a simulation model that produces estimations of end-to-end job latency. We explore this method in the design space of MapReduce applications to estimate the performance for different sizes of input data. Our approach provides a model-to-simulation-to-prediction method for observing the performance behaviour of MapReduce applications given a configuration of a MapReduce platform. Hanieh Alipour, Yan Liu 0001, Abdelwahab Hamou-Lhadj, Ian Gorton |
MiSE@ICSE | 3 |
| 2016 | Key Elements Extraction and Traces Comprehension Using Gestalt Theory and the Helmholtz PrincipleabstractTrace analysis techniques are used by software engineers to understand the behaviour of large systems. This understanding can facilitate various software maintenance activities including debugging and feature enhancement. However, traces usually tend to be very large, which makes it difficult for software engineers to unveil the key logic and functionalities embedded in a program's execution. Hence, it is necessary to develop methods and tools that can efficiently identify the important information contained in a large trace. In this paper, we propose an approach that builds on the concept of trace segmentation to extract the major components of a traced scenario. Our approach is based on Gestalt theory and the Helmholtz principle. We show the effectiveness of our approach by applying it to a dataset of large traces. Raphaël Khoury, Abdelwahab Hamou-Lhadj |
ICSME | 3 |
| 2016 | A Controlled Experiment for Evaluating the Comprehensibility of UML Action LanguagesabstractAction Languages represent an emerging paradigm where modeling abstractions are embedded in code to bridge the gap with visual models, such as UML models. The paradigm is gaining momentum, evident by the growing number of tools and standards that support this paradigm. In this paper, we report on a controlled experiment to assess the comprehensibility of those languages and compare it to that of object-oriented (OO) programming languages. We further report on the impact of also having access to the UML notation on the comprehensibility of those languages. Results suggest that action languages are significantly more comprehensible than traditional OO languages. Furthermore, there was not a significant improvement in comprehensibility when the UML notation was used along with both OO and action language code. We conclude that action languages are a promising alternative to traditional OO languages for specifying details, yet seem to be as comprehensible as high-level visual models. Omar Bahy Badreddin, Maged Elaasar, Abdelwahab Hamou-Lhadj |
MODELSWARD | 3 |
| 2016 | BUMPER: A Tool for Coping with Natural Language Searches of Millions of Bugs and FixesabstractIn recent years, mining bug report (BR) repositories has perhaps been one of the most active software engineering research fields. There exist many open source bug tracking and version control systems that developers and researchers can use to examine bug reports so as to reason about software quality. The issue is that these repositories use different interfaces and ways to access and represent data, which hinders productivity and reuse. To address this, we introduce BUMPER (BUg Metarepository for dEvelopers and Researchers), a common infrastructure for developers and researchers interested in mining data from many (heterogeneous) repositories. BUMPER is an open source web-based environment that extracts information from a variety of BR repositories and version control systems. It is equipped with a powerful search engine to help users rapidly query the repositories using a single point of access. To demonstrate the effectiveness of BUMPER, we use it to build a large dataset from a variety of repositories. The dataset contains more than one million bug reports and fixes. Both BUMPER and the dataset are publicly available at https://bumper-app.com. Mathieu Nayrolles, Abdelwahab Hamou-Lhadj |
SANER | 2 |
| 2016 | Segmenting large traces of inter-process communication with a focus on high performance computing systems
Luay Alawneh, Abdelwahab Hamou-Lhadj, Jameleddine Hassine |
J. Syst. Softw. | 2 |
| 2016 | Mining trends and patterns of software vulnerabilities
Syed Shariyar Murtaza, Wael Khreich, Abdelwahab Hamou-Lhadj, Ayse Basar Bener |
J. Syst. Softw. | 3 |
| 2015 | A trace abstraction approach for host-based anomaly detectionabstractHigh false alarm rates and execution times are among the key issues in host-based anomaly detection systems. In this paper, we investigate the use of trace abstraction techniques for reducing the execution time of anomaly detectors while keeping the same accuracy. The key idea is to represent system call traces as traces of kernel module interactions and use the resulting abstract traces as input to known anomaly detection techniques, such as STIDE (the Sequence Time-Delay Embedding) and HMM (Hidden Markov Models). We performed experiments on three datasets, namely, the traditional UNM dataset as well as two modern datasets, Firefox and ADFA-LD. The results show that kernel module traces can lead to similar or fewer false alarms and considerably smaller execution times compared to raw system call traces for host-based anomaly detection systems. Syed Shariyar Murtaza, Wael Khreich, Abdelwahab Hamou-Lhadj, Stéphane Gagnon |
CISDA | 3 |
| 2015 | An empirical study on the handling of crash reports in a large software company: An experience reportabstractIn this paper, we report on an empirical study we have conducted at Ericsson to understand the handling of crash reports (CRs). The study was performed on a dataset of CRs spanning over two years of activities on one of Ericsson's largest systems (+4 Million LOC). CRs at Ericsson are divided into two types: Internal and External. Internal CRs are reported within the organization after the integration and system testing phase. External CRs are submitted by customers and caused mainly by field failures. We examine the proportion and severity of internal CRs and that of external CRs. A large number of external (and severe) CRs could indicate flaws in the testing phase. Failing to react quickly to external CRs, on the other hand, may expose Ericsson to fines and penalties due to the Working Level Agreements (WLA) that Ericsson has with its customers. Moreover, we contrast the time it takes to handle each type of CRs with the dual aim to understand the similarities and differences as well as the factors that impact the handling of each type of CRs. Our results show that (a) it takes more time to fix external CRs compared to internal CRs, (b) the severity attribute is used inconsistently through organizational units, (c) assignment time of internal CRs is less than that of external CRs, (d) More than 50% of CRs are not answered within the organization's fixing time requirements defined in WLA. Abdou Maiga, Abdelwahab Hamou-Lhadj, Mathieu Nayrolles, Korosh Koochekian Sabor, Alf Larsson |
ICSME | 2 |
| 2015 | An Anomaly Detection System Based on Ensemble of Detectors with Effective Pruning TechniquesabstractAnomaly detection systems rely on machine learning techniques to model the normal behavior of the system. This model is used during operation to detect anomalies due to attacks or design faults. Ensemble methods have been used to improve the overall detection accuracy by combining the outputs of several accurate and diverse models. Existing Boolean combination techniques either require an exponential number of combinations or sequential combinations that grow linearly with the number of iterations, which make them difficult to scale up and analyze. In this paper, we propose PBC (Pruning Boolean Combination), an efficient approach for selecting and combining anomaly detectors. PBC relies on two novel pruning techniques that we have developed to aggressively prune redundant and trivial detectors. Compared to existing work, PBC reduces significantly the number of detectors to combine, while keeping similar accuracy. We show the effectiveness of PBC when applying it to a large dataset. Amirreza Soudi, Wael Khreich, Abdelwahab Hamou-Lhadj |
QRS | 3 |
| 2015 | Towards a common metamodel for traces of high performance computing systems to enable software analysis tasksabstractThere exist several tools for analyzing traces generated from HPC (High Performance Computing) applications, used by software engineers for debugging and other maintenance tasks. These tools, however, use different formats to represent HPC traces, which hinders interoperability and data exchange. At the present time, there is no standard metamodel that represents HPC trace concepts and their relations. In this paper, we argue that the lack of a common metamodel is a serious impediment for effective analysis for this class of software systems. We aim to fill this void by presenting MTF2 (MPI Trace Format2)-a metamodel for representing HPC system traces. MTF2 is built with expressiveness and scalability in mind. Scalability, an important requirement when working with large traces, is achieved by adopting graph theory concepts to compact large traces. We show through a case study that a trace represented in MTF2 can be in average 49% smaller than a trace represented in a format that does not consider compaction. Luay Alawneh, Abdelwahab Hamou-Lhadj, Jameleddine Hassine |
SANER | 2 |
| 2015 | JCHARMING: A bug reproduction approach using crash traces and directed model checkingabstractDue to their inherent complexity, software systems are pledged to be released with bugs. These bugs manifest themselves on client's computers, causing crashes and undesired behaviors. Field crashes, in particular, are challenging to understand and fix as the information provided by the impacted customers are often scarce and inaccurate. To address this issue, there is a need to find ways for automatically reproducing the crash in a lab environment in order to fully understand its root causes. Crash reproduction is also an important step towards developing adequate patches. In this paper, we propose a novel crash reproduction approach, called JCHARMING (Java CrasH Automatic Reproduction by directed Model checkING). JCHARMING uses crash traces and model checking to identify program statements needed to reproduce a crash. Our approach takes advantage of the completeness provided by model checking while ignoring unneeded system states by means of information found in crash traces combined with static slices. We show the effectiveness of JCHARMING by applying it to seven different open source programs cumulating more than one million lines of code scattered in around 7000 classes. Overall, JCHARMING was able to reproduce 85% of the submitted bugs. Mathieu Nayrolles, Abdelwahab Hamou-Lhadj, Sofiène Tahar, Alf Larsson |
SANER | 2 |
| 2015 | Identifying Recurring Faulty Functions in Field Traces of a Large Industrial Software SystemabstractSoftware maintainers use the traces of field failures to understand and diagnose faulty functions that cause the system to fail. Despite their usefulness, traces from the field can be quite overwhelming, especially for software systems with a vast client base. In the execution of realistic applications, many of them being millions of lines of code, there are just too many traces that are generated. In addition, traces are known to be extraordinarily large, which further complicates matters. Fortunately, not all field failures are caused by new faults. In fact, previous studies showed that 50% to 90% of field failures are due to previously known faults. In this paper, we propose a machine learning approach that automatically detects recurring faulty functions in the traces of new field failures. We achieve our goal by training decision trees on earlier resolved traces of system failures from the current and prior releases of the system. When applied to a large industrial system with 20 million lines of code and 200,000 functions, our approach was able to detect recurring faulty functions in the traces of field failures with an accuracy of 90%, to even 97% in some cases. Syed Shariyar Murtaza, Nazim H. Madhavji, Mechelle Gittens, Abdelwahab Hamou-Lhadj |
IEEE Trans. Reliab. | 4 |
| 2014 | Total ADS: Automated Software Anomaly Detection SystemabstractWhen a software system starts behaving abnormally during normal operations, system administrators resort to the use of logs, execution traces, and system scanners (e.g., anti-malwares, intrusion detectors, etc.) to diagnose the cause of the anomaly. However, the unpredictable context in which the system runs and daily emergence of new software threats makes it extremely challenging to diagnose anomalies using current tools. Host-based anomaly detection techniques can facilitate the diagnosis of unknown anomalies but there is no common platform with the implementation of such techniques. In this paper, we propose an automated anomaly detection framework (Total ADS) that automatically trains different anomaly detection techniques on a normal trace stream from a software system, raise anomalous alarms on suspicious behaviour in streams of trace data, and uses visualization to facilitate the analysis of the cause of the anomalies. Total ADS is an extensible Eclipse-based open source framework that employs a common trace format to use different types of traces, a common interface to adapt to a variety of anomaly detection techniques (e.g., HMM, sequence matching, etc.). Our case study on a modern Linux server shows that Total ADS automatically detects attacks on the server, shows anomalous paths in traces, and provides forensic insights. Syed Shariyar Murtaza, Abdelwahab Hamou-Lhadj, Wael Khreich, Mario Couture |
SCAM | 2 |
| 2014 | Taxonomy of intrusion risk assessment and response system
Alireza Shameli-Sendi, Mohamed Cheriet, Abdelwahab Hamou-Lhadj |
Comput. Secur. | 3 |
| 2014 | An empirical study on the use of mutant traces for diagnosis of faults in deployed systems
Syed Shariyar Murtaza, Abdelwahab Hamou-Lhadj, Nazim H. Madhavji, Mechelle Gittens |
J. Syst. Softw. | 2 |
| 2013 | Mining Telecom System Logs to Facilitate Debugging TasksabstractTelecommunication systems are monitored continuously to ensure quality and continuity of service. When an error or an abnormal behaviour occurs, software engineers resort to the analysis of the generated logs for troubleshooting. The problem is that, even for a small system, the log data generated after running the system for a period of time can be considerably large. There is a need to automatically mine important information from this data. There exist studies that aim to do just that, but their focus has been mainly on software applications, paying little attention to network information used by telecom systems. In this paper, we show how data mining techniques, more particularly the ones based on mining frequent itemsets, can be used to extract patterns that characterize the main behaviour of the traced scenarios. We show the effectiveness of our approach through a representative study conducted in an industrial setting. Alf Larsson, Abdelwahab Hamou-Lhadj |
ICSM | 2 |
| 2013 | A host-based anomaly detection approach by representing system calls as states of kernel modulesabstractDespite over two decades of research, high false alarm rates, large trace sizes and high processing times remain among the key issues in host-based anomaly intrusion detection systems. In an attempt to reduce the false alarm rate and processing time while increasing the detection rate, this paper presents a novel anomaly detection technique based on semantic interactions of system calls. The key concept is to represent system calls as states of kernel modules, analyze the state interactions, and identify anomalies by comparing the probabilities of occurrences of states in normal and anomalous traces. In addition, the proposed technique allows a visual understanding of system behaviour, and hence a more informed decision making. We evaluated this technique on Linux based programs of UNM datasets and a new modern Firefox dataset. We created the Firefox dataset on Linux using contemporary test suites and hacking techniques. The results show that our technique yields fewer false alarms and can handle large traces with smaller (or comparable) processing times compared against the existing techniques for the host based anomaly intrusion detection systems. Syed Shariyar Murtaza, Wael Khreich, Abdelwahab Hamou-Lhadj, Mario Couture |
ISSRE | 3 |
| 2013 | Automatic configuration generation for service high availability with load balancingabstractSUMMARY The need for highly available services is ever increasing in various domains ranging from mission‐critical systems to transaction‐based ones such as banking. The Service Availability Forum has defined a set of services and related API specifications to address the growing need of commercial off‐the‐shelf high availability solutions. Among these services, the availability management framework (AMF) is the service responsible for managing the high availability of the application services by coordinating redundant application components deployed on the AMF cluster. To achieve this task, an AMF implementation requires a specific logical view of the organization of the application's services and components, known as an AMF configuration. Developing manually such a configuration is a complex error‐prone task that requires extensive domain knowledge. In this paper, we present an approach for the automatic generation of AMF configurations and alleviate the task of configuration designers. One important aspect of the AMF configuration is ranking the service units, when it is required by the redundancy model, for the assignment of the workload by AMF at runtime. Our approach includes a technique for generating these rankings in such a way that guarantees load balancing even after the occurrence of a failure. Copyright © 2012 John Wiley & Sons, Ltd. Ali Kanso, Ferhat Khendek, Maria Toeroe, Abdelwahab Hamou-Lhadj |
Concurr. Comput. Pract. Exp. | 4 |
| 2013 | Stratified sampling of execution traces: Execution phases serving as strata
Heidar Pirzadeh, Sara Shanian, Abdelwahab Hamou-Lhadj, Luay Alawneh, Arya Shafiee |
Sci. Comput. Program. | 3 |
| 2012 | Towards a formal framework for evaluating the effectiveness of system diversity when applied to securityabstractN-version programming has been shown to be an effective way to increase the reliability of systems. In this study, we examine the possibility of extending this approach to address security, rather than reliability concerns. We focus specifically on how to evaluate the efficiency of the use of diversity for security. We show that while several key elements must be taken into account when N-version programming is used for security rather than reliability, it is nonetheless possible to devise a reasoning framework to evaluate the efficiency of this development paradigm in a security context. This framework allows us to reason about the most effective way to use diversity for security. Raphaël Khoury, Abdelwahab Hamou-Lhadj, Mario Couture |
CISDA | 2 |
| 2012 | An improved Hidden Markov Model for anomaly detection using frequent common patternsabstractHost-based intrusion detection techniques are needed to ensure the safety and security of software systems, especially, if these systems handle sensitive data. Most host-based intrusion detection systems involve building some sort of reference models offline, usually from execution traces (in the absence of the source code), to characterize the system healthy behavior. The models can later be used as a baseline for online detection of abnormal behavior. Perhaps the most popular techniques are the ones based on the use of Hidden Markov Models (HMM). These techniques, however, require long training time of the models, which makes them computationally infeasible, the main reason being the large size of typical traces. In this paper, we propose an improved HMM using the concept of frequent common patterns. In other words, we build models based on extracting the largest n-grams (patterns) in the traces instead of taking each trace event on its own. We show through a case study that our approach can reduce the training time by 31.96%-48.44% compared to the original HMM algorithms while keeping almost the same accuracy rate. Afroza Sultana, Abdelwahab Hamou-Lhadj, Mario Couture |
ICC | 2 |
| 2012 | Identifying computational phases from inter-process communication traces of HPC applicationsabstractUnderstanding the behaviour of High Performance Computing (HPC) systems is a challenging task due to the large number of processes they involve as well as the complex interactions among these processes. In this paper, we present a novel approach that aims to simplify the analysis of large execution traces generated from HPC applications. We achieve this through a technique that allows semiautomatic extraction of execution phases from large traces. These phases, which characterize the main computations of the traced scenario, can be used by software engineers to browse the content of a trace at different levels of abstraction. Our approach is based on the application of information theory principles to the analysis of sequences of communication patterns found in HPC traces. The results of the proposed approach when applied to traces of a large HPC industrial system demonstrate its effectiveness in identifying the main program phases and their corresponding sub-phases. Luay Alawneh, Abdelwahab Hamou-Lhadj |
ICPC | 2 |
| 2012 | A metamodel for the compact but lossless exchange of execution traces
Abdelwahab Hamou-Lhadj, Timothy Lethbridge |
Softw. Syst. Model. | 1 |
| 2011 | AMF configurations: Checking for service protection using heuristics
Pejman Salehi, Ferhat Khendek, Abdelwahab Hamou-Lhadj, Maria Toeroe |
CNSM | 3 |
| 2011 | A Novel Approach Based on Gestalt Psychology for Abstracting the Content of Large Execution Traces for Program ComprehensionabstractThe analysis of execution traces can reveal important information about the behavioral aspects of complex software systems, hence reducing the time and effort it takes to understand and maintain them. Traces, however, tend to be considerably large which hinders their effective analysis. Existing traces analysis tools rely on some sort of visualization techniques to help software engineers make sense of trace content. Many of these techniques have been studied and found to be limited in many ways. In this paper, we present a novel trace analysis technique that automatically divides the content of a large trace into meaningful segments that correspond to the program's main execution phases such as initializing variables, performing a specific computation, etc. These phases can simplify significantly the exploration of large traces by allowing software engineers to first understand the content of a trace at a high-level before they decide to dig into the details. Our phase detection method is inspired by Gestalt laws that characterize the proximity, similarity, and continuity of the elements of a data space. We model these concepts in the context of execution traces and show how they can be used as gravitational forces that yield the formation of dense groups of trace elements, which indicate candidate phases. We applied our approach to two software systems. The results are very promising. Heidar Pirzadeh, Abdelwahab Hamou-Lhadj |
ICECCS | 2 |
| 2011 | A software behaviour analysis framework based on the human perception systemsabstractUnderstanding software behaviour can help in a variety of software engineering tasks if one can develop effective techniques for analyzing the information generated from a system's run. These techniques often rely on tracing. Traces, however, can be considerably large and complex to process. Heidar Pirzadeh, Abdelwahab Hamou-Lhadj |
ICSE | 2 |
| 2011 | Exploiting text mining techniques in the analysis of execution tracesabstractThe analysis of execution traces can be useful in many software engineering activities including debugging, feature enhancement, performance analysis, and any other task that requires some degree of understanding of the way a system behaves. Traces, however, tend to be considerably large, which often hinders effective analysis of their content. There is a need to investigate ways to help software engineers find and understand important information conveyed in a trace despite the trace being massive. Motivated by the work done in the area of text mining, we propose, in this paper, a trace exploration approach based on examining the trace execution phases. The approach consists of automatically identifying relevant information about the phases as well as the ability to provide an efficient representation of the flow of phases by detecting redundant phases using a cosine similarity metric. We applied our approach to large traces generated from two different systems and were able to quickly understand their content and extract higher level views that characterize the essence of the information conveyed in these traces. Heidar Pirzadeh, Abdelwahab Hamou-Lhadj, Mohak Shah |
ICSM | 2 |
| 2011 | MTF: A Scalable Exchange Format for Traces of High Performance Computing SystemsabstractExecution traces generated from running high performance computing applications (HPC) may reach tens or hundreds of gigabytes. The trace data can be used for visualization, analysis of profiling information about the target system. However, in order to make the utilization of this data efficient, the trace needs to be represented in a structure that facilitates the access to its data. One important factor that should be considered when representing trace data is scalability; the trace met model should be able to represent the trace in a compact form that enables scalability of the analysis tools. Additionally, a trace file needs to be available in a format that is well-known in the software engineering area by making it open. In this paper, we propose a metamodel for representing dynamic information generated from HPC that use the MPI standard as the inter-process communication model. MPI Trace Format (MTF) is meant to meet the aforementioned requirements and is intended to facilitate the interoperability among different trace analysis tools. Luay Alawneh, Abdelwahab Hamou-Lhadj |
ICPC | 2 |
| 2011 | The Concept of Stratified Sampling of Execution TracesabstractExecution traces can be overwhelmingly large. To reduce their size, sampling techniques, especially the ones based on random sampling, have been extensively used. Random sampling, however, may result in samples that are not representative of the original trace. We propose a trace sampling framework based on stratified sampling that not only reduces the size of a trace but also results in a sample that is representative of the original trace by ensuring that the desired characteristics of an execution are distributed similarly in both the sampled and the original trace. Heidar Pirzadeh, Sara Shanian, Abdelwahab Hamou-Lhadj, Ali Mehrabian |
ICPC | 3 |
| 2011 | An exchange format for representing dynamic information generated from High Performance Computing applications
Luay Alawneh, Abdelwahab Hamou-Lhadj |
Future Gener. Comput. Syst. | 2 |
| 2011 | An approach based on citation analysis to support effective handling of regulatory compliance
Mohammad Hamdaqa, Abdelwahab Hamou-Lhadj |
Future Gener. Comput. Syst. | 2 |
| 2010 | An Approach for Detecting Execution Phases of a System for the Purpose of Program ComprehensionabstractUnderstanding the behavioural aspects of a software system is an important activity in many software engineering activities including program comprehension and reverse engineering. The behaviour of software is typically represented in the form of execution traces. Traces, however, tend to be considerably large which makes analyzing their content a complex task. There is a need for trace simplification techniques that can help software engineers make sense of the content of a trace despite the trace being massive. In this paper, we present a novel algorithm that aims to simplify the analysis of a large trace by detecting the execution phases that compose it. An example of a phase could be an initialization phase, a specific computation, etc. Our algorithm processes a trace generated from running the program under study and divides it into phases that can be later used by software engineers to understand where and why a particular computation appears. We also show the effectiveness of our approach through a case study. Heidar Pirzadeh, Akanksha Agarwal, Abdelwahab Hamou-Lhadj |
SERA | 3 |
| 2010 | Extending the UML Metamodel to Provide Support for Crosscutting ConcernsabstractAspect-orientation is a term used to describe approaches that explicitly capture, model and implement crosscutting concerns (or aspects). There is currently a number of new programming languages as well as extensions to current programming languages, the design dimensions of most of which have been influenced by the AspectJ language through three concepts and their respective constructs, namely join points, point cuts and advice which can support two principles recognized as being key concepts of aspect-oriented programming (AOP): quantification and obliviousness. At the modeling level, the reception of AOP has long been focused on the modeling of AspectJ programs, and there exists no model that is generic enough to capture non-AspectJ aspects either as a source language during forward engineering or as a target language during reverse engineering. In this paper, we present an extension to the UML metamodel to explicitly capture crosscutting concerns. The model is independent from any programming language and abstracted away from platform specific details. An instantiation of the newly created metamodel can be represented in standard XMI format, which enables current CASE tools to read and to visualize the instance models in UML. This language-independent aspectual description can support model transformations vital to software development and maintenance, such as forward engineering, reverse engineering, and reengineering. Zohreh Sharafi, Parisa Mirshams, Abdelwahab Hamou-Lhadj, Constantinos Constantinides |
SERA | 3 |
| 2009 | Generating AMF Configurations from Software Vendor Constraints and User RequirementsabstractThe service availability forum (SAF) has defined a set of service API specifications addressing the growing need of commercial-off-the-shelf high availability solutions. Among these services, the availability management framework (AMF) is the service responsible for managing the high availability of the application services by coordinating redundant application components. To achieve this task, an AMF implementation requires a specific logical view of the organization of the application's services and components known as an AMF configuration. Developing manually such a configuration is a complex, error prone, and time consuming task. In this paper, we present an approach for automatic generation of AMF configurations from a set of requirements given by the configuration designer and the description of the software as provided by the vendor. Our approach alleviates the need of configuration designers dealing with a large number of AMF entities and their relations. Ali Kanso, Maria Toeroe, Abdelwahab Hamou-Lhadj, Ferhat Khendek |
ARES | 3 |
| 2009 | A Tool Suite for the Generation and Validation of Configurations for Software AvailabilityabstractThe Availability Management Framework (AMF) is a service responsible for managing the availability of services provided by applications that run under its control. Standardized by the Service Availability Forum (SAF), AMF requires for its operations a complete and compliant AMF configuration of the applications to be managed. In this paper, we describe two complementary and integrated tools for AMF configurations generation and validation. Indeed, writing manually an AMF configuration is a tedious and error prone task as a large number of requirements defined in the standard have to be taken into consideration during the process. One solution for ensuring compliance with the standard is the validation of the configurations against all the AMF requirements. For this, we have designed and implemented a domain model for AMF configurations and use it as a basis for an AMF configuration validator. To further ease the task of a configuration designer, we have devised and implemented a method for generating automatically AMF configurations. Abdelouahed Gherbi, Ali Kanso, Ferhat Khendek, Abdelwahab Hamou-Lhadj, Maria Toeroe |
ASE | 4 |
| 2008 | An Approach for Mapping Features to Code Based on Static and Dynamic AnalysisabstractSystem evolution depends greatly on the ability of a maintainer to locate source code that is specific to feature implementation. Existing feature location techniques require either exercising several features of the system, or rely heavily on domain experts to guide the feature location process. In this paper, we present a novel approach for feature location that combines static and dynamic analysis techniques. An execution trace is generated by exercising the feature under study (dynamic analysis). A component dependency graph (static analysis) is used to rank the components invoked in the trace according to their relevance to the feature. Our ranking technique is based on the impact of a component modification on the rest of the system. The proposed approach is automatic to a large extent relieving users from any decision that would otherwise require extensive domain knowledge of the system. A case study is presented to support and evaluate the applicability of our approach. Abhishek Rohatgi, Abdelwahab Hamou-Lhadj, Juergen Rilling |
ICPC | 2 |
| 2008 | Introduction to the special issue on program comprehension through dynamic analysis (PCODA)abstractThis special issue on program comprehension through dynamic analysis is tightly related to the international workshop on program comprehension through dynamic analysis (PCODA) series. The aim of PCODA is to bring together researchers and practitioners using dynamic analysis as a basis for their program comprehension and reverse engineering technique(s). Within the reverse engineering community much attention is focused on static analysis, the analysis of the source code of a software system, while dynamic analysis focusing on runtime properties of software systems has often been neglected. Nevertheless, dynamic analysis is recognized as yielding a more precise analysis in the face of polymorphism, a language feature widely used in object-oriented software systems. PCODA was first co-located with WCRE 2005, the 12th Working Conference on Reverse Engineering, in Pittsburgh, and over the past three years has proved to be a very successful event, attracting a constant number of attendees and high-quality submissions. We chose to adopt a unique format for the half day PCODA workshop. The authors do not present their own papers; each paper is assigned in advance to another participant who then presents a summary of the paper at the workshop. Experience has shown that the key advantage of adopting this format is that authors are given the opportunity to see an external interpretation of their work, which in turn leads to interesting and lively discussions among the authors, the presenter and the audience. Thus, it is not surprising that during the PCODA workshop an equal amount of time is devoted to discussion and presentation. Three highly successful PCODA workshops have led to this special issue of the Journal of Software Maintenance and Evolution: Research and Practice devoted to the topic of program comprehension through dynamic analysis. The authors of two papers from the PCODA 2007 workshop were invited to submit extended versions of their papers. An additional 10 papers were submitted through an open call. Each paper was subjected to a rigorous reviewing process, involving three independent reviewers with expert knowledge in the area and two rounds of revisions. Subsequently four papers were accepted for inclusion in this special issue. In the paper ‘Mining Temporal Rules for Software Maintenance’ Lo, Khoo and Liu describe a technique to mine statistically significant temporal rules of arbitrary length to describe system behavior from sets of traces. They represent their rules as temporal logic expressions that then serve as input to formal analysis toolkits supporting program comprehension, program verification, debugging and specification mining. They demonstrate the scalability of their technique by applying it to two open-source case study applications. A key contribution of this work is that each condition of a rule can consist of multiple events. In the paper ‘An Automated Approach for Abstracting Execution Logs to Execution Events’ Jiang, Hassan, Hamann and Flora present a technique for abstracting execution logs generated from output statements inserted by developers in the source code. They apply clone detection techniques, the static and dynamic part of each log line is detected, and subsequently each log line is abstracted to its corresponding execution event. In the paper ‘Improving Dynamic Software Analysis by Applying Grammar Inference Principles’ Walkinshaw, Bogdanov, Holcombe and Salahuddin propose to improving dynamic analysis by applying grammar inference principles. The authors' approach aims at tackling the fact that dynamic analysis can provide only a partial view of the system. Large traces can be analyzed to infer properties about a program's behavior but it is difficult to obtain a complete state of traces which covers all possible execution paths. Grammar inference is an example of an other field that suffers from the problem of incomplete samples for inferring grammar rules. This paper argues that many solutions in grammar inference, which produce reliably accurate approximations of regular grammars, can be applied with similar effect to improve dynamic analysis techniques. The authors perform three experiments that show the effect of adopting particular grammar inference principles on the accuracy of dynamic analysis techniques. In the paper ‘A Survey and Evaluation of Tool Features for Understanding Reverse Engineered Sequence Diagrams’, Bennett, Myers, Ouellet, Storey, Salois, German and Charland present a thorough study of the features supported by several tools that focus on the exploration of sequence diagrams, reverse engineered from large execution traces. They have developed a prototype tool, called OASIS, to evaluate the usefulness of these features in understanding the behavior of a software system. The result of the experiment confirms that most existing features are indeed useful in a variety of reverse engineering tasks. In addition, the paper presents the results of a user study that focuses on understanding how existing tools can be improved. Several improvements have been proposed such as the ability to save the state of the session, the ability to navigate between the source code and the extracted sequence diagram, etc. Another important contribution of the paper consists of a rich discussion on how to improve cognitive support in reverse-engineered sequence diagram tools. We hope that readers will enjoy this special issue and through these papers gain useful insights into the domain of dynamic analysis. We would like to thank all the authors who submitted their papers to the PCODA workshop series and to this special issue. A special thank you goes to the external reviewers who helped in making this special issue a highly qualitative one. Finally, we would like to thank Aniello Cimitile, the editor in chief of the Journal of Software Maintenance and Evolution: Research and Practice, and the publisher Wiley for providing us with the opportunity to devote an issue of this journal to PCODA. The organization of the PCODA workshop series and this special issue has been sponsored by the Netherlands Organisation for Scientific Research (NWO) through the ‘Jacquard RECONSTRUCTOR’ project (2005–2009), the Swiss National Science Foundation through the project ‘Analyzing, capturing and taming software change’ (SNF Project No. 200020-113342, October 2006–September 2008) and the Natural Sciences and Engineering Research Council of Canada (NSERC) for the project ‘Program Comprehension through Dynamic Analysis’ (April 2007–March 2012) led by the DASS (Dynamic Analysis of Software Systems) research group at ECE, Concordia University, Canada. Andy Zaidman, Abdelwahab Hamou-Lhadj, Orla Greevy |
J. Softw. Maintenance Res. Pract. | 2 |
| 2006 | Summarizing the Content of Large Traces to Facilitate the Understanding of the Behaviour of a Software SystemabstractIn this paper, we present a semi-automatic approach for summarizing the content of large execution traces. Similar to text summarization, where abstracts can be extracted from large documents, the aim of trace summarization is to take an execution trace as input and return a summary of its main content as output. The resulting summary can then be converted into a UML sequence diagram and used by software engineers to understand the main behavioural aspects of the system. Our approach to trace summarization is based on the removal of implementation details such as utilities from execution traces. To achieve our goal, we have developed a metric based on fan-in and fan-out to rank the system components according to whether they implement key system concepts or they are mere implementation details. We applied our approach to a trace generated from an object-oriented system called Weka that initially contains 97413 method calls. We succeeded to extract a summary from this trace that contains 453 calls. According to the developers of the Weka system, the resulting summary is an adequate high-level representation of the main interactions of the traced scenario Abdelwahab Hamou-Lhadj, Timothy Lethbridge |
ICPC | 1 |
| 2005 | Measuring Various Properties of Execution Traces to Help Build Better Trace Analysis ToolsabstractUnderstanding the behavior of a software system by studying its execution traces can be extremely difficult due to the sheer size and complexity of typical traces. In this paper, we propose that if various aspects that contribute to a trace's complexity could be measured and if this information could be used by tools, then trace analysis could be facilitated. For this purpose, we present a set of simple and practical metrics that aim at measuring various properties of execution traces. We also show the results of applying these metrics to traces of three software systems and suggest how the results could be used to improve existing trace analysis tools. Abdelwahab Hamou-Lhadj, Timothy Lethbridge |
ICECCS | 1 |