EDBT 2026 Demo / reviewers in the wild / expert
Ettore Merlo
dblp:47/4667
· DBLP profile ↗
65ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0002-1436-6076ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 53 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorSecurity and privacy · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Story About Cohesion and Separation: Label-Free Metric for Log Parser Evaluation
Qiaolin Qin, Jianchen Zhao, Heng Li 0007, Weiyi Shang, Ettore Merlo |
SANER | 5 |
| 2026 | An empirical study on logging evolution on stack overflow: trends, topics, and challenges
Patrick Loic Foalem, Andre Nguimbous, Foutse Khomh, Heng Li 0007, Ettore Merlo |
Empir. Softw. Eng. | 5 |
| 2026 | Plug it and Play on Logs: A configuration-free statistic-based log parser
Qiaolin Qin, Xingfang Wu, Heng Li 0007, Ettore Merlo |
Empir. Softw. Eng. | 4 |
| 2026 | Unsupervised, robust, and lightweight detection of data pattern anomalies and outliersabstractContext: As a current consensus, data quality strongly impacts the process of building software and AI systems. Hence, practitioners must detect the anomalies in data and repair these underlying problems. When dealing with big data in the industry, statistic-based unsupervised anomaly detectors come in handy since they do not require labels and are highly scalable. However, we noticed that these tools unsupervised, always require data-dependent parameters, which can largely affect the detection performance and are effort-consuming to configure. Objectives: In this work, we propose a fully unsupervised, statistic-based cell-level data anomaly detector, LUCARIO (Learning Unsupervised, Cell-level Anomaly-detector for Regex Incompatibilities and Outliers). Our approach aims to detect common cell-level data anomalies (pattern violations and outliers) without manual efforts in data annotations or parameter configurations, yet providing a robust performance for different data across diverse domains. Methods: According to previous studies, we categorized cell anomalies into two categories: pattern violations and outliers (categorical and numerical). We proposed three detection approaches based on heuristics and statistical theories to identify these anomalies. To evaluate LUCARIO’s effectiveness and usability, we conducted experiments on six open-source datasets and a real-life industrial dataset from our industrial partner CompanyX . Results: According to our experiment on six open-source datasets in various domains, LUCARIO can stably detect cell-level data issues (pattern violations and outliers) regardless of the dataset’s size and anomaly rate. LUCARIO reached an average F1 score of 0.54, higher than all baseline unsupervised anomaly detectors, including GPT-5 with few-shot prompting. Practitioners from CompanyX generally agree that LUCARIO can benefit their data quality by detecting critical data issues and providing reliable suggestions. Conclusion: The experimental results show that LUCARIO has the potential to improve the data used for both software and AI system construction in real-life applications, suggesting its practicality in data management. Qiaolin Qin, Heng Li 0007, Ettore Merlo |
Inf. Softw. Technol. | 3 |
| 2026 | Prune bias from the root: Bias removal and fairness estimation by pruning sensitive attributes in pre-trained DNN modelsabstractDeep learning models (DNNs) are widely used in high-stakes decision-making domains, but they often inherit and amplify biases present in training data, leading to unfair predictions. Given this context, fairness estimation metrics and bias removal methods are required to select and enhance fair models. However, we found that existing metrics lack robustness in estimating multi-attribute group fairness. Further, existing post-processing bias removal methods often focus on group fairness and fail to address individual fairness or optimize along multiple sensitive attributes. In this study, we explore the effectiveness of attribute pruning (i.e., zeroing out sensitive attribute weights in a pre-trained DNN’s input layer) in both bias removal and multi-attribute group fairness estimation. To study attribute pruning’s impact on bias removal, we conducted experiments on 32 models and 4 widely used datasets, and compared its effect in single-attribute group bias removal and accuracy preservation with 3 baseline post-processing methods. We then leveraged 3 datasets with multiple sensitive attributes to demonstrate how to use attribute pruning for multi-attribute group fairness estimation. Single-attribute pruning can better preserve model accuracy than conventional post-processing methods in 23 out of 32 cases, and enforces individual fairness by design. However, since individual fairness and group fairness are fundamentally different objectives, attribute pruning’s effect on group fairness metrics is often inconsistent. We also extend our approach to a multi-attribute setting, demonstrating its potential for improving individual fairness jointly across sensitive attributes and for enabling multi-attribute fairness-aware model selection. Attribute pruning is a practical post-processing approach for enforcing individual fairness, with limited and data-dependent impact on group fairness. These limitations reflect the inherent trade-off between individual and group fairness objectives. In addition, attribute pruning provides a useful mechanism for bias estimation, particularly in multi-attribute contexts. We advocate for its adoption as a comparison baseline in fairness-aware AI development and encourage further exploration. • Prove the effectiveness of single-attribute pruning for bias removal on DNNs. • Propose the use of attribute pruning on multiple features. • Introduce a robust multi-attribute fairness estimation method. • Explicitly discuss the limits of algorithmic bias removal and provide suggestions. Qiaolin Qin, Ettore Merlo |
Inf. Softw. Technol. | 2 |
| 2025 | Automated, Unsupervised, and Auto-Parameterized Inference of Data Patterns and Anomaly DetectionabstractWith the advent of data-centric and machine learning (ML) systems, data quality is playing an increasingly critical role for ensuring the overall quality of software systems. Data preparation, an essential step towards high data quality, is known to be a highly effort-intensive process. Although prior studies have dealt with one of the most impacting issues, data pattern violations, these studies usually require data-specific configurations (i.e., parameterized) or use carefully curated data as learning examples (i.e., supervised), relying on domain knowledge and deep understanding of the data, or demanding significant manual effort. In this paper, we introduce RIOLU: Regex Inferencer autO-parameterized Learning with Uncleaned data. RIOLU is fully automated, automatically parameterized, and does not need labeled samples. RIOLU can generate precise patterns from datasets in various domains, with a high F1 score of 97.2 %, exceeding the state-of-the-art baseline. In addition, according to our experiment on five datasets with anomalies, RIOLU can automatically estimate a data column's error rate, draw normal patterns, and predict anomalies from unlabeled data with higher performance (up to$\mathbf{8 0 0. 4 \%}$improvement in terms of F1) than the state-of-the-art baseline, even outperforming ChatGPT in terms of both accuracy (12.3 % higher F1) and efficiency (10 % less inference time). A variant of RIOLU, with user guidance, can further boost its precision, with up to$\mathbf{3 7. 4 \%}$improvement in terms of F1. Our evaluation in an industrial setting further demonstrates the practical benefits of RIOLU. Qiaolin Qin, Heng Li 0007, Ettore Merlo, Maxime Lamothe |
ICSE | 3 |
| 2025 | Effective, Efficient, and Environmentally Friendly Out-of-Model-Scope Detection MethodologyabstractIntegrating deep neural networks (DNNs) in safety-critical systems is widespread, but their reliability depends on accurate performance in real-world environments. Capturing all scenarios in training data is impractical. One solution is to use DNNs within their known range and alert a human operator when encountering unreliable outputs, i.e., out-of-model-scope (OMS) outputs. However, current unsupervised OMS detection methods monitor all neurons, are computationally expensive, and are not robust enough to avoid neuron noises. In this paper, we propose an effective, efficient, and environmentally friendly methodology, EFOMS, that automatically filters unreliable outputs by extending existing OMS detection methods to focus only on significant neurons, thereby filtering out noise from unimportant neurons. EFOMS achieves comparable or better OMS detection quality with significantly reduced computational costs: 45% faster, consuming 30% less energy, producing 30% fewer carbon emissions, and using up to 21% less peak memory. Ettore Merlo, Clément Benesse, Lina Marsso |
ISSRE | 2 |
| 2025 | Preprocessing is All You Need: Boosting the Performance of Log Parsers with a General Preprocessing FrameworkabstractLog parsing has been a long-studied area in software engineering due to its importance in identifying dynamic vari-ables and constructing log templates. Prior work has proposed many statistic-based log parsers (e.g., Drain), which are highly efficient; they, unfortunately, met the bottleneck of parsing performance in comparison to semantic-based log parsers, which require labeling and more computational resources. Meanwhile, we noticed that previous studies mainly focused on parsing and often treated preprocessing as an ad hoc step (e.g., masking numbers). However, we argue that both preprocessing and parsing are essential for log parsers to identify dynamic variables: the lack of understanding of preprocessing may hinder the optimal use of parsers and future research. Therefore, our work studied existing log preprocessing approaches based on Loghub, a popular log parsing benchmark. We developed a general preprocessing framework with our findings and evaluated its impact on existing parsers. Our experiments show that the preprocessing framework significantly boosts the performance of four state-of-the-art statistic-based parsers. Drain, the best statistic-based parser, obtained improvements across all four parsing metrics (e.g., Fl score of template accuracy, FTA, increased by 108.9%). Compared to semantic-based parsers, it achieved a 28.3% improvement in grouping accuracy (GA), 38.1 % in FGA, and an 18.6% increase in FTA. Our work pioneers log preprocessing and provides a generalizable framework to enhance log parsing. Qiaolin Qin, Roozbeh Aghili, Heng Li 0007, Ettore Merlo |
SANER | 4 |
| 2025 | Fairness Evaluation of Neural Networks Through Computational Profile LikelihoodabstractABSTRACT Despite high predictive performance, machine learning models can be unfair towards specific demographic subgroups characterized by sensitive attributes such as gender or race. This paper presents a novel approach using Computational Profile Likelihood (CPL) to assess potential bias in neural network decisions with respect to sensitive attributes. CPL estimates the conditional probability of a network's internal neuron excitation levels during predictions. To assess the impact of sensitive attributes on predictions, the CPL distribution of individuals sharing a particular value of a sensitive attribute and a specific outcome (e.g., “women” and “high income”) is compared to a subgroup sharing another value of the sensitive attribute but with the same outcome (e.g., “men” and “high income”). The resulting disparities between distributions can be used to quantify the bias with respect to the sensitive attribute and the outcome class. We also assess the efficacy of bias reduction techniques through their influence on the resulting disparities. Experimental results on three widely used datasets indicate that the CPL of the trained models can be used to characterize significant differences between multiple protected groups, highlighting that these models display quantifiable biases. Furthermore, after applying bias mitigation methods, the gaps in CPL distributions are reduced, indicating a more similar internal representation for profiles of different protected groups. Benjamin Djian, Ettore Merlo, Sébastien Gambs, Rosin Claude Ngueveu |
Comput. Intell. | 2 |
| 2025 | Logging requirement for continuous auditing of responsible machine learning-based applications
Patrick Loic Foalem, Léuson M. P. da Silva, Foutse Khomh, Heng Li 0007, Ettore Merlo |
Empir. Softw. Eng. | 5 |
| 2025 | Representation-based fairness evaluation and bias correction robustness assessment in neural networksabstractContext: While machine learning has achieved high predictive performance in many domains, decisions may still be biased and unfair regarding specific demographic groups characterized by sensitive attributes such as gender, age, or race. Objectives: In this paper, we introduce a novel approach to assess model fairness and bias correction robustness based on Computational Profile Distance (CPD) analysis with respect to sensitive attributes. Methods: To study model fairness, we quantify the model’s representation difference using the computational profile learned from different subgroups (e.g., male and female) on the individual and group level. To analyze the robustness of bias correction outcomes, we compare the correction suggestions provided based on confidence (i.e., softmax score) and likelihood (i.e., CPD). Results: To demonstrate the potential of the proposed approach, experiments have been performed using 24 models targeting 3 datasets used in previous fairness studies. Our experiments showed that computational profile distributions can effectively address model fairness from a representation perspective. Further, the experiments indicated that confidence-based bias correction decisions can vary largely from likelihood-based ones, and we should take both suggestions into account to obtain robust outcomes. Conclusion: Demonstrated with a set of experiments, our CPD-based approaches can help users build their trust in fairness assessment and bias mitigation of AI decisions, in ethically sensitive domains such as human resources, finance, health, and more. Qiaolin Qin, Benjamin Djian, Ettore Merlo, Heng Li 0007, Sébastien Gambs |
Inf. Softw. Technol. | 3 |
| 2025 | Recovering Traceability Links Between Code and Documentation: A RetrospectiveabstractSoftware system documentation is almost always expressed informally in natural language and free text. Examples include requirement specifications, design documents, manual pages, system development journals, error logs, and related maintenance reports. In our 2002 seminal paper we proposed a method based on information retrieval to recover traceability links between source code and free text documents. A premise of our work was that programmers use meaningful names for program items, such as functions, variables, types, classes, and methods. The paper paved the way to the adoption of IR in software engineering opening a new perspective. Reflecting on the past twenty years we briefly overview the many results that have been achieved, however, the emergence of new technologies, such as AI, pose unprecedented challenges. Giuliano Antoniol, Gerardo Canfora, Gerardo Casazza, Andrea De Lucia, Ettore Merlo |
IEEE Trans. Software Eng. | 5 |
| 2023 | Unsupervised Graph Neural Networks for Source Code Similarity Detection
Julien Cassagne, Ettore Merlo, Paula Branco, Guy-Vincent Jourdan, Iosif-Viorel Onut |
DS | 2 |
| 2022 | Identification of out-of-distribution cases of CNN using class-based surprise adequacyabstractMachine learning is vulnerable to possible incorrect classification of cases that are out of the distribution observed during training and calibration. Mira Marhaba, Ettore Merlo, Foutse Khomh, Giuliano Antoniol |
CAIN | 2 |
| 2022 | Phishing Kits Source Code Similarity Distribution: A Case StudyabstractAttackers (“phishers”) typically deploy source code in some host website to impersonate a brand or in general a situation in which a user is expected to provide some personal information of interest to phishers (e.g. credentials, credit card number). Phishing kits are ready-to-deploy sets of files that can be simply copied on a web server and used almost as they are. In this paper, we consider the static similarity analysis of the source code of 20871 phishing kits totalling over 182 million lines of PHP, Javascript and HTML code, that have been collected during phishing attacks and recovered by forensics teams. Reported experimental results show that as much as 90% of the analyzed kits share 90% or more of their source code with at least another kit. Differences are small, less than about 1000 programming words – identifiers, constants, strings and so on – in 40% of cases. A plausible lineage of phishing kits is presented by connecting together kits with the highest similarity. Obtained results show a very different reconstructed lineage for phishing kits when compared to a publicly available application such as Wordpress. Observed kits similarity distribution is consistent with the assumed hypothesis that kit propagation is often based on identical or near-identical copies at low cost changes. The proposed approach may help classifying new incoming phishing kits as “near-copy” or “intellectual leaps” from known and already encountered kits. This could facilitate the identification and classification of new kits as derived from older known kits. Ettore Merlo, Mathieu Margier, Guy-Vincent Jourdan, Iosif-Viorel Onut |
SANER | 1 |
| 2022 | How to certify machine learning based safety-critical systems? A systematic literature review
Florian Tambon, Gabriel Laberge, Amin Nikanjam, Paulina Stevia Nouwou Mindom, Yann Pequignot, Foutse Khomh, Giuliano Antoniol, Ettore Merlo, François Laviolette |
Autom. Softw. Eng. | 9 |
| 2021 | RBAC protection-impacting changes identification: A case study of the security evolution of two PHP applications
Marc-André Laverdière, Karl Julien, Ettore Merlo |
Inf. Softw. Technol. | 3 |
| 2020 | An Effective Evolutionary Analysis Scheme for Industrial Software Access Control ModelsabstractAccess control is an essential feature of industrial software systems security mechanisms. Role-based access control (RBAC), which is likely the most popular access-control technique, specifies “user roles” and associates each role with “permissions” to access distinct system functionalities. These role-permissions assignment rules, as well as the types of system users and system functionalities, evolve over time. In this paper, we describe a methodology for analyzing and understanding the RBAC-configuration evolution, its relation to the overall evolutionary lifecycle of industrial systems, and its impact on security vulnerabilities from which the system may suffer. Our methodology considers two different sources of information regarding the RBAC-configuration evolution: 1) the role-permissions matrices of subsequent system versions; and 2) the corresponding concept lattices, implied by these matrices. By examining the evolution of these two system properties, developers can easily notice which versions involve more and more complex RBAC-configuration changes that may indicate higher security risks. We demonstrate our methodology using a study of four popular real-world systems: 1) MediaWiki; 2) Moodle; 3) Joomla; and 4) WordPress. Our findings show that the proposed metrics have strong, positive linear correlations with the security vulnerabilities' properties. Zhuobing Han, Xiaohong Li 0001, Guangquan Xu, Naixue Xiong, Ettore Merlo, Eleni Stroulia |
IEEE Trans. Ind. Informatics | 5 |
| 2018 | Detection of protection-impacting changes during software evolutionabstractRole-Based Access Control (RBAC) is often used in web applications to restrict operations and protect security sensitive information and resources. Web applications regularly undergo maintenance and evolution and their security may be affected by source code changes between releases. To prevent security regression and vulnerabilities, developers have to take re-validation actions before deploying new releases. This may become a significant undertaking, especially when quick and repeated releases are sought. We define protection-impacting changes as those changed statements during evolution that alter privilege protection of some code. We propose an automated method that identifies protection-impacting changes within all changed statements between two versions. The proposed approach compares statically computed security protection models and repository information corresponding to different releases of a system to identify protection-impacting changes. Results of experiments present the occurrence of protection-impacting changes over 210 release pairs of WordPress, a PHP content management web application. First, we show that only 41% of the release pairs present protection-impacting changes. Second, for these affected release pairs, protection-impacting changes can be identified and represent a median of 47.00 lines of code, that is 27.41% of the total changed lines of code. Over all investigated releases in WordPress, protection-impacting changes amounted to 10.89% of changed lines of code. Conversely, an average of about 89% of changed source code have no impact on RBAC security and thus need no re-validation nor investigation. The proposed method reduces the amount of candidate causes of protection changes that developers need to investigate. This information could help developers re-validate application security, identify causes of negative security changes, and perform repairs in a more effective way. Marc-André Laverdière, Ettore Merlo |
SANER | 2 |
| 2017 | Classification and Distribution of RBAC Privilege Protection Changes in Wordpress Evolution (Short Paper)abstractRole-Based Access Control (RBAC) is commonly used in web applications to protect information and restrict operations. Their security may be affected by source code changes between releases in unexpected ways. To prevent regression and vulnerabilities, developers need to validate them prior to each release, which may be a major undertaking. We automatically and statically determine privilege-level security impacts of code changes using privilege protection changes and apply a set-theoretic classification to them. To do so, we analyze code and determine the security privilege protection models of Web applications written in PHP using Pattern Traversal Flow Analysis (PTFA). We present the distribution of both privilege protection changes and their classification over 147 release pairs of WordPress, spanning from 2.0 to 4.5.1. We found that code changes had no impact on privilege protection in the 82 (56%) release pairs. The remaining 65 (44%) release pairs are affected by privilege protection changes. For the latter release pairs, only 0.30% of code is affected by privilege protection changes. We also found that the most common change categories are complete gains (40.81%), complete losses (17.99%) and substitution (20.10%). The automated identification and classification of privilege protection changes may help developers to more efficiently focus their effort during security reviews, verification, validation, testing, and repairs. Marc-André Laverdière, Ettore Merlo |
PST | 2 |
| 2017 | Computing counter-examples for privilege protection losses using security modelsabstractRole-Based Access Control (RBAC) is commonly used in web applications to protect information and restrict operations. Code changes may affect the security of the application and need to be validated, in order to avoid security vulnerabilities, which is a major undertaking. A statement suffers from privilege protection loss in a release pair when it was definitely protected on all execution paths in the previous release and is now reachable by some execution paths with an inferior privilege protection. Because the code change and the resulting privilege protection loss may be distant (e.g. in different functions or files), developers may find it difficult to diagnose and correct the issue. We use Pattern Traversal Flow Analysis (PTFA) to statically analyze code-derived formal models. Our analysis automatically computes counter-examples of definite protection properties and privilege protection losses. We computed privilege protections and their changes for 147 release pairs of WordPress. We computed counter-examples for a total of 14,116 privilege protection losses we found spread in 31 release pairs.We present the distribution of counter-examples' lengths, as well as their spread across function and file boundaries. Our results show that counter-examples are typically short and localized. The median example spans 88 statements, crosses a single function boundary, and is contained in the same file. The 90thcentile example measures 174 statements and spans 3 function boundaries over 3 files. We believe that the privilege protection counter-examples' characteristics would be helpful to focus developers' attention for security reviews. These counter-examples are also a first step toward explanations. Marc-André Laverdière, Ettore Merlo |
SANER | 2 |
| 2017 | A case study of TTCN-3 test scripts clone analysis in an industrial telecommunication setting
Thierry Lavoie, Mathieu Mérineau, Ettore Merlo, Pascal Potvin |
Inf. Softw. Technol. | 3 |
| 2015 | Taint analysis of manual service compositions using Cross-Application Call GraphsabstractWe propose an extension over the traditional call graph to incorporate edges representing control flow between web services, named the Cross-Application Call Graph (CACG). We introduce a construction algorithm for applications built on the Jax-WS standard and validate its effectiveness on sample applications from Apache CXF and JBossWS. Then, we demonstrate its applicability for taint analysis over a sample application of our making. Our CACG construction algorithm accurately identifies service call targets 81.07% of the time on average. Our taint analysis obtains a F-Measure of 95.60% over a benchmark. The use of a CACG, compared to a naive approach, improves the F-Measure of a taint analysis from 66.67% to 100.00% for our sample application. Marc-André Laverdière, Bernhard J. Berger, Ettore Merlo |
SANER | 3 |
| 2014 | Supporting Maintenance and Evolution of Access Control Models in Web ApplicationsabstractThis paper presents an approach to support the maintenance and evolution of Role-Based Access Control (RBAC) models with reverse-engineered Secure UML models. Starting from the Policy Decision Points (PDP) and Policy Enforcement Points (PEP) of an application, our approach statically reverse-engineers the implemented Secure UML model of an application. The secure UML model is then stored in an RDF triple store for easy querying and exploration. In the context of this study, we extracted the Secure UML model of the GRAND Forum, a web-based forum for the members of the GRAND (Graphics, Animation and New Media) NCE (Networks of Centers of Excellence), that is developed and maintained at the University of Alberta. Using three real use-case scenarios, we illustrate how simple queries to the extracted Secure UML can save developers significant amounts of manual work and support them in their access control related maintenance and evolution tasks. François Gauthier 0001, Ettore Merlo, Eleni Stroulia |
ICSME | 2 |
| 2013 | Uncovering access control weaknesses and flaws with security-discordant software clonesabstractSoftware clone detection techniques identify fragments of code that share some level of syntactic similarity. In this study, we investigate security-sensitive clone clusters: clusters of syntactically similar fragments of code that are protected by some privileges. From a security perspective, security-sensitive clone clusters can help reason about the implemented security model: given syntactically similar fragments of code, it is expected that they are protected by similar privileges. We hypothesize that clones that violate this assumption, defined as security-discordant clones, are likely to reveal weaknesses and flaws in access control models. François Gauthier 0001, Thierry Lavoie, Ettore Merlo |
ACSAC | 3 |
| 2013 | Semantic smells and errors in access control models: a case study in PHPabstractAccess control models implement mechanisms to restrict access to sensitive data from unprivileged users. Access controls typically check privileges that capture the semantics of the operations they protect. Semantic smells and errors in access control models stem from privileges that are partially or totally unrelated to the action they protect. This paper presents a novel approach, partly based on static analysis and information retrieval techniques, for the automatic detection of semantic smells and errors in access control models. Investigation of the case study application revealed 31 smells and 2 errors. Errors were reported to developers who quickly confirmed their relevance and took actions to correct them. Based on the obtained results, we also propose three categories of semantic smells and errors to lay the foundations for further research on access control smells in other systems and domains. François Gauthier 0001, Ettore Merlo |
ICSE | 2 |
| 2012 | Locating features in dynamically configured avionics softwareabstractLocating features in software is an important activity for program comprehension and to support software reengineering. We present a novel automated approach to locate features in source code based on static analysis and model checking. The technique is aimed at dynamically configured software, which is software in which the activation of specific features is controlled by configuration variables. The approach is evaluated on an industrial avionics system. Maxime Ouellet, Ettore Merlo, Neset Sozen, Martin Gagnon |
ICSE | 2 |
| 2012 | Alias-Aware Propagation of Simple Pattern-Based Properties in PHP ApplicationsabstractIn this paper, we present novel algorithms for the propagation of pattern-based properties in PHP applications. Intuitively, pattern-based properties designate those properties that are intrinsically associated to syntactic patterns in the source code. Security checks in access control models are an example of pattern-based properties. At the source code level, permissions are typically verified with stereotyped constructs, called security checks, that can be detected with syntactic patterns. Depending on the program, pattern-based properties can be a liased to variables that are propagated through the application. In that context, support from data-flow approaches is needed to track the propagation of patterns through the application. In the context of this paper, we focus on the alias-aware propagation of security checks through PHP applications. Specifically, we investigated the propagation of security checks in 8 PHP applications that implement access control models. We show how, using the Data log language, one can implement conceptually complex data-flow algorithms in an incremental, intuitive and compact manner. From the results perspective, we show how our algorithm identifies security checks and security check a liased variables in a precise way. The reported false positive rate varies between 0% and 4% for the investigated applications. François Gauthier 0001, Ettore Merlo |
SCAM | 2 |
| 2011 | Security Model Evolution of PHP Web ApplicationsabstractWeb sites are often a mixture of static sites and programs that integrate relational databases as a back-end. As they evolve to meet ever-changing user needs, new versions of programs, interactions and functionalities may be added and existing ones may be removed or modified. Web sites require configuration and programming attention to assure security, confidentiality, and trust of the published information. During evolution of Web software, from one version to the next one, security properties may change and possible changes may include new flaws or corrections. Changes to security properties, including access control privileges, can be monitored by observing and analyzing changes between security models extracted from different versions of an application. This paper defines Property Satisfaction Profiles (PSP) as the satisfaction values of properties computed on the extracted models. This paper presents also an investigation of the evolution of the changes in the PSP computed on security models of different versions of a Web application. Model extraction and PSP computation can be performed in linear time on one version. Comparison between two versions is also linear and practical performance is fast. This paper reports results about experiments performed on 31 versions of phpBB, that is a publicly available bulletin board written in PHP. Version 1.0.0 (9547 LOC) to version 2.0.22 (40663 LOC) have been considered as a case study. Results show that the proposed approach can be used to observe and monitor the evolution of PSP in successive versions of the same software package. Suggestions for further research are also presented. Dominic Letarte, François Gauthier 0001, Ettore Merlo |
ICST | 3 |
| 2011 | Extraction and comprehension of moodle's access control model: A case studyabstractWhether for development, maintenance or refactoring, multiple steps in software development cycle require comprehension of a program's access control model (AC model). In this paper, we present a novel approach to reverse-engineer AC model structure from PHP source code. Using an hybrid approach combining static analysis and model checking techniques, we are able to extract AC model structure in a fast and precise way. An experimental tool was developed to evaluate the presented approach and report AC models using source code coloring. For this case study, Moodle, a medium-scale (approx. 625K lines of code), open-source PHP application with a rich AC model was investigated. Results revealed that, although very complex by design, implemented AC models may comparatively be very simple, suggesting that developers tend to maintain a low complexity level when implementing ACs. Detailed figures and distributions are reported. We believe the presented tool and approach may help in understanding and evaluating the implemented AC models in Web systems. Discussion of findings, limitations, and further research are presented. François Gauthier 0001, Dominic Letarte, Thierry Lavoie, Ettore Merlo |
PST | 4 |
| 2007 | Mining the Lexicon Used by Programmers during Sofware EvolutionabstractIdentifiers represent an important source of information for programmers understanding and maintaining a system. Self-documenting identifiers reduce the time and effort necessary to obtain the level of understanding appropriate for the task at hand. While the role of the lexicon in program comprehension has long been recognized, only a few works have studied the quality and enhancement of the identifiers and no works have studied the evolution of the lexicon. In this paper, we characterize the evolution of program identifiers in terms of stability metrics and occurrences of renaming. We assess whether an evolution process similar to the one occurring for the program structure exists for identifiers. We report data and results about the evolution of three large systems, for which several releases are available. We have found evidence that the evolution of the lexicon is more limited and constrained than the evolution of the structure. We argue that the different evolution results from several factors including the lack of advanced tool support for lexicon construction, documentation, and evolution. Giuliano Antoniol, Yann-Gaël Guéhéneuc, Ettore Merlo, Paolo Tonella |
ICSM | 3 |
| 2007 | Comparison and Evaluation of Clone Detection ToolsabstractMany techniques for detecting duplicated source code (software clones) have been proposed in the past. However, it is not yet clear how these techniques compare in terms of recall and precision as well as space and time requirements. This paper presents an experiment that evaluates six clone detectors based on eight large C and Java programs (altogether almost 850 KLOC). Their clone candidates were evaluated by one of the authors as independent third party. The selected techniques cover the whole spectrum of the state-of-the-art in clone detection. The techniques work on text, lexical and syntactic information, software metrics, and program dependency graphs. Stefan Bellon, Rainer Koschke, Giuliano Antoniol, Jens Krinke, Ettore Merlo |
IEEE Trans. Software Eng. | 5 |
| 2006 | A novel approach to optimize clone refactoring activityabstractSoftware evolution and software quality are ever changing phenomena. As software evolves, evolution impacts software quality. On the other hand, software quality needs may drive software evolution strategies.This paper presents an approach to schedule quality improvement under constraints and priority. The general problem of scheduling quality improvement has been instantiated into the concrete problem of planning duplicated code removal in a geographical information system developed in C throughout the last 20 years. Priority and constraints arise from development team and from the adopted development process. The developer team long term goal is to get rid of duplicated code, improve software structure, decrease coupling, and improve cohesion.We present our problem formulation, the adopted approach, including a model of clone removal effort and preliminary results obtained on a real world application. Salah Bouktif, Giuliano Antoniol, Ettore Merlo, Markus Neteler |
GECCO | 3 |
| 2006 | A Feedback Based Quality Assessment to Support Open Source Software Evolution: the GRASS Case StudyabstractManaging the software evolution for large open source software is a major challenge. Some factors that make software hard to maintain are geographically distributed development teams, frequent and rapid turnover of volunteers, absence of a formal means, and lack of documentation and explicit project planning. In this paper we propose remote and continuous analysis of open source software to monitor evolution using available resources such as CVS code repository, commitment log files and exchanged mail. Evolution monitoring relies on three principal services. The first service analyzes and monitors the increase in complexity and the decline in quality; the second supports distributed developers by sending them a feedback report after each contribution; the third allows developers to gain insight into the "big picture" of software by providing a dashboard of project evolution. Besides the description of provided services, the paper presents a prototype environment for continuous analysis of the evolution of GRASS, an open source software. Salah Bouktif, Giuliano Antoniol, Ettore Merlo |
ICSM | 3 |
| 2005 | Improving network applications security: a new heuristic to generate stress testing dataabstractBuffer overflows cause serious problems in different categories of software systems. For example, if present in network or security applications, they can be exploited to gain unauthorized grant or access to the system. In embedded systems, such as avionics or automotive systems, they can be the cause of serious accidents.This paper proposes to combine static analysis and program slicing with evolutionary testing, to detect buffer overflow threats. Static analysis identifies vulnerable statements, while slicing and data dependency analysis identify the relationship between these statements and program or function inputs, thus reducing the search space.To guide the search towards discovering buffer overflow in this work we define three multi-objective fitness functions and compare them on two open-source systems. These functions account for terms such as the statement coverage, the coverage of vulnerable statements, the distance form buffer boundaries and the coverage of unconstrained nodes of the control flow graph. Concettina Del Grosso, Giuliano Antoniol, Massimiliano Di Penta, Philippe Galinier, Ettore Merlo |
GECCO | 5 |
| 2005 | A language-independent software renovation framework
Massimiliano Di Penta, Markus Neteler, Giuliano Antoniol, Ettore Merlo |
J. Syst. Softw. | 4 |
| 2004 | Linear Complexity Object-Oriented Similarity for Clone Detection and Software Evolution AnalysesabstractWith the widespread adoption of object-oriented technologies, the lack of computationally efficient and scalable approaches is limiting the ability to model and analyze the history of large object-oriented software systems. This paper proposes an approximate representation of object-oriented code characteristics, inspired by pattern recognition centroids for clustering. An interesting application of such a representation is a linear-time complexity algorithm to detect duplicate or nearly duplicated code in object-oriented systems. The algorithm accuracy and time complexity were assessed on 11 releases of a large software system, the Eclipse framework. Ettore Merlo, Giuliano Antoniol, Massimiliano Di Penta, Vincenzo Fabio Rollo |
ICSM | 1 |
| 2003 | Investigating Java Type Analyses for the Receiver-Classes Testing CriterionabstractThis paper investigates the precision of three linear-complexity type analyses for Java software: Class Hierarchy Analysis (CHA), Rapid Type Analysis (RTA) and Variable Type Analysis (VTA). Precision is measured relative to class targets. Class targets results are useful in the context of the receiver-classes criterion, which is an object-oriented testing strategy that aims to exercise every possible class binding of the receiver object reference at each dynamic call site. In this context, using a more precise analysis decreases the number of infeasible bindings to cover, thus it reduces the time spent on conceiving test data sets. This paper also introduces two novel variations to VTA, called the iteration and intersection variants. We present experimental results about the precision of CHA, RTA and VTA on a set of 17 Java programs, corresponding to a total of 600 kLOC of source code. Results show that, on average, RTA suggests 13% less bindings than CHA, standard VTA suggests 23% less bindings than CHAt and VTA with the two variations together suggests 32% less bindings than CHA. Pierre-Luc Brunelle, Ettore Merlo, Giuliano Antoniol |
ISSRE | 2 |
| 2003 | Feed-forward and recurrent neural networks for source code informal information analysisabstractAbstract Design recovery, which is a part of the reverse engineering process of source code, must supply programmers with all the information they need to fully understand a program or a system. In this paper, a connectionist method that can be used for design recovery in conjunction with more traditional approaches is proposed for analyzing the informal information (comments and mnemonics) in programs. An approach based on artificial neural networks (ANNs) was chosen because of its property of being robust (capable of tolerating noisy inputs), because of its associative memory ability (capable of retrieving a concept given only the context of the input word that originally fired the concept), and because of its generalization power (ability to learn conceptually relevant micro‐features of the domain). The proposed approach uses a combination of top down domain analysis (i.e., the creation of a concept hierarchy by a domain expert, to be used in the construction of the training set) and a bottom up approach (i.e., the analysis of the informal information using ANNs). A preprocessing system that extracts the relevant comments and identifier names and transforms them into an input for the ANNs has been developed. Feed‐forward neural networks (FNNs) and recurrent neural networks (RNNs) were tried. RNN architectures are capable of learning sequences and are able to make use of the word ordering of the sentence. The networks were trained on part of the source code of an existing system and tested on a different portion of the system code. Test results, consisting of coverage and evaluation figures, are presented. They show a remarkably higher accuracy when ANNs, in general, are used as opposed to simple lexical methods. RNNs, in particular, also show higher coverage and accuracy than FNNs. Copyright © 2003 John Wiley & Sons, Ltd. Ettore Merlo, Ian McAdam, Renato De Mori |
J. Softw. Maintenance Res. Pract. | 1 |
| 2002 | Investigating Large Software System Evolution: The Linux KernelabstractLarge multi-platform, multi-million lines of codes software systems evolve to cope with new platform or to meet user ever changing needs. While there has been several studies focused on the similarity of code fragments or modules, few studies addressed the need to monitor the overall system evolution. Meanwhile, the decision to evolve or to re-factor a large software system needs to be supported by high level information, representing the system overall picture, abstracting from unnecessary details. This paper proposes to extend the concept of similarity of code fragments to quantify similarities at the release/system level. Similarities are captured by four software metrics representative of the commonalities and differences within and among software artifacts. To show the feasibility of characterizing large software system with the new metrics, 365 releases of the Linux kernel were analyzed. The metrics, the experimental results as well as the lessons learned are presented in the paper. Ettore Merlo, Michel R. Dagenais, P. Bachand, J. S. Sormani, Sara Gradara, Giuliano Antoniol |
COMPSAC | 1 |
| 2002 | Analyzing cloning evolution in the Linux kernel
Giuliano Antoniol, Umberto Villano, Ettore Merlo, Massimiliano Di Penta |
Inf. Softw. Technol. | 3 |
| 2002 | Recovering Traceability Links between Code and DocumentationabstractSoftware system documentation is almost always expressed informally in natural language and free text. Examples include requirement specifications, design documents, manual pages, system development journals, error logs, and related maintenance reports. We propose a method based on information retrieval to recover traceability links between source code and free text documents. A premise of our work is that programmers use meaningful names for program items, such as functions, variables, types, classes, and methods. We believe that the application-domain knowledge that programmers process when writing the code is often captured by the mnemonics for identifiers; therefore, the analysis of these mnemonics can help to associate high-level concepts with program concepts and vice-versa. We apply both a probabilistic and a vector space information retrieval model in two case studies to trace C++ source code onto manual pages and Java code to functional requirements. We compare the results of applying the two models, discuss the benefits and limitations, and describe directions for improvements. Giuliano Antoniol, Gerardo Canfora, Gerardo Casazza, Andrea De Lucia, Ettore Merlo |
IEEE Trans. Software Eng. | 5 |
| 2001 | Modeling Clones Evolution through Time SeriesabstractThe actual effort to evolve and maintain a software system is likely to vary depending on the amount of clones (i.e., duplicated or slightly different code fragments) present in the system. This paper presents a method for monitoring and predicting clones evolution across subsequent versions of a software system. Clones are firstly identified using a metric-based approach, then they are modeled in terms of time series identifying a predictive model. The proposed method has been validated with an experimental activity performed on 27 subsequent versions of mSQL, a medium-size software system written in C. The time span period of the analyzed mSQL releases covers four years, from May 1995 (mSQL 1.0.6) to May 1999 (mSQL 2. 0. 10). For any given software release, the identified models was able to predict the clone percentage of the subsequent release with an average error below 4 %. A higher prediction error was observed only in correspondence of major system redesign. Giuliano Antoniol, Gerardo Casazza, Massimiliano Di Penta, Ettore Merlo |
ICSM | 4 |
| 2001 | Flow Analysis to Detect Blocked StatementsabstractIn the context of software quality assessment, the paper proposes two new kinds of data which can be extracted from source code. The first, definitely blocked statements, can never be executed because preceding code prevents the execution of the program. The other data, called possibly blocked statements, may be blocked by blocking code. The paper presents original flow equations to compute definitely and possibly blocked statements in source code. The experimental context is described and results are shown and discussed. Suggestions for further research are also presented. Bruno Malenfant, Giuliano Antoniol, Ettore Merlo, Michel R. Dagenais |
ICSM | 3 |
| 2000 | C/C++ Conditional Compilation Analysis using Symbolic ExecutionabstractConditional compilation is one of the most powerful parts of a C/C++ environment available for building software for different platforms with different feature sets. Although conditional compilation is powerful, it can be difficult to understand and is error-prone. In large software systems, file inclusion, conditional compilation and macro substitution are closely related and are often largely interleaved. Without adequate tools, understanding complex header files is a tedious task. This practice may even be complicated as the hierarchies of header files grow with projects. This paper presents our experiences of studying conditional compilation based on the symbolic execution of preprocessing directives. Our two concrete goals are: for any given preprocessor directive or C/C++ source code line, finding the simplest sufficient condition to reach/compile it, and finding the full condition to reach/compile that code line. Two different strategies were used to achieve these two goals. A series of experiments conducted on the Linux kernel are presented. Ettore Merlo, Michel R. Dagenais, Bruno Laguë |
ICSM | 2 |
| 1999 | Automatic Unit Test Data Generation Using Mixed-Integer Linear Programming and Execution TreesabstractThis paper presents an approach to automatic unit test data generation for branch coverage using mixed-integer linear programming, execution trees, and symbolic execution. This approach can be useful to both general testing and regression testing after software maintenance and reengineering activities. Several strategies, including original algorithms, to move towards practical test data generation have been investigated in this paper. Methods include: the analysis of minimum path-length partial execution trees for unconstrained arcs, thus increasing the generation performance and reducing the difficulties originated by infeasible paths the reduction of the difficulties originated by nonlinear path conditions by considering alternative linear paths the reduction of the number of test cases, which are needed to achieve the desired coverage, based on the concept of unconstrained arcs in a control flow graph the extension of symbolic execution to deal with dynamic memory allocation and deallocation, pointers and pointers to functions system. Preliminary results are encouraging and show that a high percentage of the program branches can be covered by the test data automatically produced. The approach is flexible to branch selection criteria coming from general testing as well as regression testing. Sébastien Lapierre, Ettore Merlo, Gilles Savard, Giuliano Antoniol, Roberto Fiutem, Paolo Tonella |
ICSM | 2 |
| 1999 | Points-to analysis for program understanding
Roberto Fiutem, Paolo Tonella, Giuliano Antoniol, Ettore Merlo |
J. Syst. Softw. | 4 |
| 1999 | ART: an architectural reverse engineering environmentabstractWhen programmers perform maintenance tasks, program understanding is often required. One of the first activities in understanding a software system is identifying its subsystems and their relations, i.e., its software architecture. Since a large part of the effort is spent in creating a mental model of the system under study, tools can help maintainers in managing the evolution of legacy systems by showing them architectural information. This paper describes an environment for the architectural recovery of software systems called the architectural recovery tool (ART). The environment is based on a hierarchical architectural model that drives the application of a set of recognizers, each producing a different architectural view of a system or of some of its parts. Recognizers embody knowledge about architectural clichés and use flow analysis techniques to make their output more accurate. To test the accuracy and effectiveness of the ART, a suite of public domain applications containing interesting architectural organizations was selected as a benchmark. Results are presented by showing ART performance in terms of precision and recall of the architectural concept retrieval process. The results obtained show that cliché-based architectural recovery is feasible and the recovered information can be valuable support in reengineering and maintenance activities. Copyright © 1999 John Wiley & Sons, Ltd. Roberto Fiutem, Giuliano Antoniol, Paolo Tonella, Ettore Merlo |
J. Softw. Maintenance Res. Pract. | 4 |
| 1999 | Variable-precision reaching definitions analysisabstractAscertaining the reaching definitions from the source code can give views of the linkages in that source code. These views can aid source code analyses, such as impact analysis and program slicing, and can assist in the reverse engineering and re-engineering of large legacy systems. Maintainers like to do such activities interactively and value fast responses from program analysis tools. Therefore the control of the trade-off between accuracy and efficiency should be given to the maintainer. Since some real world programs, especially in languages like C, make much use of pointers, and efficient points-to analysis should be integrated within the computation of the data dependencies during the process of ascertaining the reaching definitions. This paper proposes three different approaches to the analysis of the reaching definitions based on different levels of precision, reflecting differences in their sensitivity to the calling context and the control flow. The least precise approach produces an overestimate by an average of 41% of data dependencies compared to the approach with the highest degree of precision. The result for the least precise approach is conservative because all detectable data dependencies are included, and is far faster than the more precise approaches. Runs on a test suite show an almost 2000 to 1 reduction in execution time by the least precise approach compared with the most precise approach. The intermediate approach is more than 30 times faster than the most precise approach, and much more precise than the least precise one (an average of 2% extra dependencies compared to the most precise approach). Therefore, while on medium size systems the intermediate approach could be a good compromise, on large systems the least precise approach becomes extremely valuable, being the only one feasible. Copyright © 1999 John Wiley & Sons, Ltd. Paolo Tonella, Giuliano Antoniol, Roberto Fiutem, Ettore Merlo |
J. Softw. Maintenance Res. Pract. | 4 |
| 1998 | Process Assurance Audits: Lessons LearnedabstractDuring 1997, a large Information System (IS) Division of a Canadian phone company implemented formal process assurance in its Quality Assurance group. The status report presents a new perspective on the measurement of process assurance and the lessons learned after one year of assessing the individual conformance of software development projects to the corporate software development process (CSDP) of the organization. The status report presents the assurance process overview, goals. Benefits and scope, as well as the 1997 results overview, followed by the lessons learned for the 1998 audit program. Alain April, Alain Abran, Ettore Merlo |
ICSE | 3 |
| 1997 | Flow Insensitive C++ Pointers and Polymorphism Analysis and its Application to SlicingabstractLarge software systems are difficult to understand and maintain.Code analysis tools can provide programmers with different views of the software which may help their understanding activity.To be applicable to real programs written in modern programming languages, these tools need to efficiently handle pointers.In the case of C++ analysis, object oriented peculiarities (like, e.g., polymorphism) have to be accounted for as well.We propose a flow insensitive, context insensitive points-to analysis capable of dealing with the features of the object oriented code.It is extremely promising because of the positive trade-off between complexity and accuracy.The integration of the points-to results with other analyses, such as reaching definitions and slicing, is also discussed in the context of our program understanding environment. Paolo Tonella, Giuliano Antoniol, Roberto Fiutem, Ettore Merlo |
ICSE | 4 |
| 1997 | Program Understanding and Maintenance with the CANTO EnvironmentabstractDuring maintenance activities, the availability of integrated conceptual views that present software at different levels of abstraction, from software architecture to control and data flow relations at code level, is fundamental to understand and modify legacy systems. This paper presents CANTO (Code and Architecture Analysis Tool), a comprehensive program understanding and maintenance environment which integrates fine grained information with architectural views extracted from source code, giving the user control of what is being computed by analyses. The capabilities and usefulness of CANTO are illustrated with reference to a real understanding and maintenance task Giuliano Antoniol, Roberto Fiutem, G. Lutteri, Paolo Tonella, S. Zanfei, Ettore Merlo |
ICSM | 6 |
| 1997 | Assessing the Benefits of Incorporating Function Clone Detection in a Development ProcessabstractThe aim of the experiment presented in this paper is to present an insight into the evaluation of the potential benefits of introducing a function clone detection technology in an industrial software development process. To take advantage of function clone detection, two modifications to the software development process are presented. Our experiment consists of evaluating the impact that these proposed changes would have had on a specific software system if they had been applied over a 3 year period (involving 10000 person-months), where 6 subsequent versions of the software under study were released. The software under study is a large telecommunication system. In total 89 million lines of code have been analyzed. A first result showed that, against our expectations, a significant number of clones are being removed from the system over time. However, this effort is insufficient to prevent the growth of the overall number of clones in the system. In this context the first process change would have added value. We have also found that the second process change would have provided programmers with a significant number of opportunities for correcting problems before customers experienced them. This result shows a potential for improving the software system quality and customer satisfaction Bruno Laguë, Daniel Proulx, Jean Mayrand, Ettore Merlo, John P. Hudepohl |
ICSM | 4 |
| 1996 | A Cliche'-Based Environment to Support Architectural Reverse EngineeringabstractWhen programmers perform maintenance tasks, program understanding is required. One of the first activities in understanding a software system is identifying its subsystems and their relations, i.e. its software architecture. Since a large part of the effort is spent in creating a mental model of the system under study, tools can help maintainers in managing the evolution of legacy systems, by showing them architectural information. An environment for the architectural analysis of software systems is described. The environment is based on a hierarchical architectural model that drives the application of a set of recognizers, each producing a different architectural view of the system or of some of its parts. Recognizers embody knowledge about architectural cliches and use flow analysis techniques to make their output more accurate. Roberto Fiutem, Paolo Tonella, Giuliano Antoniol, Ettore Merlo |
ICSM | 4 |
| 1996 | Experiment on the Automatic Detection of Function Clones in a Software System Using MetricsabstractThis paper presents a technique to automatically identify duplicate and near duplicate functions in a large software system. The identification technique is based on metrics extracted from the source code using the tool Datrix/sup TM/. This clone identification technique uses 21 function metrics grouped into four points of comparison. Each point of comparison is used to compare functions and determine their cloning level. An ordinal scale of eight cloning levels is defined. The levels range from an exact copy to distinct functions. The metrics, the thresholds and the process used are fully described. The results of applying the clone detection technique to two telecommunication monitoring systems totaling one million lines of source code are provided as examples. The information provided by this study is useful in monitoring the maintainability of large software systems. Jean Mayrand, Claude Leblanc, Ettore Merlo |
ICSM | 3 |
| 1996 | Pattern Matching for Clone and Concept Detection
Kostas Kontogiannis, Renato De Mori, Ettore Merlo, Michael Galler, Morris Bernstein |
Autom. Softw. Eng. | 3 |
| 1995 | Application and user interface migration from BASIC to Visual C++abstractAn approach to reengineer BASIC PC legacy code into modern graphical systems is proposed. BASIC has historically been one of the first languages available on PCs. Based on it, small or medium size companies have developed systems that represent valuable company assets to be preserved. Our goal is the automatic migration from the BASIC character oriented user interface to a graphical environment which includes a GUI builder, and compiles event driven C/C++ code. For this purpose a conceptual representation in terms of abstract graphical objects and call-backs has been inferred from the original code, and a translator from BASIC to C has been developed. Moreover the GUI builder internal representation has been generated, so that the user interface can be interactively fine-tuned by the programmer. We present and discuss BASIC peculiarities, with preliminary results on code translation. For the explanation of our approach to user interface migration an example is used throughout the text. Giuliano Antoniol, Roberto Fiutem, Ettore Merlo, Paolo Tonella |
ICSM | 3 |
| 1995 | Multi-Valued Constant Propagation Analysis for User Interface ReengineeringabstractThe definition and use of multi-valued constant propagation analysis (MVCP), which is an extension of simple constant propagation analysis, is presented in this paper in the context of a user interface reengineering process. A brief description of the adopted COBOL/CICS user interface reengineering model, which makes use of an Abstract User Interface Description Language (AUIDL) to represent user interface structures and behavior, is also given. The experimental context is described and results are shown and discussed. Suggestions for further directions of research and investigation are also presented. Ettore Merlo, Jean-Francois Girard, Laurie J. Hendren, Renato De Mori |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 1994 | Localization of Design Concepts in Legacy SystemsabstractComplete automation of design recovery of large systems is a desirable but impractical goal due to complexity and size issues, so current research efforts focus on redocumentation and partial design recovery. Pattern matching lies at the center of any design recovery system. In the context of a larger project to develop an integrated reverse engineering environment, we are developing a framework for performing clone detection, code localization, and plan recognition. This paper discusses a plan localization and selection strategy based on a dynamic programming function that records the matching process and identifies parts of the plan and code fragment that are most "similar". Program features used for matching are currently based on data flow, control flow, and structural properties. The matching model uses a transition network and allows for the detection of insertions and deletions, and it is targeted for legacy C-based systems.> Kostas Kontogiannis, Renato De Mori, Morris Bernstein, Ettore Merlo |
ICSM | 4 |
| 1994 | Inference of Graphical AUIDL Specifications for the Reverse Engineering of User InterfacesabstractA method is presented for inferring abstract user interface specifications from structural and behavioural information which is extracted from the source code. Original algorithms for inference and transformation of user interface specifications are presented together with examples of user interface reverse engineering taken from a COBOL/CICS environment. Current limits and directions for further research are also discussed.> Ettore Merlo, Pierre-Yves Gagné, Alain Thiboutôt |
ICSM | 1 |
| 1993 | Multi-Valued Constant Propagation for the Reengineering of User InterfacesabstractAn extension of simple constant propagation analysis is presented in the context of the ongoing Macroscope project on the reengineering of user interfaces. Multi-valued constant propagation analysis (MVCP) is needed to extract user interface behavioral specifications from source code. Structural and behavioral specifications are used to generate new user interface code that will be integrated into the original system. The flow analysis aspects involved in the MVCP approach are described together with an overall view of the user interface reengineering project. The motivations and advantages of such an analysis method are presented.> Ettore Merlo, Jean-Francois Girard, Laurie J. Hendren, Renato De Mori |
ICSM | 1 |
| 1993 | Source Code Informal Information Analysis Using Connectionist Models
Ettore Merlo, Ian McAdam, Renato De Mori |
IJCAI | 1 |
| 1988 | A network of actions for automatic speech recognition
Renato De Mori, Régis Cardin, Ettore Merlo, Mathew J. Palakal, Jean Rouat |
Speech Commun. | 3 |
| 1987 | Use of Procedural Knowledge for Automatic Speech Recognition
Renato De Mori, Ettore Merlo, Mathew J. Palakal, Jean Rouat |
IJCAI | 2 |
| 1986 | A continuous parameter and frequency domain based Markov modelabstractIn most of the existing Automatic Speech Recognition Systems which make use of Markov Models, the output of the Markov Chain are strings whose symbols belong to a finite alphabet and are generated sequentially over the time domain. We propose a Markov Model System in which symbols are substituted by Spectral Lines which are sequentially generated over the frequency domain. Each spectral line is represented by Continuous Distribution of Parameters. Switching from time-domain to frequency domain drastically reduces the number of states on the Markov chain and the use of continuous parameters eliminates quantization error completely. An application will be presented with experimental results in a multi-speaker environment. Ettore Merlo, Renato De Mori, Mathew J. Palakal, Guy Mercier |
ICASSP | 1 |