Anna Corazza

dblp:02/2002 · DBLP profile ↗
← Back
37ranked-venue papers
20as first author
5since 2021 · last 2024
0000-0002-9156-5079ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 19 · 10 first-author · 5 since 2021Artificial intelligence and machine learning · 15 · 9 first-authorGraphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Regression test prioritization leveraging source code similarity with tree kernels
abstract
Abstract Regression test prioritization (RTP) is an active research field, aiming at re‐ordering the tests in a test suite to maximize the rate at which faults are detected. A number of RTP strategies have been proposed, leveraging different factors to reorder tests. Some techniques include an analysis of changed source code, to assign higher priority to tests stressing modified parts of the codebase. Still, most of these change‐based solutions focus on simple text‐level comparisons among versions. We believe that measuring source code changes in a more refined way, capable of discriminating between mere textual changes (e.g., renaming of a local variable) and more structural changes (e.g., changes in the control flow), could lead to significant benefits in RTP, under the assumption that major structural changes are also more likely to introduce faults. To this end, we propose two novel RTP techniques that leverage tree kernels (TK), a class of similarity functions largely used in Natural Language Processing on tree‐structured data. In particular, we apply TKs to abstract syntax trees of source code, to more precisely quantify the extent of structural changes in the source code, and prioritize tests accordingly. We assessed the effectiveness of the proposals by conducting an empirical study on five real‐world Java projects, also used in a number of RTP‐related papers. We automatically generated, for each considered pair of software versions (i.e., old version, new version) in the evolution of the involved projects, 100 variations with artificially injected faults, leading to over 5k different software evolution scenarios overall. We compared the proposed prioritization approaches against well‐known prioritization techniques, evaluating both their effectiveness and their execution times. Our findings show that leveraging more refined code change analysis techniques to quantify the extent of changes in source code can lead to relevant improvements in prioritization effectiveness, while typically introducing negligible overheads due to their execution.
Francesco Altiero, Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi L. L. Starace
J. Softw. Evol. Process.2
2023 AI-based Fault-proneness Metrics for Source Code Changes
Francesco Altiero, Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi L. L. Starace
IWSM-Mensura2
2022 Change-Aware Regression Test Prioritization using Genetic Algorithms
abstract
Regression testing is a practice aimed at providing confidence that, within software maintenance, the changes in the code base have introduced no faults in previously validated functionalities. With the software industry shifting towards iterative and incremental development with shorter release cycles, the straightforward approach of re-executing the entire test suite on each new version of the software is often unfeasible due to time and resource constraints. In such scenarios, Test Case Prioritization (TCP) strategies aim at providing an effective ordering of the test suite, so that the tests that are more likely to expose faults are executed earlier and fault detection is maximised even when test execution needs to be abruptly terminated due to external constraints. In this work, we propose Genetic-Diff, a TCP strategy based on a genetic algorithm featuring a specifically-designed crossover operator and a novel objective function that combines code coverage metrics with an analysis of changes in the code base. We empirically evaluate the proposed algorithm on several releases of three heterogeneous real-world, open source Java projects, in which we artificially injected faults, and compare the results with other state-of-the-art TCP techniques using fault-detection rate metrics. Findings show that the proposed technique performs generally better than the baselines, especially when there is a limited amount of code changes, which is a common scenario in modern development practices.
Francesco Altiero, Giovanni Colella, Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi L. L. Starace
SEAA3
2022 ReCover: a Curated Dataset for Regression Testing Research
abstract
It is recognized in the literature that finding representative data to conduct regression testing research is non-trivial. In our experience within this field, existing datasets are often affected by issues that limit their applicability. Indeed, these datasets often lack fine-grained coverage information, reference software repositories that are not available anymore, or do not allow researchers to readily build and run the software projects, e.g., to obtain additional information. As a step towards better replicability and data-availability in regression testing research, we introduce ReCover, a dataset of 114 pairs of subsequent versions from 28 open source Java projects from GitHub. In particular, ReCover is intended as a consolidation and enrichment of recent dedicated regression testing datasets proposed in the literature, to overcome some of the above described issues, and to make them ready to use with a broader number of regression testing techniques. To this end, we developed a custom mining tool, that we make available as well, to automatically process two recent, massive regression testing datasets, retaining pairs of software versions for which we were able to (1) retrieve the full source code; (2) build the software in a general-purpose Java/Maven environment (which we provide as a Docker container for ease of replication); and (3) compute fine-grained test coverage metrics. ReCover can be readily employed in regression testing studies, as it bundles in a single package full, buildable source code and detailed coverage reports for all the projects. We envision that its use could foster regression testing research, improving replicability and long-term data availability.
Francesco Altiero, Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi L. L. Starace
MSR2
2021 Web Application Testing: Using Tree Kernels to Detect Near-duplicate States in Automated Model Inference
abstract
In the context of End-to-End testing of web applications, automated exploration techniques (a.k.a. crawling) are widely used to infer state-based models of the site under test. These models, in which states represent features of the web application and transitions represent reachability relationships, can be used for several model-based testing tasks, such as test case generation. However, current exploration techniques often lead to models containing many near-duplicate states, i.e., states representing slightly different pages that are in fact instances of the same feature. This has a negative impact on the subsequent model-based testing tasks, adversely affecting, for example, size, running time, and achieved coverage of generated test suites. As a web page can be naturally represented by its tree-structured DOM representation, we propose a novel near-duplicate detection technique to improve the model inference of web applications, based on Tree Kernel (TK) functions. TKs are a class of functions that compute similarity between tree-structured objects, largely investigated and successfully applied in the Natural Language Processing domain. To evaluate the capability of the proposed approach in detecting near-duplicate web pages, we conducted preliminary classification experiments on a freely-available massive dataset of about 100k manually annotated web page pairs. We compared the classification performance of the proposed approach with other state-of-the-art near-duplicate detection techniques. Preliminary results show that our approach performs better than state-of-the-art techniques in the near-duplicate detection classification task. These promising results show that TKs can be applied to near-duplicate detection in the context of web application model inference, and motivate further research in this direction.
Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi L. L. Starace
ESEM1
2020 Inspecting Code Churns to Prioritize Test Cases
Francesco Altiero, Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi L. L. Starace
ICTSS2
2020 Adequate vs. inadequate test suite reduction approaches
Carmen Coviello, Simone Romano 0001, Giuseppe Scanniello, Alessandro Marchetto 0001, Anna Corazza, Giuliano Antoniol
Inf. Softw. Technol.5
2019 Word Embeddings for Comment Coherence
abstract
During the evolution of software, it could happen that the information in the comments and in the associated source code are not aligned, so hampering the execution of software evolution and maintenance tasks. This kind of misalignment is known as lack of coherence and it can happen for several reasons, e.g., programmers modify the intent of source code while executing a maintenance task without updating its comment accordingly. We study the problem of detecting a lack of coherence between comments and source code by exploiting Word Embeddings (WEs). We present four models based on WE and tested these models using six different WE variants through an experiment conducted on a publicly available dataset. Results are compared against a baseline. The most important outcome is: the considered models and WE variants are more efficient in terms of execution time while maintaining performance very close to the baseline. The explanation for such an improvement is that WEs are able to concentrate the important information in a more compact input representation.
Alfonso Cimasa, Anna Corazza, Carmen Coviello, Giuseppe Scanniello
SEAA2
2018 Clustering support for inadequate test suite reduction
abstract
Regression testing is an important activity that can be expensive (e.g., for large test suites). Test suite reduction approaches speed up regression testing by removing redundant test cases. These approaches can be classified as adequate or inadequate. Adequate approaches reduce test suites so that they completely preserve the test requirements (e.g., code coverage) of the original test suites. Inadequate approaches produce reduced test suites that only partially preserve the test requirements. An inadequate approach is appealing when it leads to a greater reduction in test suite size at the expense of a small loss in fault-detection capability. We investigate a clustering-based approach for inadequate test suite reduction and compare it with well-known adequate approaches. Our investigation is founded on a public dataset and allows an exploration of trade-offs in test suite reduction. Results help a more informed decision, using guidelines defined in this research, to balance size, coverage, and fault-detection loss of reduced test suites when using clustering.
Carmen Coviello, Simone Romano 0001, Giuseppe Scanniello, Alessandro Marchetto 0001, Giuliano Antoniol, Anna Corazza
SANER6
2018 Coherence of comments and method implementations: a dataset and an empirical investigation
Anna Corazza, Valerio Maggio, Giuseppe Scanniello
Softw. Qual. J.1
2017 Integrating a Priori Probabilistic Knowledge into Classification for Image Description
abstract
This paper discusses a possible implementation of the integration of knowledge from a probabilistic ontology in the automatic description of images. This combination not only provides the relations existing between the different segments, but also improve the classification accuracy, as the context often gives cues suggesting the correct class of the segment.
Andrea Apicella 0001, Anna Corazza, Francesco Isgrò, Giuseppe Vettigli
WETICE2
2016 Weighing lexical information for software clustering in the context of architecture recovery
Anna Corazza, Sergio Di Martino, Valerio Maggio, Giuseppe Scanniello
Empir. Softw. Eng.1
2015 From Function Points to COSMIC - A Transfer Learning Approach for Effort Estimation
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro
PROFES1
2013 Using tabu search to configure support vector regression for effort estimation
abstract
Recent studies have reported that Support Vector Regression (SVR) has the potential as a technique for software development effort estimation. However, its prediction accuracy is heavily influenced by the setting of parameters that needs to be done when employing it. No general guidelines are available to select these parameters, whose choice also depends on the characteristics of the dataset being used. This motivated the work described in (Corazza et al. 2010 ), extended herein. In order to automatically select suitable SVR parameters we proposed an approach based on the use of the meta-heuristics Tabu Search (TS). We designed TS to search for the parameters of both the support vector algorithm and of the employed kernel function, namely RBF. We empirically assessed the effectiveness of the approach using different types of datasets (single and cross-company datasets, Web and not Web projects) from the PROMISE repository and from the Tukutuku database. A total of 21 datasets were employed to perform a 10-fold or a leave-one-out cross-validation, depending on the size of the dataset. Several benchmarks were taken into account to assess both the effectiveness of TS to set SVR parameters and the prediction accuracy of the proposed approach with respect to widely used effort estimation techniques. The use of TS allowed us to automatically obtain suitable parameters’ choices required to run SVR. Moreover, the combination of TS and SVR significantly outperformed all the other techniques. The proposed approach represents a suitable technique for software development effort estimation.
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Federica Sarro, Emilia Mendes
Empir. Softw. Eng.1
2013 Supporting concept location through identifier parsing and ontology extraction
Surafel Lemma Abebe, Anita Alicante, Anna Corazza, Paolo Tonella
J. Syst. Softw.3
2012 LINSEN: An efficient approach to split identifiers and expand abbreviations
abstract
Information Retrieval (IR) techniques are being exploited by an increasing number of tools supporting Software Maintenance activities. Indeed the lexical information embedded in the source code can be valuable for tasks such as concept location, clustering or recovery of traceability links. The application of such IR-based techniques relies on the consistency of the lexicon available in the different artifacts, and their effectiveness can worsen if programmers introduce abbreviations (e.g: rect) and/or do not strictly follow naming conventions such as Camel Case (e.g: UTFtoASCII). In this paper we propose an approach to automatically split identifiers in their composing words, and expand abbreviations. The solution is based on a graph model and performs in linear time with respect to the size of the dictionary, taking advantage of an approximate string matching technique. The proposed technique exploits a number of different dictionaries, referring to increasingly broader contexts, in order to achieve a disambiguation strategy based on the knowledge gathered from the most appropriate domain. The approach has been compared to other splitting and expansion techniques, using freely available oracles for the identifiers extracted from 24 C/C++ and Java open source systems. Results show an improvement in both splitting and expanding performance, in addition to a strong enhancement in the computational efficiency.
Anna Corazza, Sergio Di Martino, Valerio Maggio
ICSM1
2012 A treebank-based study on the influence of Italian word order on parsing performance
Anita Alicante, Cristina Bosco, Anna Corazza, Alberto Lavelli
LREC3
2011 Investigating the use of Support Vector Regression for web effort estimation
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes
Empir. Softw. Eng.1
2010 A Tree Kernel based approach for clone detection
abstract
Reusing software by copying and pasting is a common practice in software development. This phenomenon is widely known as code cloning. Problems with clones are mainly due to the need of managing each duplication, thus increasing the effort to maintain software systems. Clone detection approaches generally take into account either the syntactic structure (e.g., Abstract Syntax Tree) or lexical elements (e.g., the signature of a function). In this paper we propose an approach to detect code clones, based on syntactic information enriched by lexical elements. To this end, we have defined a Tree Kernel function to compare Abstract Syntax Trees. A preliminary investigation has been also conducted to assess the validity of the proposed approach.
Anna Corazza, Sergio Di Martino, Valerio Maggio, Giuseppe Scanniello
ICSM1
2009 Applying support vector regression for web effort estimation using a cross-company dataset
abstract
Support vector regression (SVR) is a new generation of machine learning algorithms, suitable for predictive data modeling problems. The objective of this paper is to investigate the effectiveness of SVR for Web effort estimation, in particular when dealing with a cross-company dataset. To gain a deeper insight on the method, we carried out an empirical study using four kernels for SVR, namely linear, polynomial, Gaussian, and sigmoid. Moreover, we used two variables' preprocessing strategies (normalization and logarithmic), and two different dependent variables (effort and inverse effort). As a result, SVR was applied using six different configurations for each kernel. As for the dataset, we employed the Tukutuku database, which is widely adopted in Web effort estimation studies. A hold-out approach was adopted to evaluate the prediction accuracy for all the configurations, using two training sets, each containing data on 130 projects randomly selected, and two test sets, each containing the remaining 65 projects. As benchmark, SVR-based predictions were also compared to predictions obtained using manual stepwise regression, case-based reasoning, and Bayesian networks. Our results suggest that SVR performed well, since on the first hold-out, the linear kernel with a logarithmic transformation of variables provided significantly superior prediction accuracy than all the other techniques, while for the second hold-out, the Gaussian kernel achieved significantly superior predictions than all other techniques, except for manual stepwise regression.
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes
ESEM1
2009 Using Support Vector Regression for Web Development Effort Estimation
Anna Corazza, Sergio Di Martino, Filomena Ferrucci, Carmine Gravino, Emilia Mendes
IWSM/Mensura1
2008 Comparing Italian parsers on a common Treebank: the EVALITA experience
Cristina Bosco, Alessandro Mazzei, Vincenzo Lombardo, Giuseppe Attardi, Anna Corazza, Alberto Lavelli, Leonardo Lesmo, Giorgio Satta, Maria Simi
LREC5
2007 Probabilistic Context-Free Grammars Estimated from Infinite Distributions
abstract
In this paper, we consider probabilistic context-free grammars, a class of generative devices that has been successfully exploited in several applications of syntactic pattern matching, especially in statistical natural language parsing. We investigate the problem of training probabilistic context-free grammars on the basis of distributions defined over an infinite set of trees or an infinite set of sentences by minimizing the cross-entropy. This problem has applications in cases of context-free approximation of distributions generated by more expressive statistical models. We show several interesting theoretical properties of probabilistic context-free grammars that are estimated in this way, including the previously unknown equivalence between the grammar cross-entropy with the input distribution and the so-called derivational entropy of the grammar itself. We discuss important consequences of these results involving the standard application of the maximum-likelihood estimator on finite tree and sentence samples, as well as other finite-state models such as Hidden Markov Models and probabilistic finite automata.
Anna Corazza, Giorgio Satta
IEEE Trans. Pattern Anal. Mach. Intell.1
2006 Cross-Entropy and Estimation of Probabilistic Context-Free Grammars
Anna Corazza, Giorgio Satta
HLT-NAACL1
2003 Learning rule ranking by dynamic construction of context-free grammars using AND/OR graphs
Anna Corazza, Louis ten Bosch
INTERSPEECH1
2002 Integration of two stochastic context-free grammars
abstract
Some problems in speech and natural language processing involve combining two information sources each modeled by a stochastic context-free grammar. Such cases include parsing the output of a speech recognizer by using a contextfree language model, finding the best solution among all the possible ones in language generation, preserving ambiguity in machine translation. In these cases usually at least one of the two grammars is non-recursive. In order to find the best solution while taking into account both grammars, the two probabilities must be integrated. One of the most important advantages of using a non-recursive context-free model is its compactness. Therefore, it is important to exploit this property when searching for the solution. In this paper, an algorithm aiming to this goal is presented, based on a recent work [1] in which the non probabilistic case is considered. 1.
Anna Corazza
INTERSPEECH1
1999 An inter-domain portable approach to interchange format construction
abstract
We present a novel efficient algorithm for time-scale \nmodification (TSM) of speech which gives output \nquality equal to that of a conventional TSM algorithm, \nbut having computational load an order of magnitude \nless. The algorithm presented uses a fixed length \nrectangular stepping window and a simple peak \nalignment criterion to track the local natural scaling \nfactor and adapt the window step size. The desired \nTSM factor is realised by the appropriate number of \napplications of the constantly varying local natural \nscaling factor. The local natural scaling factor estimate \nis updated at sub-pitch period intervals giving accurate \npitch tracking and high quality in the output scaled \nsignal.
Anna Corazza
EUROSPEECH1
1999 Semantic boundaries in multiple languages
abstract
This paper presents the results obtained for the task of detecting Semantic Boundaries (SBs) in spoken language using two different methods on the same data set. Hence we first introduce the two approaches developed by ITC-Irst in Trento (Italy) and the LME of the University Erlangen (Germany) and discuss the individually obtained results. The basis for the decision upon SBs in both cases are textual and prosodic features. The LME has already worked for several years on the computation and application of prosodic features in automatic speech processing within the Verbmobil project. The approaches developed in that project were adapted to work on the data collected at IRST in the Italian language. Finally we compare the results we obtain with the German SB detection against the Italian result with regard to precision and recall. 1. INTRODUCTION For robust spoken language processing it is not always necessary to analyse a user's utterance completely as one coherent segment. Often i...
Volker Warnke, Heinrich Niemann, Mauro Cettolo, Anna Corazza, Daniele Falavigna, Gianni Lazzari
EUROSPEECH5
1998 Language portability of a speech understanding system
Mauro Cettolo, Anna Corazza, Renato De Mori
Comput. Speech Lang.2
1997 Multilingual person to person communication at IRST
abstract
This paper refers to a machine-mediated person-to-person multilingual communication system. Stress is put on robustness, that is the ability of the system to preserve communication even in presence of the variability and errors typical of spoken language systems. The statistical approach is adopted not only at the acoustic level, but also for the linguistic processing. Therefore, while an overview of the global architecture is briefly introduced, the focus is put on the acoustic recognizer and the understanding module. Experimental evaluations complete the presentation.
Bianca Angelini, Mauro Cettolo, Anna Corazza, Daniele Falavigna, Gianni Lazzari
ICASSP3
1997 Automatic detection of semantic boundaries
abstract
In spoken language systems, the segmentation of utterances into coherent linguistic/semantic units is very useful, as it makes easier processing after the speech recognition phase. In this paper, a methodology for semantic boundary prediction is presented and tested on a corpus of person-to-person dialogues. The approach is based on binary decision trees and uses text context, including broad classes of silent pauses, filled pauses and human noises. Best results give more than 90% precision, almost 80% recall and about 3% false alarms. 1. INTRODUCTION This work focuses on the automatic segmentation of dialogue turns into homogeneous Semantic Units (SUs) [7]. The approach described below is evaluated in the domain of appointment scheduling, where a system able to deal with this kind of interaction between two persons speaking different languages is being developed [1]. As a working hypothesis, it is assumed that each turn can be represented as a "flat" sequence of concepts, i.e. no nes...
Mauro Cettolo, Anna Corazza
EUROSPEECH2
1996 A mixed approach to speech understanding
Mauro Cettolo, Anna Corazza, Renato De Mori
ICSLP2
1994 Optimal Probabilistic Evaluation Functions for Search Controlled by Stochastic Context-Free Grammars
abstract
The possibility of using stochastic context-free grammars (SCFG's) in language modeling (LM) has been considered previously. When these grammars are used, search can be directed by evaluation functions based on the probabilities that a SCFG generates a sentence, given only some words in it. Expressions for computing the evaluation function have been proposed by Jelinek and Lafferty (1991) for the recognition of word sequences in the case in which only the prefix of a sequence is known. Corazza et al. (1991) have proposed methods for probability computation in the more general case in which partial word sequences interleaved by gaps are known. This computation is too complex in practice unless the lengths of the gaps are known. This paper proposes a method for computing the probability of the best parse tree that can generate a sentence only part of which (consisting of islands and gaps) is known. This probability is the minimum possible, and thus the most informative, upper-bound that can be used in the evaluation function. The computation of the proposed upper-bound has cubic time complexity even if the lengths of the gaps are unknown. This makes possible the practical use of SCFG for driving interpretations of sentences in natural language processing.>
Anna Corazza, Renato De Mori, Roberto Gretter, Giorgio Satta
IEEE Trans. Pattern Anal. Mach. Intell.1
1993 Language modeling using stochastic context-free grammars
Anna Corazza, Renato De Mori, Roberto Gretter, Giorgio Satta
Speech Commun.1
1992 Computation of Upper-Bounds for Stochastic Context-Free Languages
Anna Corazza, Renato De Mori, Giorgio Satta
AAAI1
1991 Computation of upper-bounds for island-driven stochastic parsers
abstract
Automatic speech understanding is the process of deriving a complete sentence interpretation of an acoustic signal. Stochastic language models can be of considerable help for the solution of this problem. In this paper we present a new method to apply stochastic context-free grammar models to the search of the most likely syntactic interpretation of the signal. The problem is discussed both theoretically and computationally. The analysis is also extended to cases in which the underlying parsing process is carried out in a bidirectional way. Introduction Automatic Speech Understanding (ASU) differs from Automatic Speech Recognition (ASR) because it has to produce a conceptual representation of a spoken message rather than a sequence of recognized words. The structure of a conceptual representation depends on the use it has to be made of it. Examples of actions based on conceptual representations are data-base query, inference, robot planning or replanning. In all these cases...
Anna Corazza, Renato De Mori, Roberto Gretter, Giorgio Satta
EUROSPEECH1
1991 Computation of Probabilities for an Island-Driven Parser
abstract
The authors describe an effort to adapt island-driven parsers to handle stochastic context-free grammars. These grammars could be used as language models (LMs) by a language processor (LP) to computer the probability of a linguistic interpretation. As different islands may compete for growth, it is important to compute the probability that an LM generates a sentence containing islands and gaps between them. Algorithms for computing these probabilities are introduced. The complexity of these algorithms is analyzed both from theoretical and practical points of view. It is shown that the computation of probabilities in the presence of gaps of unknown length requires the impractical solution of a nonlinear system of equations, whereas the computation of probabilities for cases with gaps containing a known number of unknown words has polynomial time complexity and is practically feasible. The use of the results obtained in automatic speech understanding systems is discussed.>
Anna Corazza, Renato De Mori, Roberto Gretter, Giorgio Satta
IEEE Trans. Pattern Anal. Mach. Intell.1