Fayola Peters

dblp:52/7880 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
0since 2021 · last 2019
0000-0001-6150-856XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 5 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
7 papers
Empirical software engineering · 95% Requirements engineering and software design · 5%
Network and information security
3 papers
Privacy and data protection · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Empirical software engineering › mining software repositories
defect prediction
0.942019
Text Filtering and Ranking for Security Bug Report Prediction · IEEE Trans. Software Eng. 2019
LACE2: Better Privacy-Preserving Data Sharing for Cross Project Defect Prediction · ICSE (1) 2015
Balancing Privacy and Utility in Cross-Company Defect Prediction · IEEE Trans. Software Eng. 2013
Empirical software engineering
mining software repositories
0.422015
The Art and Science of Analyzing Software Data; Quantitative Methods · ICSE (2) 2015
Data science for software engineering · ICSE 2013
Empirical software engineering › mining software repositories › defect prediction
cross-project defect prediction
0.422015
LACE2: Better Privacy-Preserving Data Sharing for Cross Project Defect Prediction · ICSE (1) 2015
Privacy and utility for defect prediction: Experiments with MORPH · ICSE 2012
Privacy and data protection › anonymization
data obfuscation
0.212015
LACE2: Better Privacy-Preserving Data Sharing for Cross Project Defect Prediction · ICSE (1) 2015
Privacy and data protection
privacy-preserving data sharing
0.212015
LACE2: Better Privacy-Preserving Data Sharing for Cross Project Defect Prediction · ICSE (1) 2015
Empirical software engineering
data science for software engineering
0.212013
Data science for software engineering · ICSE 2013
Privacy and data protection
anonymization
0.112012
Privacy and utility for defect prediction: Experiments with MORPH · ICSE 2012
Data mining › text mining
text classification
0.112019
Text Filtering and Ranking for Security Bug Report Prediction · IEEE Trans. Software Eng. 2019
Requirements engineering and software design
software process
0.112009
Applications of Simulation and AI Search: Assessing the Relative Merits of Agile vs Traditional Software Development · ASE 2009
Privacy and data protection
privacy-preserving data analysis
0.012013
Balancing Privacy and Utility in Cross-Company Defect Prediction · IEEE Trans. Software Eng. 2013
Algorithms and data structures
search algorithms
0.012009
Applications of Simulation and AI Search: Assessing the Relative Merits of Agile vs Traditional Software Development · ASE 2009

Methods — techniques the papers use, named apart from their topics

text-based prediction · 0.8ranking · 0.8filtering · 0.8multi-party data sharing · 0.4data minimization · 0.4ensemble methods · 0.2data mining · 0.2privatization algorithms · 0.2machine learning · 0.2instance pruning · 0.2data privatization · 0.2data mutation · 0.2clustering · 0.2random forest · 0.1naive bayes · 0.1logistic regression · 0.1simulation · 0.1AI search · 0.1
YearPublicationVenuePosition
2019 Text Filtering and Ranking for Security Bug Report Prediction
abstract
Security bug reports can describe security critical vulnerabilities in software products. Bug tracking systems may contain thousands of bug reports, where relatively few of them are security related. Therefore finding unlabelled security bugs among them can be challenging. To help security engineers identify these reports quickly and accurately, text-based prediction models have been proposed. These can often mislabel security bug reports due to a number of reasons such as class imbalance, where the ratio of non-security to security bug reports is very high. More critically, we have observed that the presence of security related keywords in both security and non-security bug reports can lead to the mislabelling of security bug reports. This paper proposes FARSEC, a framework for filtering and ranking bug reports for reducing the presence of security related keywords. Before building prediction models, our framework identifies and removes non-security bug reports with security related keywords. We demonstrate that FARSEC improves the performance of text-based prediction models for security bug reports in 90 percent of cases. Specifically, we evaluate it with 45,940 bug reports from Chromium and four Apache projects. With our framework, we mitigate the class imbalance issue and reduce the number of mislabelled security bug reports by 38 percent.
Fayola Peters, Thein Than Tun, Yijun Yu 0001, Bashar Nuseibeh
IEEE Trans. Software Eng.1
2015 The Art and Science of Analyzing Software Data; Quantitative Methods
abstract
Using the tools of quantitative data science, software engineers that can predict useful information on new projects based on past projects. This tutorial reflects on the state-of-the-art in quantitative reasoning in this important field. This tutorial discusses the following: (a) when local data is scarce, we show how to adapt data from other organizations to local problems; (b) when working with data of dubious quality, we show how to prune spurious information; (c) when data or models seem too complex, we show how to simplify data mining results; (d) when the world changes, and old models need to be updated, we show how to handle those updates; (e) when the effect is too complex for one model, we show to how reason over ensembles.
Tim Menzies, Leandro L. Minku, Fayola Peters
ICSE (2)3
2015 LACE2: Better Privacy-Preserving Data Sharing for Cross Project Defect Prediction
abstract
Before a community can learn general principles, it must share individual experiences. Data sharing is the fundamental step of cross project defect prediction, i.e. the process of using data from one project to predict for defects in another. Prior work on secure data sharing allowed data owners to share their data on a single-party basis for defect prediction via data minimization and obfuscation. However the studied method did not consider that bigger data required the data owner to share more of their data. In this paper, we extend previous work with LACE2 which reduces the amount of data shared by using multi-party data sharing. Here data owners incrementally add data to a cache passed among them and contribute "interesting" data that are not similar to the current content of the cache. Also, before data owner i passes the cache to data owner j, privacy is preserved by applying obfuscation algorithms to hide project details. The experiments of this paper show that (a) LACE2 is comparatively less expensive than the single-party approach and (b) the multi-party approach of LACE2 yields higher privacy than the prior approach without damaging predictive efficacy (indeed, in some cases, LACE2 leads to better defect predictors).
Fayola Peters, Tim Menzies, Lucas Layman
ICSE (1)1
2013 Learning from Open-Source Projects: An Empirical Study on Defect Prediction
abstract
The fundamental issue in cross project defect prediction is selecting the most appropriate training data for creating quality defect predictors. Another concern is whether historical data of open-source projects can be used to create quality predictors for proprietary projects from a practical point-of-view. Current studies have proposed statistical approaches to finding these training data, however, thus far no apparent effort has been made to study their success on proprietary data. Also these methods apply brute force techniques which are computationally expensive. In this work we introduce a novel data selection procedure which takes into account the similarities between the distribution of the test and potential training data. Additionally we use feature subset selection to increase the similarity between the test and training sets. Our procedure provides a comparable and scalable means of solving the cross project defect prediction problem for creating quality defect predictors. To evaluate our procedure we conducted empirical studies with comparisons to the within company defect prediction and a relevancy filtering method. We found that our proposed method performs relatively better than the filtering method in terms of both computation cost and prediction performance.
Fayola Peters, Tim Menzies
ESEM2
2013 Data science for software engineering
abstract
Target audience: Software practitioners and researchers wanting to understand the state of the art in using data science for software engineering (SE). Content: In the age of big data, data science (the knowledge of deriving meaningful outcomes from data) is an essential skill that should be equipped by software engineers. It can be used to predict useful information on new projects based on completed projects. This tutorial offers core insights about the state-of-the-art in this important field. What participants will learn: Before data science: this tutorial discusses the tasks needed to deploy machine-learning algorithms to organizations (Part 1: Organization Issues). During data science: from discretization to clustering to dichotomization and statistical analysis. And the rest: When local data is scarce, we show how to adapt data from other organizations to local problems. When privacy concerns block access, we show how to privatize data while still being able to mine it. When working with data of dubious quality, we show how to prune spurious information. When data or models seem too complex, we show how to simplify data mining results. When data is too scarce to support intricate models, we show methods for generating predictions. When the world changes, and old models need to be updated, we show how to handle those updates. When the effect is too complex for one model, we show how to reason across ensembles of models. Pre-requisites: This tutorial makes minimal use of maths of advanced algorithms and would be understandable by developers and technical managers.
Tim Menzies, Ekrem Kocaguneli, Fayola Peters, Burak Turhan, Leandro L. Minku
ICSE3
2013 Better cross company defect prediction
abstract
How can we find data for quality prediction? Early in the life cycle, projects may lack the data needed to build such predictors. Prior work assumed that relevant training data was found nearest to the local project. But is this the best approach? This paper introduces the Peters filter which is based on the following conjecture: When local data is scarce, more information exists in other projects. Accordingly, this filter selects training data via the structure of other projects. To assess the performance of the Peters filter, we compare it with two other approaches for quality prediction. Within-company learning and cross-company learning with the Burak filter (the state-of-the-art relevancy filter). This paper finds that: 1) within-company predictors are weak for small data-sets; 2) the Peters filter+cross-company builds better predictors than both within-company and the Burak filter+cross-company; and 3) the Peters filter builds 64% more useful predictors than both within-company and the Burak filter+cross-company approaches. Hence, we recommend the Peters filter for cross-company learning.
Fayola Peters, Tim Menzies, Andrian Marcus
MSR1
2013 Balancing Privacy and Utility in Cross-Company Defect Prediction
abstract
Background: Cross-company defect prediction (CCDP) is a field of study where an organization lacking enough local data can use data from other organizations for building defect predictors. To support CCDP, data must be shared. Such shared data must be privatized, but that privatization could severely damage the utility of the data. Aim: To enable effective defect prediction from shared data while preserving privacy. Method: We explore privatization algorithms that maintain class boundaries in a dataset. CLIFF is an instance pruner that deletes irrelevant examples. MORPH is a data mutator that moves the data a random distance, taking care not to cross class boundaries. CLIFF+MORPH are tested in a CCDP study among 10 defect datasets from the PROMISE data repository. Results: We find: 1) The CLIFFed+MORPHed algorithms provide more privacy than the state-of-the-art privacy algorithms; 2) in terms of utility measured by defect prediction, we find that CLIFF+MORPH performs significantly better. Conclusions: For the OO defect data studied here, data can be privatized and shared without a significant degradation in utility. To the best of our knowledge, this is the first published result where privatization does not compromise defect prediction.
Fayola Peters, Tim Menzies, Hongyu Zhang 0002
IEEE Trans. Software Eng.1
2012 Privacy and utility for defect prediction: Experiments with MORPH
abstract
Ideally, we can learn lessons from software projects across multiple organizations. However, a major impediment to such knowledge sharing are the privacy concerns of software development organizations. This paper aims to provide defect data-set owners with an effective means of privatizing their data prior to release. We explore MORPH which understands how to maintain class boundaries in a data-set. MORPH is a data mutator that moves the data a random distance, taking care not to cross class boundaries. The value of training on this MORPHed data is tested via a 10-way within learning study and a cross learning study using Random Forests, Naive Bayes, and Logistic Regression for ten object-oriented defect datasets from the PROMISE data repository. Measured in terms of exposure of sensitive attributes, the MORPHed data was four times more private than the unMORPHed data. Also, in terms of the f-measures, there was little difference between the MORPHed and unMORPHed data (original data and data privatized by data-swapping) for both the cross and within study. We conclude that at least for the kinds of OO defect data studied in this project, data can be privatized without concerns for inference efficacy.
Fayola Peters, Tim Menzies
ICSE1
2009 Applications of Simulation and AI Search: Assessing the Relative Merits of Agile vs Traditional Software Development
abstract
This paper augments Boehm-Turner's model of agile and plan-based software development augmented with an AI search algorithm. The AI search finds the key factors that predict for the success of agile or traditional plan-based software developments. According to our simulations and AI search algorithm: (1) in no case did agile methods perform worse than plan-based approaches; (2) in some cases, agile performed best. Hence, we recommend that the default development practice for organizations be an agile method. The simplicity of this style of analysis begs the question: why is so much time wasted on evidence-less debates on software process when a simple combination of simulation plus automatic search can mature the dialog much faster?
Bryan Lemon, Aaron Riesbeck, Tim Menzies, Justin Price, Joseph D'Alessandro, Rikard Carlsson, Tomi Prifiti, Fayola Peters, Huihua Lu, Daniel Port
ASE8