Tianpei Xia

dblp:218/5181 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2022
0000-0002-6340-8041ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2022 Methods for Stabilizing Models Across Large Samples of Projects (with case studies on Predicting Defect and Project Health)
abstract
Despite decades of research, Software Engineering (SE) lacks widely accepted models (that offer precise quantitative stable predictions) about what factors most influence software quality. This paper provides a promising result showing such stable models can be generated using a new transfer learning framework called "STABILIZER". Given a tree of recursively clustered projects (using project meta-data), STABILIZER promotes a model upwards if it performs best in the lower clusters (stopping when the promoted model performs worse than the models seen at a lower level).
Suvodeep Majumder, Tianpei Xia, Rahul Krishna, Tim Menzies
MSR2
2022 Dazzle: Using Optimized Generative Adversarial Networks to Address Security Data Class Imbalance Issue
abstract
Background: Machine learning techniques have been widely used and demonstrate promising performance in many software security tasks such as software vulnerability prediction. However, the class ratio within software vulnerability datasets is often highly imbalanced (since the percentage of observed vulnerability is usually very low). Goal: To help security practitioners address software security data class imbalanced issues and further help build better prediction models with resampled datasets. Method: We introduce an approach called Dazzle which is an optimized version of conditional Wasserstein Generative Adversarial Networks with gradient penalty (cWGAN-GP). Dazzle explores the architecture hyperparameters of cWGAN-GP with a novel optimizer called Bayesian Optimization. We use Dazzle to generate minority class samples to resample the original imbalanced training dataset. Results: We evaluate Dazzle with three software security datasets, i.e., Moodle vulnerable files, Ambari bug reports, and JavaScript function code. We show that Dazzle is practical to use and demonstrates promising improvement over existing state-of-the-art oversampling techniques such as SMOTE (e.g., with an average of about 60% improvement rate over SMOTE in recall among all datasets). Conclusion: Based on this study, we would suggest the use of optimized GANs as an alternative method for security vulnerability data class imbalanced issues.
Tianpei Xia, Laurie A. Williams, Tim Menzies
MSR2
2022 Omni: automated ensemble with unexpected models against adversarial evasion attack
Tianpei Xia, Laurie A. Williams, Tim Menzies
Empir. Softw. Eng.2
2022 Predicting health indicators for open source projects (using hyperparameter optimization)
Tianpei Xia, Wei Fu 0002, Tim Menzies
Empir. Softw. Eng.1
2022 Sequential Model Optimization for Software Effort Estimation
abstract
Many methods have been proposed to estimate how much effort is required to build and maintain software. Much of that research tries to recommend a single method – an approach that makes the dubious assumption that one method can handle the diversity of software project data. To address this drawback, we apply a configuration technique called “ROME” (Rapid Optimizing Methods for Estimation), which uses sequential model-based optimization (SMO) to find what configuration settings of effort estimation techniques work best for a particular data set. We test this method using data from 1161 traditional waterfall projects and 120 contemporary projects (from GitHub). In terms of magnitude of relative error and standardized accuracy, we find that ROME achieves better performance than the state-of-the-art methods for both traditional waterfall and contemporary projects. In addition, we conclude that we should not recommendonemethod for estimation. Rather, it is better to search through a wide range of different methods to find what works best for the local data. To the best of our knowledge, this is the largest effort estimation experiment yet attempted and the only one to test its methods on traditional waterfall and contemporary projects.
Tianpei Xia, Xipeng Shen, Tim Menzies
IEEE Trans. Software Eng.1
2021 How to Better Distinguish Security Bug Reports (Using Dual Hyperparameter Optimization)
Tianpei Xia, Laurie A. Williams, Tim Menzies
Empir. Softw. Eng.2