Ning Li 0022

dblp:14/5410-22 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
3since 2021 · last 2024
0000-0001-7394-0640ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Improving classifier-based effort-aware software defect prediction by reducing ranking errors
abstract
Context: Software defect prediction utilizes historical data to direct software quality assurance resources to potentially problematic components. Effort-aware (EA) defect prediction prioritizes more bug-like components by taking cost-effectiveness into account. In other words, it is a ranking problem, however, existing ranking strategies based on classification, give limited consideration to ranking errors. Objective: Improve the performance of classifier-based EA ranking methods by focusing on ranking errors. Method: We propose a ranking score calculation strategy called EA-Z which sets a lower bound to avoid near-zero ranking errors. We investigate four primary EA ranking strategies with 16 classification learners, and conduct the experiments for EA-Z and the other four existing strategies. Results: Experimental results from 72 data sets show EA-Z is the best ranking score calculation strategy in terms of Recall@20% and Popt when considering all 16 learners. For particular learners, imbalanced ensemble learner UBag-svm and UBst-rf achieve top performance with EA-Z. Conclusion: Our study indicates the effectiveness of reducing ranking errors for classifier-based effort-aware defect prediction. We recommend using EA-Z with imbalanced ensemble learning.
Martin J. Shepperd, Ning Li 0022
EASE3
2021 A Semi-structured Data Classification Model with Integrating Tag Sequence and Ngram
Lijun Zhang 0003, Ning Li 0022, Wei Pan 0007, Zhanhuai Li
DASFAA (2)2
2021 An Overview on Supervised Semi-structured Data Classification
abstract
Many collaboratively building resources, such as Wikipedia, Weibo and Quora, exist in the form of semi-structured data. The semi-structured data has been widely used in areas such as data integration, data distribution, data storage, data management, information retrieval and knowledge management. For large volumes of semi-structured data on the Web, semi-structured data classification technique can group them into different categories by their structure and/or content information. Supervised semi-structured data classification plays an important role in many applications. This paper provides an overview of the literature in the area of supervised semi-structured data classification. A general framework for semi-structured data classification is presented, which is mainly composed of two steps: feature extraction and model building. Several different representation models of semi-structured data are discussed, mainly including rooted labeled tree model, feature vector space model and feature set model. A large selection of semi-structured data classification approaches are reviewed in detail from two aspects: based on structure only and based on both structure and content. Finally, several future research directions for semistructured data classification are presented.
Lijun Zhang 0003, Ning Li 0022, Zhanhuai Li
DSAA2
2020 How Well Just-In-Time Defect Prediction Techniques Enhance Software Reliability?
abstract
Many Just-In-Time defect prediction (JIT) techniques, which anticipate defect-prone software changes, have been proposed in recent years. Researchers have evaluated these techniques from different perspectives and have drawn inconsistent conclusions about which JIT defect prediction techniques are the most effective and efficient. This paper evaluates JIT techniques from a reliability perspective. For short-term early evaluation, we measure JIT predictive performance on early exposed defects. While for long-term evaluation, we quantify the overall reliability improvement resulted from JIT. A case study applying 11 state-of-the-art JIT methods on 18 large open-source projects has shown: 1) Different JIT methods have their own individual strengths for different purposes, 2) in general, RandomForest is the most effective method in short-term software reliability improvement, and CBS+ performs best in long-term reliability improvement; 3) JIT prediction accuracy is highly correlated to overall reliability improvement.
Yuli Tian, Ning Li 0022, Jeff Tian, Wei Zheng 0006
QRS2
2020 A systematic review of unsupervised learning techniques for software defect prediction
Ning Li 0022, Martin J. Shepperd
Inf. Softw. Technol.1
2020 Cloud reliability and efficiency improvement via failure risk based proactive actions
Yuli Tian, Jeff Tian, Ning Li 0022
J. Syst. Softw.3
2019 The Prevalence of Errors in Machine Learning Experiments
Martin J. Shepperd, Ning Li 0022, Mahir Arzoky, Andrea Capiluppi, Steve Counsell, Giuseppe Destefanis, Stephen Swift, Allan Tucker, Leila Yousefi
IDEAL (1)3
2012 Mining Frequent Association Tag Sequences for Clustering XML Documents
Lijun Zhang 0003, Zhanhuai Li, Qun Chen 0001, Ning Li 0022, Ying Lou
APWeb5