Yulei Pang

dblp:144/3274 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Machine Learning-Based Assessment of Body Roundness Index and Cardiovascular Disease Risk in Older Adults: U.S.-China Evidence
Paicheng Liu, Qianyun Xu, Yulei Pang
COMPSAC5
2024 Identifying Which Clients Will Cancel Their Insurance Policy Machine Learning Methods
abstract
Using the 2013 data set provided by Insurance Inc., logistic regression and linear discriminant analysis models were created along with data visualizations to find out which factors recorded in the data set and the state of those factors causes a client to cancel their policy. The factors that impact whether a client will cancel are those that directly pertain to the policy. For example, the coverage type and the premium the client is paying for the policy impacts the probability the client will cancel their policy. Factors that go into forming the policy and have a relationship between one another such as age and premium, also impact the probability that a client will cancel their policy. The credit status of a client, whether it is low, medium, or high, and the type of coverage they have, has the most impact on a client's inevitability to cancel. If a client's credit score is classified as low, then that client is has a high probability of cancelling their policy according to the LDA (Linear Discriminant Analysis) classifier and logistic regression model. Likewise, if a client has coverage type$\mathbf{B}$, the probability that they will cancel their policy is higher. The sales channel used to sell a client a policy also impacts the probability they will cancel. According to the LDA classifier and the logistic regression model, if a client was sold a policy over the phone, they are more likely to cancel.
Kate Narain, Yulei Pang
COMPSAC2
2023 Identification of Adipogenic and Osteogenic Differentiation using transfer learning of ResNet-18
abstract
Human mesenchymal stem cells (hMSCs) have great potential in cell-based therapies and regenerative medicine due to their self-renewal and multipotency. hMSCs can be differentiated into several cell types, including adipocytes and osteblast. Conventional approaches for determining adipocyte formation include staining of lipid droplets (i.e., oil-red-O) during adipogenesis, which is time-consuming and uneconomical. Thus, there is an emerging need for a more effective and accurate approach to the prediction of adipogenic differentiation. Here, by combining live-cell imaging with a deep learning method, we developed a convolutional neural network-based approach to precisely predict lipid droplet formation during adipogenic differentiation of hMSCs.
Maxwell Mai, Shue Wang, Yulei Pang
ICMLA4
2022 Application of Time Series Analysis Methods to Assess the impact of COVID-19 on Carbon Dioxide Emission Reduction in Connecticut
abstract
The primary goals of this study are to determine if the datasets of positive COVID-19 test cases and CO2 emissions from Connecticut over the span of March 24th, 2020-October 31, 2021 are in any ways correlated. With climate change a prominent issue facing the entire world today, it is important to explore methods of providing records of past patterns of greenhouse gas emissions in order to inform decision making that could reduce future ones. Autoregressive integrated moving average (ARIMA) modeling is also implemented in this paper to provide forecasting based on CO2 emissions in CT starting from 2019. The most significant results from this paper are as follows: the CO2 emission data of transportation sectors including ground transportation, domestics aviation, and international aviation and weekly COVID-19 positive test cases data has a strong relationship during the first 28 weeks of the pandemic with a correlation of -86.34%. The CO2 emissions experienced on average a -22.96% change of pre-pandemic vs during initial quarantine conditions and at most a - 44.48% change when comparing the pre-pandemic mean to the during initial quarantine minimum value. Lastly, the ARIMA model found to have the lowest Akaike information criterion (AIC) was ARIMA (4,0,4). In conclusion, in the event of a collective global pandemic and lockdown conditions, less traveling resulting in a correlated decrease of CO2 emissions. This means that perhaps concentrated efforts on reducing unnecessary travel could help mitigate the levels of carbon dioxide emissions as a more long-term solution to climate change opposed to the pandemic’s short-term example.
Katherine A. Kiernan, Raymond Mugno, Yulei Pang
IEEE Big Data4
2019 Comparison of Support Vector Machine and Gradient Boosting Regression Tree for Predicting Spatially Explicit Life Cycle Global Warming and Eutrophication Impacts: A case study in corn production
abstract
Agriculture ranks one of the top contributors to global warming and nutrient pollution. Quantifying life cycle environmental impacts from agricultural production serves as scientific foundation for forming effective remediation strategies. However, the methods capable of accurately and efficiently calculating spatially explicit life cycle global warming and eutrophication impacts at a fine spatial scale over a geographic region are lacking. The objective of this study was to compare two regression models for estimating spatially explicit life cycle global warming and eutrophication, with corn production in the Midwest region as a demonstrating example. The results indicated that the gradient boosting regression tree model built with monthly weather features yielded higher predictive accuracy for life cycle global warming impact and life cycle EU. Moreover, predictive accuracy was improved at the cost of simulation time. The gradient boosting regression tree model required longer training time. Additionally, all machine learning models were million times faster than the traditional process-based model and were suitable for use in computationally-intensive applications like optimization and predication.
Xiaobo Xue Romeiko, Zhijian Guo, Yulei Pang
IEEE BigData3
2017 Predicting students' graduation outcomes through support vector machines
abstract
Low graduation rate is a significant and growing problem in U.S. higher education systems. Although previous studies have demonstrated the usefulness of building statistical models for predicting students' graduation outcomes, advanced machine learning models promise to improve the effectiveness of these models, and hone in on the “difference that makes a difference” not only on the group level, but also on the level of the individual student. In this paper we propose an ensemble support vector machines based model for predicting students' graduation. Up to about 100 features, including a set of psychological-educational factors, were employed to construct the predicting model. We evaluated the proposed model using data taken from a state university's longitudinal, cohort data sets from the incoming classes of students from 2011-2012 (n=350). The experimental results demonstrated the effectiveness of the model, with considerable accuracy, precision, and recall. This paper presents the results of analysis that were conducted in order to gauge the predictive capability of a machine learning algorithm to predict on-time graduation that took into consideration students' learning and development.
Yulei Pang, Nicolas Judd, Joseph O'Brien, Michael Ben-Avie
FIE1
2017 Fault Localizations Through Feature Selections
abstract
We introduce a novel application of feature ranking methods to the fault localization problem. We envision the problem of localizing causes of failures as instances of ranking program’s elements where elements are conceptualized as features. In this paper, we define features as program’s statements. However, in its fine-grained definition, the idea of program’s features can refer to any traits of programs. This paper proposes feature ranking-based algorithms. The algorithms analyze execution traces of both passing and failing test cases, and extract the bug signatures from the failing test cases. The proposed procedure extracts possible combinations of program’s elements when executed together from bug signatures. The feature ranking-based algorithms then order statements according to the suspiciousness of the combinations. When viewed as sequences, the combination of program’s elements produced and traced in bug signatures can be utilized to reason about the common longest subsequence. The common longest subsequence of bug signatures represents the common statements executed by all failing test cases and thus provides a means for identifying statements that contain possible faults. Our evaluation indicates that the proposed feature-based fault localization outperforms existing fault localization ranking schemes.
Yulei Pang, Xiaozhen Xue, Akbar Siami Namin
Int. J. Softw. Eng. Knowl. Eng.1
2016 Early Identification of Vulnerable Software Components via Ensemble Learning
abstract
Software components, which are vulnerable to being exploited, need to be identified and patched. Employing any prevention techniques designed for the purpose of detecting vulnerable software components in early stages can reduce the expenses associated with the software testing process significantly and thus help building a more reliable and robust software system. Although previous studies have demonstrated the effectiveness of adapting prediction techniques in vulnerability detection, the feasibility of those techniques is limited mainly because of insufficient training data sets. This paper proposes a prediction technique targeting at early identification of potentially vulnerable software components. In the proposed scheme, the potentially vulnerable components are viewed as mislabeled data that may contain true but not yet observed vulnerabilities. The proposed hybrid technique combines the supports vector machine algorithm and ensemble learning strategy to better identify potential vulnerable components. The proposed vulnerability detection scheme is evaluated using some Java Android applications. The results demonstrated that the proposed hybrid technique could identify potentially vulnerable classes with high precision and relatively acceptable accuracy and recall.
Yulei Pang, Xiaozhen Xue, Akbar Siami Namin
ICMLA1
2015 Diversity and anxiety: A case study on collaborative testing
abstract
In order to reduce students' test anxiety, collaborative testing was suggested as an evaluation strategy. However, few studies have focused on testing group construction, especially when an important factor, i.e., group diversity is taken into consideration. In this paper we conducted a case study to assess the association between group diversity and test anxiety in collaborative testing. The results observed may indicate that: 1) around 20% of students suffered from test anxiety to some extent in either an individual test or a collaborative test; 2) collaborative testing could alleviate test anxiety, whereas the effect is not statistically significant; 3) there exists a moderate positive correlation between group diversity and test anxiety in collaborative testing. The results of the study may suggest limiting group diversity in collaborative testing in order to alleviate test anxiety.
Yulei Pang, Raymond Mugno, Xiaozhen Xue
FIE1
2015 Constructing Collaborative Learning Groups with Maximum Diversity Requirements
abstract
Due to the considerable advantages of collaborative learning, group work is widely used in tertiary institutions. Previous studies demonstrated that group diversity had positive influence on group work achievement. Therefore, an interesting question that arises is how to achieve maximum group diversity effectively and automatically, especially when the features to be considered are numerous and the number of students is large. In this paper we apply a multi-start algorithm composed by a greedy constructive and strategic oscillation improvement to group students. We evaluated the technique based on a small-scale case study. The results observed indicate that the multi-start algorithm-based grouping model is feasible. It improved the overall and average students diversity within group significantly, and it also enhanced students' collaborative learning outcomes compared to random grouping model. However, we did not find any evidence on monotonic positive relationship between diversity and students' learning outcomes.
Yulei Pang, Raymond Mugno, Xiaozhen Xue, Huaying Wang
ICALT1
2015 Predicting Vulnerable Software Components through N-Gram Analysis and Statistical Feature Selection
abstract
Vulnerabilities need to be detected and removed from software. Although previous studies demonstrated the usefulness of employing prediction techniques in deciding about vulnerabilities of software components, the accuracy and improvement of effectiveness of these prediction techniques is still a grand challenging research question. This paper proposes a hybrid technique based on combining N-gram analysis and feature selection algorithms for predicting vulnerable software components where features are defined as continuous sequences of token in source code files, i.e., Java class file. Machine learning-based feature selection algorithms are then employed to reduce the feature and search space. We evaluated the proposed technique based on some Java Android applications, and the results demonstrated that the proposed technique could predict vulnerable classes, i.e., software components, with high precision, accuracy and recall.
Yulei Pang, Xiaozhen Xue, Akbar Siami Namin
ICMLA1
2014 Trimming Test Suites with Coincidentally Correct Test Cases for Enhancing Fault Localizations
abstract
Although empirical studies have demonstrated the usefulness of statistical fault localizations based on code coverage, the effectiveness of these techniques may be deteriorated due to the presence of some undesired circumstances such as the existence of coincidental correctness where one or more passing test cases exercise a faulty statement and thus causing some confusion to decide whether the underlying exercised statement is faulty or not. Fault localizations based on coverage can be improved if all possible instances of coincidental correctness are identified and proper strategies are employed to deal with these troublesome test cases. We introduce a technique to effectively identify coincidentally correct test cases. The proposed technique combines support vector machines and ensemble learning to detect mislabeled test cases, i.e. Coincidentally correct test cases. The ensemble-based support vector machine then can be used to trim a test suite or flip the test status of the coincidental correctness test cases and thus improving the effectiveness of fault localizations.
Xiaozhen Xue, Yulei Pang, Akbar Siami Namin
COMPSAC2
2014 A Clustering-Based Grouping Model for Enhancing Collaborative Learning
abstract
Group work is widely used in tertiary institutions due to the considerable advantages of collaborative learning. Previous studies indicated that the group diversity had positive influence on the group work achievement. Therefore, how to achieve diversity within a group effectively and automatically is an interesting question. In this paper we propose a novel clustering-based grouping model. The proposed technique first employs balanced K-means algorithm to divide the students into several size-balanced clusters, such that the students within the same cluster are more similar (in some sense) to each other than to those in other clusters, then adopts one-sample-each-cluster strategy to construct the groups1. We evaluated the proposed technique based on two small-scale case studies. The result observed may indicate that the clustering-based grouping model is feasible and effective.
Yulei Pang, Feiya Xiao, Huaying Wang, Xiaozhen Xue
ICMLA1
2014 Feature Selections for Effectively Localizing Faulty Events in GUI Applications
abstract
Due to the complex causality of failure and the special characteristics of test cases, the faults in GUI (Graphic User Interface) applications are difficult to localize. This paper adapts feature selection algorithms to localize GUI-related faults in a given program. Features are defined as the subsequences of events executed. By employing statistical feature ranking techniques, the events can be ranked by the suspiciousness of events being responsible to exhibit faulty behavior. The features defined in a given source code implementing (event handle) the underlying event are then ranked in suspiciousness order. The evaluation of the proposed technique based on some open source Java projects verified the effectiveness of this feature selection based fault localization technique for GUI applications.
Xiaozhen Xue, Yulei Pang, Akbar Siami Namin
ICMLA2
2013 Identifying Effective Test Cases through K-Means Clustering for Enhancing Regression Testing
abstract
Testing is the most time consuming and expensive process in the software development life cycle. In order to reduce the cost of regression testing, we propose a test case classification methodology based on k-means clustering with the purpose of classifying test cases into two groups of effective and non-effective test cases. The clustering strategy is based on Hamming distances measured over the differences between coverage information obtained for current and the previous releases of the program under test. Our empirical study shows that the clustering-based test case classification can identify effective test cases with high recall ratio and considerable accuracy percentage. The paper also investigates and compares the performance of the proposed clustering-based approach with various factors including coverage criteria and the weights factor used in measuring distances.
Yulei Pang, Xiaozhen Xue, Akbar Siami Namin
ICMLA (2)1