Zhe Yu 0002

dblp:32/9128-2 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0002-6841-1725ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 12 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Differential Parity: Relative Fairness Between Two Sets of Decisions
abstract
Background: With AI systems increasingly being applied to assist humans in decision-making processes such as talent hiring, school admissions, and loan approvals, there is a growing need to ensure that the resulting decisions are fair. A major challenge in analyzing fairness is that standards are highly subjective and context-dependent —- there is no consensus on what absolute fairness means in every scenario. Moreover, different standards of fairness often conflict with each other. Objectives: To address this issue, this work aims to evaluate the relative fairness between decisions. Methods: Instead of defining what constitutes “absolutely” fair decisions, we propose assessing the relative fairness of one decision set against another using differential parity —- two sets of decisions are considered relatively fair with respect to each other if and only if the difference between them is independent of a given sensitive attribute. The proposed notion of differential parity fairness offers three key benefits: (1) it avoids the ambiguity and contradictions inherent in defining “absolutely” fair decisions; (2) it reveals relative preferences and biases between two decision sets; and (3) it can serve as a new notion of group fairness when a reference set of decisions (e.g., ground truth) is available. One limitation of differential parity is that the two sets of decisions being compared must be made on the same data subjects. To overcome this limitation, we propose to utilize a machine learning model to bridge the gap between the two sets of decisions made on different data and approximate the differential parity metrics. In addition to differential parity and inspired by the statistical parity fairness notion, we also define relative statistical parity – the difference between the means of two sets of decisions is required to be independent of the sensitive attribute – as a weaker notion of relative fairness compared to differential parity. Results: Theoretically, we show how the proposed metrics statistically evaluate differential parity and relative statistical parity. We also proved the feasibility of using the proposed biased bridge algorithm to approximate differential parity metrics between decisions made on different data. Empirically, we evaluated the Type I and Type II error rates of differential parity and relative statistical parity both between decisions made on the same data and on different data. Experimental results suggest that differential parity outperforms relative statistical parity by having a much lower Type II error rate in both scenarios. Conclusions: With lower than 0.1 Type I and Type II error rates in both scenarios, the effectiveness of differential parity demonstrated in this article suggests that it is feasible and beneficial to evaluate relative bias between decisions made by different entities. We expect this to pave the way for the analysis of relative fairness in AI and beyond.
Zhe Yu 0002, Xiaoyin Xi, Pranam Prakash Shetty
J. Artif. Intell. Res.1
2025 Approaching code search for python as a translation retrieval problem with dual encoders
Monoshiz Mahbub Khan, Zhe Yu 0002
Empir. Softw. Eng.2
2024 FairBalance: How to Achieve Equalized Odds With Data Pre-Processing
abstract
This research seeks to benefit the software engineering society by providing a simple yet effective pre-processing approach to achieve equalized odds fairness in machine learning software. Fairness issues have attracted increasing attention since machine learning software is increasingly used for high-stakes and high-risk decisions. It is the responsibility of all software developers to make their software accountable by ensuring that the machine learning software do not perform differently on different sensitive demographic groups—satisfying equalized odds. Different from prior works which either optimize for an equalized odds related metric during the learning process like a black-box, or manipulate the training data following some intuition; this work studies the root cause of the violation of equalized odds and how to tackle it. We found that equalizing the class distribution in each demographic group with sample weights is a necessary condition for achieving equalized odds without modifying the normal training process. In addition, an important partial condition for equalized odds (zero average odds difference) can be guaranteed when the class distributions are weighted to be not only equal but also balanced (1:1). Based on these analyses, we proposed FairBalance, a pre-processing algorithm which balances the class distribution in each demographic group by assigning calculated weights to the training data. On eight real-world datasets, our empirical results show that, at low computational overhead, the proposed pre-processing algorithm FairBalance can significantly improve equalized odds without much, if any damage to the utility. FairBalance also outperforms existing state-of-the-art approaches in terms of equalized odds. To facilitate reuse, reproduction, and validation, we made our scripts available athttps://github.com/hil-se/FairBalance.
Zhe Yu 0002, Joymallya Chakraborty, Tim Menzies
IEEE Trans. Software Eng.1
2022 Assessing expert system-assisted literature reviews with a case study
Zhe Yu 0002, Jeffrey C. Carver, Gregg Rothermel, Tim Menzies
Expert Syst. Appl.1
2022 Better Data Labelling With EMBLEM (and how that Impacts Defect Prediction)
abstract
Standard automatic methods for recognizing problematic development commits can be greatly improved via the incremental application of human+artificial expertise. In this approach, called EMBLEM, an AI tool first explore the software development process to label commits that are most problematic. Humans then apply their expertise to check those labels (perhaps resulting in the AI updating the support vectors within their SVM learner). We recommend this human+AI partnership, for several reasons. When a new domain is encountered, EMBLEM can learn better ways to label which comments refer to real problems. Also, in studies with 9 open source software projects, labelling via EMBLEM's incremental application of human+AI is at least an order of magnitude cheaper than existing methods ($\approx$eight times). Further, EMBLEM is very effective. For the data sets explored here, EMBLEM better labelling methods significantly improved$P_{opt}20$and G-scores performance in nearly all the projects studied here.
Huy Tu, Zhe Yu 0002, Tim Menzies
IEEE Trans. Software Eng.2
2022 Identifying Self-Admitted Technical Debts With Jitterbug: A Two-Step Approach
abstract
Keeping track of and managing Self-Admitted Technical Debts (SATDs) are important to maintaining a healthy software project. This requires much time and effort from human experts to identify the SATDs manually. The current automated solutions do not have satisfactory precision and recall in identifying SATDs to fully automate the process. To solve the above problems, we propose a two-step framework calledJitterbugfor identifying SATDs.Jitterbugfirst identifies the “easy to find” SATDs automatically with close to 100 percent precision using a novel pattern recognition technique. Subsequently, machine learning techniques are applied to assist human experts in manually identifying the remaining “hard to find” SATDs with reduced human effort. Our simulation studies on ten software projects show thatJitterbugcan identify SATDs more efficiently (with less human effort) than the prior state-of-the-art methods.
Zhe Yu 0002, Fahmid M. Fahid, Huy Tu, Tim Menzies
IEEE Trans. Software Eng.1
2021 Learning to recognize actionable static code warnings (is intrinsically easy)
Xueqi Yang, Rahul Yedida, Zhe Yu 0002, Tim Menzies
Empir. Softw. Eng.4
2021 Understanding static code warnings: An incremental AI approach
Xueqi Yang, Zhe Yu 0002, Junjie Wang 0001, Tim Menzies
Expert Syst. Appl.2
2021 Improving Vulnerability Inspection Efficiency Using Active Learning
abstract
Software engineers can find vulnerabilities with less effort if they are directed towards code that might contain more vulnerabilities. HARMLESS is an incremental support vector machine tool that builds a vulnerability prediction model from the source code inspected to date, then suggests what source code files should be inspected next. In this way, HARMLESS can reduce the time and effort required to achieve some desired level of recall for finding vulnerabilities. The tool also provides feedback on when to stop (at that desired level of recall) while at the same time, correcting human errors by double-checking suspicious files. This paper evaluates HARMLESS on Mozilla Firefox vulnerability data. HARMLESS found 80, 90, 95, 99 percent of the vulnerabilities by inspecting 10, 16, 20, 34 percent of the source code files. When targeting 90, 95, 99 percent recall, HARMLESS could stop after inspecting 23, 30, 47 percent of the source code files. Even when human reviewers fail to identify half of the vulnerabilities (50 percent false negative rate), HARMLESS could detect 96 percent of the missing vulnerabilities by double-checking half of the inspected files. Our results serve to highlight the very steep cost of protecting software from vulnerabilities (in our case study that cost is, for example, the human effort of inspecting 28,750 × 20% = 5,750 source code files to identify 95 percent of the vulnerabilities). While this result could benefit the mission-critical projects where human resources are available for inspecting thousands of source code files, the research challenge for future work is how to further reduce that cost. The conclusion of this paper discusses various ways that goal might be achieved.
Zhe Yu 0002, Christopher Theisen, Laurie A. Williams, Tim Menzies
IEEE Trans. Software Eng.1
2020 Fairway: a way to build fair ML software
abstract
Machine learning software is increasingly being used to make decisions that affect people's lives. But sometimes, the core part of this software (the learned model), behaves in a biased manner that gives undue advantages to a specific group of people (where those groups are determined by sex, race, etc.). This "algorithmic discrimination" in the AI software systems has become a matter of serious concern in the machine learning and software engineering community. There have been works done to find "algorithmic bias" or "ethical bias" in the software system. Once the bias is detected in the AI software system, the mitigation of bias is extremely important. In this work, we a)explain how ground-truth bias in training data affects machine learning model fairness and how to find that bias in AI software,b)propose a method Fairway which combines pre-processing and in-processing approach to remove ethical bias from training data and trained model. Our results show that we can find bias and mitigate bias in a learned model, without much damaging the predictive performance of that model. We propose that (1) testing for bias and (2) bias mitigation should be a routine part of the machine learning software development life cycle. Fairway offers much support for these two purposes.
Joymallya Chakraborty, Suvodeep Majumder, Zhe Yu 0002, Tim Menzies
ESEC/SIGSOFT FSE3
2020 Better software analytics via "DUO": Data mining algorithms using/used-by optimizers
Amritanshu Agrawal, Tim Menzies, Leandro L. Minku, Markus Wagner 0007, Zhe Yu 0002
Empir. Softw. Eng.5
2020 Finding Faster Configurations Using FLASH
abstract
Finding good configurations of a software system is often challenging since the number of configuration options can be large. Software engineers often make poor choices about configuration or, even worse, they usually use a sub-optimal configuration in production, which leads to inadequate performance. To assist engineers in finding the better configuration, this article introduces Flash, a sequential model-based method that sequentially explores the configuration space by reflecting on the configurations evaluated so far to determine the next best configuration to explore. Flash scales up to software systems that defeat the prior state-of-the-art model-based methods in this area. Flash runs much faster than existing methods and can solve both single-objective and multi-objective optimization problems. The central insight of this article is to use the prior knowledge of the configuration space (gained from prior runs) to choose the next promising configuration. This strategy reduces the effort (i.e., number of measurements) required to find the better configuration. We evaluate Flash using 30 scenarios based on 7 software systems to demonstrate that Flash saves effort in 100 and 80 percent of cases in single-objective and multi-objective problems respectively by up to several orders of magnitude compared to state-of-the-art techniques.
Vivek Nair, Zhe Yu 0002, Tim Menzies, Norbert Siegmund, Sven Apel
IEEE Trans. Software Eng.2
2019 TERMINATOR: better automated UI test case prioritization
abstract
Automated UI testing is an important component of the continuous integration process of software development. A modern web-based UI is an amalgam of reports from dozens of microservices written by multiple teams. Queries on a page that opens up another will fail if any of that page's microservices fails. As a result, the overall cost for automated UI testing is high since the UI elements cannot be tested in isolation. For example, the entire automated UI testing suite at LexisNexis takes around 30 hours (3-5 hours on the cloud) to execute, which slows down the continuous integration process.
Zhe Yu 0002, Fahmid M. Fahid, Tim Menzies, Gregg Rothermel, Kyle Patrick, Snehit Cherian
ESEC/SIGSOFT FSE1
2019 FAST2: An intelligent assistant for finding relevant papers
Zhe Yu 0002, Tim Menzies
Expert Syst. Appl.1
2018 Data-driven search-based software engineering
abstract
This paper introduces Data-Driven Search-based Software Engineering (DSE), which combines insights from Mining Software Repositories (MSR) and Search-based Software Engineering (SBSE). While MSR formulates software engineering problems as data mining problems, SBSE reformulate Software Engineering (SE) problems as optimization problems and use meta-heuristic algorithms to solve them. Both MSR and SBSE share the common goal of providing insights to improve software engineering. The algorithms used in these two areas also have intrinsic relationships. We, therefore, argue that combining these two fields is useful for situations (a) which require learning from a large data source or (b) when optimizers need to know the lay of the land to find better solutions, faster.
Vivek Nair, Amritanshu Agrawal, Wei Fu 0002, George Mathew, Tim Menzies, Leandro L. Minku, Markus Wagner 0007, Zhe Yu 0002
MSR9
2018 Finding better active learners for faster literature reviews
Zhe Yu 0002, Nicholas A. Kraft, Tim Menzies
Empir. Softw. Eng.1