VLDB 2026 Research / reviewers in the wild / expert
Leor Fishman
dblp:341/5390
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Trustworthy machine learning · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › dataset bias
dataset bias analysis |
0.8 | 1 | 2024 | Feature Importance Disparities for Data Bias Investigations · ICML 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Feature Importance Disparities for Data Bias Investigations · ICML 2024 |
Machine learning › Trustworthy machine learning › interpretability
feature importance |
0.8 | 1 | 2024 | Feature Importance Disparities for Data Bias Investigations · ICML 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Feature Importance Disparities for Data Bias Investigations · ICML 2024 |
Data mining › pattern mining
subgroup discovery |
0.2 | 1 | 2024 | Feature Importance Disparities for Data Bias Investigations · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
feature importance methods · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Feature Importance Disparities for Data Bias InvestigationsabstractIt is widely held that one cause of downstream bias in classifiers is bias present in the training data. Rectifying such biases may involve context-dependent interventions such as training separate models on subgroups, removing features with bias in the collection process, or even conducting real-world experiments to ascertain sources of bias. Despite the need for such data bias investigations, few automated methods exist to assist practitioners in these efforts. In this paper, we present one such method that given a dataset $X$ consisting of protected and unprotected features, outcomes $y$, and a regressor $h$ that predicts $y$ given $X$, outputs a tuple $(f_j, g)$, with the following property: $g$ corresponds to a subset of the training dataset $(X, y)$, such that the $j^{th}$ feature $f_j$ has much larger (or smaller) *influence* in the subgroup $g$, than on the dataset overall, which we call *feature importance disparity* (FID). We show across $4$ datasets and $4$ common feature importance methods of broad interest to the machine learning community that we can efficiently find subgroups with large FID values even over exponentially large subgroup classes and in practice these groups correspond to subgroups with potentially serious bias issues as measured by standard fairness metrics. Peter W. Chang, Leor Fishman, Seth Neel |
ICML | 2 |