VLDB 2026 Research / reviewers in the wild / expert
Simone Fabbrizzi
dblp:297/4647
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0003-4374-5806ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Adversarial Reweighting Guided by Wasserstein Distance to Achieve Demographic ParityabstractTo address bias issues, fair machine learning usually jointly optimizes two (or more) metrics aiming at predictive utility and fairness. However, the inherent under-representation of minorities in the data often makes the disparate impact of subpopulations less noticeable and difficult to deal with during learning. In this paper, we propose a novel adversarial reweighting method to address such disparate impact. To balance the data distribution between the majority and the minority groups, our approach prefers samples from the majority group that are closer to the minority group as evaluated by the Wasserstein distance. Theoretical analysis shows the effectiveness of our adversarial reweighting approach. Experiments demonstrate that our approach mitigates disparate impact without sacrificing classification accuracy, outperforming related state-of-the-art methods on image and tabular benchmark datasets. Code is available at https://github.com/zhaoxuan00707/wasserstein_reweight. Xuan Zhao 0025, Simone Fabbrizzi, Paula Reyero Lobo, S. Siamak Ghodsi, Klaus Broelemann, Steffen Staab, Gjergji Kasneci |
IEEE Big Data | 2 |
| 2024 | Studying bias in visual features through the lens of optimal transportabstractAbstract Computer vision systems are employed in a variety of high-impact applications. However, making them trustworthy requires methods for the detection of potential biases in their training data, before models learn to harm already disadvantaged groups in downstream applications. Image data are typically represented via extracted features, which can be hand-crafted or pre-trained neural network embeddings. In this work, we introduce a framework for bias discovery given such features that is based on optimal transport theory; it uses the (quadratic) Wasserstein distance to quantify disparity between the feature distributions of two demographic groups (e.g., women vs men). In this context, we show that the Kantorovich potentials of the images, which are a byproduct of computing the Wasserstein distance and act as “transportation prices", can serve as bias scores by indicating which images might exhibit distinct biased characteristics. We thus introduce a visual dataset exploration pipeline that helps auditors identify common characteristics across high- or low-scored images as potential sources of bias. We conduct a case study to identify prospective gender biases and demonstrate theoretically-derived properties with experiments on the CelebA and Biased MNIST datasets. Simone Fabbrizzi, Xuan Zhao 0025, Emmanouil Krasanakis, Symeon Papadopoulos, Eirini Ntoutsi |
Data Min. Knowl. Discov. | 1 |
| 2024 | Correction to: Studying bias in visual features through the lens of optimal transportabstractIn this article the statement after Equation 1 had an error in the published version. Please refer the correction as follows: “where ν = T#µ and T# is the push-forward of µ along the function T : X → Y” was incorrectly written as “where T# is the push-forward of µ along the function T : X → Y. Furthermore, Equation 1 itself was incorrectly formulated. Namely, the integral should have been over X and not over X × Y. The original article has been corrected. Simone Fabbrizzi, Xuan Zhao 0025, Emmanouil Krasanakis, Symeon Papadopoulos, Eirini Ntoutsi |
Data Min. Knowl. Discov. | 1 |
| 2022 | A survey on bias in visual datasetsabstractComputer Vision (CV) has achieved remarkable results, outperforming humans in several tasks. Nonetheless, it may result in significant discrimination if not handled properly. Indeed, CV systems highly depend on training datasets and can learn and amplify biases that such datasets may carry. Thus, the problem of understanding and discovering bias in visual datasets is of utmost importance; yet, it has not been studied in a systematic way to date. Hence, this work aims to: (i) describe the different kinds of bias that may manifest in visual datasets; (ii) review the literature on methods for bias discovery and quantification in visual datasets; (iii) discuss existing attempts to collect visual datasets in a bias-aware manner. A key conclusion of our study is that the problem of bias discovery and quantification in visual datasets is still open, and there is room for improvement in terms of both methods and the range of biases that can be addressed. Moreover, there is no such thing as a bias-free dataset, so scientists and practitioners must become aware of the biases in their datasets and make them explicit. To this end, we propose a checklist to spot different types of bias during visual dataset collection. Simone Fabbrizzi, Symeon Papadopoulos, Eirini Ntoutsi, Ioannis Kompatsiaris |
Comput. Vis. Image Underst. | 1 |