VLDB 2026 Research / reviewers in the wild / expert
Nima Shahbazi
dblp:40/4462
· DBLP profile ↗
10ranked-venue papers in the field
7as first author
8since 2021 · last 2024
0000-0001-7016-3807ORCID · corroborated
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 8 (6 first)Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Data Coverage for Detecting Representation Bias in Image Datasets: A Crowdsourcing Approach
Melika Mousavi, Nima Shahbazi, Abolfazl Asudeh |
EDBT | 2 |
| 2024 | Fairness-Aware Data Preparation for Entity MatchingabstractEntity matching is a crucial task in many real applications. Despite the substantial body of research that focuses on improving the effectiveness of entity matching, enhancing its fairness has received scant attention. To fill this gap, this paper introduces a new problem of preparing fairness-aware datasets for entity matching. We formally outline the problem, drawing upon the principles of group fairness and statistical parity. We devise three highly efficient algorithms to accelerate the process of identifying an unbiased dataset from the vast search space. Our experiments on four real-world datasets show that our proposed algorithms can significantly improve fairness in the results while achieving comparable effectiveness to existing fairness-agnostic methods. Furthermore, we conduct case studies to demonstrate that our proposed techniques can be seamlessly integrated into end-to-end entity matching pipelines to support fairness requirements in real-world applications. Nima Shahbazi, Jin Wang 0007, Zhengjie Miao, Nikita Bhutani |
ICDE | 1 |
| 2024 | FairHash: A Fair and Memory/Time-efficient HashmapabstractHashmap is a fundamental data structure in computer science. There has been extensive research on constructing hashmaps that minimize the number of collisions leading to efficient lookup query time. Recently, the data-dependant approaches, construct hashmaps tailored for a target data distribution that guarantee to uniformly distribute data across different buckets and hence minimize the collisions. Still, to the best of our knowledge, none of the existing technique guarantees group fairness among different groups of items stored in the hashmap. Therefore, in this paper, we introduce FairHash, a data-dependant hashmap that guarantees uniform distribution at the group-level across hash buckets, and hence, satisfies the statistical parity notion of group fairness. We formally define, three notions of fairness and, unlike existing work, FairHash satisfies all three of them simultaneously. We propose three families of algorithms to design fair hashmaps, suitable for different settings. Our ranking-based algorithms reduce the unfairness of data-dependant hashmaps without any memory-overhead. The cut-based algorithms guarantee zero-unfairness in all cases, irrespective of how the data is distributed, but those introduce an extra memory-overhead. Last but not least, the discrepancy-based algorithms enable trading off between various fairness notions. In addition to the theoretical analysis, we perform extensive experiments to evaluate the efficiency and efficacy of our algorithms on real datasets. Our results verify the superiority of FairHash compared to the other baselines on fairness at almost no performance cost. Nima Shahbazi, Stavros Sintos, Abolfazl Asudeh |
Proc. ACM Manag. Data | 1 |
| 2024 | FairEM360: A Suite for Responsible Entity MatchingabstractEntity matching is one of the earliest tasks that occur in the big data pipeline and is alarmingly exposed to unintentional biases that affect the quality of data. Identifying and mitigating the biases that exist in the data or are introduced by the matcher at this stage can contribute to promoting fairness in downstream tasks. This demonstration showcases FairEM360, a framework for 1) auditing the output of entity matchers across a wide range of fairness measures and paradigms, 2) providing potential explanations for the underlying reasons for unfairness, and 3) providing resolutions for the unfairness issues through an exploratory process with human-in-the-loop feedback, utilizing an ensemble of matchers. We aspire for FairEM360 to contribute to the prioritization of fairness as a key consideration in the evaluation of EM pipelines. Nima Shahbazi, Mahdi Erfanian, Abolfazl Asudeh, Fatemeh Nargesian, Divesh Srivastava |
Proc. VLDB Endow. | 1 |
| 2024 | Reliability evaluation of individual predictions: a data-centric approach
Nima Shahbazi, Abolfazl Asudeh |
VLDB J. | 1 |
| 2023 | Through the Fairness Lens: Experimental Analysis and Evaluation of Entity MatchingabstractEntity matching (EM) is a challenging problem studied by different communities for over half a century. Algorithmic fairness has also become a timely topic to address machine bias and its societal impacts. Despite extensive research on these two topics, little attention has been paid to the fairness of entity matching. Towards addressing this gap, we perform an extensive experimental evaluation of a variety of EM techniques in this paper. We generated two social datasets from publicly available datasets for the purpose of auditing EM through the lens of fairness. Our findings underscore potential unfairness under two common conditions in real-world societies: (i) when some demographic groups are over-represented, and (ii) when names are more similar in some groups compared to others. Among our many findings, it is noteworthy to mention that while various fairness definitions are valuable for different settings, due to EM's class imbalance nature, measures such as positive predictive value parity and true positive rate parity are, in general, more capable of revealing EM unfairness. Nima Shahbazi, Nikola Danevski, Fatemeh Nargesian, Abolfazl Asudeh, Divesh Srivastava |
Proc. VLDB Endow. | 1 |
| 2022 | Upper bounds for can-tree and FP-tree
Nima Shahbazi, Jarek Gryz |
J. Intell. Inf. Syst. | 1 |
| 2021 | Identifying Insufficient Data Coverage for Ordinal Continuous-Valued AttributesabstractAppropriate training data is a requirement for building good machine-learned models. In this paper, we study the notion of coverage for ordinal and continuous-valued attributes, by formalizing the intuition that the learned model can accurately predict only at data points for which there are "enough" similar data points in the training data set. Abolfazl Asudeh, Nima Shahbazi, Zhongjun Jin, H. V. Jagadish |
SIGMOD Conference | 2 |
| 2020 | Upper Bound on the Size of FP-Tree
Nima Shahbazi, Jarek Gryz |
ADBIS | 1 |
| 2018 | Machine learning and BIM visualization for maintenance issue classification and enhanced data collection
J. J. McArthur, Nima Shahbazi, Ricky Fok, Christopher Raghubar, Brandon Bortoluzzi, Aijun An |
Adv. Eng. Informatics | 2 |