VLDB 2026 Research / reviewers in the wild / expert
Mostafa Milani
dblp:139/0871
· DBLP profile ↗
23ranked-venue papers in the field
3as first author
16since 2021 · last 2026
0000-0002-3386-7079ORCID · corroborated
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 16 (3 first)Big Data, Cloud & Distributed Data Systems · 6Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Heterogeneity in entity matching: A survey and experimental analysisabstractEntity matching (EM) is a fundamental task in data integration and analytics, essential for identifying records that refer to the same real-world entity across diverse sources. In practice, datasets often differ widely in structure, format, schema, and semantics, creating substantial challenges for EM. We refer to this setting as Heterogeneous EM (HEM) . This survey offers a unified perspective on HEM by introducing a taxonomy, grounded in prior work, that distinguishes two primary categories– representation and semantic heterogeneity –and their subtypes. The taxonomy provides a systematic lens for understanding how variations in data form and meaning shape the complexity of matching tasks. We then connect this framework to the FAIR principles – Findability , Accessibility , Interoperability , and Reusability –demonstrating how they both reveal the challenges of HEM and suggest strategies for mitigating them. Building on this foundation, we critically review recent EM methods, examining their ability to address different heterogeneity types, and conduct targeted experiments on state-of-the-art models to evaluate their robustness and adaptability under semantic heterogeneity. Our analysis uncovers persistent limitations in current approaches and points to promising directions for future research, including multimodal matching, human-in-the-loop workflows, deeper integration with large language models and knowledge graphs, and fairness-aware evaluation in heterogeneous settings. Mohammad Hossein Moslemi, Amir Mousavi, Behshid Behkamal, Mostafa Milani |
Data Knowl. Eng. | 4 |
| 2026 | Semi-Oblivious Chase Termination for Linear Existential Rules and Beyond: An Experimental AnalysisabstractThe chase procedure is a fundamental algorithmic tool in databases that allows us to reason with constraints, such as existential rules, with a plethora of applications. It takes a database and a set of constraints as input and iteratively completes the database as dictated by the constraints. A key challenge, though, is the fact that the chase may not terminate, which leads to the problem of checking whether it terminates given a database and a set of constraints. In this work, we focus on the semi-oblivious version of the chase, which is well-suited for practical implementations, and linear existential rules, a central class of constraints with several applications. In this setting, there is a mature body of theoretical work that provides syntactic characterizations of when the chase terminates, algorithms for checking chase termination, precise complexity results, and worst-case optimal bounds on the size of the result of the chase (whenever it is finite). Our main objective is to experimentally evaluate the existing chase termination algorithms with the aim of understanding which input parameters affect their performance, clarifying whether they can be used in practice, and revealing their performance limitations. Concerning guarded existential rules, a natural generalization of linear existential rules, one can reuse the machinery for linear existential rules by first applying the so-called linearization technique, that is, the technique of converting guarded existential rules into linear existential rules without affecting the termination of the chase. A secondary objective of this work is to understand how realistic is the use of the linearization technique in the context of the semi-oblivious chase termination problem. Amirhossein Alizad, Marco Calautti, Mostafa Milani, Andreas Pieris |
ACM Trans. Database Syst. | 3 |
| 2025 | Relation-Stratified Sampling for Shapley Values Estimation in Relational Databases
Amirhossein Alizad, Mostafa Milani |
IEEE Big Data | 2 |
| 2025 | Beyond Accuracy: An Empirical Study of Uncertainty Estimation in Imputation
Zarin Tahia Hossain, Mostafa Milani |
IEEE Big Data | 2 |
| 2025 | Evaluating SQL Understanding in Large Language Models
Ananya Rahaman, Anny Zheng, Mostafa Milani, Fei Chiang, Rachel Pottinger |
EDBT | 3 |
| 2024 | Evaluating Blocking Biases in Entity MatchingabstractEntity Matching (EM) is crucial for identifying equivalent data entities across different sources, a task that becomes increasingly challenging with the growth and heterogeneity of data. Blocking techniques, which reduce the computational complexity of EM, play a vital role in making this process scalable. Despite advancements in blocking methods, the issue of fairness—where blocking may inadvertently favor certain demographic groups—has been largely overlooked. This study extends traditional blocking metrics to incorporate fairness, providing a framework for assessing bias in blocking techniques. Through experimental analysis, we evaluate the effectiveness and fairness of various blocking methods, offering insights into their potential biases. Our findings highlight the importance of considering fairness in EM, particularly in the blocking phase, to ensure equitable outcomes in data integration tasks. Mohammad Hossein Moslemi, Harini Balamurugan, Mostafa Milani |
IEEE Big Data | 3 |
| 2024 | OTClean: Data Cleaning for Conditional Independence Violations using Optimal TransportabstractEnsuring Conditional Independence (CI) constraints is pivotal for the development of fair and trustworthy machine learning models. In this paper, we introduce OTClean, a framework that harnesses optimal transport theory for data repair under CI constraints. Optimal transport theory provides a rigorous framework for measuring the discrepancy between probability distributions, thereby ensuring control over data utility. We formulate the data repair problem concerning CIs as a Quadratically Constrained Linear Program (QCLP) and propose an alternating method for its solution. However, this approach faces scalability issues due to the computational cost associated with computing optimal transport distances, such as the Wasserstein distance. To overcome these scalability challenges, we reframe our problem as a regularized optimization problem, enabling us to develop an iterative algorithm inspired by Sinkhorn's matrix scaling algorithm, which efficiently addresses high-dimensional and large-scale data. Through extensive experiments, we demonstrate the efficacy and efficiency of our proposed methods, showcasing their practical utility in real-world data cleaning and preprocessing tasks. Furthermore, we provide comparisons with traditional approaches, highlighting the superiority of our techniques in terms of preserving data utility while ensuring adherence to the desired CI constraints. Alireza Pirhadi, Mohammad Hossein Moslemi, Alexander Cloninger, Mostafa Milani, Babak Salimi |
Proc. ACM Manag. Data | 4 |
| 2023 | Leveraging Knowledge Graphs for Matching Heterogeneous Entities and ExplanationabstractEntity matching (EM), also known as record linkage, is crucial in data integration, cleaning, and knowledge base construction. Modern matching techniques leverage deep learning and pre-trained language models (PLMs) to effectively identify matching records, showcasing significant advancements over traditional methods. However, certain critical matching aspects have received limited attention in these techniques. They heavily rely on PLMs’ encodings and face challenges in integrating external sources of knowledge to enhance matching accuracy. Additionally, these techniques often lack transparency, impeding users’ understanding of the underlying rationale for matching decisions. Furthermore, they exhibit limitations and decreased performance in handling heterogeneous records from datasets with diverse schemas. This paper presents EXKG, a novel technique that addresses these challenges and effectively matches heterogeneous records with varying attributes. EXKG combines the power of knowledge graphs (KGs) and PLMs to perform record linkage while offering explanatory insights into the matching results. We demonstrate that EXKG achieves competitive performance through experimental studies compared to state-of-the-art matching techniques. As a by-product, our solution generates explanations that give end users a comprehensive understanding of the matching process. We evaluate the quality of these explanations by using a user study and show they empower end users to make informed decisions Sahar Ghassabi, Behshid Behkamal, Mostafa Milani |
IEEE Big Data | 3 |
| 2023 | Workload-Aware Query Recommendation Using Deep Learning
Eugenie Y. Lai, Zainab Zolaktaf, Mostafa Milani, Omar AlOmeir, Jianhao Cao 0001, Rachel Pottinger |
EDBT | 3 |
| 2023 | Extending sticky-Datalog± via finite-position selection functions: Tractability, algorithms, and optimization
Leo Bertossi, Mostafa Milani |
Inf. Syst. | 2 |
| 2023 | Semi-Oblivious Chase Termination for Linear Existential Rules: An Experimental StudyabstractThe chase procedure is a fundamental algorithmic tool in databases that allows us to reason with constraints, such as existential rules, with a plethora of applications. It takes as input a database and a set of constraints, and iteratively completes the database as dictated by the constraints. A key challenge, though, is the fact that it may not terminate, which leads to the problem of checking whether it terminates given a database and a set of constraints. In this work, we focus on the semi-oblivious version of the chase, which is well-suited for practical implementations, and linear existential rules, a central class of constraints with several applications. In this setting, there is a mature body of theoretical work that provides syntactic characterizations of when the chase terminates, algorithms for checking chase termination, and precise complexity results. Our main objective is to experimentally evaluate the existing chase termination algorithms with the aim of understanding which input parameters affect their performance, clarifying whether they can be used in practice, and revealing their performance limitations. Marco Calautti, Mostafa Milani, Andreas Pieris |
Proc. VLDB Endow. | 2 |
| 2023 | Summarizing Provenance of Aggregate Query Results in Relational DatabasesabstractData provenance is any information about the origin of a piece of data and the process that led to its creation. Most database provenance work has focused on creating models and semantics to query and generate this provenance information. While comprehensive, provenance information remains large and overwhelming, making it hard for data provenance systems to support data exploration. We present a new approach to provenance exploration that builds on data summarization techniques. We contribute novel summarization schemes for the provenance of aggregation queries and techniques for the fast generation of these summarization schemes. We introduce two types of summaries for aggregate queries.Impact summariestake into account the impact of specific groups of tuples in the provenance of the query on an aggregate result, andcomparative summariesallow users to compare the provenance of two aggregate results. We also present algorithms for efficient computation of these summaries, implement optimizations using data sampling and feature selection, and conduct experiments and a user survey to show the feasibility and relevance of our approaches. Omar AlOmeir, Eugenie Y. Lai, Mostafa Milani, Rachel Pottinger |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Data Anonymization With Diversity ConstraintsabstractRecent privacy legislation has aimed to restrict and control the amount of personal data published by companies and shared with third parties. Much of this real data is not only sensitive requiring anonymization but also contains characteristic details from a variety of individuals. This diversity is desirable in many applications ranging from Web search to drug and product development. Unfortunately, data anonymization techniques have largely ignored diversity in its published result. This inadvertently propagates underlying bias in subsequent data analysis. We study the problem of finding a diverse anonymized data instance where diversity is measured via a set of diversity constraints. We formalize diversity constraints, and study their fundamental problems of satisfiability, implication, and validation. We show that determining the existence of a diverse, anonymized instance can be done in PTIME, and we present a clustering-based algorithm, along with optimizations to improve performance. We conduct extensive experiments using real and synthetic data showing the effectiveness of our techniques, and improvement over existing baselines. Our work aligns with recent trends towards responsible data science by coupling diversity with privacy-preserving data publishing. Mostafa Milani, Yu Huang 0019, Fei Chiang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Entity Matching with AUC-Based FairnessabstractThe research on fair machine learning (ML) has been growing due to the high demand for unbiased and fair ML models for objective decision-making. Most of this research has been focused on training and tuning the ML model, and less effort has been made to study biases in the processes that clean and prepare data for these models. This paper studies fairness in entity matching (EM), a.k.a. record matching and entity resolution, a primary task in a data cleaning pipeline that can significantly impact ML models’ performance. We introduce a new metric for measuring bias in EM based on Area Under the Curve (AUC) and the risk of record matching between and within subpopulations. We use this metric and real-world data to show biases in a state-of-the-art EM technique. We introduce a debiasing algorithm based on data augmentation (DA) to mitigate bias and conduct experiments to show the algorithm’s effectiveness. Soudeh Nilforoushan, Qianfan Wu, Mostafa Milani |
IEEE Big Data | 3 |
| 2021 | Preserving Diversity in Anonymized Data
Mostafa Milani, Yu Huang 0019, Fei Chiang |
EDBT | 1 |
| 2021 | Summarizing Provenance of Aggregate Query Results in Relational DatabasesabstractData provenance is any information about the origin of a piece of data and the process that led to its creation. Most database provenance work has focused on creating models and semantics to query and generate this information. While comprehensive, provenance information remains large and overwhelming, which can make it hard for provenance systems to support data exploration. We present a new approach to provenance exploration that builds on data summarization techniques. We contribute two novel summarization schemes for the provenance of aggregation queries: Impact summaries, and comparative summaries. We show with experiments that our techniques incur little overhead compared to basic summaries. We conduct a survey to show that our approaches are useful to users. Omar AlOmeir, Eugenie Y. Lai, Mostafa Milani, Rachel Pottinger |
ICDE | 3 |
| 2020 | The Pastwatch: On the usability of provenance data in relational databasesabstractProvenance information can be large and overwhelming to users. We present a set of criteria that any provenance exploration tool must have and introduce Pastwatch, a provenance exploration system that adheres to those criteria. We also address the issues associated with provenance of aggregation queries, including the creation of a summarization method that makes provenance of aggregation queries manageable for users. Finally, we conduct a quantitative user study to show statistically significant results that Pastwatch makes provenance information more efficient and easier to use than standard approaches. Omar AlOmeir, Eugenie Y. Lai, Mostafa Milani, Rachel Pottinger |
ICDE | 3 |
| 2020 | Facilitating SQL Query Composition and AnalysisabstractFormulating efficient SQL queries requires several cycles of tuning and execution. We examine methods that can accelerate and improve this interaction by providing insights about SQL queries prior to execution. We achieve this by predicting properties such as the query answer size, its run-time, and error class. Unlike existing approaches, our approach does not rely on any statistics from the database instance or query execution plans. Our approach is based on using data-driven machine learning techniques that rely on large query workloads to model SQL queries and their properties. Empirical results show that the neural network models are more accurate in predicting several query properties. Zainab Zolaktaf, Mostafa Milani, Rachel Pottinger |
SIGMOD Conference | 2 |
| 2020 | Privacy-aware data cleaning-as-a-service
Yu Huang 0019, Mostafa Milani, Fei Chiang |
Inf. Syst. | 2 |
| 2019 | CurrentClean: Interactive Change Exploration and Cleaning of Stale DataabstractEnterprises often assume their data is up-to-date, where the presence of a timestamp in the recent past qualifies the data as current. However, entities modeled in the data experience varying rates of change that influence data currency. We argue that data currency is a relative notion based on individual spatio-temporal update patterns, and these patterns can be learned and predicted. We develop CurrentClean, a probabilistic system for identifying and cleaning stale values, and enables a user to interactively explore change in her data. Our system provides a Web-based user-interface, and a backend infrastructure that learns update correlations among cell values in a database to infer and repair stale values. Our demonstration provides two motivating scenarios that highlight change exploration, and cleaning features using clinical, and sensor data from a data centre enterprise. Zheng Zheng 0005, Tri Minh Quach, Ziyi Jin, Fei Chiang, Mostafa Milani |
CIKM | 5 |
| 2019 | CurrentClean: Spatio-Temporal Cleaning of Stale DataabstractData currency is imperative towards achieving up-to-date and accurate data analysis. Data is considered current if changes in real world entities are reflected in the database. When this does not occur, stale data arises. Identifying and repairing stale data goes beyond simply having timestamps. Individual entities each have their own update patterns in both space and time. These update patterns can be learned and predicted given available query logs. In this paper, we present CurrentClean, a probabilistic system for identifying and cleaning stale values. We introduce a spatio-temporal probabilistic model that captures the database update patterns to infer stale values, and propose a set of inference rules that model spatio-temporal update patterns commonly seen in real data. We recommend repairs to clean stale values by learning from past update values over cells. Our evaluation shows CurrentClean's effectiveness to identify stale values over real data, and achieves improved error detection and repair accuracy over state-of-the-art techniques. Mostafa Milani, Zheng Zheng 0005, Fei Chiang |
ICDE | 1 |
| 2019 | Improvement of SQL Recommendation on Scientific DatabaseabstractQuery recommendation is critical to assisting first-time users, who may not have the knowledge necessary to know how to issue effective SQL queries, especially in scientific databases. To help users learn how to issue SQL queries, we turn to recommendation. In particular, we consider how to recommend SQL queries and other relevant features to users using nearest-neighbor collaborative filtering. This paper presents Skyrec Summary, an improved summary method which represents the number of resulting tuples queried from each session as an importance rating. The paper adopts SkyServer, an astronomy log dataset, and presents SkyServer Surfliner, a recommender system that supports interactive database exploration, for comparative evaluation of different summary methods. A user study on SkyServer Surfliner shows that Skyrec summary method can quickly provide helpful query recommendations. Zainab Zolaktaf, Rachel Pottinger, Mostafa Milani |
SSDBM | 4 |
| 2018 | PACAS: Privacy-Aware, Data Cleaning-as-a-ServiceabstractData cleaning consumes up to 80% of the data analysis pipeline. This is a significant overhead for organizations where data cleaning is still a manually driven process requiring domain expertise. Recent advances have fueled a new computing paradigm called Database-as-a-Service, where data management tasks are outsourced to large service providers. We propose a new Data Cleaning-as-a-Service model that allows a client to interact with a data cleaning provider who hosts curated, and sensitive data. We present PACAS: a Privacy-Aware data Cleaning-As-a-Service framework that facilitates communication between the client and the service provider via a data pricing scheme where clients issue queries, and the service provider returns clean answers for a price while protecting her data. We propose a practical privacy model in such interactive settings called (X,Y,L)-anonymity that extends existing data publishing techniques to consider the data semantics while protecting sensitive values. Our evaluation over real data shows that PACAS effectively safeguards semantically related sensitive values, and provides improved accuracy over existing privacy-aware cleaning techniques. Yu Huang 0019, Mostafa Milani, Fei Chiang |
IEEE BigData | 2 |