VLDB 2026 Research / reviewers in the wild / expert
Xinyue Wang 0003
dblp:29/5027-3
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-5837-5361ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cafe: Improved Federated Data Imputation by Leveraging Missing Data HeterogeneityabstractFederated learning (FL), a decentralized machine learning approach, offers great performance while alleviating autonomy and confidentiality concerns. Despite FL's popularity, how to deal with missing values in a federated manner is not well understood. In this work, we initiate a study of federated imputation of missing values, particularly in complex scenarios, where missing data heterogeneity exists and the state-of-the-art (SOTA) approaches for federated imputation suffer from significant loss in imputation quality. We propose Cafe, a personalized FL approach for missing data imputation. Cafe is inspired from the observation that heterogeneity can induce differences in observable and missing data distribution across clients, and that these differences can be leveraged to improve the imputation quality. Cafe computes personalized weights that are automatically calibrated for the level of heterogeneity, which can remain unknown, to develop personalized imputation models for each client. An extensive empirical evaluation over a variety of settings demonstrates that Cafe matches the performance of SOTA baselines in homogeneous settings while significantly outperforming the baselines in heterogeneous settings. Sitao Min, Hafiz Salman Asif, Xinyue Wang 0003, Jaideep Vaidya |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Data Synthesis Reinvented: Preserving Missing Patterns for Enhanced AnalysisabstractSynthetic data is being widely used as a replacement or enhancement for real data in fields as diverse as healthcare, telecommunications, and finance. Unlike real data, which represents actual people and objects, synthetic data is generated from an estimated distribution that retains key statistical properties of the real data. This makes synthetic data attractive for sharing while addressing privacy, confidentiality, and autonomy concerns. Real data often contains missing values that hold important information about individual, system, or organizational behavior. Standard synthetic data generation methods eliminate missing values as part of their pre-processing steps and thus completely ignore this valuable source of information. Instead, we propose methods to generate synthetic data that preserve both the observable and missing data distributions; consequently, retaining the valuable information encoded in the missing patterns of the real data. Our approach handles various missing data scenarios and can easily integrate with existing data generation methods. Extensive empirical evaluations on diverse datasets demonstrate the effectiveness of our approach as well as the value of preserving missing data distribution in synthetic data. Xinyue Wang 0003, Hafiz Salman Asif, Jaideep Vaidya |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Preserving Missing Data Distribution in Synthetic DataabstractData from Web artifacts and from the Web is often sensitive and cannot be directly shared for data analysis. Therefore, synthetic data generated from the real data is increasingly used as a privacy-preserving substitute. In many cases, real data from the web has missing values where the missingness itself possesses important informational content, which domain experts leverage to improve their analysis. However, this information content is lost if either imputation or deletion is used before synthetic data generation. In this paper, we propose several methods to generate synthetic data that preserve both the observable and the missing data distributions. An extensive empirical evaluation over a range of carefully fabricated and real world datasets demonstrates the effectiveness of our approach. Xinyue Wang 0003, Hafiz Salman Asif, Jaideep Vaidya |
WWW | 1 |
| 2023 | Privacy-preserving federated genome-wide association studies via dynamic samplingabstractMOTIVATION: Genome-wide association studies (GWAS) benefit from the increasing availability of genomic data and cross-institution collaborations. However, sharing data across institutional boundaries jeopardizes medical data confidentiality and patient privacy. While modern cryptographic techniques provide formal secure guarantees, the substantial communication and computational overheads hinder the practical application of large-scale collaborative GWAS. RESULTS: This work introduces an efficient framework for conducting collaborative GWAS on distributed datasets, maintaining data privacy without compromising the accuracy of the results. We propose a novel two-step strategy aimed at reducing communication and computational overheads, and we employ iterative and sampling techniques to ensure accurate results. We instantiate our approach using logistic regression, a commonly used statistical method for identifying associations between genetic markers and the phenotype of interest. We evaluate our proposed methods using two real genomic datasets and demonstrate their robustness in the presence of between-study heterogeneity and skewed phenotype distributions using a variety of experimental settings. The empirical results show the efficiency and applicability of the proposed method and the promise for its application for large-scale collaborative GWAS. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/amioamo/TDS. Xinyue Wang 0003, Leonard Dervishi, Erman Ayday, Xiaoqian Jiang, Jaideep Vaidya |
Bioinform. | 1 |
| 2022 | Facilitating Federated Genomic Data Analysis by Identifying Record Correlations while Ensuring Privacy
Leonard Dervishi, Xinyue Wang 0003, Anisa Halimi, Jaideep Vaidya, Xiaoqian Jiang, Erman Ayday |
AMIA | 2 |
| 2021 | Efficient verification for outsourced genome-wide association studies
Xinyue Wang 0003, Xiaoqian Jiang, Jaideep Vaidya |
J. Biomed. Informatics | 1 |