EDBT 2026 Demo / reviewers in the wild / expert
Gamze Gürsoy
dblp:46/10044
· DBLP profile ↗
10ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-1352-8686ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Universal Metric of Dataset Similarity for Cross-Silo Federated LearningabstractFederated Learning (FL) enables collaborative model training across institutions without sharing raw data, making it valuable for privacy-sensitive domains like healthcare. However, FL performance deteriorates significantly when client datasets are non-IID. While dataset similarity metrics could guide collaboration decisions, existing approaches have critical limitations: unbounded costs that lack interpretability across domains, requirements for direct data access that violate FL's privacy constraints, and poor sample efficiency. We propose a novel metric that extracts model representations after a single federated training round to predict whether collaboration will improve performance. Our approach formulates similarity assessment as an optimal transport problem with a hybrid cost function that captures both feature-level differences and label distribution divergence between clients. We ensure privacy through a careful composition of Secure Multiparty Computation (SMC) and Differential Privacy (DP) mechanisms. Our theoretical analysis establishes a formal connection between the proposed metric and weight divergence in federated training, explaining why early-round activations can predict long-term collaboration outcomes. Empirically, the metric remains tightly correlated with weight divergence throughout training, reinforcing the validity of our single-round probe. Extensive experiments on synthetic benchmarks and real-world medical imaging tasks demonstrate that our metric reliably identifies beneficial collaborations providing practitioners with an actionable tool for participant selection in cross-silo FL. Ahmed Elhussein, Gamze Gürsoy |
ICDM | 2 |
| 2025 | ScatTR: Estimating the Size of Long Tandem Repeat Expansions from Short-Reads
Rashid Al-Abri, Gamze Gürsoy |
RECOMB | 2 |
| 2024 | Privacy-preserving model evaluation for logistic and linear regression using homomorphically encrypted genotype dataabstractOBJECTIVE: Linear and logistic regression are widely used statistical techniques in population genetics for analyzing genetic data and uncovering patterns and associations in large genetic datasets, such as identifying genetic variations linked to specific diseases or traits. However, obtaining statistically significant results from these studies requires large amounts of sensitive genotype and phenotype information from thousands of patients, which raises privacy concerns. Although cryptographic techniques such as homomorphic encryption offers a potential solution to the privacy concerns as it allows computations on encrypted data, previous methods leveraging homomorphic encryption have not addressed the confidentiality of shared models, which can leak information about the training data. METHODS: In this work, we present a secure model evaluation method for linear and logistic regression using homomorphic encryption for six prediction tasks, where input genotypes, output phenotypes, and model parameters are all encrypted. RESULTS: Our method ensures no private information leakage during inference and achieves high accuracy (≥93% for all outcomes) with each inference taking less than ten seconds for ∼200 genomes. CONCLUSION: Our study demonstrates that it is possible to perform linear and logistic regression model evaluation while protecting patient confidentiality with theoretical security guarantees. Our implementation and test data are available at https://github.com/G2Lab/privateML/. Seungwan Hong 0001, Yoolim A. Choi, Daniel S. Joo, Gamze Gürsoy |
J. Biomed. Informatics | 4 |
| 2023 | A generalizable physiological model for detection of Delayed Cerebral Ischemia using Federated LearningabstractDelayed cerebral ischemia (DCI) is a complication seen in patients with subarachnoid hemorrhage stroke. It is a major predictor of poor outcomes and is detected late. Machine learning models are shown to be useful for early detection, however training such models suffers from small sample sizes due to rarity of the condition. Here we propose a Federated Learning approach to train a DCI classifier across three institutions to overcome challenges of sharing data across hospitals. We developed a framework for federated feature selection and built a federated ensemble classifier. We compared the performance of FL model to that obtained by training separate models at each site. FL significantly improved performance at only two sites. We found that this was due to feature distribution differences across sites. FL improves performance in sites with similar feature distributions, however, FL can worsen performance in sites with heterogeneous distributions. The results highlight both the benefit of FL and the need to assess dataset distribution similarity before conducting FL. Ahmed Elhussein, Jude P. Savarraj, E. Sander Connolly, Murad Megjhani, Soon Bin Kwon, Angela Velazquez, Daniel Nametz, David J. Roh, Jan Claassen, Miriam Weiss, Sachin Agarwal 0005, Huimahn A. Choi, Gerrit A. Schubert, Soojin Park, Gamze Gürsoy |
BIBM | 15 |
| 2023 | LDmat: efficiently queryable compression of linkage disequilibrium matricesabstractMOTIVATION: Linkage disequilibrium (LD) matrices derived from large populations are widely used in population genetics in fine-mapping, LD score regression, and linear mixed models for Genome-wide Association Studies (GWAS). However, these matrices can reach large sizes when they are derived from millions of individuals; hence, moving, sharing and extracting granular information from this large amount of data can be cumbersome. RESULTS: We sought to address the need for compressing and easily querying large LD matrices by developing LDmat. LDmat is a standalone tool to compress large LD matrices in an HDF5 file format and query these compressed matrices. It can extract submatrices corresponding to a sub-region of the genome, a list of select loci, and loci within a minor allele frequency range. LDmat can also rebuild the original file formats from the compressed files. AVAILABILITY AND IMPLEMENTATION: LDmat is implemented in python, and can be installed on Unix systems with the command 'pip install ldmat'. It can also be accessed through https://github.com/G2Lab/ldmat and https://pypi.org/project/ldmat/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rockwell J. Weiner, Chirag M. Lakhani, David A. Knowles, Gamze Gürsoy |
Bioinform. | 4 |
| 2022 | A consortium blockchain for secure, unified, and efficient sharing of EHR and genomics data
Ahmed Elhussein, Gamze Gürsoy |
AMIA | 2 |
| 2022 | How to Achieve Privacy in Large Diverse Health Systems
Gamze Gürsoy, Bradley A. Malin, Erman Ayday, Ellen Wright Clayton |
AMIA | 1 |
| 2021 | FANCY: fast estimation of privacy risk in functional genomics dataabstractMOTIVATION: Functional genomics data are becoming clinically actionable, raising privacy concerns. However, quantifying privacy leakage via genotyping is difficult due to the heterogeneous nature of sequencing techniques. Thus, we present FANCY, a tool that rapidly estimates the number of leaking variants from raw RNA-Seq, ATAC-Seq and ChIP-Seq reads, without explicit genotyping. FANCY employs supervised regression using overall sequencing statistics as features and provides an estimate of the overall privacy risk before data release. RESULTS: FANCY can predict the cumulative number of leaking SNVs with an average 0.95 R2 for all independent test sets. We realize the importance of accurate prediction when the number of leaked variants is low. Thus, we develop a special version of the model, which can make predictions with higher accuracy when the number of leaking variants is low. AVAILABILITY AND IMPLEMENTATION: A python and MATLAB implementation of FANCY, as well as custom scripts to generate the features can be found at https://github.com/gersteinlab/FANCY. We also provide jupyter notebooks so that users can optimize the parameters in the regression model based on their own data. An easy-to-use webserver that takes inputs and displays results can be found at fancy.gersteinlab.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gamze Gürsoy, Charlotte M. Brannon, Fabio C. P. Navarro, Mark Gerstein |
Bioinform. | 1 |
| 2020 | DiNeR: a Differential graphical model for analysis of co-regulation Network RewiringabstractBACKGROUND: During transcription, numerous transcription factors (TFs) bind to targets in a highly coordinated manner to control the gene expression. Alterations in groups of TF-binding profiles (i.e. "co-binding changes") can affect the co-regulating associations between TFs (i.e. "rewiring the co-regulator network"). This, in turn, can potentially drive downstream expression changes, phenotypic variation, and even disease. However, quantification of co-regulatory network rewiring has not been comprehensively studied. RESULTS: To address this, we propose DiNeR, a computational method to directly construct a differential TF co-regulation network from paired disease-to-normal ChIP-seq data. Specifically, DiNeR uses a graphical model to capture the gained and lost edges in the co-regulation network. Then, it adopts a stability-based, sparsity-tuning criterion -- by sub-sampling the complete binding profiles to remove spurious edges -- to report only significant co-regulation alterations. Finally, DiNeR highlights hubs in the resultant differential network as key TFs associated with disease. We assembled genome-wide binding profiles of 104 TFs in the K562 and GM12878 cell lines, which loosely model the transition between normal and cancerous states in chronic myeloid leukemia (CML). In total, we identified 351 significantly altered TF co-regulation pairs. In particular, we found that the co-binding of the tumor suppressor BRCA1 and RNA polymerase II, a well-known transcriptional pair in healthy cells, was disrupted in tumors. Thus, DiNeR successfully extracted hub regulators and discovered well-known risk genes. CONCLUSIONS: Our method DiNeR makes it possible to quantify changes in co-regulatory networks and identify alterations to TF co-binding patterns, highlighting key disease regulators. Our method DiNeR makes it possible to quantify changes in co-regulatory networks and identify alterations to TF co-binding patterns, highlighting key disease regulators. Jing Zhang 0062, Jason Liu 0003, Donghoon Lee 0006, Shaoke Lou, Zhanlin Chen, Gamze Gürsoy, Mark Gerstein |
BMC Bioinform. | 6 |
| 2017 | Spatial organization of the budding yeast genome in the cell nucleus and identification of specific chromatin interactions from multi-chromosome constrained chromatin modelabstractNuclear landmarks and biochemical factors play important roles in the organization of the yeast genome. The interaction pattern of budding yeast as measured from genome-wide 3C studies are largely recapitulated by model polymer genomes subject to landmark constraints. However, the origin of inter-chromosomal interactions, specific roles of individual landmarks, and the roles of biochemical factors in yeast genome organization remain unclear. Here we describe a multi-chromosome constrained self-avoiding chromatin model (mC-SAC) to gain understanding of the budding yeast genome organization. With significantly improved sampling of genome structures, both intra- and inter-chromosomal interaction patterns from genome-wide 3C studies are accurately captured in our model at higher resolution than previous studies. We show that nuclear confinement is a key determinant of the intra-chromosomal interactions, and centromere tethering is responsible for the inter-chromosomal interactions. In addition, important genomic elements such as fragile sites and tRNA genes are found to be clustered spatially, largely due to centromere tethering. We uncovered previously unknown interactions that were not captured by genome-wide 3C studies, which are found to be enriched with tRNA genes, RNAPIII and TFIIS binding. Moreover, we identified specific high-frequency genome-wide 3C interactions that are unaccounted for by polymer effects under landmark constraints. These interactions are enriched with important genes and likely play biological roles. Gamze Gürsoy, Jie Liang 0002 |
PLoS Comput. Biol. | 1 |