EDBT 2026 Demo / reviewers in the wild / expert
Payam Dibaeinia
dblp:301/4950
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 81% Medical and health informatics · 19% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › genomics › machine learning for genomics
deep learning for genomics |
0.9 | 1 | 2025 | PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025 |
Medical and health informatics › clinical prediction
disease risk prediction |
0.9 | 1 | 2025 | PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025 |
Bioinformatics and computational biology
genomics |
0.9 | 1 | 2025 | PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025 |
Bioinformatics and computational biology › statistical genetics › genomic prediction
polygenic risk prediction |
0.9 | 1 | 2025 | PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025 |
Bioinformatics and computational biology › phylogenetics
phylogenomics |
0.5 | 1 | 2021 | FASTRAL: improving scalability of phylogenomic analysis · Bioinform. 2021 |
Bioinformatics and computational biology › phylogenetics
species tree estimation |
0.5 | 1 | 2021 | FASTRAL: improving scalability of phylogenomic analysis · Bioinform. 2021 |
Bioinformatics and computational biology › phylogenetics
incomplete lineage sorting |
0.1 | 1 | 2021 | FASTRAL: improving scalability of phylogenomic analysis · Bioinform. 2021 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9neighborhood attention · 0.9multi-task learning · 0.9multi-locus coalescent model · 0.5dynamic programming · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PRSformer: Disease Prediction from Million-Scale Individual GenotypesabstractPredicting disease risk from DNA presents an unprecedented emerging challenge as biobanks approach population scale sizes ($N>10^6$ individuals) with ultra-high-dimensional features ($L>10^5$ genotypes). Current methods, often linear and reliant on summary statistics, fail to capture complex genetic interactions and discard valuable individual-level information. We introduce **PRSformer**, a scalable deep learning architecture designed for end-to-end, multitask disease prediction directly from million-scale individual genotypes. PRSformer employs neighborhood attention, achieving linear $O(L)$ complexity per layer, making Transformers tractable for genome-scale inputs. Crucially, PRSformer utilizes a stacking of these efficient attention layers, progressively increasing the effective receptive field to model local dependencies (e.g., within linkage disequilibrium blocks) before integrating information across wider genomic regions. This design, tailored for genomics, allows PRSformer to learn complex, potentially non-linear and long-range interactions directly from raw genotypes. We demonstrate PRSformer's effectiveness using a unique large private cohort ($N \approx 5$M) for predicting 18 autoimmune and inflammatory conditions using $L \approx 140$k variants. PRSformer significantly outperforms highly optimized linear models trained on the *same individual-level data* and state-of-the-art summary-statistic-based methods (LDPred2) derived from the *same cohort*, quantifying the benefits of non-linear modeling and multitask learning at scale. Furthermore, experiments reveal that the advantage of non-linearity emerges primarily at large sample sizes ($N > 1$M), and that a multi-ancestry trained model improves generalization, establishing PRSformer as a new framework for deep learning in population-scale genomics. Payam Dibaeinia, Chris German, Suyash Shringarpure, Adam Auton, Aly Azeem Khan |
NeurIPS | 1 |
| 2021 | FASTRAL: improving scalability of phylogenomic analysisabstractMOTIVATION: ASTRAL is the current leading method for species tree estimation from phylogenomic datasets (i.e. hundreds to thousands of genes) that addresses gene tree discord resulting from incomplete lineage sorting (ILS). ASTRAL is statistically consistent under the multi-locus coalescent model (MSC), runs in polynomial time, and is able to run on large datasets. Key to ASTRAL's algorithm is the use of dynamic programming to find an optimal solution to the MQSST (maximum quartet support supertree) within a constraint space that it computes from the input. Yet, ASTRAL can fail to complete within reasonable timeframes on large datasets with many genes and species, because in these cases the constraint space it computes is too large. RESULTS: Here, we introduce FASTRAL, a phylogenomic estimation method. FASTRAL is based on ASTRAL, but uses a different technique for constructing the constraint space. The technique we use to define the constraint space maintains statistical consistency and is polynomial time; thus we prove that FASTRAL is a polynomial time algorithm that is statistically consistent under the MSC. Our performance study on both biological and simulated datasets demonstrates that FASTRAL matches or improves on ASTRAL with respect to species tree topology accuracy (and under high ILS conditions it is statistically significantly more accurate), while being dramatically faster-especially on datasets with large numbers of genes and high ILS-due to using a significantly smaller constraint space. AVAILABILITYAND IMPLEMENTATION: FASTRAL is available in open-source form at https://github.com/PayamDiba/FASTRAL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Payam Dibaeinia, Shayan Tabe-Bordbar, Tandy J. Warnow |
Bioinform. | 1 |