Suyash Shringarpure

dblp:39/6034 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
1since 2021 · last 2025
0000-0001-6464-2668ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 78% Medical and health informatics · 22%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › genomics › machine learning for genomics
deep learning for genomics
0.912025
PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025
Medical and health informatics › clinical prediction
disease risk prediction
0.912025
PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025
Bioinformatics and computational biology
genomics
0.912025
PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025
Bioinformatics and computational biology › statistical genetics › genomic prediction
polygenic risk prediction
0.912025
PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025
Bioinformatics and computational biology › population genetics
ancestry inference
0.222011
StructHDP: automatic inference of number of clusters and population structure from admixed genotype data · Bioinform. 2011
mStruct: a new admixture model for inference of population structure in light of both genetic admixing and allele mutations · ICML 2008
Bioinformatics and computational biology
population genetics
0.222011
StructHDP: automatic inference of number of clusters and population structure from admixed genotype data · Bioinform. 2011
mStruct: a new admixture model for inference of population structure in light of both genetic admixing and allele mutations · ICML 2008
Bioinformatics and computational biology › population genetics
admixture analysis
0.012011
StructHDP: automatic inference of number of clusters and population structure from admixed genotype data · Bioinform. 2011

Methods — techniques the papers use, named apart from their topics

transformer · 0.9neighborhood attention · 0.9multi-task learning · 0.9hierarchical dirichlet process · 0.1gibbs sampling · 0.1variational inference · 0.1hierarchical bayesian model · 0.1admixture model · 0.1
YearPublicationVenuePosition
2025 PRSformer: Disease Prediction from Million-Scale Individual Genotypes
abstract
Predicting disease risk from DNA presents an unprecedented emerging challenge as biobanks approach population scale sizes ($N>10^6$ individuals) with ultra-high-dimensional features ($L>10^5$ genotypes). Current methods, often linear and reliant on summary statistics, fail to capture complex genetic interactions and discard valuable individual-level information. We introduce **PRSformer**, a scalable deep learning architecture designed for end-to-end, multitask disease prediction directly from million-scale individual genotypes. PRSformer employs neighborhood attention, achieving linear $O(L)$ complexity per layer, making Transformers tractable for genome-scale inputs. Crucially, PRSformer utilizes a stacking of these efficient attention layers, progressively increasing the effective receptive field to model local dependencies (e.g., within linkage disequilibrium blocks) before integrating information across wider genomic regions. This design, tailored for genomics, allows PRSformer to learn complex, potentially non-linear and long-range interactions directly from raw genotypes. We demonstrate PRSformer's effectiveness using a unique large private cohort ($N \approx 5$M) for predicting 18 autoimmune and inflammatory conditions using $L \approx 140$k variants. PRSformer significantly outperforms highly optimized linear models trained on the *same individual-level data* and state-of-the-art summary-statistic-based methods (LDPred2) derived from the *same cohort*, quantifying the benefits of non-linear modeling and multitask learning at scale. Furthermore, experiments reveal that the advantage of non-linearity emerges primarily at large sample sizes ($N > 1$M), and that a multi-ancestry trained model improves generalization, establishing PRSformer as a new framework for deep learning in population-scale genomics.
Payam Dibaeinia, Chris German, Suyash Shringarpure, Adam Auton, Aly Azeem Khan
NeurIPS3
2011 StructHDP: automatic inference of number of clusters and population structure from admixed genotype data
abstract
MOTIVATION: Clustering of genotype data is an important way of understanding similarities and differences between populations. A summary of populations through clustering allows us to make inferences about the evolutionary history of the populations. Many methods have been proposed to perform clustering on multilocus genotype data. However, most of these methods do not directly address the question of how many clusters the data should be divided into and leave that choice to the user. METHODS: We present StructHDP, which is a method for automatically inferring the number of clusters from genotype data in the presence of admixture. Our method is an extension of two existing methods, Structure and Structurama. Using a Hierarchical Dirichlet Process (HDP), we model the presence of admixture of an unknown number of ancestral populations in a given sample of genotype data. We use a Gibbs sampler to perform inference on the resulting model and infer the ancestry proportions and the number of clusters that best explain the data. RESULTS: To demonstrate our method, we simulated data from an island model using the neutral coalescent. Comparing the results of StructHDP with Structurama shows the utility of combining HDPs with the Structure model. We used StructHDP to analyze a dataset of 155 Taita thrush, Turdus helleri, which has been previously analyzed using Structure and Structurama. StructHDP correctly picks the optimal number of populations to cluster the data. The clustering based on the inferred ancestry proportions also agrees with that inferred using Structure for the optimal number of populations. We also analyzed data from 1048 individuals from the Human Genome Diversity project from 53 world populations. We found that the clusters obtained correspond with major geographical divisions of the world, which is in agreement with previous analyses of the dataset. AVAILABILITY: StructHDP is written in C++. The code will be available for download at http://www.sailing.cs.cmu.edu/structhdp. CONTACT: [email protected]; [email protected].
Suyash Shringarpure, Daegun Won, Eric P. Xing
Bioinform.1
2008 mStruct: a new admixture model for inference of population structure in light of both genetic admixing and allele mutations
abstract
Traditional methods for analyzing population structure, such as the Structure program, ignore the influence of mutational effects. We propose mStruct, an admixture of population-specific mixtures of inheritance models, that addresses the task of structure inference and mutation estimation jointly through a hierarchical Bayesian framework, and a variational algorithm for inference. We validated our method on synthetic data, and used it to analyze the HGDP-CEPH cell line panel of microsatellites used in (Rosenberg et al., 2002) and the HGDP SNP data used in (Conrad et al., 2006). A comparison of the structural maps of world populations estimated by mStruct and Structure is presented, and we also report potentially interesting mutation patterns in world populations estimated by mStruct, which is not possible by Structure.
Suyash Shringarpure, Eric P. Xing
ICML1
2008 CSMET: Comparative Genomic Motif Detection via Multi-Resolution Phylogenetic Shadowing
abstract
Functional turnover of transcription factor binding sites (TFBSs), such as whole-motif loss or gain, are common events during genome evolution. Conventional probabilistic phylogenetic shadowing methods model the evolution of genomes only at nucleotide level, and lack the ability to capture the evolutionary dynamics of functional turnover of aligned sequence entities. As a result, comparative genomic search of non-conserved motifs across evolutionarily related taxa remains a difficult challenge, especially in higher eukaryotes, where the cis-regulatory regions containing motifs can be long and divergent; existing methods rely heavily on specialized pattern-driven heuristic search or sampling algorithms, which can be difficult to generalize and hard to interpret based on phylogenetic principles. We propose a new method: Conditional Shadowing via Multi-resolution Evolutionary Trees, or CSMET, which uses a context-dependent probabilistic graphical model that allows aligned sites from different taxa in a multiple alignment to be modeled by either a background or an appropriate motif phylogeny conditioning on the functional specifications of each taxon. The functional specifications themselves are the output of a phylogeny which models the evolution not of individual nucleotides, but of the overall functionality (e.g., functional retention or loss) of the aligned sequence segments over lineages. Combining this method with a hidden Markov model that autocorrelates evolutionary rates on successive sites in the genome, CSMET offers a principled way to take into consideration lineage-specific evolution of TFBSs during motif detection, and a readily computable analytical form of the posterior distribution of motifs under TFBS turnover. On both simulated and real Drosophila cis-regulatory modules, CSMET outperforms other state-of-the-art comparative genomic motif finders.
Pradipta Ray, Suyash Shringarpure, Mladen Kolar, Eric P. Xing
PLoS Comput. Biol.2