Adam Auton

dblp:85/9915 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
1since 2021 · last 2025
0000-0002-1630-1225ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 76% Medical and health informatics · 24%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › genomics › machine learning for genomics
deep learning for genomics
0.912025
PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025
Medical and health informatics › clinical prediction
disease risk prediction
0.912025
PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025
Bioinformatics and computational biology
genomics
0.912025
PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025
Bioinformatics and computational biology › statistical genetics › genomic prediction
polygenic risk prediction
0.912025
PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025
Bioinformatics and computational biology › genomics › genomic data management
variant call format
0.112011
The variant call format and VCFtools · Bioinform. 2011

Methods — techniques the papers use, named apart from their topics

transformer · 0.9neighborhood attention · 0.9multi-task learning · 0.9indexing · 0.1data compression · 0.1
YearPublicationVenuePosition
2025 PRSformer: Disease Prediction from Million-Scale Individual Genotypes
abstract
Predicting disease risk from DNA presents an unprecedented emerging challenge as biobanks approach population scale sizes ($N>10^6$ individuals) with ultra-high-dimensional features ($L>10^5$ genotypes). Current methods, often linear and reliant on summary statistics, fail to capture complex genetic interactions and discard valuable individual-level information. We introduce **PRSformer**, a scalable deep learning architecture designed for end-to-end, multitask disease prediction directly from million-scale individual genotypes. PRSformer employs neighborhood attention, achieving linear $O(L)$ complexity per layer, making Transformers tractable for genome-scale inputs. Crucially, PRSformer utilizes a stacking of these efficient attention layers, progressively increasing the effective receptive field to model local dependencies (e.g., within linkage disequilibrium blocks) before integrating information across wider genomic regions. This design, tailored for genomics, allows PRSformer to learn complex, potentially non-linear and long-range interactions directly from raw genotypes. We demonstrate PRSformer's effectiveness using a unique large private cohort ($N \approx 5$M) for predicting 18 autoimmune and inflammatory conditions using $L \approx 140$k variants. PRSformer significantly outperforms highly optimized linear models trained on the *same individual-level data* and state-of-the-art summary-statistic-based methods (LDPred2) derived from the *same cohort*, quantifying the benefits of non-linear modeling and multitask learning at scale. Furthermore, experiments reveal that the advantage of non-linearity emerges primarily at large sample sizes ($N > 1$M), and that a multi-ancestry trained model improves generalization, establishing PRSformer as a new framework for deep learning in population-scale genomics.
Payam Dibaeinia, Chris German, Suyash Shringarpure, Adam Auton, Aly Azeem Khan
NeurIPS4
2011 The variant call format and VCFtools
abstract
SUMMARY: The variant call format (VCF) is a generic format for storing DNA polymorphism data such as SNPs, insertions, deletions and structural variants, together with rich annotations. VCF is usually stored in a compressed manner and can be indexed for fast data retrieval of variants from a range of positions on the reference genome. The format was developed for the 1000 Genomes Project, and has also been adopted by other projects such as UK10K, dbSNP and the NHLBI Exome Project. VCFtools is a software suite that implements various utilities for processing VCF files, including validation, merging, comparing and also provides a general Perl API. AVAILABILITY: http://vcftools.sourceforge.net
Petr Danecek, Adam Auton, Gonçalo R. Abecasis, Cornelis A. Albers, Eric Banks, Mark A. DePristo, Robert E. Handsaker, Gerton Lunter, Gabor T. Marth, Stephen T. Sherry, Gil McVean, Richard Durbin
Bioinform.2