EDBT 2026 Demo / reviewers in the wild / expert
Alba Gutiérrez-Sacristán
dblp:168/1184
· DBLP profile ↗
13ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0002-1245-198XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 6 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VarPPUD: Pinpointing diagnostic variants from sets of prioritized, strong candidate variantsabstractRare and ultra-rare genetic conditions are estimated to impact nearly 1 in 17 people worldwide, yet accurately pinpointing the diagnostic variants underlying each of these conditions remains a formidable challenge. Because comprehensive, in vivo functional assessment of all possible genetic variants is infeasible, clinicians instead consider in silico variant pathogenicity predictions to distinguish plausibly disease-causing from benign variants across the genome. However, in the most difficult undiagnosed cases, such as those accepted to the Undiagnosed Diseases Network (UDN), existing pathogenicity predictions cannot reliably discern true etiological variant(s) from other deleterious candidate variants that were prioritized through case- or family-level analyses. Pinpointing the disease-causing variant from a small pool of plausible candidates remains a largely manual effort requiring extensive clinical workups, functional and experimental assays, and eventual identification of genotype- and phenotype-matched individuals. Here, we introduce VarPPUD, a tool trained on prioritized variants from UDN cases, that leverages gene-, amino acid-, and nucleotide-level features to discern pathogenic (disease causative) variants from other damaging or deleterious variants that are unlikely to be confirmed as relevant to the disease. VarPPUD achieves a cross-validated accuracy of 79.3% and precision of 77.5% on a held-out subset of uniquely challenging UDN cases, respectively representing an average 18.6% and 23.4% improvement over nine existing state-of-the-art pathogenicity prediction tools on this task. We validate VarPPUD's ability to discriminate likely from unlikely pathogenic variants using both synthetic data generated via a GAN-based framework and a temporally held-out set of UDN patients evaluated between 2022 and 2024. The model was trained exclusively on data available through 2021 and applied without retraining to the post-2021 cohort, demonstrating strong generalizability to newly accrued cases. Finally, we show how VarPPUD can be probed to evaluate each input feature's importance and contribution toward prediction-an essential step toward understanding the distinct characteristics of newly-uncovered disease-causing variants. Rui Yin 0002, Alba Gutiérrez-Sacristán, Shilpa Nadimpalli Kobren, Paul Avillach |
PLoS Comput. Biol. | 2 |
| 2023 | Building a collaborative cloud platform to accelerate heart, lung, blood, and sleep researchabstractResearch increasingly relies on interrogating large-scale data resources. The NIH National Heart, Lung, and Blood Institute developed the NHLBI BioData CatalystⓇ (BDC), a community-driven ecosystem where researchers, including bench and clinical scientists, statisticians, and algorithm developers, find, access, share, store, and compute on large-scale datasets. This ecosystem provides secure, cloud-based workspaces, user authentication and authorization, search, tools and workflows, applications, and new innovative features to address community needs, including exploratory data analysis, genomic and imaging tools, tools for reproducibility, and improved interoperability with other NIH data science platforms. BDC offers straightforward access to large-scale datasets and computational resources that support precision medicine for heart, lung, blood, and sleep conditions, leveraging separately developed and managed platforms to maximize flexibility based on researcher needs, expertise, and backgrounds. Through the NHLBI BioData Catalyst Fellows Program, BDC facilitates scientific discoveries and technological advances. BDC also facilitated accelerated research on the coronavirus disease-2019 (COVID-19) pandemic. Stanley C. Ahalt, Paul Avillach, Rebecca R. Boyles, Kira Bradford, Steven Cox 0001, Brandi Davis-Dusenbery, Robert L. Grossman, Ashok K. Krishnamurthy 0001, Alisa Manning, Benedict Paten, Anthony Philippakis, Ingrid Borecki, Shu Hui Chen, Jon Kaltman, Sweta Ladwa, Chip Schwartz, Alastair Thomson, Sarah Davis, Alison Leaf, Jessica Lyons, Elizabeth Sheets, Joshua C. Bis, Matthew P. Conomos, Alessandro Culotti, Thomas N. Desain, Jack DiGiovanna, Milan Domazet, Stephanie M. Gogarten, Alba Gutiérrez-Sacristán, Tim Harris 0003, Benjamin D. Heavner, Deepti Jain, Brian O'Connor, Kevin Osborn, Danielle Pillion, Jacob Pleiness, Ken Rice, Garrett Rupp, Arnaud Serret-Larmande, Albert Smith, Jason Stedman, Adrienne Stilp, Teresa Barsanti, John B. Cheadle, Christopher Erdmann, Brandy Farlow, Allie Gartland-Gray, Julie Hayes, Hannah Hiles, Paul Kerr, W. Christopher Lenhardt, Tom Madden, Joanna O. Mieczkowska, Amanda Miller, Patrick Patton, Marcie Rathbun, Stephanie Suber, Joe Asare |
J. Am. Medical Informatics Assoc. | 29 |
| 2023 | Informative missingness: What can we learn from patterns in missing laboratory data in the electronic health record?
Amelia L. M. Tan, Emily J. Getzen, Meghan Hutch, Zachary H. Strasser, Alba Gutiérrez-Sacristán, Trang T. Le, Arianna Dagliati, Michele Morris, David A. Hanauer, Bertrand Moal, Clara-Lea Bonzel, William Yuan, Lorenzo Chiudinelli, Priyam Das, Harrison G. Zhang, Bruce J. Aronow, Paul Avillach, Gabriel A. Brat, Tianxi Cai, Chuan Hong, William G. La Cava, He Hooi Will Loh, Yuan Luo 0001, Shawn N. Murphy, Kee Yuan Hgiam, Gilbert S. Omenn, Lav P. Patel, Malarkodi J. Samayamuthu, Emily R. Shriver, Zahra Shakeri Hossein Abad, Byorn W. L. Tan, Shyam Visweswaran, Griffin M. Weber, Zongqi Xia, Bertrand Verdy, Qi Long, Danielle L. Mowery, John H. Holmes |
J. Biomed. Informatics | 5 |
| 2022 | Multi-PheWAS intersection approach to identify sex differences across comorbidities in 59 140 pediatric patients with autism spectrum disorderabstractOBJECTIVE: To identify differences related to sex and define autism spectrum disorder (ASD) comorbidities female-enriched through a comprehensive multi-PheWAS intersection approach on big, real-world data. Although sex difference is a consistent and recognized feature of ASD, additional clinical correlates could help to identify potential disease subgroups, based on sex and age. MATERIALS AND METHODS: We performed a systematic comorbidity analysis on 1860 groups of comorbidities exploring all spectrum of known disease, in 59 140 individuals (11 440 females) with ASD from 4 age groups. We explored ASD sex differences in 2 independent real-world datasets, across all potential comorbidities by comparing (1) females with ASD vs males with ASD and (2) females with ASD vs females without ASD. RESULTS: We identified 27 different comorbidities that appeared significantly more frequently in females with ASD. The comorbidities were mostly neurological (eg, epilepsy, odds ratio [OR] > 1.8, 3-18 years of age), congenital (eg, chromosomal anomalies, OR > 2, 3-18 years of age), and mental disorders (eg, intellectual disability, OR > 1.7, 6-18 years of age). Novel comorbidities included endocrine metabolic diseases (eg, failure to thrive, OR = 2.5, ages 0-2), digestive disorders (gastroesophageal reflux disease: OR = 1.7, 6-11 years of age; and constipation: OR > 1.6, 3-11 years of age), and sense organs (strabismus: OR > 1.8, 3-18 years of age). DISCUSSION: A multi-PheWAS intersection approach on real-world data as presented in this study uniquely contributes to the growing body of research regarding sex-based comorbidity analysis in ASD population. CONCLUSIONS: Our findings provide insights into female-enriched ASD comorbidities that are potentially important in diagnosis, as well as the identification of distinct comorbidity patterns influencing anticipatory treatment or referrals. The code is publicly available (https://github.com/hms-dbmi/sexDifferenceInASD). Alba Gutiérrez-Sacristán, Carlos Sáez 0001, Carlos De Niz, Niloofar Jalali, Thomas N. Desain, Ranjay Kumar, Joany M. Zachariasse, Kathe P. Fox, Nathan P. Palmer, Isaac S. Kohane, Paul Avillach |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai |
J. Biomed. Informatics | 12 |
| 2021 | GenoPheno: cataloging large-scale phenotypic and next-generation sequencing data within human datasetsabstractPrecision medicine promises to revolutionize treatment, shifting therapeutic approaches from the classical one-size-fits-all to those more tailored to the patient's individual genomic profile, lifestyle and environmental exposures. Yet, to advance precision medicine's main objective-ensuring the optimum diagnosis, treatment and prognosis for each individual-investigators need access to large-scale clinical and genomic data repositories. Despite the vast proliferation of these datasets, locating and obtaining access to many remains a challenge. We sought to provide an overview of available patient-level datasets that contain both genotypic data, obtained by next-generation sequencing, and phenotypic data-and to create a dynamic, online catalog for consultation, contribution and revision by the research community. Datasets included in this review conform to six specific inclusion parameters that are: (i) contain data from more than 500 human subjects; (ii) contain both genotypic and phenotypic data from the same subjects; (iii) include whole genome sequencing or whole exome sequencing data; (iv) include at least 100 recorded phenotypic variables per subject; (v) accessible through a website or collaboration with investigators and (vi) make access information available in English. Using these criteria, we identified 30 datasets, reviewed them and provided results in the release version of a catalog, which is publicly available through a dynamic Web application and on GitHub. Users can review as well as contribute new datasets for inclusion (Web: https://avillachlab.shinyapps.io/genophenocatalog/; GitHub: https://github.com/hms-dbmi/GenoPheno-CatalogShiny). Alba Gutiérrez-Sacristán, Carlos De Niz, Cartik Kothari, Sek Won Kong, Kenneth D. Mandl, Paul Avillach |
Briefings Bioinform. | 1 |
| 2021 | Finding commonalities in rare diseases through the undiagnosed diseases networkabstractOBJECTIVE: When studying any specific rare disease, heterogeneity and scarcity of affected individuals has historically hindered investigators from discerning on what to focus to understand and diagnose a disease. New nongenomic methodologies must be developed that identify similarities in seemingly dissimilar conditions. MATERIALS AND METHODS: This observational study analyzes 1042 patients from the Undiagnosed Diseases Network (2015-2019), a multicenter, nationwide research study using phenotypic data annotated by specialized staff using Human Phenotype Ontology terms. We used Louvain community detection to cluster patients linked by Jaccard pairwise similarity and 2 support vector classifier to assign new cases. We further validated the clusters' most representative comorbidities using a national claims database (67 million patients). RESULTS: Patients were divided into 2 groups: those with symptom onset before 18 years of age (n = 810) and at 18 years of age or older (n = 232) (average symptom onset age: 10 [interquartile range, 0-14] years). For 810 pediatric patients, we identified 4 statistically significant clusters. Two clusters were characterized by growth disorders, and developmental delay enriched for hypotonia presented a higher likelihood of diagnosis. Support vector classifier showed 0.89 balanced accuracy (0.83 for Human Phenotype Ontology terms only) on test data. DISCUSSIONS: To set the framework for future discovery, we chose as our endpoint the successful grouping of patients by phenotypic similarity and provide a classification tool to assign new patients to those clusters. CONCLUSION: This study shows that despite the scarcity and heterogeneity of patients, we can still find commonalities that can potentially be harnessed to uncover new insights and targets for therapy. Josephine Yates, Alba Gutiérrez-Sacristán, Vianney Jouhet, Kimberly Leblanc, Cecilia Esteves, Thomas N. Desain, Nick Benik, Jason Stedman, Nathan P. Palmer, Guillaume Mellon, Isaac S. Kohane, Paul Avillach |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Scalability and cost-effectiveness analysis of whole genome-wide association studies on Google Cloud Platform and Amazon Web ServicesabstractOBJECTIVE: Advancements in human genomics have generated a surge of available data, fueling the growth and accessibility of databases for more comprehensive, in-depth genetic studies. METHODS: We provide a straightforward and innovative methodology to optimize cloud configuration in order to conduct genome-wide association studies. We utilized Spark clusters on both Google Cloud Platform and Amazon Web Services, as well as Hail (http://doi.org/10.5281/zenodo.2646680) for analysis and exploration of genomic variants dataset. RESULTS: Comparative evaluation of numerous cloud-based cluster configurations demonstrate a successful and unprecedented compromise between speed and cost for performing genome-wide association studies on 4 distinct whole-genome sequencing datasets. Results are consistent across the 2 cloud providers and could be highly useful for accelerating research in genetics. CONCLUSIONS: We present a timely piece for one of the most frequently asked questions when moving to the cloud: what is the trade-off between speed and cost? Inès Krissaane, Carlos De Niz, Alba Gutiérrez-Sacristán, Gabor Korodi, Nneka Ede, Ranjay Kumar, Jessica Lyons, Arjun K. Manrai, Chirag J. Patel, Isaac S. Kohane, Paul Avillach |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Comorbidity4j: a tool for interactive analysis of disease comorbidities over large patient datasetsabstractSUMMARY: Pushed by the growing availability of Electronic Health Records for data mining, the identification of relevant patterns of co-occurring diseases over a population of individuals-referred to as comorbidity analysis-has become a common practice due to its great impact on life expectancy, quality of life and healthcare costs. In this scenario, the availability of scalable, easy-to-use software frameworks tailored to support the study of comorbidities over large datasets of patients is essential. We introduce Comorbidity4j, an open-source Java tool to perform systematic analyses of comorbidities by generating interactive Web visualizations to explore and refine results. Comorbidity4j processes user-provided clinical data by identifying significant disease co-occurrences and computing a comprehensive set of comorbidity indices. Patients can be stratified by sex, age and user-defined criteria. Comorbidity4j supports the analysis of the temporal directionality and the sex ratio of diseases. The incremental upload and validation of clinical input data and the customization of comorbidity analyses are performed by an interactive Web interface. With a Web browser, the results of such analyses can be filtered with respect to comorbidity indexes and disease names and explored by means of heat maps and network charts of disease associations. Comorbidity4j is optimized to efficiently process large datasets of clinical data. Besides a software tool for local execution, we provide Comorbidity4j as a Web service to enable users to perform online comorbidity analyses. AVAILABILITY AND IMPLEMENTATION: Doc: http://comorbidity4j.readthedocs.io/; Source code: https://github.com/fra82/comorbidity4j, Web tool: http://comorbidity.eu/comorbidity4web/. Francesco Ronzano, Alba Gutiérrez-Sacristán, Laura Inés Furlong |
Bioinform. | 2 |
| 2018 | Rcupcake: an R package for querying and analyzing biomedical data through the BD2K PIC-SURE RESTful APIabstractMotivation: In the era of big data and precision medicine, the number of databases containing clinical, environmental, self-reported and biochemical variables is increasing exponentially. Enabling the experts to focus on their research questions rather than on computational data management, access and analysis is one of the most significant challenges nowadays. Results: We present Rcupcake, an R package that contains a variety of functions for leveraging different databases through the BD2K PIC-SURE RESTful API and facilitating its query, analysis and interpretation. The package offers a variety of analysis and visualization tools, including the study of the phenotype co-occurrence and prevalence, according to multiple layers of data, such as phenome, exposome or genome. Availability and implementation: The package is implemented in R and is available under Mozilla v2 license from GitHub (https://github.com/hms-dbmi/Rcupcake). Two reproducible case studies are also available (https://github.com/hms-dbmi/Rcupcake-case-studies/blob/master/SSCcaseStudy_v01.ipynb, https://github.com/hms-dbmi/Rcupcake-case-studies/blob/master/NHANEScaseStudy_v01.ipynb). Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Alba Gutiérrez-Sacristán, Romain Guedj, Gabor Korodi, Jason Stedman, Laura Inés Furlong, Chirag J. Patel, Isaac S. Kohane, Paul Avillach |
Bioinform. | 1 |
| 2018 | comoRbidity: an R package for the systematic analysis of disease comorbiditiesabstractMotivation: The study of comorbidities is a major priority due to their impact on life expectancy, quality of life and healthcare cost. The availability of electronic health records (EHRs) for data mining offers the opportunity to discover disease associations and comorbidity patterns from the clinical history of patients gathered during routine medical care. This opens the need for analytical tools for detection of disease comorbidities, including the investigation of their underlying genetic basis. Results: We present comoRbidity, an R package aimed at providing a systematic and comprehensive analysis of disease comorbidities from both the clinical and molecular perspectives. comoRbidity leverages from (i) user provided clinical data from EHR databases (the clinical comorbidity analysis) and (ii) genotype-phenotype information of the diseases under study (the molecular comorbidity analysis) for a comprehensive analysis of disease comorbidities. The clinical comorbidity analysis enables identifying significant disease comorbidities from clinical data, including sex and age stratification and temporal directionality analyses, while the molecular comorbidity analysis supports the generation of hypothesis on the underlying mechanisms of the disease comorbidities by exploring shared genes among disorders. The open-source comoRbidity package is a software tool aimed at expediting the integrative analysis of disease comorbidities by incorporating several analytical and visualization functions. Availability and implementation: https://bitbucket.org/ibi_group/comorbidity. Supplementary information: Supplementary data are available at Bioinformatics online. Alba Gutiérrez-Sacristán, Àlex Bravo, Alexia Giannoula, Miguel Angel Mayer, Ferran Sanz, Laura Inés Furlong |
Bioinform. | 1 |
| 2017 | psygenet2r: a R/Bioconductor package for the analysis of psychiatric disease genesabstractMOTIVATION: Psychiatric disorders have a great impact on morbidity and mortality. Genotype-phenotype resources for psychiatric diseases are key to enable the translation of research findings to a better care of patients. PsyGeNET is a knowledge resource on psychiatric diseases and their genes, developed by text mining and curated by domain experts. RESULTS: We present psygenet2r, an R package that contains a variety of functions for leveraging PsyGeNET database and facilitating its analysis and interpretation. The package offers different types of queries to the database along with variety of analysis and visualization tools, including the study of the anatomical structures in which the genes are expressed and gaining insight of gene's molecular function. Psygenet2r is especially suited for network medicine analysis of psychiatric disorders. AVAILABILITY AND IMPLEMENTATION: The package is implemented in R and is available under MIT license from Bioconductor (http://bioconductor.org/packages/release/bioc/html/psygenet2r.html). CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alba Gutiérrez-Sacristán, Carles Hernandez-Ferrer, Juan R. González, Laura Inés Furlong |
Bioinform. | 1 |
| 2015 | PsyGeNET: a knowledge platform on psychiatric disorders and their genesabstractUNLABELLED: PsyGeNET (Psychiatric disorders and Genes association NETwork) is a knowledge platform for the exploratory analysis of psychiatric diseases and their associated genes. PsyGeNET is composed of a database and a web interface supporting data search, visualization, filtering and sharing. PsyGeNET integrates information from DisGeNET and data extracted from the literature by text mining, which has been curated by domain experts. It currently contains 2642 associations between 1271 genes and 37 psychiatric disease concepts. In its first release, PsyGeNET is focused on three psychiatric disorders: major depression, alcohol and cocaine use disorders. PsyGeNET represents a comprehensive, open access resource for the analysis of the molecular mechanisms underpinning psychiatric disorders and their comorbidities. AVAILABILITY AND IMPLEMENTATION: The PysGeNET platform is freely available at http://www.psygenet.org/. The PsyGeNET database is made available under the Open Database License (http://opendatacommons.org/licenses/odbl/1.0/). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alba Gutiérrez-Sacristán, Solène Grosdidier, Olga Valverde, Marta Torrens, Àlex Bravo, Janet Piñero González, Ferran Sanz, Laura Inés Furlong |
Bioinform. | 1 |