Paul Avillach

dblp:36/7332 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-0235-7543ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 VarPPUD: Pinpointing diagnostic variants from sets of prioritized, strong candidate variants
abstract
Rare and ultra-rare genetic conditions are estimated to impact nearly 1 in 17 people worldwide, yet accurately pinpointing the diagnostic variants underlying each of these conditions remains a formidable challenge. Because comprehensive, in vivo functional assessment of all possible genetic variants is infeasible, clinicians instead consider in silico variant pathogenicity predictions to distinguish plausibly disease-causing from benign variants across the genome. However, in the most difficult undiagnosed cases, such as those accepted to the Undiagnosed Diseases Network (UDN), existing pathogenicity predictions cannot reliably discern true etiological variant(s) from other deleterious candidate variants that were prioritized through case- or family-level analyses. Pinpointing the disease-causing variant from a small pool of plausible candidates remains a largely manual effort requiring extensive clinical workups, functional and experimental assays, and eventual identification of genotype- and phenotype-matched individuals. Here, we introduce VarPPUD, a tool trained on prioritized variants from UDN cases, that leverages gene-, amino acid-, and nucleotide-level features to discern pathogenic (disease causative) variants from other damaging or deleterious variants that are unlikely to be confirmed as relevant to the disease. VarPPUD achieves a cross-validated accuracy of 79.3% and precision of 77.5% on a held-out subset of uniquely challenging UDN cases, respectively representing an average 18.6% and 23.4% improvement over nine existing state-of-the-art pathogenicity prediction tools on this task. We validate VarPPUD's ability to discriminate likely from unlikely pathogenic variants using both synthetic data generated via a GAN-based framework and a temporally held-out set of UDN patients evaluated between 2022 and 2024. The model was trained exclusively on data available through 2021 and applied without retraining to the post-2021 cohort, demonstrating strong generalizability to newly accrued cases. Finally, we show how VarPPUD can be probed to evaluate each input feature's importance and contribution toward prediction-an essential step toward understanding the distinct characteristics of newly-uncovered disease-causing variants.
Rui Yin 0002, Alba Gutiérrez-Sacristán, Shilpa Nadimpalli Kobren, Paul Avillach
PLoS Comput. Biol.4
2023 Building a collaborative cloud platform to accelerate heart, lung, blood, and sleep research
abstract
Research increasingly relies on interrogating large-scale data resources. The NIH National Heart, Lung, and Blood Institute developed the NHLBI BioData CatalystⓇ (BDC), a community-driven ecosystem where researchers, including bench and clinical scientists, statisticians, and algorithm developers, find, access, share, store, and compute on large-scale datasets. This ecosystem provides secure, cloud-based workspaces, user authentication and authorization, search, tools and workflows, applications, and new innovative features to address community needs, including exploratory data analysis, genomic and imaging tools, tools for reproducibility, and improved interoperability with other NIH data science platforms. BDC offers straightforward access to large-scale datasets and computational resources that support precision medicine for heart, lung, blood, and sleep conditions, leveraging separately developed and managed platforms to maximize flexibility based on researcher needs, expertise, and backgrounds. Through the NHLBI BioData Catalyst Fellows Program, BDC facilitates scientific discoveries and technological advances. BDC also facilitated accelerated research on the coronavirus disease-2019 (COVID-19) pandemic.
Stanley C. Ahalt, Paul Avillach, Rebecca R. Boyles, Kira Bradford, Steven Cox 0001, Brandi Davis-Dusenbery, Robert L. Grossman, Ashok K. Krishnamurthy 0001, Alisa Manning, Benedict Paten, Anthony Philippakis, Ingrid Borecki, Shu Hui Chen, Jon Kaltman, Sweta Ladwa, Chip Schwartz, Alastair Thomson, Sarah Davis, Alison Leaf, Jessica Lyons, Elizabeth Sheets, Joshua C. Bis, Matthew P. Conomos, Alessandro Culotti, Thomas N. Desain, Jack DiGiovanna, Milan Domazet, Stephanie M. Gogarten, Alba Gutiérrez-Sacristán, Tim Harris 0003, Benjamin D. Heavner, Deepti Jain, Brian O'Connor, Kevin Osborn, Danielle Pillion, Jacob Pleiness, Ken Rice, Garrett Rupp, Arnaud Serret-Larmande, Albert Smith, Jason Stedman, Adrienne Stilp, Teresa Barsanti, John B. Cheadle, Christopher Erdmann, Brandy Farlow, Allie Gartland-Gray, Julie Hayes, Hannah Hiles, Paul Kerr, W. Christopher Lenhardt, Tom Madden, Joanna O. Mieczkowska, Amanda Miller, Patrick Patton, Marcie Rathbun, Stephanie Suber, Joe Asare
J. Am. Medical Informatics Assoc.2
2023 Informative missingness: What can we learn from patterns in missing laboratory data in the electronic health record?
Amelia L. M. Tan, Emily J. Getzen, Meghan Hutch, Zachary H. Strasser, Alba Gutiérrez-Sacristán, Trang T. Le, Arianna Dagliati, Michele Morris, David A. Hanauer, Bertrand Moal, Clara-Lea Bonzel, William Yuan, Lorenzo Chiudinelli, Priyam Das, Harrison G. Zhang, Bruce J. Aronow, Paul Avillach, Gabriel A. Brat, Tianxi Cai, Chuan Hong, William G. La Cava, He Hooi Will Loh, Yuan Luo 0001, Shawn N. Murphy, Kee Yuan Hgiam, Gilbert S. Omenn, Lav P. Patel, Malarkodi J. Samayamuthu, Emily R. Shriver, Zahra Shakeri Hossein Abad, Byorn W. L. Tan, Shyam Visweswaran, Griffin M. Weber, Zongqi Xia, Bertrand Verdy, Qi Long, Danielle L. Mowery, John H. Holmes
J. Biomed. Informatics17
2022 Multi-PheWAS intersection approach to identify sex differences across comorbidities in 59 140 pediatric patients with autism spectrum disorder
abstract
OBJECTIVE: To identify differences related to sex and define autism spectrum disorder (ASD) comorbidities female-enriched through a comprehensive multi-PheWAS intersection approach on big, real-world data. Although sex difference is a consistent and recognized feature of ASD, additional clinical correlates could help to identify potential disease subgroups, based on sex and age. MATERIALS AND METHODS: We performed a systematic comorbidity analysis on 1860 groups of comorbidities exploring all spectrum of known disease, in 59 140 individuals (11 440 females) with ASD from 4 age groups. We explored ASD sex differences in 2 independent real-world datasets, across all potential comorbidities by comparing (1) females with ASD vs males with ASD and (2) females with ASD vs females without ASD. RESULTS: We identified 27 different comorbidities that appeared significantly more frequently in females with ASD. The comorbidities were mostly neurological (eg, epilepsy, odds ratio [OR] > 1.8, 3-18 years of age), congenital (eg, chromosomal anomalies, OR > 2, 3-18 years of age), and mental disorders (eg, intellectual disability, OR > 1.7, 6-18 years of age). Novel comorbidities included endocrine metabolic diseases (eg, failure to thrive, OR = 2.5, ages 0-2), digestive disorders (gastroesophageal reflux disease: OR = 1.7, 6-11 years of age; and constipation: OR > 1.6, 3-11 years of age), and sense organs (strabismus: OR > 1.8, 3-18 years of age). DISCUSSION: A multi-PheWAS intersection approach on real-world data as presented in this study uniquely contributes to the growing body of research regarding sex-based comorbidity analysis in ASD population. CONCLUSIONS: Our findings provide insights into female-enriched ASD comorbidities that are potentially important in diagnosis, as well as the identification of distinct comorbidity patterns influencing anticipatory treatment or referrals. The code is publicly available (https://github.com/hms-dbmi/sexDifferenceInASD).
Alba Gutiérrez-Sacristán, Carlos Sáez 0001, Carlos De Niz, Niloofar Jalali, Thomas N. Desain, Ranjay Kumar, Joany M. Zachariasse, Kathe P. Fox, Nathan P. Palmer, Isaac S. Kohane, Paul Avillach
J. Am. Medical Informatics Assoc.11
2022 SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai
J. Biomed. Informatics56
2022 Multiview Incomplete Knowledge Graph Integration with application to cross-institutional EHR data harmonization
abstract
OBJECTIVE: The growing availability of electronic health records (EHR) data opens opportunities for integrative analysis of multi-institutional EHR to produce generalizable knowledge. A key barrier to such integrative analyses is the lack of semantic interoperability across different institutions due to coding differences. We propose a Multiview Incomplete Knowledge Graph Integration (MIKGI) algorithm to integrate information from multiple sources with partially overlapping EHR concept codes to enable translations between healthcare systems. METHODS: The MIKGI algorithm combines knowledge graph information from (i) embeddings trained from the co-occurrence patterns of medical codes within each EHR system and (ii) semantic embeddings of the textual strings of all medical codes obtained from the Self-Aligning Pretrained BERT (SAPBERT) algorithm. Due to the heterogeneity in the coding across healthcare systems, each EHR source provides partial coverage of the available codes. MIKGI synthesizes the incomplete knowledge graphs derived from these multi-source embeddings by minimizing a spherical loss function that combines the pairwise directional similarities of embeddings computed from all available sources. MIKGI outputs harmonized semantic embedding vectors for all EHR codes, which improves the quality of the embeddings and enables direct assessment of both similarity and relatedness between any pair of codes from multiple healthcare systems. RESULTS: With EHR co-occurrence data from Veteran Affairs (VA) healthcare and Mass General Brigham (MGB), MIKGI algorithm produces high quality embeddings for a variety of downstream tasks including detecting known similar or related entity pairs and mapping VA local codes to the relevant EHR codes used at MGB. Based on the cosine similarity of the MIKGI trained embeddings, the AUC was 0.918 for detecting similar entity pairs and 0.809 for detecting related pairs. For cross-institutional medical code mapping, the top 1 and top 5 accuracy were 91.0% and 97.5% when mapping medication codes at VA to RxNorm medication codes at MGB; 59.1% and 75.8% when mapping VA local laboratory codes to LOINC hierarchy. When trained with 500 labels, the lab code mapping attained top 1 and 5 accuracy at 77.7% and 87.9%. MIKGI also attained best performance in selecting VA local lab codes for desired laboratory tests and COVID-19 related features for COVID EHR studies. Compared to existing methods, MIKGI attained the most robust performance with accuracy the highest or near the highest across all tasks. CONCLUSIONS: The proposed MIKGI algorithm can effectively integrate incomplete summary data from biomedical text and EHR data to generate harmonized embeddings for EHR codes for knowledge graph modeling and cross-institutional translation of EHR codes.
Doudou Zhou, Ziming Gan, Alina Patwari, Everett Neil Rush, Clara-Lea Bonzel, Vidul Ayakulangara Panickan, Chuan Hong, Yuk-Lam Ho, Tianrun A. Cai, Lauren Costa, Victor M. Castro, Shawn N. Murphy, Gabriel A. Brat, Griffin M. Weber, Paul Avillach, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai
J. Biomed. Informatics17
2021 GenoPheno: cataloging large-scale phenotypic and next-generation sequencing data within human datasets
abstract
Precision medicine promises to revolutionize treatment, shifting therapeutic approaches from the classical one-size-fits-all to those more tailored to the patient's individual genomic profile, lifestyle and environmental exposures. Yet, to advance precision medicine's main objective-ensuring the optimum diagnosis, treatment and prognosis for each individual-investigators need access to large-scale clinical and genomic data repositories. Despite the vast proliferation of these datasets, locating and obtaining access to many remains a challenge. We sought to provide an overview of available patient-level datasets that contain both genotypic data, obtained by next-generation sequencing, and phenotypic data-and to create a dynamic, online catalog for consultation, contribution and revision by the research community. Datasets included in this review conform to six specific inclusion parameters that are: (i) contain data from more than 500 human subjects; (ii) contain both genotypic and phenotypic data from the same subjects; (iii) include whole genome sequencing or whole exome sequencing data; (iv) include at least 100 recorded phenotypic variables per subject; (v) accessible through a website or collaboration with investigators and (vi) make access information available in English. Using these criteria, we identified 30 datasets, reviewed them and provided results in the release version of a catalog, which is publicly available through a dynamic Web application and on GitHub. Users can review as well as contribute new datasets for inclusion (Web: https://avillachlab.shinyapps.io/genophenocatalog/; GitHub: https://github.com/hms-dbmi/GenoPheno-CatalogShiny).
Alba Gutiérrez-Sacristán, Carlos De Niz, Cartik Kothari, Sek Won Kong, Kenneth D. Mandl, Paul Avillach
Briefings Bioinform.6
2021 A high-throughput phenotyping algorithm is portable from adult to pediatric populations
abstract
OBJECTIVE: Multimodal automated phenotyping (MAP) is a scalable, high-throughput phenotyping method, developed using electronic health record (EHR) data from an adult population. We tested transportability of MAP to a pediatric population. MATERIALS AND METHODS: Without additional feature engineering or supervised training, we applied MAP to a pediatric population enrolled in a biobank and evaluated performance against physician-reviewed medical records. We also compared performance of MAP at the pediatric institution and the original adult institution where MAP was developed, including for 6 phenotypes validated at both institutions against physician-reviewed medical records. RESULTS: MAP performed equally well in the pediatric setting (average AUC 0.98) as it did at the general adult hospital system (average AUC 0.96). MAP's performance in the pediatric sample was similar across the 6 specific phenotypes also validated against gold-standard labels in the adult biobank. CONCLUSIONS: MAP is highly transportable across diverse populations and has potential for wide-scale use.
Alon Geva, Molei Liu, Vidul Ayakulangara Panickan, Paul Avillach, Tianxi Cai, Kenneth D. Mandl
J. Am. Medical Informatics Assoc.4
2021 Validation of an internationally derived patient severity phenotype to support COVID-19 analytics from electronic health record data
abstract
OBJECTIVE: The Consortium for Clinical Characterization of COVID-19 by EHR (4CE) is an international collaboration addressing coronavirus disease 2019 (COVID-19) with federated analyses of electronic health record (EHR) data. We sought to develop and validate a computable phenotype for COVID-19 severity. MATERIALS AND METHODS: Twelve 4CE sites participated. First, we developed an EHR-based severity phenotype consisting of 6 code classes, and we validated it on patient hospitalization data from the 12 4CE clinical sites against the outcomes of intensive care unit (ICU) admission and/or death. We also piloted an alternative machine learning approach and compared selected predictors of severity with the 4CE phenotype at 1 site. RESULTS: The full 4CE severity phenotype had pooled sensitivity of 0.73 and specificity 0.83 for the combined outcome of ICU admission and/or death. The sensitivity of individual code categories for acuity had high variability-up to 0.65 across sites. At one pilot site, the expert-derived phenotype had mean area under the curve of 0.903 (95% confidence interval, 0.886-0.921), compared with an area under the curve of 0.956 (95% confidence interval, 0.952-0.959) for the machine learning approach. Billing codes were poor proxies of ICU admission, with as low as 49% precision and recall compared with chart review. DISCUSSION: We developed a severity phenotype using 6 code classes that proved resilient to coding variability across international institutions. In contrast, machine learning approaches may overfit hospital-specific orders. Manual chart review revealed discrepancies even in the gold-standard outcomes, possibly owing to heterogeneous pandemic conditions. CONCLUSIONS: We developed an EHR-based severity phenotype for COVID-19 in hospitalized patients and validated it at 12 international sites.
Jeffrey G. Klann, Hossein Estiri, Griffin M. Weber, Bertrand Moal, Paul Avillach, Chuan Hong, Amelia L. M. Tan, Brett K. Beaulieu-Jones, Victor M. Castro, Thomas Maulhardt, Alon Geva, Alberto Malovini, Andrew M. South, Shyam Visweswaran, Michele Morris, Malarkodi J. Samayamuthu, Gilbert S. Omenn, Kee Yuan Ngiam, Kenneth D. Mandl, Martin Boeker, Karen L. Olson, Danielle L. Mowery, Robert W. Follett, David A. Hanauer, Riccardo Bellazzi, Jason H. Moore, Ne-Hooi Will Loh, Douglas S. Bell, Kavishwar B. Wagholikar, Luca Chiovato, Valentina Tibollo, Siegbert Rieg, Anthony L. L. J. Li, Vianney Jouhet, Emily Schriver, Zongqi Xia, Meghan Hutch, Yuan Luo 0001, Isaac S. Kohane, Gabriel A. Brat, Shawn N. Murphy
J. Am. Medical Informatics Assoc.5
2021 Finding commonalities in rare diseases through the undiagnosed diseases network
abstract
OBJECTIVE: When studying any specific rare disease, heterogeneity and scarcity of affected individuals has historically hindered investigators from discerning on what to focus to understand and diagnose a disease. New nongenomic methodologies must be developed that identify similarities in seemingly dissimilar conditions. MATERIALS AND METHODS: This observational study analyzes 1042 patients from the Undiagnosed Diseases Network (2015-2019), a multicenter, nationwide research study using phenotypic data annotated by specialized staff using Human Phenotype Ontology terms. We used Louvain community detection to cluster patients linked by Jaccard pairwise similarity and 2 support vector classifier to assign new cases. We further validated the clusters' most representative comorbidities using a national claims database (67 million patients). RESULTS: Patients were divided into 2 groups: those with symptom onset before 18 years of age (n = 810) and at 18 years of age or older (n = 232) (average symptom onset age: 10 [interquartile range, 0-14] years). For 810 pediatric patients, we identified 4 statistically significant clusters. Two clusters were characterized by growth disorders, and developmental delay enriched for hypotonia presented a higher likelihood of diagnosis. Support vector classifier showed 0.89 balanced accuracy (0.83 for Human Phenotype Ontology terms only) on test data. DISCUSSIONS: To set the framework for future discovery, we chose as our endpoint the successful grouping of patients by phenotypic similarity and provide a classification tool to assign new patients to those clusters. CONCLUSION: This study shows that despite the scarcity and heterogeneity of patients, we can still find commonalities that can potentially be harnessed to uncover new insights and targets for therapy.
Josephine Yates, Alba Gutiérrez-Sacristán, Vianney Jouhet, Kimberly Leblanc, Cecilia Esteves, Thomas N. Desain, Nick Benik, Jason Stedman, Nathan P. Palmer, Guillaume Mellon, Isaac S. Kohane, Paul Avillach
J. Am. Medical Informatics Assoc.12
2020 dbgap2x: an R package to explore and extract data from the database of Genotypes and Phenotypes (dbGaP)
abstract
SUMMARY: Based on the Genomic Data Sharing Policy issued in August 2007, the National Institutes of Health (NIH) has supported several repositories such as the database of Genotypes and Phenotypes (dbGaP). dbGaP is an online repository that provides access to large-scale genetic and phenotypic datasets with more than 1000 studies. However, navigating the website and understanding the relationship between the studies are not easy tasks. Moreover, the decryption of the files is a complex procedure. In this study we propose the dbgap2x R package that covers a broad range of functions for searching dbGaP studies, exploring the characteristics of a study and easily decrypting the files from dbGaP. AVAILABILITY AND IMPLEMENTATION: dbgap2x is an R package with the code available at https://github.com/gversmee/dbgap2x. A containerized version including the package, a Jupyter server and with a Notebook example is available at https://hub.docker.com/r/gversmee/dbgap2x. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Grégoire Versmée, Laura Versmée, Mikaël Dusenne, Niloofar Jalali, Paul Avillach
Bioinform.5
2020 Scalability and cost-effectiveness analysis of whole genome-wide association studies on Google Cloud Platform and Amazon Web Services
abstract
OBJECTIVE: Advancements in human genomics have generated a surge of available data, fueling the growth and accessibility of databases for more comprehensive, in-depth genetic studies. METHODS: We provide a straightforward and innovative methodology to optimize cloud configuration in order to conduct genome-wide association studies. We utilized Spark clusters on both Google Cloud Platform and Amazon Web Services, as well as Hail (http://doi.org/10.5281/zenodo.2646680) for analysis and exploration of genomic variants dataset. RESULTS: Comparative evaluation of numerous cloud-based cluster configurations demonstrate a successful and unprecedented compromise between speed and cost for performing genome-wide association studies on 4 distinct whole-genome sequencing datasets. Results are consistent across the 2 cloud providers and could be highly useful for accelerating research in genetics. CONCLUSIONS: We present a timely piece for one of the most frequently asked questions when moving to the cloud: what is the trade-off between speed and cost?
Inès Krissaane, Carlos De Niz, Alba Gutiérrez-Sacristán, Gabor Korodi, Nneka Ede, Ranjay Kumar, Jessica Lyons, Arjun K. Manrai, Chirag J. Patel, Isaac S. Kohane, Paul Avillach
J. Am. Medical Informatics Assoc.11
2018 ADEPT: An End-to-End System for High-Throughput Pharmacovigilance
Alon Geva, Jason Stedman, Guergana K. Savova, Chen Lin 0002, Shannon F. Manzi, Paul Avillach, Kenneth D. Mandl
AMIA6
2018 Rcupcake: an R package for querying and analyzing biomedical data through the BD2K PIC-SURE RESTful API
abstract
Motivation: In the era of big data and precision medicine, the number of databases containing clinical, environmental, self-reported and biochemical variables is increasing exponentially. Enabling the experts to focus on their research questions rather than on computational data management, access and analysis is one of the most significant challenges nowadays. Results: We present Rcupcake, an R package that contains a variety of functions for leveraging different databases through the BD2K PIC-SURE RESTful API and facilitating its query, analysis and interpretation. The package offers a variety of analysis and visualization tools, including the study of the phenotype co-occurrence and prevalence, according to multiple layers of data, such as phenome, exposome or genome. Availability and implementation: The package is implemented in R and is available under Mozilla v2 license from GitHub (https://github.com/hms-dbmi/Rcupcake). Two reproducible case studies are also available (https://github.com/hms-dbmi/Rcupcake-case-studies/blob/master/SSCcaseStudy_v01.ipynb, https://github.com/hms-dbmi/Rcupcake-case-studies/blob/master/NHANEScaseStudy_v01.ipynb). Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Alba Gutiérrez-Sacristán, Romain Guedj, Gabor Korodi, Jason Stedman, Laura Inés Furlong, Chirag J. Patel, Isaac S. Kohane, Paul Avillach
Bioinform.8
2017 smartAPI: Towards a More Intelligent Network of Web APIs
Amrapali Zaveri, Shima Dastgheib, Chunlei Wu, Trish Whetzel, Ruben Verborgh, Paul Avillach, Gabor Korodi, Raymond Terryn, Kathleen M. Jagodnik, Pedro Assis, Michel Dumontier
ESWC (2)6
2016 An informatics research agenda to support precision medicine: seven key areas
abstract
The recent announcement of the Precision Medicine Initiative by President Obama has brought precision medicine (PM) to the forefront for healthcare providers, researchers, regulators, innovators, and funders alike. As technologies continue to evolve and datasets grow in magnitude, a strong computational infrastructure will be essential to realize PM's vision of improved healthcare derived from personal data. In addition, informatics research and innovation affords a tremendous opportunity to drive the science underlying PM. The informatics community must lead the development of technologies and methodologies that will increase the discovery and application of biomedical knowledge through close collaboration between researchers, clinicians, and patients. This perspective highlights seven key areas that are in need of further informatics research and innovation to support the realization of PM.
Jessica D. Tenenbaum, Paul Avillach, Marge M. Benham-Hutchins, Matthew K. Breitenstein, Erin L. Crowgey, Mark A. Hoffman, Xia Jiang, Subha Madhavan, John E. Mattison, Radhakrishnan Nagarajan, Bisakha Ray, Dmitriy Shin, Shyam Visweswaran, Zhongming Zhao, Robert R. Freimuth
J. Am. Medical Informatics Assoc.2
2015 Translational research platforms integrating clinical and omics data: a review of publicly available solutions
abstract
The rise of personalized medicine and the availability of high-throughput molecular analyses in the context of clinical care have increased the need for adequate tools for translational researchers to manage and explore these data. We reviewed the biomedical literature for translational platforms allowing the management and exploration of clinical and omics data, and identified seven platforms: BRISK, caTRIP, cBio Cancer Portal, G-DOC, iCOD, iDASH and tranSMART. We analyzed these platforms along seven major axes. (1) The community axis regrouped information regarding initiators and funders of the project, as well as availability status and references. (2) We regrouped under the information content axis the nature of the clinical and omics data handled by each system. (3) The privacy management environment axis encompassed functionalities allowing control over data privacy. (4) In the analysis support axis, we detailed the analytical and statistical tools provided by the platforms. We also explored (5) interoperability support and (6) system requirements. The final axis (7) platform support listed the availability of documentation and installation procedures. A large heterogeneity was observed in regard to the capability to manage phenotype information in addition to omics data, their security and interoperability features. The analytical and visualization features strongly depend on the considered platform. Similarly, the availability of the systems is variable. This review aims at providing the reader with the background to choose the platform best suited to their needs. To conclude, we discuss the desiderata for optimal translational research platforms, in terms of privacy, interoperability and technical features.
Vincent Canuel, Bastien Rance, Paul Avillach, Patrice Degoulet, Anita Burgun-Parenthoine
Briefings Bioinform.3
2014 Handling Clinical and Next Generation Sequencing data: new strategies in i2b2 and tranSMART
Shawn N. Murphy, Riccardo Bellazzi, Matteo Gabetta, Paul Avillach, Lori C. Phillips
AMIA4
2013 Preliminary Evaluation of a Literature-based Discovery Method for Explaining Drug Adverse Effects
Dimitar Hristovski, Anita Burgun-Parenthoine, Paul Avillach, Thomas C. Rindflesch
AMIA3
2013 Harmonization process for the identification of medical events in eight European healthcare databases: the experience from the EU-ADR project
abstract
OBJECTIVE: Data from electronic healthcare records (EHR) can be used to monitor drug safety, but in order to compare and pool data from different EHR databases, the extraction of potential adverse events must be harmonized. In this paper, we describe the procedure used for harmonizing the extraction from eight European EHR databases of five events of interest deemed to be important in pharmacovigilance: acute myocardial infarction (AMI); acute renal failure (ARF); anaphylactic shock (AS); bullous eruption (BE); and rhabdomyolysis (RHABD). DESIGN: The participating databases comprise general practitioners' medical records and claims for hospitalization and other healthcare services. Clinical information is collected using four different disease terminologies and free text in two different languages. The Unified Medical Language System was used to identify concepts and corresponding codes in each terminology. A common database model was used to share and pool data and verify the semantic basis of the event extraction queries. Feedback from the database holders was obtained at various stages to refine the extraction queries. MEASUREMENTS: Standardized and age specific incidence rates (IRs) were calculated to facilitate benchmarking and harmonization of event data extraction across the databases. This was an iterative process. RESULTS: The study population comprised overall 19 647 445 individuals with a follow-up of 59 929 690 person-years (PYs). Age adjusted IRs for the five events of interest across the databases were as follows: (1) AMI: 60-148/100 000 PYs; (2) ARF: 3-49/100 000 PYs; (3) AS: 2-12/100 000 PYs; (4) BE: 2-17/100 000 PYs; and (5) RHABD: 0.1-8/100 000 PYs. CONCLUSIONS: The iterative harmonization process enabled a more homogeneous identification of events across differently structured databases using different coding based algorithms. This workflow can facilitate transparent and reproducible event extractions and understanding of differences between databases.
Paul Avillach, Preciosa M. Coloma, Rosa Gini, Martijn J. Schuemie, Fleur Mougin, Jean-Charles Dufour, Giampiero Mazzaglia, Carlo Giaquinto, Carla Fornari, Ron Herings, Mariam Molokhia, Lars Pedersen, Annie Fourrier-Réglat, Marius Fieschi, Miriam C. J. M. Sturkenboom, Johan van der Lei, Antoine Pariente, Gianluca Trifirò
J. Am. Medical Informatics Assoc.1
2013 Design and validation of an automated method to detect known adverse drug reactions in MEDLINE: a contribution from the EU-ADR project
abstract
OBJECTIVES: The aim of this research was to automate the search of publications concerning adverse drug reactions (ADR) by defining the queries used to search MEDLINE and by determining the required threshold for the number of extracted publications to confirm the drug/event association in the literature. METHODS: We defined an approach based on the medical subject headings (MeSH) 'descriptor records' and 'supplementary concept records' thesaurus, using the subheadings 'chemically induced' and 'adverse effects' with the 'pharmacological action' knowledge. An expert-built validation set of true positive and true negative drug/adverse event associations (n=61) was used to validate our method. RESULTS: Using a threshold of three of more extracted publications, the automated search method presented a sensitivity of 90% and a specificity of 100%. For nine different drug/event pairs selected, the recall of the automated search ranged from 24% to 64% and the precision from 93% to 48%. CONCLUSIONS: This work presents a method to find previously established relationships between drugs and adverse events in the literature. Using MEDLINE, following a MeSH approach to filter the signals, is a valid option. Our contribution is available as a web service that will be integrated in the final European EU-ADR project (Exploring and Understanding Adverse Drug Reactions by integrative mining of clinical records and biomedical knowledge) automated system.
Paul Avillach, Jean-Charles Dufour, Gayo Diallo, Francesco Salvo, Michel Joubert, Frantz Thiessard, Fleur Mougin, Gianluca Trifirò, Annie Fourrier-Réglat, Antoine Pariente, Marius Fieschi
J. Am. Medical Informatics Assoc.1
2013 Phenome-Wide Association Studies on a Quantitative Trait: Application to TPMT Enzyme Activity and Thiopurine Therapy in Pharmacogenomics
abstract
Phenome-Wide Association Studies (PheWAS) investigate whether genetic polymorphisms associated with a phenotype are also associated with other diagnoses. In this study, we have developed new methods to perform a PheWAS based on ICD-10 codes and biological test results, and to use a quantitative trait as the selection criterion. We tested our approach on thiopurine S-methyltransferase (TPMT) activity in patients treated by thiopurine drugs. We developed 2 aggregation methods for the ICD-10 codes: an ICD-10 hierarchy and a mapping to existing ICD-9-CM based PheWAS codes. Eleven biological test results were also analyzed using discretization algorithms. We applied these methods in patients having a TPMT activity assessment from the clinical data warehouse of a French academic hospital between January 2000 and July 2013. Data after initiation of thiopurine treatment were analyzed and patient groups were compared according to their TPMT activity level. A total of 442 patient records were analyzed representing 10,252 ICD-10 codes and 72,711 biological test results. The results from the ICD-9-CM based PheWAS codes and ICD-10 hierarchy codes were concordant. Cross-validation with the biological test results allowed us to validate the ICD phenotypes. Iron-deficiency anemia and diabetes mellitus were associated with a very high TPMT activity (p = 0.0004 and p = 0.0015, respectively). We describe here an original method to perform PheWAS on a quantitative trait using both ICD-10 diagnosis codes and biological test results to identify associated phenotypes. In the field of pharmacogenomics, PheWAS allow for the identification of new subgroups of patients who require personalized clinical and therapeutic management.
Antoine Neuraz, Laurent Chouchana, Georgia Malamut, Christine Le Beller, Denis Roche, Philippe Beaune, Patrice Degoulet, Anita Burgun-Parenthoine, Marie-Anne Loriot, Paul Avillach
PLoS Comput. Biol.10
2012 Using Literature-based Discovery to Explain Drug Adverse Effects
Dimitar Hristovski, Anita Burgun-Parenthoine, Paul Avillach, Thomas C. Rindflesch
AMIA3
2012 Automatic Filtering and Substantiation of Drug Safety Signals
abstract
Drug safety issues pose serious health threats to the population and constitute a major cause of mortality worldwide. Due to the prominent implications to both public health and the pharmaceutical industry, it is of great importance to unravel the molecular mechanisms by which an adverse drug reaction can be potentially elicited. These mechanisms can be investigated by placing the pharmaco-epidemiologically detected adverse drug reaction in an information-rich context and by exploiting all currently available biomedical knowledge to substantiate it. We present a computational framework for the biological annotation of potential adverse drug reactions. First, the proposed framework investigates previous evidences on the drug-event association in the context of biomedical literature (signal filtering). Then, it seeks to provide a biological explanation (signal substantiation) by exploring mechanistic connections that might explain why a drug produces a specific adverse reaction. The mechanistic connections include the activity of the drug, related compounds and drug metabolites on protein targets, the association of protein targets to clinical events, and the annotation of proteins (both protein targets and proteins associated with clinical events) to biological pathways. Hence, the workflows for signal filtering and substantiation integrate modules for literature and database mining, in silico drug-target profiling, and analyses based on gene-disease networks and biological pathways. Application examples of these workflows carried out on selected cases of drug safety signals are discussed. The methodology and workflows presented offer a novel approach to explore the molecular mechanisms underlying adverse drug reactions.
Anna Bauer-Mehren, Erik M. van Mulligen, Paul Avillach, María del Carmen Carrascosa, Ricard García-Serna, Janet Piñero González, Pedro Lopes 0002, José Luís Oliveira, Gayo Diallo, Ernst Ahlberg Helgee, Scott Boyer, Jordi Mestres, Ferran Sanz, Jan A. Kors, Laura Inés Furlong
PLoS Comput. Biol.3
2007 A Model for Indexing Medical Documents Combining Statistical and Symbolic Knowledge
Paul Avillach, Michel Joubert, Marius Fieschi
AMIA1