Mattia Prosperi

dblp:98/1006 · also Mattia C. F. Prosperi · DBLP profile ↗
← Back
47ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0002-9021-5595ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 44 · 9 first-author · 25 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSystems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Multi-site analysis of COVID-19 and new-onset diabetes reveals need for improved sensitivity of EHR-based COVID-19 phenotypes - a DiCAYA Network analysis
abstract
OBJECTIVE: We discuss implications of potential ascertainment biases for studies examining diabetes risk following SARS-CoV-2 infection using electronic health records (EHRs). We quantitatively explore sensitivity of results to misclassification of COVID-19 status using data from the U.S.-based Diabetes in Children, Adolescents and Young Adults (DiCAYA) Network on children (≤17 years) and young adults (18-44 years). MATERIALS AND METHODS: In our retrospective case study from the DiCAYA Network, SARS-CoV-2 was identified using labs and diagnoses from June 1, 2020 to December 31, 2021. Patients were followed through December 31, 2022 for new diabetes diagnoses. Sites examined incident diabetes by COVID-19 status using Cox proportional hazards models. Results were pooled in meta-analyses. A bias analysis examined potential impact of COVID-19 misclassification scenarios on results, guided by hypotheses that sensitivity would be <50% and would be higher among those who developed diabetes. RESULTS: Prevalence of documented COVID-19 was low overall and variable across sites (children: 4.4%-7.7%, young adults: 6.2%-22.7%). Individuals with documented COVID-19 were at higher risk of incident diabetes compared to those with no documented infection, but results were heterogeneous across sites. Findings were highly sensitive to COVID-19 misclassification assumptions. Observed results could be biased away from the null under several differential misclassification scenarios. DISCUSSION: Although EHR-based documentation of COVID-19 was associated with incident diabetes, COVID-19 phenotypes likely had low sensitivity, with considerable variation across sites. Misclassification assumptions strongly impacted interpretation of results. CONCLUSION: Given the potential for low phenotype sensitivity and misclassification, caution is warranted when interpreting analyses of COVID-19 and incident diabetes using clinical or administrative databases.
Lorna E. Thorpe, Jasmin Divers, Annemarie Hirsch, Brian S. Schwartz, Jihad S. Obeid, Angela Liese, Tessa L. Crume, Anna Bellatorre, Jiang Bian 0001, Yi Guo 0005, Sarah Bost, Tianchen Lyu, Matthew T. Mefford, Matt Zhou, Eva Lustigova, Levon Utidjian, Mitchell Maltenfort, Patrick Hanley, Meda E. Pavkov, Marc B. Rosenman, Andrea R. Titus, L. Charles Bailey, Christopher B. Forrest, Mitch Maltenfort, Amy Shah, Eneida A. Mendonça, G. Todd Alonso, Sara J. Deakyne Davies, H. Timothy Bunnell, Anne Kazak, Melody Kitzmiller, Manmohan Kamboj, Dimitri A. Christakis, Daksha Ranade, Annemarie G. Hirsch, Joseph J. Dewalle, H. Lester Kirchner, Meredith Lewis, Dione G. Mercer, Cara M. Nordberg, Amy Poissant, Brian E. Dixon, Shaun J. Grannis, Katie Allen, Anna Roberts, Nimish Valvi, Jeff Warvel, Ashley Wiensch, Tamara S. Hannon, Kristi Reynolds, John Chang, Don McCarthy, Rong Wei, Marc Rosenman, George Lales, Anthony Wong, Allison Zelinski, Yuan Luo 0001, Mark Weiner, Pedro Rivera, Thomas Carton, Elizabeth Nauman, Harold P. Lehmann, Meredith Akerman, Rebecca Anthopolos, Stefanie Bendik, Sarah Conderino, Andrew Fair, Jessica Guillaume, Shahidul Islam, Alan Jacobson, David C. Lee, Chinyere Okpara, Anand Rajan, Andrea Titus, Dana Dabelea, Theresa Anderson, Rebecca Conway, Toan Ong, Jack Pattee, Shawna Burgett, Elizabeth Shenkman, William T. Donahoo, William R. Hogan, Piaopiao Li, Mattia Prosperi, Yonghui Wu 0001, Angela D. Liese, Lisa Knight, Caroline Rudisill, Jessica Stucker, Deborah Bowlby, Elaine Apperson, Alex Ewing, Giuseppina Imperatore, Deborah Rolka, Ibrahim Zaganjor
J. Am. Medical Informatics Assoc.92
2025 Comparative Evaluation of Clinical Large Language Models and Machine Learning to Predict Antimicrobial Resistance in Hospital-Onset Sepsis
Scott A. Cohen, Jiang Bian 0001, Christina Boucher 0001, Yonghui Wu 0001, Mattia Prosperi
AIME (1)6
2025 Enabling Transcriptome Classification in Micro-cohorts with Pathway-Anchoring and Single-Subject Studies
Mahdieh Shabanian, Nima Pouladi, Liam S. Wilson, Mattia Prosperi, Yves A. Lussier
AIME (2)4
2025 SARITA: a large language model for generating the S1 subunit of the SARS-CoV-2 spike protein
abstract
BACKGROUND: The COVID-19 pandemic has caused over 776 million infections and 7 million deaths globally between December 2019 and November 2024. Since the emergence of the original Wuhan strain, SARS-CoV-2 has evolved into multiple variants-including Alpha, Delta, and Omicron-primarily through mutations in the Spike glycoprotein. The S1 subunit, which binds the human angiotensin-converting enzyme 2 (ACE2) receptor, mutates frequently and plays a key role in infectivity and immune escape, while the more conserved S2 subunit mediates membrane fusion. Anticipating future mutations is essential for guiding vaccine design and therapeutic strategies. Generative Large Language Models (LLMs) have shown promise in protein sequence modeling due to their capacity to produce realistic and functional synthetic sequences. Here, we introduce SARITA, a GPT-3-based LLM with up to 1.2 billion parameters, fine-tuned via continual learning on the protein model RITA trained on 107 017 high-quality SARS-CoV-2 Spike sequences (up to March 1st 2021) to generate high-quality synthetic SARS-CoV-2 Spike S1 subunits. RESULTS: SARITA is able to generate realistic, full-length synthetic S1 subunits starting from a 14-amino-acid prompt. When evaluated on unseen sequences collected between March 2021 and November 2023-including major Variants of Concern (VOCs) such as Delta and Omicron, and Variants of Interest such as Iota-SARITA outperforms baseline and state-of-the-art LLMs in terms of sequence quality, biological plausibility, and similarity to real-world viral evolution. SARITA generates high-quality sequences in over 97% of cases, with markedly lower False Mutation Rate and higher similarity scores (PAM30, Levenshtein distance) compared to alternative approaches. It also accurately reproduces key mutations characteristic of future variants-such as L212I, R158L, T95P, and E406K-which were not present in the training data but emerged later in VOCs like Omicron and Delta. Structure-based analysis confirms the functional plausibility of these substitutions, with ΔΔG values within experimentally supported thresholds for ACE2 and antibody binding. Furthermore, SARITA anticipates immune-evasive mutations and accurately captures the positional and statistical distribution of mutations found in post- March 1st 2021 variants, highlighting its potential as a predictive tool for viral evolution. CONCLUSION: These results indicate the potential of SARITA to predict future SARS-CoV-2 S1 evolution, potentially aiding in the development of adaptable vaccines and treatments.
Simone Rancati, Giovanna Nicora, Laura Bergomi, Tommaso Mario Buonocore, Daniel M. Czyz, Enea Parimbelli, Riccardo Bellazzi, Marco Salemi, Mattia Prosperi, Simone Marini
Briefings Bioinform.9
2025 Leveraging large language models to predict antibiotic resistance in Mycobacterium tuberculosis
abstract
MOTIVATION: Antibiotic resistance in Mycobacterium tuberculosis (MTB) poses a significant challenge to global public health. Rapid and accurate prediction of antibiotic resistance can inform treatment strategies and mitigate the spread of resistant strains. In this study, we present a novel approach leveraging large language models (LLMs) to predict antibiotic resistance in MTB (LLMTB). Our model is trained and evaluated on genomic data from 12 185 CRyPTIC isolates and their associated resistance profiles, utilizing natural language processing techniques to capture patterns and mutations linked to resistance. The model's architecture integrates state-of-the-art transformer-based LLMs, enabling the analysis of complex genomic sequences and the extraction of critical features relevant to antibiotic resistance. RESULTS: We evaluate our model's performance using a comprehensive dataset of MTB strains, demonstrating its ability to achieve high performance in predicting resistance to various antibiotics. Unlike traditional machine learning methods, fine-tuning or few-shot learning opens avenues for LLMs to adapt to new or emerging drugs, thereby reducing reliance on extensive data curation. Beyond predictive accuracy, LLMTB uncovers deeper biological insights, identifying critical genes, intergenic regions, and novel resistance mechanisms. This method marks a transformative shift in resistance prediction and offers significant potential for enhancing diagnostic capabilities and guiding personalized treatment plans, ultimately contributing to the global effort to combat tuberculosis and antibiotic resistance. AVAILABILITY AND IMPLEMENTATION: All source code is publicly available at https://github.com/ctestagrose/LLMTB.
Conrad Testagrose, Sakshi Pandey, Mohammadali Serajian, Simone Marini, Mattia Prosperi, Christina Boucher 0001
Bioinform.5
2025 Variational temporal deconfounder network for individualized treatment effect estimation with longitudinal observational data
Yu Huang 0018, Yuxi Liu 0003, Xing He 0003, Jingchuan Guo, Mattia Prosperi, Jiang Bian 0001
J. Biomed. Informatics6
2024 Predicting HIV Diagnosis Among Emerging Adults Using Electronic Health Records and Health Survey Data in All of Us Research Program
abstract
The global decline in HIV incidence has not been mirrored in the United States, where young adults (ages 18-29) continue to account for a significant portion of new infections. In this study, we leverage the All of Us (AoU) Research Program's extensive electronic health records (EHRs) and health survey data to develop machine learning models capable of predicting HIV diagnoses at least three months before clinical identification. Among various models tested, the Support Vector Machine (SVM) model demonstrated a balanced performance, integrating clinically relevant features with robust predictive accuracy (AUC = 0.91). Risky drinking behaviors emerged as consistent top predictors across models, highlighting the importance of targeted interventions in this age group. Our findings underscore the potential of predictive analytics in enhancing HIV prevention strategies and informing public health efforts aimed at reducing HIV transmission among emerging adults.
Balu Bhasuran, Mattia Prosperi, Karen MacDonell, Sylvie Naar, Zhe He 0001
BIBM3
2024 Forecasting dominance of SARS-CoV-2 lineages by anomaly detection using deep AutoEncoders
abstract
The COVID-19 pandemic is marked by the successive emergence of new SARS-CoV-2 variants, lineages, and sublineages that outcompete earlier strains, largely due to factors like increased transmissibility and immune escape. We propose DeepAutoCoV, an unsupervised deep learning anomaly detection system, to predict future dominant lineages (FDLs). We define FDLs as viral (sub)lineages that will constitute >10% of all the viral sequences added to the GISAID, a public database supporting viral genetic sequence sharing, in a given week. DeepAutoCoV is trained and validated by assembling global and country-specific data sets from over 16 million Spike protein sequences sampled over a period of ~4 years. DeepAutoCoV successfully flags FDLs at very low frequencies (0.01%-3%), with median lead times of 4-17 weeks, and predicts FDLs between ~5 and ~25 times better than a baseline approach. For example, the B.1.617.2 vaccine reference strain was flagged as FDL when its frequency was only 0.01%, more than a year before it was considered for an updated COVID-19 vaccine. Furthermore, DeepAutoCoV outputs interpretable results by pinpointing specific mutations potentially linked to increased fitness and may provide significant insights for the optimization of public health 'pre-emptive' intervention strategies.
Simone Rancati, Giovanna Nicora, Mattia Prosperi, Riccardo Bellazzi, Marco Salemi, Simone Marini
Briefings Bioinform.3
2024 Scalable de novo classification of antibiotic resistance of Mycobacterium tuberculosis
abstract
MOTIVATION: World Health Organization estimates that there were over 10 million cases of tuberculosis (TB) worldwide in 2019, resulting in over 1.4 million deaths, with a worrisome increasing trend yearly. The disease is caused by Mycobacterium tuberculosis (MTB) through airborne transmission. Treatment of TB is estimated to be 85% successful, however, this drops to 57% if MTB exhibits multiple antimicrobial resistance (AMR), for which fewer treatment options are available. RESULTS: We develop a robust machine-learning classifier using both linear and nonlinear models (i.e. LASSO logistic regression (LR) and random forests (RF)) to predict the phenotypic resistance of Mycobacterium tuberculosis (MTB) for a broad range of antibiotic drugs. We use data from the CRyPTIC consortium to train our classifier, which consists of whole genome sequencing and antibiotic susceptibility testing (AST) phenotypic data for 13 different antibiotics. To train our model, we assemble the sequence data into genomic contigs, identify all unique 31-mers in the set of contigs, and build a feature matrix M, where M[i, j] is equal to the number of times the ith 31-mer occurs in the jth genome. Due to the size of this feature matrix (over 350 million unique 31-mers), we build and use a sparse matrix representation. Our method, which we refer to as MTB++, leverages compact data structures and iterative methods to allow for the screening of all the 31-mers in the development of both LASSO LR and RF. MTB++ is able to achieve high discrimination (F-1 >80%) for the first-line antibiotics. Moreover, MTB++ had the highest F-1 score in all but three classes and was the most comprehensive since it had an F-1 score >75% in all but four (rare) antibiotic drugs. We use our feature selection to contextualize the 31-mers that are used for the prediction of phenotypic resistance, leading to some insights about sequence similarity to genes in MEGARes. Lastly, we give an estimate of the amount of data that is needed in order to provide accurate predictions. AVAILABILITY: The models and source code are publicly available on Github at https://github.com/M-Serajian/MTB-Pipeline.
Mohammadali Serajian, Simone Marini, Jarno Alanko, Noelle R. Noyes, Mattia Prosperi, Christina Boucher 0001
Bioinform.5
2024 DeepDynaForecast: Phylogenetic-informed graph deep learning for epidemic transmission dynamic prediction
abstract
In the midst of an outbreak or sustained epidemic, reliable prediction of transmission risks and patterns of spread is critical to inform public health programs. Projections of transmission growth or decline among specific risk groups can aid in optimizing interventions, particularly when resources are limited. Phylogenetic trees have been widely used in the detection of transmission chains and high-risk populations. Moreover, tree topology and the incorporation of population parameters (phylodynamics) can be useful in reconstructing the evolutionary dynamics of an epidemic across space and time among individuals. We now demonstrate the utility of phylodynamic trees for transmission modeling and forecasting, developing a phylogeny-based deep learning system, referred to as DeepDynaForecast. Our approach leverages a primal-dual graph learning structure with shortcut multi-layer aggregation, which is suited for the early identification and prediction of transmission dynamics in emerging high-risk groups. We demonstrate the accuracy of DeepDynaForecast using simulated outbreak data and the utility of the learned model using empirical, large-scale data from the human immunodeficiency virus epidemic in Florida between 2012 and 2020. Our framework is available as open-source software (MIT license) at github.com/lab-smile/DeepDynaForcast.
Chaoyue Sun, Ruogu Fang, Marco Salemi, Mattia Prosperi, Brittany Rife Magalis
PLoS Comput. Biol.4
2023 The role of health system penetration rate in estimating the prevalence of type 1 diabetes in children and adolescents using electronic health records
abstract
OBJECTIVE: Having sufficient population coverage from the electronic health records (EHRs)-connected health system is essential for building a comprehensive EHR-based diabetes surveillance system. This study aimed to establish an EHR-based type 1 diabetes (T1D) surveillance system for children and adolescents across racial and ethnic groups by identifying the minimum population coverage from EHR-connected health systems to accurately estimate T1D prevalence. MATERIALS AND METHODS: We conducted a retrospective, cross-sectional analysis involving children and adolescents <20 years old identified from the OneFlorida+ Clinical Research Network (2018-2020). T1D cases were identified using a previously validated computable phenotyping algorithm. The T1D prevalence for each ZIP Code Tabulation Area (ZCTA, 5 digits), defined as the number of T1D cases divided by the total number of residents in the corresponding ZCTA, was calculated. Population coverage for each ZCTA was measured using observed health system penetration rates (HSPR), which was calculated as the ratio of residents in the corresponding ZTCA and captured by OneFlorida+ to the overall population in the same ZCTA reported by the Census. We used a recursive partitioning algorithm to identify the minimum required observed HSPR to estimate T1D prevalence and compare our estimate with the reported T1D prevalence from the SEARCH study. RESULTS: Observed HSPRs of 55%, 55%, and 60% were identified as the minimum thresholds for the non-Hispanic White, non-Hispanic Black, and Hispanic populations. The estimated T1D prevalence for non-Hispanic White and non-Hispanic Black were 2.87 and 2.29 per 1000 youth, which are comparable to the reference study's estimation. The estimated prevalence of T1D for Hispanics (2.76 per 1000 youth) was higher than the reference study's estimation (1.48-1.64 per 1000 youth). The standardized T1D prevalence in the overall Florida population was 2.81 per 1000 youth in 2019. CONCLUSION: Our study provides a method to estimate T1D prevalence in children and adolescents using EHRs and reports the estimated HSPRs and prevalence of T1D for different race and ethnicity groups to facilitate EHR-based diabetes surveillance.
Piaopiao Li, Tianchen Lyu, Khalid Alkhuzam, Eliot Spector, William T. Donahoo, Sarah Bost, Yonghui Wu 0001, William R. Hogan, Mattia Prosperi, Desmond A. Schatz, Mark A. Atkinson, Michael J. Haller, Elizabeth Shenkman, Yi Guo 0005, Jiang Bian 0001
J. Am. Medical Informatics Assoc.9
2022 DR-VIDAL - Doubly Robust Variational Information-theoretic Deep Adversarial Learning for Counterfactual Prediction and Treatment Effect Estimation
Shantanu Ghosh, Zheng Feng, Jiang Bian 0001, Kevin Butler, Mattia Prosperi
AMIA5
2022 Complementing ICD Codes with Nurses' Assessment Data Can Improve the Identification of Patients with Hearing and Visual Impairments
Urszula Alina Snigurska, Sarah E. Ser, Mattia Prosperi, Ragnhildur I. Bjarnadottir, Robert James Lucero
AMIA3
2022 An interleaved hardware-accelerated k-mer parser
abstract
Advances in next-generation sequencing (NGS) have not only increased the overall throughput of genomic content (e.g. Illumina NovaSeq up to 6, 000GB), but also provided technology miniaturization (e.g. Oxford Nanopore MinION) enabling real-time, mobile experiments. Single Instruction/Multiple Data (SIMD) hardware acceleration is increasingly used to improve performance of NGS data processing tools, while generic template programming libraries are advantageous to adapt to the fast changes in sequencing and computing platforms. We here present a novel k-mer parser written in ISO C++ that exploits an interleaved, non-sequential, hardware accelerated SIMD implementation within a generic programming framework called libseq. We benchmarked our k-mer parser using different NGS experimental datasets comparing with other two popular k-mer counting tools (DSK and KMC3). On an Intel machine with AVX2 (Quad-Core Intel Core i5 CPU, 32 GB RAM), using simulated in-memory reads, DSK and KMC3 were on average 3. 6x and 1. 03x times slower than our parser across k value ranges of 35-63. On real sequencing experiments, DSK and KMC3 were on average 8. 3x and 28. 8x times slower in file/read parsing and k-mer building than ours. Since our tool uses generic programming, other methods that rely on k-mers (e.g. de Bruijn graphs) can directly benefit from its SIMD acceleration. Our k-mer parser and libseq 2.0 are released under the BSD license and available at https://zenodo.org/record/7015294.
Franco Milicchio, Marco Oliva, Mattia Prosperi
BIBM3
2022 Transmission cluster characteristics of global, regional, and lineage-specific SARS-CoV-2 phylogenies
abstract
The SARS-CoV-2 pandemic has been presenting in periodic waves and multiple variants, of which some dominated over time with increased transmissibility. SARS-CoV-2 is still adapting in the human population, thus it is crucial to understand its evolutionary patterns and dynamics ahead of time. In this work, we analyzed transmission clusters and topology of SARS-CoV-2 phylogenies at the global, regional (North America) and clade-specific (Delta and Omicron) epidemic scales. We used the Nextstrain's nCov open global all-time phylogeny (September 2022, 2,698 strains, 2,243 for North America, 499 for Delta21A, and 543 for Omicron20M), with Nextstrain's clade annotation and Pango lineages. Transmission clusters were identified using Phylopart, DYNAMITE, and several tree imbalance measures were calculated, including staircase-ness, Sackin and Colless index. We found that the phylogenetic clustering profiles of the global epidemic have highest diversification at a distance threshold of 3% (divergence of 10, where the tree sampled median is 49). Phylopart and DYNAMITE clusters moderately-to-highly agree with the Pango nomenclature and the Nextstrain's clade. At the regional and clade-specific scale, transmission clustering profiles tend to flatten and similar clusters are found at distance thresholds between 0.05% and 25%. All the considered phylogenies exhibit high tree imbalance with respect to what expected in random phylogenies, suggesting short infection times and antigenic drift, perhaps due to progressive transition from innate to adaptive immunity in the population.
Mattia Prosperi, Brittany Rife Magalis, Simone Marini, Marco Salemi
BIBM1
2022 Identification of Social and Racial Disparities in Risk of HIV Infection in Florida using Causal AI Methods
abstract
most populous state in the USA-has the highest rates of Human Immunodeficiency Virus (HIV) infections and of unfavorable HIV outcomes, with marked social and racial disparities. In this work, we leveraged large-scale, real-world data, i.e., statewide surveillance records and publicly available data resources encoding social determinants of health (SDoH), to identify social and racial disparities contributing to individuals' risk of HIV infection. We used the Florida Department of Health's Syndromic Tracking and Reporting System (STARS) database (including 100,000+ individuals screened for HIV infection and their partners), and a novel algorithmic fairness assessment method -the Fairness-Aware Causal paThs decompoSition (FACTS)- merging causal inference and artificial intelligence. FACTS deconstructs disparities based on SDoH and individuals' characteristics, and can discover novel mechanisms of inequity, quantifying to what extent they could be reduced by interventions. We paired the deidentified demographic information (age, gender, drug use) of 44,350 individuals in STARS -with non-missing data on interview year, county of residence, and infection status- to eight SDoH, including access to healthcare facilities, % uninsured, median household income, and violent crime rate. Using an expert-reviewed causal graph, we found that the risk of HIV infection for African Americans was higher than for non- African Americans (both in terms of direct and total effect), although a null effect could not be ruled out. FACTS identified several paths leading to racial disparity in HIV risk, including multiple SDoH: education, income, violent crime, drinking, smoking, and rurality.
Mattia Prosperi, Jie Xu 0012, Jingchuan Serena Guo, Jiang Bian 0001, Wei-Han William Chen, Shantrel S. Canidate, Simone Marini
BIBM1
2022 Assessing putative bias in prediction of anti-microbial resistance from real-world genotyping data under explicit causal assumptions
Mattia Prosperi, Christina Boucher 0001, Jiang Bian 0001, Simone Marini
Artif. Intell. Medicine1
2022 Towards routine employment of computational tools for antimicrobial resistance determination via high-throughput sequencing
abstract
Antimicrobial resistance (AMR) is a growing threat to public health and farming at large. In clinical and veterinary practice, timely characterization of the antibiotic susceptibility profile of bacterial infections is a crucial step in optimizing treatment. High-throughput sequencing is a promising option for clinical point-of-care and ecological surveillance, opening the opportunity to develop genotyping-based AMR determination as a possibly faster alternative to phenotypic testing. In the present work, we compare the performance of state-of-the-art methods for detection of AMR using high-throughput sequencing data from clinical settings. We consider five computational approaches based on alignment (AMRPlusPlus), deep learning (DeepARG), k-mer genomic signatures (KARGA, ResFinder) or hidden Markov models (Meta-MARC). We use an extensive collection of 585 isolates with available AMR resistance profiles determined by phenotypic tests across nine antibiotic classes. We show how the prediction landscape of AMR classifiers is highly heterogeneous, with balanced accuracy varying from 0.40 to 0.92. Although some algorithms-ResFinder, KARGA and AMRPlusPlus-exhibit overall better balanced accuracy than others, the high per-AMR-class variance and related findings suggest that: (1) all algorithms might be subject to sampling bias both in data repositories used for training and experimental/clinical settings; and (2) a portion of clinical samples might contain uncharacterized AMR genes that the algorithms-mostly trained on known AMR genes-fail to generalize upon. These results lead us to formulate practical advice for software configuration and application, and give suggestions for future study designs to further develop AMR prediction tools from proof-of-concept to bedside.
Simone Marini, Rodrigo A. Mora, Christina Boucher 0001, Noelle R. Noyes, Mattia Prosperi
Briefings Bioinform.5
2022 Optimizing viral genome subsampling by genetic diversity and temporal distribution (TARDiS) for phylogenetics
abstract
SUMMARY: TARDiS is a novel phylogenetic tool for optimal genetic subsampling. It optimizes both genetic diversity and temporal distribution through a genetic algorithm. AVAILABILITY AND IMPLEMENTATION: TARDiS, along with example datasets and a user manual, is available at https://github.com/smarini/tardis-phylogenetics.
Simone Marini, Carla Mavian, Alberto Riva, Mattia Prosperi, Marco Salemi, Brittany Rife Magalis
Bioinform.4
2022 Finding Overlapping Rmaps via Clustering
abstract
Optical mapping is a method for creating high resolution restriction maps of an entire genome. Optical mapping has been largely automated, and first produces single molecule restriction maps, called Rmaps, which are assembled to generate genome wide optical maps. Since the location and orientation of each Rmap is unknown, the first problem in the analysis of this data is finding related Rmaps, i.e., pairs of Rmaps that share the same orientation and have significant overlap in their genomic location. Although heuristics for identifying related Rmaps exist, they all require quantization of the data which leads to a loss in the precision. In this paper, we propose a Gaussian mixture modelling clustering based method, which we refer to as OMclust, that finds overlapping Rmaps without quantization. Using both simulated and real datasets, we show that OMclust substantially improves the precision (from 48.3% to 73.3%) over the state-of-the art methods while also reducing CPU time and memory consumption. Further, we integrated OMclust into the error correction methods (Elmeri and cOMet) to demonstrate the increase in the performance of these methods. When OMclust was combined with cOMet to error correct Rmap data generated from human DNA, it was able to error correct close to 3x more Rmaps, and reduced the CPU time by more than 35x. Our software is written in C++ and is publicly available under GNU General Public License at https://github.com/kingufl/OMclust.
Kingshuk Mukherjee, Daniel Dole-Muinos, Massimiliano Rossi 0001, Ayomide Ajayi, Mattia Prosperi, Christina Boucher 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2021 Data and Model Biases in Social Media Analyses: A Case Study of COVID-19 Tweets
Pengfei Yin, Yongqiu Li, Xing He 0003, Jingcheng Du, Cui Tao, Yi Guo 0005, Mattia Prosperi, Pierangelo Veltri, Xi Yang 0015, Yonghui Wu 0001, Jiang Bian 0001
AMIA8
2021 Experimental Survey on Power Dissipation of k-mer-Handling Data Structures for Mobile Bioinformatics
abstract
Mobile sequencing technologies, including Oxford Nanopore’s MinION, MklC, and SmidgION, are bringing genomics in the palm of a hand, opening unprecedented new opportunities in clinical and ecological research and translational applications. While sequencers now need only a USB outlet and provide on-board preprocessing (e.g., base calling), the main data analysis phases are tied to an available broadband Internet connection and cloud computing. Yet the ubiquity of tablets and smartphones, along with their increase in computational power, makes them a perfect candidate for enabling mobile/edge mobile bioinformatics analytics. Also, in on site experimental settings tablets and smartphones are preferable to standard computers due to resilience to humidity or spills, and ease of sterilization. We here present an experimental study on power dissipation, aiming at reducing the battery consumption that currently impedes the execution of intensive bioinformatics analytics pipelines. In particular, we investigated the effects of assorted data structures (including hash tables, vectors, balanced trees, tries) employed in some of the most common tasks of a bioinformatics pipeline, the k- mer representation and counting. By employing a thermal camera, we show how different k-mer-handling data structures impact the power dissipation on a smartphone, finding that a cache-oblivious data structure reduces power dissipation (up to 26% better than others). In conclusion, the choice of data structures in mobile bioinformatics must consider not only computing efficiency (e.g., succinct data structures to reduce RAM usage), but also power consumption of mobile devices that heavily rely on batteries in order to function.
Franco Milicchio, Mattia Prosperi
BIBM2
2021 Fast and exact quantification of motif occurrences in biological sequences
abstract
BACKGROUND: Identification of motifs and quantification of their occurrences are important for the study of genetic diseases, gene evolution, transcription sites, and other biological mechanisms. Exact formulae for estimating count distributions of motifs under Markovian assumptions have high computational complexity and are impractical to be used on large motif sets. Approximated formulae, e.g. based on compound Poisson, are faster, but reliable p value calculation remains challenging. Here, we introduce 'motif_prob', a fast implementation of an exact formula for motif count distribution through progressive approximation with arbitrary precision. Our implementation speeds up the exact calculation, usually impractical, making it feasible and posit to substitute currently employed heuristics. RESULTS: We implement motif_prob in both Perl and C+ + languages, using an efficient error-bound iterative process for the exact formula, providing comparison with state-of-the-art tools (e.g. MoSDi) in terms of precision, run time benchmarks, along with a real-world use case on bacterial motif characterization. Our software is able to process a million of motifs (13-31 bases) over genome lengths of 5 million bases within the minute on a regular laptop, and the run times for both the Perl and C+ + code are several orders of magnitude smaller (50-1000× faster) than MoSDi, even when using their fast compound Poisson approximation (60-120× faster). In the real-world use cases, we first show the consistency of motif_prob with MoSDi, and then how the p-value quantification is crucial for enrichment quantification when bacteria have different GC content, using motifs found in antimicrobial resistance genes. The software and the code sources are available under the MIT license at https://github.com/DataIntellSystLab/motif_prob . CONCLUSIONS: The motif_prob software is a multi-platform and efficient open source solution for calculating exact frequency distributions of motifs. It can be integrated with motif discovery/characterization tools for quantifying enrichment and deviation from expected frequency ranges with exact p values, without loss in data processing efficiency.
Mattia Prosperi, Simone Marini, Christina Boucher 0001
BMC Bioinform.1
2021 The application of artificial intelligence and data integration in COVID-19 studies: a scoping review
abstract
OBJECTIVE: To summarize how artificial intelligence (AI) is being applied in COVID-19 research and determine whether these AI applications integrated heterogenous data from different sources for modeling. MATERIALS AND METHODS: We searched 2 major COVID-19 literature databases, the National Institutes of Health's LitCovid and the World Health Organization's COVID-19 database on March 9, 2021. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guideline, 2 reviewers independently reviewed all the articles in 2 rounds of screening. RESULTS: In the 794 studies included in the final qualitative analysis, we identified 7 key COVID-19 research areas in which AI was applied, including disease forecasting, medical imaging-based diagnosis and prognosis, early detection and prognosis (non-imaging), drug repurposing and early drug discovery, social media data analysis, genomic, transcriptomic, and proteomic data analysis, and other COVID-19 research topics. We also found that there was a lack of heterogenous data integration in these AI applications. DISCUSSION: Risk factors relevant to COVID-19 outcomes exist in heterogeneous data sources, including electronic health records, surveillance systems, sociodemographic datasets, and many more. However, most AI applications in COVID-19 research adopted a single-sourced approach that could omit important risk factors and thus lead to biased algorithms. Integrating heterogeneous data for modeling will help realize the full potential of AI algorithms, improve precision, and reduce bias. CONCLUSION: There is a lack of data integration in the AI applications in COVID-19 research and a need for a multilevel AI framework that supports the analysis of heterogeneous data from different sources.
Yi Guo 0005, Yahan Zhang, Tianchen Lyu, Mattia Prosperi, Fei Wang 0001, Hua Xu 0001, Jiang Bian 0001
J. Am. Medical Informatics Assoc.4
2021 Deep propensity network using a sparse autoencoder for estimation of treatment effects
abstract
OBJECTIVE: Drawing causal estimates from observational data is problematic, because datasets often contain underlying bias (eg, discrimination in treatment assignment). To examine causal effects, it is important to evaluate what-if scenarios-the so-called "counterfactuals." We propose a novel deep learning architecture for propensity score matching and counterfactual prediction-the deep propensity network using a sparse autoencoder (DPN-SA)-to tackle the problems of high dimensionality, nonlinear/nonparallel treatment assignment, and residual confounding when estimating treatment effects. MATERIALS AND METHODS: We used 2 randomized prospective datasets, a semisynthetic one with nonlinear/nonparallel treatment selection bias and simulated counterfactual outcomes from the Infant Health and Development Program and a real-world dataset from the LaLonde's employment training program. We compared different configurations of the DPN-SA against logistic regression and LASSO as well as deep counterfactual networks with propensity dropout (DCN-PD). Models' performances were assessed in terms of average treatment effects, mean squared error in precision on effect's heterogeneity, and average treatment effect on the treated, over multiple training/test runs. RESULTS: The DPN-SA outperformed logistic regression and LASSO by 36%-63%, and DCN-PD by 6%-10% across all datasets. All deep learning architectures yielded average treatment effects close to the true ones with low variance. Results were also robust to noise-injection and addition of correlated variables. Code is publicly available at https://github.com/Shantanu48114860/DPN-SAz. DISCUSSION AND CONCLUSION: Deep sparse autoencoders are particularly suited for treatment effect estimation studies using electronic health records because they can handle high-dimensional covariate sets, large sample sizes, and complex heterogeneity in treatment assignments.
Shantanu Ghosh, Jiang Bian 0001, Yi Guo 0005, Mattia Prosperi
J. Am. Medical Informatics Assoc.4
2021 Bagged random causal networks for interventional queries on observational biomedical datasets
Mattia Prosperi, Yi Guo 0005, Jiang Bian 0001
J. Biomed. Informatics1
2020 Multivariate Independence Set Search via Progressive Addition for Conditional Markov Acyclic Networks
abstract
Estimation of conditional dependencies over a joint multivariate probability distribution is a difficult task for big data, e.g. -omics datasets, and it becomes quickly intractable when the number of variables involved grows large. For instance, structure learning in Bayesian networks has super-exponential complexity. Dimension reduction techniques such as principal component analysis can be useful but often transform the original space and can still pose problems with scalability. This substantially limits characterization of joint probability: in general, only pairwise or k-level correlations can be analyzed efficiently. We introduce the Multivariate Independence Set Search via Progressive Addition for Conditional Markov Acyclic Networks (MISS-PACMAN), which operates a greedy selection of jointly independent feature sets from a larger set of covariates, given a variable ordering. The method is non-parametric and can be used with any kernel function. MISS-PACMAN is therefore suitable for heterogeneous big data, as it combines flexibility and scalability. In our tests on both simulated and real-world data, using random forests as kernel, MISS-PACMAN was able to select independence feature sets linearly with the number of features. Further, by combining multiple independence sets, MISS-PACMAN well-approximates the underlying conditional structure of data variables (according to the generating Bayesian network), and compares favorably with other network structure discovery algorithms, such as the Peter-Clark and the fast incremental association Markov blanket.
Mattia Prosperi, Jiang Bian 0001
BIBM1
2020 On the use of clinical based infection data for pandemic case studies
abstract
Epidemiological models are relevant to study and analyze clinical as well as environmental and behavioural data, useful to support health studies. The target is to perform epidemiological analysis producing fast and reliable data access useful to guide prevention and curing processes. This is currently true in pandemic emergency as the current Covid-19 context. Epidemiological models should support in the early identification of pandemic phenomena and in making available data set for studying more accurate drug-based strategy for vaccines or virus containment.In this contribution we present an epidemiology database which integrates different types of clinical data to support research, follow-up and patient monitoring. The idea starts from an hospital databases cooperation integration where virus available data have been integrated to support statistical based studies. Starting from an available database containing 5 years data of infection related viruses (such as HPC, hepatitis) and patient anonymous data, the proposed system provide an integrated data access able to (i) extracting data filtered by means of clinical hypothesis based on patient profiles, environment and drugs and (ii) allowing to build large scale geographical data mappings in order to study correlations among chronic infection diseases and their relations with upcoming pandemic phenomena. Even if the application is in its infancy, the application is relevant with high very important applications.
Giuseppe Tradigo, Patrizia Vizza, Gabriel Gabriele, Maria Mazzitelli, Carlo Torti, Mattia Prosperi, Pietro H. Guzzi, Pierangelo Veltri
BIBM6
2020 Integrating Crowdsourcing and Active Learning for Classification of Work-Life Events from Tweets
Mattia Prosperi, Tianchen Lyu, Yi Guo 0005, Jiang Bian 0001
IEA/AIE2
2020 Portable nanopore analytics: are we there yet?
abstract
MOTIVATION: Oxford Nanopore technologies (ONT) add miniaturization and real time to high-throughput sequencing. All available software for ONT data analytics run on cloud/clusters or personal computers. Instead, a linchpin to true portability is software that works on mobile devices of internet connections. Smartphones' and tablets' chipset/memory/operating systems differ from desktop computers, but software can be recompiled. We sought to understand how portable current ONT analysis methods are. RESULTS: Several tools, from base-calling to genome assembly, were ported and benchmarked on an Android smartphone. Out of 23 programs, 11 succeeded. Recompilation failures included lack of standard headers and unsupported instruction sets. Only DSK, BCALM2 and Kraken were able to process files up to 16 GB, with linearly scaling CPU-times. However, peak CPU temperatures were high. In conclusion, the portability scenario is not favorable. Given the fast market growth, attention of developers to ARM chipsets and Android/iOS is warranted, as well as initiatives to implement mobile-specific libraries. AVAILABILITY AND IMPLEMENTATION: The source code is freely available at: https://github.com/marco-oliva/portable-nanopore-analytics.
Marco Oliva, Franco Milicchio, Kaden King, Grace Benson, Christina Boucher 0001, Mattia Prosperi, Inanç Birol
Bioinform.6
2020 Assessing the practice of data quality evaluation in a national clinical data research network through a systematic scoping review in the era of real-world data
abstract
OBJECTIVE: To synthesize data quality (DQ) dimensions and assessment methods of real-world data, especially electronic health records, through a systematic scoping review and to assess the practice of DQ assessment in the national Patient-centered Clinical Research Network (PCORnet). MATERIALS AND METHODS: We started with 3 widely cited DQ literature-2 reviews from Chan et al (2010) and Weiskopf et al (2013a) and 1 DQ framework from Kahn et al (2016)-and expanded our review systematically to cover relevant articles published up to February 2020. We extracted DQ dimensions and assessment methods from these studies, mapped their relationships, and organized a synthesized summarization of existing DQ dimensions and assessment methods. We reviewed the data checks employed by the PCORnet and mapped them to the synthesized DQ dimensions and methods. RESULTS: We analyzed a total of 3 reviews, 20 DQ frameworks, and 226 DQ studies and extracted 14 DQ dimensions and 10 assessment methods. We found that completeness, concordance, and correctness/accuracy were commonly assessed. Element presence, validity check, and conformance were commonly used DQ assessment methods and were the main focuses of the PCORnet data checks. DISCUSSION: Definitions of DQ dimensions and methods were not consistent in the literature, and the DQ assessment practice was not evenly distributed (eg, usability and ease-of-use were rarely discussed). Challenges in DQ assessments, given the complex and heterogeneous nature of real-world data, exist. CONCLUSION: The practice of DQ assessment is still limited in scope. Future work is warranted to generate understandable, executable, and reusable DQ measures.
Jiang Bian 0001, Tianchen Lyu, Alexander T. Loiacono, Tonatiuh Mendoza Viramontes, Gloria P. Lipori, Yi Guo 0005, Yonghui Wu 0001, Mattia Prosperi, Thomas J. George, Christopher A. Harle, Elizabeth Shenkman, William R. Hogan
J. Am. Medical Informatics Assoc.8
2020 Mining Twitter to assess the determinants of health behavior toward human papillomavirus vaccination in the United States
abstract
OBJECTIVES: The study sought to test the feasibility of using Twitter data to assess determinants of consumers' health behavior toward human papillomavirus (HPV) vaccination informed by the Integrated Behavior Model (IBM). MATERIALS AND METHODS: We used 3 Twitter datasets spanning from 2014 to 2018. We preprocessed and geocoded the tweets, and then built a rule-based model that classified each tweet into either promotional information or consumers' discussions. We applied topic modeling to discover major themes and subsequently explored the associations between the topics learned from consumers' discussions and the responses of HPV-related questions in the Health Information National Trends Survey (HINTS). RESULTS: We collected 2 846 495 tweets and analyzed 335 681 geocoded tweets. Through topic modeling, we identified 122 high-quality topics. The most discussed consumer topic is "cervical cancer screening"; while in promotional tweets, the most popular topic is to increase awareness of "HPV causes cancer." A total of 87 of the 122 topics are correlated between promotional information and consumers' discussions. Guided by IBM, we examined the alignment between our Twitter findings and the results obtained from HINTS. Thirty-five topics can be mapped to HINTS questions by keywords, 112 topics can be mapped to IBM constructs, and 45 topics have statistically significant correlations with HINTS responses in terms of geographic distributions. CONCLUSIONS: Mining Twitter to assess consumers' health behaviors can not only obtain results comparable to surveys, but also yield additional insights via a theory-driven approach. Limitations exist; nevertheless, these encouraging results impel us to develop innovative ways of leveraging social media in the changing health communication landscape.
Hansi Zhang, Christopher Wheldon, Adam G. Dunn, Cui Tao, Jinhai Huo, Rui Zhang 0028, Mattia Prosperi, Yi Guo 0005, Jiang Bian 0001
J. Am. Medical Informatics Assoc.7
2019 Flexible design of multiple metagenomics classification pipelines with UGENE
abstract
SUMMARY: UGENE is a free, open-source, cross-platform bioinformatics software. UGENE deploys pre-defined pipelines and a flexible instrument to design new workflows and visually build multi-step analytics pipelines. The new UGENE v.1.31 release offers graphical, user-friendly wrapping of a number of popular command-line metagenomics classification programs (Kraken, CLARK, DIAMOND), combinable serially and in parallel through the workflow designer, with multiple, customizable reference databases. Ensemble classification voting is available through the WEVOTE algorithm, with augmented output in the form of detailed table reports. Pre-built workflows (which include all steps from data cleaning to summaries) are included with the installation and a tutorial is available on the UGENE website. Further expansion with multiple visualization tools for reports is planned. AVAILABILITY AND IMPLEMENTATION: UGENE is available at http://ugene.net/, implemented in C++ and Qt, and released under GNU General Public License (GPL) version 2.
Rebecca Rose, Olga Golosova, Dmitrii Sukhomlinov, Aleksey Tiunov, Mattia Prosperi
Bioinform.5
2018 Computable Eligibility Criteria through Ontology-driven Data Access: A Case Study of Hepatitis C Virus Trials
Hansi Zhang, Zhe He 0001, Xing He 0003, Yi Guo 0005, David R. Nelson, François Modave, Yonghui Wu 0001, William R. Hogan, Mattia Prosperi, Jiang Bian 0001
AMIA9
2018 Prototyping an Interactive Visualization of Dietary Supplement Knowledge Graph
Xing He 0003, Rui Zhang 0028, Rubina F. Rizvi, Jake Vasilakes, Xi Yang 0015, Yi Guo 0005, Zhe He 0001, Mattia Prosperi, Jiang Bian 0001
BIBM8
2016 On the identification of long non-coding RNAs from RNA-seq
abstract
Long non-coding RNAs (lncRNAs) are molecules more than 200 nucleotides involved in several biological processes. Next Generation Sequencing allows to identify transcripts containing both coding and non-coding RNAs, but no strategies have been identified so far to discover ncRNA (non-coding RNA) biological functions; thus, most of the ncRNA functionalities are still unknown. We propose a new approach to detect putative lncRNAs transcripts starting from an RNA-seq analysis performed by a reference-based assembly. The extracted transcripts are then analyzed to filter out protein transcripts, detecting putative, thus interesting, lncRNAs submitted to biologists for further validations.
Francesca Cristiano, Pierangelo Veltri, Mattia Prosperi, Giuseppe Tradigo
BIBM3
2016 A* fast and scalable high-throughput sequencing data error correction via oligomers
abstract
Next-generation sequencing (NGS) technologies have superseded traditional Sanger sequencing approach in many experimental settings, given their tremendous yield and affordable cost. Nowadays it is possible to sequence any microbial organism or meta-genomic sample within hours, and to obtain a whole human genome in weeks. Nonetheless, NGS technologies are error-prone. Correcting errors is a challenge due to multiple factors, including the data sizes, the machine-specific and non-at-random characteristics of errors, and the error distributions. Errors in NGS experiments can hamper the subsequent data analysis and inference. This work proposes an error correction method based on the de Bruijn graph that permits its execution on Gigabyte-sized data sets using normal desktop/laptop computers, ideal for genome sizes in the Megabase range, e.g. bacteria. The implementation makes extensive use of hashing techniques, and implements an A* algorithm for optimal error correction, minimizing the distance between an erroneous read and its possible replacement with the Needleman-Wunsch score. Our approach outperforms other popular methods both in terms of random access memory usage and computing times.
Franco Milicchio, Iain E. Buchan, Mattia Prosperi
CIBCB3
2014 Using String Metrics to Identify Patient Journeys through Care Pathways
Richard Williams 0001, Iain E. Buchan, Mattia Prosperi, John D. Ainsworth
AMIA3
2014 HErCoOl: High-Throughput Error Correction by Oligomers
abstract
Next-generation sequencing (NGS) technologies are marking the foundations for a new paradigm in genomics and transcriptomics. Nowadays is possible to sequence any microbial organism or meta-genomic sample within hours, and to obtain a whole human genome in less than a month. The sequencing prices are decreasing dramatically, opening to actual personalised medicine. NGS technologies however are error-prone, and correcting errors is a challenge due to multiple factors, including the data sizes (gigabyte scale) and the machine-specific, non-at-random, characteristics of errors and error distributions. Several approaches have been proposed, but yet the problem is a challenge, especially when analysing mixtures of (closely related) species, e.g., highly variable viruses infecting in a host as a swarm, like hepatitis C or human immunodeficiency virus. This work presents a novel error correction algorithm based on k-mer strings with their associated overlap graph, along with an open-source, multi-threaded, implementation. The algorithm, named Her Cool (High-throughput Error Correction by Oligomers), needs minimal tuning, only an overall error rate and -optionally- information about the genome sizes. Her Cool was compared against other state-of-the art methods, using empirical NGS data obtained with Roche 454 technology, focusing the benchmarks on mixtures of related species. Results show that Her Cool improves significantly over the current methods, and the parallelisation scales well with the size of input NGS genome producing long sequence reads, such as Roche 454 or Ion Torrent. Her Cool provides a fast and efficient error correction of NGS data, especially for mixed samples. Its platform-independent, open-source, multi-threaded implementation assures flexibility for being employed and integrated in any NGS data analysis software.
Franco Milicchio, Mattia Prosperi
CBMS2
2014 Baseline CD4+ T Cell Counts Correlates with HIV-1 Synonymous Rate in HLA-B*5701 Subjects with Different Risk of Disease Progression
abstract
HLA-B*5701 is the host factor most strongly associated with slow HIV-1 disease progression, although risk of progression may vary among patients carrying this allele. The interplay between HIV-1 evolutionary rate variation and risk of progression to AIDS in HLA-B*5701 subjects was studied using longitudinal viral sequences from high-risk progressors (HRPs) and low-risk progressors (LRPs). Posterior distributions of HIV-1 genealogies assuming a Bayesian relaxed molecular clock were used to estimate the absolute rates of nonsynonymous and synonymous substitutions for different set of branches. Rates of viral evolution, as well as in vitro viral replication capacity assessed using a novel phenotypic assay, were correlated with various clinical parameters. HIV-1 synonymous substitution rates were significantly lower in LRPs than HRPs, especially for sets of internal branches. The viral population infecting LRPs was also characterized by a slower increase in synonymous divergence over time. This pattern did not correlate to differences in viral fitness, as measured by in vitro replication capacity, nor could be explained by differences among subjects in T cell activation or selection pressure. Interestingly, a significant inverse correlation was found between baseline CD4+ T cell counts and mean HIV-1 synonymous rate (which is proportional to the viral replication rate) along branches representing viral lineages successfully propagating through time up to the last sampled time point. The observed lower replication rate in HLA-B*5701 subjects with higher baseline CD4+ T cell counts provides a potential model to explain differences in risk of disease progression among individuals carrying this allele.
Melissa M. Norström, Nazle M. Veras, Mattia Prosperi, Jennifer Cook, Wendy Hartogensis, Frederick M. Hecht, Annika C. Karlsoon, Marco Salemi
PLoS Comput. Biol.4
2012 QuRe: software for viral quasispecies reconstruction from next-generation sequencing data
abstract
SUMMARY: Next-generation sequencing (NGS) is an ideal framework for the characterization of highly variable pathogens, with a deep resolution able to capture minority variants. However, the reconstruction of all variants of a viral population infecting a host is a challenging task for genome regions larger than the average NGS read length. QuRe is a program for viral quasispecies reconstruction, specifically developed to analyze long read (>100 bp) NGS data. The software performs alignments of sequence fragments against a reference genome, finds an optimal division of the genome into sliding windows based on coverage and diversity and attempts to reconstruct all the individual sequences of the viral quasispecies--along with their prevalence--using a heuristic algorithm, which matches multinomial distributions of distinct viral variants overlapping across the genome division. QuRe comes with a built-in Poisson error correction method and a post-reconstruction probabilistic clustering, both parameterized on given error rates in homopolymeric and non-homopolymeric regions. AVAILABILITY: QuRe is platform-independent, multi-threaded software implemented in Java. It is distributed under the GNU General Public License, available at https://sourceforge.net/projects/qure/. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mattia Prosperi, Marco Salemi
Bioinform.1
2011 Combinatorial analysis and algorithms for quasispecies reconstruction using next-generation sequencing
abstract
BACKGROUND: Next-generation sequencing (NGS) offers a unique opportunity for high-throughput genomics and has potential to replace Sanger sequencing in many fields, including de-novo sequencing, re-sequencing, meta-genomics, and characterisation of infectious pathogens, such as viral quasispecies. Although methodologies and software for whole genome assembly and genome variation analysis have been developed and refined for NGS data, reconstructing a viral quasispecies using NGS data remains a challenge. This application would be useful for analysing intra-host evolutionary pathways in relation to immune responses and antiretroviral therapy exposures. Here we introduce a set of formulae for the combinatorial analysis of a quasispecies, given a NGS re-sequencing experiment and an algorithm for quasispecies reconstruction. We require that sequenced fragments are aligned against a reference genome, and that the reference genome is partitioned into a set of sliding windows (amplicons). The reconstruction algorithm is based on combinations of multinomial distributions and is designed to minimise the reconstruction of false variants, called in-silico recombinants. RESULTS: The reconstruction algorithm was applied to error-free simulated data and reconstructed a high percentage of true variants, even at a low genetic diversity, where the chance to obtain in-silico recombinants is high. Results on empirical NGS data from patients infected with hepatitis B virus, confirmed its ability to characterise different viral variants from distinct patients. CONCLUSIONS: The combinatorial analysis provided a description of the difficulty to reconstruct a quasispecies, given a determined amplicon partition and a measure of population diversity. The reconstruction algorithm showed good performance both considering simulated data and real data, even in presence of sequencing errors.
Mattia Prosperi, Luciano Prosperi, Alessandro Bruselles, Isabella Abbate, Gabriella Rozera, Donatella Vincenti, Maria Carmela Solmone, Maria Rosaria Capobianchi, Giovanni Ulivi
BMC Bioinform.1
2009 Stochastic modelling of genotypic drug-resistance for human immunodeficiency virus towards long-term combination therapy optimization
abstract
MOTIVATION: Several mathematical models have been investigated for the description of viral dynamics in the human body: HIV-1 infection is a particular and interesting scenario, because the virus attacks cells of the immune system that have a role in the antibody production and its high mutation rate permits to escape both the immune response and, in some cases, the drug pressure. The viral genetic evolution is intrinsically a stochastic process, eventually driven by the drug pressure, dependent on the drug combinations and concentration: in this article the viral genotypic drug resistance onset is the main focus addressed. The theoretical basis is the modelling of HIV-1 population dynamics as a predator-prey system of differential equations with a time-dependent therapy efficacy term, while the viral genome mutation evolution follows a Poisson distribution. The instant probabilities of drug resistance are estimated by means of functions trained from in vitro phenotypes, with a roulette-wheel-based mechanisms of resistant selection. Simulations have been designed for treatments made of one and two drugs as well as for combination antiretroviral therapies. The effect of limited adherence to therapy was also analyzed. Sequential treatment change episodes were also exploited with the aim to evaluate optimal synoptic treatment scenarios. RESULTS: The stochastic predator-prey modelling usefully predicted long-term virologic outcomes of evolved HIV-1 strains for selected antiretroviral therapy combinations. For a set of widely used combination therapies, results were consistent with findings reported in literature and with estimates coming from analysis on a large retrospective data base (EuResist).
Mattia Prosperi, Roberto D'Autilia, Francesca Incardona, Andrea De Luca, Maurizio Zazzi, Giovanni Ulivi
Bioinform.1
2008 A bacterial colony growth framework for collaborative multi-robot localization
abstract
In this paper the multi-robot localization problem is addressed. A new biology-inspired approach is proposed and implemented: the bacterial colony growth framework (BCGF). It takes advantage of the models of species reproduction to provide a suitable framework for carrying on the multi-hypothesis, along with proper policies for both autonomous and collaborative contexts. Collaboration among robots is obtained by exchanging sensory data and their relative distance and orientation. This information is integrated into the framework in such a way that the convergence aptitude is enhanced. Several simulations in different environments have been performed, comparing autonomous and collaborative localization, along with proper statistical analysis for performance assessment.
Andrea Gasparri, Mattia Prosperi
ICRA2
2008 Selecting anti-HIV therapies based on a variety of genomic and clinical factors
abstract
MOTIVATION: Optimizing HIV therapies is crucial since the virus rapidly develops mutations to evade drug pressure. Recent studies have shown that genotypic information might not be sufficient for the design of therapies and that other clinical and demographical factors may play a role in therapy failure. This study is designed to assess the improvement in prediction achieved when such information is taken into account. We use these factors to generate a prediction engine using a variety of machine learning methods and to determine which clinical conditions are most misleading in terms of predicting the outcome of a therapy. RESULTS: Three different machine learning techniques were used: generative-discriminative method, regression with derived evolutionary features, and regression with a mixture of effects. All three methods had similar performances with an area under the receiver operating characteristic curve (AUC) of 0.77. A set of three similar engines limited to genotypic information only achieved an AUC of 0.75. A straightforward combination of the three engines consistently improves the prediction, with significantly better prediction when the full set of features is employed. The combined engine improves on predictions obtained from an online state-of-the-art resistance interpretation system. Moreover, engines tend to disagree more on the outcome of failure therapies than regarding successful ones. Careful analysis of the differences between the engines revealed those mutations and drugs most closely associated with uncertainty of the therapy outcome. AVAILABILITY: The combined prediction engine will be available from July 2008, see http://engine.euresist.org.
Michal Rosen-Zvi, André Altmann, Mattia Prosperi, Ehud Aharoni, Hani Neuvirth, Anders Sönnerborg, Eugen Schülter, Daniel Struck, Yardena Peres, Francesca Incardona, Rolf Kaiser, Maurizio Zazzi, Thomas Lengauer
ISMB3
2007 HIV-1 Coreceptor Usage Prediction via Indexed Local Kernel Smoothing Methods and Grid-Based Multiple Statistical Validation
abstract
Human immunodeficiency virus type 1 (HIV-1) isolates differ in their use of coreceptors to enter target cells. This has important implications for both viral pathogenicity and susceptibility to entry inhibitors under development. Predicting HIV-1 coreceptor usage on the basis of sequence information is a challenging task due to the high variability of the HIV-1 genome. We present an efficient local smoothing kernel method, enhanced with a BLAST-based distance function, implemented by usage of multithreading grid procedures and indexing. Robust validation of the model is achieved through multiple cross-validation, along with statistical comparisons of results for performance assessment.
Iuri Fanti, Mattia Prosperi, Giovanni Ulivi, Alessandro Micarelli
CBMS2
2007 Statistical Comparison of Machine Learning Techniques for Treatment Optimisation of Drug-Resistant HIV-1
abstract
Predicting the in-vivo effect of genotypic drug resistance of Human Immunodeficiency Virus type-1 (HIV-1) on response to antiretroviral therapies represents a major clinical issue. Different machine learning and feature selection methods are applied for the classification of treatment success, based on viral genotype, therapy and derived input features. The robustness of results is assessed through statistical validation. The procedures described are intended to be a general methodology in the challenging context of biology and medical science data mining.
Mattia Prosperi, Giovanni Ulivi, Maurizio Zazzi
CBMS1