EDBT 2026 Demo / reviewers in the wild / expert
Marco Salemi
dblp:26/4350
· DBLP profile ↗
13ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-0136-2102ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 8 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CholeraSeq: a comprehensive genomic pipeline for cholera surveillance and near real-time outbreak investigationabstractSUMMARY: Next Generation Sequencing is widely deployed in cholera-endemic regions, yet an end-to-end reproducible pipeline that unifies read QC, filtering, reference mapping, variant calling/annotation, recombination screening, and extraction of parsimony informative sites/variant codons, phylogenetic inference for downstream phylodynamic and epidemiological analyses have been lacking, slowing outbreak investigation and public health response. CholeraSeq is a high-throughput genomics pipeline for cholera genomic surveillance. It ingests consensus genomes, short read sequence data, draft assemblies, and scales seamlessly from local to cloud environments. To accelerate epidemiological context placement of new outbreak strains, we provide a curated ready-to-use core genome alignment compiled from public data, enabling flexible, fast, integration of new samples for outbreak investigations. AVAILABILITY AND IMPLEMENTATION: CholeraSeq is freely available on the GitHub platform https://github.com/CERI-KRISP/CholeraSeq. CholeraSeq is implemented in Nextflow with a modular design building upon the nf-core community standards. Massimiliano S. Tagliamonte, Alberto Riva, Monika Moir, Marco Salemi, Cheryl Baxter, Tulio de Oliveira, Carla Mavian, Eduan Wilkinson |
Bioinform. | 5 |
| 2025 | SARITA: a large language model for generating the S1 subunit of the SARS-CoV-2 spike proteinabstractBACKGROUND: The COVID-19 pandemic has caused over 776 million infections and 7 million deaths globally between December 2019 and November 2024. Since the emergence of the original Wuhan strain, SARS-CoV-2 has evolved into multiple variants-including Alpha, Delta, and Omicron-primarily through mutations in the Spike glycoprotein. The S1 subunit, which binds the human angiotensin-converting enzyme 2 (ACE2) receptor, mutates frequently and plays a key role in infectivity and immune escape, while the more conserved S2 subunit mediates membrane fusion. Anticipating future mutations is essential for guiding vaccine design and therapeutic strategies. Generative Large Language Models (LLMs) have shown promise in protein sequence modeling due to their capacity to produce realistic and functional synthetic sequences. Here, we introduce SARITA, a GPT-3-based LLM with up to 1.2 billion parameters, fine-tuned via continual learning on the protein model RITA trained on 107 017 high-quality SARS-CoV-2 Spike sequences (up to March 1st 2021) to generate high-quality synthetic SARS-CoV-2 Spike S1 subunits. RESULTS: SARITA is able to generate realistic, full-length synthetic S1 subunits starting from a 14-amino-acid prompt. When evaluated on unseen sequences collected between March 2021 and November 2023-including major Variants of Concern (VOCs) such as Delta and Omicron, and Variants of Interest such as Iota-SARITA outperforms baseline and state-of-the-art LLMs in terms of sequence quality, biological plausibility, and similarity to real-world viral evolution. SARITA generates high-quality sequences in over 97% of cases, with markedly lower False Mutation Rate and higher similarity scores (PAM30, Levenshtein distance) compared to alternative approaches. It also accurately reproduces key mutations characteristic of future variants-such as L212I, R158L, T95P, and E406K-which were not present in the training data but emerged later in VOCs like Omicron and Delta. Structure-based analysis confirms the functional plausibility of these substitutions, with ΔΔG values within experimentally supported thresholds for ACE2 and antibody binding. Furthermore, SARITA anticipates immune-evasive mutations and accurately captures the positional and statistical distribution of mutations found in post- March 1st 2021 variants, highlighting its potential as a predictive tool for viral evolution. CONCLUSION: These results indicate the potential of SARITA to predict future SARS-CoV-2 S1 evolution, potentially aiding in the development of adaptable vaccines and treatments. Simone Rancati, Giovanna Nicora, Laura Bergomi, Tommaso Mario Buonocore, Daniel M. Czyz, Enea Parimbelli, Riccardo Bellazzi, Marco Salemi, Mattia Prosperi, Simone Marini |
Briefings Bioinform. | 8 |
| 2024 | Sequencing Efforts and Epidemiological Trends: Analyzing SARS-CoV-2 Dynamics Across European NationsabstractThe COVID-19 pandemic has profoundly impacted global health, leading to millions of deaths and overwhelming healthcare systems worldwide. This study investigates the relationship between SARS-CoV-2 sequencing rates and critical epidemiological parameters, such as cases, deaths, and ICU admissions, across 25 European countries from January 2020 to November 2023. By analyzing these relationships, we aim to determine whether sequencing efforts were reactive—in response to epidemiological pressures—or proactive, guided by public health strategies. The analysis used publicly available data from GISAID, OxCGRT, and ECDC, and included weekly aggregation, correlation analysis, and the application of TimeGPT for predictive modeling. Results show that sequencing rates were significantly correlated with ICU admissions, hospitalizations, case numbers, and deaths, though with variability between countries and over different pandemic phases. TimeGPT analysis revealed that sequencing rates were often the most informative feature for predicting future COVID-19 cases in many countries. These findings highlight the potential of sequencing rates to serve as early indicators for severe pandemic outcomes and underscore the importance of context-specific approaches for managing future health crises. Simone Rancati, Daniele Pala, Simone Marini, Marco Salemi, Riccardo Bellazzi, Giovanna Nicora |
BIBM | 4 |
| 2024 | Forecasting dominance of SARS-CoV-2 lineages by anomaly detection using deep AutoEncodersabstractThe COVID-19 pandemic is marked by the successive emergence of new SARS-CoV-2 variants, lineages, and sublineages that outcompete earlier strains, largely due to factors like increased transmissibility and immune escape. We propose DeepAutoCoV, an unsupervised deep learning anomaly detection system, to predict future dominant lineages (FDLs). We define FDLs as viral (sub)lineages that will constitute >10% of all the viral sequences added to the GISAID, a public database supporting viral genetic sequence sharing, in a given week. DeepAutoCoV is trained and validated by assembling global and country-specific data sets from over 16 million Spike protein sequences sampled over a period of ~4 years. DeepAutoCoV successfully flags FDLs at very low frequencies (0.01%-3%), with median lead times of 4-17 weeks, and predicts FDLs between ~5 and ~25 times better than a baseline approach. For example, the B.1.617.2 vaccine reference strain was flagged as FDL when its frequency was only 0.01%, more than a year before it was considered for an updated COVID-19 vaccine. Furthermore, DeepAutoCoV outputs interpretable results by pinpointing specific mutations potentially linked to increased fitness and may provide significant insights for the optimization of public health 'pre-emptive' intervention strategies. Simone Rancati, Giovanna Nicora, Mattia Prosperi, Riccardo Bellazzi, Marco Salemi, Simone Marini |
Briefings Bioinform. | 5 |
| 2024 | DeepDynaForecast: Phylogenetic-informed graph deep learning for epidemic transmission dynamic predictionabstractIn the midst of an outbreak or sustained epidemic, reliable prediction of transmission risks and patterns of spread is critical to inform public health programs. Projections of transmission growth or decline among specific risk groups can aid in optimizing interventions, particularly when resources are limited. Phylogenetic trees have been widely used in the detection of transmission chains and high-risk populations. Moreover, tree topology and the incorporation of population parameters (phylodynamics) can be useful in reconstructing the evolutionary dynamics of an epidemic across space and time among individuals. We now demonstrate the utility of phylodynamic trees for transmission modeling and forecasting, developing a phylogeny-based deep learning system, referred to as DeepDynaForecast. Our approach leverages a primal-dual graph learning structure with shortcut multi-layer aggregation, which is suited for the early identification and prediction of transmission dynamics in emerging high-risk groups. We demonstrate the accuracy of DeepDynaForecast using simulated outbreak data and the utility of the learned model using empirical, large-scale data from the human immunodeficiency virus epidemic in Florida between 2012 and 2020. Our framework is available as open-source software (MIT license) at github.com/lab-smile/DeepDynaForcast. Chaoyue Sun, Ruogu Fang, Marco Salemi, Mattia Prosperi, Brittany Rife Magalis |
PLoS Comput. Biol. | 3 |
| 2023 | ARCA: the interactive database for arbovirus reported cases in the AmericasabstractBACKGROUND: Accurate case report data are essential to understand arbovirus dynamics, including spread and evolution of arboviruses such as Zika, dengue and chikungunya viruses. Giving the multi-country nature of arbovirus epidemics in the Americas, these data are not often accessible or are reported at different time scales (weekly, monthly) from different sources. RESULTS: We developed a publicly available and user-friendly database for arboviral case data in the Americas: ARCA. ARCA is a relational database that is hosted on the ARCA website. Users can interact with the database through the website by submitting queries through the website, which generates displays results and allows users to download these results in different, convenient file formats. Users can choose to view arboviral case data through a table which containscontaining the number of cases for a particular week, a plot, or through a map. CONCLUSION: Our ARCA database is a useful tool for arboviral epidemiology research allowing for complex queries, data visualization, integration, and formatting. Maria V. Meneses, Alberto Riva, Marco Salemi, Carla Mavian |
BMC Bioinform. | 3 |
| 2022 | Transmission cluster characteristics of global, regional, and lineage-specific SARS-CoV-2 phylogeniesabstractThe SARS-CoV-2 pandemic has been presenting in periodic waves and multiple variants, of which some dominated over time with increased transmissibility. SARS-CoV-2 is still adapting in the human population, thus it is crucial to understand its evolutionary patterns and dynamics ahead of time. In this work, we analyzed transmission clusters and topology of SARS-CoV-2 phylogenies at the global, regional (North America) and clade-specific (Delta and Omicron) epidemic scales. We used the Nextstrain's nCov open global all-time phylogeny (September 2022, 2,698 strains, 2,243 for North America, 499 for Delta21A, and 543 for Omicron20M), with Nextstrain's clade annotation and Pango lineages. Transmission clusters were identified using Phylopart, DYNAMITE, and several tree imbalance measures were calculated, including staircase-ness, Sackin and Colless index. We found that the phylogenetic clustering profiles of the global epidemic have highest diversification at a distance threshold of 3% (divergence of 10, where the tree sampled median is 49). Phylopart and DYNAMITE clusters moderately-to-highly agree with the Pango nomenclature and the Nextstrain's clade. At the regional and clade-specific scale, transmission clustering profiles tend to flatten and similar clusters are found at distance thresholds between 0.05% and 25%. All the considered phylogenies exhibit high tree imbalance with respect to what expected in random phylogenies, suggesting short infection times and antigenic drift, perhaps due to progressive transition from innate to adaptive immunity in the population. Mattia Prosperi, Brittany Rife Magalis, Simone Marini, Marco Salemi |
BIBM | 4 |
| 2022 | Optimizing viral genome subsampling by genetic diversity and temporal distribution (TARDiS) for phylogeneticsabstractSUMMARY: TARDiS is a novel phylogenetic tool for optimal genetic subsampling. It optimizes both genetic diversity and temporal distribution through a genetic algorithm. AVAILABILITY AND IMPLEMENTATION: TARDiS, along with example datasets and a user manual, is available at https://github.com/smarini/tardis-phylogenetics. Simone Marini, Carla Mavian, Alberto Riva, Mattia Prosperi, Marco Salemi, Brittany Rife Magalis |
Bioinform. | 5 |
| 2014 | Evolved neural networks for HIV-1 co-receptor identificationabstractHIV-1 infects a variety of cell types such as macrophages, T-cells and dendritic cells by expressing different chemokine receptors. R5 HIV-1 viruses use the CCR5 co-receptor for entry, X4 viruses use the CXCR4 co-receptor, and several viral strains make use of both co-receptors (a so-called “dual tropic” or R5X4 virus). Both X4 and R5X4 viruses are associated with late stage rapid progression to AIDS. It remains difficult to identify viral co-receptor type in advance of treatment, especially the R5X4 variety. In this paper we extended previous work to classify HIV-1 tropism using evolved neural networks and a larger set of HIV-1 sequences and features to improve overall classification accuracy. Gary B. Fogel, Enoch S. Liu, Marco Salemi, Susanna L. Lamers, Michael S. McGrath |
IEEE Congress on Evolutionary Computation | 3 |
| 2014 | Baseline CD4+ T Cell Counts Correlates with HIV-1 Synonymous Rate in HLA-B*5701 Subjects with Different Risk of Disease ProgressionabstractHLA-B*5701 is the host factor most strongly associated with slow HIV-1 disease progression, although risk of progression may vary among patients carrying this allele. The interplay between HIV-1 evolutionary rate variation and risk of progression to AIDS in HLA-B*5701 subjects was studied using longitudinal viral sequences from high-risk progressors (HRPs) and low-risk progressors (LRPs). Posterior distributions of HIV-1 genealogies assuming a Bayesian relaxed molecular clock were used to estimate the absolute rates of nonsynonymous and synonymous substitutions for different set of branches. Rates of viral evolution, as well as in vitro viral replication capacity assessed using a novel phenotypic assay, were correlated with various clinical parameters. HIV-1 synonymous substitution rates were significantly lower in LRPs than HRPs, especially for sets of internal branches. The viral population infecting LRPs was also characterized by a slower increase in synonymous divergence over time. This pattern did not correlate to differences in viral fitness, as measured by in vitro replication capacity, nor could be explained by differences among subjects in T cell activation or selection pressure. Interestingly, a significant inverse correlation was found between baseline CD4+ T cell counts and mean HIV-1 synonymous rate (which is proportional to the viral replication rate) along branches representing viral lineages successfully propagating through time up to the last sampled time point. The observed lower replication rate in HLA-B*5701 subjects with higher baseline CD4+ T cell counts provides a potential model to explain differences in risk of disease progression among individuals carrying this allele. Melissa M. Norström, Nazle M. Veras, Mattia Prosperi, Jennifer Cook, Wendy Hartogensis, Frederick M. Hecht, Annika C. Karlsoon, Marco Salemi |
PLoS Comput. Biol. | 9 |
| 2012 | QuRe: software for viral quasispecies reconstruction from next-generation sequencing dataabstractSUMMARY: Next-generation sequencing (NGS) is an ideal framework for the characterization of highly variable pathogens, with a deep resolution able to capture minority variants. However, the reconstruction of all variants of a viral population infecting a host is a challenging task for genome regions larger than the average NGS read length. QuRe is a program for viral quasispecies reconstruction, specifically developed to analyze long read (>100 bp) NGS data. The software performs alignments of sequence fragments against a reference genome, finds an optimal division of the genome into sliding windows based on coverage and diversity and attempts to reconstruct all the individual sequences of the viral quasispecies--along with their prevalence--using a heuristic algorithm, which matches multinomial distributions of distinct viral variants overlapping across the genome division. QuRe comes with a built-in Poisson error correction method and a post-reconstruction probabilistic clustering, both parameterized on given error rates in homopolymeric and non-homopolymeric regions. AVAILABILITY: QuRe is platform-independent, multi-threaded software implemented in Java. It is distributed under the GNU General Public License, available at https://sourceforge.net/projects/qure/. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mattia Prosperi, Marco Salemi |
Bioinform. | 2 |
| 2008 | Prediction of R5, X4, and R5X4 HIV-1 Coreceptor Usage with Evolved Neural NetworksabstractThe HIV-1 genome is highly heterogeneous. This variation affords the virus a wide range of molecular properties, including the ability to infect cell types, such as macrophages and lymphocytes, expressing different chemokine receptors on the cell surface. In particular, R5 HIV-1 viruses use CCR5 as co-receptor for viral entry, X4 viruses use CXCR4, whereas some viral strains, known as R5X4 or D-tropic, have the ability to utilize both co-receptors. X4 and R5X4 viruses are associated with rapid disease progression to AIDS. R5X4 viruses differ in that they have yet to be characterized by the examination of the genetic sequence of HIV-1 alone. In this study, a series of experiments was performed to evaluate different strategies of feature selection and neural network optimization. We demonstrate the use of artificial neural networks trained via evolutionary computation to predict viral co-receptor usage. The results indicate identification of R5X4 viruses with predictive accuracy of 75.5%. Susanna L. Lamers, Marco Salemi, Michael S. McGrath, Gary B. Fogel |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2004 | HIVbase: a PC/Windows-based software offering storage and querying power for locally held HIV-1 genetic, experimental and clinical dataabstractBACKGROUND: Human immunodeficiency virus (HIV) research involves ongoing, repetitious sequencing of the HIV genome and the massive accumulation of associated investigational data. As a result, the storage of annotated DNA and/or protein sequences, as well as information retrieval, have become increasingly difficult tasks, with scientists extracting less information from their collected data than they should. OBJECTIVES: Our objective was to design and develop a software package to aid researchers in the storage, analysis and exploration of their HIV-associated data. RESULTS: HIVbase contains familiar, easy-to-use interfaces and functionality for integrating many types of disparate data. The software contains tools that allow for the mass import of raw genetic data, eliminate repetitious sequence translations, have the ability to identify automatically and store HIV regions of interest from nucleic acid or protein sequences, allow for the export of data in commonly used analysis-ready formats, and for unique querying approaches. Susanna L. Lamers, Scott Beason, Luke Dunlap, Robert Compton, Marco Salemi |
Bioinform. | 5 |