VLDB 2026 Research / reviewers in the wild / expert
Simone Rancati
dblp:382/7797
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2026
0009-0002-4405-1697ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Biological Plausibility Assessment of Viral Sequences Generated by a Genomic Language Model
Pablo Arozarena Donelli, Simone Rancati, Giovanna Nicora, Riccardo Bellazzi, Enea Parimbelli, Luigi Portinale |
AIME (2) | 2 |
| 2026 | Epistemologically Guided LLM Reasoning for Differential Diagnosis
Simone Rancati, Laura Bergomi, Enea Parimbelli, Giovanna Nicora, Riccardo Bellazzi |
AIME (1) | 1 |
| 2025 | SARITA: a large language model for generating the S1 subunit of the SARS-CoV-2 spike proteinabstractBACKGROUND: The COVID-19 pandemic has caused over 776 million infections and 7 million deaths globally between December 2019 and November 2024. Since the emergence of the original Wuhan strain, SARS-CoV-2 has evolved into multiple variants-including Alpha, Delta, and Omicron-primarily through mutations in the Spike glycoprotein. The S1 subunit, which binds the human angiotensin-converting enzyme 2 (ACE2) receptor, mutates frequently and plays a key role in infectivity and immune escape, while the more conserved S2 subunit mediates membrane fusion. Anticipating future mutations is essential for guiding vaccine design and therapeutic strategies. Generative Large Language Models (LLMs) have shown promise in protein sequence modeling due to their capacity to produce realistic and functional synthetic sequences. Here, we introduce SARITA, a GPT-3-based LLM with up to 1.2 billion parameters, fine-tuned via continual learning on the protein model RITA trained on 107 017 high-quality SARS-CoV-2 Spike sequences (up to March 1st 2021) to generate high-quality synthetic SARS-CoV-2 Spike S1 subunits. RESULTS: SARITA is able to generate realistic, full-length synthetic S1 subunits starting from a 14-amino-acid prompt. When evaluated on unseen sequences collected between March 2021 and November 2023-including major Variants of Concern (VOCs) such as Delta and Omicron, and Variants of Interest such as Iota-SARITA outperforms baseline and state-of-the-art LLMs in terms of sequence quality, biological plausibility, and similarity to real-world viral evolution. SARITA generates high-quality sequences in over 97% of cases, with markedly lower False Mutation Rate and higher similarity scores (PAM30, Levenshtein distance) compared to alternative approaches. It also accurately reproduces key mutations characteristic of future variants-such as L212I, R158L, T95P, and E406K-which were not present in the training data but emerged later in VOCs like Omicron and Delta. Structure-based analysis confirms the functional plausibility of these substitutions, with ΔΔG values within experimentally supported thresholds for ACE2 and antibody binding. Furthermore, SARITA anticipates immune-evasive mutations and accurately captures the positional and statistical distribution of mutations found in post- March 1st 2021 variants, highlighting its potential as a predictive tool for viral evolution. CONCLUSION: These results indicate the potential of SARITA to predict future SARS-CoV-2 S1 evolution, potentially aiding in the development of adaptable vaccines and treatments. Simone Rancati, Giovanna Nicora, Laura Bergomi, Tommaso Mario Buonocore, Daniel M. Czyz, Enea Parimbelli, Riccardo Bellazzi, Marco Salemi, Mattia Prosperi, Simone Marini |
Briefings Bioinform. | 1 |
| 2024 | Sequencing Efforts and Epidemiological Trends: Analyzing SARS-CoV-2 Dynamics Across European NationsabstractThe COVID-19 pandemic has profoundly impacted global health, leading to millions of deaths and overwhelming healthcare systems worldwide. This study investigates the relationship between SARS-CoV-2 sequencing rates and critical epidemiological parameters, such as cases, deaths, and ICU admissions, across 25 European countries from January 2020 to November 2023. By analyzing these relationships, we aim to determine whether sequencing efforts were reactive—in response to epidemiological pressures—or proactive, guided by public health strategies. The analysis used publicly available data from GISAID, OxCGRT, and ECDC, and included weekly aggregation, correlation analysis, and the application of TimeGPT for predictive modeling. Results show that sequencing rates were significantly correlated with ICU admissions, hospitalizations, case numbers, and deaths, though with variability between countries and over different pandemic phases. TimeGPT analysis revealed that sequencing rates were often the most informative feature for predicting future COVID-19 cases in many countries. These findings highlight the potential of sequencing rates to serve as early indicators for severe pandemic outcomes and underscore the importance of context-specific approaches for managing future health crises. Simone Rancati, Daniele Pala, Simone Marini, Marco Salemi, Riccardo Bellazzi, Giovanna Nicora |
BIBM | 1 |
| 2024 | Forecasting dominance of SARS-CoV-2 lineages by anomaly detection using deep AutoEncodersabstractThe COVID-19 pandemic is marked by the successive emergence of new SARS-CoV-2 variants, lineages, and sublineages that outcompete earlier strains, largely due to factors like increased transmissibility and immune escape. We propose DeepAutoCoV, an unsupervised deep learning anomaly detection system, to predict future dominant lineages (FDLs). We define FDLs as viral (sub)lineages that will constitute >10% of all the viral sequences added to the GISAID, a public database supporting viral genetic sequence sharing, in a given week. DeepAutoCoV is trained and validated by assembling global and country-specific data sets from over 16 million Spike protein sequences sampled over a period of ~4 years. DeepAutoCoV successfully flags FDLs at very low frequencies (0.01%-3%), with median lead times of 4-17 weeks, and predicts FDLs between ~5 and ~25 times better than a baseline approach. For example, the B.1.617.2 vaccine reference strain was flagged as FDL when its frequency was only 0.01%, more than a year before it was considered for an updated COVID-19 vaccine. Furthermore, DeepAutoCoV outputs interpretable results by pinpointing specific mutations potentially linked to increased fitness and may provide significant insights for the optimization of public health 'pre-emptive' intervention strategies. Simone Rancati, Giovanna Nicora, Mattia Prosperi, Riccardo Bellazzi, Marco Salemi, Simone Marini |
Briefings Bioinform. | 1 |