EDBT 2026 Demo / reviewers in the wild / expert
Simone Marini
dblp:58/2285
· DBLP profile ↗
29ranked-venue papers
7as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorArtificial intelligence and machine learning · 4 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SARITA: a large language model for generating the S1 subunit of the SARS-CoV-2 spike proteinabstractBACKGROUND: The COVID-19 pandemic has caused over 776 million infections and 7 million deaths globally between December 2019 and November 2024. Since the emergence of the original Wuhan strain, SARS-CoV-2 has evolved into multiple variants-including Alpha, Delta, and Omicron-primarily through mutations in the Spike glycoprotein. The S1 subunit, which binds the human angiotensin-converting enzyme 2 (ACE2) receptor, mutates frequently and plays a key role in infectivity and immune escape, while the more conserved S2 subunit mediates membrane fusion. Anticipating future mutations is essential for guiding vaccine design and therapeutic strategies. Generative Large Language Models (LLMs) have shown promise in protein sequence modeling due to their capacity to produce realistic and functional synthetic sequences. Here, we introduce SARITA, a GPT-3-based LLM with up to 1.2 billion parameters, fine-tuned via continual learning on the protein model RITA trained on 107 017 high-quality SARS-CoV-2 Spike sequences (up to March 1st 2021) to generate high-quality synthetic SARS-CoV-2 Spike S1 subunits. RESULTS: SARITA is able to generate realistic, full-length synthetic S1 subunits starting from a 14-amino-acid prompt. When evaluated on unseen sequences collected between March 2021 and November 2023-including major Variants of Concern (VOCs) such as Delta and Omicron, and Variants of Interest such as Iota-SARITA outperforms baseline and state-of-the-art LLMs in terms of sequence quality, biological plausibility, and similarity to real-world viral evolution. SARITA generates high-quality sequences in over 97% of cases, with markedly lower False Mutation Rate and higher similarity scores (PAM30, Levenshtein distance) compared to alternative approaches. It also accurately reproduces key mutations characteristic of future variants-such as L212I, R158L, T95P, and E406K-which were not present in the training data but emerged later in VOCs like Omicron and Delta. Structure-based analysis confirms the functional plausibility of these substitutions, with ΔΔG values within experimentally supported thresholds for ACE2 and antibody binding. Furthermore, SARITA anticipates immune-evasive mutations and accurately captures the positional and statistical distribution of mutations found in post- March 1st 2021 variants, highlighting its potential as a predictive tool for viral evolution. CONCLUSION: These results indicate the potential of SARITA to predict future SARS-CoV-2 S1 evolution, potentially aiding in the development of adaptable vaccines and treatments. Simone Rancati, Giovanna Nicora, Laura Bergomi, Tommaso Mario Buonocore, Daniel M. Czyz, Enea Parimbelli, Riccardo Bellazzi, Marco Salemi, Mattia Prosperi, Simone Marini |
Briefings Bioinform. | 10 |
| 2025 | Leveraging large language models to predict antibiotic resistance in Mycobacterium tuberculosisabstractMOTIVATION: Antibiotic resistance in Mycobacterium tuberculosis (MTB) poses a significant challenge to global public health. Rapid and accurate prediction of antibiotic resistance can inform treatment strategies and mitigate the spread of resistant strains. In this study, we present a novel approach leveraging large language models (LLMs) to predict antibiotic resistance in MTB (LLMTB). Our model is trained and evaluated on genomic data from 12 185 CRyPTIC isolates and their associated resistance profiles, utilizing natural language processing techniques to capture patterns and mutations linked to resistance. The model's architecture integrates state-of-the-art transformer-based LLMs, enabling the analysis of complex genomic sequences and the extraction of critical features relevant to antibiotic resistance. RESULTS: We evaluate our model's performance using a comprehensive dataset of MTB strains, demonstrating its ability to achieve high performance in predicting resistance to various antibiotics. Unlike traditional machine learning methods, fine-tuning or few-shot learning opens avenues for LLMs to adapt to new or emerging drugs, thereby reducing reliance on extensive data curation. Beyond predictive accuracy, LLMTB uncovers deeper biological insights, identifying critical genes, intergenic regions, and novel resistance mechanisms. This method marks a transformative shift in resistance prediction and offers significant potential for enhancing diagnostic capabilities and guiding personalized treatment plans, ultimately contributing to the global effort to combat tuberculosis and antibiotic resistance. AVAILABILITY AND IMPLEMENTATION: All source code is publicly available at https://github.com/ctestagrose/LLMTB. Conrad Testagrose, Sakshi Pandey, Mohammadali Serajian, Simone Marini, Mattia Prosperi, Christina Boucher 0001 |
Bioinform. | 4 |
| 2025 | Underwater Mediterranean image analysis based on the compute continuum paradigmabstractHuman activity depends on the oceans for food, transportation, leisure, and many more purposes. Oceans cover 70% of the Earth’s surface, but most of them are unknown to humankind. This is the reason why underwater imaging is a valuable resource asset to Marine Science. Images are acquired with observing systems, e.g. autonomous underwater vehicles or underwater observatories, that presently transmit all the raw data to land stations. However, the transfer of such an amount of data could be challenging, considering the limited power supply and transmission bandwidth of these systems. In this paper, we discuss these aspects, and in particular how it is possible to couple Edge and Cloud computing for effective management of the full processing pipeline according to the Compute Continuum paradigm. Michele Ferrari, Daniele D'Agostino, Jacopo Aguzzi, Simone Marini |
Future Gener. Comput. Syst. | 4 |
| 2024 | Sequencing Efforts and Epidemiological Trends: Analyzing SARS-CoV-2 Dynamics Across European NationsabstractThe COVID-19 pandemic has profoundly impacted global health, leading to millions of deaths and overwhelming healthcare systems worldwide. This study investigates the relationship between SARS-CoV-2 sequencing rates and critical epidemiological parameters, such as cases, deaths, and ICU admissions, across 25 European countries from January 2020 to November 2023. By analyzing these relationships, we aim to determine whether sequencing efforts were reactive—in response to epidemiological pressures—or proactive, guided by public health strategies. The analysis used publicly available data from GISAID, OxCGRT, and ECDC, and included weekly aggregation, correlation analysis, and the application of TimeGPT for predictive modeling. Results show that sequencing rates were significantly correlated with ICU admissions, hospitalizations, case numbers, and deaths, though with variability between countries and over different pandemic phases. TimeGPT analysis revealed that sequencing rates were often the most informative feature for predicting future COVID-19 cases in many countries. These findings highlight the potential of sequencing rates to serve as early indicators for severe pandemic outcomes and underscore the importance of context-specific approaches for managing future health crises. Simone Rancati, Daniele Pala, Simone Marini, Marco Salemi, Riccardo Bellazzi, Giovanna Nicora |
BIBM | 3 |
| 2024 | Forecasting dominance of SARS-CoV-2 lineages by anomaly detection using deep AutoEncodersabstractThe COVID-19 pandemic is marked by the successive emergence of new SARS-CoV-2 variants, lineages, and sublineages that outcompete earlier strains, largely due to factors like increased transmissibility and immune escape. We propose DeepAutoCoV, an unsupervised deep learning anomaly detection system, to predict future dominant lineages (FDLs). We define FDLs as viral (sub)lineages that will constitute >10% of all the viral sequences added to the GISAID, a public database supporting viral genetic sequence sharing, in a given week. DeepAutoCoV is trained and validated by assembling global and country-specific data sets from over 16 million Spike protein sequences sampled over a period of ~4 years. DeepAutoCoV successfully flags FDLs at very low frequencies (0.01%-3%), with median lead times of 4-17 weeks, and predicts FDLs between ~5 and ~25 times better than a baseline approach. For example, the B.1.617.2 vaccine reference strain was flagged as FDL when its frequency was only 0.01%, more than a year before it was considered for an updated COVID-19 vaccine. Furthermore, DeepAutoCoV outputs interpretable results by pinpointing specific mutations potentially linked to increased fitness and may provide significant insights for the optimization of public health 'pre-emptive' intervention strategies. Simone Rancati, Giovanna Nicora, Mattia Prosperi, Riccardo Bellazzi, Marco Salemi, Simone Marini |
Briefings Bioinform. | 6 |
| 2024 | Scalable de novo classification of antibiotic resistance of Mycobacterium tuberculosisabstractMOTIVATION: World Health Organization estimates that there were over 10 million cases of tuberculosis (TB) worldwide in 2019, resulting in over 1.4 million deaths, with a worrisome increasing trend yearly. The disease is caused by Mycobacterium tuberculosis (MTB) through airborne transmission. Treatment of TB is estimated to be 85% successful, however, this drops to 57% if MTB exhibits multiple antimicrobial resistance (AMR), for which fewer treatment options are available. RESULTS: We develop a robust machine-learning classifier using both linear and nonlinear models (i.e. LASSO logistic regression (LR) and random forests (RF)) to predict the phenotypic resistance of Mycobacterium tuberculosis (MTB) for a broad range of antibiotic drugs. We use data from the CRyPTIC consortium to train our classifier, which consists of whole genome sequencing and antibiotic susceptibility testing (AST) phenotypic data for 13 different antibiotics. To train our model, we assemble the sequence data into genomic contigs, identify all unique 31-mers in the set of contigs, and build a feature matrix M, where M[i, j] is equal to the number of times the ith 31-mer occurs in the jth genome. Due to the size of this feature matrix (over 350 million unique 31-mers), we build and use a sparse matrix representation. Our method, which we refer to as MTB++, leverages compact data structures and iterative methods to allow for the screening of all the 31-mers in the development of both LASSO LR and RF. MTB++ is able to achieve high discrimination (F-1 >80%) for the first-line antibiotics. Moreover, MTB++ had the highest F-1 score in all but three classes and was the most comprehensive since it had an F-1 score >75% in all but four (rare) antibiotic drugs. We use our feature selection to contextualize the 31-mers that are used for the prediction of phenotypic resistance, leading to some insights about sequence similarity to genes in MEGARes. Lastly, we give an estimate of the amount of data that is needed in order to provide accurate predictions. AVAILABILITY: The models and source code are publicly available on Github at https://github.com/M-Serajian/MTB-Pipeline. Mohammadali Serajian, Simone Marini, Jarno Alanko, Noelle R. Noyes, Mattia Prosperi, Christina Boucher 0001 |
Bioinform. | 2 |
| 2022 | Transmission cluster characteristics of global, regional, and lineage-specific SARS-CoV-2 phylogeniesabstractThe SARS-CoV-2 pandemic has been presenting in periodic waves and multiple variants, of which some dominated over time with increased transmissibility. SARS-CoV-2 is still adapting in the human population, thus it is crucial to understand its evolutionary patterns and dynamics ahead of time. In this work, we analyzed transmission clusters and topology of SARS-CoV-2 phylogenies at the global, regional (North America) and clade-specific (Delta and Omicron) epidemic scales. We used the Nextstrain's nCov open global all-time phylogeny (September 2022, 2,698 strains, 2,243 for North America, 499 for Delta21A, and 543 for Omicron20M), with Nextstrain's clade annotation and Pango lineages. Transmission clusters were identified using Phylopart, DYNAMITE, and several tree imbalance measures were calculated, including staircase-ness, Sackin and Colless index. We found that the phylogenetic clustering profiles of the global epidemic have highest diversification at a distance threshold of 3% (divergence of 10, where the tree sampled median is 49). Phylopart and DYNAMITE clusters moderately-to-highly agree with the Pango nomenclature and the Nextstrain's clade. At the regional and clade-specific scale, transmission clustering profiles tend to flatten and similar clusters are found at distance thresholds between 0.05% and 25%. All the considered phylogenies exhibit high tree imbalance with respect to what expected in random phylogenies, suggesting short infection times and antigenic drift, perhaps due to progressive transition from innate to adaptive immunity in the population. Mattia Prosperi, Brittany Rife Magalis, Simone Marini, Marco Salemi |
BIBM | 3 |
| 2022 | Identification of Social and Racial Disparities in Risk of HIV Infection in Florida using Causal AI Methodsabstractmost populous state in the USA-has the highest rates of Human Immunodeficiency Virus (HIV) infections and of unfavorable HIV outcomes, with marked social and racial disparities. In this work, we leveraged large-scale, real-world data, i.e., statewide surveillance records and publicly available data resources encoding social determinants of health (SDoH), to identify social and racial disparities contributing to individuals' risk of HIV infection. We used the Florida Department of Health's Syndromic Tracking and Reporting System (STARS) database (including 100,000+ individuals screened for HIV infection and their partners), and a novel algorithmic fairness assessment method -the Fairness-Aware Causal paThs decompoSition (FACTS)- merging causal inference and artificial intelligence. FACTS deconstructs disparities based on SDoH and individuals' characteristics, and can discover novel mechanisms of inequity, quantifying to what extent they could be reduced by interventions. We paired the deidentified demographic information (age, gender, drug use) of 44,350 individuals in STARS -with non-missing data on interview year, county of residence, and infection status- to eight SDoH, including access to healthcare facilities, % uninsured, median household income, and violent crime rate. Using an expert-reviewed causal graph, we found that the risk of HIV infection for African Americans was higher than for non- African Americans (both in terms of direct and total effect), although a null effect could not be ruled out. FACTS identified several paths leading to racial disparity in HIV risk, including multiple SDoH: education, income, violent crime, drinking, smoking, and rurality. Mattia Prosperi, Jie Xu 0012, Jingchuan Serena Guo, Jiang Bian 0001, Wei-Han William Chen, Shantrel S. Canidate, Simone Marini |
BIBM | 7 |
| 2022 | Assessing putative bias in prediction of anti-microbial resistance from real-world genotyping data under explicit causal assumptions
Mattia Prosperi, Christina Boucher 0001, Jiang Bian 0001, Simone Marini |
Artif. Intell. Medicine | 4 |
| 2022 | Towards routine employment of computational tools for antimicrobial resistance determination via high-throughput sequencingabstractAntimicrobial resistance (AMR) is a growing threat to public health and farming at large. In clinical and veterinary practice, timely characterization of the antibiotic susceptibility profile of bacterial infections is a crucial step in optimizing treatment. High-throughput sequencing is a promising option for clinical point-of-care and ecological surveillance, opening the opportunity to develop genotyping-based AMR determination as a possibly faster alternative to phenotypic testing. In the present work, we compare the performance of state-of-the-art methods for detection of AMR using high-throughput sequencing data from clinical settings. We consider five computational approaches based on alignment (AMRPlusPlus), deep learning (DeepARG), k-mer genomic signatures (KARGA, ResFinder) or hidden Markov models (Meta-MARC). We use an extensive collection of 585 isolates with available AMR resistance profiles determined by phenotypic tests across nine antibiotic classes. We show how the prediction landscape of AMR classifiers is highly heterogeneous, with balanced accuracy varying from 0.40 to 0.92. Although some algorithms-ResFinder, KARGA and AMRPlusPlus-exhibit overall better balanced accuracy than others, the high per-AMR-class variance and related findings suggest that: (1) all algorithms might be subject to sampling bias both in data repositories used for training and experimental/clinical settings; and (2) a portion of clinical samples might contain uncharacterized AMR genes that the algorithms-mostly trained on known AMR genes-fail to generalize upon. These results lead us to formulate practical advice for software configuration and application, and give suggestions for future study designs to further develop AMR prediction tools from proof-of-concept to bedside. Simone Marini, Rodrigo A. Mora, Christina Boucher 0001, Noelle R. Noyes, Mattia Prosperi |
Briefings Bioinform. | 1 |
| 2022 | Optimizing viral genome subsampling by genetic diversity and temporal distribution (TARDiS) for phylogeneticsabstractSUMMARY: TARDiS is a novel phylogenetic tool for optimal genetic subsampling. It optimizes both genetic diversity and temporal distribution through a genetic algorithm. AVAILABILITY AND IMPLEMENTATION: TARDiS, along with example datasets and a user manual, is available at https://github.com/smarini/tardis-phylogenetics. Simone Marini, Carla Mavian, Alberto Riva, Mattia Prosperi, Marco Salemi, Brittany Rife Magalis |
Bioinform. | 1 |
| 2021 | Fast and exact quantification of motif occurrences in biological sequencesabstractBACKGROUND: Identification of motifs and quantification of their occurrences are important for the study of genetic diseases, gene evolution, transcription sites, and other biological mechanisms. Exact formulae for estimating count distributions of motifs under Markovian assumptions have high computational complexity and are impractical to be used on large motif sets. Approximated formulae, e.g. based on compound Poisson, are faster, but reliable p value calculation remains challenging. Here, we introduce 'motif_prob', a fast implementation of an exact formula for motif count distribution through progressive approximation with arbitrary precision. Our implementation speeds up the exact calculation, usually impractical, making it feasible and posit to substitute currently employed heuristics. RESULTS: We implement motif_prob in both Perl and C+ + languages, using an efficient error-bound iterative process for the exact formula, providing comparison with state-of-the-art tools (e.g. MoSDi) in terms of precision, run time benchmarks, along with a real-world use case on bacterial motif characterization. Our software is able to process a million of motifs (13-31 bases) over genome lengths of 5 million bases within the minute on a regular laptop, and the run times for both the Perl and C+ + code are several orders of magnitude smaller (50-1000× faster) than MoSDi, even when using their fast compound Poisson approximation (60-120× faster). In the real-world use cases, we first show the consistency of motif_prob with MoSDi, and then how the p-value quantification is crucial for enrichment quantification when bacteria have different GC content, using motifs found in antimicrobial resistance genes. The software and the code sources are available under the MIT license at https://github.com/DataIntellSystLab/motif_prob . CONCLUSIONS: The motif_prob software is a multi-platform and efficient open source solution for calculating exact frequency distributions of motifs. It can be integrated with motif discovery/characterization tools for quantifying enrichment and deviation from expected frequency ranges with exact p values, without loss in data processing efficiency. Mattia Prosperi, Simone Marini, Christina Boucher 0001 |
BMC Bioinform. | 2 |
| 2019 | A Semi-supervised Learning Approach for Pan-Cancer Somatic Genomic Variant Classification
Giovanna Nicora, Simone Marini, Ivan Limongelli, Ettore Rizzo, Stefano Montoli, Francesca Floriana Tricomi, Riccardo Bellazzi |
AIME | 2 |
| 2019 | Underwater Fish Detection with Weak Multi-Domain SupervisionabstractGiven a sufficiently large training dataset, it is relatively easy to train a modern convolution neural network (CNN) as a required image classifier. However, for the task of fish classification and/or fish detection, if a CNN was trained to detect or classify particular fish species in particular background habitats, the same CNN exhibits much lower accuracy when applied to new/unseen fish species and/or fish habitats. Therefore, in practice, the CNN needs to be continuously fine-tuned to improve its classification accuracy to handle new project-specific fish species or habitats. In this work we present a labelling-efficient method of training a CNN-based fish-detector (the Xception CNN was used as the base) on relatively small numbers (4,000) of project-domain underwater fish/no-fish images from 20 different habitats. Additionally, 17,000 of known negative (that is, missing fish) general-domain (VOC2012) above-water images were used. Two publicly available fish-domain datasets supplied additional 27,000 of above-water and underwater positive/fish images. By using this multi-domain collection of images, the trained Xception-based binary (fish/not-fish) classifier achieved 0.17% false-positives and 0.61% false-negatives on the project's 20,000 negative and 16,000 positive holdout test images, respectively. The area under the ROC curve (AUC) was 99.94%. Dmitry A. Konovalov, Alzayat Saleh, Michael Bradley, Mangalam Sankupellay, Simone Marini, Marcus Sheaves |
IJCNN | 5 |
| 2019 | Toward more accurate prediction of caspase cleavage sites: a comprehensive review of current methods, tools and featuresabstractAs one of the few irreversible protein posttranslational modifications, proteolytic cleavage is involved in nearly all aspects of cellular activities, ranging from gene regulation to cell life-cycle regulation. Among the various protease-specific types of proteolytic cleavage, cleavages by casapses/granzyme B are considered as essential in the initiation and execution of programmed cell death and inflammation processes. Although a number of substrates for both types of proteolytic cleavage have been experimentally identified, the complete repertoire of caspases and granzyme B substrates remains to be fully characterized. To tackle this issue and complement experimental efforts for substrate identification, systematic bioinformatics studies of known cleavage sites provide important insights into caspase/granzyme B substrate specificity, and facilitate the discovery of novel substrates. In this article, we review and benchmark 12 state-of-the-art sequence-based bioinformatics approaches and tools for caspases/granzyme B cleavage prediction. We evaluate and compare these methods in terms of their input/output, algorithms used, prediction performance, validation methods and software availability and utility. In addition, we construct independent data sets consisting of caspases/granzyme B substrates from different species and accordingly assess the predictive power of these different predictors for the identification of cleavage sites. We find that the prediction results are highly variable among different predictors. Furthermore, we experimentally validate the predictions of a case study by performing caspase cleavage assay. We anticipate that this comprehensive review and survey analysis will provide an insightful resource for biologists and bioinformaticians who are interested in using and/or developing tools for caspase/granzyme B cleavage prediction. Simone Marini, Takeyuki Tamura, Mayumi Kamada, Shingo Maegawa, Hiroshi Hosokawa, Jiangning Song, Tatsuya Akutsu |
Briefings Bioinform. | 2 |
| 2019 | Protease target prediction via matrix factorizationabstractMOTIVATION: Protein cleavage is an important cellular event, involved in a myriad of processes, from apoptosis to immune response. Bioinformatics provides in silico tools, such as machine learning-based models, to guide the discovery of targets for the proteases responsible for protein cleavage. State-of-the-art models have a scope limited to specific protease families (such as Caspases), and do not explicitly include biological or medical knowledge (such as the hierarchical protein domain similarity or gene-gene interactions). To fill this gap, we present a novel approach for protease target prediction based on data integration. RESULTS: By representing protease-protein target information in the form of relational matrices, we design a model (i) that is general and not limited to a single protease family, and (b) leverages on the available knowledge, managing extremely sparse data from heterogeneous data sources, including primary sequence, pathways, domains and interactions. When compared with other algorithms on test data, our approach provides a better performance even for models specifically focusing on a single protease family. AVAILABILITY AND IMPLEMENTATION: https://gitlab.com/smarini/MaDDA/ (Matlab code and utilized data.). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Simone Marini, Francesca Vitali, Sara Rampazzi, Andrea Demartini, Tatsuya Akutsu |
Bioinform. | 1 |
| 2018 | Patient similarity for precision medicine: A systematic review
Enea Parimbelli, Simone Marini, Lucia Sacchi, Riccardo Bellazzi |
J. Biomed. Informatics | 2 |
| 2017 | Dscam1 web server: online prediction of Dscam1 self- and hetero-affinityabstractMOTIVATION: Formation of homodimers by identical Dscam1 protein isomers on cell surface is the key factor for the self-avoidance of growing neurites. Dscam1 immense diversity has a critical role in the formation of arthropod neuronal circuit, showing unique evolutionary properties when compared to other cell surface proteins. Experimental measures are available for 89 self-binding and 1722 hetero-binding protein samples, out of more than 19 thousands (self-binding) and 350 millions (hetero-binding) possible isomer combinations. RESULTS: We developed Dscam1 Web Server to quickly predict Dscam1 self- and hetero- binding affinity for batches of Dscam1 isomers. The server can help the study of Dscam1 affinity and help researchers navigate through the tens of millions of possible isomer combinations to isolate the strong-binding ones. AVAILABILITY AND IMPLEMENTATION: Dscam1 Web Server is freely available at: http://bioinformatics.tecnoparco.org/Dscam1-webserver . Web server code is available at https://gitlab.com/ne1s0n/Dscam1-binding . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Simone Marini, Nelson Nazzicari, Filippo Biscarini, Guang-Zhong Wang |
Bioinform. | 1 |
| 2015 | A Genomic Data Fusion Framework to Exploit Rare and Common Variants for Association Discovery
Simone Marini, Ivan Limongelli, Ettore Rizzo, Tan Da, Riccardo Bellazzi |
AIME | 1 |
| 2015 | Preclinical Tests for Cerebral StrokeabstractStroke is the second single highest cause of death in Europe. The low reliability of animal models in replicating the human disease is one of the most serious problems in the field of medical and pharmaceutical research about stroke. The standard models for the study of ischemic stroke are often poorly predictive as they simulate only partially the human disease. This work aims at investigating animal models with diseases typically associated with the onset of stroke in human patients. Maria Francesca Zini, Nadia Pisanti, E. Biasci, A. Podda, V. Mey, F. Piras, G. L. L'Abbate, Simone Marini, D. Fratta, S. Bonaretti, S. Trasciatti |
ASONAM | 8 |
| 2015 | PaPI: pseudo amino acid composition to score human protein-coding variantsabstractBACKGROUND: High throughput sequencing technologies are able to identify the whole genomic variation of an individual. Gene-targeted and whole-exome experiments are mainly focused on coding sequence variants related to a single or multiple nucleotides. The analysis of the biological significance of this multitude of genomic variant is challenging and computational demanding. RESULTS: We present PaPI, a new machine-learning approach to classify and score human coding variants by estimating the probability to damage their protein-related function. The novelty of this approach consists in using pseudo amino acid composition through which wild and mutated protein sequences are represented in a discrete model. A machine learning classifier has been trained on a set of known deleterious and benign coding variants with the aim to score unobserved variants by taking into account hidden sequence patterns in human genome potentially leading to diseases. We show how the combination of amphiphilic pseudo amino acid composition, evolutionary conservation and homologous proteins based methods outperforms several prediction algorithms and it is also able to score complex variants such as deletions, insertions and indels. CONCLUSIONS: This paper describes a machine-learning approach to predict the deleteriousness of human coding variants. A freely available web application (http://papi.unipv.it) has been developed with the presented method, able to score up to thousands variants in a single run. Ivan Limongelli, Simone Marini, Riccardo Bellazzi |
BMC Bioinform. | 2 |
| 2015 | A Dynamic Bayesian Network model for long-term simulation of clinical complications in type 1 diabetes
Simone Marini, Emanuele Trifoglio, Nicola Barbarini, Francesco Sambo, Barbara Di Camillo, Alberto Malovini, Marco Manfrini, Claudio Cobelli, Riccardo Bellazzi |
J. Biomed. Informatics | 1 |
| 2011 | Part-in-whole 3D shape matching and docking
Marco Attene, Simone Marini, Michela Spagnuolo, Bianca Falcidieno |
Vis. Comput. | 2 |
| 2011 | Spectral feature selection for shape characterization and classification
Simone Marini, Giuseppe Patanè 0001, Michela Spagnuolo, Bianca Falcidieno |
Vis. Comput. | 1 |
| 2010 | Thesaurus-based 3D Object Retrieval with Part-in-Whole Matching
Alfredo Ferreira, Simone Marini, Marco Attene, Manuel J. Fonseca, Michela Spagnuolo, Joaquim Jorge 0001, Bianca Falcidieno |
Int. J. Comput. Vis. | 2 |
| 2009 | A Critical Assessment of 2D and 3D Face Recognition AlgorithmsabstractWe present the results of a project aimed to evaluate 2D and 3D face recognition algorithms. In particular, we focused on the potentialities of 3D-based techniques to overcome typical limitations of 2D methods in non-controlled situations. According to the reference scenario of people identification at airport check points, we built a representative database on which we tested different face recognition algorithms. We implemented and tested an improved version of a well-known state-of-the-art 3D approach, and verified that on our dataset it performs better than a widely used commercial system. Daniela Giorgi, Marco Attene, Giuseppe Patanè 0001, Simone Marini, Corrado Pizzi, Silvia Biasotti, Michela Spagnuolo, Bianca Falcidieno, Marco Corvi, L. Usai, L. Roncarolo, Giovanni Garibotto |
AVSS | 4 |
| 2008 | SHape REtrieval contest 2008: Classification of watertight modelsabstractThis track focuses on 3D shape classification, i.e. the assignment of a query object to one of the classes in a database. A total amount of 646 watertight models have been classified, and released as training, testing and query data. Three different levels of categorization have been taken into account, from coarse to fine. Daniela Giorgi, Simone Marini |
Shape Modeling International | 2 |
| 2006 | Sub-part correspondence by structural descriptors of 3D shapes
Silvia Biasotti, Simone Marini, Michela Spagnuolo, Bianca Falcidieno |
Comput. Aided Des. | 2 |
| 2003 | An overview on properties and efficacy of topological skeletons in Shape ModellingabstractThe paper investigates the main issues related to the definition of abstraction tools for deriving high-level descriptions of complex geometric models. Among the wide range of shape descriptors, topological graph-like representations not only give a powerful and synthetic sketch of the object, but also capture its inner structure, that is how features connect together to give the overall shape. This aspect makes them useful to describe complex 3D objects in various applications like modeling, morphing, matching and recognition. The paper surveys the main properties of skeletons developed in shape modeling for representing objects. Silvia Biasotti, Simone Marini, Michela Mortara, Giuseppe Patanè 0001 |
Shape Modeling International | 2 |