EDBT 2026 Demo / reviewers in the wild / expert
Barbara Di Camillo
dblp:84/7110
· DBLP profile ↗
51ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0001-8415-4688ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 44 · 8 first-author · 20 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Network-Based Integration of Multi-omics Data for Biomarker Discovery in Acute Coronary Syndromes
Elena Marinello, Shahryar Noei, Marco Chierici, Núria Amigó, Dane Cvijanovic, Aleksandar Davidovic, Dalibor Dragisic, Claudia Giglione, Dimitris Kardassis, Marko Milanov, Jelena Munjas, Iva Perovic-Blagojevic, Aleksa Petkovic, Tamara Ratkovic, Pol Torne, Christos Tsatsanis, Sandra Vladimirov-Sopic, Luka Vukmirovic, Paolo Magni, Miron Sopic, Barbara Di Camillo, Giuseppe Jurman |
AIME (2) | 21 |
| 2025 | Exploring the Use of Projecting Conflicting Gradients in Multi-task Neural Networks with an Application to Amyotrophic Lateral Sclerosis
Davide Dei Cas, Enrico Longato, Erica Tavazzi, Umberto Manera, Adriano Chiò, Marta Gromicho, Inês Alves, Mamede de Carvalho, Barbara Di Camillo |
AIME (1) | 9 |
| 2025 | Deep Learning Model Predicts Relapse Occurrence in Multiple Sclerosis Via Sequences of Environmental DataabstractAir pollution is a known risk factor for the exacerbation of many diseases. Among these, is multiple sclerosis (MS), a chronic, autoimmune, neurological disease, characterised by transient episodes of neurological impairment known as relapses. Although the link between environmental factors and relapses has been a subject of investigation in the medical and biostatistical literature, its implications for predictive modelling are still unclear. Thus, in this work, we develop a deep learning model that is able to combine four weeks of environmental data, collected by pollutant-monitoring and weather stations, with patient information to predict an imminent relapse in the following week. Specifically, we cast the task as distinguishing between 4-week sequences followed by a relapse vs. 4-week sequences followed by another relapse-free week, the latter of which were extracted from MS patients who were never observed to have had a relapse. The 1556 sequences were collected in the context of the H2020 BRAINTEASER (”Bringing Artificial Intelligence Home for a Better Care of Amyotrophic Lateral Sclerosis and Multiple Sclerosis”) project. The best-performing model was a recurrent neural network, which yielded an encouraging test-set area under the receiveroperating characteristic curve (AUROC) of 0.70. It also performed adequately (AUROC$=0.60$) on a modified version of the test set where the 4-week relapse-free sequences followed by another relapse-free week were extracted from the same subjects from whom the test sequences followed by a relapse came. Thus, our results, albeit preliminary, suggest that the inclusion of environmental data as the basis of predictive models of MS relapses is a promising direction to obtain short-term predictions, which may be helpful for therapy and life planning. It is especially encouraging that better-than-random performance was preserved on the modified test set, where environmental factors were, by construction, the most informative predictors. Enrico Longato, Erica Tavazzi, Anna Milani, Elena Marinello, Pietro Bosoni, Arianna Dagliati, Mahin Vazifehdan, Riccardo Bellazzi, Isotta Trescato, Alessandro Guazzo, Martina Vettoretti, Eleonora Tavazzi, Lara Ahmad, Roberto Bergamaschi, Paola Cavalla, Umberto Manera, Adriano Chiò, Barbara Di Camillo |
BIBM | 18 |
| 2025 | A Dynamic Bayesian Network Approach for Generating Synthetic Longitudinal Clinical Data: A Case Study on Long-Term Diabetes OutcomesabstractSynthetic clinical data offer several advantages, including the possibility to simulate patient trajectories and investigate long-term outcomes that would otherwise require extensive time and resources to investigate through traditional clinical trials. In this work, we propose a modelling approach based on dynamic Bayesian networks (DBNs) to generate reliable synthetic data that faithfully reproduce the characteristics of original longitudinal clinical datasets. The proposed pipeline includes two main steps: i) learning a DBN model using a layered variable structure informed by domain knowledge, and ii) employing the trained model to simulate longitudinal clinical data. We applied this approach to generate a synthetic version of the LEADER trial dataset, which includes longitudinal data of patients with type 2 diabetes and high cardiovascular risk. After preprocessing, the dataset consisted of 45,556 observations across 82 variables from 8,301 patients. The general utility of the synthetic data was assessed by comparing the distribution of each variable between the synthetic and original datasets. In addition, we evaluated the similarity of Kaplan-Meier survival curves for the major adverse cardiovascular events (MACE), that was the primary endpoints of the trial, between synthetic and real data. Overall, the proposed method demonstrated satisfactory performance, with the synthetic data closely replicating both the variable distributions and time-to-event outcomes observed in the original dataset. Sara Poletto, Noemi Gonzato, Erica Tavazzi, Enrico Longato, Amanda Adler, Mari-Anne Gall, Matthias Müllenborn, Barbara Di Camillo, Martina Vettoretti |
BIBM | 8 |
| 2025 | Biologically Informed procedure for Feature Summarization in Spatial TranscriptomicsabstractModern machine learning approaches have shown remarkable success in extracting patterns from high-dimensional biological data. However, when applied to spatial transcriptomics, these methods face significant challenges due to the sparsity of spatially resolved measurements and the complex, nonlinear relationships between molecular features.To address these challenges, we propose a procedure that integrates single-cell and spatial transcriptomics by considering biologically meaningful regulatory factors as an interpretable feature space. These factors act as latent variables that encode transcriptional programmes, reducing dimensionality and preserving mechanistic relevance.This approach improves interpretability by shifting from raw gene expression to a structured representation of regulatory activity, providing a scalable and biologically interpretable framework for spatial transcriptomic analysis. Matteo Baldan, Giulia Cesaro, Giacomo Baruzzo, Barbara Di Camillo |
IJCNN | 4 |
| 2025 | quickSparseM: a library for memory- and time-efficient computation on large, sparse matrices with application to omics dataabstractOmics data have revolutionized molecular biology by introducing large-scale data analysis, pushing the field into the realm of big data and presenting substantial challenges in data storage and analysis. Despite describing distinct aspects of molecular biology, most omics data share common characteristics, such as being representable as large, sparse matrices, and requiring similar computational approaches, mainly involving embarrassing parallel tasks across rows or columns. While R is a popular choice for omics analysis, it encounters performance bottlenecks when handling large datasets due to its reliance on dense data formats and constraints like 32-bit indexing in some structures. Even when sparse representations are utilized, the inherent limitations of R lead to inefficiencies. Additionally, its lack of native support for shared-memory parallelism prevents it from fully utilizing modern parallel computing architectures. Similarly, many other data-intensive fields that rely on R face similar challenges with large, sparse data requiring fast and memory-efficient row-wise and column-wise operations. To address these challenges, we introduce quickSparseM, a time- and memory-efficient library for storing and processing large, sparse matrices, available as an R package. Developed in C++ with OpenMP for parallelism, quickSparseM achieves efficient performance while remaining compatible with existing R-based workflows. The library utilizes the R dgCMatrix format to represent sparse matrices in a compressed, column-oriented format and provide functions to compute basic statistics and operations commonly used in omics analyses. Experiments varying dataset sizes and core counts, as well as two case studies using omics data, demonstrate the library’s efficiency and scalability. The results indicate that quickSparseM outperforms state-of-the-art R packages for sparse matrix computation in terms of time, memory usage, and scalability. Giacomo Baruzzo, Giulia Cesaro, Barbara Di Camillo |
PDP | 3 |
| 2025 | Advances and challenges in cell-cell communication inference: a comprehensive review of tools, resources, and future directionsabstractRecent advancements in high-resolution and high-throughput sequencing technologies have significantly enhanced the study of cell-cell communication inference using single-cell and spatial transcriptomics data. Over the past 6 years, this growing interest has led to the development of more than 100 bioinformatics tools and nearly 50 resources, primarily in the form of ligand-receptor databases. These tools vary widely in their requirements, scoring approaches, ability to infer inter- and/or intra-cellular communication, assumptions, and limitations. Similarly, cell-cell communication resources differ in many aspects, mainly in the number of annotated interactions, species coverage, and their focus on inter-cellular signaling or both inter- and intra-cellular communication. This abundance and diversity create challenges in identifying compatible and suitable tools and resources to meet specific user needs. In this collaborative effort, we aim to provide a comprehensive report on the current state of cell-cell communication analysis derived from single-cell or spatial transcriptomics data. The report reviews existing methods and resources, addressing all relevant aspects from the user's perspective. It also explores current limitations, pitfalls, and unresolved issues in cell-cell communication inference, offering an aggregated analysis of the existing literature on the topic. Furthermore, we highlight potential future directions in the field and consolidate the collected knowledge into CCC-Catalog (https://sysbiobig.gitlab.io/ccc-catalog), a centralized web platform designed to serve as a hub for bioinformaticians and researchers interested in cell-cell communication inference. Giulia Cesaro, James Shiniti Nagai, Nicolò Gnoato, Alice Chiodi, Gaia Tussardi, Vanessa Klöker, Carmelo Vittorio Musumarra, Ettore Mosca, Ivan G. Costa, Barbara Di Camillo, Enrica Calura, Giacomo Baruzzo |
Briefings Bioinform. | 10 |
| 2025 | MOV&RSim: computational modelling of cancer-specific variants and sequencing reads characteristics for realistic tumoral sample simulationabstractBACKGROUND: Bioinformatics pipelines for variant calling have undergone significant advancements due to the decreasing costs of next-generation sequencing. Accurate mutation detection is crucial for personalised medicine in cancer, particularly in assignment of therapy. Somatic variant calling, however, remains challenging due to diverse cancer types, heterogeneity, complex mutational profiles, and unpredictable sequencing errors. A dataset of fully characterised tumoral genomes and sequencing reads, large enough to represent the variability inherent in different cancer types, is still lacking, even considering synthetic data. The lack of such datasets hampers rigorous evaluation, benchmarking and optimization of variant callers for specific cancer types. RESULTS: The contribution of this work is twofold. First, we conducted a comprehensive analysis of nine somatic sample simulators (Synggen, BAMSurgeon, SVEngine, VarSim, Xome-Blender, tHapMix, Pysim-sv, SCNVSim, HeteroGenesis) assessing their ability to control biological parameters, including variants characteristics (type, number, position, length, content, zygosity), and sample characteristics (clonality, contamination); and technical parameters, including reads characteristics (sequencing errors, coverage, base qualities). No single simulator provided complete control over both biological and technical parameters, nor guidance on tuning biological parameters for cancer-specific simulations. Consequently, we developed MOV&RSim, a novel simulator that leverages data-driven information to set variants and reads characteristics, producing realistic tumoral samples, and providing full control on biological and technical parameters. Additionally, we leveraged well-annotated variant databases to create cancer-specific presets that inform the simulator’s parameters for 21 cancer types. CONCLUSION: This new simulator, containerised with Docker and freely available for academic use, empowers users to define each biological parameter of a tumoral genome and faithfully replicates the variability of technical noise observed in real sequencing reads. The proposed simulator and presets represent the most adaptable and comprehensive framework currently available for generating tumor samples, enabling comprehensive benchmarking and, ultimately, the optimization of somatic variant callers across diverse cancer types. Francesca Longhin, Giacomo Baruzzo, Enidia Hazizaj, Diego Boscarino, Dino Paladin, Barbara Di Camillo |
BMC Bioinform. | 6 |
| 2024 | Machine Learning Models Highlight the Impact of Pollution and Weather Patterns on Relapse Occurrence in Multiple Sclerosis PatientsabstractMultiple Sclerosis (MS) is a chronic autoimmune and inflammatory neurological disorder characterised by episodes of symptom exacerbation, known as relapses. Relapses have been linked to environmental factors such as the weather and pollutant concentrations in the air, but the exact relationship between these phenomena is still unclear. In this study, we investigated the role of environmental factors in predicting imminent relapse occurrence in MS patients, leveraging clinical and environmental data collected over a period of one week preceeding the possible event, using data collected in the context of the H2020 BRAINTEASER project. To do this, we developed and tested a range of combinations of predictive models (logistic regression, LR; and random forest, RF) and feature selection schemes, both manual and data-driven. The RF model trained after a data-driven feature selection process based on the Variable Importance in Projection (VIP) metric yielded the best results, i.e., an AUC-ROC of 0.713 and an AUC-PR of 0.639. We identified several key predictors, including clinical variables such as time since MS onset, age at onset, diagnostic delay, and the Expanded Disability Status Scale (EDSS) score, and environmental variables such as wind speed, precipitation, NO2, PM10, average and maximum temperatures, and humidity. These findings suggest that environmental factors may be viable predictors of imminent relapse occurrence in MS. Elena Marinello, Erica Tavazzi, Enrico Longato, Pietro Bosoni, Arianna Dagliati, Mahin Vazifehdan, Riccardo Bellazzi, Isotta Trescato, Alessandro Guazzo, Martina Vettoretti, Eleonora Tavazzi, Lara Ahmad, Roberto Bergamaschi, Paola Cavalla, Umberto Manera, Adriano Chiò, Barbara Di Camillo |
BIBM | 17 |
| 2024 | iDPP@CLEF 2024: The Intelligent Disease Progression Prediction Challenge
Helena Aidos, Roberto Bergamaschi, Paola Cavalla, Adriano Chiò, Arianna Dagliati, Barbara Di Camillo, Mamede de Carvalho, Nicola Ferro 0001, Piero Fariselli, Jose Manuel García Dominguez, Sara C. Madeira, Eleonora Tavazzi |
ECIR (6) | 6 |
| 2024 | From translational bioinformatics computational methodologies to personalized medicine
Barbara Di Camillo, Rosalba Giugno |
J. Biomed. Informatics | 1 |
| 2024 | DYNAMITE: Integrating Archetypal Analysis and Process Mining for Interpretable Disease Progression ModellingabstractDYNAMITE, an acronym for DYNamic Archetypal analysis for MIning disease TrajEctories, is a new methodology developed specifically to model disease progression by exploiting information available in longitudinal clinical datasets. First, archetypal analysis is applied to data organised in matrix form, with the aim of finding extreme and representative disease states (archetypes) linked to the original data through convex coefficients. Then, each original observation is associated with a single archetype based on their similarity; finally, an event log is created encoding the progression of disease states for each patient in terms of archetype states. In the last stage of the procedure, archetypal analysis is coupled with process mining, which allows the event log archetypes to be visualised graphically as sequences of disease states, allowing the clinical trajectories of patients to be extracted and examined. As a proof of concept, we applied the proposed method to data from a cohort of amyotrophic lateral sclerosis patients whose progression was monitored using the 12-item ALSFRS-R questionnaire. Without any a priori knowledge, DYNAMITE identified six archetypes clearly describing different types and severity of impairment and provided reliable clinical trajectories consistent with the prognosis of amyotrophic lateral sclerosis patients. DYNAMITE offers high interpretability at every stage of the analysis, which makes it particularly suitable for use in healthcare where explainability is paramount, and enables analysis of clinical trajectories at both individual and population levels. Isotta Trescato, Erica Tavazzi, Martina Vettoretti, Roberto Gatta, Rosario Vasta, Adriano Chiò, Barbara Di Camillo |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | Dealing with Data Scarcity in Rare Diseases: Dynamic Bayesian Networks and Transfer Learning to Develop Prognostic Models of Amyotrophic Lateral Sclerosis
Enrico Longato, Erica Tavazzi, Adriano Chiò, Gabriele Mora, Giovanni Sparacino, Barbara Di Camillo |
AIME | 6 |
| 2023 | iDPP@CLEF 2023: The Intelligent Disease Progression Prediction Challenge
Helena Aidos, Roberto Bergamaschi, Paola Cavalla, Adriano Chiò, Arianna Dagliati, Barbara Di Camillo, Mamede de Carvalho, Nicola Ferro 0001, Piero Fariselli, Jose Manuel García Dominguez, Sara C. Madeira, Eleonora Tavazzi |
ECIR (3) | 6 |
| 2023 | Artificial intelligence and statistical methods for stratification and prediction of progression in amyotrophic lateral sclerosis: A systematic reviewabstractBACKGROUND: Amyotrophic Lateral Sclerosis (ALS) is a fatal neurodegenerative disorder characterised by the progressive loss of motor neurons in the brain and spinal cord. The fact that ALS's disease course is highly heterogeneous, and its determinants not fully known, combined with ALS's relatively low prevalence, renders the successful application of artificial intelligence (AI) techniques particularly arduous. OBJECTIVE: This systematic review aims at identifying areas of agreement and unanswered questions regarding two notable applications of AI in ALS, namely the automatic, data-driven stratification of patients according to their phenotype, and the prediction of ALS progression. Differently from previous works, this review is focused on the methodological landscape of AI in ALS. METHODS: We conducted a systematic search of the Scopus and PubMed databases, looking for studies on data-driven stratification methods based on unsupervised techniques resulting in (A) automatic group discovery or (B) a transformation of the feature space allowing patient subgroups to be identified; and for studies on internally or externally validated methods for the prediction of ALS progression. We described the selected studies according to the following characteristics, when applicable: variables used, methodology, splitting criteria and number of groups, prediction outcomes, validation schemes, and metrics. RESULTS: Of the starting 1604 unique reports (2837 combined hits between Scopus and PubMed), 239 were selected for thorough screening, leading to the inclusion of 15 studies on patient stratification, 28 on prediction of ALS progression, and 6 on both stratification and prediction. In terms of variables used, most stratification and prediction studies included demographics and features derived from the ALSFRS or ALSFRS-R scores, which were also the main prediction targets. The most represented stratification methods were K-means, and hierarchical and expectation-maximisation clustering; while random forests, logistic regression, the Cox proportional hazard model, and various flavours of deep learning were the most widely used prediction methods. Predictive model validation was, albeit unexpectedly, quite rarely performed in absolute terms (leading to the exclusion of 78 eligible studies), with the overwhelming majority of included studies resorting to internal validation only. CONCLUSION: This systematic review highlighted a general agreement in terms of input variable selection for both stratification and prediction of ALS progression, and in terms of prediction targets. A striking lack of validated models emerged, as well as a general difficulty in reproducing many published studies, mainly due to the absence of the corresponding parameter lists. While deep learning seems promising for prediction applications, its superiority with respect to traditional methods has not been established; there is, instead, ample room for its application in the subfield of patient stratification. Finally, an open question remains on the role of new environmental and behavioural variables collected via novel, real-time sensors. Erica Tavazzi, Enrico Longato, Martina Vettoretti, Helena Aidos, Isotta Trescato, Chiara Roversi, Andreia S. Martins, Eduardo N. Castanho, Ruben Branco, Diogo F. Soares, Alessandro Guazzo, Giovanni Birolo, Daniele Pala, Pietro Bosoni, Adriano Chiò, Umberto Manera, Mamede de Carvalho, Bruno Miranda, Marta Gromicho, Inês Alves, Riccardo Bellazzi, Arianna Dagliati, Piero Fariselli, Sara C. Madeira, Barbara Di Camillo |
Artif. Intell. Medicine | 25 |
| 2022 | Identify, quantify and characterize cellular communication from single-cell RNA sequencing data with scSeqCommabstractMOTIVATION: Recently, single-cell RNA-seq (scRNA-seq) data have been used to study cellular communication. Most bioinformatics methods infer only the intercellular signaling between groups of cells, mainly exploiting ligand-receptor expression levels. Only few methods consider the entire intercellular + intracellular signaling, mainly inferring lists/networks of signaling involved genes. RESULTS: Here, we present scSeqComm, a computational method to identify and quantify the evidence of ongoing intercellular and intracellular signaling from scRNA-seq data, and at the same time providing a functional characterization of the inferred cellular communication. The possibility to quantify the evidence of ongoing communication assists the prioritization of the results, while the combined evidence of both intercellular and intracellular signaling increase the reliability of inferred communication. The application to a scRNA-seq dataset of tumor microenvironment, the agreement with independent bioinformatics analysis, the validation using spatial transcriptomics data and the comparison with state-of-the-art intercellular scoring schemes confirmed the robustness and reliability of the proposed method. AVAILABILITY AND IMPLEMENTATION: scSeqComm R package is freely available at https://gitlab.com/sysbiobig/scseqcomm and https://sysbiobig.dei.unipd.it/software/#scSeqComm. Submitted software version and test data are available in Zenodo, at https://dx.doi.org/10.5281/zenodo.5833298. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Giacomo Baruzzo, Giulia Cesaro, Barbara Di Camillo |
Bioinform. | 3 |
| 2022 | From translational bioinformatics computational methodologies to personalized medicine
Barbara Di Camillo, Rosalba Giugno |
J. Biomed. Informatics | 1 |
| 2022 | Investigating differential abundance methods in microbiome data: A benchmark studyabstractThe development of increasingly efficient and cost-effective high throughput DNA sequencing techniques has enhanced the possibility of studying complex microbial systems. Recently, researchers have shown great interest in studying the microorganisms that characterise different ecological niches. Differential abundance analysis aims to find the differences in the abundance of each taxa between two classes of subjects or samples, assigning a significance value to each comparison. Several bioinformatic methods have been specifically developed, taking into account the challenges of microbiome data, such as sparsity, the different sequencing depth constraint between samples and compositionality. Differential abundance analysis has led to important conclusions in different fields, from health to the environment. However, the lack of a known biological truth makes it difficult to validate the results obtained. In this work we exploit metaSPARSim, a microbial sequencing count data simulator, to simulate data with differential abundance features between experimental groups. We perform a complete comparison of recently developed and established methods on a common benchmark with great effort to the reliability of both the simulated scenarios and the evaluation metrics. The performance overview includes the investigation of numerous scenarios, studying the effect on methods' results on the main covariates such as sample size, percentage of differentially abundant features, sequencing depth, feature variability, normalisation approach and ecological niches. Mainly, we find that methods show a good control of the type I error and, generally, also of the false discovery rate at high sample size, while recall seem to depend on the dataset and sample size. Marco Cappellato, Giacomo Baruzzo, Barbara Di Camillo |
PLoS Comput. Biol. | 3 |
| 2022 | Guest Editorial: Deep Learning For GenomicsabstractThe six papers in this special section focus on deep learning for genomics. Thanks to the development of high-throughput technologies, a huge amount of omics data is being produced relative to DNA and RNA sequences and (and also) abundance at individual subject or even at individual cell level. In particular, the genomics field is rich in data thanks to the rapid reduction in the cost of genetic sequencing. On the other hand, deep learning is transforming the field of many machine learning applications, such as computer vision and natural language processing, by effectively leveraging on big amount of data and is now emerging as a promising approach for many genomics modeling tasks. The scope of this special section is to discuss novel algorithms, methodologies and applications of deep learning to genomic studies with focus on their potentialities and challenges. Barbara Di Camillo, Giuseppe Nicosia |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Recurrent Neural Network to Predict Renal Function Impairment in Diabetic Patients via Longitudinal Routine Check-up Data
Enrico Longato, Gian Paolo Fadini, Giovanni Sparacino, Angelo Avogaro, Barbara Di Camillo |
AIME | 5 |
| 2021 | Comparing the Predictive Power of Heart Failure Hospitalisation Risk Scores in the Diabetic Outpatient Clinic and Primary Care SettingsabstractThe organisation of care at diabetes outpatient clinics is typically different from that delivered by general practitioners, it is thus of interest to assess whether there is also a difference in the predictive power of heart failure hospitalisation risk scores developed independently for each subpopulation. To such a purpose, a diabetes outpatient clinic dataset and a primary care dataset were considered. A Cox proportional hazard model, an accelerated failure time model, a logistic regression, a random forest, and a K-nearest neighbours model were trained in each dataset and tested on both. The UK Prospective Diabetes Study (UKPDS) risk engine was used as benchmark. Results show that models developed using primary care data performed well on the corresponding test set but poorly when used in the diabetes outpatient clinic setting (best C-Index = 0.759vs. 0.615, best AUROC = 0.757vs. 0.598). Models trained on the diabetes outpatient clinic data performed well on the corresponding test set, and their predictive power in the primary care setting was not statistically different from the one of models developed using primary care data (best C-Index = 0.814 vs 0.740, best AUROC = 0.812 vs 0.750). In both settings UKPDS had lower predictive power than the best newly-developed models. Different care setting led to a difference in the predictive power of heart failure hospitalisation risk scores that depended on both the data used for training and the methodological approach chosen. This suggests the need to consider these factors when applying risk scores to a target population where the expected incidence of the outcome and the distribution of baseline covariates differ from those of the population for which scores were proposed. Alessandro Guazzo, Alessandro Battaggia, Enrico Longato, Bruno Franco-Novelletto, Angelo Avogaro, Gian Paolo Fadini, Maurizio Cancian, Barbara Di Camillo, Giovanni Sparacino, Massimo Fusello |
BIBM | 8 |
| 2021 | Comparison of microbiome samples: methods and computational challengesabstractThe study of microbial communities crucially relies on the comparison of metagenomic next-generation sequencing data sets, for which several methods have been designed in recent years. Here, we review three key challenges in the comparison of such data sets: species identification and quantification, the efficient computation of distances between metagenomic samples and the identification of metagenomic features associated with a phenotype such as disease status. We present current solutions for such challenges, considering both reference-based methods relying on a database of reference genomes and reference-free methods working directly on all sequencing reads from the samples. Matteo Comin, Barbara Di Camillo, Cinzia Pizzi, Fabio Vandin |
Briefings Bioinform. | 2 |
| 2021 | Beware to ignore the rare: how imputing zero-values can improve the quality of 16S rRNA gene studies resultsabstractBACKGROUND: 16S rRNA-gene sequencing is a valuable approach to characterize the taxonomic content of the whole bacterial population inhabiting a metabolic and spatial niche, providing an important opportunity to study bacteria and their role in many health and environmental mechanisms. The analysis of data produced by amplicon sequencing, however, brings very specific methodological issues that need to be properly addressed to obtain reliable biological conclusions. Among these, 16S count data tend to be very sparse, with many null values reflecting species that are present but got unobserved due to the multiplexing constraints. However, current data workflows do not consider a step in which the information about unobserved species is recovered. RESULTS: In this work, we evaluate for the first time the effects of introducing in the 16S data workflow a new preprocessing step, zero-imputation, to recover this lost information. Due to the lack of published zero-imputation methods specifically designed for 16S count data, we considered a set of zero-imputation strategies available for other frameworks, and benchmarked them using in silico 16S count data reflecting different experimental designs. Additionally, we assessed the effect of combining zero-imputation and normalization, i.e. the only preprocessing step in current 16S workflow. Overall, we benchmarked 35 16S preprocessing pipelines assessing their ability to handle data sparsity, identify species presence/absence, recovery sample proportional abundance distributions, and improve typical downstream analyses such as computation of alpha and beta diversity indices and differential abundance analysis. CONCLUSIONS: The results clearly show that 16S data analysis greatly benefits from a properly-performed zero-imputation step, despite the choice of the right zero-imputation method having a pivotal role. In addition, we identify a set of best-performing pipelines that could be a valuable indication for data analysts. Giacomo Baruzzo, Ilaria Patuzzi, Barbara Di Camillo |
BMC Bioinform. | 3 |
| 2021 | Mathematical modelling of SigE regulatory network reveals new insights into bistability of mycobacterial stress responseabstractBACKGROUND: The ability to rapidly adapt to adverse environmental conditions represents the key of success of many pathogens and, in particular, of Mycobacterium tuberculosis. Upon exposition to heat shock, antibiotics or other sources of stress, appropriate responses in terms of genes transcription and proteins activity are activated leading part of a genetically identical bacterial population to express a different phenotype, namely to develop persistence. When the stress response network is mathematically described by an ordinary differential equations model, development of persistence in the bacterial population is associated with bistability of the model, since different emerging phenotypes are represented by different stable steady states. RESULTS: In this work, we develop a mathematical model of SigE stress response network that incorporates interactions not considered in mathematical models currently available in the literature. We provide, through involved analytical computations, accurate approximations of the system's nullclines, and exploit the obtained expressions to determine, in a reliable though computationally efficient way, the number of equilibrium points of the system. CONCLUSIONS: Theoretical analysis and perturbation experiments point out the crucial role played by the degradation pathway involving RseA, the anti-sigma factor of SigE, for coexistence of two stable equilibria and the emergence of bistability. Our results also indicate that a fine control on RseA concentration is a necessary requirement in order for the system to exhibit bistability. Irene Zorzan, Simone Del Favero, Alberto Giaretta 0002, Riccardo Manganelli, Barbara Di Camillo, Luca Schenato 0001 |
BMC Bioinform. | 5 |
| 2021 | A Deep Learning Approach to Predict Diabetes' Cardiovascular Complications From Administrative ClaimsabstractPeople with diabetes require lifelong access to healthcare services to delay the onset of complications. Their disease management processes generate great volumes of data across several domains, from clinical to administrative. Difficulties in accessing and processing these data hinder their secondary use in an institutional setting, even for highly desirable applications, such as the prediction of cardiovascular disease, the main driver of excess mortality in diabetes. Hence, in the present work, we propose a deep learning model for the prediction of major adverse cardiovascular events (MACE), developed and validated using the administrative claims of 214,676 diabetic patients of the Veneto region, in North East Italy. Specifically, we use a year of pharmacy and hospitalisation claims, together with basic patient's information, to predict the 4P-MACE composite endpoint, i.e., the first occurrence of death, heart failure, myocardial infarction, or stroke, with a variable prediction horizon of 1 to 5 years. Adapting to the time-to-event nature of this task, we cast our problem as a multi-outcome (4P-MACE and components), multi-label (1 to 5 years) classification task with a custom loss to account for the effect of censoring. Our model, purposefully specified to minimise data preparation costs, exhibits satisfactory performance in predicting 4P-MACE at all prediction horizons: AUROC from 0.812 (C.I.: 0.797 - 0.827) to 0.792 (C.I.: 0.781 - 0.802); C-index from 0.802 (C.I.: 0.788 - 0.816) to 0.770 (C.I.: 0.761 - 0.779). Components' prediction performance is also adequate, ranging from death's 0.877 1-year AUROC to stroke's 0.689 5-year AUROC. Enrico Longato, Gian Paolo Fadini, Giovanni Sparacino, Angelo Avogaro, Lara Tramontan, Barbara Di Camillo |
IEEE J. Biomed. Health Informatics | 6 |
| 2020 | SPARSim single cell: a count data simulator for scRNA-seq dataabstractMOTIVATION: Single cell RNA-seq (scRNA-seq) count data show many differences compared with bulk RNA-seq count data, making the application of many RNA-seq pre-processing/analysis methods not straightforward or even inappropriate. For this reason, the development of new methods for handling scRNA-seq count data is currently one of the most active research fields in bioinformatics. To help the development of such new methods, the availability of simulated data could play a pivotal role. However, only few scRNA-seq count data simulators are available, often showing poor or not demonstrated similarity with real data. RESULTS: In this article we present SPARSim, a scRNA-seq count data simulator based on a Gamma-Multivariate Hypergeometric model. We demonstrate that SPARSim allows to generate count data that resemble real data in terms of count intensity, variability and sparsity, performing comparably or better than one of the most used scRNA-seq simulator, Splat. In particular, SPARSim simulated count matrices well resemble the distribution of zeros across different expression intensities observed in real count data. AVAILABILITY AND IMPLEMENTATION: SPARSim R package is freely available at http://sysbiobig.dei.unipd.it/? q=SPARSim and at https://gitlab.com/sysbiobig/sparsim. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Giacomo Baruzzo, Ilaria Patuzzi, Barbara Di Camillo |
Bioinform. | 3 |
| 2020 | A practical perspective on the concordance index for the evaluation and selection of prognostic time-to-event models
Enrico Longato, Martina Vettoretti, Barbara Di Camillo |
J. Biomed. Informatics | 3 |
| 2020 | Guest Editorial Data Science in Smart Healthcare: Challenges and OpportunitiesabstractThe fifteen articles in this special section focus on data science used in smart healthcare applications. A shift toward a data-driven socio-economic health model is occurring. This is the result of the increased volume, velocity and variety of data collected from the public and private sector in healthcare, and biology in general. In the past five-years, there has been an impressive development of computational intelligence and informatics methods for application to health and biomedical science. However, the effective use of data to address the scale and scope of human health problems has yet to realize its full potential. The barriers limiting the impact of practical application of standard data mining and machine learning methods have been inherent to the characteristics of health data. Besides the volume of the data (‘big data’), these are challenging due to their heterogeneity, complexity, variability and dynamic nature. Finally, data management and interpretability of the results have been limited by practical challenges in implementing new and also existing standards across the different health providers and research institutions. The scope of this Special issue is to discuss some of these challenges and opportunities in health and biological data science, with particular focus on the infrastructure, software, methods and algorithms needed to analyze large datasets in biological and clinical research. Barbara Di Camillo, Giuseppe Nicosia, Francesca Buffa, Benny P. L. Lo |
IEEE J. Biomed. Health Informatics | 1 |
| 2019 | Deep convolutional neural network for survival estimation of Amyotrophic Lateral Sclerosis patients
Enrico Grisan, Alessandro Zandonà, Barbara Di Camillo |
ESANN | 3 |
| 2019 | How to design a single-cell RNA-sequencing experiment: pitfalls, challenges and perspectivesabstractThe sequencing of the transcriptome of single cells, or single-cell RNA-sequencing, has now become the dominant technology for the identification of novel cell types in heterogeneous cell populations or for the study of stochastic gene expression. In recent years, various experimental methods and computational tools for analysing single-cell RNA-sequencing data have been proposed. However, most of them are tailored to different experimental designs or biological questions, and in many cases, their performance has not been benchmarked yet, thus increasing the difficulty for a researcher to choose the optimal single-cell transcriptome sequencing (scRNA-seq) experiment and analysis workflow. In this review, we aim to provide an overview of the current available experimental and computational methods developed to handle single-cell RNA-sequencing data and, based on their peculiarities, we suggest possible analysis frameworks depending on specific experimental designs. Together, we propose an evaluation of challenges and open questions and future perspectives in the field. In particular, we go through the different steps of scRNA-seq experimental protocols such as cell isolation, messenger RNA capture, reverse transcription, amplification and use of quantitative standards such as spike-ins and Unique Molecular Identifiers (UMIs). We then analyse the current methodological challenges related to preprocessing, alignment, quantification, normalization, batch effect correction and methods to control for confounding effects. Alessandra Dal Molin, Barbara Di Camillo |
Briefings Bioinform. | 2 |
| 2019 | metaSPARSim: a 16S rRNA gene sequencing count data simulatorabstractBACKGROUND: In the last few years, 16S rRNA gene sequencing (16S rDNA-seq) has seen a surprisingly rapid increase in election rate as a methodology to perform microbial community studies. Despite the considerable popularity of this technique, an exiguous number of specific tools are currently available for proper 16S rDNA-seq count data preprocessing and simulation. Indeed, the great majority of tools have been developed adapting methodologies previously used for bulk RNA-seq data, with poor assessment of their applicability in the metagenomics field. For such tools and the few ones specifically developed for 16S rDNA-seq data, performance assessment is challenging, mainly due to the complex nature of the data and the lack of realistic simulation models. In fact, to the best of our knowledge, no software thought for data simulation are available to directly obtain synthetic 16S rDNA-seq count tables that properly model heavy sparsity and compositionality typical of these data. RESULTS: In this paper we present metaSPARSim, a sparse count matrix simulator intended for usage in development of 16S rDNA-seq metagenomic data processing pipelines. metaSPARSim implements a new generative process that models the sequencing process with a Multivariate Hypergeometric distribution in order to realistically simulate 16S rDNA-seq count table, resembling real experimental data compositionality and sparsity. It provides ready-to-use count matrices and comes with the possibility to reproduce different pre-coded scenarios and to estimate simulation parameters from real experimental data. The tool is made available at http://sysbiobig.dei.unipd.it/?q=Software#metaSPARSimand https://gitlab.com/sysbiobig/metasparsim. CONCLUSION: metaSPARSim is able to generate count matrices resembling real 16S rDNA-seq data. The availability of count data simulators is extremely valuable both for methods developers, for which a ground truth for tools validation is needed, and for users who want to assess state of the art analysis tools for choosing the most accurate one. Thus, we believe that metaSPARSim is a valuable tool for researchers involved in developing, testing and using robust and reliable data analysis methods in the context of 16S rRNA gene sequencing. Ilaria Patuzzi, Giacomo Baruzzo, Carmen Losasso, Antonia Ricci, Barbara Di Camillo |
BMC Bioinform. | 5 |
| 2019 | A Dynamic Bayesian Network model for the simulation of Amyotrophic Lateral Sclerosis progressionabstractBACKGROUND: Amyotrophic lateral sclerosis (ALS) is an adult-onset neurodegenerative disease progressively affecting upper and lower motor neurons in the brain and spinal cord. Mean life expectancy is three to five years, with paralysis of muscles, respiratory failure and loss of vital functions being the common causes of death. Clinical manifestations of ALS are heterogeneous due to the mix of anatomic regions involvement and the variability in disease course; consequently, diagnosis and prognosis at the level of individual patient is really challenging. Prediction of ALS progression and stratification of patients into meaningful subgroups have been long-standing interests to clinical practice, research and drug development. METHODS: We developed a Dynamic Bayesian Network (DBN) model on more than 4500 ALS patients included in the Pooled Resource Open-Access ALS Clinical Trials Database (PRO-ACT), in order to detect probabilistic relationships among clinical variables and identify risk factors related to survival and loss of vital functions. Furthermore, the DBN was used to simulate the temporal evolution of an ALS cohort predicting survival and the time to impairment of vital functions (communication, swallowing, gait and respiration). A first attempt to stratify patients by risk factors and simulate the progression of ALS subgroups was also implemented. RESULTS: The DBN model provided the prediction of ALS most probable trajectories over time in terms of important clinical outcomes, including survival and loss of autonomy in functional domains. Furthermore, it allowed the identification of biomarkers related to patients' clinical status as well as vital functions, and unrevealed their probabilistic relationships. For instance, DBN found that bicarbonate and calcium levels influence survival time; moreover, the model evidenced dependencies over time among phosphorus level, movement impairment and creatinine. Finally, our model provided a tool to stratify patients into subgroups of different prognosis studying the effect of specific variables, or combinations of them, on either survival time or time to loss of autonomy in specific functional domains. CONCLUSIONS: The analysis of the risk factors and the simulation allowed by our DBN model might enable better support for ALS prognosis as well as a deeper insight into disease manifestations, in a context of a personalized medicine approach. Alessandro Zandonà, Rosario Vasta, Adriano Chiò, Barbara Di Camillo |
BMC Bioinform. | 4 |
| 2018 | Measuring the diversity of the human microbiota with targeted next-generation sequencingabstractThe human microbiota is a complex ecological community of commensal, symbiotic and pathogenic microorganisms harboured by the human body. Next-generation sequencing (NGS) technologies, in particular targeted amplicon sequencing of the 16S ribosomal RNA gene (16S-seq), are enabling the identification and quantification of human-resident microorganisms at unprecedented resolution, providing novel insights into the role of the microbiota in health and disease. Once microbial abundances are quantified through NGS data analysis, diversity indices provide valuable mathematical tools to describe the ecological complexity of a single sample or to detect species differences between samples. However, diversity is not a determined physical quantity for which a consensus definition and unit of measure have been established, and several diversity indices are currently available. Furthermore, they were originally developed for macroecology and their robustness to the possible bias introduced by sequencing has not been characterized so far. To assist the reader with the selection and interpretation of diversity measures, we review a panel of broadly used indices, describing their mathematical formulations, purposes and properties, and characterize their behaviour and criticalities in dependence of the data features using simulated data as ground truth. In addition, we make available an R package, DiversitySeq, which implements in a unified framework the full panel of diversity indices and a simulator of 16S-seq data, and thus represents a valuable resource for the analysis of diversity from NGS count data and for the benchmarking of computational methods for 16S-seq. Francesca Finotello, Eleonora Mastrorilli, Barbara Di Camillo |
Briefings Bioinform. | 3 |
| 2018 | Optimizing PCR primers targeting the bacterial 16S ribosomal RNA geneabstractBACKGROUND: Targeted amplicon sequencing of the 16S ribosomal RNA gene is one of the key tools for studying microbial diversity. The accuracy of this approach strongly depends on the choice of primer pairs and, in particular, on the balance between efficiency, specificity and sensitivity in the amplification of the different bacterial 16S sequences contained in a sample. There is thus the need for computational methods to design optimal bacterial 16S primers able to take into account the knowledge provided by the new sequencing technologies. RESULTS: We propose here a computational method for optimizing the choice of primer sets, based on multi-objective optimization, which simultaneously: 1) maximizes efficiency and specificity of target amplification; 2) maximizes the number of different bacterial 16S sequences matched by at least one primer; 3) minimizes the differences in the number of primers matching each bacterial 16S sequence. Our algorithm can be applied to any desired amplicon length without affecting computational performance. The source code of the developed algorithm is released as the mopo16S software tool (Multi-Objective Primer Optimization for 16S experiments) under the GNU General Public License and is available at http://sysbiobig.dei.unipd.it/?q=Software#mopo16S . CONCLUSIONS: Results show that our strategy is able to find better primer pairs than the ones available in the literature according to all three optimization criteria. We also experimentally validated three of the primer pairs identified by our method on multiple bacterial species, belonging to different genera and phyla. Results confirm the predicted efficiency and the ability to maximize the number of different bacterial 16S sequences matched by primers. Francesco Sambo, Francesca Finotello, Enrico Lavezzo, Giacomo Baruzzo, Giulia Masi, Elektra Peta, Marco Falda, Stefano Toppo, Luisa Barzon, Barbara Di Camillo |
BMC Bioinform. | 10 |
| 2017 | bnstruct: an R package for Bayesian Network structure learning in the presence of missing dataabstractMotivation: A Bayesian Network is a probabilistic graphical model that encodes probabilistic dependencies between a set of random variables. We introduce bnstruct, an open source R package to (i) learn the structure and the parameters of a Bayesian Network from data in the presence of missing values and (ii) perform reasoning and inference on the learned Bayesian Networks. To the best of our knowledge, there is no other open source software that provides methods for all of these tasks, particularly the manipulation of missing data, which is a common situation in practice. Availability and Implementation: The software is implemented in R and C and is available on CRAN under a GPL licence. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Alberto Franzin, Francesco Sambo, Barbara Di Camillo |
Bioinform. | 3 |
| 2015 | A Bayesian Network for Probabilistic Reasoning and Imputation of Missing Risk Factors in Type 2 Diabetes
Francesco Sambo, Andrea Facchinetti, Liisa Hakaste, Jasmina Kravic, Barbara Di Camillo, Giuseppe Fico, Jaakko Tuomilehto, Leif Groop, Rafael Gabriel, Tuomi Tiinamaija, Claudio Cobelli |
AIME | 5 |
| 2015 | A Dynamic Bayesian Network model for long-term simulation of clinical complications in type 1 diabetes
Simone Marini, Emanuele Trifoglio, Nicola Barbarini, Francesco Sambo, Barbara Di Camillo, Alberto Malovini, Marco Manfrini, Claudio Cobelli, Riccardo Bellazzi |
J. Biomed. Informatics | 5 |
| 2014 | ABACUS: an entropy-based cumulative bivariate statistic robust to rare variants and different direction of genotype effectabstractMOTIVATION: In the past years, both sequencing and microarray have been widely used to search for relations between genetic variations and predisposition to complex pathologies such as diabetes or neurological disorders. These studies, however, have been able to explain only a small fraction of disease heritability, possibly because complex pathologies cannot be referred to few dysfunctional genes, but are rather heterogeneous and multicausal, as a result of a combination of rare and common variants possibly impairing multiple regulatory pathways. Rare variants, though, are difficult to detect, especially when the effects of causal variants are in different directions, i.e. with protective and detrimental effects. RESULTS: Here, we propose ABACUS, an Algorithm based on a BivAriate CUmulative Statistic to identify single nucleotide polymorphisms (SNPs) significantly associated with a disease within predefined sets of SNPs such as pathways or genomic regions. ABACUS is robust to the concurrent presence of SNPs with protective and detrimental effects and of common and rare variants; moreover, it is powerful even when few SNPs in the SNP-set are associated with the phenotype. We assessed ABACUS performance on simulated and real data and compared it with three state-of-the-art methods. When ABACUS was applied to type 1 and 2 diabetes data, besides observing a wide overlap with already known associations, we found a number of biologically sound pathways, which might shed light on diabetes mechanism and etiology. AVAILABILITY AND IMPLEMENTATION: ABACUS is available at http://www.dei.unipd.it/∼dicamill/pagine/Software.html. Barbara Di Camillo, Francesco Sambo, Gianna Toffolo, Claudio Cobelli |
Bioinform. | 1 |
| 2014 | Compression and fast retrieval of SNP dataabstractMOTIVATION: The increasing interest in rare genetic variants and epistatic genetic effects on complex phenotypic traits is currently pushing genome-wide association study design towards datasets of increasing size, both in the number of studied subjects and in the number of genotyped single nucleotide polymorphisms (SNPs). This, in turn, is leading to a compelling need for new methods for compression and fast retrieval of SNP data. RESULTS: We present a novel algorithm and file format for compressing and retrieving SNP data, specifically designed for large-scale association studies. Our algorithm is based on two main ideas: (i) compress linkage disequilibrium blocks in terms of differences with a reference SNP and (ii) compress reference SNPs exploiting information on their call rate and minor allele frequency. Tested on two SNP datasets and compared with several state-of-the-art software tools, our compression algorithm is shown to be competitive in terms of compression rate and to outperform all tools in terms of time to load compressed data. AVAILABILITY AND IMPLEMENTATION: Our compression and decompression algorithms are implemented in a C++ library, are released under the GNU General Public License and are freely downloadable from http://www.dei.unipd.it/~sambofra/snpack.html. Francesco Sambo, Barbara Di Camillo, Gianna Toffolo, Claudio Cobelli |
Bioinform. | 2 |
| 2014 | Reducing bias in RNA sequencing data: a novel approach to compute countsabstractBACKGROUND: In the last decade, Next-Generation Sequencing technologies have been extensively applied to quantitative transcriptomics, making RNA sequencing a valuable alternative to microarrays for measuring and comparing gene transcription levels. Although several methods have been proposed to provide an unbiased estimate of transcript abundances through data normalization, all of them are based on an initial count of the total number of reads mapping on each transcript. This procedure, in principle robust to random noise, is actually error-prone if reads are not uniformly distributed along sequences, as happens indeed due to sequencing errors and ambiguity in read mapping. Here we propose a new approach, called maxcounts, to quantify the expression assigned to an exon as the maximum of its per-base counts, and we assess its performance in comparison with the standard approach described above, which considers the total number of reads aligned to an exon. The two measures are compared using multiple data sets and considering several evaluation criteria: independence from gene-specific covariates, such as exon length and GC-content, accuracy and precision in the quantification of true concentrations and robustness of measurements to variations of alignments quality. RESULTS: Both measures show high accuracy and low dependency on GC-content. However, maxcounts expression quantification is less biased towards long exons with respect to the standard approach. Moreover, it shows lower technical variability at low expressions and is more robust to variations in the quality of alignments. CONCLUSIONS: In summary, we confirm that counts computed with the standard approach depend on the length of the feature they are summarized on, and are sensitive to the non-uniform distribution of reads along transcripts. On the opposite, maxcounts are robust to biases due to the non-uniformity distribution of reads and are characterized by a lower technical variability. Hence, we propose maxcounts as an alternative approach for quantitative RNA-sequencing applications. Francesca Finotello, Enrico Lavezzo, Luca Bianco, Luisa Barzon, Paolo Mazzon, Paolo Fontana, Stefano Toppo, Barbara Di Camillo |
BMC Bioinform. | 8 |
| 2012 | Discriminant functional gene groups identification with machine learning and prior knowledge
Grzegorz Zycinski, Margherita Squillario, Annalisa Barla, Tiziana Sanavia, Alessandro Verri, Barbara Di Camillo |
ESANN | 6 |
| 2012 | Comparative analysis of algorithms for whole-genome assembly of pyrosequencing dataabstractNext-generation sequencing technologies have fostered an unprecedented proliferation of high-throughput sequencing projects and a concomitant development of novel algorithms for the assembly of short reads. In this context, an important issue is the need of a careful assessment of the accuracy of the assembly process. Here, we review the efficiency of a panel of assemblers, specifically designed to handle data from GS FLX 454 platform, on three bacterial data sets with different characteristics in terms of reads coverage and repeats content. Our aim is to investigate their strengths and weaknesses in the reconstruction of the reference genomes. In our benchmarking, we assess assemblers' performance, quantifying and characterizing assembly gaps and errors, and evaluating their ability to solve complex genomic regions containing repeats. The final goal of this analysis is to highlight pros and cons of each method, in order to provide the final user with general criteria for the right choice of the appropriate assembly strategy, depending on the specific needs. A further aspect we have explored is the relationship between coverage of a sequencing project and quality of the obtained results. The final outcome suggests that, for a good tradeoff between costs and results, the planned genome coverage of an experiment should not exceed 20-30 ×. Francesca Finotello, Enrico Lavezzo, Paolo Fontana, Denis Peruzzo, Alessandro Albiero, Luisa Barzon, Marco Falda, Barbara Di Camillo, Stefano Toppo |
Briefings Bioinform. | 8 |
| 2012 | Integrating literature-constrained and data-driven inference of signalling networksabstractMOTIVATION: Recent developments in experimental methods facilitate increasingly larger signal transduction datasets. Two main approaches can be taken to derive a mathematical model from these data: training a network (obtained, e.g., from literature) to the data, or inferring the network from the data alone. Purely data-driven methods scale up poorly and have limited interpretability, whereas literature-constrained methods cannot deal with incomplete networks. RESULTS: We present an efficient approach, implemented in the R package CNORfeeder, to integrate literature-constrained and data-driven methods to infer signalling networks from perturbation experiments. Our method extends a given network with links derived from the data via various inference methods, and uses information on physical interactions of proteins to guide and validate the integration of links. We apply CNORfeeder to a network of growth and inflammatory signalling. We obtain a model with superior data fit in the human liver cancer HepG2 and propose potential missing pathways. AVAILABILITY: CNORfeeder is in the process of being submitted to Bioconductor and in the meantime available at www.cellnopt.org. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Federica Eduati, Javier De Las Rivas, Barbara Di Camillo, Gianna Toffolo, Julio Saez-Rodriguez |
Bioinform. | 3 |
| 2012 | Argot2: a large scale function prediction tool relying on semantic similarity of weighted Gene Ontology termsabstractBACKGROUND: Predicting protein function has become increasingly demanding in the era of next generation sequencing technology. The task to assign a curator-reviewed function to every single sequence is impracticable. Bioinformatics tools, easy to use and able to provide automatic and reliable annotations at a genomic scale, are necessary and urgent. In this scenario, the Gene Ontology has provided the means to standardize the annotation classification with a structured vocabulary which can be easily exploited by computational methods. RESULTS: Argot2 is a web-based function prediction tool able to annotate nucleic or protein sequences from small datasets up to entire genomes. It accepts as input a list of sequences in FASTA format, which are processed using BLAST and HMMER searches vs UniProKB and Pfam databases respectively; these sequences are then annotated with GO terms retrieved from the UniProtKB-GOA database and the terms are weighted using the e-values from BLAST and HMMER. The weighted GO terms are processed according to both their semantic similarity relations described by the Gene Ontology and their associated score. The algorithm is based on the original idea developed in a previous tool called Argot. The entire engine has been completely rewritten to improve both accuracy and computational efficiency, thus allowing for the annotation of complete genomes. CONCLUSIONS: The revised algorithm has been already employed and successfully tested during in-house genome projects of grape and apple, and has proven to have a high precision and recall in all our benchmark conditions. It has also been successfully compared with Blast2GO, one of the methods most commonly employed for sequence annotation. The server is freely accessible at http://www.medcomp.medicina.unipd.it/Argot2. Marco Falda, Stefano Toppo, Alessandro Pescarolo, Enrico Lavezzo, Barbara Di Camillo, Andrea Facchinetti, Elisa Cilia, Riccardo Velasco, Paolo Fontana |
BMC Bioinform. | 5 |
| 2012 | Bag of Naïve Bayes: biomarker selection and classification from genome-wide SNP dataabstractBACKGROUND: Multifactorial diseases arise from complex patterns of interaction between a set of genetic traits and the environment. To fully capture the genetic biomarkers that jointly explain the heritability component of a disease, thus, all SNPs from a genome-wide association study should be analyzed simultaneously. RESULTS: In this paper, we present Bag of Naïve Bayes (BoNB), an algorithm for genetic biomarker selection and subjects classification from the simultaneous analysis of genome-wide SNP data. BoNB is based on the Naïve Bayes classification framework, enriched by three main features: bootstrap aggregating of an ensemble of Naïve Bayes classifiers, a novel strategy for ranking and selecting the attributes used by each classifier in the ensemble and a permutation-based procedure for selecting significant biomarkers, based on their marginal utility in the classification process. BoNB is tested on the Wellcome Trust Case-Control study on Type 1 Diabetes and its performance is compared with the ones of both a standard Naïve Bayes algorithm and HyperLASSO, a penalized logistic regression algorithm from the state-of-the-art in simultaneous genome-wide data analysis. CONCLUSIONS: The significantly higher classification accuracy obtained by BoNB, together with the significance of the biomarkers identified from the Type 1 Diabetes dataset, prove the effectiveness of BoNB as an algorithm for both classification and biomarker selection from genome-wide SNP data. AVAILABILITY: Source code of the BoNB algorithm is released under the GNU General Public Licence and is available at http://www.dei.unipd.it/~sambofra/bonb.html. Francesco Sambo, Emanuele Trifoglio, Barbara Di Camillo, Gianna Toffolo, Claudio Cobelli |
BMC Bioinform. | 3 |
| 2012 | Improving biomarker list stability by integration of biological knowledge in the learning processabstractBACKGROUND: The identification of robust lists of molecular biomarkers related to a disease is a fundamental step for early diagnosis and treatment. However, methodologies for biomarker discovery using microarray data often provide results with limited overlap. It has been suggested that one reason for these inconsistencies may be that in complex diseases, such as cancer, multiple genes belonging to one or more physiological pathways are associated with the outcomes. Thus, a possible approach to improve list stability is to integrate biological information from genomic databases in the learning process; however, a comprehensive assessment based on different types of biological information is still lacking in the literature. In this work we have compared the effect of using different biological information in the learning process like functional annotations, protein-protein interactions and expression correlation among genes. RESULTS: Biological knowledge has been codified by means of gene similarity matrices and expression data linearly transformed in such a way that the more similar two features are, the more closely they are mapped. Two semantic similarity matrices, based on Biological Process and Molecular Function Gene Ontology annotation, and geodesic distance applied on protein-protein interaction networks, are the best performers in improving list stability maintaining almost equal prediction accuracy. CONCLUSIONS: The performed analysis supports the idea that when some features are strongly correlated to each other, for example because are close in the protein-protein interaction network, then they might have similar importance and are equally relevant for the task at hand. Obtained results can be a starting point for additional experiments on combining similarity matrices in order to obtain even more stable lists of biomarkers. The implementation of the classification algorithm is available at the link: http://www.math.unipd.it/~dasan/biomarkers.html. Tiziana Sanavia, Fabio Aiolli, Giovanni Da San Martino, Andrea Bisognin, Barbara Di Camillo |
BMC Bioinform. | 5 |
| 2012 | Qualitative Reasoning for Biological Network Inference from Systematic Perturbation ExperimentsabstractThe systematic perturbation of the components of a biological system has been proven among the most informative experimental setups for the identification of causal relations between the components. In this paper, we present Systematic Perturbation-Qualitative Reasoning (SPQR), a novel Qualitative Reasoning approach to automate the interpretation of the results of systematic perturbation experiments. Our method is based on a qualitative abstraction of the experimental data: for each perturbation experiment, measured values of the observed variables are modeled as lower, equal or higher than the measurements in the wild type condition, when no perturbation is applied. The algorithm exploits a set of IF-THEN rules to infer causal relations between the variables, analyzing the patterns of propagation of the perturbation signals through the biological network, and is specifically designed to minimize the rate of false positives among the inferred relations. Tested on both simulated and real perturbation data, SPQR indeed exhibits a significantly higher precision than the state of the art. Silvana Badaloni, Barbara Di Camillo, Francesco Sambo |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2012 | SimBioNeT: A Simulator of Biological Network TopologyabstractStudying biological networks at topological level is a major issue in computational biology studies and simulation is often used in this context, either to assess reverse engineering algorithms or to investigate how topological properties depend on network parameters. In both contexts, it is desirable for a topology simulator to reproduce the current knowledge on biological networks, to be able to generate a number of networks with the same properties and to be flexible with respect to the possibility to mimic networks of different organisms. We propose a biological network topology simulator, SimBioNeT, in which module structures of different type and size are replicated at different level of network organization and interconnected, so to obtain the desired degree distribution, e.g., scale free, and a clustering coefficient constant with the number of nodes in the network, a typical characteristic of biological networks. Empirical assessment of the ability of the simulator to reproduce characteristic properties of biological network and comparison with E. coli and S. cerevisiae transcriptional networks demonstrates the effectiveness of our proposal. Barbara Di Camillo, Marco Falda, Gianna Toffolo, Claudio Cobelli |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2012 | MORE: Mixed Optimization for Reverse Engineering - An Application to Modeling Biological Networks Response via Sparse Systems of Nonlinear Differential EquationsabstractReverse engineering is the problem of inferring the structure of a network of interactions between biological variables from a set of observations. In this paper, we propose an optimization algorithm, called MORE, for the reverse engineering of biological networks from time series data. The model inferred by MORE is a sparse system of nonlinear differential equations, complex enough to realistically describe the dynamics of a biological system. MORE tackles separately the discrete component of the problem, the determination of the biological network topology, and the continuous component of the problem, the strength of the interactions. This approach allows us both to enforce system sparsity, by globally constraining the number of edges, and to integrate a priori information about the structure of the underlying interaction network. Experimental results on simulated and real-world networks show that the mixed discrete/continuous optimization approach of MORE significantly outperforms standard continuous optimization and that MORE is competitive with the state of the art in terms of accuracy of the inferred networks. Francesco Sambo, Marco Antonio Montes de Oca, Barbara Di Camillo, Gianna Toffolo, Thomas Stützle |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2007 | Significance analysis of microarray transcript levels in time series experimentsabstractBACKGROUND: Microarray time series studies are essential to understand the dynamics of molecular events. In order to limit the analysis to those genes that change expression over time, a first necessary step is to select differentially expressed transcripts. A variety of methods have been proposed to this purpose; however, these methods are seldom applicable in practice since they require a large number of replicates, often available only for a limited number of samples. In this data-poor context, we evaluate the performance of three selection methods, using synthetic data, over a range of experimental conditions. Application to real data is also discussed. RESULTS: Three methods are considered, to assess differentially expressed genes in data-poor conditions. Method 1 uses a threshold on individual samples based on a model of the experimental error. Method 2 calculates the area of the region bounded by the time series expression profiles, and considers the gene differentially expressed if the area exceeds a threshold based on a model of the experimental error. These two methods are compared to Method 3, recently proposed in the literature, which exploits splines fit to compare time series profiles. Application of the three methods to synthetic data indicates that Method 2 outperforms the other two both in Precision and Recall when short time series are analyzed, while Method 3 outperforms the other two for long time series. CONCLUSION: These results help to address the choice of the algorithm to be used in data-poor time series expression study, depending on the length of the time series. Barbara Di Camillo, Gianna Toffolo, Sreekumaran K. Nair, Laura J. Greenlund, Claudio Cobelli |
BMC Bioinform. | 1 |
| 2005 | A quantization method based on threshold optimization for microarray short time seriesabstractBACKGROUND: Reconstructing regulatory networks from gene expression profiles is a challenging problem of functional genomics. In microarray studies the number of samples is often very limited compared to the number of genes, thus the use of discrete data may help reducing the probability of finding random associations between genes. RESULTS: A quantization method, based on a model of the experimental error and on a significance level able to compromise between false positive and false negative classifications, is presented, which can be used as a preliminary step in discrete reverse engineering methods. The method is tested on continuous synthetic data with two discrete reverse engineering methods: Reveal and Dynamic Bayesian Networks. CONCLUSION: The quantization method, evaluated in comparison with two standard methods, 5% threshold based on experimental error and rank sorting, improves the ability of Reveal and Dynamic Bayesian Networks to identify relations among genes. Barbara Di Camillo, Fátima Sánchez-Cabo, Gianna Toffolo, Sreekumaran K. Nair, Zlatko Trajanoski, Claudio Cobelli |
BMC Bioinform. | 1 |