VLDB 2026 Research / reviewers in the wild / expert
Gastone C. Castellani
dblp:30/2650
· DBLP profile ↗
16ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0003-4892-925XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Survival Model Optimization via Federated Learning: A Study Combining Simulations and ExperimentsabstractFederated Learning is an emerging, powerful approach that allows training an artificial intelligence model in distributed setting. Two survival models, Cox and DeepSurv, have been trained in a federated setting, exploiting both code simulations and real experiments on the new platform, developed by the GenoMed4All consortium. Different scenarios have been tested by splitting a Myelodysplastic Syndrome dataset into three nodes and performing feature removal. A significant gain in model performance has been observed due to federated aggregation. Francesco Casadei, Luciana Carota, Gianluca Asti, Saverio D'Amico, Davide Piscia, Santiago Zazo, Patricia A. Apellániz, Juan Parras, Claudia Sala, Cesare Rollo, Nono S. C. Merleau, Piero Fariselli, Matteo Giovanni Della Porta, Tiziana Sanavia, Federico Álvarez García, Gastone C. Castellani, Enrico Giampieri |
IEEE Big Data | 16 |
| 2024 | Covering Hierarchical Dirichlet Mixture Models on binary data to enhance genomic stratifications in onco-hematologyabstractOnco-hematological studies are increasingly adopting statistical mixture models to support the advancement of the genomically-driven classification systems for blood cancer. Targeting enhanced patients stratification based on the sole role of molecular biology attracted much interest and contributes to bring personalized medicine closer to reality. In onco-hematology, Hierarchical Dirichlet Mixture Models (HDMM) have become one of the preferred method to cluster the genomics data, that include the presence or absence of gene mutations and cytogenetics anomalies, into components. This work unfolds the standard workflow used in onco-hematology to improve patient stratification and proposes alternative approaches to characterize the components and to assign patient to them, as they are crucial tasks usually supported by a priori clinical knowledge. We propose (a) to compute the parameters of the multinomial components of the HDMM or (b) to estimate the parameters of the HDMM components as if they were Multivariate Fisher's Non-Central Hypergeometric (MFNCH) distributions. Then, our approach to perform patients assignments to the HDMM components is designed to essentially determine for each patient its most likely component. We show on simulated data that the patients assignment using the MFNCH-based approach can be superior, if not comparable, to using the multinomial-based approach. Lastly, we illustrate on real Acute Myeloid Leukemia data how the utilization of MFNCH-based approach emerges as a good trade-off between the rigorous multinomial-based characterization of the HDMM components and the common refinement of them based on a priori clinical knowledge. Daniele Dall'Olio, J. Eric Sträng, Amin T. Turki, Jesse M. Tettero, Martje Barbus, Renate Schulze-Rath, Javier Martinez Elicegui, Tommaso Matteuzzi, Alessandra Merlotti, Luciana Carota, Claudia Sala, Matteo G. Della Porta, Enrico Giampieri, Jesús María Hernández-Rivas, Lars Bullinger, Gastone C. Castellani |
PLoS Comput. Biol. | 16 |
| 2021 | Characterization and comparison of gene-centered human interactomesabstractThe complex web of macromolecular interactions occurring within cells-the interactome-is the backbone of an increasing number of studies, but a clear consensus on the exact structure of this network is still lacking. Different genome-scale maps of human interactome have been obtained through several experimental techniques and functional analyses. Moreover, these maps can be enriched through literature-mining approaches, and different combinations of various 'source' databases have been used in the literature. It is therefore unclear to which extent the various interactomes yield similar results when used in the context of interactome-based approaches in network biology. We compared a comprehensive list of human interactomes on the basis of topology, protein complexes, molecular pathways, pathway cross-talk and disease gene prediction. In a general context of relevant heterogeneity, our study provides a series of qualitative and quantitative parameters that describe the state of the art of human interactomes and guidelines for selecting interactomes in future applications. Ettore Mosca, Matteo Bersanelli, Tommaso Matteuzzi, Noemi Di Nanni, Gastone C. Castellani, Luciano Milanesi, Daniel Remondini |
Briefings Bioinform. | 5 |
| 2021 | Impact of concurrency on the performance of a whole exome sequencing pipelineabstractBACKGROUND: Current high-throughput technologies-i.e. whole genome sequencing, RNA-Seq, ChIP-Seq, etc.-generate huge amounts of data and their usage gets more widespread with each passing year. Complex analysis pipelines involving several computationally-intensive steps have to be applied on an increasing number of samples. Workflow management systems allow parallelization and a more efficient usage of computational power. Nevertheless, this mostly happens by assigning the available cores to a single or few samples' pipeline at a time. We refer to this approach as naive parallel strategy (NPS). Here, we discuss an alternative approach, which we refer to as concurrent execution strategy (CES), which equally distributes the available processors across every sample's pipeline. RESULTS: Theoretically, we show that the CES results, under loose conditions, in a substantial speedup, with an ideal gain range spanning from 1 to the number of samples. Also, we observe that the CES yields even faster executions since parallelly computable tasks scale sub-linearly. Practically, we tested both strategies on a whole exome sequencing pipeline applied to three publicly available matched tumour-normal sample pairs of gastrointestinal stromal tumour. The CES achieved speedups in latency up to 2-2.4 compared to the NPS. CONCLUSIONS: Our results hint that if resources distribution is further tailored to fit specific situations, an even greater gain in performance of multiple samples pipelines execution could be achieved. For this to be feasible, a benchmarking of the tools included in the pipeline would be necessary. It is our opinion these benchmarks should be consistently performed by the tools' developers. Finally, these results suggest that concurrent strategies might also lead to energy and cost savings by making feasible the usage of low power machine clusters. Daniele Dall'Olio, Nico Curti, Eugenio Fonzi, Claudia Sala, Daniel Remondini, Gastone C. Castellani, Enrico Giampieri |
BMC Bioinform. | 6 |
| 2021 | Correction to: Impact of concurrency on the performance of a whole exome sequencing pipelineabstractAn amendment to this paper has been published and can be accessed via the original article. Daniele Dall'Olio, Nico Curti, Eugenio Fonzi, Claudia Sala, Daniel Remondini, Gastone C. Castellani, Enrico Giampieri |
BMC Bioinform. | 6 |
| 2018 | Statistical modelling of CG interdistance across multiple organismsabstractBACKGROUND: Statistical approaches to genetic sequences have revealed helpful to gain deeper insight into biological and structural functionalities, using ideas coming from information theory and stochastic modelling of symbolic sequences. In particular, previous analyses on CG dinucleotide position along the genome allowed to highlight its epigenetic role in DNA methylation, showing a different distribution tail as compared to other dinucleotides. In this paper we extend the analysis to the whole CG distance distribution over a selected set of higher-order organisms. Then we apply the best fitting probability density function to a large range of organisms (>4400) of different complexity (from bacteria to mammals) and we characterize some emerging global features. RESULTS: We find that the Gamma distribution is optimal for the selected subset as compared to a group of several distributions, chosen for their physical meaning or because recently used in literature for similar studies. The parameters of this distribution, when applied to our larger set of organisms, allows to highlight some biologically relavant features for the considered organism classes, that can be useful also for classification purposes. CONCLUSIONS: The quantification of statistical properties of CG dinucleotide positioning along the genome is confirmed as a useful tool to characterize broad classes of organisms, spanning the whole range of biological complexity. A. Merlotti, Ìtalo Faria do Valle, Gastone C. Castellani, Daniel Remondini |
BMC Bioinform. | 3 |
| 2016 | Systems medicine of inflammagingabstractSystems Medicine (SM) can be defined as an extension of Systems Biology (SB) to Clinical-Epidemiological disciplines through a shifting paradigm, starting from a cellular, toward a patient centered framework. According to this vision, the three pillars of SM are Biomedical hypotheses, experimental data, mainly achieved by Omics technologies and tailored computational, statistical and modeling tools. The three SM pillars are highly interconnected, and their balancing is crucial. Despite the great technological progresses producing huge amount of data (Big Data) and impressive computational facilities, the Bio-Medical hypotheses are still of primary importance. A paradigmatic example of unifying Bio-Medical theory is the concept of Inflammaging. This complex phenotype is involved in a large number of pathologies and patho-physiological processes such as aging, age-related diseases and cancer, all sharing a common inflammatory pathogenesis. This Biomedical hypothesis can be mapped into an ecological perspective capable to describe by quantitative and predictive models some experimentally observed features, such as microenvironment, niche partitioning and phenotype propagation. In this article we show how this idea can be supported by computational methods useful to successfully integrate, analyze and model large data sets, combining cross-sectional and longitudinal information on clinical, environmental and omics data of healthy subjects and patients to provide new multidimensional biomarkers capable of distinguishing between different pathological conditions, e.g. healthy versus unhealthy state, physiological versus pathological aging. Gastone C. Castellani, Giulia Menichetti, Paolo Garagnani, Maria Giulia Bacalini, Chiara Pirazzini, Claudio Franceschi, Sebastiano Collino, Claudia Sala, Daniel Remondini, Enrico Giampieri, Ettore Mosca, Matteo Bersanelli, Silvia Vitali, Ìtalo Faria do Valle, Pietro Liò, Luciano Milanesi |
Briefings Bioinform. | 1 |
| 2016 | Methods for the integration of multi-omics data: mathematical aspectsabstractBACKGROUND: Methods for the integrative analysis of multi-omics data are required to draw a more complete and accurate picture of the dynamics of molecular systems. The complexity of biological systems, the technological limits, the large number of biological variables and the relatively low number of biological samples make the analysis of multi-omics datasets a non-trivial problem. RESULTS AND CONCLUSIONS: We review the most advanced strategies for integrating multi-omics datasets, focusing on mathematical and methodological aspects. Matteo Bersanelli, Ettore Mosca, Daniel Remondini, Enrico Giampieri, Claudia Sala, Gastone C. Castellani, Luciano Milanesi |
BMC Bioinform. | 6 |
| 2016 | Stochastic neutral modelling of the Gut Microbiota's relative species abundance from next generation sequencing dataabstractBACKGROUND: Interest in understanding the mechanisms that lead to a particular composition of the Gut Microbiota is highly increasing, due to the relationship between this ecosystem and the host health state. Particularly relevant is the study of the Relative Species Abundance (RSA) distribution, that is a component of biodiversity and measures the number of species having a given number of individuals. It is the universal behaviour of RSA that induced many ecologists to look for theoretical explanations. In particular, a simple stochastic neutral model was proposed by Volkov et al. relying on population dynamics and was proved to fit the coral-reefs and rain forests RSA. Our aim is to ascertain if this model also describes the Microbiota RSA and if it can help in explaining the Microbiota plasticity. RESULTS: We analyzed 16S rRNA sequencing data sampled from the Microbiota of three different animal species by Jeraldo et al. Through a clustering procedure (UCLUST), we built the Operational Taxonomic Units. These correspond to bacterial species considered at a given phylogenetic level defined by the similarity threshold used in the clustering procedure. The RSAs, plotted in the form of Preston plot, were fitted with Volkov's model. The model fits well the Microbiota RSA, except in the tail region, that shows a deviation from the neutrality assumption. Looking at the model parameters we were able to discriminate between different animal species, giving also a biological explanation. Moreover, the biodiversity estimator obtained by Volkov's model also differentiates the animal species and is in good agreement with the first and second order Hill's numbers, that are common evenness indexes simply based on the fraction of individuals per species. CONCLUSIONS: We conclude that the neutrality assumption is a good approximation for the Microbiota dynamics and the observation that Volkov's model works for this ecosystem is a further proof of the RSA universality. Moreover, the ability to separate different animals with the model parameters and biodiversity number are promising results if we think about future applications on human data, in which the Microbiota composition and biodiversity are in close relationships with a variety of diseases and life-styles. Claudia Sala, Silvia Vitali, Enrico Giampieri, Ìtalo Faria do Valle, Daniel Remondini, Paolo Garagnani, Matteo Bersanelli, Ettore Mosca, Luciano Milanesi, Gastone C. Castellani |
BMC Bioinform. | 10 |
| 2016 | Optimized pipeline of MuTect and GATK tools to improve the detection of somatic single nucleotide polymorphisms in whole-exome sequencing dataabstractDetecting somatic mutations in whole exome sequencing data of cancer samples has become a popular approach for profiling cancer development, progression and chemotherapy resistance. Several studies have proposed software packages, filters and parametrizations. However, many research groups reported low concordance among different methods. We aimed to develop a pipeline which detects a wide range of single nucleotide mutations with high validation rates. We combined two standard tools – Genome Analysis Toolkit (GATK) and MuTect – to create the GATK-LOD N method. As proof of principle, we applied our pipeline to exome sequencing data of hematological (Acute Myeloid and Acute Lymphoblastic Leukemias) and solid (Gastrointestinal Stromal Tumor and Lung Adenocarcinoma) tumors. We performed experiments on simulated data to test the sensitivity and specificity of our pipeline. The software MuTect presented the highest validation rate (90 %) for mutation detection, but limited number of somatic mutations detected. The GATK detected a high number of mutations but with low specificity. The GATK-LOD N increased the performance of the GATK variant detection (from 5 of 14 to 3 of 4 confirmed variants), while preserving mutations not detected by MuTect. However, GATK-LOD N filtered more variants in the hematological samples than in the solid tumors. Experiments in simulated data demonstrated that GATK-LOD N increased both specificity and sensitivity of GATK results. We presented a pipeline that detects a wide range of somatic single nucleotide variants, with good validation rates, from exome sequencing data of cancer samples. We also showed the advantage of combining standard algorithms to create the GATK-LOD N method, that increased specificity and sensitivity of GATK results. This pipeline can be helpful in discovery studies aimed to profile the somatic mutational landscape of cancer genomes. Ìtalo Faria do Valle, Enrico Giampieri, Giorgia Simonetti, Antonella Padella, Marco Manfrini, Anna Ferrari, Cristina Papayannidis, Isabella Zironi, Marianna Garonzi, Simona Bernardi 0002, Massimo Delledonne, Giovanni Martinelli, Daniel Remondini, Gastone C. Castellani |
BMC Bioinform. | 14 |
| 2009 | Trends in modeling Biomedical Complex Systems
Luciano Milanesi, Paolo Romano 0001, Gastone C. Castellani, Daniel Remondini, Pietro Liò |
BMC Bioinform. | 3 |
| 2008 | Reconstructing networks of pathways via significance analysis of their intersectionsabstractBACKGROUND: Significance analysis at single gene level may suffer from the limited number of samples and experimental noise that can severely limit the power of the chosen statistical test. This problem is typically approached by applying post hoc corrections to control the false discovery rate, without taking into account prior biological knowledge. Pathway or gene ontology analysis can provide an alternative way to relax the significance threshold applied to single genes and may lead to a better biological interpretation. RESULTS: Here we propose a new analysis method based on the study of networks of pathways. These networks are reconstructed considering both the significance of single pathways (network nodes) and the intersection between them (links). We apply this method for the reconstruction of networks of pathways to two gene expression datasets: the first one obtained from a c-Myc rat fibroblast cell line expressing a conditional Myc-estrogen receptor oncoprotein; the second one obtained from the comparison of Acute Myeloid Leukemia and Acute Lymphoblastic Leukemia derived from bone marrow samples. CONCLUSION: Our method extends statistical models that have been recently adopted for the significance analysis of functional groups of genes to infer links between these groups. We show that groups of genes at the interface between different pathways can be considered as relevant even if the pathways they belong to are not significant by themselves. Mirko Francesconi, Daniel Remondini, Nicola Neretti, John M. Sedivy, Leon N. Cooper, Ettore Verondini, Luciano Milanesi, Gastone C. Castellani |
BMC Bioinform. | 8 |
| 2007 | Correlation analysis reveals the emergence of coherence in the gene expression dynamics following system perturbationabstractTime course gene expression experiments are a popular means to infer co-expression. Many methods have been proposed to cluster genes or to build networks based on similarity measures of their expression dynamics. In this paper we apply a correlation based approach to network reconstruction to three datasets of time series gene expression following system perturbation: 1) Conditional, Tamoxifen dependent, activation of the cMyc proto-oncogene in rat fibroblast; 2) Genomic response to nutrition changes in D. melanogaster; 3) Patterns of gene activity as a consequence of ageing occurring over a life-span time series (25y-90y) sampled from T-cells of human donors. We show that the three datasets undergo similar transitions from an "uncorrelated" regime to a positively or negatively correlated one that is symptomatic of a shift from a "ground" or "basal" state to a "polarized" state. In addition, we show that a similar transition is conserved at the pathway level, and that this information can be used for the construction of "meta-networks" where it is possible to assess new relations among functionally distant sets of molecular functions. Nicola Neretti, Daniel Remondini, Marc Tatar, John M. Sedivy, Michela Pierini, Dawn Mazzatti, Jonathan Powell, Claudio Franceschi, Gastone C. Castellani |
BMC Bioinform. | 9 |
| 2005 | Quantifying the relevance of different mediators in the human immune cell networkabstractMOTIVATION: Immune cells coordinate their efforts for the correct and efficient functioning of the immune system (IS). Each cell type plays a distinct role and communicates with other cell types through mediators such as cytokines, chemokines and hormones, among others, that are crucial for the functioning of the IS and its fine tuning. Nevertheless, a quantitative analysis of the topological properties of an immunological network involving this complex interchange of mediators among immune cells is still lacking. RESULTS: Here we present a method for quantifying the relevance of different mediators in the immune network, which exploits a definition of centrality based on the concept of efficient communication. The analysis, applied to the human IS, indicates that its mediators differ significantly in their network relevance. We found that cytokines involved in innate immunity and inflammation and some hormones rank highest in the network, revealing that the most prominent mediators of the IS are molecules involved in these ancestral types of defence mechanisms which are highly integrated with the adaptive immune response, and at the interplay among the nervous, the endocrine and the immune systems. CONTACT: [email protected]. Paolo Tieri, Silvana Valensin, Vito Latora, Gastone C. Castellani, Massimo Marchiori, Daniel Remondini, Claudio Franceschi |
Bioinform. | 4 |
| 2003 | The Effect of Noise on a Class of Energy-Based Learning RulesabstractWestudy the selectivity properties of neurons based on BCM and kurtosis energy functions in a general case of noisy high-dimensional input space. The proposed approach, which is used for characterization of the stable states, can be generalized to a whole class of energy functions. We characterize the critical noise levels beyond which the selectivity is destroyed. We also perform a quantitative analysis of such transitions, which shows interesting dependency on data set size. We observe that the robustness to noise of the BCM neuron (Bienenstock, Cooper, & Munro, 1982; Intrator & Cooper, 1992) increases as a function of dimensionality. We explicitly compute the separability limit of BCM and kurtosis learning rules in the case of a bimodal input distribution. Numerical simulations show a stronger robustness of the BCM rule for practical data set size when compared with kurtosis. Armando Bazzani, Daniel Remondini, Nathan Intrator, Gastone C. Castellani |
Neural Comput. | 4 |
| 2002 | Optimal spontaneous activity in neural network modeling
Daniel Remondini, Nathan Intrator, Gastone C. Castellani, F. Bersani, Leon N. Cooper |
Neurocomputing | 3 |