VLDB 2026 Research / reviewers in the wild / expert
Claudia Sala
dblp:176/7284
· DBLP profile ↗
10ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0002-4889-1047ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Survival Model Optimization via Federated Learning: A Study Combining Simulations and ExperimentsabstractFederated Learning is an emerging, powerful approach that allows training an artificial intelligence model in distributed setting. Two survival models, Cox and DeepSurv, have been trained in a federated setting, exploiting both code simulations and real experiments on the new platform, developed by the GenoMed4All consortium. Different scenarios have been tested by splitting a Myelodysplastic Syndrome dataset into three nodes and performing feature removal. A significant gain in model performance has been observed due to federated aggregation. Francesco Casadei, Luciana Carota, Gianluca Asti, Saverio D'Amico, Davide Piscia, Santiago Zazo, Patricia A. Apellániz, Juan Parras, Claudia Sala, Cesare Rollo, Nono S. C. Merleau, Piero Fariselli, Matteo Giovanni Della Porta, Tiziana Sanavia, Federico Álvarez García, Gastone C. Castellani, Enrico Giampieri |
IEEE Big Data | 9 |
| 2024 | Covering Hierarchical Dirichlet Mixture Models on binary data to enhance genomic stratifications in onco-hematologyabstractOnco-hematological studies are increasingly adopting statistical mixture models to support the advancement of the genomically-driven classification systems for blood cancer. Targeting enhanced patients stratification based on the sole role of molecular biology attracted much interest and contributes to bring personalized medicine closer to reality. In onco-hematology, Hierarchical Dirichlet Mixture Models (HDMM) have become one of the preferred method to cluster the genomics data, that include the presence or absence of gene mutations and cytogenetics anomalies, into components. This work unfolds the standard workflow used in onco-hematology to improve patient stratification and proposes alternative approaches to characterize the components and to assign patient to them, as they are crucial tasks usually supported by a priori clinical knowledge. We propose (a) to compute the parameters of the multinomial components of the HDMM or (b) to estimate the parameters of the HDMM components as if they were Multivariate Fisher's Non-Central Hypergeometric (MFNCH) distributions. Then, our approach to perform patients assignments to the HDMM components is designed to essentially determine for each patient its most likely component. We show on simulated data that the patients assignment using the MFNCH-based approach can be superior, if not comparable, to using the multinomial-based approach. Lastly, we illustrate on real Acute Myeloid Leukemia data how the utilization of MFNCH-based approach emerges as a good trade-off between the rigorous multinomial-based characterization of the HDMM components and the common refinement of them based on a priori clinical knowledge. Daniele Dall'Olio, J. Eric Sträng, Amin T. Turki, Jesse M. Tettero, Martje Barbus, Renate Schulze-Rath, Javier Martinez Elicegui, Tommaso Matteuzzi, Alessandra Merlotti, Luciana Carota, Claudia Sala, Matteo G. Della Porta, Enrico Giampieri, Jesús María Hernández-Rivas, Lars Bullinger, Gastone C. Castellani |
PLoS Comput. Biol. | 11 |
| 2022 | Evaluation of different computational methods for DNA methylation-based biological ageabstractIn recent years there has been a widespread interest in researching biomarkers of aging that could predict physiological vulnerability better than chronological age. Aging, in fact, is one of the most relevant risk factors for a wide range of maladies, and molecular surrogates of this phenotype could enable better patients stratification. Among the most promising of such biomarkers is DNA methylation-based biological age. Given the potential and variety of computational implementations (epigenetic clocks), we here present a systematic review of such clocks. Furthermore, we provide a large-scale performance comparison across different tissues and diseases in terms of age prediction accuracy and age acceleration, a measure of deviance from physiology. Our analysis offers both a state-of-the-art overview of the computational techniques developed so far and a heterogeneous picture of performances, which can be helpful in orienting future research. Pietro Di Lena, Claudia Sala, Christine Nardini |
Briefings Bioinform. | 2 |
| 2021 | Impact of concurrency on the performance of a whole exome sequencing pipelineabstractBACKGROUND: Current high-throughput technologies-i.e. whole genome sequencing, RNA-Seq, ChIP-Seq, etc.-generate huge amounts of data and their usage gets more widespread with each passing year. Complex analysis pipelines involving several computationally-intensive steps have to be applied on an increasing number of samples. Workflow management systems allow parallelization and a more efficient usage of computational power. Nevertheless, this mostly happens by assigning the available cores to a single or few samples' pipeline at a time. We refer to this approach as naive parallel strategy (NPS). Here, we discuss an alternative approach, which we refer to as concurrent execution strategy (CES), which equally distributes the available processors across every sample's pipeline. RESULTS: Theoretically, we show that the CES results, under loose conditions, in a substantial speedup, with an ideal gain range spanning from 1 to the number of samples. Also, we observe that the CES yields even faster executions since parallelly computable tasks scale sub-linearly. Practically, we tested both strategies on a whole exome sequencing pipeline applied to three publicly available matched tumour-normal sample pairs of gastrointestinal stromal tumour. The CES achieved speedups in latency up to 2-2.4 compared to the NPS. CONCLUSIONS: Our results hint that if resources distribution is further tailored to fit specific situations, an even greater gain in performance of multiple samples pipelines execution could be achieved. For this to be feasible, a benchmarking of the tools included in the pipeline would be necessary. It is our opinion these benchmarks should be consistently performed by the tools' developers. Finally, these results suggest that concurrent strategies might also lead to energy and cost savings by making feasible the usage of low power machine clusters. Daniele Dall'Olio, Nico Curti, Eugenio Fonzi, Claudia Sala, Daniel Remondini, Gastone C. Castellani, Enrico Giampieri |
BMC Bioinform. | 4 |
| 2021 | Correction to: Impact of concurrency on the performance of a whole exome sequencing pipelineabstractAn amendment to this paper has been published and can be accessed via the original article. Daniele Dall'Olio, Nico Curti, Eugenio Fonzi, Claudia Sala, Daniel Remondini, Gastone C. Castellani, Enrico Giampieri |
BMC Bioinform. | 4 |
| 2020 | Methylation data imputation performances under different representations and missingness patternsabstractBACKGROUND: High-throughput technologies enable the cost-effective collection and analysis of DNA methylation data throughout the human genome. This naturally entails missing values management that can complicate the analysis of the data. Several general and specific imputation methods are suitable for DNA methylation data. However, there are no detailed studies of their performances under different missing data mechanisms -(completely) at random or not- and different representations of DNA methylation levels (β and M-value). RESULTS: We make an extensive analysis of the imputation performances of seven imputation methods on simulated missing completely at random (MCAR), missing at random (MAR) and missing not at random (MNAR) methylation data. We further consider imputation performances on the popular β- and M-value representations of methylation levels. Overall, β-values enable better imputation performances than M-values. Imputation accuracy is lower for mid-range β-values, while it is generally more accurate for values at the extremes of the β-value range. The MAR values distribution is on the average more dense in the mid-range in comparison to the expected β-value distribution. As a consequence, MAR values are on average harder to impute. CONCLUSIONS: The results of the analysis provide guidelines for the most suitable imputation approaches for DNA methylation data under different representations of DNA methylation levels and different missing data mechanisms. Pietro Di Lena, Claudia Sala, Andrea Prodi, Christine Nardini |
BMC Bioinform. | 2 |
| 2019 | Missing value estimation methods for DNA methylation dataabstractMOTIVATION: DNA methylation is a stable epigenetic mark with major implications in both physiological (development, aging) and pathological conditions (cancers and numerous diseases). Recent research involving methylation focuses on the development of molecular age estimation methods based on DNA methylation levels (mAge). An increasing number of studies indicate that divergences between mAge and chronological age may be associated to age-related diseases. Current advances in high-throughput technologies have allowed the characterization of DNA methylation levels throughout the human genome. However, experimental methylation profiles often contain multiple missing values that can affect the analysis of the data and also mAge estimation. Although several imputation methods exist, a major deficiency lies in the inability to cope with large datasets, such as DNA methylation chips. Specific methods for imputing missing methylation data are therefore needed. RESULTS: We present a simple and computationally efficient imputation method, metyhLImp, based on linear regression. The rationale of the approach lies in the observation that methylation levels show a high degree of inter-sample correlation. We performed a comparative study of our approach with other imputation methods on DNA methylation data of healthy and disease samples from different tissues. Performances have been assessed both in terms of imputation accuracy and in terms of the impact imputed values have on mAge estimation. In comparison to existing methods, our linear regression model proves to perform equally or better and with good computational efficiency. The results of our analysis provide recommendations for accurate estimation of missing methylation values. AVAILABILITY AND IMPLEMENTATION: The R-package methyLImp is freely available at https://github.com/pdilena/methyLImp. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pietro Di Lena, Claudia Sala, Andrea Prodi, Christine Nardini |
Bioinform. | 2 |
| 2016 | Systems medicine of inflammagingabstractSystems Medicine (SM) can be defined as an extension of Systems Biology (SB) to Clinical-Epidemiological disciplines through a shifting paradigm, starting from a cellular, toward a patient centered framework. According to this vision, the three pillars of SM are Biomedical hypotheses, experimental data, mainly achieved by Omics technologies and tailored computational, statistical and modeling tools. The three SM pillars are highly interconnected, and their balancing is crucial. Despite the great technological progresses producing huge amount of data (Big Data) and impressive computational facilities, the Bio-Medical hypotheses are still of primary importance. A paradigmatic example of unifying Bio-Medical theory is the concept of Inflammaging. This complex phenotype is involved in a large number of pathologies and patho-physiological processes such as aging, age-related diseases and cancer, all sharing a common inflammatory pathogenesis. This Biomedical hypothesis can be mapped into an ecological perspective capable to describe by quantitative and predictive models some experimentally observed features, such as microenvironment, niche partitioning and phenotype propagation. In this article we show how this idea can be supported by computational methods useful to successfully integrate, analyze and model large data sets, combining cross-sectional and longitudinal information on clinical, environmental and omics data of healthy subjects and patients to provide new multidimensional biomarkers capable of distinguishing between different pathological conditions, e.g. healthy versus unhealthy state, physiological versus pathological aging. Gastone C. Castellani, Giulia Menichetti, Paolo Garagnani, Maria Giulia Bacalini, Chiara Pirazzini, Claudio Franceschi, Sebastiano Collino, Claudia Sala, Daniel Remondini, Enrico Giampieri, Ettore Mosca, Matteo Bersanelli, Silvia Vitali, Ìtalo Faria do Valle, Pietro Liò, Luciano Milanesi |
Briefings Bioinform. | 8 |
| 2016 | Methods for the integration of multi-omics data: mathematical aspectsabstractBACKGROUND: Methods for the integrative analysis of multi-omics data are required to draw a more complete and accurate picture of the dynamics of molecular systems. The complexity of biological systems, the technological limits, the large number of biological variables and the relatively low number of biological samples make the analysis of multi-omics datasets a non-trivial problem. RESULTS AND CONCLUSIONS: We review the most advanced strategies for integrating multi-omics datasets, focusing on mathematical and methodological aspects. Matteo Bersanelli, Ettore Mosca, Daniel Remondini, Enrico Giampieri, Claudia Sala, Gastone C. Castellani, Luciano Milanesi |
BMC Bioinform. | 5 |
| 2016 | Stochastic neutral modelling of the Gut Microbiota's relative species abundance from next generation sequencing dataabstractBACKGROUND: Interest in understanding the mechanisms that lead to a particular composition of the Gut Microbiota is highly increasing, due to the relationship between this ecosystem and the host health state. Particularly relevant is the study of the Relative Species Abundance (RSA) distribution, that is a component of biodiversity and measures the number of species having a given number of individuals. It is the universal behaviour of RSA that induced many ecologists to look for theoretical explanations. In particular, a simple stochastic neutral model was proposed by Volkov et al. relying on population dynamics and was proved to fit the coral-reefs and rain forests RSA. Our aim is to ascertain if this model also describes the Microbiota RSA and if it can help in explaining the Microbiota plasticity. RESULTS: We analyzed 16S rRNA sequencing data sampled from the Microbiota of three different animal species by Jeraldo et al. Through a clustering procedure (UCLUST), we built the Operational Taxonomic Units. These correspond to bacterial species considered at a given phylogenetic level defined by the similarity threshold used in the clustering procedure. The RSAs, plotted in the form of Preston plot, were fitted with Volkov's model. The model fits well the Microbiota RSA, except in the tail region, that shows a deviation from the neutrality assumption. Looking at the model parameters we were able to discriminate between different animal species, giving also a biological explanation. Moreover, the biodiversity estimator obtained by Volkov's model also differentiates the animal species and is in good agreement with the first and second order Hill's numbers, that are common evenness indexes simply based on the fraction of individuals per species. CONCLUSIONS: We conclude that the neutrality assumption is a good approximation for the Microbiota dynamics and the observation that Volkov's model works for this ecosystem is a further proof of the RSA universality. Moreover, the ability to separate different animals with the model parameters and biodiversity number are promising results if we think about future applications on human data, in which the Microbiota composition and biodiversity are in close relationships with a variety of diseases and life-styles. Claudia Sala, Silvia Vitali, Enrico Giampieri, Ìtalo Faria do Valle, Daniel Remondini, Paolo Garagnani, Matteo Bersanelli, Ettore Mosca, Luciano Milanesi, Gastone C. Castellani |
BMC Bioinform. | 1 |