Elena Casiraghi

dblp:66/4510 · DBLP profile ↗
← Back
33ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0003-2024-7572ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 15 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Quantum-enhanced Representation Learning and Matching Learning for Recommendation
abstract
Quantum computing is an emerging research area. This paper investigates why and how quantum computing can be integrated into recommender systems. Although some existing recommendation methods explore quantum concepts, they either remain theoretical without empirical validation or provide limited insight into the use of quantum computing for designing core functions in recommendation. To fill these gaps, we first analyze the potential advantages of quantum computing for two key components (i.e., representation learning and matching learning) in recommender algorithms and formulate corresponding hypotheses. Then, based on our analysis and the quantum computing operations, we propose three quantum-enhanced recommendation paradigms. To show the extensibility of our paradigms, we further apply them to the graph-based and social recommendation scenarios. We conduct extensive experiments on the six real-world datasets, comparing our methods with various baselines. Experimental results not only validate our hypotheses but also show the strong performance of our proposed methods.
Anchen Li, Elena Casiraghi
WWW2
2026 A cross-attentive multi-task graph learning framework for chemical reaction modeling
abstract
Motivation: Understanding chemical reactions requires bridging fine-grained molecular edits with broader semantic context.Reaction mechanisms are determined not only by local atom-bond transformations but also by the global reaction class.However, most existing approaches treat these tasks separately or rely on external atom-mapping tools, introducing noise and limiting end-to-end learnability.We introduce MARCC (Mapping-Assisted Reaction Center and Classification), a multi-task graph neural network that jointly predicts atom mappings, reaction centers, and reaction classes within a unified architecture.Results: MARCC integrates three key innovations: (i) a mapping-guided cross-attention mechanism that aligns reactants and products for local edit detection, (ii) a dual-graph design that explicitly reasons about bond-level transformations, and (iii) pooled product embeddings for global reaction classification.On the USPTO-50K benchmark, MARCC achieves state-of-the-art results when trained with both reactants and products, including 98.2% atom mapping accuracy, 99.1% Top-1 edit localization accuracy, and 97.2% reaction classification accuracy.Even under the products-only setting, MARCC delivers competitive performance comparable to specialized baselines.Ablation studies confirm the value of mapping-guided attention and multi-task supervision, which enhance both predictive accuracy and interpretability.By unifying atom-level alignment, local reactivity, and global classification, MARCC provides a structured and interpretable framework for reaction understanding.Beyond benchmarks, MARCC has the potential to support applications in reaction annotation, template discovery, and mechanism inference; with additional domain-specific modeling and data, it could be extended to biochemical domains such as enzyme-catalyzed transformations and metabolic pathway modeling.
Maryam Astero, Anchen Li, Elena Casiraghi, Juho Rousu
Bioinform.3
2026 Tabular implicit deep neural networks ensembles for the prediction of pathogenic genetic variants in Mendelian diseases
abstract
Mendelian genetic diseases comprise approximately described disorders, yet the genetic basis is known only for about half of them, and a molecular diagnosis often remains difficult or unresolved. In this context, machine learning methods play a key role. However, identifying pathogenic variants in non-coding regions of the human genome is particularly challenging, as they are vastly outnumbered by neutral variants. This extreme imbalance causes standard machine learning approaches to exhibit a strong predictive bias towards the majority class, significantly limiting their sensitivity. Building on recent advances in imbalance-aware and ensemble learning methods, we propose two novel deep learning models for predicting pathogenic non-coding variants in Mendelian diseases. The first model, T-ResNet (Tabular Residual Neural Network), adopts a modular architecture with residual connections, along with a mini-batch balancing strategy to address class imbalance. This design simplifies hyperparameter optimization while mitigating vanishing-gradient effects. The second model, TIDE-Var (Tabular Implicit Deep neural network Ensembles for Variant prediction), leverages the TabM and BatchEnsemble models to build an implicit ensemble of deep neural networks trained jointly by minimizing a common objective function, and partially sharing learning parameters. Learner-specific adapter parameters promote diversity among the implicit base learners while keeping the ensemble computationally efficient. Genome-wide experiments show that T-ResNet achieves an average Area Under the Precision-Recall Curve (AUPRC) of , comparable to that of HyperSMURF , a state-of-the-art method. In contrast, TIDE-Var yields significantly better results than HyperSMURF (AUPRC ) and slightly better than XGBoost , one of the top-methods for the classification of tabular data. Ablation studies confirm that regularized joint ensemble learning with partially shared and base-learner specific adapter learning parameters are key factors in achieving high predictive performance, essential to improve the diagnostic yield for patients with rare genetic diseases.
Federico Stacchietti, Marco Nicolini, Leonardo Chimirri, Peter N. Robinson, Elena Casiraghi, Giorgio Valentini
Neurocomputing5
2025 An Empirical Study of Remote Homology Detection using Protein Language Models
abstract
Detecting remote homologs, proteins that share evolutionary ancestry despite low sequence similarity, remains a central challenge in computational biology. The Structural Classification of Proteins extended (SCOPe) database organizes protein domains into superfamilies based on structural and functional evidence of common origin, making it a widely used benchmark for remote homology detection. In this study, we investigate the effectiveness of protein language model (PLM) embeddings for predicting SCOPe superfamilies directly from sequence. We introduce DOMCLASS, a deep learning framework that combines supervised contrastive learning with a distance-weighted K-NN classifier to learn and exploit an embedding space aligned with SCOPe superfamily annotations. Our empirical results show that general-purpose PLM embeddings already outperform sequence similarity- and profile-based methods for remote homology detection, and that contrastive learning can further improve performance, even in challenging low sequence identity settings.
Ruben Jimenez, Aldo Galeano, Marcelo Báez, Santiago Ferreyra, Guilherme Melo, Giorgio Valentini, Elena Casiraghi, Luca Cernuzzi, Alberto Paccanaro
CLEI7
2025 Biasing second-order random walk sampling for heterogeneous graph embedding *
abstract
We present heterogeneous-node2vec, a novel method that leverages the well-known node2vec algorithm to enable the generation of random-walk samples in a heterogeneous context. Specifically, we propose a strategy to bias the random walk, enabling type-aware transitions between different node and edge types. We evaluate the proposed technique on node-label prediction tasks, applied to various real-world, complex networks. A comparison with state-of-the-art techniques for heterogeneous graph embedding demonstrates that our strategy achieves competitive results for node-label prediction. This evidences that graph representation methods based on heterogeneous random-walk sampling can attain strong performance on standard supervised tasks when the sampling procedure incorporates the semantic information defined by the type heterogeneity of entities within the graph. This approach provides an effective and scalable solution for representing and learning from complex heterogeneous graphs.
Mauricio Soto Gomez, Carlos Cano, Justin T. Reese, Peter N. Robinson, Giorgio Valentini, Elena Casiraghi
IJCNN6
2025 Intrinsic-dimension analysis for guiding dimensionality reduction and data fusion in multi-omics data processing
abstract
Multi-omics data have revolutionized biomedical research by providing a comprehensive understanding of biological systems and the molecular mechanisms of disease development. However, analyzing multi-omics data is challenging due to high dimensionality and limited sample sizes, necessitating proper data-reduction pipelines to ensure reliable analyses. Additionally, its multimodal nature requires effective data-integration pipelines. While several dimensionality reduction and data fusion algorithms have been proposed, crucial aspects are often overlooked. Specifically, the choice of projection space dimension is typically heuristic and uniformly applied across all omics, neglecting the unique high dimension small sample size challenges faced by individual omics. This paper introduces a novel multi-modal dimensionality reduction pipeline tailored to individual views. By leveraging intrinsic dimensionality estimators, we assess the curse-of-dimensionality impact on each view and propose a two-step reduction strategy for significantly affected views, combining feature selection with feature extraction. Compared to traditional uniform reduction pipelines in a crucial and supervised multi-omics analysis setting, our approach shows significant improvement. Additionally, we explore three effective unsupervised multi-omics data fusion methods rooted in the main data fusion strategies to gain insights into their performance under crucial, yet overlooked, settings.
Jessica Gliozzo, Mauricio Soto Gomez, Valentina Guarino, Arturo Bonometti, Alberto Cabri, Emanuele Cavalleri, Justin T. Reese, Peter N. Robinson, Marco Mesiti, Giorgio Valentini, Elena Casiraghi
Artif. Intell. Medicine11
2025 miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources
abstract
MOTIVATION: Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. RESULTS: We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is "task agnostic", in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer's disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. AVAILABILITY AND IMPLEMENTATION: miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.
Jessica Gliozzo, Mauricio Soto Gomez, Arturo Bonometti, Alex Patak, Elena Casiraghi, Giorgio Valentini
Bioinform.5
2025 CSGL: chemical synthesis graph learning for molecule representation
abstract
MOTIVATION: Molecule representation learning (MRL) translates molecules into a real vector space, serving as input to downstream tasks in biology, chemistry, and computer science. This article introduces a chemical synthesis graph learning (CSGL) framework, which enhances MRL by considering both the atomic structures of molecules and their roles in chemical reactions through a hierarchical graph representation. Specifically, molecules are first modeled based on their molecular graphs, which capture atomic-level structural information. They are then further refined using a chemical synthesis graph, where nodes represent reactant and product molecule sets, and edges encode chemical transformations between reactants and products (e.g. changes in molecular structures). CSGL optimizes molecular embeddings of reactant and product nodes in a fashion that ensures the embeddings conform to a chemical balance constraint. RESULTS: Experimental results show that our method CSGL achieves strong performance on a variety of tasks, including product prediction, reaction classification, and molecular property prediction. AVAILABILITY AND IMPLEMENTATION: https://github.com/li-2023/CSGL.
Anchen Li, Elena Casiraghi, Juho Rousu
Bioinform.2
2024 Chemical reaction enhanced graph learning for molecule representation
abstract
MOTIVATION: Molecular representation learning (MRL) models molecules with low-dimensional vectors to support biological and chemical applications. Current methods primarily rely on intrinsic molecular information to learn molecular representations, but they often overlook effectively integrating domain knowledge into MRL. RESULTS: In this article, we develop a reaction-enhanced graph learning (RXGL) framework for MRL, utilizing chemical reactions as domain knowledge. RXGL introduces dual graph learning modules to model molecule representation. One module employs graph convolutions on molecular graphs to capture molecule structures. The other module constructs a reaction-aware graph from chemical reactions and designs a novel graph attention network on this graph to integrate reaction-level relations into molecular modeling. To refine molecule representations, we design a reaction-based relation learning task, which considers the relations between the reactant and product sides in reactions. In addition, we introduce a cross-view contrastive task to strengthen the cooperative associations between molecular and reaction-aware graph learning. Experiment results show that our RXGL achieves strong performance in various downstream tasks, including product prediction, reaction classification, and molecular property prediction. AVAILABILITY AND IMPLEMENTATION: The code is publicly available at https://github.com/coder-ACAC/RLM.
Anchen Li, Elena Casiraghi, Juho Rousu
Bioinform.2
2023 Enhancing Fairness and Accuracy in Machine Learning Through Similarity Networks
Samira Maghool, Elena Casiraghi, Paolo Ceravolo
CoopIS2
2023 An expectation-maximization framework for comprehensive prediction of isoform-specific functions
abstract
MOTIVATION: Advances in RNA sequencing technologies have achieved an unprecedented accuracy in the quantification of mRNA isoforms, but our knowledge of isoform-specific functions has lagged behind. There is a need to understand the functional consequences of differential splicing, which could be supported by the generation of accurate and comprehensive isoform-specific gene ontology annotations. RESULTS: We present isoform interpretation, a method that uses expectation-maximization to infer isoform-specific functions based on the relationship between sequence and functional isoform similarity. We predicted isoform-specific functional annotations for 85 617 isoforms of 17 900 protein-coding human genes spanning a range of 17 430 distinct gene ontology terms. Comparison with a gold-standard corpus of manually annotated human isoform functions showed that isoform interpretation significantly outperforms state-of-the-art competing methods. We provide experimental evidence that functionally related isoforms predicted by isoform interpretation show a higher degree of domain sharing and expression correlation than functionally related genes. We also show that isoform sequence similarity correlates better with inferred isoform function than with gene-level function. AVAILABILITY AND IMPLEMENTATION: Source code, documentation, and resource files are freely available under a GNU3 license at https://github.com/TheJacksonLaboratory/isopretEM and https://zenodo.org/record/7594321.
Guy Karlebach, Leigh Carmody, Jagadish Chandrabose Sundaramurthi, Elena Casiraghi, Justin T. Reese, Chris Mungall, Giorgio Valentini, Peter N. Robinson
Bioinform.4
2023 A method for comparing multiple imputation techniques: A case study on the U.S. national COVID cohort collaborative
abstract
Healthcare datasets obtained from Electronic Health Records have proven to be extremely useful for assessing associations between patients' predictors and outcomes of interest. However, these datasets often suffer from missing values in a high proportion of cases, whose removal may introduce severe bias. Several multiple imputation algorithms have been proposed to attempt to recover the missing information under an assumed missingness mechanism. Each algorithm presents strengths and weaknesses, and there is currently no consensus on which multiple imputation algorithm works best in a given scenario. Furthermore, the selection of each algorithm's parameters and data-related modeling choices are also both crucial and challenging. In this paper we propose a novel framework to numerically evaluate strategies for handling missing data in the context of statistical analysis, with a particular focus on multiple imputation techniques. We demonstrate the feasibility of our approach on a large cohort of type-2 diabetes patients provided by the National COVID Cohort Collaborative (N3C) Enclave, where we explored the influence of various patient characteristics on outcomes related to COVID-19. Our analysis included classic multiple imputation techniques as well as simple complete-case Inverse Probability Weighted models. Extensive experiments show that our approach can effectively highlight the most promising and performant missing-data handling strategy for our case study. Moreover, our methodology allowed a better understanding of the behavior of the different models and of how it changed as we modified their parameters. Our method is general and can be applied to different research fields and on datasets containing heterogeneous types.
Elena Casiraghi, Rachel Wong, Margaret Hall, Ben D. Coleman, Marco Notaro, Michael D. Evans, Jena S. Tronieri, Hannah Blau, Bryan Laraway, Tiffany Callahan, Lauren E. Chan, Carolyn T. Bramante, John B. Buse, Richard A. Moffitt, Til Sturmer, Steven G. Johnson, Yu Raymond Shao, Justin T. Reese, Peter N. Robinson, Alberto Paccanaro, Giorgio Valentini, Jared D. Huling, Kenneth Wilkins
J. Biomed. Informatics1
2022 Heterogeneous data integration methods for patient similarity networks
abstract
Patient similarity networks (PSNs), where patients are represented as nodes and their similarities as weighted edges, are being increasingly used in clinical research. These networks provide an insightful summary of the relationships among patients and can be exploited by inductive or transductive learning algorithms for the prediction of patient outcome, phenotype and disease risk. PSNs can also be easily visualized, thus offering a natural way to inspect complex heterogeneous patient data and providing some level of explainability of the predictions obtained by machine learning algorithms. The advent of high-throughput technologies, enabling us to acquire high-dimensional views of the same patients (e.g. omics data, laboratory data, imaging data), calls for the development of data fusion techniques for PSNs in order to leverage this rich heterogeneous information. In this article, we review existing methods for integrating multiple biomedical data views to construct PSNs, together with the different patient similarity measures that have been proposed. We also review methods that have appeared in the machine learning literature but have not yet been applied to PSNs, thus providing a resource to navigate the vast machine learning literature existing on this topic. In particular, we focus on methods that could be used to integrate very heterogeneous datasets, including multi-omics data as well as data derived from clinical information and medical imaging.
Jessica Gliozzo, Marco Mesiti, Marco Notaro, Alessandro Petrini, Alex Patak, Antonio Puertas Gallardo, Alberto Paccanaro, Giorgio Valentini, Elena Casiraghi
Briefings Bioinform.9
2022 Boosting tissue-specific prediction of active cis-regulatory regions through deep learning and Bayesian optimization techniques
abstract
BACKGROUND: Cis-regulatory regions (CRRs) are non-coding regions of the DNA that fine control the spatio-temporal pattern of transcription; they are involved in a wide range of pivotal processes such as the development of specific cell-lines/tissues and the dynamic cell response to physiological stimuli. Recent studies showed that genetic variants occurring in CRRs are strongly correlated with pathogenicity or deleteriousness. Considering the central role of CRRs in the regulation of physiological and pathological conditions, the correct identification of CRRs and of their tissue-specific activity status through Machine Learning methods plays a major role in dissecting the impact of genetic variants on human diseases. Unfortunately, the problem is still open, though some promising results have been already reported by (deep) machine-learning based methods that predict active promoters and enhancers in specific tissues or cell lines by encoding epigenetic or spectral features directly extracted from DNA sequences. RESULTS: We present the experiments we performed to compare two Deep Neural Networks, a Feed-Forward Neural Network model working on epigenomic features, and a Convolutional Neural Network model working only on genomic sequence, targeted to the identification of enhancer- and promoter-activity in specific cell lines. While performing experiments to understand how the experimental setup influences the prediction performance of the methods, we particularly focused on (1) automatic model selection performed by Bayesian optimization and (2) exploring different data rebalancing setups for reducing negative unbalancing effects. CONCLUSIONS: Results show that (1) automatic model selection by Bayesian optimization improves the quality of the learner; (2) data rebalancing considerably impacts the prediction performance of the models; test set rebalancing may provide over-optimistic results, and should therefore be cautiously applied; (3) despite working on sequence data, convolutional models obtain performance close to those of feed forward models working on epigenomic information, which suggests that also sequence data carries informative content for CRR-activity prediction. We therefore suggest combining both models/data types in future works.
Luca Cappelletti, Alessandro Petrini, Jessica Gliozzo, Elena Casiraghi, Max Schubach, Martin Kircher, Giorgio Valentini
BMC Bioinform.4
2021 HEMDAG: a family of modular and scalable hierarchical ensemble methods to improve Gene Ontology term prediction
abstract
MOTIVATION: Automated protein function prediction is a complex multi-class, multi-label, structured classification problem in which protein functions are organized in a controlled vocabulary, according to the Gene Ontology (GO). 'Hierarchy-unaware' classifiers, also known as 'flat' methods, predict GO terms without exploiting the inherent structure of the ontology, potentially violating the True-Path-Rule (TPR) that governs the GO, while 'hierarchy-aware' approaches, even if they obey the TPR, do not always show clear improvements with respect to flat methods, or do not scale well when applied to the full GO. RESULTS: To overcome these limitations, we propose Hierarchical Ensemble Methods for Directed Acyclic Graphs (HEMDAG), a family of highly modular hierarchical ensembles of classifiers, able to build upon any flat method and to provide 'TPR-safe' predictions, by leveraging a combination of isotonic regression and TPR learning strategies. Extensive experiments on synthetic and real data across several organisms firstly show that HEMDAG can be used as a general tool to improve the predictions of flat classifiers, and secondly that HEMDAG is competitive versus state-of-the-art hierarchy-aware learning methods proposed in the last CAFA international challenges. AVAILABILITY AND IMPLEMENTATION: Fully tested R code freely available at https://anaconda.org/bioconda/r-hemdag. Tutorial and documentation at https://hemdag.readthedocs.io. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marco Notaro, Marco Frasca 0001, Alessandro Petrini, Jessica Gliozzo, Elena Casiraghi, Peter N. Robinson, Giorgio Valentini
Bioinform.5
2020 A cockpit of multiple measures for assessing film restoration quality
Barbara Rita Barricelli, Elena Casiraghi, Michela Lecca, Alice Plutino, Alessandro Rizzi
Pattern Recognit. Lett.2
2019 ki67 nuclei detection and ki67-index estimation: a novel automatic approach based on human vision modeling
abstract
BACKGROUND: The protein ki67 (pki67) is a marker of tumor aggressiveness, and its expression has been proven to be useful in the prognostic and predictive evaluation of several types of tumors. To numerically quantify the pki67 presence in cancerous tissue areas, pathologists generally analyze histochemical images to count the number of tumor nuclei marked for pki67. This allows estimating the ki67-index, that is the percentage of tumor nuclei positive for pki67 over all the tumor nuclei. Given the high image resolution and dimensions, its estimation by expert clinicians is particularly laborious and time consuming. Though automatic cell counting techniques have been presented so far, the problem is still open. RESULTS: In this paper we present a novel automatic approach for the estimations of the ki67-index. The method starts by exploiting the STRESS algorithm to produce a color enhanced image where all pixels belonging to nuclei are easily identified by thresholding, and then separated into positive (i.e. pixels belonging to nuclei marked for pki67) and negative by a binary classification tree. Next, positive and negative nuclei pixels are processed separately by two multiscale procedures identifying isolated nuclei and separating adjoining nuclei. The multiscale procedures exploit two Bayesian classification trees to recognize positive and negative nuclei-shaped regions. CONCLUSIONS: The evaluation of the computed results, both through experts' visual assessments and through the comparison of the computed indexes with those of experts, proved that the prototype is promising, so that experts believe in its potential as a tool to be exploited in the clinical practice as a valid aid for clinicians estimating the ki67-index. The MATLAB source code is open source for research purposes.
Barbara Rita Barricelli, Elena Casiraghi, Jessica Gliozzo, Veronica Huber, Biagio Eugenio Leone, Alessandro Rizzi, Barbara Vergani
BMC Bioinform.2
2019 UNIPred-Web: a web tool for the integration and visualization of biomolecular networks for protein function prediction
abstract
BACKGROUND: One of the main issues in the automated protein function prediction (AFP) problem is the integration of multiple networked data sources. The UNIPred algorithm was thereby proposed to efficiently integrate -in a function-specific fashion- the protein networks by taking into account the imbalance that characterizes protein annotations, and to subsequently predict novel hypotheses about unannotated proteins. UNIPred is publicly available as R code, which might result of limited usage for non-expert users. Moreover, its application requires efforts in the acquisition and preparation of the networks to be integrated. Finally, the UNIPred source code does not handle the visualization of the resulting consensus network, whereas suitable views of the network topology are necessary to explore and interpret existing protein relationships. RESULTS: We address the aforementioned issues by proposing UNIPred-Web, a user-friendly Web tool for the application of the UNIPred algorithm to a variety of biomolecular networks, already supplied by the system, and for the visualization and exploration of protein networks. We support different organisms and different types of networks -e.g., co-expression, shared domains and physical interaction networks. Users are supported in the different phases of the process, ranging from the selection of the networks and the protein function to be predicted, to the navigation of the integrated network. The system also supports the upload of user-defined protein networks. The vertex-centric and the highly interactive approach of UNIPred-Web allow a narrow exploration of specific proteins, and an interactive analysis of large sub-networks with only a few mouse clicks. CONCLUSIONS: UNIPred-Web offers a practical and intuitive (visual) guidance to biologists interested in gaining insights into protein biomolecular functions. UNIPred-Web provides facilities for the integration of networks, and supplies a framework for the imbalance-aware protein network integration of nine organisms, the prediction of thousands of GO protein functions, and a easy-to-use graphical interface for the visual analysis, navigation and interpretation of the integrated networks and of the functional predictions.
Paolo Perlasca, Marco Frasca 0001, Cheick Tidiane Ba, Marco Notaro, Alessandro Petrini, Elena Casiraghi, Giuliano Grossi, Jessica Gliozzo, Giorgio Valentini, Marco Mesiti
BMC Bioinform.6
2018 A novel computational method for automatic segmentation, quantification and comparative analysis of immunohistochemically labeled tissue sections
abstract
BACKGROUND: In the clinical practice, the objective quantification of histological results is essential not only to define objective and well-established protocols for diagnosis, treatment, and assessment, but also to ameliorate disease comprehension. SOFTWARE: The software MIAQuant_Learn presented in this work segments, quantifies and analyzes markers in histochemical and immunohistochemical images obtained by different biological procedures and imaging tools. MIAQuant_Learn employs supervised learning techniques to customize the marker segmentation process with respect to any marker color appearance. Our software expresses the location of the segmented markers with respect to regions of interest by mean-distance histograms, which are numerically compared by measuring their intersection. When contiguous tissue sections stained by different markers are available, MIAQuant_Learn aligns them and overlaps the segmented markers in a unique image enabling a visual comparative analysis of the spatial distribution of each marker (markers' relative location). Additionally, it computes novel measures of markers' co-existence in tissue volumes depending on their density. CONCLUSIONS: Applications of MIAQuant_Learn in clinical research studies have proven its effectiveness as a fast and efficient tool for the automatic extraction, quantification and analysis of histological sections. It is robust with respect to several deficits caused by image acquisition systems and produces objective and reproducible results. Thanks to its flexibility, MIAQuant_Learn represents an important tool to be exploited in basic research where needs are constantly changing.
Elena Casiraghi, Veronica Huber, Marco Frasca 0001, Mara Cossa, Matteo Tozzi, Licia Rivoltini, Biagio Eugenio Leone, Antonello Villa, Barbara Vergani
BMC Bioinform.1
2014 DANCo: An intrinsic dimensionality estimator exploiting angle and norm concentration
Claudio Ceruti, Simone Bassis, Alessandro Rozza, Gabriele Lombardi, Elena Casiraghi, Paola Campadelli
Pattern Recognit.5
2012 Novel high intrinsic dimensionality estimators
Alessandro Rozza, Gabriele Lombardi, Claudio Ceruti, Elena Casiraghi, Paola Campadelli
Mach. Learn.4
2012 Novel Fisher discriminant classifiers
Alessandro Rozza, Gabriele Lombardi, Elena Casiraghi, Paola Campadelli
Pattern Recognit.3
2011 Minimum Neighbor Distance Estimators of Intrinsic Dimension
Gabriele Lombardi, Alessandro Rozza, Claudio Ceruti, Elena Casiraghi, Paola Campadelli
ECML/PKDD (2)4
2010 A segmentation framework for abdominal organs from CT scans
Paola Campadelli, Elena Casiraghi, Stella Pratissoli
Artif. Intell. Medicine2
2009 Novel IPCA-Based Classifiers and Their Application to Spam Filtering
abstract
This paper proposes a novel two-class classifier, called IPCAC, based on the isotropic principal component analysis technique; it allows to deal with training data drawn from mixture of Gaussian distributions, by projecting the data on the Fisher subspace that separates the two classes. The obtained results demonstrate that IPCAC is a promising technique; furthermore, to cope with training datasets being dynamically supplied, and to work with non-linearly separable classes, two improvements of this classifier are defined: a model merging algorithm, and a kernel version of IPCAC. The effectiveness of the proposed methods is shown by their application to the spam classification problem, and by the comparison of the achieved results with those obtained by support vector machines SVM, and K-nearest neighbors KNN.
Alessandro Rozza, Gabriele Lombardi, Elena Casiraghi
ISDA3
2009 Liver segmentation from computed tomography scans: A survey and a new algorithm
Paola Campadelli, Elena Casiraghi, Alessandro Andrea Esposito
Artif. Intell. Medicine2
2008 Curvature Estimation and Curve Inference with Tensor Voting: A New Approach
Gabriele Lombardi, Elena Casiraghi, Paola Campadelli
ACIVS2
2008 Fully Automatic Segmentation of Abdominal Organs from CT Images Using Fast Marching Methods
abstract
Computed tomography (CT) images are becoming an invaluable mean for abdominal organ investigation. In the field of medical image processing, some of the current interests are the automatic diagnosis of liver, spleen, and kidney pathologies and the 3D volume rendering of the abdominal organs. The first and fundamental step in all these studies is the automatic organs segmentation, that is still an open problem. In this paper we propose a fully automatic gray level based segmentation framework that employs a fast marching technique; the proposed segmentation scheme is general, and employs only established and not critical anatomical knowledge. For this reason, it can be easily adapted to separately segment different abdominal organs, by overcoming problems due to the high inter and intra patient gray level and shape variabilities; the extracted volumes are then combined to achieve robust results. The system performance has been evaluated on the data of 40 patients, by comparing the automatically detected organ volumes to the organ boundaries manually traced by three experts. The good quality of the achieved results is proved by the fact that they are comparable to the inter and intra personal variability of the manual segmentation produced by experts.
Paola Campadelli, Elena Casiraghi, Stella Pratissoli
CBMS2
2007 Automatic Segmentation of Abdominal Organs from CT Scans
abstract
In the field of medical image processing, one of the current interests is the automatic segmentation of abdominal CT images, which is the first and fundamental step both before any automatic pathology detection and to compute the 3D abdominal organ volumes. In this paper we propose our fully automatic system that segments liver and its blood vessels, kidneys, spleen, ribs and spine. The overall system has been evaluated on the data of 40 patients, obtaining a good assessment both by visual inspection by three experts, and by comparing the computed results to the manually traced boundaries.
Paola Campadelli, Elena Casiraghi, Stella Pratissoli
ICTAI (1)2
2006 A Fully Automated Method for Lung Nodule Detection From Postero-Anterior Chest Radiographs
abstract
In the past decades, a great deal of research work has been devoted to the development of systems that could improve radiologists' accuracy in detecting lung nodules. Despite the great efforts, the problem is still open. In this paper, we present a fully automated system processing digital postero-anterior (PA) chest radiographs, that starts by producing an accurate segmentation of the lung field area. The segmented lung area includes even those parts of the lungs hidden behind the heart, the spine, and the diaphragm, which are usually excluded from the methods presented in the literature. This decision is motivated by the fact that lung nodules may be found also in these areas. The segmented area is processed with a simple multiscale method that enhances the visibility of the nodules, and an extraction scheme is then applied to select potential nodules. To reduce the high number of false positives extracted, cost-sensitive support vector machines (SVMs) are trained to recognize the true nodules. Different learning experiments were performed on two different data sets, created by means of feature selection, and employing Gaussian and polynomial SVMs trained with different parameters; the results are reported and compared. With the best SVM models, we obtain about 1.5 false positives per image (fp/image) when sensitivity is approximately equal to 0.71; this number increases to about 2.5 and 4 fp/image when sensitivity is = 0.78 and = 0.85, respectively. For the highest sensitivity (= 0.92 and 1.0), we get 7 or 8 fp/image.
Paola Campadelli, Elena Casiraghi, Diana Artioli
IEEE Trans. Medical Imaging2
2005 Lung nodules detection and classification
abstract
Image processing techniques and computer aided diagnosis (CAD) systems have proved to be effective for the improvement of radiologists' diagnosis. In this paper an automatic system detecting lung nodules from postero anterior chest radiographs is presented. The system extracts a set of candidate regions by applying to the radiograph three different and consecutive multi-scale schemes. The comparison of the results obtained with those presented in the literature show the efficacy of our multi-scale framework. Learning systems using as input different sets of features have been experimented for candidates classification, showing that support vector machines (SVMs) can be successfully applied for this task.
Paola Campadelli, Elena Casiraghi, Giorgio Valentini
ICIP (1)2
2005 Support vector machines for candidate nodules classification
Paola Campadelli, Elena Casiraghi, Giorgio Valentini
Neurocomputing2
2004 Nodule Detection in Postero Anterior Chest Radiographs
Paola Campadelli, Elena Casiraghi
MICCAI (2)2