Giuseppe Jurman

dblp:12/6150 · DBLP profile ↗
← Back
25ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-2705-5728ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Network-Based Integration of Multi-omics Data for Biomarker Discovery in Acute Coronary Syndromes
Elena Marinello, Shahryar Noei, Marco Chierici, Núria Amigó, Dane Cvijanovic, Aleksandar Davidovic, Dalibor Dragisic, Claudia Giglione, Dimitris Kardassis, Marko Milanov, Jelena Munjas, Iva Perovic-Blagojevic, Aleksa Petkovic, Tamara Ratkovic, Pol Torne, Christos Tsatsanis, Sandra Vladimirov-Sopic, Luka Vukmirovic, Paolo Magni, Miron Sopic, Barbara Di Camillo, Giuseppe Jurman
AIME (2)22
2025 Explainable deep neural networks for predicting sample phenotypes from single-cell transcriptomics
abstract
Recent advances in single-cell RNA-Sequencing (scRNA-Seq) technologies have revolutionized our ability to gather molecular insights into different phenotypes at the level of individual cells. The analysis of the resulting data poses significant challenges, and proper statistical methods are required to analyze and extract information from scRNA-Seq datasets. Sample classification based on gene expression data has proven effective and valuable for precision medicine applications. However, standard classification schemas are often not suitable for scRNA-Seq due to their unique characteristics, and new algorithms are required to effectively analyze and classify samples at the single-cell level. Furthermore, existing methods for this purpose have limitations in their usability. Those reasons motivated us to develop singleDeep, an end-to-end pipeline that streamlines the analysis of scRNA-Seq data training deep neural networks, enabling robust prediction and characterization of sample phenotypes. We used singleDeep to make predictions on scRNA-Seq datasets from different conditions, including systemic lupus erythematosus, Alzheimer's disease and coronavirus disease 2019. Our results demonstrate strong diagnostic performance, validated both internally and externally. Moreover, singleDeep outperformed traditional machine learning methods and alternative single-cell approaches. In addition to prediction accuracy, singleDeep provides valuable insights into cell types and gene importance estimation for phenotypic characterization. This functionality provided additional and valuable information in our use cases. For instance, we corroborated that some interferon signature genes are consistently relevant for autoimmunity across all immune cell types in lupus. On the other hand, we discovered that genes linked to dementia have relevant roles in specific brain cell populations, such as APOE in astrocytes.
Jordi Martorell-Marugan, Raúl López-Domínguez, Juan Antonio Villatoro-García, Daniel Toro-Domínguez, Marco Chierici, Giuseppe Jurman, Pedro Carmona-Saez
Briefings Bioinform.6
2025 Bottlenecks in advancing and applying multiomic data integration - common data resources as rate-limiting drivers - the high-impact use case of atherosclerotic cardiovascular disease
abstract
Despite striking successes in identifying novel biomarkers for improved patient stratification and predicting disease progression, numerous challenges remain in the effective integration and exploitation of multiomic data in biomedical applications beyond cancer, for which most bioinformatics strategies are developed and validated. That focus on cancer severely limits the effective development and advancement of algorithms in machine learning and artificial intelligence that do not suffer degraded out-of-domain performance. Generalizability and interpretability of models, however, are also required for robust insights that may translate into clinical practice. Work across different independent datasets is critical for establishing models robust towards unwanted variation in assays, protocols, and cohort populations. Disease-specific context like ethnicity, socioeconomic background, sex, lifestyle, disease phase, and tissue type also strongly affect molecular profiles. We here discuss atherosclerotic cardiovascular disease (ASCVD) as a high-impact non-cancer use case for the challenges remaining in the development and application of the latest bioinformatics approaches to multiomics data integration. ASCVD remains the leading cause of death globally. Disease aetiology, progression, and therapy outcome depend on a complex interplay of genetic, environmental, and lifestyle factors. Integrating these diverse data types effectively remains a challenge but holds transformative potential for personalized medicine. Discovery and access to data of sufficient diversity and extent form key bottlenecks. We here compile a first comprehensive overview of key data sets in ASCVD to complement the established cancer-focused resources as a foundation for future effective development and application of state-of-the-art bioinformatics tools for multiomic data integration.
Stephanie Bezzina Wettinger, Kanita Karaduzovic-Hadziabdic, Ritienne Attard, Rosienne Farrugia, Brooke N. Wolford, Marco Chierici, Giuseppe Jurman, Panagiotis Alexiou, José L. Peñalvo, Rafael S. Costa, José Basilio, Frantisek Sabovcik, Rui Vitorino, Johannes A. Schmid, Rajesh Shigdel, Baiba Vilne, Artemis G. Hatzigeorgiou, Miron Sopic, Yvan Devaux, Paolo Magni, Maria Tellez-Plaza, David P. Kreil, Aleksandra Gruca
Briefings Bioinform.7
2025 Comment on "Using genomic data and machine learning to predict antibiotic resistance: A tutorial paper"
abstract
A recent study by Faye Orcales and colleagues proposes a teaching curriculum on supervised machine learning applied to genomics data aimed at predicting antibiotic resistance. The article describes a traditional machine learning pipeline step-by-step in a way that is accessible to anyone, including novices. However, the authors provide a misleading piece of advice in the "Evaluating model performance" section, where they recommend that readers use accuracy and the F1 score for binary classification. We write this short formal comment on that article to reaffirm and explain why accuracy and the F1 score should be avoided in the evaluation of binary classification and why the Matthews correlation coefficient (MCC) should be employed instead. We also take this opportunity to warn readers about the dangers of k-fold cross-validation, which is suggested as a standard method for dividing data into training set and test set, but has several flaws and pitfalls.
Davide Chicco, Giuseppe Jurman
PLoS Comput. Biol.2
2025 Generative AI mitigates representation bias and improves model fairness through synthetic health data
abstract
Representation bias in health data can lead to unfair decisions and compromise the generalisability of research findings. As a consequence, underrepresented subpopulations, such as those from specific ethnic backgrounds or genders, do not benefit equally from clinical discoveries. Several approaches have been developed to mitigate representation bias, ranging from simple resampling methods, such as SMOTE, to recent approaches based on generative adversarial networks (GAN). However, generating high-dimensional time-series synthetic health data remains a significant challenge. In response, we devised a novel architecture (CA-GAN) that synthesises authentic, high-dimensional time series data. CA-GAN outperforms state-of-the-art methods in a qualitative and a quantitative evaluation while avoiding mode collapse, a serious GAN failure. We perform evaluation using 7535 patients with hypotension and sepsis from two diverse, real-world clinical datasets. We show that synthetic data generated by our CA-GAN improves model fairness in Black patients as well as female patients when evaluated separately for each subpopulation. Furthermore, CA-GAN generates authentic data of the minority class while faithfully maintaining the original distribution of data, resulting in improved performance in a downstream predictive task.
Raffaele Marchesi, Nicolo Micheletti, Nicholas I-Hsien Kuo, Sebastiano Barbieri, Giuseppe Jurman, Venet Osmani
PLoS Comput. Biol.5
2023 A statistical comparison between Matthews correlation coefficient (MCC), prevalence threshold, and Fowlkes-Mallows index
Davide Chicco, Giuseppe Jurman
J. Biomed. Informatics2
2021 Cyst segmentation on kidney tubules by means of U-Net deep-learning models
abstract
Autosomal dominant polycystic kidney disease (ADPKD) is one of the most widespread genetic disorders affecting the kidney. Nevertheless, there is still no cure for ADPKD. Domain experts test the effectiveness of different treatments by investigating how they can reduce the number and dimension of cysts on kidney tissues. Image processing of the microscope acquisitions is then an expensive but necessary operation currently performed by operators to determine and compare cyst size and quantity. In this work, we propose a deep learning algorithm for fast and accurate cysts detection in sequential 2-D images. Experiments on 507 RGB immunofluorescence images of 8 kidney tubules show that the proposed U-Net-based deep-learning solution can automatically segment images with increasing performance at larger cyst dimensions (Pr > 0.8, Re > 0.75 for cysts larger than 32 µm2). Such a reliable method performing an accurate cyst segmentation can be a valid support for researchers in optimising the effort to find new effective treatments for ADPKD.
Simone Monaco, Nicole Bussola, Sara Buttò, Diego Sona, Daniele Apiletti, Giuseppe Jurman, Elisa Viola, Marco Chierici, Christodoulos Xinaris, Vincenzo Viola
IEEE BigData6
2019 Integrating deep and radiomics features in cancer bioimaging
abstract
Almost every clinical specialty will use artificial intelligence in the future. The first area of practical impact is expected to be the rapid and accurate interpretation of image streams such as radiology scans, histo-pathology slides, ophthalmic imaging, and any other bioimaging diagnostic systems, enriched by clinical phenotypes used as outcome labels or additional descriptors. In this study, we introduce a machine learning framework for automatic image interpretation that combines the current pattern recognition approach (“radiomics”) with Deep Learning (DL). As a first application in cancer bioimaging, we apply the framework for prognosis of locoregional recurrence in head and neck squamous cell carcinoma (N=298)from Computed Tomography (CT)and Positron Emission Tomography (PET)imaging. The DL architecture is composed of two parallel cascades of Convolutional Neural Network (CNN)layers merging in a softmax classification layer. The network is first pretrained on head and neck tumor stage diagnosis, then fine-tuned on the prognostic task by internal transfer learning. In parallel, radiomics features (e.g., shape of the tumor mass, texture and pixels intensity statistics)are derived by predefined feature extractors on the CT/PET pairs. We compare and mix deep learning and radiomics features into a unifying classification pipeline (RADLER), where model selection and evaluation are based on a data analysis plan developed in the MAQC initiative for reproducible biomarkers. On the multimodal CT/PET cancer dataset, the mixed deep learning/radiomics approach is more accurate than using only one feature type, or image mode. Further, RADLER significantly improves over published results on the same data.
Andrea Bizzego, Nicole Bussola, D. Salvalai, Marco Chierici, Valerio Maggio, Giuseppe Jurman, Cesare Furlanello
CIBCB6
2019 Evaluating reproducibility of AI algorithms in digital pathology with DAPPER
abstract
Artificial Intelligence is exponentially increasing its impact on healthcare. As deep learning is mastering computer vision tasks, its application to digital pathology is natural, with the promise of aiding in routine reporting and standardizing results across trials. Deep learning features inferred from digital pathology scans can improve validity and robustness of current clinico-pathological features, up to identifying novel histological patterns, e.g., from tumor infiltrating lymphocytes. In this study, we examine the issue of evaluating accuracy of predictive models from deep learning features in digital pathology, as an hallmark of reproducibility. We introduce the DAPPER framework for validation based on a rigorous Data Analysis Plan derived from the FDA's MAQC project, designed to analyze causes of variability in predictive biomarkers. We apply the framework on models that identify tissue of origin on 787 Whole Slide Images from the Genotype-Tissue Expression (GTEx) project. We test three different deep learning architectures (VGG, ResNet, Inception) as feature extractors and three classifiers (a fully connected multilayer, Support Vector Machine and Random Forests) and work with four datasets (5, 10, 20 or 30 classes), for a total of 53, 000 tiles at 512 × 512 resolution. We analyze accuracy and feature stability of the machine learning classifiers, also demonstrating the need for diagnostic tests (e.g., random labels) to identify selection bias and risks for reproducibility. Further, we use the deep features from the VGG model from GTEx on the KIMIA24 dataset for identification of slide of origin (24 classes) to train a classifier on 1, 060 annotated tiles and validated on 265 unseen ones. The DAPPER software, including its deep learning pipeline and the Histological Imaging-Newsy Tiles (HINT) benchmark dataset derived from GTEx, is released as a basis for standardization and validation initiatives in AI for digital pathology.
Andrea Bizzego, Nicole Bussola, Marco Chierici, Valerio Maggio, Margherita Francescatto, Luca Cima, Marco Cristoforetti, Giuseppe Jurman, Cesare Furlanello
PLoS Comput. Biol.8
2018 Phylogenetic convolutional neural networks in metagenomics
abstract
BACKGROUND: Convolutional Neural Networks can be effectively used only when data are endowed with an intrinsic concept of neighbourhood in the input space, as is the case of pixels in images. We introduce here Ph-CNN, a novel deep learning architecture for the classification of metagenomics data based on the Convolutional Neural Networks, with the patristic distance defined on the phylogenetic tree being used as the proximity measure. The patristic distance between variables is used together with a sparsified version of MultiDimensional Scaling to embed the phylogenetic tree in a Euclidean space. RESULTS: Ph-CNN is tested with a domain adaptation approach on synthetic data and on a metagenomics collection of gut microbiota of 38 healthy subjects and 222 Inflammatory Bowel Disease patients, divided in 6 subclasses. Classification performance is promising when compared to classical algorithms like Support Vector Machines and Random Forest and a baseline fully connected neural network, e.g. the Multi-Layer Perceptron. CONCLUSION: Ph-CNN represents a novel deep learning approach for the classification of metagenomics data. Operatively, the algorithm has been implemented as a custom Keras layer taking care of passing to the following convolutional layer not only the data but also the ranked list of neighbourhood of each sample, thus mimicking the case of image data, transparently to the user.
Diego Fioravanti, Ylenia Giarratano, Valerio Maggio, Claudio Agostinelli, Marco Chierici, Giuseppe Jurman, Cesare Furlanello
BMC Bioinform.6
2018 Deep learning for automatic stereotypical motor movement detection using wearable sensors in autism spectrum disorders
Nastaran Mohammadian Rad, Seyed Mostafa Kia, Calogero Zarbo, Twan van Laarhoven, Giuseppe Jurman, Paola Venuti, Elena Marchiori, Cesare Furlanello
Signal Process.5
2016 Efficient randomization of biological networks while preserving functional characterization of individual nodes
abstract
BACKGROUND: Networks are popular and powerful tools to describe and model biological processes. Many computational methods have been developed to infer biological networks from literature, high-throughput experiments, and combinations of both. Additionally, a wide range of tools has been developed to map experimental data onto reference biological networks, in order to extract meaningful modules. Many of these methods assess results' significance against null distributions of randomized networks. However, these standard unconstrained randomizations do not preserve the functional characterization of the nodes in the reference networks (i.e. their degrees and connection signs), hence including potential biases in the assessment. RESULTS: Building on our previous work about rewiring bipartite networks, we propose a method for rewiring any type of unweighted networks. In particular we formally demonstrate that the problem of rewiring a signed and directed network preserving its functional connectivity (F-rewiring) reduces to the problem of rewiring two induced bipartite networks. Additionally, we reformulate the lower bound to the iterations' number of the switching-algorithm to make it suitable for the F-rewiring of networks of any size. Finally, we present BiRewire3, an open-source Bioconductor package enabling the F-rewiring of any type of unweighted network. We illustrate its application to a case study about the identification of modules from gene expression data mapped on protein interaction networks, and a second one focused on building logic models from more complex signed-directed reference signaling networks and phosphoproteomic data. CONCLUSIONS: BiRewire3 it is freely available at https://www.bioconductor.org/packages/BiRewire/ , and it should have a broad application as it allows an efficient and analytically derived statistical assessment of results from any network biology tool.
Francesco Iorio, Marti Bernardo-Faura, Andrea Gobbi, Thomas Cokelaer, Giuseppe Jurman, Julio Saez-Rodriguez
BMC Bioinform.5
2015 The HIM glocal metric and kernel for network comparison and classification
abstract
Comparing and classifying graphs represent two essential steps for network analysis, across different scientific and applicative domains. Here we deal with both operations by introducing the Hamming-Ipsen-Mikhailov (HIM) distance, a novel metric to quantitatively measure the difference between two graphs sharing the same vertices. The new measure combines the local Hamming edit distance and the global Ipsen-Mikhailov spectral distance so to overcome the drawbacks affecting the two components when considered separately. Building the kernel function derived from the HIM distance makes possible to move from network comparison to network classification via the Support Vector Machine (SVM) algorithm. Applications of HIM-based methods on synthetic dynamical networks as well as in trade economy and diplomacy datasets demonstrate the effectiveness of HIM as a general purpose solution. An Open Source implementation is provided by the R package nettools, (already configured for High Performance Computing) and the Django-Celery web interface ReNette http://renette.fbk.eu.
Giuseppe Jurman, Roberto Visintainer, Michele Filosi, Samantha Riccadonna, Cesare Furlanello
DSAA1
2014 Fast randomization of large genomic datasets while preserving alteration counts
abstract
MOTIVATION: Studying combinatorial patterns in cancer genomic datasets has recently emerged as a tool for identifying novel cancer driver networks. Approaches have been devised to quantify, for example, the tendency of a set of genes to be mutated in a 'mutually exclusive' manner. The significance of the proposed metrics is usually evaluated by computing P-values under appropriate null models. To this end, a Monte Carlo method (the switching-algorithm) is used to sample simulated datasets under a null model that preserves patient- and gene-wise mutation rates. In this method, a genomic dataset is represented as a bipartite network, to which Markov chain updates (switching-steps) are applied. These steps modify the network topology, and a minimal number of them must be executed to draw simulated datasets independently under the null model. This number has previously been deducted empirically to be a linear function of the total number of variants, making this process computationally expensive. RESULTS: We present a novel approximate lower bound for the number of switching-steps, derived analytically. Additionally, we have developed the R package BiRewire, including new efficient implementations of the switching-algorithm. We illustrate the performances of BiRewire by applying it to large real cancer genomics datasets. We report vast reductions in time requirement, with respect to existing implementations/bounds and equivalent P-value computations. Thus, we propose BiRewire to study statistical properties in genomic datasets, and other data that can be modeled as bipartite networks. AVAILABILITY AND IMPLEMENTATION: BiRewire is available on BioConductor at http://www.bioconductor.org/packages/2.13/bioc/html/BiRewire.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Andrea Gobbi, Francesco Iorio, Kevin J. Dawson, David C. Wedge, David Tamborero, Ludmil B. Alexandrov, Núria López-Bigas, Mathew Garnett, Giuseppe Jurman, Julio Saez-Rodriguez
Bioinform.9
2013 minerva and minepy: a C engine for the MINE suite and its R, Python and MATLAB wrappers
abstract
UNLABELLED: We introduce a novel implementation in ANSI C of the MINE family of algorithms for computing maximal information-based measures of dependence between two variables in large datasets, with the aim of a low memory footprint and ease of integration within bioinformatics pipelines. We provide the libraries minerva (with the R interface) and minepy for Python, MATLAB, Octave and C++. The C solution reduces the large memory requirement of the original Java implementation, has good upscaling properties and offers a native parallelization for the R interface. Low memory requirements are demonstrated on the MINE benchmarks as well as on large ( = 1340) microarray and Illumina GAII RNA-seq transcriptomics datasets. AVAILABILITY AND IMPLEMENTATION: Source code and binaries are freely available for download under GPL3 licence at http://minepy.sourceforge.net for minepy and through the CRAN repository http://cran.r-project.org for the R package minerva. All software is multiplatform (MS Windows, Linux and OSX).
Davide Albanese, Michele Filosi, Roberto Visintainer, Samantha Riccadonna, Giuseppe Jurman, Cesare Furlanello
Bioinform.5
2010 A machine learning pipeline for quantitative phenotype prediction from genotype data
abstract
BACKGROUND: Quantitative phenotypes emerge everywhere in systems biology and biomedicine due to a direct interest for quantitative traits, or to high individual variability that makes hard or impossible to classify samples into distinct categories, often the case with complex common diseases. Machine learning approaches to genotype-phenotype mapping may significantly improve Genome-Wide Association Studies (GWAS) results by explicitly focusing on predictivity and optimal feature selection in a multivariate setting. It is however essential that stringent and well documented Data Analysis Protocols (DAP) are used to control sources of variability and ensure reproducibility of results. We present a genome-to-phenotype pipeline of machine learning modules for quantitative phenotype prediction. The pipeline can be applied for the direct use of whole-genome information in functional studies. As a realistic example, the problem of fitting complex phenotypic traits in heterogeneous stock mice from single nucleotide polymorphims (SNPs) is here considered. METHODS: The core element in the pipeline is the L1L2 regularization method based on the naïve elastic net. The method gives at the same time a regression model and a dimensionality reduction procedure suitable for correlated features. Model and SNP markers are selected through a DAP originally developed in the MAQC-II collaborative initiative of the U.S. FDA for the identification of clinical biomarkers from microarray data. The L1L2 approach is compared with standard Support Vector Regression (SVR) and with Recursive Jump Monte Carlo Markov Chain (MCMC). Algebraic indicators of stability of partial lists are used for model selection; the final panel of markers is obtained by a procedure at the chromosome scale, termed 'saturation', to recover SNPs in Linkage Disequilibrium with those selected. RESULTS: With respect to both MCMC and SVR, comparable accuracies are obtained by the L1L2 pipeline. Good agreement is also found between SNPs selected by the L1L2 algorithms and candidate loci previously identified by a standard GWAS. The combination of L1L2-based feature selection with a saturation procedure tackles the issue of neglecting highly correlated features that affects many feature selection algorithms. CONCLUSIONS: The L1L2 pipeline has proven effective in terms of marker selection and prediction accuracy. This study indicates that machine learning techniques may support quantitative phenotype prediction, provided that adequate DAPs are employed to control bias in model selection.
Giorgio Guzzetta, Giuseppe Jurman, Cesare Furlanello
BMC Bioinform.2
2008 Machine learning methods for predictive proteomics
abstract
The search for predictive biomarkers of disease from high-throughput mass spectrometry (MS) data requires a complex analysis path. Preprocessing and machine-learning modules are pipelined, starting from raw spectra, to set up a predictive classifier based on a shortlist of candidate features. As a machine-learning problem, proteomic profiling on MS data needs caution like the microarray case. The risk of overfitting and of selection bias effects is pervasive: not only potential features easily outnumber samples by 10(3) times, but it is easy to neglect information-leakage effects during preprocessing from spectra to peaks. The aim of this review is to explain how to build a general purpose design analysis protocol (DAP) for predictive proteomic profiling: we show how to limit leakage due to parameter tuning and how to organize classification and ranking on large numbers of replicate versions of the original data to avoid selection bias. The DAP can be used with alternative components, i.e. with different preprocessing methods (peak clustering or wavelet based), classifiers e.g. Support Vector Machine (SVM) or feature ranking methods (recursive feature elimination or I-Relief). A procedure for assessing stability and predictive value of the resulting biomarkers' list is also provided. The approach is exemplified with experiments on synthetic datasets (from the Cromwell MS simulator) and with publicly available datasets from cancer studies.
Annalisa Barla, Giuseppe Jurman, Samantha Riccadonna, Stefano Merler, Marco Chierici, Cesare Furlanello
Briefings Bioinform.2
2008 Algebraic stability indicators for ranked lists in molecular profiling
abstract
MOTIVATION: We propose a method for studying the stability of biomarker lists obtained from functional genomics studies. It is common to adopt resampling methods to tune and evaluate marker-based diagnostic and prognostic systems in order to prevent selection bias. Such caution promotes honest estimation of class prediction, but leads to alternative sets of solutions. In microarray studies, the difference in lists may be bewildering, also due to the presence of modules of functionally related genes. Methods for assessing stability understand the dependency of the markers on the data or on the predictor's type and help selecting solutions. RESULTS: A computational framework for comparing sets of ranked biomarker lists is presented. Notions and algorithms are based on concepts from permutation group theory. We introduce several algebraic indicators and metric methods for symmetric groups, including the Canberra distance, a weighted version of Spearman's footrule. We also consider distances between partial lists and an aggregation of sets of lists into an optimal list based on voting theory (Borda count). The stability indicators are applied in practical situations to several synthetic, cancer microarray and proteomics datasets. The addressed issues are predictive classification, presence of modules, comparison of alternative biomarker lists, outlier removal, control of selection bias by randomization techniques and enrichment analysis. AVAILABILITY: Supplementary Material and software are available at the address http://biodcv.fbk.eu/listspy.html
Giuseppe Jurman, Stefano Merler, Annalisa Barla, Silvano Paoli, Antonio Galea, Cesare Furlanello
Bioinform.1
2008 Integrating gene expression profiling and clinical data
Silvano Paoli, Giuseppe Jurman, Davide Albanese, Stefano Merler, Cesare Furlanello
Int. J. Approx. Reason.2
2006 Proteome Profiling without Selection Bias
abstract
In this paper, we present a method for predictive profiling of mass spectrometry data. The method integrates a spectra preprocessing pipeline with a complete validation setup aimed at identifying the discriminating peaks and at providing an unbiased estimate of the predictive classification error, based on SVM classifiers and on Entropy-based RFE procedure. A particular emphasis is placed upon avoiding selection bias effects throughout all the analysis steps, from preprocessing to peak importance ranking.
Annalisa Barla, Bettina Irler, Stefano Merler, Giuseppe Jurman, Silvano Paoli, Cesare Furlanello
CBMS4
2006 Terminated Ramp-Support Vector Machines: A nonparametric data dependent kernel
Stefano Merler, Giuseppe Jurman
Neural Networks2
2005 Semisupervised Learning for Molecular Profiling
abstract
Class prediction and feature selection are two learning tasks that are strictly paired in the search of molecular profiles from microarray data. Researchers have become aware how easy it is to incur a selection bias effect, and complex validation setups are required to avoid overly optimistic estimates of the predictive accuracy of the models and incorrect gene selections. This paper describes a semisupervised pattern discovery approach that uses the by-products of complete validation studies on experimental setups for gene profiling. In particular, we introduce the study of the patterns of single sample responses (sample-tracking profiles) to the gene selection process induced by typical supervised learning tasks in microarray studies. We originate sample-tracking profiles as the aggregated off-training evaluation of SVM models of increasing gene panel sizes. Genes are ranked by E-RFE, an entropy-based variant of the recursive feature elimination for support vector machines (RFE-SVM). A Dynamic Time Warping (DTW) algorithm is then applied to define a metric between sample-tracking profiles. An unsupervised clustering based on the DTW metric allows automating the discovery of outliers and of subtypes of different molecular profiles. Applications are described on synthetic data and in two gene expression studies.
Cesare Furlanello, Maria Serafini, Stefano Merler, Giuseppe Jurman
IEEE ACM Trans. Comput. Biol. Bioinform.4
2003 Gene selection and classification by entropy-based recursive feature elimination
abstract
We analyse E-RFE (entropy-based recursive feature elimination), a new wrapper algorithm for fast feature ranking in classification problems. The E-RFE method operates the elimination of chunks of uninteresting features according to the entropy of the weights distribution of a SVM classifier. The method is designed to support computationally intensive model selection in classification problems in which the number of features is much larger than the number of samples. We test the elimination procedure on synthetic data sets, and we demonstrate the applicability of E-RFE for the identification of biomedically important genes in predictive classification of microarray data.
Cesare Furlanello, Maria Serafini, Stefano Merler, Giuseppe Jurman
IJCNN4
2003 Entropy-based gene ranking without selection bias for the predictive classification of microarray data
abstract
BACKGROUND: We describe the E-RFE method for gene ranking, which is useful for the identification of markers in the predictive classification of array data. The method supports a practical modeling scheme designed to avoid the construction of classification rules based on the selection of too small gene subsets (an effect known as the selection bias, in which the estimated predictive errors are too optimistic due to testing on samples already considered in the feature selection process). RESULTS: With E-RFE, we speed up the recursive feature elimination (RFE) with SVM classifiers by eliminating chunks of uninteresting genes using an entropy measure of the SVM weights distribution. An optimal subset of genes is selected according to a two-strata model evaluation procedure: modeling is replicated by an external stratified-partition resampling scheme, and, within each run, an internal K-fold cross-validation is used for E-RFE ranking. Also, the optimal number of genes can be estimated according to the saturation of Zipf's law profiles. CONCLUSIONS: Without a decrease of classification accuracy, E-RFE allows a speed-up factor of 100 with respect to standard RFE, while improving on alternative parametric RFE reduction strategies. Thus, a process for gene selection and error estimation is made practical, ensuring control of the selection bias, and providing additional diagnostic indicators of gene importance.
Cesare Furlanello, Maria Serafini, Stefano Merler, Giuseppe Jurman
BMC Bioinform.4
2003 An accelerated procedure for recursive feature ranking on microarray data
Cesare Furlanello, Maria Serafini, Stefano Merler, Giuseppe Jurman
Neural Networks4