Giorgio Valentini

dblp:17/6574 · DBLP profile ↗
← Back
63ranked-venue papers
15as first author
11since 2021 · last 2026
0000-0002-5694-3919ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 8 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Tabular implicit deep neural networks ensembles for the prediction of pathogenic genetic variants in Mendelian diseases
abstract
Mendelian genetic diseases comprise approximately described disorders, yet the genetic basis is known only for about half of them, and a molecular diagnosis often remains difficult or unresolved. In this context, machine learning methods play a key role. However, identifying pathogenic variants in non-coding regions of the human genome is particularly challenging, as they are vastly outnumbered by neutral variants. This extreme imbalance causes standard machine learning approaches to exhibit a strong predictive bias towards the majority class, significantly limiting their sensitivity. Building on recent advances in imbalance-aware and ensemble learning methods, we propose two novel deep learning models for predicting pathogenic non-coding variants in Mendelian diseases. The first model, T-ResNet (Tabular Residual Neural Network), adopts a modular architecture with residual connections, along with a mini-batch balancing strategy to address class imbalance. This design simplifies hyperparameter optimization while mitigating vanishing-gradient effects. The second model, TIDE-Var (Tabular Implicit Deep neural network Ensembles for Variant prediction), leverages the TabM and BatchEnsemble models to build an implicit ensemble of deep neural networks trained jointly by minimizing a common objective function, and partially sharing learning parameters. Learner-specific adapter parameters promote diversity among the implicit base learners while keeping the ensemble computationally efficient. Genome-wide experiments show that T-ResNet achieves an average Area Under the Precision-Recall Curve (AUPRC) of , comparable to that of HyperSMURF , a state-of-the-art method. In contrast, TIDE-Var yields significantly better results than HyperSMURF (AUPRC ) and slightly better than XGBoost , one of the top-methods for the classification of tabular data. Ablation studies confirm that regularized joint ensemble learning with partially shared and base-learner specific adapter learning parameters are key factors in achieving high predictive performance, essential to improve the diagnostic yield for patients with rare genetic diseases.
Federico Stacchietti, Marco Nicolini, Leonardo Chimirri, Peter N. Robinson, Elena Casiraghi, Giorgio Valentini
Neurocomputing6
2025 An Empirical Study of Remote Homology Detection using Protein Language Models
abstract
Detecting remote homologs, proteins that share evolutionary ancestry despite low sequence similarity, remains a central challenge in computational biology. The Structural Classification of Proteins extended (SCOPe) database organizes protein domains into superfamilies based on structural and functional evidence of common origin, making it a widely used benchmark for remote homology detection. In this study, we investigate the effectiveness of protein language model (PLM) embeddings for predicting SCOPe superfamilies directly from sequence. We introduce DOMCLASS, a deep learning framework that combines supervised contrastive learning with a distance-weighted K-NN classifier to learn and exploit an embedding space aligned with SCOPe superfamily annotations. Our empirical results show that general-purpose PLM embeddings already outperform sequence similarity- and profile-based methods for remote homology detection, and that contrastive learning can further improve performance, even in challenging low sequence identity settings.
Ruben Jimenez, Aldo Galeano, Marcelo Báez, Santiago Ferreyra, Guilherme Melo, Giorgio Valentini, Elena Casiraghi, Luca Cernuzzi, Alberto Paccanaro
CLEI6
2025 Biasing second-order random walk sampling for heterogeneous graph embedding *
abstract
We present heterogeneous-node2vec, a novel method that leverages the well-known node2vec algorithm to enable the generation of random-walk samples in a heterogeneous context. Specifically, we propose a strategy to bias the random walk, enabling type-aware transitions between different node and edge types. We evaluate the proposed technique on node-label prediction tasks, applied to various real-world, complex networks. A comparison with state-of-the-art techniques for heterogeneous graph embedding demonstrates that our strategy achieves competitive results for node-label prediction. This evidences that graph representation methods based on heterogeneous random-walk sampling can attain strong performance on standard supervised tasks when the sampling procedure incorporates the semantic information defined by the type heterogeneity of entities within the graph. This approach provides an effective and scalable solution for representing and learning from complex heterogeneous graphs.
Mauricio Soto Gomez, Carlos Cano, Justin T. Reese, Peter N. Robinson, Giorgio Valentini, Elena Casiraghi
IJCNN5
2025 Intrinsic-dimension analysis for guiding dimensionality reduction and data fusion in multi-omics data processing
abstract
Multi-omics data have revolutionized biomedical research by providing a comprehensive understanding of biological systems and the molecular mechanisms of disease development. However, analyzing multi-omics data is challenging due to high dimensionality and limited sample sizes, necessitating proper data-reduction pipelines to ensure reliable analyses. Additionally, its multimodal nature requires effective data-integration pipelines. While several dimensionality reduction and data fusion algorithms have been proposed, crucial aspects are often overlooked. Specifically, the choice of projection space dimension is typically heuristic and uniformly applied across all omics, neglecting the unique high dimension small sample size challenges faced by individual omics. This paper introduces a novel multi-modal dimensionality reduction pipeline tailored to individual views. By leveraging intrinsic dimensionality estimators, we assess the curse-of-dimensionality impact on each view and propose a two-step reduction strategy for significantly affected views, combining feature selection with feature extraction. Compared to traditional uniform reduction pipelines in a crucial and supervised multi-omics analysis setting, our approach shows significant improvement. Additionally, we explore three effective unsupervised multi-omics data fusion methods rooted in the main data fusion strategies to gain insights into their performance under crucial, yet overlooked, settings.
Jessica Gliozzo, Mauricio Soto Gomez, Valentina Guarino, Arturo Bonometti, Alberto Cabri, Emanuele Cavalleri, Justin T. Reese, Peter N. Robinson, Marco Mesiti, Giorgio Valentini, Elena Casiraghi
Artif. Intell. Medicine10
2025 miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources
abstract
MOTIVATION: Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. RESULTS: We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is "task agnostic", in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer's disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. AVAILABILITY AND IMPLEMENTATION: miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.
Jessica Gliozzo, Mauricio Soto Gomez, Arturo Bonometti, Alex Patak, Elena Casiraghi, Giorgio Valentini
Bioinform.6
2023 An expectation-maximization framework for comprehensive prediction of isoform-specific functions
abstract
MOTIVATION: Advances in RNA sequencing technologies have achieved an unprecedented accuracy in the quantification of mRNA isoforms, but our knowledge of isoform-specific functions has lagged behind. There is a need to understand the functional consequences of differential splicing, which could be supported by the generation of accurate and comprehensive isoform-specific gene ontology annotations. RESULTS: We present isoform interpretation, a method that uses expectation-maximization to infer isoform-specific functions based on the relationship between sequence and functional isoform similarity. We predicted isoform-specific functional annotations for 85 617 isoforms of 17 900 protein-coding human genes spanning a range of 17 430 distinct gene ontology terms. Comparison with a gold-standard corpus of manually annotated human isoform functions showed that isoform interpretation significantly outperforms state-of-the-art competing methods. We provide experimental evidence that functionally related isoforms predicted by isoform interpretation show a higher degree of domain sharing and expression correlation than functionally related genes. We also show that isoform sequence similarity correlates better with inferred isoform function than with gene-level function. AVAILABILITY AND IMPLEMENTATION: Source code, documentation, and resource files are freely available under a GNU3 license at https://github.com/TheJacksonLaboratory/isopretEM and https://zenodo.org/record/7594321.
Guy Karlebach, Leigh Carmody, Jagadish Chandrabose Sundaramurthi, Elena Casiraghi, Justin T. Reese, Chris Mungall, Giorgio Valentini, Peter N. Robinson
Bioinform.8
2023 A method for comparing multiple imputation techniques: A case study on the U.S. national COVID cohort collaborative
abstract
Healthcare datasets obtained from Electronic Health Records have proven to be extremely useful for assessing associations between patients' predictors and outcomes of interest. However, these datasets often suffer from missing values in a high proportion of cases, whose removal may introduce severe bias. Several multiple imputation algorithms have been proposed to attempt to recover the missing information under an assumed missingness mechanism. Each algorithm presents strengths and weaknesses, and there is currently no consensus on which multiple imputation algorithm works best in a given scenario. Furthermore, the selection of each algorithm's parameters and data-related modeling choices are also both crucial and challenging. In this paper we propose a novel framework to numerically evaluate strategies for handling missing data in the context of statistical analysis, with a particular focus on multiple imputation techniques. We demonstrate the feasibility of our approach on a large cohort of type-2 diabetes patients provided by the National COVID Cohort Collaborative (N3C) Enclave, where we explored the influence of various patient characteristics on outcomes related to COVID-19. Our analysis included classic multiple imputation techniques as well as simple complete-case Inverse Probability Weighted models. Extensive experiments show that our approach can effectively highlight the most promising and performant missing-data handling strategy for our case study. Moreover, our methodology allowed a better understanding of the behavior of the different models and of how it changed as we modified their parameters. Our method is general and can be applied to different research fields and on datasets containing heterogeneous types.
Elena Casiraghi, Rachel Wong, Margaret Hall, Ben D. Coleman, Marco Notaro, Michael D. Evans, Jena S. Tronieri, Hannah Blau, Bryan Laraway, Tiffany Callahan, Lauren E. Chan, Carolyn T. Bramante, John B. Buse, Richard A. Moffitt, Til Sturmer, Steven G. Johnson, Yu Raymond Shao, Justin T. Reese, Peter N. Robinson, Alberto Paccanaro, Giorgio Valentini, Jared D. Huling, Kenneth Wilkins
J. Biomed. Informatics21
2022 Heterogeneous data integration methods for patient similarity networks
abstract
Patient similarity networks (PSNs), where patients are represented as nodes and their similarities as weighted edges, are being increasingly used in clinical research. These networks provide an insightful summary of the relationships among patients and can be exploited by inductive or transductive learning algorithms for the prediction of patient outcome, phenotype and disease risk. PSNs can also be easily visualized, thus offering a natural way to inspect complex heterogeneous patient data and providing some level of explainability of the predictions obtained by machine learning algorithms. The advent of high-throughput technologies, enabling us to acquire high-dimensional views of the same patients (e.g. omics data, laboratory data, imaging data), calls for the development of data fusion techniques for PSNs in order to leverage this rich heterogeneous information. In this article, we review existing methods for integrating multiple biomedical data views to construct PSNs, together with the different patient similarity measures that have been proposed. We also review methods that have appeared in the machine learning literature but have not yet been applied to PSNs, thus providing a resource to navigate the vast machine learning literature existing on this topic. In particular, we focus on methods that could be used to integrate very heterogeneous datasets, including multi-omics data as well as data derived from clinical information and medical imaging.
Jessica Gliozzo, Marco Mesiti, Marco Notaro, Alessandro Petrini, Alex Patak, Antonio Puertas Gallardo, Alberto Paccanaro, Giorgio Valentini, Elena Casiraghi
Briefings Bioinform.8
2022 Boosting tissue-specific prediction of active cis-regulatory regions through deep learning and Bayesian optimization techniques
abstract
BACKGROUND: Cis-regulatory regions (CRRs) are non-coding regions of the DNA that fine control the spatio-temporal pattern of transcription; they are involved in a wide range of pivotal processes such as the development of specific cell-lines/tissues and the dynamic cell response to physiological stimuli. Recent studies showed that genetic variants occurring in CRRs are strongly correlated with pathogenicity or deleteriousness. Considering the central role of CRRs in the regulation of physiological and pathological conditions, the correct identification of CRRs and of their tissue-specific activity status through Machine Learning methods plays a major role in dissecting the impact of genetic variants on human diseases. Unfortunately, the problem is still open, though some promising results have been already reported by (deep) machine-learning based methods that predict active promoters and enhancers in specific tissues or cell lines by encoding epigenetic or spectral features directly extracted from DNA sequences. RESULTS: We present the experiments we performed to compare two Deep Neural Networks, a Feed-Forward Neural Network model working on epigenomic features, and a Convolutional Neural Network model working only on genomic sequence, targeted to the identification of enhancer- and promoter-activity in specific cell lines. While performing experiments to understand how the experimental setup influences the prediction performance of the methods, we particularly focused on (1) automatic model selection performed by Bayesian optimization and (2) exploring different data rebalancing setups for reducing negative unbalancing effects. CONCLUSIONS: Results show that (1) automatic model selection by Bayesian optimization improves the quality of the learner; (2) data rebalancing considerably impacts the prediction performance of the models; test set rebalancing may provide over-optimistic results, and should therefore be cautiously applied; (3) despite working on sequence data, convolutional models obtain performance close to those of feed forward models working on epigenomic information, which suggests that also sequence data carries informative content for CRR-activity prediction. We therefore suggest combining both models/data types in future works.
Luca Cappelletti, Alessandro Petrini, Jessica Gliozzo, Elena Casiraghi, Max Schubach, Martin Kircher, Giorgio Valentini
BMC Bioinform.7
2021 Semi-automatic Column Type Inference for CSV Table Understanding
Sara Bonfitto, Luca Cappelletti, Fabrizio Trovato, Giorgio Valentini, Marco Mesiti
SOFSEM4
2021 HEMDAG: a family of modular and scalable hierarchical ensemble methods to improve Gene Ontology term prediction
abstract
MOTIVATION: Automated protein function prediction is a complex multi-class, multi-label, structured classification problem in which protein functions are organized in a controlled vocabulary, according to the Gene Ontology (GO). 'Hierarchy-unaware' classifiers, also known as 'flat' methods, predict GO terms without exploiting the inherent structure of the ontology, potentially violating the True-Path-Rule (TPR) that governs the GO, while 'hierarchy-aware' approaches, even if they obey the TPR, do not always show clear improvements with respect to flat methods, or do not scale well when applied to the full GO. RESULTS: To overcome these limitations, we propose Hierarchical Ensemble Methods for Directed Acyclic Graphs (HEMDAG), a family of highly modular hierarchical ensembles of classifiers, able to build upon any flat method and to provide 'TPR-safe' predictions, by leveraging a combination of isotonic regression and TPR learning strategies. Extensive experiments on synthetic and real data across several organisms firstly show that HEMDAG can be used as a general tool to improve the predictions of flat classifiers, and secondly that HEMDAG is competitive versus state-of-the-art hierarchy-aware learning methods proposed in the last CAFA international challenges. AVAILABILITY AND IMPLEMENTATION: Fully tested R code freely available at https://anaconda.org/bioconda/r-hemdag. Tutorial and documentation at https://hemdag.readthedocs.io. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marco Notaro, Marco Frasca 0001, Alessandro Petrini, Jessica Gliozzo, Elena Casiraghi, Peter N. Robinson, Giorgio Valentini
Bioinform.7
2020 Protein function prediction as a graph-transduction game
Sebastiano Vascon, Marco Frasca 0001, Rocco Tripodi, Giorgio Valentini, Marcello Pelillo
Pattern Recognit. Lett.4
2019 On the Quality of Classification Models for Inferring ABAC Policies from Access Logs
abstract
The attribute-based access control (ABAC) model has been gaining popularity in recent years because of its advantages in granularity, flexibility, and usability. Few approaches based on association rules mining have been proposed for the automatic generation of ABAC policies from access logs. Their aim is the identification of policies that do not overfit over training data, are not too general and thus does to disclose sensitive resources to everyone, and are interpretable by humans. The large ABAC privilege space along with the sparsity and unbalance distribution of the available logs make the solution of this task particularly complex and current approaches have different limitations. In this paper we compare different symbolic and nonsymbolic machine learning (ML) techniques for inferring ABAC policies and discuss their pros and cons. Based on experimental results on a toy dataset and on a real dataset, we argue that which is the best technique depends on the characteristics of the considered data. When the data are highly separable according to PCA and t-SNE decomposition, the quality of the obtained ABAC policies is higher and also policies are easily interpretable. By contrast, when this property does not hold, the quality of the obtained policies is low; in this case, non-symbolic ML techniques show better results than the symbolic ones.
Luca Cappelletti, Stefano Valtolina, Giorgio Valentini, Marco Mesiti, Elisa Bertino
IEEE BigData3
2019 Multitask Hopfield Networks
Marco Frasca 0001, Giuliano Grossi, Giorgio Valentini
ECML/PKDD (2)3
2019 UNIPred-Web: a web tool for the integration and visualization of biomolecular networks for protein function prediction
abstract
BACKGROUND: One of the main issues in the automated protein function prediction (AFP) problem is the integration of multiple networked data sources. The UNIPred algorithm was thereby proposed to efficiently integrate -in a function-specific fashion- the protein networks by taking into account the imbalance that characterizes protein annotations, and to subsequently predict novel hypotheses about unannotated proteins. UNIPred is publicly available as R code, which might result of limited usage for non-expert users. Moreover, its application requires efforts in the acquisition and preparation of the networks to be integrated. Finally, the UNIPred source code does not handle the visualization of the resulting consensus network, whereas suitable views of the network topology are necessary to explore and interpret existing protein relationships. RESULTS: We address the aforementioned issues by proposing UNIPred-Web, a user-friendly Web tool for the application of the UNIPred algorithm to a variety of biomolecular networks, already supplied by the system, and for the visualization and exploration of protein networks. We support different organisms and different types of networks -e.g., co-expression, shared domains and physical interaction networks. Users are supported in the different phases of the process, ranging from the selection of the networks and the protein function to be predicted, to the navigation of the integrated network. The system also supports the upload of user-defined protein networks. The vertex-centric and the highly interactive approach of UNIPred-Web allow a narrow exploration of specific proteins, and an interactive analysis of large sub-networks with only a few mouse clicks. CONCLUSIONS: UNIPred-Web offers a practical and intuitive (visual) guidance to biologists interested in gaining insights into protein biomolecular functions. UNIPred-Web provides facilities for the integration of networks, and supplies a framework for the imbalance-aware protein network integration of nine organisms, the prediction of thousands of GO protein functions, and a easy-to-use graphical interface for the visual analysis, navigation and interpretation of the integrated networks and of the functional predictions.
Paolo Perlasca, Marco Frasca 0001, Cheick Tidiane Ba, Marco Notaro, Alessandro Petrini, Elena Casiraghi, Giuliano Grossi, Jessica Gliozzo, Giorgio Valentini, Marco Mesiti
BMC Bioinform.9
2018 A GPU-based algorithm for fast node label learning in large and unbalanced biomolecular networks
abstract
BACKGROUND: Several problems in network biology and medicine can be cast into a framework where entities are represented through partially labeled networks, and the aim is inferring the labels (usually binary) of the unlabeled part. Connections represent functional or genetic similarity between entities, while the labellings often are highly unbalanced, that is one class is largely under-represented: for instance in the automated protein function prediction (AFP) for most Gene Ontology terms only few proteins are annotated, or in the disease-gene prioritization problem only few genes are actually known to be involved in the etiology of a given disease. Imbalance-aware approaches to accurately predict node labels in biological networks are thereby required. Furthermore, such methods must be scalable, since input data can be large-sized as, for instance, in the context of multi-species protein networks. RESULTS: We propose a novel semi-supervised parallel enhancement of COSNET, an imbalance-aware algorithm build on Hopfield neural model recently suggested to solve the AFP problem. By adopting an efficient representation of the graph and assuming a sparse network topology, we empirically show that it can be efficiently applied to networks with millions of nodes. The key strategy to speed up the computations is to partition nodes into independent sets so as to process each set in parallel by exploiting the power of GPU accelerators. This parallel technique ensures the convergence to asymptotically stable attractors, while preserving the asynchronous dynamics of the original model. Detailed experiments on real data and artificial big instances of the problem highlight scalability and efficiency of the proposed method. CONCLUSIONS: By parallelizing COSNET we achieved on average a speed-up of 180x in solving the AFP problem in the S. cerevisiae, Mus musculus and Homo sapiens organisms, while lowering memory requirements. In addition, to show the potential applicability of the method to huge biomolecular networks, we predicted node labels in artificially generated sparse networks involving hundreds of thousands to millions of nodes.
Marco Frasca 0001, Giuliano Grossi, Jessica Gliozzo, Marco Mesiti, Marco Notaro, Paolo Perlasca, Alessandro Petrini, Giorgio Valentini
BMC Bioinform.8
2017 Prediction of Human Phenotype Ontology terms by means of hierarchical ensemble methods
abstract
BACKGROUND: The prediction of human gene-abnormal phenotype associations is a fundamental step toward the discovery of novel genes associated with human disorders, especially when no genes are known to be associated with a specific disease. In this context the Human Phenotype Ontology (HPO) provides a standard categorization of the abnormalities associated with human diseases. While the problem of the prediction of gene-disease associations has been widely investigated, the related problem of gene-phenotypic feature (i.e., HPO term) associations has been largely overlooked, even if for most human genes no HPO term associations are known and despite the increasing application of the HPO to relevant medical problems. Moreover most of the methods proposed in literature are not able to capture the hierarchical relationships between HPO terms, thus resulting in inconsistent and relatively inaccurate predictions. RESULTS: We present two hierarchical ensemble methods that we formally prove to provide biologically consistent predictions according to the hierarchical structure of the HPO. The modular structure of the proposed methods, that consists in a "flat" learning first step and a hierarchical combination of the predictions in the second step, allows the predictions of virtually any flat learning method to be enhanced. The experimental results show that hierarchical ensemble methods are able to predict novel associations between genes and abnormal phenotypes with results that are competitive with state-of-the-art algorithms and with a significant reduction of the computational complexity. CONCLUSIONS: Hierarchical ensembles are efficient computational methods that guarantee biologically meaningful predictions that obey the true path rule, and can be used as a tool to improve and make consistent the HPO terms predictions starting from virtually any flat learning method. The implementation of the proposed methods is available as an R package from the CRAN repository.
Marco Notaro, Max Schubach, Peter N. Robinson, Giorgio Valentini
BMC Bioinform.4
2017 COSNet: An R package for label prediction in unbalanced biological networks
Marco Frasca 0001, Giorgio Valentini
Neurocomputing2
2016 Multi-species protein function prediction: towards web-based visual analytics
abstract
The visualization and analysis of big bio-molecular networks is a key feature for the investigation and prediction of protein functions in a multi-species context. In this paper we present the design of a system that integrates data management, machine learning and visualization facilities to make effective the visual analysis of big networks by means of web-based interfaces.
Paolo Perlasca, Giorgio Valentini, Marco Frasca 0001, Marco Mesiti
iiWAS2
2016 RANKS: a flexible tool for node label ranking and classification in biological networks
abstract
UNLABELLED: RANKS is a flexible software package that can be easily applied to any bioinformatics task formalizable as ranking of nodes with respect to a property given as a label, such as automated protein function prediction, gene disease prioritization and drug repositioning. To this end RANKS provides an efficient and easy-to-use implementation of kernelized score functions, a semi-supervised algorithmic scheme embedding both local and global learning strategies for the analysis of biomolecular networks. To facilitate comparative assessment, baseline network-based methods, e.g. label propagation and random walk algorithms, have also been implemented. AVAILABILITY AND IMPLEMENTATION: The package is available from CRAN: https://cran.r-project.org/ The package is written in R, except for the most computationally intensive functionalities which are implemented in C. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Giorgio Valentini, Giuliano Armano, Marco Frasca 0001, Jianyi Lin, Marco Mesiti, Matteo Ré
Bioinform.1
2016 Learning node labels with multi-category Hopfield networks
Marco Frasca 0001, Simone Bassis, Giorgio Valentini
Neural Comput. Appl.3
2014 An extensive analysis of disease-gene associations using network integration and fast kernel-based gene prioritization methods
abstract
OBJECTIVE: In the context of "network medicine", gene prioritization methods represent one of the main tools to discover candidate disease genes by exploiting the large amount of data covering different types of functional relationships between genes. Several works proposed to integrate multiple sources of data to improve disease gene prioritization, but to our knowledge no systematic studies focused on the quantitative evaluation of the impact of network integration on gene prioritization. In this paper, we aim at providing an extensive analysis of gene-disease associations not limited to genetic disorders, and a systematic comparison of different network integration methods for gene prioritization. MATERIALS AND METHODS: We collected nine different functional networks representing different functional relationships between genes, and we combined them through both unweighted and weighted network integration methods. We then prioritized genes with respect to each of the considered 708 medical subject headings (MeSH) diseases by applying classical guilt-by-association, random walk and random walk with restart algorithms, and the recently proposed kernelized score functions. RESULTS: The results obtained with classical random walk algorithms and the best single network achieved an average area under the curve (AUC) across the 708 MeSH diseases of about 0.82, while kernelized score functions and network integration boosted the average AUC to about 0.89. Weighted integration, by exploiting the different "informativeness" embedded in different functional networks, outperforms unweighted integration at 0.01 significance level, according to the Wilcoxon signed rank sum test. For each MeSH disease we provide the top-ranked unannotated candidate genes, available for further bio-medical investigation. CONCLUSIONS: Network integration is necessary to boost the performances of gene prioritization methods. Moreover the methods based on kernelized score functions can further enhance disease gene ranking results, by adopting both local and global learning strategies, able to exploit the overall topology of the network.
Giorgio Valentini, Alberto Paccanaro, Horacio Caniza, Alfonso E. Romero, Matteo Ré
Artif. Intell. Medicine1
2014 GOssTo: a stand-alone application and a web tool for calculating semantic similarities on the Gene Ontology
abstract
SUMMARY: We present GOssTo, the Gene Ontology semantic similarity Tool, a user-friendly software system for calculating semantic similarities between gene products according to the Gene Ontology. GOssTo is bundled with six semantic similarity measures, including both term- and graph-based measures, and has extension capabilities to allow the user to add new similarities. Importantly, for any measure, GOssTo can also calculate the Random Walk Contribution that has been shown to greatly improve the accuracy of similarity measures. GOssTo is very fast, easy to use, and it allows the calculation of similarities on a genomic scale in a few minutes on a regular desktop machine. CONTACT: [email protected] AVAILABILITY: GOssTo is available both as a stand-alone application running on GNU/Linux, Windows and MacOS from www.paccanarolab.org/gossto and as a web application from www.paccanarolab.org/gosstoweb. The stand-alone application features a simple and concise command line interface for easy integration into high-throughput data processing pipelines.
Horacio Caniza, Alfonso E. Romero, Samuel Heron, Haixuan Yang, Alessandra Devoto, Marco Frasca 0001, Marco Mesiti, Giorgio Valentini, Alberto Paccanaro
Bioinform.8
2013 A neural network algorithm for semi-supervised node label learning from unbalanced data
Marco Frasca 0001, Alberto Bertoni, Matteo Ré, Giorgio Valentini
Neural Networks4
2013 Network-Based Drug Ranking and Repositioning with Respect to DrugBank Therapeutic Categories
abstract
Drug repositioning is a challenging computational problem involving the integration of heterogeneous sources of biomolecular data and the design of label ranking algorithms able to exploit the overall topology of the underlying pharmacological network. In this context, we propose a novel semisupervised drug ranking problem: prioritizing drugs in integrated biochemical networks according to specific DrugBank therapeutic categories. Algorithms for drug repositioning usually perform the inference step into an inhomogeneous similarity space induced by the relationships existing between drugs and a second type of entity (e.g., disease, target, ligand set), thus making unfeasible a drug ranking within a homogeneous pharmacological space. To deal with this problem, we designed a general framework based on bipartite network projections by which homogeneous pharmacological networks can be constructed and integrated from heterogeneous and complementary sources of chemical, biomolecular and clinical information. Moreover, we present a novel algorithmic scheme based on kernelized score functions that adopts both local and global learning strategies to effectively rank drugs in the integrated pharmacological space using different network combination methods. Detailed experiments with more than 80 DrugBank therapeutic categories involving about 1,300 FDA-approved drugs show the effectiveness of the proposed approach.
Matteo Ré, Giorgio Valentini
IEEE ACM Trans. Comput. Biol. Bioinform.2
2013 A Novel Approach to the Problem of Non-uniqueness of the Solution in Hierarchical Clustering
abstract
The existence of multiple solutions in clustering, and in hierarchical clustering in particular, is often ignored in practical applications. However, this is a non-trivial problem, as different data orderings can result in different cluster sets that, in turns, may lead to different interpretations of the same data. The method presented here offers a solution to this issue. It is based on the definition of an equivalence relation over dendrograms that allows developing all and only the significantly different dendrograms for the same dataset, thus reducing the computational complexity to polynomial from the exponential obtained when all possible dendrograms are considered. Experimental results in the neuroimaging and bioinformatics domains show the effectiveness of the proposed method.
Isabella Cattinelli, Giorgio Valentini, Eraldo Paulesu, N. Alberto Borghese
IEEE Trans. Neural Networks Learn. Syst.2
2012 Large Scale Ranking and Repositioning of Drugs with Respect to DrugBank Therapeutic Categories
Matteo Ré, Giorgio Valentini
ISBRA2
2012 Cancer module genes ranking using kernelized score functions
abstract
BACKGROUND: Co-expression based Cancer Modules (CMs) are sets of genes that act in concert to carry out specific functions in different cancer types, and are constructed by exploiting gene expression profiles related to specific clinical conditions or expression signatures associated to specific processes altered in cancer. Unfortunately, genes involved in cancer are not always detectable using only expression signatures or co-expressed sets of genes, and in principle other types of functional interactions should be exploited to obtain a comprehensive picture of the molecular mechanisms underlying the onset and progression of cancer. RESULTS: We propose a novel semi-supervised method to rank genes with respect to CMs using networks constructed from different sources of functional information, not limited to gene expression data. It exploits on the one hand local learning strategies through score functions that extend the guilt-by-association approach, and on the other hand global learning strategies through graph kernels embedded in the score functions, able to take into account the overall topology of the network. The proposed kernelized score functions compare favorably with other state-of-the-art semi-supervised machine learning methods for gene ranking in biological networks and scales well with the number of genes, thus allowing fast processing of very large gene networks. CONCLUSIONS: The modular nature of kernelized score functions provides an algorithmic scheme from which different gene ranking algorithms can be derived, and the results show that using integrated functional networks we can successfully predict CMs defined mainly through expression signatures obtained from gene expression data profiling. A preliminary analysis of top ranked "false positive" genes shows that our approach could be in perspective applied to discover novel genes involved in the onset and progression of tumors related to specific CMs.
Matteo Ré, Giorgio Valentini
BMC Bioinform.2
2012 Synergy of multi-label hierarchical ensembles, data fusion, and cost-sensitive methods for gene functional inference
Nicolò Cesa-Bianchi, Matteo Ré, Giorgio Valentini
Mach. Learn.3
2012 A Fast Ranking Algorithm for Predicting Gene Functions in Biomolecular Networks
abstract
Ranking genes in functional networks according to a specific biological function is a challenging task raising relevant performance and computational complexity problems. To cope with both these problems we developed a transductive gene ranking method based on kernelized score functions able to fully exploit the topology and the graph structure of biomolecular networks and to capture significant functional relationships between genes. We run the method on a network constructed by integrating multiple biomolecular data sources in the yeast model organism, achieving significantly better results than the compared state-of-the-art network-based algorithms for gene function prediction, and with relevant savings in computational time. The proposed approach is general and fast enough to be in perspective applied to other relevant node ranking problems in large and complex biological networks.
Matteo Ré, Marco Mesiti, Giorgio Valentini
IEEE ACM Trans. Comput. Biol. Bioinform.3
2012 Optimisation of the enhanced distance based broadcasting protocol for MANETs
Patricia Ruiz, Bernabé Dorronsoro, Giorgio Valentini, Frédéric Pinel, Pascal Bouvry
J. Supercomput.3
2011 COSNet: A Cost Sensitive Neural Network for Semi-supervised Learning in Graphs
Alberto Bertoni, Marco Frasca 0001, Giorgio Valentini
ECML/PKDD (1)3
2011 A Mathematical Model for the Validation of Gene Selection Methods
abstract
Gene selection methods aim at determining biologically relevant subsets of genes in DNA microarray experiments. However, their assessment and validation represent a major difficulty since the subset of biologically relevant genes is usually unknown. To solve this problem a novel procedure for generating biologically plausible synthetic gene expression data is proposed. It is based on a proper mathematical model representing gene expression signatures and expression profiles through Boolean threshold functions. The results show that the proposed procedure can be successfully adopted to analyze the quality of statistical and machine learning-based gene selection algorithms.
Marco Muselli, Alberto Bertoni, Marco Frasca 0001, Alessandro Beghini, Francesca Ruffino, Giorgio Valentini
IEEE ACM Trans. Comput. Biol. Bioinform.6
2011 True Path Rule Hierarchical Ensembles for Genome-Wide Gene Function Prediction
abstract
Gene function prediction is a complex computational problem, characterized by several items: the number of functional classes is large, and a gene may belong to multiple classes; functional classes are structured according to a hierarchy; classes are usually unbalanced, with more negative than positive examples; class labels can be uncertain and the annotations largely incomplete; to improve the predictions, multiple sources of data need to be properly integrated. In this contribution, we focus on the first three items, and, in particular, on the development of a new method for the hierarchical genome-wide and ontology-wide gene function prediction. The proposed algorithm is inspired by the “true path rule” (TPR) that governs both the Gene Ontology and FunCat taxonomies. According to this rule, the proposed TPR ensemble method is characterized by a two-way asymmetric flow of information that traverses the graph-structured ensemble: positive predictions for a node influence in a recursive way its ancestors, while negative predictions influence its offsprings. Cross-validated results with the model organism S. Crevisiae, using seven different sources of biomolecular data, and a theoretical analysis of the the TPR algorithm show the effectiveness and the drawbacks of the proposed approach.
Giorgio Valentini
IEEE ACM Trans. Comput. Biol. Bioinform.1
2010 Dynamic multi-objective routing algorithm: a multi-objective routing algorithm for the simple hybrid routing protocol on wireless sensor networks
abstract
This study describes a non-dominated algorithm which we call the dynamic multi-objective routing algorithm (DyMORA) developed to improve the simple hybrid routing protocol (SHRP) in choosing the best route towards the Sink node. The multi-objective approach presented allows simultaneous analysis of the four metrics used in the protocol and generates a Pareto-optimal solution. The performance of SHRP concerned to time convergence and reliability with and without DyMORA was analysed via simulation tool NS-2. The performance of SHRP with DyMORA proved to have closed performance to the original SHRP protocol and in many cases superior performance, despite the use of a more complex election algorithm.
Giorgio Valentini, Cláudia Jacy Barenco Abbas, Luis Javier García Villalba, Luis Astorga
IET Commun.1
2010 Integration of heterogeneous data sources for gene function prediction using decision templates and ensembles of learning machines
Matteo Ré, Giorgio Valentini
Neurocomputing2
2009 Fuzzy ensemble clustering based on random projections for DNA microarray data analysis
Roberto Avogadri, Giorgio Valentini
Artif. Intell. Medicine2
2009 Computational intelligence and machine learning in bioinformatics
Giorgio Valentini, Roberto Tagliaferri, Francesco Masulli
Artif. Intell. Medicine1
2009 XML-based approaches for the integration of heterogeneous bio-molecular data
abstract
BACKGROUND: The today's public database infrastructure spans a very large collection of heterogeneous biological data, opening new opportunities for molecular biology, bio-medical and bioinformatics research, but raising also new problems for their integration and computational processing. RESULTS: In this paper we survey the most interesting and novel approaches for the representation, integration and management of different kinds of biological data by exploiting XML and the related recommendations and approaches. Moreover, we present new and interesting cutting edge approaches for the appropriate management of heterogeneous biological data represented through XML. CONCLUSION: XML has succeeded in the integration of heterogeneous biomolecular information, and has established itself as the syntactic glue for biological data sources. Nevertheless, a large variety of XML-based data formats have been proposed, thus resulting in a difficult effective integration of bioinformatics data schemes. The adoption of a few semantic-rich standard formats is urgent to achieve a seamless integration of the current biological resources.
Marco Mesiti, Ernesto Jiménez-Ruiz, Ismael Sanz, Rafael Berlanga Llavori, Paolo Perlasca, Giorgio Valentini, David Manset
BMC Bioinform.6
2008 Dataset complexity can help to generate accurate ensembles of k-nearest neighbors
abstract
Gene expression based cancer classification using classifier ensembles is the main focus of this work. A new ensemble method is proposed that combines predictions of a small number of k-nearest neighbor (k-NN) classifiers with majority vote. Diversity of predictions is guaranteed by assigning a separate feature subset, randomly sampled from the original set of features, to each classifier. Accuracy of k-NNs is ensured by the statistically confirmed dependence between dataset complexity, determining how difficult is a dataset for classification, and classification error. Experiments carried out on three gene expression datasets containing different types of cancer show that our ensemble method is superior to 1) a single best classifier in the ensemble, 2) the nearest shrunken centroids method originally proposed for gene expression data, and 3) the traditional ensemble construction scheme that does not take into account dataset complexity.
Oleg Okun, Giorgio Valentini
IJCNN2
2008 An Algorithm to Assess the Reliability of Hierarchical Clusters in Gene Expression Data
Roberto Avogadri, Matteo Brioschi, Francesca Ruffino, Fulvia Ferrazzi, Alessandro Beghini, Giorgio Valentini
KES (3)6
2008 HCGene: a software tool to support the hierarchical classification of genes
abstract
Abstract Summary: The R package HCGene (Hierarchical Classification of Genes) implements methods to process and analyze the Gene Ontology and the FunCat taxonomy in order to support the functional classification of genes. HCGene allows the extraction of subgraphs and subtrees related to specific biological problems, the labeling of genes and gene products with multiple and hierarchical functional classes, and the association of different types of bio-molecular data to genes for learning to predict their functions. Availability: http://homes.dsi.unimi.it/~valenti/SW/hcgene/download/hcgene_1.0.tar.gz Contact: [email protected] Supplementary information: Supplementary data are available at http://homes.dsi.unimi.it/~valenti/SW/hcgene
Giorgio Valentini, Nicolò Cesa-Bianchi
Bioinform.1
2008 Discovering multi-level structures in bio-molecular data through the Bernstein inequality
abstract
BACKGROUND: The unsupervised discovery of structures (i.e. clusterings) underlying data is a central issue in several branches of bioinformatics. Methods based on the concept of stability have been recently proposed to assess the reliability of a clustering procedure and to estimate the "optimal" number of clusters in bio-molecular data. A major problem with stability-based methods is the detection of multi-level structures (e.g. hierarchical functional classes of genes), and the assessment of their statistical significance. In this context, a chi-square based statistical test of hypothesis has been proposed; however, to assure the correctness of this technique some assumptions about the distribution of the data are needed. RESULTS: To assess the statistical significance and to discover multi-level structures in bio-molecular data, a new method based on Bernstein's inequality is proposed. This approach makes no assumptions about the distribution of the data, thus assuring a reliable application to a large range of bioinformatics problems. Results with synthetic and DNA microarray data show the effectiveness of the proposed method. CONCLUSIONS: The Bernstein test, due to its loose assumptions, is more sensitive than the chi-square test to the detection of multiple structures simultaneously present in the data. Nevertheless it is less selective, that is subject to more false positives, but adding independence assumptions, a more selective variant of the Bernstein inequality-based test is also presented. The proposed methods can be applied to discover multiple structures and to assess their significance in different types of bio-molecular data.
Alberto Bertoni, Giorgio Valentini
BMC Bioinform.2
2008 Gene expression modeling through positive boolean functions
Francesca Ruffino, Marco Muselli, Giorgio Valentini
Int. J. Approx. Reason.3
2007 Mosclust: a software library for discovering significant structures in bio-molecular data
abstract
UNLABELLED: The R package mosclust (model order selection for clustering problems) implements algorithms based on the concept of stability for discovering significant structures in bio-molecular data. The software library provides stability indices obtained through different data perturbations methods (resampling, random projections, noise injection), as well as statistical tests to assess the significance of multi-level structures singled out from the data. AVAILABILITY: http://homes.dsi.unimi.it/~valenti/SW/mosclust/download/mosclust_1.0.tar.gz. SUPPLEMENTARY INFORMATION: http://homes.dsi.unimi.it/~valenti/SW/mosclust.
Giorgio Valentini
Bioinform.1
2007 Model order selection for bio-molecular data clustering
abstract
BACKGROUND: Cluster analysis has been widely applied for investigating structure in bio-molecular data. A drawback of most clustering algorithms is that they cannot automatically detect the "natural" number of clusters underlying the data, and in many cases we have no enough "a priori" biological knowledge to evaluate both the number of clusters as well as their validity. Recently several methods based on the concept of stability have been proposed to estimate the "optimal" number of clusters, but despite their successful application to the analysis of complex bio-molecular data, the assessment of the statistical significance of the discovered clustering solutions and the detection of multiple structures simultaneously present in high-dimensional bio-molecular data are still major problems. RESULTS: We propose a stability method based on randomized maps that exploits the high-dimensionality and relatively low cardinality that characterize bio-molecular data, by selecting subsets of randomized linear combinations of the input variables, and by using stability indices based on the overall distribution of similarity measures between multiple pairs of clusterings performed on the randomly projected data. A chi2-based statistical test is proposed to assess the significance of the clustering solutions and to detect significant and if possible multi-level structures simultaneously present in the data (e.g. hierarchical structures). CONCLUSION: The experimental results show that our model order selection methods are competitive with other state-of-the-art stability based algorithms and are able to detect multiple levels of structure underlying both synthetic and gene expression data.
Alberto Bertoni, Giorgio Valentini
BMC Bioinform.2
2006 Randomized maps for assessing the reliability of patients clusters in DNA microarray data analyses
Alberto Bertoni, Giorgio Valentini
Artif. Intell. Medicine2
2006 Clusterv: a tool for assessing the reliability of clusters discovered in DNA microarray data
abstract
Abstract Summary: We present a new R package for the assessment of the reliability of clusters discovered in high-dimensional DNA microarray data. The package implements methods based on random projections that approximately preserve distances between examples in the projected subspaces. Availability: Contact: [email protected] Supplementary information:
Giorgio Valentini
Bioinform.1
2005 Lung nodules detection and classification
abstract
Image processing techniques and computer aided diagnosis (CAD) systems have proved to be effective for the improvement of radiologists' diagnosis. In this paper an automatic system detecting lung nodules from postero anterior chest radiographs is presented. The system extracts a set of candidate regions by applying to the radiograph three different and consecutive multi-scale schemes. The comparison of the results obtained with those presented in the literature show the efficacy of our multi-scale framework. Learning systems using as input different sets of features have been experimented for candidates classification, showing that support vector machines (SVMs) can be successfully applied for this task.
Paola Campadelli, Elena Casiraghi, Giorgio Valentini
ICIP (1)3
2005 Random projections for assessing gene expression cluster stability
abstract
Clustering analysis of gene expression is characterized by the very high dimensionality and low cardinality of the data, and two important related topics are the validation and the estimate of the number of the obtained clusters. In this paper we focus on the estimate of the stability of the clusters. Our approach to this problem is based on random projections obeying the Johnson-Lindenstrauss lemma, by which gene expression data may be projected into randomly selected low dimensional suhspaces, approximately preserving pairwise distances between examples. We experiment with different types of random projections, comparing empirical and theoretical distortions induced by randomized embeddings between Euclidean metric spaces, and we present cluster-stability measures that may be used to validate and to quantitatively assess the reliability of the clusters obtained by a large class of clustering algorithms. Experimental results with high dimensional synthetic and DNA microarray data show the effectiveness of the proposed approach.
Alberto Bertoni, Giorgio Valentini
IJCNN2
2005 Bio-molecular cancer prediction with random subspace ensembles of support vector machines
Alberto Bertoni, Raffaella Folgieri, Giorgio Valentini
Neurocomputing3
2005 Support vector machines for candidate nodules classification
Paola Campadelli, Elena Casiraghi, Giorgio Valentini
Neurocomputing3
2005 An experimental bias-variance analysis of SVM ensembles based on resampling techniques
abstract
Recently, bias-variance decomposition of error has been used as a tool to study the behavior of learning algorithms and to develop new ensemble methods well suited to the bias-variance characteristics of base learners. We propose methods and procedures, based on Domingo's unified bias-variance theory, to evaluate and quantitatively measure the bias-variance decomposition of error in ensembles of learning machines. We apply these methods to study and compare the bias-variance characteristics of single support vector machines (SVMs) and ensembles of SVMs based on resampling techniques, and their relationships with the cardinality of the training samples. In particular, we present an experimental bias-variance analysis of bagged and random aggregated ensembles of SVMs in order to verify their theoretical variance reduction properties. The experimental bias-variance analysis quantitatively characterizes the relationships between bagging and random aggregating, and explains the reasons why ensembles built on small subsamples of the data work with large databases. Our analysis also suggests new directions for research to improve on classical bagging.
Giorgio Valentini
IEEE Trans. Syst. Man Cybern. Part B1
2004 An experimental analysis of the dependence among codeword bit errors in ECOC learning machines
Francesco Masulli, Giorgio Valentini
Neurocomputing2
2004 Cancer recognition with bagged ensembles of support vector machines
Giorgio Valentini, Marco Muselli, Francesca Ruffino
Neurocomputing1
2004 Bias-Variance Analysis of Support Vector Machines for the Development of SVM-Based Ensemble Methods
Giorgio Valentini, Thomas G. Dietterich
J. Mach. Learn. Res.1
2004 Effectiveness of error correcting output coding methods in ensemble and monolithic learning machines
Francesco Masulli, Giorgio Valentini
Pattern Anal. Appl.2
2003 Low Bias Bagged Support Vector Machines
Giorgio Valentini, Thomas G. Dietterich
ICML1
2003 Bagged ensembles of Support Vector Machines for gene expression data analysis
abstract
Extracting information from gene expression data is a difficult task, as these data are characterized by very high dimensional, small sized, samples and large degree of biological variability. However, a possible way of dealing with the curse of dimensionality is offered by feature selection algorithms, while variance problems arising from small samples and biological variability can be addressed through ensemble methods based on resampling techniques. These two approaches have been combined to improve the accuracy of Support Vector Machines (SVM) in the classification of malignant tissues from DNA microarray data. To assess the accuracy and the confidence of the predictions performed proper measures have been introduced. Presented results show that bagged ensembles of SVM are more reliable and achieve equal or better classification accuracy with respect to single SVM, whereas feature selection methods can further enhance classification accuracy.
Giorgio Valentini, Marco Muselli, Francesca Ruffino
IJCNN1
2002 Gene expression data analysis of human lymphoma using support vector machines and output coding ensembles
Giorgio Valentini
Artif. Intell. Medicine1
2002 NEURObjects: an object-oriented library for neural network development
Giorgio Valentini, Francesco Masulli
Neurocomputing1
2000 Parallel Non Linear Dichotomizers
abstract
We present a new learning machine model for classification problems, based on decompositions of multiclass classification problems in sets of two-class subproblems, assigned to nonlinear dichotomizers that learn their task independently of each other. The experimentation performed on classical data sets, shows that this learning machine model achieves significant performance improvements over MLP, and previous classifiers models based on decomposition of polychotomies into dichotomies. The theoretical reasons of the good properties of generalization of the proposed learning machine model are explained in the framework of the statistical learning theory.
Francesco Masulli, Giorgio Valentini
IJCNN (2)2
2000 Comparing decomposition methods for classification
abstract
Decomposition methods for multiclass classification problems constitute a powerful framework to improve generalization capabilities of a large set of learning machines, including support vector machines and multi-layer perceptrons. We present a review of the main decomposition approach to classification and an experimental comparison of One-Per-Class (OPC), Correcting Classifiers (CC) and Error Correcting Output Codes (ECOC) decomposition methods implemented using multi-layer perceptrons as dichotomizers. The results show that CC and ECOC outperform OPC over the considered data sets.
Francesco Masulli, Giorgio Valentini
KES2