Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Fabio Fabris

dblp:237/7172 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
1since 2021 · last 2021
0000-0001-7159-4668ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
2 papers
Trustworthy machine learning · 43% Deep learning architectures and training · 28% Knowledge representation and reasoning · 28%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
modular neural network
0.412020
Using deep learning to associate human genes with age-related diseases · Bioinform. 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning › information fusion
multisource information fusion
0.412020
Using deep learning to associate human genes with age-related diseases · Bioinform. 2020
Bioinformatics and computational biology
deep learning classification
0.412020
Using deep learning to associate human genes with age-related diseases · Bioinform. 2020
Bioinformatics and computational biology › biomarker discovery
disease-gene association
0.412020
Using deep learning to associate human genes with age-related diseases · Bioinform. 2020
Machine learning › Trustworthy machine learning › interpretability
feature importance
0.312018
A new approach for interpreting Random Forest models and its application to the biology of ageing · Bioinform. 2018
Machine learning › Trustworthy machine learning
interpretability
0.312018
A new approach for interpreting Random Forest models and its application to the biology of ageing · Bioinform. 2018
Bioinformatics and computational biology
gene expression analysis
0.312018
A new approach for interpreting Random Forest models and its application to the biology of ageing · Bioinform. 2018
Bioinformatics and computational biology › systems bioinformatics
pathway analysis
0.212016
New KEGG pathway-based interpretable features for classifying ageing-related mouse proteins · Bioinform. 2016
Bioinformatics and computational biology › protein function prediction
protein classification
0.212016
New KEGG pathway-based interpretable features for classifying ageing-related mouse proteins · Bioinform. 2016

Methods — techniques the papers use, named apart from their topics

logistic regression · 0.9gradient boosted trees · 0.9deep neural network · 0.9random forest · 0.7feature importance measures · 0.3feature importance measure · 0.3supervised machine learning · 0.2hierarchical classification · 0.2
YearPublicationVenuePosition
2021 A Novel Feature Selection Method for Uncertain Features: An Application to the Prediction of Pro-/Anti-Longevity Genes
abstract
Understanding the ageing process is a very challenging problem for biologists. To help in this task, there has been a growing use of classification methods (from machine learning) to learn models that predict whether a gene influences the process of ageing or promotes longevity. One type of predictive feature often used for learning such classification models is Protein-Protein Interaction (PPI) features. One important property of PPI features is their uncertainty, i.e., a given feature (PPI annotation) is often associated with a confidence score, which is usually ignored by conventional classification methods. Hence, we propose the Lazy Feature Selection for Uncertain Features (LFSUF) method, which is tailored for coping with the uncertainty in PPI confidence scores. In addition, following the lazy learning paradigm, LFSUF selects features for each instance to be classified, making the feature selection process more flexible. We show that our LFSUF method achieves better predictive accuracy when compared to other feature selection methods that either do not explicitly take PPI confidence scores into account or deal with uncertainty globally rather than using a per-instance approach. Also, we interpret the results of the classification process using the features selected by LFSUF, showing that the number of selected features is significantly reduced, assisting the interpretability of the results. The datasets used in the experiments and the program code of the LFSUF method are freely available on the web at http://github.com/pablonsilva/FSforUncertainFeatureSpaces.
Pablo Nascimento da Silva, Alexandre Plastino 0001, Fabio Fabris, Alex Alves Freitas
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 Comparing enrichment analysis and machine learning for identifying gene properties that discriminate between gene classes
abstract
Biologists very often use enrichment methods based on statistical hypothesis tests to identify gene properties that are significantly over-represented in a given set of genes of interest, by comparison with a 'background' set of genes. These enrichment methods, although based on rigorous statistical foundations, are not always the best single option to identify patterns in biological data. In many cases, one can also use classification algorithms from the machine-learning field. Unlike enrichment methods, classification algorithms are designed to maximize measures of predictive performance and are capable of analysing combinations of gene properties, instead of one property at a time. In practice, however, the majority of studies use either enrichment or classification methods (rather than both), and there is a lack of literature discussing the pros and cons of both types of method. The goal of this paper is to compare and contrast enrichment and classification methods, offering two contributions. First, we discuss the (to some extent complementary) advantages and disadvantages of both types of methods for identifying gene properties that discriminate between gene classes. Second, we provide a set of high-level recommendations for using enrichment and classification methods. Overall, by highlighting the strengths and the weaknesses of both types of methods we argue that both should be used in bioinformatics analyses.
Fabio Fabris, Daniel Palmer, João Pedro de Magalhães, Alex Alves Freitas
Briefings Bioinform.1
2020 Using deep learning to associate human genes with age-related diseases
abstract
MOTIVATION: One way to identify genes possibly associated with ageing is to build a classification model (from the machine learning field) capable of classifying genes as associated with multiple age-related diseases. To build this model, we use a pre-compiled list of human genes associated with age-related diseases and apply a novel Deep Neural Network (DNN) method to find associations between gene descriptors (e.g. Gene Ontology terms, protein-protein interaction data and biological pathway information) and age-related diseases. RESULTS: The novelty of our new DNN method is its modular architecture, which has the capability of combining several sources of biological data to predict which ageing-related diseases a gene is associated with (if any). Our DNN method achieves better predictive performance than standard DNN approaches, a Gradient Boosted Tree classifier (a strong baseline method) and a Logistic Regression classifier. Given the DNN model produced by our method, we use two approaches to identify human genes that are not known to be associated with age-related diseases according to our dataset. First, we investigate genes that are close to other disease-associated genes in a complex multi-dimensional feature space learned by the DNN algorithm. Second, using the class label probabilities output by our DNN approach, we identify genes with a high probability of being associated with age-related diseases according to the model. We provide evidence of these putative associations retrieved from the DNN model with literature support. AVAILABILITY AND IMPLEMENTATION: The source code and datasets can be found at: https://github.com/fabiofabris/Bioinfo2019. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fabio Fabris, Daniel Palmer, Khalid M. Salama, João Pedro de Magalhães, Alex Alves Freitas
Bioinform.1
2018 A new approach for interpreting Random Forest models and its application to the biology of ageing
abstract
Motivation: This work uses the Random Forest (RF) classification algorithm to predict if a gene is over-expressed, under-expressed or has no change in expression with age in the brain. RFs have high predictive power, and RF models can be interpreted using a feature (variable) importance measure. However, current feature importance measures evaluate a feature as a whole (all feature values). We show that, for a popular type of biological data (Gene Ontology-based), usually only one value of a feature is particularly important for classification and the interpretation of the RF model. Hence, we propose a new algorithm for identifying the most important and most informative feature values in an RF model. Results: The new feature importance measure identified highly relevant Gene Ontology terms for the aforementioned gene classification task, producing a feature ranking that is much more informative to biologists than an alternative, state-of-the-art feature importance measure. Availability and implementation: The dataset and source codes used in this paper are available as 'Supplementary Material' and the description of the data can be found at: https://fabiofabris.github.io/bioinfo2018/web/. Supplementary information: Supplementary data are available at Bioinformatics online.
Fabio Fabris, Aoife Doherty, Daniel Palmer, João Pedro de Magalhães, Alex Alves Freitas
Bioinform.1
2017 A Situation-Aware Fear Learning (SAFEL) model for robots
abstract
This work proposes a novel Situation-Aware FEar Learning (SAFEL) model for robots. SAFEL combines concepts of situation-aware expert systems with well-known neuroscientific findings on the brain fear-learning mechanism to allow companion robots to predict undesirable or threatening situations based on past experiences. One of the main objectives is to allow robots to learn complex temporal patterns of sensed environmental stimuli and create a representation of these patterns. This memory can be later associated with a negative or positive “emotion”, analogous to fear and confidence. Experiments with a real robot demonstrated SAFEL's success in generating contextual fear conditioning behavior with predictive capabilities based on situational information.
Caroline Rizzi Raymundo, Colin G. Johnson, Fabio Fabris, Patrícia Amâncio Vargas
Neurocomputing3
2016 New KEGG pathway-based interpretable features for classifying ageing-related mouse proteins
abstract
MOTIVATION: The incidence of ageing-related diseases has been constantly increasing in the last decades, raising the need for creating effective methods to analyze ageing-related protein data. These methods should have high predictive accuracy and be easily interpretable by ageing experts. To enable this, one needs interpretable classification models (supervised machine learning) and features with rich biological meaning. In this paper we propose two interpretable feature types based on Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways and compare them with traditional feature types in hierarchical classification (a more challenging classification task regarding predictive performance) and binary classification (a classification task producing easier to interpret classification models). As far as we know, this work is the first to: (i) explore the potential of the KEGG pathway data in the hierarchical classification setting, (i) use the graph structure of KEGG pathways to create a feature type that quantifies the influence of a current protein on another specific protein within a KEGG pathway graph and (iii) propose a method for interpreting the classification models induced using KEGG features. RESULTS: We performed tests measuring predictive accuracy considering hierarchical and binary class labels extracted from the Mouse Phenotype Ontology. One of the KEGG feature types leads to the highest predictive accuracy among five individual feature types across three hierarchical classification algorithms. Additionally, the combination of the two KEGG feature types proposed in this work results in one of the best predictive accuracies when using the binary class version of our datasets, at the same time enabling the extraction of knowledge from ageing-related data using quantitative influence information. AVAILABILITY AND IMPLEMENTATION: The datasets created in this paper will be freely available after publication. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fabio Fabris, Alex Alves Freitas
Bioinform.1
2016 An Extensive Empirical Comparison of Probabilistic Hierarchical Classifiers in Datasets of Ageing-Related Genes
abstract
This study comprehensively evaluates the performance of five types of probabilistic hierarchical classification methods used for predicting Gene Ontology (GO) terms related to ageing. Of those tested, a new hybrid of a Local Hierarchical Classifier (LHC) and the Predictive Clustering Tree algorithm (LHC-PCT) had the best predictive accuracy results. We also tested the impact of two types of variations in most hierarchical classification algorithms, namely: (a) changing the base algorithm (we tested Naive Bayes and Support Vector Machines), and the impact of (b) using or not the Correlation based Feature Selection (CFS) algorithm in a pre-processing step. In total, we evaluated the predictive performance of 17 variations of hierarchical classifiers across 15 datasets of ageing and longevity-related genes. We conclude that the LHC-PCT algorithm ranks better across several tests (seven out of 12). In addition, we interpreted the models generated by the PCT algorithm to show how hierarchical classification algorithms can be used to extract biological insights out of the ageing-related datasets that we compiled.
Fabio Fabris, Alex Alves Freitas, Jennifer M. A. Tullet
IEEE ACM Trans. Comput. Biol. Bioinform.1
2015 A Novel Extended Hierarchical Dependence Network Method Based on Non-hierarchical Predictive Classes and Applications to Ageing-Related Data
abstract
We propose a novel algorithm for hierarchical classification, the Hierarchical Dependence Network based on non-Hierarchical Predictive Classes (HDN-nHPC) algorithm. HDN-nHPC uses relationships among predictive classes that are not descendants or ancestors of each other to improve classification performance and, at the same time, provide insights to non-obvious predictive class relationships. To test our algorithm and baselines, we have used hierarchical ageing-related datasets where the classes are terms in the Gene Ontology. We have concluded, based on our experiments, that using non-hierarchical predictive class relationships improves the performance of the classification algorithm and that, considering one out of three accuracy measures, the HDN-nHPC is statistically significantly better than the other three algorithms that we have tested, while no statistical significant differences were found on the other two measures.
Fabio Fabris, Alex Alves Freitas
ICTAI1
2014 Dependency network methods for Hierarchical Multi-label Classification of gene functions
abstract
Hierarchical Multi-label Classification (HMC) is a challenging real-world problem that naturally emerges in several areas. This work proposes two new algorithms using a Probabilistic Graphical Model based on Dependency Networks (DN) to solve the HMC problem of classifying gene functions into pre-established class hierarchies. DNs are especially attractive for their capability of using traditional, “out-of-the-shelf”, classification algorithms to model the relationship among classes and for their ability to cope with cyclic dependencies, resulting in greater flexibility with respect to Bayesian Networks. We tested our two algorithms: the first is a stand-alone Hierarchical Dependency Network (HDN) algorithm, and the second is a hybrid between the HDN and the Predictive Clustering Tree (PCT) algorithm, a well-known classifier for HMC. Based on our experiments, the hybrid classifier, using SVMs as base classifiers, obtained higher predictive accuracy than both the standard PCT algorithm and the HDN algorithm, considering 22 bioinformatics datasets and two out of three predictive accuracy measures specific for hierarchical classification (AU(PRC) and AUPRCw).
Fabio Fabris, Alex Alves Freitas
CIDM1