Alberto Paccanaro

dblp:94/3076 · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
5since 2021 · last 2025
0000-0001-8059-1346ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 6 · 2 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 An Empirical Study of Remote Homology Detection using Protein Language Models
abstract
Detecting remote homologs, proteins that share evolutionary ancestry despite low sequence similarity, remains a central challenge in computational biology. The Structural Classification of Proteins extended (SCOPe) database organizes protein domains into superfamilies based on structural and functional evidence of common origin, making it a widely used benchmark for remote homology detection. In this study, we investigate the effectiveness of protein language model (PLM) embeddings for predicting SCOPe superfamilies directly from sequence. We introduce DOMCLASS, a deep learning framework that combines supervised contrastive learning with a distance-weighted K-NN classifier to learn and exploit an embedding space aligned with SCOPe superfamily annotations. Our empirical results show that general-purpose PLM embeddings already outperform sequence similarity- and profile-based methods for remote homology detection, and that contrastive learning can further improve performance, even in challenging low sequence identity settings.
Ruben Jimenez, Aldo Galeano, Marcelo Báez, Santiago Ferreyra, Guilherme Melo, Giorgio Valentini, Elena Casiraghi, Luca Cernuzzi, Alberto Paccanaro
CLEI9
2025 Host centric drug repurposing for viral diseases
abstract
Computational approaches for drug repurposing for viral diseases have mainly focused on a small number of antivirals that directly target pathogens (virus centric therapies). In this work, we combine ideas from collaborative filtering and network medicine for making predictions on a much larger set of drugs that could be repurposed for host centric therapies, that are aimed at interfering with host cell factors required by a pathogen. Our idea is to create matrices quantifying the perturbation that drugs and viruses induce on human protein interaction networks. Then, we decompose these matrices to learn embeddings of drugs, viruses, and proteins in a low dimensional space. Predictions of host-centric antivirals are obtained by taking the dot product between the corresponding drug and virus representations. Our approach is general and can be applied systematically to any compound with known targets and any virus whose host proteins are known. We show that our predictions have high accuracy and that the embeddings contain meaningful biological information that may provide insights into the underlying biology of viral infections. Our approach can integrate different types of information, does not rely on known drug-virus associations and can be applied to new viral diseases and drugs.
Suzana de Siqueira Santos, Haixuan Yang, Aldo Galeano, Alberto Paccanaro
PLoS Comput. Biol.4
2023 A method for comparing multiple imputation techniques: A case study on the U.S. national COVID cohort collaborative
abstract
Healthcare datasets obtained from Electronic Health Records have proven to be extremely useful for assessing associations between patients' predictors and outcomes of interest. However, these datasets often suffer from missing values in a high proportion of cases, whose removal may introduce severe bias. Several multiple imputation algorithms have been proposed to attempt to recover the missing information under an assumed missingness mechanism. Each algorithm presents strengths and weaknesses, and there is currently no consensus on which multiple imputation algorithm works best in a given scenario. Furthermore, the selection of each algorithm's parameters and data-related modeling choices are also both crucial and challenging. In this paper we propose a novel framework to numerically evaluate strategies for handling missing data in the context of statistical analysis, with a particular focus on multiple imputation techniques. We demonstrate the feasibility of our approach on a large cohort of type-2 diabetes patients provided by the National COVID Cohort Collaborative (N3C) Enclave, where we explored the influence of various patient characteristics on outcomes related to COVID-19. Our analysis included classic multiple imputation techniques as well as simple complete-case Inverse Probability Weighted models. Extensive experiments show that our approach can effectively highlight the most promising and performant missing-data handling strategy for our case study. Moreover, our methodology allowed a better understanding of the behavior of the different models and of how it changed as we modified their parameters. Our method is general and can be applied to different research fields and on datasets containing heterogeneous types.
Elena Casiraghi, Rachel Wong, Margaret Hall, Ben D. Coleman, Marco Notaro, Michael D. Evans, Jena S. Tronieri, Hannah Blau, Bryan Laraway, Tiffany Callahan, Lauren E. Chan, Carolyn T. Bramante, John B. Buse, Richard A. Moffitt, Til Sturmer, Steven G. Johnson, Yu Raymond Shao, Justin T. Reese, Peter N. Robinson, Alberto Paccanaro, Giorgio Valentini, Jared D. Huling, Kenneth Wilkins
J. Biomed. Informatics20
2022 Heterogeneous data integration methods for patient similarity networks
abstract
Patient similarity networks (PSNs), where patients are represented as nodes and their similarities as weighted edges, are being increasingly used in clinical research. These networks provide an insightful summary of the relationships among patients and can be exploited by inductive or transductive learning algorithms for the prediction of patient outcome, phenotype and disease risk. PSNs can also be easily visualized, thus offering a natural way to inspect complex heterogeneous patient data and providing some level of explainability of the predictions obtained by machine learning algorithms. The advent of high-throughput technologies, enabling us to acquire high-dimensional views of the same patients (e.g. omics data, laboratory data, imaging data), calls for the development of data fusion techniques for PSNs in order to leverage this rich heterogeneous information. In this article, we review existing methods for integrating multiple biomedical data views to construct PSNs, together with the different patient similarity measures that have been proposed. We also review methods that have appeared in the machine learning literature but have not yet been applied to PSNs, thus providing a resource to navigate the vast machine learning literature existing on this topic. In particular, we focus on methods that could be used to integrate very heterogeneous datasets, including multi-omics data as well as data derived from clinical information and medical imaging.
Jessica Gliozzo, Marco Mesiti, Marco Notaro, Alessandro Petrini, Alex Patak, Antonio Puertas Gallardo, Alberto Paccanaro, Giorgio Valentini, Elena Casiraghi
Briefings Bioinform.7
2021 A Recommender System Approach for Predicting Effective Antivirals
abstract
Emerging infectious diseases such as COVID-19, caused by the SARS-CoV-2 virus, require systematic strategies to assist in the discovery of effective treatments. Drug repositioning, the process of finding new therapeutic indications for commercialized drugs, is a promising alternative to the development of new drugs, with lower costs and shorter development times. In this paper, we propose a recommendation system called geometric confidence non-negative matrix factorization (GcNMF) to assist in the repositioning of 126 broad spectrum antiviral drugs for 80 viruses, including SARS-CoV-2. GcNMF models the non-Euclidean structure of the space using graphs, and produces a ranked list of drugs for each virus. Our experiments reveal that GcNMF significanlty outperforms other matrix decomposition methods at predicting missing drug-virus associations. Our analysis suggests that GcNMF could assist pharmacological experts in the search for effective drugs against viral diseases.
Rafael Adorno, Diego Galeano, Diego H. Stalder, Luca Cernuzzi, Alberto Paccanaro
CLEI5
2020 The corrected gene proximity map for analyzing the 3D genome organization using Hi-C data
abstract
BACKGROUND: Genome-wide ligation-based assays such as Hi-C provide us with an unprecedented opportunity to investigate the spatial organization of the genome. Results of a typical Hi-C experiment are often summarized in a chromosomal contact map, a matrix whose elements reflect the co-location frequencies of genomic loci. To elucidate the complex structural and functional interactions between those genomic loci, networks offer a natural and powerful framework. RESULTS: We propose a novel graph-theoretical framework, the Corrected Gene Proximity (CGP) map to study the effect of the 3D spatial organization of genes in transcriptional regulation. The starting point of the CGP map is a weighted network, the gene proximity map, whose weights are based on the contact frequencies between genes extracted from genome-wide Hi-C data. We derive a null model for the network based on the signal contributed by the 1D genomic distance and use it to "correct" the gene proximity for cell type 3D specific arrangements. The CGP map, therefore, provides a network framework for the 3D structure of the genome on a global scale. On human cell lines, we show that the CGP map can detect and quantify gene co-regulation and co-localization more effectively than the map obtained by raw contact frequencies. Analyzing the expression pattern of metabolic pathways of two hematopoietic cell lines, we find that the relative positioning of the genes, as captured and quantified by the CGP, is highly correlated with their expression change. We further show that the CGP map can be used to form an inter-chromosomal proximity map that allows large-scale abnormalities, such as chromosomal translocations, to be identified. CONCLUSIONS: The Corrected Gene Proximity map is a map of the 3D structure of the genome on a global scale. It allows the simultaneous analysis of intra- and inter- chromosomal interactions and of gene co-regulation and co-localization more effectively than the map obtained by raw contact frequencies, thus revealing hidden associations between global spatial positioning and gene expression. The flexible graph-based formalism of the CGP map can be easily generalized to study any existing Hi-C datasets.
Alberto Paccanaro, Mark Gerstein, Koon-Kiu Yan
BMC Bioinform.2
2019 Disease gene prediction for molecularly uncharacterized diseases
abstract
Network medicine approaches have been largely successful at increasing our knowledge of molecularly characterized diseases. Given a set of disease genes associated with a disease, neighbourhood-based methods and random walkers exploit the interactome allowing the prediction of further genes for that disease. In general, however, diseases with no known molecular basis constitute a challenge. Here we present a novel network approach to prioritize gene-disease associations that is able to also predict genes for diseases with no known molecular basis. Our method, which we have called Cardigan (ChARting DIsease Gene AssociatioNs), uses semi-supervised learning and exploits a measure of similarity between disease phenotypes. We evaluated its performance at predicting genes for both molecularly characterized and uncharacterized diseases in OMIM, using both weighted and binary interactomes, and compared it with state-of-the-art methods. Our tests, which use datasets collected at different points in time to replicate the dynamics of the disease gene discovery process, prove that Cardigan is able to accurately predict disease genes for molecularly uncharacterized diseases. Additionally, standard leave-one-out cross validation tests show how our approach outperforms state-of-the-art methods at predicting genes for molecularly characterized diseases by 14%-65%. Cardigan can also be used for disease module prediction, where it outperforms state-of-the-art methods by 87%-299%.
Juan J. Caceres, Alberto Paccanaro
PLoS Comput. Biol.2
2018 A Recommender System Approach for Predicting Drug Side Effects
abstract
The accurate identification of drug side effects represents a major concern for public health. We propose a collaborative filtering model for large-scale prediction of drug side effects. Our approach provides side effects recommendations for drugs to safety professionals. The proposed latent factor model relies solely on the public drug-side effect relationships from safety data. Applied to 1,525 marketed drugs and 2,050 side effect terms, we achieved an AUPRC (area under the precision- recall curve) of 0.342 in a test set, with a sensitivity of 0.73 given a specificity of 0.95, providing state-of-the-art performance in side effect prediction. We analyze the performance of the method on drug-specific Anatomical Therapeutic and Chemical (ATC) category and side effect- specific medical category of disorders. Our findings suggest that latent factor models can be useful for the early and accurate detection of unknown adverse drug events.
Diego Galeano, Alberto Paccanaro
IJCNN2
2017 Mining the biomedical literature to predict shared drug targets in DrugBank
abstract
The current drug development pipelines are characterised by long processes with high attrition rates and elevated costs. More than 80% of new compounds fail in the later stages of testing due to severe side-effects caused by unknown biomolecular targets of the compounds. In this work, we present a measure that can predict shared targets for drugs in DrugBank through large scale analysis of the biomedical literature. We show that using MeSH ontology terms can accurately describe the drugs and that appropriate use of the MeSH ontological structure can determine pairwise drug similarity.
Horacio Caniza, Diego Galeano, Alberto Paccanaro
CLEI3
2017 Drug cocktail selection for the treatment of chagas disease: A multi-objective approach
abstract
Chagas disease is a parasitic disease, endemic in South America. As of today, there is no effective treatment in its chronic stage. We have recently identified 134 FDA approved drugs with potential antitrypanosomal activity. In this paper, we propose a novel method for selecting combinations of drugs (drug cocktails), to provide a more effective treatment against Chagas disease. We define three measures to evaluate the predicted performance of a cocktail, establishing in this way a mathematical foundation for its analysis. This allows us to model the drug cocktail selection as a multi-objective optimisation problem, that we show can be solved efficiently with state-of-the-art evolutionary algorithms. Our analysis retrieves 57 drug cocktails containing between 2 and 6 drugs. We discuss the improvement of the cocktail selection given by our method, and the application of this approach to the identification of cocktails against other parasitic diseases.
Mateo Torres, Juan J. Caceres, Ruben Jimenez, Victor Yubero, Celeste Vega, Miriam Rolon, Luca Cernuzzi, Benjamín Barán, Alberto Paccanaro
CLEI9
2016 Combining interactomes from multiple organisms: A case study on human-mouse
abstract
The amount and quality of available data on different organisms varies greatly. While model organisms benefit from extensive experimental studies, there is often a lack of detailed experimental data for more specific organisms. Additionally, even among model organisms there are noticeable differences in the amount and type of data available, due to the different suitability of experiments in different organisms. The combination of interactomes for closely related species, represents a viable tool to increase the amount of protein-protein interaction data for a given organism. The Human-Mouse case of study is particularly relevant, as many experiments cannot be carried out on humans. This paper describes a general method to construct a combined interactome from different organisms. The construction is achieved through the integration of data from different sources and formats, including gene-protein relations, protein homology classes and protein-protein interactions. We show that the Human-Mouse combined interactome increase the mouse gene coverage by over 150% and the interaction coverage by over 430%. We also provide a novel mathematical formalisation for the interactome combination.
Juan J. Caceres, Alberto Paccanaro
CLEI2
2016 Drug targets prediction using chemical similarity
abstract
The growing productivity gap between investment in drug research and development (R&D) and the number of new medicines approved by the US Food and Drug Administration (FDA) in the past decade is concerning. This productivity problem raises the need for innovative approaches for drug-target prediction and a deeper understanding of the interplay between drugs and their target proteins. Chemogenomics is the interdisciplinary field which aims to predict gene/protein/ligand relationships. The predictions are based on the assumption that chemically similar compounds should share common targets. Here, we exploit our understanding of the network-based representation of the protein-protein interaction (PPI network) to introduce a distance between drug-targets and could verify whether it correlates with their chemical similarity. We build a fully connected graph composed of US Food and Drug Administration (FDA) - approved drugs using the Tanimoto 2D similarity based on fingerprints from the SMILES representation of the chemical structure. Our analysis of 1165 FDA-approved drugs indicates that the chemical similarity of drugs predicts closeness of their targets in the human interactome.
Diego Galeano, Alberto Paccanaro
CLEI2
2014 An extensive analysis of disease-gene associations using network integration and fast kernel-based gene prioritization methods
abstract
OBJECTIVE: In the context of "network medicine", gene prioritization methods represent one of the main tools to discover candidate disease genes by exploiting the large amount of data covering different types of functional relationships between genes. Several works proposed to integrate multiple sources of data to improve disease gene prioritization, but to our knowledge no systematic studies focused on the quantitative evaluation of the impact of network integration on gene prioritization. In this paper, we aim at providing an extensive analysis of gene-disease associations not limited to genetic disorders, and a systematic comparison of different network integration methods for gene prioritization. MATERIALS AND METHODS: We collected nine different functional networks representing different functional relationships between genes, and we combined them through both unweighted and weighted network integration methods. We then prioritized genes with respect to each of the considered 708 medical subject headings (MeSH) diseases by applying classical guilt-by-association, random walk and random walk with restart algorithms, and the recently proposed kernelized score functions. RESULTS: The results obtained with classical random walk algorithms and the best single network achieved an average area under the curve (AUC) across the 708 MeSH diseases of about 0.82, while kernelized score functions and network integration boosted the average AUC to about 0.89. Weighted integration, by exploiting the different "informativeness" embedded in different functional networks, outperforms unweighted integration at 0.01 significance level, according to the Wilcoxon signed rank sum test. For each MeSH disease we provide the top-ranked unannotated candidate genes, available for further bio-medical investigation. CONCLUSIONS: Network integration is necessary to boost the performances of gene prioritization methods. Moreover the methods based on kernelized score functions can further enhance disease gene ranking results, by adopting both local and global learning strategies, able to exploit the overall topology of the network.
Giorgio Valentini, Alberto Paccanaro, Horacio Caniza, Alfonso E. Romero, Matteo Ré
Artif. Intell. Medicine2
2014 GOssTo: a stand-alone application and a web tool for calculating semantic similarities on the Gene Ontology
abstract
SUMMARY: We present GOssTo, the Gene Ontology semantic similarity Tool, a user-friendly software system for calculating semantic similarities between gene products according to the Gene Ontology. GOssTo is bundled with six semantic similarity measures, including both term- and graph-based measures, and has extension capabilities to allow the user to add new similarities. Importantly, for any measure, GOssTo can also calculate the Random Walk Contribution that has been shown to greatly improve the accuracy of similarity measures. GOssTo is very fast, easy to use, and it allows the calculation of similarities on a genomic scale in a few minutes on a regular desktop machine. CONTACT: [email protected] AVAILABILITY: GOssTo is available both as a stand-alone application running on GNU/Linux, Windows and MacOS from www.paccanarolab.org/gossto and as a web application from www.paccanarolab.org/gosstoweb. The stand-alone application features a simple and concise command line interface for easy integration into high-throughput data processing pipelines.
Horacio Caniza, Alfonso E. Romero, Samuel Heron, Haixuan Yang, Alessandra Devoto, Marco Frasca 0001, Marco Mesiti, Giorgio Valentini, Alberto Paccanaro
Bioinform.9
2012 Improving GO semantic similarity measures by exploring the ontology beneath the terms and modelling uncertainty
abstract
Abstract Motivation: Several measures have been recently proposed for quantifying the functional similarity between gene products according to well-structured controlled vocabularies where biological terms are organized in a tree or in a directed acyclic graph (DAG) structure. However, existing semantic similarity measures ignore two important facts. First, when calculating the similarity between two terms, they disregard the descendants of these terms. While this makes no difference when the ontology is a tree, we shall show that it has important consequences when the ontology is a DAG—this is the case, for example, with the Gene Ontology (GO). Second, existing similarity measures do not model the inherent uncertainty which comes from the fact that our current knowledge of the gene annotation and of the ontology structure is incomplete. Here, we propose a novel approach based on downward random walks that can be used to improve any of the existing similarity measures to exhibit these two properties. The approach is computationally efficient—random walks do not need to be simulated as we provide formulas to calculate their stationary distributions. Results: To show that our approach can potentially improve any semantic similarity measure, we test it on six different semantic similarity measures: three commonly used measures by Resnik (1999), Lin (1998), and Jiang and Conrath (1997); and three recently proposed measures: simUI, simGIC by Pesquita et al. (2008); GraSM by Couto et al. (2007); and Couto and Silva (2011). We applied these improved measures to the GO annotations of the yeast Saccharomyces cerevisiae, and tested how they correlate with sequence similarity, mRNA co-expression and protein–protein interaction data. Our results consistently show that the use of downward random walks leads to more reliable similarity measures. Availability: We have developed a suite of tools that implement existing semantic similarity measures and our improved measures based on random walks. The tools are implemented in Matlab and are freely available from: http://www.paccanarolab.org/papers/GOsim/ Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Haixuan Yang, Tamás Nepusz, Alberto Paccanaro
Bioinform.3
2010 SCPS: a fast implementation of a spectral method for detecting protein families on a genome-wide scale
abstract
BACKGROUND: An important problem in genomics is the automatic inference of groups of homologous proteins from pairwise sequence similarities. Several approaches have been proposed for this task which are "local" in the sense that they assign a protein to a cluster based only on the distances between that protein and the other proteins in the set. It was shown recently that global methods such as spectral clustering have better performance on a wide variety of datasets. However, currently available implementations of spectral clustering methods mostly consist of a few loosely coupled Matlab scripts that assume a fair amount of familiarity with Matlab programming and hence they are inaccessible for large parts of the research community. RESULTS: SCPS (Spectral Clustering of Protein Sequences) is an efficient and user-friendly implementation of a spectral method for inferring protein families. The method uses only pairwise sequence similarities, and is therefore practical when only sequence information is available. SCPS was tested on difficult sets of proteins whose relationships were extracted from the SCOP database, and its results were extensively compared with those obtained using other popular protein clustering algorithms such as TribeMCL, hierarchical clustering and connected component analysis. We show that SCPS is able to identify many of the family/superfamily relationships correctly and that the quality of the obtained clusters as indicated by their F-scores is consistently better than all the other methods we compared it with. We also demonstrate the scalability of SCPS by clustering the entire SCOP database (14,183 sequences) and the complete genome of the yeast Saccharomyces cerevisiae (6,690 sequences). CONCLUSIONS: Besides the spectral method, SCPS also implements connected component analysis and hierarchical clustering, it integrates TribeMCL, it provides different cluster quality tools, it can extract human-readable protein descriptions using GI numbers from NCBI, it interfaces with external tools such as BLAST and Cytoscape, and it can produce publication-quality graphical representations of the clusters obtained, thus constituting a comprehensive and effective tool for practical research in computational biology. Source code and precompiled executables for Windows, Linux and Mac OS X are freely available at http://www.paccanarolab.org/software/scps.
Tamás Nepusz, Rajkumar Sasidharan, Alberto Paccanaro
BMC Bioinform.3
2006 Predicting interactions in protein networks by completing defective cliques
abstract
UNLABELLED: Datasets obtained by large-scale, high-throughput methods for detecting protein-protein interactions typically suffer from a relatively high level of noise. We describe a novel method for improving the quality of these datasets by predicting missed protein-protein interactions, using only the topology of the protein interaction network observed by the large-scale experiment. The central idea of the method is to search the protein interaction network for defective cliques (nearly complete complexes of pairwise interacting proteins), and predict the interactions that complete them. We formulate an algorithm for applying this method to large-scale networks, and show that in practice it is efficient and has good predictive performance. More information can be found on our website http://topnet.gersteinlab.org/clique/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary Materials are available at Bioinformatics online.
Haiyuan Yu, Alberto Paccanaro, Valery Trifonov, Mark Gerstein
Bioinform.2
2005 Inferring protein-protein interactions using interaction network topologies
abstract
We describe two novel methods for predicting protein interactions, using only the topology of an observed protein interaction network. The first method searches the protein interaction network for defective cliques (i.e. nearly complete complexes of pair wise interacting proteins), and predicts the interactions that complete them. The second method computes the diffusion distance between each pair of proteins and then infers an interaction when such distance is below a given threshold. We show that both methods have a good predictive performance and compare their results.
Alberto Paccanaro, Valery Trifonov, Haiyuan Yu, Mark Gerstein
IJCNN1
2003 Learning Distributed Representations of High-Arity Relational Data with Non-linear Relational Embedding
Alberto Paccanaro
ICANN1
2003 Spectral clustering of protein sequences
abstract
An important problem in genomics is automatically clustering homologous proteins when only sequence information is available. Most methods for clustering proteins are local, and are based on simply thresholding a measure related to sequence distance. We first show how locality limits the performance of such methods by analysing the distribution of distances between protein sequences. We then present a global method based on spectral clustering and provide theoretical justification of why it will have a remarkable improvement over local methods. We extensively tested our method and compared its performance with other local methods on several subsets of the SCOP (Structural Classification of Proteins) database, a gold standard for protein structure classification. We consistently observed that, the number of clusters that we obtain for a given set of proteins is close to the number of superfamilies in that set; there are fewer singletons; and the method correctly groups most remote homologs. In our experiments, the quality of the clusters as quantified by a measure that combines sensitivity and specificity was consistently better [on average, improvements were 84 % over hierarchical clustering, 34 % over Connected Component Analysis (CCA) (similar to GeneRAGE) and 72 % over another global method, TribeMCL].
Alberto Paccanaro, S. Chakra Chennubhotla, James A. Casbon, Mansoor A. S. Saqi
IJCNN1
2001 Learning Hierarchical Structures with Linear Relational Embedding
abstract
We present Linear Relational Embedding (LRE), a new method of learn- ing a distributed representation of concepts from data consisting of in- stances of relations between given concepts. Its final goal is to be able to generalize, i.e. infer new instances of these relations among the con- cepts. On a task involving family relationships we show that LRE can generalize better than any previously published method. We then show how LRE can be used effectively to find compact distributed representa- tions for variable-sized recursive data structures, such as trees and lists. 1 Linear Relational Embedding Our aim is to take a large set of facts about a domain expressed as tuples of arbitrary sym- bols in a simple and rigid syntactic format and to be able to infer other “common-sense” facts without having any prior knowledge about the domain. Let us imagine a situation in which we have a set of concepts and a set of relations among these concepts, and that our data consists of few instances of these relations that hold among the concepts. We want to be able to infer other instances of these relations. For example, if the concepts are the people in a certain family, the relations are kinship relations, and we are given the facts ”Alberto has-father Pietro” and ”Pietro has-brother Giovanni”, we would like to be able to infer ”Alberto has-uncle Giovanni”. Our approach is to learn appropriate distributed rep- resentations of the entities in the data, and then exploit the generalization properties of the distributed representations [2] to make the inferences. In this paper we present a method, which we have called Linear Relational Embedding (LRE), which learns a distributed rep- resentation for the concepts by embedding them in a space where the relations between concepts are linear transformations of their distributed representations. involve two concepts. Let us consider the case in which all the relations are binary, i.e. , and the problem In this case our data consists of triplets we are trying to solve is to infer missing triplets when we are given only few of them. Inferring a triplet is equivalent to being able to complete it, that is to come up with one of its elements, given the other two. Here we shall always try to complete the third element of the triplets 1. LRE will then represent each concept in the data as a learned vector in a
Alberto Paccanaro, Geoffrey E. Hinton
NIPS1
2001 Learning Distributed Representations of Concepts Using Linear Relational Embedding
abstract
We introduce linear relational embedding as a means of learning a distributed representation of concepts from data consisting of binary relations between these concepts. The key idea is to represent concepts as vectors, binary relations as matrices, and the operation of applying a relation to a concept as a matrix-vector multiplication that produces an approximation to the related concept. A representation for concepts and relations is learned by maximizing an appropriate discriminative goodness function using gradient ascent. On a task involving family relationships, learning is fast and leads to good generalization.
Alberto Paccanaro, Geoffrey E. Hinton
IEEE Trans. Knowl. Data Eng.1
2000 Learning Distributed Representations by Mapping Concepts and Relations into a Linear Space
Alberto Paccanaro, Geoffrey E. Hinton
ICML1
2000 Extracting Distributed Representations of Concepts and Relations from Positive and Negative Propositions
abstract
Linear relational embedding (LRE) was introduced previously by the authors (1999) as a means of extracting a distributed representation of concepts from relational data. The original formulation cannot use negative information and cannot properly handle data in which there are multiple correct answers. In this paper we propose an extended formulation of LRE that solves both these problems. We present results in two simple domains, which show that learning leads to good generalization.
Alberto Paccanaro, Geoffrey E. Hinton
IJCNN (2)1
1995 Guiding Term Reduction Through a Neural Network: Some Prelimanary Results for the Group Theory
Alberto Paccanaro
RTA1