Paolo Frasconi

dblp:73/32 · DBLP profile ↗
← Back
80ranked-venue papers
19as first author
1since 2021 · last 2021
0000-0003-3117-9245ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 60 · 14 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-authorDatabases, data management, data science and information retrieval · 7 · 4 first-authorTheory of computation · 5 · 2 first-authorComputer networks · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
21 papers
Optimization for machine learning · 30% Transfer learning and domain adaptation · 14% Knowledge representation and reasoning · 13%
Interdisciplinary, comprehensive, and emerging computing
12 papers
Bioinformatics and computational biology · 79% Computational science and engineering · 12% Medical and health informatics · 9%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%
Theoretical computer science
3 papers
Graph algorithms and graph theory · 68% Algorithms and data structures · 32%

Topics — the 30 heaviest of 57, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
hyperparameter optimization
1.132020
Marthe: Scheduling the Learning Rate Via Online Hypergradients · IJCAI 2020
Bilevel Programming for Hyperparameter Optimization and Meta-Learning · ICML 2018
Forward and Reverse Gradient-Based Hyperparameter Optimization · ICML 2017
Machine learning › Transfer learning and domain adaptation
meta-learning
0.622018
Bilevel Programming for Hyperparameter Optimization and Meta-Learning · ICML 2018
Forward and Reverse Gradient-Based Hyperparameter Optimization · ICML 2017
Computer vision › 3D vision › geometric deep learning
set learning
0.512021
Learning Aggregation Functions · IJCAI 2021
Machine learning › Trustworthy machine learning
interpretability
0.412020
Learning and Interpreting Multi-Multi-Instance Learning Networks · J. Mach. Learn. Res. 2020
Machine learning › Trustworthy machine learning › interpretability › explainable AI
interpretable neural network
0.412020
Learning and Interpreting Multi-Multi-Instance Learning Networks · J. Mach. Learn. Res. 2020
Machine learning › Optimization for machine learning
learning rate schedule
0.412020
Marthe: Scheduling the Learning Rate Via Online Hypergradients · IJCAI 2020
Machine learning › Learning paradigms
multiple instance learning
0.412020
Learning and Interpreting Multi-Multi-Instance Learning Networks · J. Mach. Learn. Res. 2020
Machine learning › Deep learning architectures and training
training optimization
0.412020
Marthe: Scheduling the Learning Rate Via Online Hypergradients · IJCAI 2020
Audio and music processing
music generation
0.412020
Pattern-Based Music Generation with Wasserstein Autoencoders and PRC Descriptions · IJCAI 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
statistical relational learning
0.432014
kLog: A language for logical and relational learning with kernels · Artif. Intell. 2014
A Statistical Relational Learning Approach to Identifying Evidence Based Medicine Categories · EMNLP-CoNLL 2012
kFOIL: Learning Simple Relational Kernels · AAAI 2006
Bioinformatics and computational biology
protein structure prediction
0.462009
Prediction of protein beta-residue contacts by Markov logic networks with grounding-specific weights · Bioinform. 2009
MetalDetector: a web server for predicting metal-binding sites and disulfide bridges in proteins from sequence · Bioinform. 2008
Predicting the Geometry of Metal Binding Sites from Protein Sequence · NIPS 2008
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.332015
kLog: A Language for Logical and Relational Learning with Kernels (Extended Abstract) · IJCAI 2015
kFOIL: Learning Simple Relational Kernels · AAAI 2006
Weighted decomposition kernels · ICML 2005
Machine learning › Optimization for machine learning
bilevel optimization
0.312018
Bilevel Programming for Hyperparameter Optimization and Meta-Learning · ICML 2018
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.312018
Bilevel Programming for Hyperparameter Optimization and Meta-Learning · ICML 2018
Machine learning › Optimization for machine learning › hyperparameter optimization
gradient-based hyperparameter optimization
0.312017
Forward and Reverse Gradient-Based Hyperparameter Optimization · ICML 2017
Computational science and engineering › information retrieval
recommender systems
0.212016
RNAcommender: genome-wide recommendation of RNA-protein interactions · Bioinform. 2016
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
RNA-protein interaction prediction
0.212016
RNAcommender: genome-wide recommendation of RNA-protein interactions · Bioinform. 2016
Bioinformatics and computational biology › protein function prediction › functional site prediction
metal-binding site prediction
0.232008
MetalDetector: a web server for predicting metal-binding sites and disulfide bridges in proteins from sequence · Bioinform. 2008
Predicting the Geometry of Metal Binding Sites from Protein Sequence · NIPS 2008
Improving Prediction of Zinc Binding Sites by Modeling the Linkage Between Residues Close in Sequence · RECOMB 2006
Graph algorithms and graph theory › graph theory › graph similarity
graph kernels
0.212015
Graph Invariant Kernels · IJCAI 2015
Bioinformatics and computational biology › neuroscience
neuroanatomy
0.212014
Large-scale automated identification of mouse brain cells in confocal light sheet microscopy images · Bioinform. 2014
Machine learning › Representation and self-supervised learning › feature transformation
feature construction
0.212013
Type Extension Trees for feature construction and learning in relational domains · Artif. Intell. 2013
Knowledge, reasoning and agents › Knowledge representation and reasoning
relational learning
0.212013
Type Extension Trees for feature construction and learning in relational domains · Artif. Intell. 2013
Data mining › structured data mining
relational data mining
0.212013
Type Extension Trees for feature construction and learning in relational domains · Artif. Intell. 2013
Medical and health informatics
evidence-based medicine
0.112012
A Statistical Relational Learning Approach to Identifying Evidence Based Medicine Categories · EMNLP-CoNLL 2012
Bioinformatics and computational biology › protein structure prediction
residue contact prediction
0.122009
Prediction of protein beta-residue contacts by Markov logic networks with grounding-specific weights · Bioinform. 2009
Prediction of Protein Topologies Using Generalized IOHMMS and RNNs · NIPS 2002
Bioinformatics and computational biology › protein structure prediction
disulfide connectivity prediction
0.122008
MetalDetector: a web server for predicting metal-binding sites and disulfide bridges in proteins from sequence · Bioinform. 2008
Disulfide connectivity prediction using recursive neural networks and evolutionary information · Bioinform. 2004
Bioinformatics and computational biology › molecular informatics
cheminformatics
0.112007
Classification of small molecules by two- and three-dimensional decomposition kernels · Bioinform. 2007
Bioinformatics and computational biology › molecular informatics
molecular classification
0.112007
Classification of small molecules by two- and three-dimensional decomposition kernels · Bioinform. 2007
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic programming
inductive logic programming
0.112006
Kernels on Prolog Proof Trees: Statistical Learning in the ILP Setting · J. Mach. Learn. Res. 2006
Bioinformatics and computational biology
protein function prediction
0.112006
Improving Prediction of Zinc Binding Sites by Modeling the Linkage Between Residues Close in Sequence · RECOMB 2006

Methods — techniques the papers use, named apart from their topics

kernel methods · 0.6deep sets · 0.5attention mechanism · 0.5wasserstein autoencoder · 0.4online hypergradient · 0.4hypergradient approximation · 0.4boolean function learning · 0.4bag-layer · 0.4DCGAN-style convolutional architecture · 0.4logical and relational learning · 0.4gradient-based optimization · 0.3bi-level optimization · 0.3secondary structure features · 0.2recommender system · 0.2domain composition features · 0.2statistical relational learning · 0.2machine learning · 0.2image analysis · 0.2
YearPublicationVenuePosition
2021 Learning Aggregation Functions
abstract
Learning on sets is increasingly gaining attention in the machine learning community, due to its widespread applicability. Typically, representations over sets are computed by using fixed aggregation functions such as sum or maximum. However, recent results showed that universal function representation by sum- (or max-) decomposition requires either highly discontinuous (and thus poorly learnable) mappings, or a latent dimension equal to the maximum number of elements in the set. To mitigate this problem, we introduce LAF (Learning Aggregation Function), a learnable aggregator for sets of arbitrary cardinality. LAF can approximate several extensively used aggregators (such as average, sum, maximum) as well as more complex functions (e.g. variance and skewness). We report experiments on semi-synthetic and real data showing that LAF outperforms state-of-the-art sum- (max-) decomposition architectures such as DeepSets and library-based architectures like Principal Neighborhood Aggregation, and can be effectively combined with attention-based architectures.
Giovanni Pellegrini, Alessandro Tibo, Paolo Frasconi, Andrea Passerini, Manfred Jaeger
IJCAI3
2020 Pattern-Based Music Generation with Wasserstein Autoencoders and PRC Descriptions
abstract
We demonstrate a pattern-based MIDI music generation system with a generation strategy based on Wasserstein autoencoders and a novel variant of pianoroll descriptions of patterns which employs separate channels for note velocities and note durations and can be fed into classic DCGAN-style convolutional architectures. We trained the system on two new datasets (in the acid-jazz and high-pop genres) composed by musicians in our team with music generation in mind. Our demonstration shows that moving smoothly in the latent space allows us to generate meaningful sequences of four-bars patterns.
Tijn Borghuis, Luca Angioloni, Lorenzo Brusci, Paolo Frasconi
IJCAI4
2020 Marthe: Scheduling the Learning Rate Via Online Hypergradients
abstract
We study the problem of fitting task-specific learning rate schedules from the perspective of hyperparameter optimization, aiming at good generalization. We describe the structure of the gradient of a validation error w.r.t. the learning rate schedule -- the hypergradient. Based on this, we introduce MARTHE, a novel online algorithm guided by cheap approximations of the hypergradient that uses past information from the optimization trajectory to simulate future behaviour. It interpolates between two recent techniques, RTHO (Franceschi et al., 2017) and HD (Baydin et al. 2018), and is able to produce learning rate schedules that are more stable leading to models that generalize better.
Michele Donini, Luca Franceschi 0001, Orchid Majumder, Massimiliano Pontil, Paolo Frasconi
IJCAI5
2020 Learning and Interpreting Multi-Multi-Instance Learning Networks
abstract
We introduce an extension of the multi-instance learning problem where examples are organized as nested bags of instances (e.g., a document could be represented as a bag of sentences, which in turn are bags of words). This framework can be useful in various scenarios, such as text and image classification, but also supervised learning over graphs. As a further advantage, multi-multi instance learning enables a particular way of interpreting predictions and the decision function. Our approach is based on a special neural network layer, called bag-layer, whose units aggregate bags of inputs of arbitrary size. We prove theoretically that the associated class of functions contains all Boolean functions over sets of sets of instances and we provide empirical evidence that functions of this kind can be actually learned on semi-synthetic datasets. We finally present experiments on text classification, on citation graphs, and social graph data, which show that our model obtains competitive results with respect to accuracy when compared to other approaches such as convolutional networks on graphs, while at the same time it supports a general approach to interpret the learnt model, as well as explain individual predictions.
Alessandro Tibo, Manfred Jaeger, Paolo Frasconi
J. Mach. Learn. Res.3
2020 Classification of Cancer Pathology Reports: A Large-Scale Comparative Study
abstract
We report about the application of state-of-the-art deep learning techniques to the automatic and interpretable assignment of ICD-O3 topography and morphology codes to free-text cancer reports. We present results on a large dataset (more than 80 000 labeled and 1 500 000 unlabeled anonymized reports written in Italian and collected from hospitals in Tuscany over more than a decade) and with a large number of classes (134 morphological classes and 61 topographical classes). We compare alternative architectures in terms of prediction accuracy and interpretability and show that our best model achieves a multiclass accuracy of 90.3% on topography site assignment and 84.8% on morphology type assignment. We found that in this context hierarchical models are not better than flat models and that an element-wise maximum aggregator is slightly better than attentive models on site classification. Moreover, the maximum aggregator offers a way to interpret the classification process.
Stefano Martina, Leonardo Ventura, Paolo Frasconi
IEEE J. Biomed. Health Informatics3
2018 Bilevel Programming for Hyperparameter Optimization and Meta-Learning
abstract
We introduce a framework based on bilevel programming that unifies gradient-based hyperparameter optimization and meta-learning. We show that an approximate version of the bilevel problem can be solved by taking into explicit account the optimization dynamics for the inner objective. Depending on the specific setting, the outer variables take either the meaning of hyperparameters in a supervised learning problem or parameters of a meta-learner. We provide sufficient conditions under which solutions of the approximate problem converge to those of the exact problem. We instantiate our approach for meta-learning in the case of deep learning where representation layers are treated as hyperparameters shared across a set of training episodes. In experiments, we confirm our theoretical findings, present encouraging results for few-shot learning and contrast the bilevel approach against classical approaches for learning-to-learn.
Luca Franceschi 0001, Paolo Frasconi, Saverio Salzo, Riccardo Grazzi, Massimiliano Pontil
ICML2
2017 Forward and Reverse Gradient-Based Hyperparameter Optimization
abstract
We study two procedures (reverse-mode and forward-mode) for computing the gradient of the validation error with respect to the hyperparameters of any iterative learning algorithm such as stochastic gradient descent. These procedures mirror two ways of computing gradients for recurrent neural networks and have different trade-offs in terms of running time and space requirements. Our formulation of the reverse-mode procedure is linked to previous work by Maclaurin et al (2015) but does not require reversible dynamics. Additionally, we explore the use of constraints on the hyperparameters. The forward-mode procedure is suitable for real-time hyperparameter updates, which may significantly speedup hyperparameter optimization on large datasets. We present a series of experiments on image and phone classification tasks. In the second task, previous gradient-based approaches are prohibitive. We show that our real-time algorithm yields state-of-the-art results in affordable time.
Luca Franceschi 0001, Michele Donini, Paolo Frasconi, Massimiliano Pontil
ICML3
2017 A Network Architecture for Multi-Multi-Instance Learning
Alessandro Tibo, Paolo Frasconi, Manfred Jaeger
ECML/PKDD (1)2
2017 kProbLog: an algebraic Prolog for machine learning
Francesco Orsini, Paolo Frasconi, Luc De Raedt
Mach. Learn.2
2016 RNAcommender: genome-wide recommendation of RNA-protein interactions
abstract
MOTIVATION: Information about RNA-protein interactions is a vital pre-requisite to tackle the dissection of RNA regulatory processes. Despite the recent advances of the experimental techniques, the currently available RNA interactome involves a small portion of the known RNA binding proteins. The importance of determining RNA-protein interactions, coupled with the scarcity of the available information, calls for in silico prediction of such interactions. RESULTS: We present RNAcommender, a recommender system capable of suggesting RNA targets to unexplored RNA binding proteins, by propagating the available interaction information taking into account the protein domain composition and the RNA predicted secondary structure. Our results show that RNAcommender is able to successfully suggest RNA interactors for RNA binding proteins using little or no interaction evidence. RNAcommender was tested on a large dataset of human RBP-RNA interactions, showing a good ranking performance (average AUC ROC of 0.75) and significant enrichment of correct recommendations for 75% of the tested RBPs. RNAcommender can be a valid tool to assist researchers in identifying potential interacting candidates for the majority of RBPs with uncharacterized binding preferences. AVAILABILITY AND IMPLEMENTATION: The software is freely available at http://rnacommender.disi.unitn.it CONTACT: [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online.
Gianluca Corrado, Toma Tebaldi, Fabrizio Costa, Paolo Frasconi, Andrea Passerini
Bioinform.4
2015 kLog: A Language for Logical and Relational Learning with Kernels (Extended Abstract)
Paolo Frasconi, Fabrizio Costa, Luc De Raedt, Kurt De Grave
IJCAI1
2015 Graph Invariant Kernels
Francesco Orsini, Paolo Frasconi, Luc De Raedt
IJCAI2
2015 kProbLog: An Algebraic Prolog for Kernel Programming
Francesco Orsini, Paolo Frasconi, Luc De Raedt
ILP2
2014 kLog: A language for logical and relational learning with kernels
Paolo Frasconi, Fabrizio Costa, Luc De Raedt, Kurt De Grave
Artif. Intell.1
2014 Large-scale automated identification of mouse brain cells in confocal light sheet microscopy images
abstract
MOTIVATION: Recently, confocal light sheet microscopy has enabled high-throughput acquisition of whole mouse brain 3D images at the micron scale resolution. This poses the unprecedented challenge of creating accurate digital maps of the whole set of cells in a brain. RESULTS: We introduce a fast and scalable algorithm for fully automated cell identification. We obtained the whole digital map of Purkinje cells in mouse cerebellum consisting of a set of 3D cell center coordinates. The method is accurate and we estimated an F1 measure of 0.96 using 56 representative volumes, totaling 1.09 GVoxel and containing 4138 manually annotated soma centers. AVAILABILITY AND IMPLEMENTATION: Source code and its documentation are available at http://bcfind.dinfo.unifi.it/. The whole pipeline of methods is implemented in Python and makes use of Pylearn2 and modified parts of Scikit-learn. Brain images are available on request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Paolo Frasconi, Ludovico Silvestri, Paolo Soda, Roberto Cortini, Francesco Pavone, Giulio Iannello
Bioinform.1
2013 A relational kernel-based approach to scene classification
abstract
Real-world scenes involve many objects that interact with each other in complex semantic patterns. For example, a bar scene can be naturally described as having a variable number of chairs of similar size, close to each other and aligned horizontally. This high-level interpretation of a scene relies on semantically meaningful entities and is most generally described using relational representations or (hyper-) graphs. Popular in early work on syntactic and structural pattern recognition, relational representations are rarely used in computer vision due to their pure symbolic nature. Yet, today recent successes in combining them with statistical learning principles motivates us to reinvestigate their use. In this paper we show that relational techniques can also improve scene classification. More specifically, we employ a new relational language for learning with kernels, called kLog. With this language we define higher-order spatial relations among semantic objects. When applied to a particular image, they characterize a particular object arrangement and provide discriminative cues for the scene category. The kernel allows us to tractably learn from such complex features. Thus, our contribution is a principled and interpretable approach to learn from symbolic relations how to classify scenes in a statistical framework. We obtain results comparable to state-of-the-art methods on 15 Scenes and a subset of the MIT indoor dataset.
Laura Antanas, McElory Hoffmann, Paolo Frasconi, Tinne Tuytelaars, Luc De Raedt
WACV3
2013 Type Extension Trees for feature construction and learning in relational domains
Manfred Jaeger, Marco Lippi 0001, Andrea Passerini, Paolo Frasconi
Artif. Intell.4
2013 Short-Term Traffic Flow Forecasting: An Experimental Comparison of Time-Series Analysis and Supervised Learning
abstract
The literature on short-term traffic flow forecasting has undergone great development recently. Many works, describing a wide variety of different approaches, which very often share similar features and ideas, have been published. However, publications presenting new prediction algorithms usually employ different settings, data sets, and performance measurements, making it difficult to infer a clear picture of the advantages and limitations of each model. The aim of this paper is twofold. First, we review existing approaches to short-term traffic flow forecasting methods under the common view of probabilistic graphical models, presenting an extensive experimental comparison, which proposes a common baseline for their performance analysis and provides the infrastructure to operate on a publicly available data set. Second, we present two new support vector regression models, which are specifically devised to benefit from typical traffic flow seasonality and are shown to represent an interesting compromise between prediction accuracy and computational efficiency. The SARIMA model coupled with a Kalman filter is the most accurate model; however, the proposed seasonal support vector regressor turns out to be highly competitive when performing forecasts during the most congested periods.
Marco Lippi 0001, Matteo Bertini, Paolo Frasconi
IEEE Trans. Intell. Transp. Syst.3
2012 A Statistical Relational Learning Approach to Identifying Evidence Based Medicine Categories
Mathias Verbeke, Vincent Van Asch, Roser Morante, Paolo Frasconi, Walter Daelemans, Luc De Raedt
EMNLP-CoNLL4
2012 Metal Binding in Proteins: Machine Learning Complements X-Ray Absorption Spectroscopy
Marco Lippi 0001, Andrea Passerini, Marco Punta, Paolo Frasconi
ECML/PKDD (2)4
2012 Guest Editors' introduction - Special issue on inductive logic programming (ILP 2010)
Paolo Frasconi, Francesca A. Lisi
Mach. Learn.1
2012 Predicting Metal-Binding Sites from Protein Sequence
abstract
Prediction of binding sites from sequence can significantly help toward determining the function of uncharacterized proteins on a genomic scale. The task is highly challenging due to the enormous amount of alternative candidate configurations. Previous research has only considered this prediction problem starting from 3D information. When starting from sequence alone, only methods that predict the bonding state of selected residues are available. The sole exception consists of pattern-based approaches, which rely on very specific motifs and cannot be applied to discover truly novel sites. We develop new algorithmic ideas based on structured-output learning for determining transition-metal-binding sites coordinated by cysteines and histidines. The inference step (retrieving the best scoring output) is intractable for general output types (i.e., general graphs). However, under the assumption that no residue can coordinate more than one metal ion, we prove that metal binding has the algebraic structure of a matroid, allowing us to employ a very efficient greedy algorithm. We test our predictor in a highly stringent setting where the training set consists of protein chains belonging to SCOP folds different from the ones used for accuracy estimation. In this setting, our predictor achieves 56 percent precision and 60 percent recall in the identification of ligand-ion bonds.
Andrea Passerini, Marco Lippi 0001, Paolo Frasconi
IEEE ACM Trans. Comput. Biol. Bioinform.3
2011 Relational Learning for Spatial Relation Extraction from Natural Language
Parisa Kordjamshidi, Paolo Frasconi, Martijn van Otterlo, Marie-Francine Moens, Luc De Raedt
ILP2
2011 Kernel-Based Logical and Relational Learning with kLog for Hedge Cue Detection
Mathias Verbeke, Paolo Frasconi, Vincent Van Asch, Roser Morante, Walter Daelemans, Luc De Raedt
ILP2
2011 Relational information gain
Marco Lippi 0001, Manfred Jaeger, Paolo Frasconi, Andrea Passerini
Mach. Learn.3
2010 Collective Traffic Forecasting
Marco Lippi 0001, Matteo Bertini, Paolo Frasconi
ECML/PKDD (2)3
2010 Fast learning of relational kernels
Niels Landwehr, Andrea Passerini, Luc De Raedt, Paolo Frasconi
Mach. Learn.4
2009 Prediction of protein beta-residue contacts by Markov logic networks with grounding-specific weights
abstract
MOTIVATION: Accurate prediction of contacts between beta-strand residues can significantly contribute towards ab initio prediction of the 3D structure of many proteins. Contacts in the same protein are highly interdependent. Therefore, significant improvements can be expected by applying statistical relational learners that overcome the usual machine learning assumption that examples are independent and identically distributed. Furthermore, the dependencies among beta-residue contacts are subject to strong regularities, many of which are known a priori. In this article, we take advantage of Markov logic, a statistical relational learning framework that is able to capture dependencies between contacts, and constrain the solution according to domain knowledge expressed by means of weighted rules in a logical language. RESULTS: We introduce a novel hybrid architecture based on neural and Markov logic networks with grounding-specific weights. On a non-redundant dataset, our method achieves 44.9% F(1) measure, with 47.3% precision and 42.7% recall, which is significantly better (P < 0.01) than previously reported performance obtained by 2D recursive neural networks. Our approach also significantly improves the number of chains for which beta-strands are nearly perfectly paired (36% of the chains are predicted with F(1) >or= 70% on coarse map). It also outperforms more general contact predictors on recent CASP 2008 targets.
Marco Lippi 0001, Paolo Frasconi
Bioinform.2
2008 Feature Discovery with Type Extension Trees
Paolo Frasconi, Manfred Jaeger, Andrea Passerini
ILP1
2008 Predicting the Geometry of Metal Binding Sites from Protein Sequence
abstract
Metal binding is important for the structural and functional characterization of proteins. Previous prediction efforts have only focused on bonding state, i.e. deciding which protein residues act as metal ligands in some binding site. Identifying the geometry of metal-binding sites, i.e. deciding which residues are jointly involved in the coordination of a metal ion is a new prediction problem that has been never attempted before from protein sequence alone. In this paper, we formulate it in the framework of learning with structured outputs. Our solution relies on the fact that, from a graph theoretical perspective, metal binding has the algebraic properties of a matroid, enabling the application of greedy algorithms for learning structured outputs. On a data set of 199 non-redundant metalloproteins, we obtained precision/recall levels of 75\%/46\% correct ligand-ion assignments, which improves to 88\%/88\% in the setting where the metal binding state is known.
Paolo Frasconi, Andrea Passerini
NIPS1
2008 MetalDetector: a web server for predicting metal-binding sites and disulfide bridges in proteins from sequence
abstract
UNLABELLED: The web server MetalDetector classifies histidine residues in proteins into one of two states (free or metal bound) and cysteines into one of three states (free, metal bound or disulfide bridged). A decision tree integrates predictions from two previously developed methods (DISULFIND and Metal Ligand Predictor). Cross-validated performance assessment indicates that our server predicts disulfide bonding state at 88.6% precision and 85.1% recall, while it identifies cysteines and histidines in transition metal-binding sites at 79.9% precision and 76.8% recall, and at 60.8% precision and 40.7% recall, respectively. AVAILABILITY: Freely available at http://metaldetector.dsi.unifi.it. SUPPLEMENTARY INFORMATION: Details and data can be found at http://metaldetector.dsi.unifi.it/help.php.
Marco Lippi 0001, Andrea Passerini, Marco Punta, Burkhard Rost, Paolo Frasconi
Bioinform.5
2008 A simplified approach to disulfide connectivity prediction from protein sequences
abstract
BACKGROUND: Prediction of disulfide bridges from protein sequences is useful for characterizing structural and functional properties of proteins. Several methods based on different machine learning algorithms have been applied to solve this problem and public domain prediction services exist. These methods are however still potentially subject to significant improvements both in terms of prediction accuracy and overall architectural complexity. RESULTS: We introduce new methods for predicting disulfide bridges from protein sequences. The methods take advantage of two new decomposition kernels for measuring the similarity between protein sequences according to the amino acid environments around cysteines. Disulfide connectivity is predicted in two passes. First, a binary classifier is trained to predict whether a given protein chain has at least one intra-chain disulfide bridge. Second, a multiclass classifier (plemented by 1-nearest neighbor) is trained to predict connectivity patterns. The two passes can be easily cascaded to obtain connectivity prediction from sequence alone. We report an extensive experimental comparison on several data sets that have been previously employed in the literature to assess the accuracy of cysteine bonding state and disulfide connectivity predictors. CONCLUSION: We reach state-of-the-art results on bonding state prediction with a simple method that classifies chains rather than individual residues. The prediction accuracy reached by our connectivity prediction method compares favorably with respect to all but the most complex other approaches. On the other hand, our method does not need any model selection or hyperparameter tuning, a property that makes it less prone to overfitting and prediction accuracy overestimation.
Marc Vincent, Andrea Passerini, Matthieu Labbé, Paolo Frasconi
BMC Bioinform.4
2007 Learning with Kernels and Logical Representations
Paolo Frasconi
ILP1
2007 Classification of small molecules by two- and three-dimensional decomposition kernels
abstract
MOTIVATION: Several kernel-based methods have been recently introduced for the classification of small molecules. Most available kernels on molecules are based on 2D representations obtained from chemical structures, but far less work has focused so far on the definition of effective kernels that can also exploit 3D information. RESULTS: We introduce new ideas for building kernels on small molecules that can effectively use and combine 2D and 3D information. We tested these kernels in conjunction with support vector machines for binary classification on the 60 NCI cancer screening datasets as well as on the NCI HIV data set. Our results show that 3D information leveraged by these kernels can consistently improve prediction accuracy in all datasets. AVAILABILITY: An implementation of the small molecule classifier is available from http://www.dsi.unifi.it/neural/src/3DDK.
Alessio Ceroni, Fabrizio Costa, Paolo Frasconi
Bioinform.3
2007 Predicting zinc binding at the proteome level
abstract
BACKGROUND: Metalloproteins are proteins capable of binding one or more metal ions, which may be required for their biological function, for regulation of their activities or for structural purposes. Metal-binding properties remain difficult to predict as well as to investigate experimentally at the whole-proteome level. Consequently, the current knowledge about metalloproteins is only partial. RESULTS: The present work reports on the development of a machine learning method for the prediction of the zinc-binding state of pairs of nearby amino-acids, using predictors based on support vector machines. The predictor was trained using chains containing zinc-binding sites and non-metalloproteins in order to provide positive and negative examples. Results based on strong non-redundancy tests prove that (1) zinc-binding residues can be predicted and (2) modelling the correlation between the binding state of nearby residues significantly improves performance. The trained predictor was then applied to the human proteome. The present results were in good agreement with the outcomes of previous, highly manually curated, efforts for the identification of human zinc-binding proteins. Some unprecedented zinc-binding sites could be identified, and were further validated through structural modelling. The software implementing the predictor is freely available at: http://zincfinder.dsi.unifi.it CONCLUSION: The proposed approach constitutes a highly automated tool for the identification of metalloproteins, which provides results of comparable quality with respect to highly manually refined predictions. The ability to model correlations between pairwise residues allows it to obtain a significant improvement over standard 1D based approaches. In addition, the method permits the identification of unprecedented metal sites, providing important hints for the work of experimentalists.
Andrea Passerini, Claudia Andreini, Sauro Menchetti, Antonio Rosato, Paolo Frasconi
BMC Bioinform.5
2006 kFOIL: Learning Simple Relational Kernels
Niels Landwehr, Andrea Passerini, Luc De Raedt, Paolo Frasconi
AAAI4
2006 Improving Prediction of Zinc Binding Sites by Modeling the Linkage Between Residues Close in Sequence
Sauro Menchetti, Andrea Passerini, Paolo Frasconi, Claudia Andreini, Antonio Rosato
RECOMB3
2006 Kernels on Prolog Proof Trees: Statistical Learning in the ILP Setting
abstract
We develop kernels for measuring the similarity between relational instances using background knowledge expressed in first-order logic. The method allows us to bridge the gap between traditional inductive logic programming (ILP) representations and statistical approaches to supervised learning. Logic programs are first used to generate proofs of given visitor programs that use predicates declared in the available background knowledge. A kernel is then defined over pairs of proof trees. The method can be used for supervised learning tasks and is suitable for classification as well as regression. We report positive empirical results on Bongard-like and M-of-N problems that are difficult or impossible to solve with traditional ILP techniques, as well as on real bioinformatics and chemoinformatics data sets.
Andrea Passerini, Paolo Frasconi, Luc De Raedt
J. Mach. Learn. Res.2
2005 Weighted decomposition kernels
abstract
We introduce a family of kernels on discrete data structures within the general class of decomposition kernels. A weighted decomposition kernel (WDK) is computed by dividing objects into substructures indexed by a selector. Two substructures are then matched if their selectors satisfy an equality predicate, while the importance of the match is determined by a probability kernel on local distributions fitted on the substructures. Under reasonable assumptions, a WDK can be computed efficiently and can avoid combinatorial explosion of the feature space. We report experimental evidence that the proposed kernel is highly competitive with respect to more complex state-of-the-art methods on a set of problems in bioinformatics.
Sauro Menchetti, Fabrizio Costa, Paolo Frasconi
ICML3
2005 Kernels on Prolog Ground Terms
Andrea Passerini, Paolo Frasconi
IJCAI2
2005 Learning protein secondary structure from sequential and relational data
Alessio Ceroni, Paolo Frasconi, Gianluca Pollastri
Neural Networks2
2005 Wide coverage natural language processing using kernel methods and neural networks for structured data
Sauro Menchetti, Fabrizio Costa, Paolo Frasconi, Massimiliano Pontil
Pattern Recognit. Lett.3
2005 Ambiguity resolution analysis in incremental parsing of natural language
abstract
Incremental parsing gains its importance in natural language processing and psycholinguistics because of its cognitive plausibility. Modeling the associated cognitive data structures, and their dynamics, can lead to a better understanding of the human parser. In earlier work, we have introduced a recursive neural network (RNN) capable of performing syntactic ambiguity resolution in incremental parsing. In this paper, we report a systematic analysis of the behavior of the network that allows us to gain important insights about the kind of information that is exploited to resolve different forms of ambiguity. In attachment ambiguities, in which a new phrase can be attached at more than one point in the syntactic left context, we found that learning from examples allows us to predict the location of the attachment point with high accuracy, while the discrimination amongst alternative syntactic structures with the same attachment point is slightly better than making a decision purely based on frequencies. We also introduce several new ideas to enhance the architectural design, obtaining significant improvements of prediction accuracy, up to 25% error reduction on the same dataset used in previous work. Finally, we report large scale experiments on the entire Wall Street Journal section of the Penn Treebank. The best prediction accuracy of the model on this large dataset is 87.6%, a relative error reduction larger than 50% compared to previous results.
Fabrizio Costa, Paolo Frasconi, Vincenzo Lombardo, Patrick Sturt, Giovanni Soda
IEEE Trans. Neural Networks2
2004 On the role of long-range dependencies in learning protein secondary structure
abstract
Accuracy of protein secondary structure predictors has been slowly growing during the last decade. Although it is clear that a relatively large fraction of current errors is due to long-range interactions, current predictors are not able to exploit such information. We present a solution based on a generalized bidirectional neural network that learns from sequences and associated interaction graphs to improve secondary structure prediction.
Alessio Ceroni, Paolo Frasconi
IJCNN2
2004 Disulfide connectivity prediction using recursive neural networks and evolutionary information
abstract
MOTIVATION: We focus on the prediction of disulfide bridges in proteins starting from their amino acid sequence and from the knowledge of the disulfide bonding state of each cysteine. The location of disulfide bridges is a structural feature that conveys important information about the protein main chain conformation and can therefore help towards the solution of the folding problem. Existing approaches based on weighted graph matching algorithms do not take advantage of evolutionary information. Recursive neural networks (RNN), on the other hand, can handle in a natural way complex data structures such as graphs whose vertices are labeled by real vectors, allowing us to incorporate multiple alignment profiles in the graphical representation of disulfide connectivity patterns. RESULTS: The core of the method is the use of machine learning tools to rank alternative disulfide connectivity patterns. We develop an ad-hoc RNN architecture for scoring labeled undirected graphs that represent connectivity patterns. In order to compare our algorithm with previous methods, we report experimental results on the SWISS-PROT 39 dataset. We find that using multiple alignment profiles allows us to obtain significant prediction accuracy improvements, clearly demonstrating the important role played by evolutionary information. AVAILABILITY: The Web interface of the predictor is available at http://neural.dsi.unifi.it/cysteines
Alessandro Vullo, Paolo Frasconi
Bioinform.2
2004 New results on error correcting output codes of kernel machines
abstract
We study the problem of multiclass classification within the framework of error correcting output codes (ECOC) using margin-based binary classifiers. Specifically, we address two important open problems in this context: decoding and model selection. The decoding problem concerns how to map the outputs of the classifiers into class codewords. In this paper we introduce a new decoding function that combines the margins through an estimate of their class conditional probabilities. Concerning model selection, we present new theoretical results bounding the leave-one-out (LOO) error of ECOC of kernel machines, which can be used to tune kernel hyperparameters. We report experiments using support vector machines as the base binary classifiers, showing the advantage of the proposed decoding function over other functions of the margin commonly used in practice. Moreover, our empirical evaluations on model selection indicate that the bound leads to good estimates of kernel parameters.
Andrea Passerini, Massimiliano Pontil, Paolo Frasconi
IEEE Trans. Neural Networks3
2004 Guest editorial: Machine learning for the Internet
abstract
The World Wide Web has been at the center of a revolution in how algorithms are designed with massive amounts of data in mind. The essence of this revo- lution is conceptually very simple: real-world massive data sets are, more often than not, highly structured and regular. Regularities can be used in two com- plementary ways. First, systematic regularities within massive data sets can be used to craft algorithms that are potentially suboptimal in the worst-case, but highly effective for expected cases. Second, nonsystematic regularities—those that are too subtle to be encoded within an algorithm—can be discovered by automated methods so that the solutions are actually determined by the un- derlying data. In both cases, the existence of enormous problem instances that arise from a highly regular source is key to building more effective methods.
Gary William Flake, Paolo Frasconi, C. Lee Giles, Marco Maggini
ACM Trans. Internet Techn.2
2004 Guest editorial: Machine learning for the Internet
abstract
The Internet and the Web are continuously evolving giving rise to a rich and extremely dynamic environment where an increasing number of users require and expect new and more sophisticated services. Because of this, a field called “ Web Intelligence” is starting to receive interest from the Artificial Intelligence community. The Web and Internet pose new challenges to AI algorithms, which have been successfully applied in many other fields, at the same time stimu- lating the development of new techniques. In particular, as pointed out in the introduction to the first part of this special issue (Vol. 4, no. 2, May 2004), ma- chine learning methods have been extensively studied and have been applied to create intelligent systems that are actively involved with the Internet and Web.
Gary William Flake, Paolo Frasconi, C. Lee Giles, Marco Maggini
ACM Trans. Internet Techn.2
2003 Towards Incremental Parsing of Natural Language Using Recursive Neural Networks
Fabrizio Costa, Paolo Frasconi, Vincenzo Lombardo, Giovanni Soda
Appl. Intell.2
2003 Hidden Tree Markov Models for Document Image Classification
abstract
Classification is an important problem in image document processing and is often a preliminary step toward recognition, understanding, and information extraction. In this paper, the problem is formulated in the framework of concept learning and each category corresponds to the set of image documents with similar physical structure. We propose a solution based on two algorithmic ideas. First, we obtain a structured representation of images based on labeled XY-trees (this representation informs the learner about important relationships between image subconstituents). Second, we propose a probabilistic architecture that extends hidden Markov models for learning probability distributions defined on spaces of labeled trees. Finally, a successful application of this method to the categorization of commercial invoices is presented.
Michelangelo Diligenti, Paolo Frasconi, Marco Gori
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Combining flat and structured representations for fingerprint classification with recursive neural networks and support vector machines
Gian Luca Marcialis, Massimiliano Pontil, Paolo Frasconi, Fabio Roli
Pattern Recognit.4
2002 Enhancing First-Pass Attachment Prediction
Fabrizio Costa, Paolo Frasconi, Vincenzo Lombardo, Patrick Sturt, Giovanni Soda
ECAI2
2002 From Margins to Probabilities in Multiclass Learning Problems
Andrea Passerini, Massimiliano Pontil, Paolo Frasconi
ECAI3
2002 Prediction of Protein Topologies Using Generalized IOHMMS and RNNs
abstract
We develop and test new machine learning methods for the predic- tion of topological representations of protein structures in the form of coarse- or (cid:12)ne-grained contact or distance maps that are transla- tion and rotation invariant. The methods are based on generalized input-output hidden Markov models (GIOHMMs) and generalized recursive neural networks (GRNNs). The methods are used to pre- dict topology directly in the (cid:12)ne-grained case and, in the coarse- grained case, indirectly by (cid:12)rst learning how to score candidate graphs and then using the scoring function to search the space of possible con(cid:12)gurations. Computer simulations show that the pre- dictors achieve state-of-the-art performance. 1 Introduction: Protein Topology Prediction Predicting the 3D structure of protein chains from the linear sequence of amino acids is a fundamental open problem in computational molecular biology [1]. Any approach to the problem must deal with the basic fact that protein structures are translation and rotation invariant. To address this invariance, we have proposed a machine learning approach to protein structure prediction [4] based on the predic- tion of topological representations of proteins, in the form of contact or distance maps. The contact or distance map is a 2D representation of neighborhood rela- tionships consisting of an adjacency matrix at some distance cuto(cid:11) (typically in the range of 6 to 12 (cid:23)A), or a matrix of pairwise Euclidean distances. Fine-grained maps are derived at the amino acid or even atomic level. Coarse maps are obtained by looking at secondary structure elements, such as helices, and the distance between their centers of gravity or, as in the simulations below, the minimal distances be- tween their C(cid:11) atoms. Reasonable methods for reconstructing 3D coordinates from contact/distance maps have been developed in the NMR literature and elsewhere
Gianluca Pollastri, Pierre Baldi, Alessandro Vullo, Paolo Frasconi
NIPS4
2002 Hidden Markov Models for Text Categorization in Multi-Page Documents
Paolo Frasconi, Giovanni Soda, Alessandro Vullo
J. Intell. Inf. Syst.1
2001 Guest Editors' Introduction: Special Section on Connectionist Models for Learning in Structured Domains
abstract
Guest Editors' Introduction to the Special Section on Connectionist Models for Learning in Structured Domains
Paolo Frasconi, Marco Gori, Alessandro Sperduti
IEEE Trans. Knowl. Data Eng.1
2000 Learning Efficiently with Neural Networks: A Theoretical Comparison between Structured and Flat Representations
Marco Gori, Paolo Frasconi, Alessandro Sperduti
ECAI2
2000 Learning incremental syntactic structures with recursive neural networks
abstract
We develop novel algorithmic ideas for building a natural language parser grounded upon the hypothesis of incrementality, which is widely supported by experimental data as a model of human parsing. Our proposal relies on a machine learning technique for predicting the correctness of partial syntactic structures that are built during the parsing process. A recursive neural network architecture is employed for computing predictions after a training phase on examples drawn from a corpus of parsed sentences, the Penn Treebank. Our results indicate the viability of the approach and lay out the premises for a novel generation of algorithms for natural language processing which more closely model human parsing. These algorithms may prove very useful in the development of efficient parsers and have an immediate application in the construction of semiautomatic annotation tools.
Fabrizio Costa, Paolo Frasconi, Vincenzo Lombardo, Giovanni Soda
KES2
2000 Competitive radial basis functions training for phone classification
Piero Cosi, Paolo Frasconi, Marco Gori, Luca Lastrucci, Giovanni Soda
Neurocomputing2
1999 A topological transformation for hidden recursive modelsarchitecture networks
Fabrizio Costa, Paolo Frasconi, Giovanni Soda
ESANN2
1999 Exploiting the past and the future in protein secondary structure prediction
abstract
MOTIVATION: Predicting the secondary structure of a protein (alpha-helix, beta-sheet, coil) is an important step towards elucidating its three-dimensional structure, as well as its function. Presently, the best predictors are based on machine learning approaches, in particular neural network architectures with a fixed, and relatively short, input window of amino acids, centered at the prediction site. Although a fixed small window avoids overfitting problems, it does not permit capturing variable long-rang information. RESULTS: We introduce a family of novel architectures which can learn to make predictions based on variable ranges of dependencies. These architectures extend recurrent neural networks, introducing non-causal bidirectional dynamics to capture both upstream and downstream information. The prediction algorithm is completed by the use of mixtures of estimators that leverage evolutionary information, expressed in terms of multiple alignments, both at the input and output levels. While our system currently achieves an overall performance close to 76% correct prediction--at least comparable to the best existing systems--the main emphasis here is on the development of new algorithmic ideas. AVAILABILITY: The executable program for predicting protein secondary structure is available from the authors free of charge. CONTACT: [email protected], [email protected], [email protected], [email protected].
Pierre Baldi, Søren Brunak, Paolo Frasconi, Giovanni Soda, Gianluca Pollastri
Bioinform.3
1999 Data Categorization Using Decision Trellises
abstract
We introduce a probabilistic graphical model for supervised learning on databases with categorical attributes. The proposed belief network contains hidden variables that play a role similar to nodes in decision trees and each of their states either corresponds to a class label or to a single attribute test. As a major difference with respect to decision trees, the selection of the attribute to be tested is probabilistic. Thus, the model can be used to assess the probability that a tuple belongs to some class, given the predictive attributes. Unfolding the network along the hidden states dimension yields a trellis structure having a signal flow similar to second order connectionist networks. The network encodes context specific probabilistic independencies to reduce parametric complexity. We present a custom tailored inference algorithm and derive a learning procedure based on the expectation-maximization algorithm. We propose decision trellises as an alternative to decision trees in the context of tuple categorization in databases, which is an important step for building data mining systems. Preliminary experiments on standard machine learning databases are reported, comparing the classification accuracy of decision trellises and decision trees induced by C4.5. In particular, we show that the proposed model can offer significant advantages for sparse databases in which many predictive attributes are missing.
Paolo Frasconi, Marco Gori, Giovanni Soda
IEEE Trans. Knowl. Data Eng.1
1998 A general framework for adaptive processing of data structures
abstract
A structured organization of information is typically required by symbolic processing. On the other hand, most connectionist models assume that data are organized according to relatively poor structures, like arrays or sequences. The framework described in this paper is an attempt to unify adaptive models like artificial neural nets and belief nets for the problem of processing structured information. In particular, relations between data variables are expressed by directed acyclic graphs, where both numerical and categorical values coexist. The general framework proposed in this paper can be regarded as an extension of both recurrent neural networks and hidden Markov models to the case of acyclic graphs. In particular we study the supervised learning problem as the problem of learning transductions from an input structured space to an output structured space, where transductions are assumed to admit a recursive hidden statespace representation. We introduce a graphical formalism for representing this class of adaptive transductions by means of recursive networks, i.e., cyclic graphs where nodes are labeled by variables and edges are labeled by generalized delay elements. This representation makes it possible to incorporate the symbolic and subsymbolic nature of data. Structures are processed by unfolding the recursive network into an acyclic graph called encoding network. In so doing, inference and learning algorithms can be easily inherited from the corresponding algorithms for artificial neural networks or probabilistic graphical model.
Paolo Frasconi, Marco Gori, Alessandro Sperduti
IEEE Trans. Neural Networks1
1997 On the Efficient Classification of Data Structures by Neural Networks
Paolo Frasconi, Marco Gori, Alessandro Sperduti
IJCAI1
1997 Links between LVQ and Backpropagation
Paolo Frasconi, Marco Gori, Giovanni Soda
Pattern Recognit. Lett.1
1996 Representation of Finite State Automata in Recurrent Radial Basis Function Networks
Paolo Frasconi, Marco Gori, Marco Maggini, Giovanni Soda
Mach. Learn.1
1996 Input-output HMMs for sequence processing
abstract
We consider problems of sequence processing and propose a solution based on a discrete-state model in order to represent past context. We introduce a recurrent connectionist architecture having a modular structure that associates a subnetwork to each state. The model has a statistical interpretation we call input-output hidden Markov model (IOHMM). It can be trained by the estimation-maximization (EM) or generalized EM (GEM) algorithms, considering state trajectories as missing data, which decouples temporal credit assignment and actual parameter estimation. The model presents similarities to hidden Markov models (HMMs), but allows us to map input sequences to output sequences, using the same processing style as recurrent neural networks. IOHMMs are trained using a more discriminant learning paradigm than HMMs, while potentially taking advantage of the EM algorithm. We demonstrate that IOHMMs are well suited for solving grammatical inference problems on a benchmark problem. Experimental results are presented for the seven Tomita grammars, showing that these adaptive models can attain excellent generalization.
Yoshua Bengio, Paolo Frasconi
IEEE Trans. Neural Networks2
1996 Computational capabilities of local-feedback recurrent networks acting as finite-state machines
abstract
In this paper we explore the expressive power of recurrent networks with local feedback connections for symbolic data streams. We rely on the analysis of the maximal set of strings that can be shattered by the concept class associated to these networks (i.e. strings that can be arbitrarily classified as positive or negative), and find that their expressive power is inherently limited, since there are sets of strings that cannot be shattered, regardless of the number of hidden units. Although the analysis holds for networks with hard threshold units, we claim that the incremental computational capabilities gained when using sigmoidal units are severely paid in terms of robustness of the corresponding representation.
Paolo Frasconi, Marco Gori
IEEE Trans. Neural Networks1
1995 Diffusion of Context and Credit Information in Markovian Models
abstract
This paper studies the problem of ergodicity of transition probability matrices in Markovian models, such as hidden Markov models (HMMs), and how it makes very difficult the task of learning to represent long-term context for sequential data. This phenomenon hurts the forward propagation of long-term context information, as well as learning a hidden state representation to represent long-term context, which depends on propagating credit information backwards in time. Using results from Markov chain theory, we show that this problem of diffusion of context and credit is reduced when the transition probabilities approach 0 or 1, i.e., the transition probability matrices are sparse and the model essentially deterministic. The results found in this paper apply to learning approaches based on continuous optimization, such as gradient descent and the Baum-Welch algorithm.
Yoshua Bengio, Paolo Frasconi
J. Artif. Intell. Res.2
1995 Recurrent neural networks and prior knowledge for sequence processing: a constrained nondeterministic approach
Paolo Frasconi, Marco Gori, Giovanni Soda
Knowl. Based Syst.1
1995 Unified Integration of Explicit Knowledge and Learning by Example in Recurrent Networks
abstract
Proposes a novel unified approach for integrating explicit knowledge and learning by example in recurrent networks. The explicit knowledge is represented by automaton rules, which are directly injected into the connections of a network. This can be accomplished by using a technique based on linear programming, instead of learning from random initial weights. Learning is conceived as a refinement process and is mainly responsible for uncertain information management. We present preliminary results for problems of automatic speech recognition.>
Paolo Frasconi, Marco Gori, Marco Maggini, Giovanni Soda
IEEE Trans. Knowl. Data Eng.1
1995 Learning in multilayered networks used as autoassociators
abstract
Gradient descent learning algorithms may get stuck in local minima, thus making the learning suboptimal. In this paper, we focus attention on multilayered networks used as autoassociators and show some relationships with classical linear autoassociators. In addition, by using the theoretical framework of our previous research, we derive a condition which is met at the end of the learning process and show that this condition has a very intriguing geometrical meaning in the pattern space.
Monica Bianchini, Paolo Frasconi, Marco Gori
IEEE Trans. Neural Networks2
1995 Learning without local minima in radial basis function networks
abstract
Learning from examples plays a central role in artificial neural networks. The success of many learning schemes is not guaranteed, however, since algorithms like backpropagation may get stuck in local minima, thus providing suboptimal solutions. For feedforward networks, optimal learning can be achieved provided that certain conditions on the network and the learning environment are met. This principle is investigated for the case of networks using radial basis functions (RBF). It is assumed that the patterns of the learning environment are separable by hyperspheres. In that case, we prove that the attached cost function is local minima free with respect to all the weights. This provides us with some theoretical foundations for a massive application of RBF in pattern recognition.
Monica Bianchini, Paolo Frasconi, Marco Gori
IEEE Trans. Neural Networks2
1994 An EM approach to grammatical inference: input/output HMMs
abstract
Proposes a modular recurrent connectionist architecture for adaptive temporal processing. The model is given, a probabilistic interpretation and is trained using the estimation-maximisation (EM) algorithm. This model can also be seen as an input/output hidden Markov model. The focus of this paper is on sequence classification tasks. The authors demonstrate that EM supervised learning is well suited for solving grammatical inference problems. Experimental benchmark results are presented for the seven Tomita grammars, showing that these adaptive models can, attain excellent generalization.
Paolo Frasconi, Yoshua Bengio
ICPR (2)1
1994 An Input Output HMM Architecture
abstract
We introduce a recurrent architecture having a modular structure and we formulate a training procedure based on the EM algorithm. The resulting model has similarities to hidden Markov models, but supports recurrent networks processing style and allows to exploit the supervised learning paradigm while using maximum likelihood estimation.
Yoshua Bengio, Paolo Frasconi
NIPS2
1994 Diffusion of Credit in Markovian Models
abstract
This paper studies the problem of diffusion in Markovian models, such as hidden Markov models (HMMs) and how it makes very difficult the task of learning of long-term dependencies in sequences. Using results from Markov chain theory, we show that the problem of diffusion is reduced if the transition probabilities approach 0 or 1. Under this condition, standard HMMs have very limited modeling capabilities, but input/output HMMs can still perform interesting computations.
Yoshua Bengio, Paolo Frasconi
NIPS2
1994 Learning long-term dependencies with gradient descent is difficult
abstract
Recurrent neural networks can be used to map input sequences to output sequences, such as for recognition, production or prediction problems. However, practical difficulties have been reported in training recurrent neural networks to perform tasks in which the temporal contingencies present in the input/output sequences span long intervals. We show why gradient based learning algorithms face an increasingly difficult problem as the duration of the dependencies to be captured increases. These results expose a trade-off between efficient learning by gradient descent and latching on information for long periods. Based on an understanding of this problem, alternatives to standard gradient descent are considered.
Yoshua Bengio, Patrice Y. Simard, Paolo Frasconi
IEEE Trans. Neural Networks3
1993 Credit Assignment through Time: Alternatives to Backpropagation
Yoshua Bengio, Paolo Frasconi
NIPS2
1992 Phonetic recognition experiments with recurrent neural networks
Piero Cosi, Paolo Frasconi, Marco Gori, N. Griggio
ICSLP2
1992 Local Feedback Multilayered Networks
abstract
In this paper, we investigate the capabilities of local feedback multilayered networks, a particular class of recurrent networks, in which feedback connections are only allowed from neurons to themselves. In this class, learning can be accomplished by an algorithm that is local in both space and time. We describe the limits and properties of these networks and give some insights on their use for solving practical problems.
Paolo Frasconi, Marco Gori, Giovanni Soda
Neural Comput.1