VLDB 2026 Research / reviewers in the wild / expert
Marco Gori
dblp:g/MarcoGori
· DBLP profile ↗
199ranked-venue papers
31as first author
33since 2021 · last 2026
0000-0001-6337-5430ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 180 · 27 first-author · 32 since 2021Databases, data management, data science and information retrieval · 33 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 4 first-author · 6 since 2021Theory of computation · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeepProofLog: Efficient Proving in Deep Stochastic Logic ProgramsabstractNeurosymbolic (NeSy) AI combines neural architectures and symbolic reasoning to improve accuracy, interpretability, and generalization. While logic inference on top of subsymbolic modules has been shown to effectively guarantee these properties, this often comes at the cost of reduced scalability, which can severely limit the usability of NeSy models. This paper introduces DeepProofLog (DPrL), a novel NeSy system based on stochastic logic programs, which addresses the scalability limitations of previous methods. DPrL parameterizes all derivation steps with neural networks, allowing efficient neural guidance over the proving system. Additionally, we establish a formal mapping between the resolution process of our deep stochastic logic programs and Markov Decision Processes, enabling the application of dynamic programming and reinforcement learning techniques for efficient inference and learning. This theoretical connection improves scalability for complex proof spaces and large knowledge bases. Our experiments on standard NeSy benchmarks and knowledge graph reasoning tasks demonstrate that DPrL outperforms existing state-of-the-art NeSy systems, advancing scalability to larger and more complex settings than previously possible. Ying Jiao, Rodrigo Castellano Ontiveros, Luc De Raedt, Marco Gori, Francesco Giannini, Michelangelo Diligenti, Giuseppe Marra |
AAAI | 4 |
| 2026 | A systematic literature review of spatio-temporal graph neural network models for time series forecasting and classificationabstractIn recent years, spatio-temporal graph neural networks (GNNs) have attracted considerable interest in the field of time series analysis, due to their ability to capture, at once, dependencies among variables and across time points. The objective of this systematic literature review is hence to provide a comprehensive overview of the various modeling approaches and application domains of GNNs for time series classification and forecasting. A database search was conducted, and 366 papers were selected for a detailed examination of the current state-of-the-art in the field. This examination is intended to offer to the reader a comprehensive review of proposed models, links to related source code, available datasets, benchmark models, and fitting results. All this information is hoped to assist researchers in their studies. To the best of our knowledge, this is the first and broadest systematic literature review presenting a detailed comparison of results from current spatio-temporal GNN models applied to different domains. In its final part, this review discusses current limitations and challenges in the application of spatio-temporal GNNs, such as comparability, reproducibility, explainability, poor information capacity, and scalability. This paper is complemented by a GitHub repository at https://github.com/FlaGer99/SLR-Spatio-Temporal-GNN.git providing additional interactive tools to further explore the presented findings. Flavio Corradini, Flavio Gerosa, Marco Gori, Carlo Lucheroni, Marco Piangerelli, Martina Zannotti |
Neural Networks | 3 |
| 2026 | State-space modeling in long sequence processing: a survey on recurrence in the transformer eraabstractEffectively learning from sequential data is a longstanding goal of Artificial Intelligence, especially in the case of long sequences. From the dawn of Machine Learning, several researchers have pursued algorithms and architectures capable of processing sequences of patterns, retaining information about past inputs while still leveraging future data, without losing precious long-term dependencies and correlations. While such an ultimate goal is inspired by the human hallmark of continuous real-time processing of sensory information, several solutions have simplified the learning paradigm by artificially limiting the processed context or dealing with sequences of limited length, given in advance. These solutions were further emphasized by the ubiquity of Transformers, which initially overshadowed the role of Recurrent Neural Nets. However, recurrent networks are currently experiencing a strong recent revival due to the growing popularity of (deep) State-Space models and novel instances of large-context Transformers, which are both based on recurrent computations that aim to go beyond several limits of currently ubiquitous technologies. The fast development of Large Language Models has renewed the interest in efficient solutions to process data over time. This survey provides an in-depth summary of the latest approaches that are based on recurrent models for sequential data processing. A complete taxonomy of recent trends in architectural and algorithmic solutions is reported and discussed, guiding researchers in this appealing research field. The emerging picture suggests that there is room for exploring novel routes, constituted by learning algorithms that depart from the standard Backpropagation Through Time, towards a more realistic scenario where patterns are effectively processed online, leveraging local-forward computations, and opening new directions for research on this topic. Matteo Tiezzi, Michele Casoni, Alessandro Betti, Marco Gori, Stefano Melacci |
Neural Networks | 4 |
| 2025 | Stability of State and Costate Dynamics in Continuous Time Recurrent Neural NetworksabstractThe notion of stability plays a crucial role in ensuring the safe development of a model in a lifelong learning context.This paper investigates the fundamental aspects of stability in a class of continuous-time recurrent neural networks which include both state and costate variables.The latter are directly inherited from optimal control theory, and they act as adjoint variables closely related to gradient terms.Stability is investigated both in terms of state and of costate dynamics, showing the key conditions that must be satisfied to produce bounded dynamics in the forward and learning stages.* This work was Alessandro Betti, Marco Gori, Stefano Melacci |
ESANN | 2 |
| 2025 | Perpetual Generation: Online Learning of Linear State-Space Models from a Single Stream
Michele Casoni, Tommaso Guidi, Stefano Melacci, Alessandro Betti, Marco Gori |
ICANN (1) | 5 |
| 2025 | Grounding Methods for Neural-Symbolic AIabstractA large class of Neural-Symbolic (NeSy) methods employs a machine learner to process the input entities, while relying on a reasoner based on First-Order Logic to represent and process more complex relationships among the entities. A fundamental role for these methods is played by the process of logic grounding, which determines the relevant substitutions for the logic rules using a (sub)set of entities. Some NeSy methods use an exhaustive derivation of all possible substitutions, preserving the full expressive power of the logic knowledge, but leading to a combinatorial explosion of the number of ground formulas to consider and, therefore, strongly limiting their scalability. Other methods rely on heuristic-based selective derivations, which are generally more computationally efficient, but lack a justification and provide no guarantees of preserving the information provided to and returned by the reasoner. Taking inspiration from multi-hop symbolic reasoning, this paper proposes a parametrized family of grounding methods generalizing classic Backward Chaining. Different selections within this family allow to obtain commonly employed grounding methods as special cases, and to control the trade-off between expressiveness and scalability of the reasoner. The experimental results show that the selection of the grounding criterion is often as important as the NeSy method itself. Rodrigo Castellano Ontiveros, Francesco Giannini, Marco Gori, Giuseppe Marra, Michelangelo Diligenti |
IJCAI | 3 |
| 2025 | Generative System Dynamics in Recurrent Neural NetworksabstractIn this study, we investigate the continuous time dynamics of Recurrent Neural Networks (RNNs), focusing on systems with nonlinear activation functions. The objective of this work is to identify conditions under which RNNs exhibit perpetual oscillatory behavior, without converging to static fixed points. We establish that skew-symmetric weight matrices are fundamental to enable stable limit cycles in both linear and nonlinear configurations. We further demonstrate that hyperbolic tangent-like activation functions (odd, bounded, and continuous) preserve these oscillatory dynamics by ensuring motion invariants in state space. Numerical simulations showcase how nonlinear activation functions not only maintain limit cycles, but also enhance the numerical stability of the system integration process, mitigating those instabilities that are commonly associated with the forward Euler method. The experimental results of this analysis highlight practical considerations for designing neural architectures capable of capturing complex temporal dependencies, i.e., strategies for enhancing memorization skills in recurrent models. Michele Casoni, Tommaso Guidi, Alessandro Betti, Stefano Melacci, Marco Gori |
IJCNN | 5 |
| 2025 | Position Paper: Collectionless Artificial IntelligenceabstractStoring and handling huge data collections has become the fundamental player in the progress of Machine Learning and of its spectacular results. However, learning from such data collections introduces risks related to data centralization, privacy, energy efficiency, limited customizability, and control. This paper sustains the position that the time has come for thinking of new learning protocols where machines conquer cognitive skills by online learning from potentially lifelong streams of sensory data, without the privilege of recording the temporal stream. The perspective of what we refer to as "Collectionless AI" pushes towards interactions with the environment, including humans and other artificial agents, to favor dynamic adaptation, customizability, control. At each time instant, data acquired from the environment is only processed with the purpose of contributing to update the current agent-internal representation of the environment, promoting the development of self-organized memorization skills. The goal of this paper is not to introduce new algorithms, but to present an extreme perspective which recovers largely known notions out of the current mainstream, with the goal of stimulating the development of new foundations on computational processes of learning and reasoning. This might open the doors to a truly orthogonal competitive track on AI technologies that avoid data accumulation by design, thus offering a framework which is better suited concerning privacy issues, control and customizability. Pushing towards massively distributed computation, the collectionless approach to AI might reduce the concentration of power in companies and governments, better facing geopolitical issues. Marco Gori, Stefano Melacci |
IJCNN | 1 |
| 2025 | Position Paper: A new Perspective on Online Continual LearningabstractWe consider the setting of online/streaming continual learning. In particular, we analyze the implications of the consolidated continual learning evaluation setting in an online scenario. As the data stream potentially extends infinitely, the evaluation set also grows boundlessly, necessitating a redefinition of model evaluation. In the streaming learning literature, the model is usually evaluated with the test-then-train method that does not require a separate evaluation set at all, or with continuous reevaluation that considers partially delayed labels [1]. In this work, we propose to model the data-generating process as random walks on concept graphs, i.e. graphs that define the possible transitions between the concepts to learn. This offers a precise formalization of the problem of learning on (possibly) infinitely long data streams with concept drifts and recurring concepts. In such a setting, we also discuss alternative possibilities that blend the continual learning setting with the traditional online learning one. We propose a framework that generalizes the commonly considered online continual learning scenario (i.e. corresponds to a specific topology of the concept graph). Finally, we discuss how different topologies of data streams, derived from our framework, could inspire new research directions in continual learning. Nicolò Navarin, Alessandro Betti, Marco Gori |
IJCNN | 3 |
| 2025 | Continual learning of conjugated visual representations through higher-order motion flowsabstractLearning with neural networks from a continuous stream of visual information presents several challenges due to the non-i.i.d. nature of the data. However, it also offers novel opportunities to develop representations that are consistent with the information flow. In this paper we investigate the case of unsupervised continual learning of pixel-wise features subject to multiple motion-induced constraints, therefore named motion-conjugated feature representations. Differently from existing approaches, motion is not a given signal (either ground-truth or estimated by external modules), but is the outcome of a progressive and autonomous learning process, occurring at various levels of the feature hierarchy. Multiple motion flows are estimated with neural networks and characterized by different levels of abstractions, spanning from traditional optical flow to other latent signals originating from higher-level features, hence called higher-order motions. Continuously learning to develop consistent multi-order flows and representations is prone to trivial solutions, which we counteract by introducing a self-supervised contrastive loss, spatially-aware and based on flow-induced similarity. We assess our model on photorealistic synthetic streams and real-world videos, comparing to pre-trained state-of-the art feature extractors (also based on Transformers) and to recent unsupervised learning models, significantly outperforming these alternatives. Simone Marullo, Matteo Tiezzi, Marco Gori, Stefano Melacci |
Neural Networks | 3 |
| 2024 | Neural Time-Reversed Generalized Riccati EquationabstractOptimal control deals with optimization problems in which variables steer a dynamical system, and its outcome contributes to the objective function. Two classical approaches to solving these problems are Dynamic Programming and the Pontryagin Maximum Principle. In both approaches, Hamiltonian equations offer an interpretation of optimality through auxiliary variables known as costates. However, Hamiltonian equations are rarely used due to their reliance on forward-backward algorithms across the entire temporal domain. This paper introduces a novel neural-based approach to optimal control. Neural networks are employed not only for implementing state dynamics but also for estimating costate variables. The parameters of the latter network are determined at each time step using a newly introduced local policy referred to as the time-reversed generalized Riccati equation. This policy is inspired by a result discussed in the Linear Quadratic (LQ) problem, which we conjecture stabilizes state dynamics. We support this conjecture by discussing experimental results from a range of optimal control case studies. Alessandro Betti, Michele Casoni, Marco Gori, Simone Marullo, Stefano Melacci, Matteo Tiezzi |
AAAI | 3 |
| 2024 | Clue-Instruct: Text-Based Clue Generation for Educational Crossword PuzzlesabstractCrossword puzzles are popular linguistic games often used as tools to engage students in learning. Educational crosswords are characterized by less cryptic and more factual clues that distinguish them from traditional crossword puzzles. Despite there exist several publicly available clue-answer pair databases for traditional crosswords, educational clue-answer pairs datasets are missing. In this article, we propose a methodology to build educational clue generation datasets that can be used to instruct Large Language Models (LLMs). By gathering from Wikipedia pages informative content associated with relevant keywords, we use Large Language Models to automatically generate pedagogical clues related to the given input keyword and its context. With such an approach, we created clue-instruct, a dataset containing 44,075 unique examples with text-keyword pairs associated with three distinct crossword clues. We used clue-instruct to instruct different LLMs to generate educational clues from a given input content and keyword. Both human and automatic evaluations confirmed the quality of the generated clues, thus validating the effectiveness of our approach. Andrea Zugarini, Kamyar Zeinalipour, Surya Sai Kadali, Marco Maggini, Marco Gori, Leonardo Rigutini |
LREC/COLING | 5 |
| 2024 | Agricultural Data Space: the METRIQA Platform and a Case Study in the CODECS projectabstractThis work describes the ongoing design and development of the METRIQA platform, hosting the Italian agrifood data space.Both are key components that the Italian National Research Centre for Agricultural Technologies is putting forward in its activities.We present a high-level description of the platform, which is designed to provide web-like access to digital resources and services following an approach called Web of Agri-Food, to support the digital transformation of the sector in Italy.To show its potential, we also present a real case study demonstrating both the benefits and impacts of the proposed architecture, connecting stakeholders and authorities at different levels. Manlio Bacco, Alexander Kocian, Antonino Crivello, Marco Gori, Giovanna Maria Dimitri, Paolo Barsocchi, Gianluca Brunori, Stefano Chessa |
FedCSIS | 4 |
| 2024 | Nature-Inspired Local PropagationabstractThe spectacular results achieved in machine learning, including the recent advances in generative AI, rely on large data collections. On the opposite, intelligent processes in nature arises without the need for such collections, but simply by on-line processing of the environmental information. In particular, natural learning processes rely on mechanisms where data representation and learning are intertwined in such a way to respect spatiotemporal locality. This paper shows that such a feature arises from a pre-algorithmic view of learning that is inspired by related studies in Theoretical Physics. We show that the algorithmic interpretation of the derived “laws of learning”, which takes the structure of Hamiltonian equations, reduces to Backpropagation when the speed of propagation goes to infinity. This opens the doors to machine learning studies based on full on-line information processing that are based on the replacement of Backpropagation with the proposed spatiotemporal local algorithm. Alessandro Betti, Marco Gori |
NeurIPS | 2 |
| 2024 | Graph Neural Networks for Graph DrawingabstractGraph drawing techniques have been developed in the last few years with the purpose of producing esthetically pleasing node-link layouts. Recently, the employment of differentiable loss functions has paved the road to the massive usage of gradient descent and related optimization algorithms. In this article, we propose a novel framework for the development of Graph Neural Drawers (GNDs), machines that rely on neural computation for constructing efficient and complex maps. GND is Graph Neural Networks (GNNs) whose learning process can be driven by any provided loss function, such as the ones commonly employed in Graph Drawing. Moreover, we prove that this mechanism can be guided by loss functions computed by means of feedforward neural networks, on the basis of supervision hints that express beauty properties, like the minimization of crossing edges. In this context, we show that GNNs can nicely be enriched by positional features to deal also with unlabeled vertexes. We provide a proof-of-concept by constructing a loss function for the edge crossing and provide quantitative and qualitative comparisons among different GNN models working under the proposed framework. Matteo Tiezzi, Gabriele Ciravegna, Marco Gori |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Building Bridges of Knowledge: Innovating Education with Automated Crossword GenerationabstractEducational crossword puzzles enhance critical thinking, vocabulary development, and concept reinforcement. They encourage independent learning, improve memorization, and foster problem-solving skills. With their multisensory approach, crossword puzzles offer a valuable educational experience. With the help of AI technology, creating high-quality, diverse crosswords is now easier, promoting enjoyable and effective learning experiences. In this endeavor, we harnessed the power of multiple language models, including GPT3, GPT2-XL, and BERT, to construct a comprehensive system that generates and verifies crossword clues. Our ultimate aim is to employ this system in the creation of educational crosswords. To achieve this, we compiled an extensive dataset consisting of over seven million clue-answer pairs spanning the years 1913 to mid-2021. By leveraging this dataset, we aimed to generate original yet challenging clues that engage solvers. Our generator underwent fine-tuning using this large collection of clues and corresponding answers, covering a wide range of themes. Additionally, we implemented a few/zero-shot learning techniques, such as prompt engineering, to generate clues based on given texts. To guarantee the quality of the generated clue-answer pairs, we utilized diverse classifiers, by fine-tuning pre-existing language models on a labeled dataset and additionally, we harnessed the power of the zero-shot learning approach to validate the generated clue-answer pairs effectively. This classifier effectively filters out nonsensical or subpar pairings. The evaluation results are highly encouraging, reinforcing the efficacy of the proposed approach. Kamyar Zeinalipour, Tommaso Iaquinta, Giovanni Angelini, Leonardo Rigutini, Marco Maggini, Marco Gori |
ICMLA | 6 |
| 2023 | Continual Learning with Pretrained Backbones by Tuning in the Input SpaceabstractThe intrinsic difficulty in adapting deep learning models to non-stationary environments limits the applicability of neural networks to real-world tasks. This issue is critical in practical supervised learning settings, such as the ones in which a pre-trained model computes projections toward a latent space where different task predictors are sequentially learned over time. As a matter of fact, incrementally fine-tuning the whole model to better adapt to new tasks usually results in catastrophic forgetting, with decreasing performance over the past experiences and losing valuable knowledge from the pretraining stage. In this paper, we propose a novel strategy to make the fine-tuning procedure more effective, by avoiding to update the pre-trained part of the network and learning not only the usual classification head, but also a set of newly-introduced learnable parameters that are responsible for transforming the input data. This process allows the network to effectively leverage the pre-training knowledge and find a good trade-off between plasticity and stability with modest computational efforts, thus especially suitable for on-the-edge settings. Our experiments on four image classification problems in a continual learning setting confirm the quality of the proposed approach when compared to several fine-tuning procedures and to popular continual learning methods. Simone Marullo, Matteo Tiezzi, Marco Gori, Stefano Melacci, Tinne Tuytelaars |
IJCNN | 3 |
| 2023 | Knowledge-Driven Active LearningabstractIn the last few years, Deep Learning models have become increasingly popular. However, their deployment is still precluded in those contexts where the amount of supervised data is limited and manual labelling expensive. Active learning strategies aim at solving this problem by requiring supervision only on few unlabelled samples, which improve the most model performances after adding them to the training set. Most strategies are based on uncertain sample selection, and even often restricted to samples lying close to the decision boundary. Here we propose a very different approach, taking into consideration domain knowledge. Indeed, in the case of multi-label classification, the relationships among classes offer a way to spot incoherent predictions, i.e., predictions where the model may most likely need supervision. We have developed a framework where first-order-logic knowledge is converted into constraints and their violation is checked as a natural guide for sample selection. We empirically demonstrate that knowledge-driven strategy outperforms standard strategies, particularly on those datasets where domain knowledge is complete. Furthermore, we show how the proposed approach enables discovering data distributions lying far from training data. Finally, the proposed knowledge-driven strategy can be also easily used in object-detection problems where standard uncertainty-based techniques are difficult to apply. Gabriele Ciravegna, Frédéric Precioso, Alessandro Betti, Kevin Mottin, Marco Gori |
ECML/PKDD (1) | 5 |
| 2023 | Logic Explained Networks
Gabriele Ciravegna, Pietro Barbiero, Francesco Giannini, Marco Gori, Pietro Liò, Marco Maggini, Stefano Melacci |
Artif. Intell. | 4 |
| 2023 | T-norms driven loss functions for machine learningabstractAbstract Injecting prior knowledge into the learning process of a neural architecture is one of the main challenges currently faced by the artificial intelligence community, which also motivated the emergence of neural-symbolic models. One of the main advantages of these approaches is their capacity to learn competitive solutions with a significant reduction of the amount of supervised data. In this regard, a commonly adopted solution consists of representing the prior knowledge via first-order logic formulas, then relaxing the formulas into a set of differentiable constraints by using a t-norm fuzzy logic. This paper shows that this relaxation, together with the choice of the penalty terms enforcing the constraint satisfaction, can be unambiguously determined by the selection of a t-norm generator, providing numerical simplification properties and a tighter integration between the logic knowledge and the learning objective. When restricted to supervised learning, the presented theoretical framework provides a straight derivation of the popular cross-entropy loss, which has been shown to provide faster convergence and to reduce the vanishing gradient problem in very deep structures. However, the proposed learning formulation extends the advantages of the cross-entropy loss to the general knowledge that can be represented by neural-symbolic methods. In addition, the presented methodology allows the development of novel classes of loss functions, which are shown in the experimental results to lead to faster convergence rates than the approaches previously proposed in the literature. Francesco Giannini, Michelangelo Diligenti, Marco Maggini, Marco Gori, Giuseppe Marra |
Appl. Intell. | 4 |
| 2023 | Local propagation of visual stimuli in focus of attentionabstractFast reactions to changes in the surrounding visual environment require efficient attention mechanisms to reallocate computational resources to most relevant locations in the visual field. While current computational models keep improving their predictive ability thanks to the increasing availability of data, they still struggle approximating the effectiveness and efficiency exhibited by foveated animals. In this paper, we present a biologically-plausible computational model of focus of attention that exhibits spatiotemporal locality and that is very well-suited for parallel and distributed implementations. Attention emerges as a wave propagation process originated by visual stimuli corresponding to details and motion information. The resulting field obeys the principle of "inhibition of return" so as not to get stuck in potential holes. An accurate experimentation of the model shows that it achieves top level performance in scanpath prediction tasks. This can easily be understood at the light of a theoretical result that we establish in the paper, where we prove that as the velocity of wave propagation goes to infinity, the proposed model reduces to recently proposed state of the art gravitational models of focus of attention. Lapo Faggi, Alessandro Betti, Dario Zanca, Stefano Melacci, Marco Gori |
Neurocomputing | 5 |
| 2022 | Entropy-Based Logic Explanations of Neural NetworksabstractExplainable artificial intelligence has rapidly emerged since lawmakers have started requiring interpretable models for safety-critical domains. Concept-based neural networks have arisen as explainable-by-design methods as they leverage human-understandable symbols (i.e. concepts) to predict class memberships. However, most of these approaches focus on the identification of the most relevant concepts but do not provide concise, formal explanations of how such concepts are leveraged by the classifier to make predictions. In this paper, we propose a novel end-to-end differentiable approach enabling the extraction of logic explanations from neural networks using the formalism of First-Order Logic. The method relies on an entropy-based criterion which automatically identifies the most relevant concepts. We consider four different case studies to demonstrate that: (i) this entropy-based criterion enables the distillation of concise logic explanations in safety-critical domains from clinical data to computer vision; (ii) the proposed approach outperforms state-of-the-art white-box models in terms of classification accuracy. Pietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Pietro Liò, Marco Gori, Stefano Melacci |
AAAI | 5 |
| 2022 | Being Friends Instead of Adversaries: Deep Networks Learn from Data Simplified by Other Networks
Simone Marullo, Matteo Tiezzi, Marco Gori, Stefano Melacci |
AAAI | 3 |
| 2022 | PARTIME: Scalable and Parallel Processing Over Time with Deep Neural NetworksabstractIn this paper, we present PARTIME, a software library written in Python and based on PyTorch, designed specifically to speed up neural networks whenever data is continuously streamed over time, for both learning and inference. Existing libraries are designed to exploit data-level parallelism, assuming that samples are batched, a condition that is not naturally met in applications that are based on streamed data. Differently, PARTIME starts processing each data sample at the time in which it becomes available from the stream. PARTIME wraps the code that implements a feed-forward multi-layer network and it distributes the layer-wise processing among multiple devices, such as Graphics Processing Units (GPUs). Thanks to its pipeline-based computational scheme, PARTIME allows the devices to perform computations in parallel. At inference time this results in scaling capabilities that are theoretically linear with respect to the number of devices. During the learning stage, PARTIME can leverage the non-i.i.d. nature of the streamed data with samples that are smoothly evolving over time for efficient gradient computations. Experiments are performed in order to empirically compare PARTIME with classic non-parallel neural computations in online learning, distributing operations on up to 8 NVIDIA GPUs, showing significant speedups that are almost linear in the number of devices, mitigating the impact of the data transfer overhead. Enrico Meloni, Lapo Faggi, Simone Marullo, Alessandro Betti, Matteo Tiezzi, Marco Gori, Stefano Melacci |
ICMLA | 6 |
| 2022 | Foveated Neural Computation
Matteo Tiezzi, Simone Marullo, Alessandro Betti, Enrico Meloni, Lapo Faggi, Marco Gori, Stefano Melacci |
ECML/PKDD (3) | 6 |
| 2022 | Domain Knowledge Alleviates Adversarial Attacks in Multi-Label ClassifiersabstractAdversarial attacks on machine learning-based classifiers, along with defense mechanisms, have been widely studied in the context of single-label classification problems. In this paper, we shift the attention to multi-label classification, where the availability of domain knowledge on the relationships among the considered classes may offer a natural way to spot incoherent predictions, i.e., predictions associated to adversarial examples lying outside of the training data distribution. We explore this intuition in a framework in which first-order logic knowledge is converted into constraints and injected into a semi-supervised learning problem. Within this setting, the constrained classifier learns to fulfill the domain knowledge over the marginal distribution, and can naturally reject samples with incoherent predictions. Even though our method does not exploit any knowledge of attacks during training, our experimental analysis surprisingly unveils that domain-knowledge constraints can help detect adversarial examples effectively, especially if such constraints are not known to the attacker. We show how to implement an adaptive attack exploiting knowledge of the constraints and, in a specifically-designed setting, we provide experimental comparisons with popular state-of-the-art attacks. We believe that our approach may provide a significant step towards designing more robust multi-label classifiers. Stefano Melacci, Gabriele Ciravegna, Angelo Sotgiu, Ambra Demontis, Battista Biggio, Marco Gori, Fabio Roli |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Guest Editorial: Non-Euclidean Machine LearningabstractOver the past decade, deep learning has had a revolutionary impact on a broad range of fields such as computer vision and image processing, computational photography, medical imaging and speech and language analysis and synthesis etc. Deep learning technologies are estimated to have added billions in business value, created new markets, and transformed entire industrial segments. Most of today’s successful deep learning methods such as Convolutional Neural Networks (CNNs) rely on classical signal processing models that limit their applicability to data with underlying Euclidean grid-like structure, e.g., images or acoustic signals. Yet, many applications deal with non-Euclidean (graph- or manifold-structured) data. For example, in social network analysis the users and their attributes are generally modeled as signals on the vertices of graphs. In biology protein-to-protein interactions are modeled as graphs. In computer vision & graphics 3D objects are modeled as meshes or point clouds. Furthermore, a graph representation is a very natural way to describe interactions between objects or signals. The classical deep learning paradigm on Euclidean domains falls short in providing appropriate tools for such kind of data. Until recently, the lack of deep learning models capable of correctly dealing with non-Euclidean data has been a major obstacle in these fields. This special section addresses the need to bring together leading efforts in non-Euclidean deep learning across all communities. From the papers that the special received twelve were selected for publication. The selected papers can naturally fall in three distinct categories: (a) methodologies that advance machine learning on data that are represented as graphs, (b) methodologies that advance machine learning on manifold-valued data, and (c) applications of machine learning methodologies on non-Euclidean spaces in computer vision and medical imaging. We briefly review the accepted papers in each of the groups. Stefanos Zafeiriou, Michael M. Bronstein, Taco Cohen, Oriol Vinyals, Jure Leskovec, Pietro Liò, Joan Bruna, Marco Gori |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2021 | Discrete and Continuous Deep Residual Learning over GraphsabstractIn this paper we propose the use of continuous residual modules for graph kernels in Graph Neural Networks. We show how both discrete and continuous residual layers allow for more robust training, being that continuous residual layers are those which are applied by integrating through an Ordinary Differential Equation (ODE) solver to produce their output. We experimentally show that these residuals achieve better results than the ones with non-residual modules when multiple layers are used, mitigating the low-pass filtering effect of GCN-based models. Finally, we apply and analyse the behaviour of these techniques and give pointers to how this technique can be useful in other domains by allowing more predictable behaviour under dynamic times of computation. Pedro H. C. Avelar, Anderson R. Tavares, Marco Gori, Luís C. Lamb |
ICAART (2) | 3 |
| 2021 | Messing Up 3D Virtual Environments: Transferable Adversarial 3D ObjectsabstractIn the last few years, the scientific community showed a remarkable and increasing interest towards 3D Virtual Environments, training and testing Machine Learning-based models in realistic virtual worlds. On one hand, these environments could also become a mean to study the weaknesses of Machine Learning algorithms, or to simulate training settings that allow Machine Learning models to gain robustness to 3D adversarial attacks. On the other hand, their growing popularity might also attract those that aim at creating adversarial conditions to invalidate the benchmarking process, especially in the case of public environments that allow the contribution from a large community of people. Most of the existing Adversarial Machine Learning approaches are focused on static images, and little work has been done in studying how to deal with 3D environments and how a 3D object should be altered to fool a classifier that observes it. In this paper, we study how to craft adversarial 3D objects by altering their textures, using a tool chain composed of easily accessible elements. We show that it is possible, and indeed simple, to create adversarial objects using off-the-shelf limited surrogate renderers that can compute gradients with respect to the parameters of the rendering process, and, to a certain extent, to transfer the attacks to more advanced 3D engines. We propose a saliency-based attack that intersects the two classes of renderers in order to focus the alteration to those texture elements that are estimated to be effective in the target engine, evaluating its impact in popular neural classifiers. Enrico Meloni, Matteo Tiezzi, Luca Pasqualini, Marco Gori, Stefano Melacci |
ICMLA | 4 |
| 2021 | The Principle of Least Cognitive Action in Vision
Marco Gori |
ICPRAM | 1 |
| 2021 | Friendly Training: Neural Networks Can Adapt Data To Make Learning EasierabstractIn the last decade, motivated by the success of Deep Learning, the scientific community proposed several approaches to make the learning procedure of Neural Networks more effective. When focussing on the way in which the training data are provided to the learning machine, we can distinguish between the classic random selection of stochastic gradient-based optimization and more involved techniques that devise curricula to organize data, and progressively increase the complexity of the training set. In this paper, we propose a novel training procedure named Friendly Training that, differently from the aforementioned approaches, involves altering the training examples in order to help the model to better fulfil its learning criterion. The model is allowed to “simplify” those examples that are too hard to be classified at a certain stage of the training procedure. The data transformation is controlled by a developmental plan that progressively reduces its impact during training, until it completely vanishes. In a sense, this is the opposite of what is commonly done in order to increase robustness against adversarial examples, i.e., Adversarial Training. Experiments on multiple datasets are provided, showing that Friendly Training yields improvements with respect to informed data sub-selection routines and random selection, especially in deep convolutional architectures. Results suggest that adapting the input data is a feasible way to stabilize learning and improve the generalization skills of the network. Simone Marullo, Matteo Tiezzi, Marco Gori, Stefano Melacci |
IJCNN | 3 |
| 2021 | Regularizing deep networks with prior knowledge: A constraint-based approachabstractDeep Learning architectures can develop feature representations and classification models in an integrated way during training. This joint learning process requires large networks with many parameters, and it is successful when a large amount of training data is available. Instead of making the learner develop its entire understanding of the world from scratch from the input examples, the injection of prior knowledge into the learner seems to be a principled way to reduce the amount of require training data, as the learner does not need to induce the rules from the data. This paper presents a general framework to integrate arbitrary prior knowledge into learning. The domain knowledge is provided as a collection of first-order logic (FOL) clauses, where each task to be learned corresponds to a predicate in the knowledge base. The logic statements are translated into a set of differentiable constraints, which can be integrated into the learning process to distill the knowledge into the network, or used during inference to enforce the consistency of the predictions with the prior knowledge. The experimental results have been carried out on multiple image datasets and show that the integration of the prior knowledge boosts the accuracy of several state-of-the-art deep architectures on image classification tasks. Soumali Roychowdhury, Michelangelo Diligenti, Marco Gori |
Knowl. Based Syst. | 3 |
| 2021 | A language modeling-like approach to sketching
Lisa Graziani, Marco Gori, Stefano Melacci |
Neural Networks | 2 |
| 2020 | A Constraint-Based Approach to Learning and ExplanationabstractIn the last few years we have seen a remarkable progress from the cultivation of the idea of expressing domain knowledge by the mathematical notion of constraint. However, the progress has mostly involved the process of providing consistent solutions with a given set of constraints, whereas learning “new” constraints, that express new knowledge, is still an open challenge. In this paper we propose a novel approach to learning of constraints which is based on information theoretic principles. The basic idea consists in maximizing the transfer of information between task functions and a set of learnable constraints, implemented using neural networks subject to L1 regularization. This process leads to the unsupervised development of new constraints that are fulfilled in different sub-portions of the input domain. In addition, we define a simple procedure that can explain the behaviour of the newly devised constraints in terms of First-Order Logic formulas, thus extracting novel knowledge on the relationships between the original tasks. An experimental evaluation is provided to support the proposed approach, in which we also explore the regularization effects introduced by the proposed Information-Based Learning of Constraint (IBLC) algorithm. Gabriele Ciravegna, Francesco Giannini, Stefano Melacci, Marco Maggini, Marco Gori |
AAAI | 5 |
| 2020 | Relational Neural MachinesabstractDeep learning has been shown to achieve impressive results in several tasks where a large amount of training data is available. However, deep learning solely focuses on the accuracy of the predictions, neglecting the reasoning process leading to a decision, which is a major issue in life-critical applications. Probabilistic logic reasoning allows to exploit both statistical regularities and specific domain expertise to perform reasoning under uncertainty, but its scalability and brittle integration with the layers processing the sensory data have greatly limited its applications. For these reasons, combining deep architectures and probabilistic logic reasoning is a fundamental goal towards the development of intelligent agents operating in complex environments. This paper presents Relational Neural Machines, a novel framework allowing to jointly train the parameters of the learners and of a First-Order Logic based reasoner. A Relational Neural Machine is able to recover both classical learning from supervised data in case of pure sub-symbolic learning, and Markov Logic Networks in case of pure symbolic reasoning, while allowing to jointly train and perform inference in hybrid learning tasks. Proper algorithmic solutions are devised to make learning and inference tractable in large-scale problems. The experiments show promising results in different relational tasks. Giuseppe Marra, Michelangelo Diligenti, Francesco Giannini, Marco Gori, Marco Maggini |
ECAI | 4 |
| 2020 | A Lagrangian Approach to Information Propagation in Graph Neural NetworksabstractIn many real world applications, data are characterized by a complex structure, that can be naturally encoded as a graph.In the last years, the popularity of deep learning techniques has renewed the interest in neural models able to process complex patterns.In particular, inspired by the Graph Neural Network (GNN) model, different architectures have been proposed to extend the original GNN scheme.GNNs exploit a set of state variables, each assigned to a graph node, and a diffusion mechanism of the states among neighbor nodes, to implement an iterative procedure to compute the fixed point of the (learnable) state transition function.In this paper, we propose a novel approach to the state computation and the learning algorithm for GNNs, based on a constraint optimisation task solved in the Lagrangian framework.The state convergence procedure is implicitly expressed by the constraint satisfaction mechanism and does not require a separate iterative phase for each epoch of the learning procedure.In fact, the computational structure is based on the search for saddle points of the Lagrangian in the adjoint space composed of weights, neural outputs (node states), and Lagrange multipliers.The proposed approach is compared experimentally with other popular models for processing graphs. Matteo Tiezzi, Giuseppe Marra, Stefano Melacci, Marco Maggini, Marco Gori |
ECAI | 5 |
| 2020 | Generating Facial Expressions Associated with Text
Lisa Graziani, Stefano Melacci, Marco Gori |
ICANN (1) | 3 |
| 2020 | SAILenv: Learning in Virtual Visual Environments Made SimpleabstractRecently, researchers in Machine Learning algorithms, Computer Vision scientists, engineers and others, showed a growing interest in 3D simulators as a mean to artificially create experimental settings that are very close to those in the real world. However, most of the existing platforms to interface algorithms with 3D environments are often designed to setup navigation-related experiments, to study physical interactions, or to handle ad-hoc cases that are not thought to be customized, sometimes lacking a strong photorealistic appearance and an easy-to-use software interface. In this paper, we present a novel platform, SAILenv, that is specifically designed to be simple and customizable, and that allows researchers to experiment visual recognition in virtual 3D scenes. A few lines of code are needed to interface every algorithm with the virtual world, and non-3D-graphics experts can easily customize the 3D environment itself, exploiting a collection of photorealistic objects. Our framework yields pixel-level semantic and instance labeling, depth, and, to the best of our knowledge, it is the only one that provides motion-related information directly inherited from the 3D engine. The client-server communication operates at a low level, avoiding the overhead of HTTP-based data exchanges. We perform experiments using a state-of-the-art object detector trained on real-world images, showing that it is able to recognize the photorealistic 3D objects of our environment. The computational burden of the optical flow compares favourably with the estimation performed using modern GPU-based convolutional networks or more classic implementations. We believe that the scientific community will benefit from the easiness and high-quality of our framework to evaluate newly proposed algorithms in their own customized realistic conditions. Enrico Meloni, Luca Pasqualini, Matteo Tiezzi, Marco Gori, Stefano Melacci |
ICPR | 4 |
| 2020 | Human-Driven FOL Explanations of Deep LearningabstractDeep neural networks are usually considered black-boxes due to their complex internal architecture, that cannot straightforwardly provide human-understandable explanations on how they behave. Indeed, Deep Learning is still viewed with skepticism in those real-world domains in which incorrect predictions may produce critical effects. This is one of the reasons why in the last few years Explainable Artificial Intelligence (XAI) techniques have gained a lot of attention in the scientific community. In this paper, we focus on the case of multi-label classification, proposing a neural network that learns the relationships among the predictors associated to each class, yielding First-Order Logic (FOL)-based descriptions. Both the explanation-related network and the classification-related network are jointly learned, thus implicitly introducing a latent dependency between the development of the explanation mechanism and the development of the classifiers. Our model can integrate human-driven preferences that guide the learning-to-explain process, and it is presented in a unified framework. Different typologies of explanations are evaluated in distinct experiments, showing that the proposed approach discovers new knowledge and can improve the classifier performance. Gabriele Ciravegna, Francesco Giannini, Marco Gori, Marco Maggini, Stefano Melacci |
IJCAI | 3 |
| 2020 | Graph Neural Networks Meet Neural-Symbolic Computing: A Survey and PerspectiveabstractNeural-symbolic computing has now become the subject of interest of both academic and industry research laboratories. Graph Neural Networks (GNNs) have been widely used in relational and symbolic domains, with widespread application of GNNs in combinatorial optimization, constraint satisfaction, relational reasoning and other scientific domains. The need for improved explainability, interpretability and trust of AI systems in general demands principled methodologies, as suggested by neural-symbolic computing. In this paper, we review the state-of-the-art on the use of GNNs as a model of neural-symbolic computing. This includes the application of GNNs in several domains as well as their relationship to current developments in neural-symbolic computing. Luís C. Lamb, Artur S. d'Avila Garcez, Marco Gori, Marcelo O. R. Prates, Pedro H. C. Avelar, Moshe Y. Vardi |
IJCAI | 3 |
| 2020 | Developing Constrained Neural Units Over TimeabstractIn this paper we present a foundational study on a constrained method that defines learning problems with Neural Networks in the context of the principle of least cognitive action, which very much resembles the principle of least action in mechanics. Starting from a general approach to enforce constraints into the dynamical laws of learning, this work focuses on an alternative way of defining Neural Networks, that is different from the majority of existing approaches. In particular, the structure of the neural architecture is defined by means of a special class of constraints that are extended also to the interaction with data, leading to "architectural" and "input-related" constraints, respectively. The proposed theory is cast into the time domain, in which data are presented to the network in an ordered manner, that makes this study an important step toward alternative ways of processing continuous streams of data with Neural Networks. The connection with the classic Backpropagation-based update rule of the weights of networks is discussed, showing that there are conditions under which our approach degenerates to Backpropagation. Moreover, the theory is experimentally evaluated on a simple problem that allows us to deeply study several aspects of the theory itself and to show the soundness of the model. Alessandro Betti, Marco Gori, Simone Marullo, Stefano Melacci |
IJCNN | 2 |
| 2020 | Local Propagation in Constraint-based Neural NetworksabstractIn this paper we study a constraint-based representation of neural network architectures. We cast the learning problem in the Lagrangian framework and we investigate a simple optimization procedure that is well suited to fulfil the so-called architectural constraints, learning from the available supervisions. The computational structure of the proposed Local Propagation (LP) algorithm is based on the search for saddle points in the adjoint space composed of weights, neural outputs, and Lagrange multipliers. All the updates of the model variables are locally performed, so that LP is fully parallelizable over the neural units, circumventing the classic problem of gradient vanishing in deep networks. The implementation of popular neural models is described in the context of LP, together with those conditions that trace a natural connection with Backpropagation. We also investigate the setting in which we tolerate bounded violations of the architectural constraints, and we provide experimental evidence that LP is a feasible approach to train shallow and deep networks, opening the road to further investigations on more complex architectures, easily describable by constraints. Giuseppe Marra, Matteo Tiezzi, Stefano Melacci, Alessandro Betti, Marco Maggini, Marco Gori |
IJCNN | 6 |
| 2020 | Toward Improving the Evaluation of Visual Attention Models: a Crowdsourcing ApproachabstractHuman visual attention is a complex phenomenon. A computational modeling of this phenomenon must take into account where people look in order to evaluate which are the salient locations (spatial distribution of the fixations), when they look in those locations to understand the temporal development of the exploration (temporal order of the fixations), and how they move from one location to another with respect to the dynamics of the scene and the mechanics of the eyes (dynamics). State-of-the-art models focus on learning saliency maps from human data, a process that only takes into account the spatial component of the phenomenon and ignore its temporal and dynamical counterparts. In this work we focus on the evaluation methodology of models of human visual attention. We underline the limits of the current metrics for saliency prediction and scanpath similarity, and we introduce a statistical measure for the evaluation of the dynamics of the simulated eye movements. While deep learning models achieve astonishing performance in saliency prediction, our analysis shows their limitations in capturing the dynamics of the process. We find that unsupervised gravitational models, despite of their simplicity, outperform all competitors. Finally, exploiting a crowd-sourcing platform, we present a study aimed at evaluating how strongly the scanpaths generated with the unsupervised gravitational models appear plausible to naive and expert human observers. Dario Zanca, Stefano Melacci, Marco Gori |
IJCNN | 3 |
| 2020 | Focus of Attention Improves Information Transfer in Visual FeaturesabstractUnsupervised learning from continuous visual streams is a challenging problem that cannot be naturally and efficiently managed in the classic batch-mode setting of computation. The information stream must be carefully processed accordingly to an appropriate spatio-temporal distribution of the visual data, while most approaches of learning commonly assume uniform probability density. In this paper we focus on unsupervised learning for transferring visual information in a truly online setting by using a computational model that is inspired to the principle of least action in physics. The maximization of the mutual information is carried out by a temporal process which yields online estimation of the entropy terms. The model, which is based on second-order differential equations, maximizes the information transfer from the input to a discrete space of symbols related to the visual features of the input, whose computation is supported by hidden neurons. In order to better structure the input probability distribution, we use a human-like focus of attention model that, coherently with the information maximization model, is also based on second-order differential equations. We provide experimental results to support the theory by showing that the spatio-temporal filtering induced by the focus of attention allows the system to globally transfer more information from the input stream over the focused areas and, in some contexts, over the whole frames with respect to the unfiltered case that yields uniform probability distributions. Matteo Tiezzi, Stefano Melacci, Alessandro Betti, Marco Maggini, Marco Gori |
NeurIPS | 5 |
| 2020 | Learning visual features under motion invarianceabstractHumans are continuously exposed to a stream of visual data with a natural temporal structure. However, most successful computer vision algorithms work at image level, completely discarding the precious information carried by motion. In this paper, we claim that processing visual streams naturally leads to formulate the motion invariance principle, which enables the construction of a new theory of learning that originates from variational principles, just like in physics. Such principled approach is well suited for a discussion on a number of interesting questions that arise in vision, and it offers a well-posed computational scheme for the discovery of convolutional filters over the retina. Differently from traditional convolutional networks, which need massive supervision, the proposed theory offers a truly new scenario for the unsupervised processing of video signals, where features are extracted in a multi-layer architecture with motion invariance. While the theory enables the implementation of novel computer vision systems, it also sheds light on the role of information-based principles to drive possible biological solutions. Alessandro Betti, Marco Gori, Stefano Melacci |
Neural Networks | 2 |
| 2020 | Gravitational Laws of Focus of AttentionabstractThe understanding of the mechanisms behind focus of attention in a visual scene is a problem of great interest in visual perception and computer vision. In this paper, we describe a model of scanpath as a dynamic process which can be interpreted as a variational law somehow related to mechanics, where the focus of attention is subject to a gravitational field. The distributed virtual mass that drives eye movements is associated with the presence of details and motion in the video. Unlike most current models, the proposed approach does not estimate directly the saliency map, but the prediction of eye movements allows us to integrate over time the positions of interest. The process of inhibition-of-return is also supported in the same dynamic model with the purpose of simulating fixations and saccades. The differential equations of motion of the proposed model are numerically integrated to simulate scanpaths on both images and videos. Experimental results for the tasks of saliency and scanpath prediction on a wide collection of datasets are presented to support the theory. Top level performances are achieved especially in the prediction of scanpaths, which is the primary purpose of the proposed model. Dario Zanca, Stefano Melacci, Marco Gori |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Pattern recognition and beyond: Alfredo Petrosino's scientific results
Lucia Maddalena 0001, Marco Gori, Sankar K. Pal |
Pattern Recognit. Lett. | 2 |
| 2020 | Cognitive Action Laws: The Case of Visual Features
Alessandro Betti, Marco Gori, Stefano Melacci |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Jointly Learning to Detect Emotions and Predict Facebook Reactions
Lisa Graziani, Stefano Melacci, Marco Gori |
ICANN (4) | 3 |
| 2019 | Constraint-Based Visual Generation
Giuseppe Marra, Francesco Giannini, Michelangelo Diligenti, Marco Gori |
ICANN (3) | 4 |
| 2019 | Motion Invariance in Visual EnvironmentsabstractThe puzzle of computer vision might find new challenging solutions when we realize that most successful methods are working at image level, which is remarkably more difficult than processing directly visual streams, just as it happens in nature. In this paper, we claim that the processing of a stream of frames naturally leads to formulate the motion invariance principle, which enables the construction of a new theory of visual learning based on convolutional features. The theory addresses a number of intriguing questions that arise in natural vision, and offers a well-posed computational scheme for the discovery of convolutional filters over the retina. They are driven by the Euler- Lagrange differential equations derived from the principle of least cognitive action, that parallels the laws of mechanics. Unlike traditional convolutional networks, which need massive supervision, the proposed theory offers a truly new scenario in which feature learning takes place by unsupervised processing of video signals. An experimental report of the theory is presented where we show that features extracted under motion invariance yield an improvement that can be assessed by measuring information-based indexes. Alessandro Betti, Marco Gori, Stefano Melacci |
IJCAI | 2 |
| 2019 | On the Relation Between Loss Functions and T-Norms
Francesco Giannini, Giuseppe Marra, Michelangelo Diligenti, Marco Maggini, Marco Gori |
ILP | 5 |
| 2019 | LYRICS: A General Interface Layer to Integrate Logic Inference and Deep Learning
Giuseppe Marra, Francesco Giannini, Michelangelo Diligenti, Marco Gori |
ECML/PKDD (2) | 4 |
| 2019 | Integrating Learning and Reasoning with Deep Logic Models
Giuseppe Marra, Francesco Giannini, Michelangelo Diligenti, Marco Gori |
ECML/PKDD (2) | 4 |
| 2019 | On a Convex Logic Fragment for Learning and ReasoningabstractIn this paper, we introduce the convex fragment of Łukasiewicz logic and discuss its possible applications in different learning schemes. The provided theoretical results are highly general because they can be exploited in any learning framework involving logical constraints. The method is of particular interest since the fragment guarantees to deal with convex constraints, which are shown to be equivalent to a set of linear constraints. Within this framework, we are able to formulate learning with kernel machines as well as collective classification as a quadratic programming problem. Francesco Giannini, Michelangelo Diligenti, Marco Gori, Marco Maggini |
IEEE Trans. Fuzzy Syst. | 3 |
| 2018 | Characterization of the Convex Łukasiewicz Fragment for Learning From Constraints
Francesco Giannini, Michelangelo Diligenti, Marco Gori, Marco Maggini |
AAAI | 3 |
| 2017 | Integrating Prior Knowledge into Deep LearningabstractDeep learning allows to develop feature representations and train classification models in a fully integrated way. However, learning deep networks is quite hard and it improves over shallow architectures only if a large number of training data is available. Injecting prior knowledge into the learner is a principled way to reduce the amount of required training data, as the learner does not need to induce the knowledge from the data itself. In this paper we propose a general and principled way to integrate prior knowledge when training deep networks. Semantic Based Regularization (SBR) is used as underlying framework to represent the prior knowledge, expressed as a collection of first-order logic clauses (FOL), and where each task to be learned corresponds to a predicate in the knowledge base. The knowledge base correlates the tasks to be learned and it is translated into a set of constraints which are integrated into the learning process via backpropagation. The experimental results show how the integration of the prior knowledge boosts the accuracy of a state-of-the-art deep network on an image classification task. Michelangelo Diligenti, Soumali Roychowdhury, Marco Gori |
ICMLA | 3 |
| 2017 | Variational Laws of Visual Attention for Dynamic ScenesabstractComputational models of visual attention are at the crossroad of disciplines like cognitive science, computational neuroscience, and computer vision. This paper proposes a model of attentional scanpath that is based on the principle that there are foundational laws that drive the emergence of visual attention. We devise variational laws of the eye-movement that rely on a generalized view of the Least Action Principle in physics. The potential energy captures details as well as peripheral visual features, while the kinetic energy corresponds with the classic interpretation in analytic mechanics. In addition, the Lagrangian contains a brightness invariance term, which characterizes significantly the scanpath trajectories. We obtain differential equations of visual attention as the stationary point of the generalized action, and we propose an algorithm to estimate the model parameters. Finally, we report experimental results to validate the model in tasks of saliency detection. Dario Zanca, Marco Gori |
NIPS | 2 |
| 2017 | Learning Łukasiewicz Logic Fragments by Quadratic Programming
Francesco Giannini, Michelangelo Diligenti, Marco Gori, Marco Maggini |
ECML/PKDD (1) | 3 |
| 2017 | Learning Semantic-based Structures from Textual Sources
Marco Gori |
WEBIST | 1 |
| 2017 | Semantic-based regularization for learning and inference
Michelangelo Diligenti, Marco Gori, Claudio Saccà |
Artif. Intell. | 2 |
| 2017 | LQG Online LearningabstractOptimal control theory and machine learning techniques are combined to formulate and solve in closed form an optimal control formulation of online learning from supervised examples with regularization of the updates. The connections with the classical linear quadratic gaussian (LQG) optimal control problem, of which the proposed learning paradigm is a nontrivial variation as it involves random matrices, are investigated. The obtained optimal solutions are compared with the Kalman filter estimate of the parameter vector to be learned. It is shown that the proposed algorithm is less sensitive to outliers with respect to the Kalman estimate (thanks to the presence of the regularization term), thus providing smoother estimates with respect to time. The basic formulation of the proposed online learning framework refers to a discrete-time setting with a finite learning horizon and a linear model. Various extensions are investigated, including the infinite learning horizon and, via the so-called kernel trick, the case of nonlinear models. Giorgio Gnecco, Alberto Bemporad, Marco Gori, Marcello Sanguineti |
Neural Comput. | 3 |
| 2016 | Learning with hard constraints as a limit case of learning with soft constraints
Giorgio Gnecco, Marco Gori, Stefano Melacci, Marcello Sanguineti |
ESANN | 2 |
| 2016 | Learning Efficiently in Semantic Based Regularization
Michelangelo Diligenti, Marco Gori, Vincenzo Scoca |
ECML/PKDD (2) | 2 |
| 2016 | Semantic video labeling by developmental visual agents
Marco Gori, Marco Lippi 0001, Marco Maggini, Stefano Melacci |
Comput. Vis. Image Underst. | 1 |
| 2016 | Neural network training as a dissipative process
Marco Gori, Marco Maggini, Alessandro Rossi 0002 |
Neural Networks | 1 |
| 2016 | The principle of least cognitive action
Alessandro Betti, Marco Gori |
Theor. Comput. Sci. | 2 |
| 2016 | Learning in Variable-Dimensional SpacesabstractThis paper proposes a unified approach to learning in environments in which patterns can be represented in variable-dimension domains, which nicely includes the case in which there are missing features. The proposal is based on the representation of the environment by pointwise constraints that are shown to model naturally pattern relationships that come out in problems of information retrieval, computer vision, and related fields. The given interpretation of learning leads to capturing the truly different aspects of similarity coming from the content at different dimensions and the pattern links. It turns out that functions that process real-valued features and functions that operate on symbolic entities are learned within a unified framework of regularization that can also be expressed using the kernel machines mathematical and algorithmic apparatus. Interestingly, in the extreme cases in which only the content or only the links are available, our theory returns classic kernel machines or graph regularization, respectively. We show experimental results that provide clear evidence of the remarkable improvements that are obtained when both types of similarities are exploited on artificial and real-world benchmarks. Michelangelo Diligenti, Marco Gori, Claudio Saccà |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Foundations of Support Constraint MachinesabstractThe mathematical foundations of a new theory for the design of intelligent agents are presented. The proposed learning paradigm is centered around the concept of constraint, representing the interactions with the environment, and the parsimony principle. The classical regularization framework of kernel machines is naturally extended to the case in which the agents interact with a richer environment, where abstract granules of knowledge, compactly described by different linguistic formalisms, can be translated into the unified notion of constraint for defining the hypothesis set. Constrained variational calculus is exploited to derive general representation theorems that provide a description of the optimal body of the agent (i.e., the functional structure of the optimal solution to the learning problem), which is the basis for devising new learning algorithms. We show that regardless of the kind of constraints, the optimal body of the agent is a support constraint machine (SCM) based on representer theorems that extend classical results for kernel machines and provide new representations. In a sense, the expressiveness of constraints yields a semantic-based regularization theory, which strongly restricts the hypothesis set of classical regularization. Some guidelines to unify continuous and discrete computational mechanisms are given so as to accommodate in the same framework various kinds of stimuli, for example, supervised examples and logic predicates. The proposed view of learning from constraints incorporates classical learning from examples and extends naturally to the case in which the examples are subsets of the input space, which is related to learning propositional logic clauses. Giorgio Gnecco, Marco Gori, Stefano Melacci, Marcello Sanguineti |
Neural Comput. | 2 |
| 2015 | Learning With Mixed Hard/Soft Pointwise ConstraintsabstractA learning paradigm is proposed and investigated, in which the classical framework of learning from examples is enhanced by the introduction of hard pointwise constraints, i.e., constraints imposed on a finite set of examples that cannot be violated. Such constraints arise, e.g., when requiring coherent decisions of classifiers acting on different views of the same pattern. The classical examples of supervised learning, which can be violated at the cost of some penalization (quantified by the choice of a suitable loss function) play the role of soft pointwise constraints. Constrained variational calculus is exploited to derive a representer theorem that provides a description of the functional structure of the optimal solution to the proposed learning paradigm. It is shown that such an optimal solution can be represented in terms of a set of support constraints, which generalize the concept of support vectors and open the doors to a novel learning paradigm, called support constraint machines. The general theory is applied to derive the representation of the optimal solution to the problem of learning from hard linear pointwise constraints combined with soft pointwise constraints induced by supervised examples. In some cases, closed-form optimal solutions are obtained. Giorgio Gnecco, Marco Gori, Stefano Melacci, Marcello Sanguineti |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | A theoretical framework for supervised learning from regions
Giorgio Gnecco, Marco Gori, Stefano Melacci, Marcello Sanguineti |
Neurocomputing | 2 |
| 2013 | Variational Foundations of Online Backpropagation
Salvatore Frandina, Marco Gori, Marco Lippi 0001, Marco Maggini, Stefano Melacci |
ICANN | 2 |
| 2013 | Learning with Hard Constraints
Giorgio Gnecco, Marco Gori, Stefano Melacci, Marcello Sanguineti |
ICANN | 2 |
| 2013 | Collective Classification Using Semantic Based RegularizationabstractSemantic Based Regularization (SBR) is a framework for injecting prior knowledge expressed as FOL clauses into a semi-supervised learning problem. The prior knowledge is converted into a set of continuous constraints, which are enforced during training. SBR employs the prior knowledge only at training time, hoping that the learning process is able to encode the knowledge via the training data into its parameters. This paper defines a collective classification approach employing the prior knowledge at test time, naturally reusing most of the mathematical apparatus developed for standard SBR. The experimental results show that the presented method outperforms state-of-the-art classification methods on multiple text categorization tasks. Claudio Saccà, Michelangelo Diligenti, Marco Gori |
ICMLA (1) | 3 |
| 2013 | Graph and Manifold Co-regularizationabstractClassical foundations of Statistical Learning Theory rely on the assumption that the input patterns are independently and identically distributed. However, in many applications, the inputs, represented as feature vectors, are also embedded into a network of pair wise relations. Transductive approaches like graph regularization rely on the network topology without considering the feature vectors. Semi-supervised approaches like Manifold Regularization learn a function taking the feature vectors as input, while being smooth over the network connections. In this latter case, the connectivity information is processed at training time, but is still neglected during generalization, as the final classification decision takes only the feature vector representations as input. This paper presents and evaluates a model merging the advantages of graph regularization and kernel machines for transductive classification problems. Claudio Saccà, Michelangelo Diligenti, Marco Gori |
ICMLA (1) | 3 |
| 2013 | Learning with Boundary ConditionsabstractKernel machines traditionally arise from an elegant formulation based on measuring the smoothness of the admissible solutions by the norm in the reproducing kernel Hilbert space (RKHS) generated by the chosen kernel. It was pointed out that they can be formulated in a related functional framework, in which the Green's function of suitable differential operators is thought of as a kernel. In this letter, our own picture of this intriguing connection is given by emphasizing some relevant distinctions between these different ways of measuring the smoothness of admissible solutions. In particular, we show that for some kernels, there is no associated differential operator. The crucial relevance of boundary conditions is especially emphasized, which is in fact the truly distinguishing feature of the approach based on differential operators. We provide a general solution to the problem of learning from data and boundary conditions and illustrate the significant role played by boundary conditions with examples. It turns out that the degree of freedom that arises in the traditional formulation of kernel machines is indeed a limitation, which is partly overcome when incorporating the boundary conditions. This likely holds true in many real-world applications in which there is prior knowledge about the expected behavior of classifiers and regressors on the boundary. Giorgio Gnecco, Marco Gori, Marcello Sanguineti |
Neural Comput. | 2 |
| 2013 | Learning with Box KernelsabstractSupervised examples and prior knowledge on regions of the input space have been profitably integrated in kernel machines to improve the performance of classifiers in different real-world contexts. The proposed solutions, which rely on the unified supervision of points and sets, have been mostly based on specific optimization schemes in which, as usual, the kernel function operates on points only. In this paper, arguments from variational calculus are used to support the choice of a special class of kernels, referred to as box kernels, which emerges directly from the choice of the kernel function associated with a regularization operator. It is proven that there is no need to search for kernels to incorporate the structure deriving from the supervision of regions of the input space, because the optimal kernel arises as a consequence of the chosen regularization operator. Although most of the given results hold for sets, we focus attention on boxes, whose labeling is associated with their propositional description. Based on different assumptions, some representer theorems are given that dictate the structure of the solution in terms of box kernel expansion. Successful results are given for problems of medical diagnosis, image, and text categorization. Stefano Melacci, Marco Gori |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Constraint Verification With Kernel MachinesabstractBased on a recently proposed framework of learning from constraints using kernel-based representations, in this brief, we naturally extend its application to the case of inferences on new constraints. We give examples for polynomials and first-order logic by showing how new constraints can be checked on the basis of given premises and data samples. Interestingly, this gives rise to a perceptual logic scheme in which the inference mechanisms do not rely only on formal schemes, but also on the data probability distribution. It is claimed that when using a properly relaxed computational checking approach, the complementary role of data samples makes it possible to break the complexity barriers of related formal checking mechanisms. Marco Gori, Stefano Melacci |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | Information Theoretic Learning for Pixel-Based Visual Agents
Marco Gori, Stefano Melacci, Marco Lippi 0001, Marco Maggini |
ECCV (6) | 1 |
| 2012 | Bridging logic and kernel machines
Michelangelo Diligenti, Marco Gori, Marco Maggini, Leonardo Rigutini |
Mach. Learn. | 2 |
| 2012 | Unsupervised Learning by Minimal Entropy EncodingabstractFollowing basic principles of information-theoretic learning, in this paper, we propose a novel approach to data clustering, referred to as minimal entropy encoding (MEE), which is based on a set of functions (features) projecting each input onto a minimum entropy configuration (code). Inspired by traditional parsimony principles, we seek solutions in reproducing kernel Hilbert spaces and then we prove that the encoding functions are expressed in terms of kernel expansion. In order to avoid trivial solutions, the developed features must be as different as possible by means of a soft constraint on the empirical estimation of the entropy associated with the encoding functions. This leads to an unconstrained optimization problem that can be efficiently solved by conjugate gradient. We also investigate an optimization strategy based on concave-convex algorithms. The relationships with maximum margin clustering are studied, showing that MEE overcomes some of its critical issues, such as the lack of a multiclass extension and the need to face problems with a large number of constraints. A massive evaluation on several benchmarks of the proposed approach shows improvements over state-of-the-art techniques, both in terms of accuracy and computational complexity. Stefano Melacci, Marco Gori |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2011 | Support Constraint Machines
Marco Gori, Stefano Melacci |
ICONIP (1) | 1 |
| 2011 | Learning with Box Kernels
Stefano Melacci, Marco Gori |
ICONIP (2) | 2 |
| 2011 | Learning from Constraints
Marco Gori |
ECML/PKDD (1) | 1 |
| 2011 | A unified representation of web logs for mining applications
Michelangelo Diligenti, Marco Gori, Marco Maggini |
Inf. Retr. | 2 |
| 2011 | Neural networks for relational learning: an experimental comparison
Werner Uwents, Gabriele Monfardini, Hendrik Blockeel, Marco Gori, Franco Scarselli |
Mach. Learn. | 4 |
| 2010 | Multitask Kernel-based Learning with Logic ConstraintsabstractThis paper presents a general framework to integrate prior knowledge in the form of logic constraints among a set of task functions into kernel machines. The logic propositions provide a partial representation of the environment, in which the learner operates, that is exploited by the learning algorithm together with the information available in the supervised examples. In particular, we consider a multi-task learning scheme, where multiple unary predicates on the feature space are to be learned by kernel machines and a higher level abstract representation consists of logic clauses on these predicates, known to hold for any input. A general approach is presented to convert the logic clauses into a continuous implementation, that processes the outputs computed by the kernel-based predicates. The learning task is formulated as a primal optimization problem of a loss function that combines a term measuring the fitting of the supervised examples, a regularization term, and a penalty term that enforces the constraints on both supervised and unsupervised examples. The proposed semi-supervised learning framework is particularly suited for learning in high dimensionality feature spaces, where the supervised training examples tend to be sparse and generalization difficult. Unlike for standard kernel machines, the cost function to optimize is not generally guaranteed to be convex. However, the experimental results show that it is still possible to find good solutions using a two stage learning schema, in which first the supervised examples are learned until convergence and then the logic constraints are forced. Some promising experimental results on artificial multi-task learning tasks are reported, showing how the classification accuracy can be effectively improved by exploiting the a priori rules and the unsupervised examples. Michelangelo Diligenti, Marco Gori, Marco Maggini, Leonardo Rigutini |
ECAI | 2 |
| 2010 | Kernel-Based Hybrid Random Fields for Nonparametric Density EstimationabstractHybrid random fields are a recently proposed graphical model for pseudo-likelihood estimation in discrete domains. In this paper, we develop a continuous version of the model for nonparametric density estimation. To this aim, Nadaraya-Watson kernel estimators are used to model the local conditional densities within hybrid random fields. First, we introduce a heuristic algorithm for tuning the kernel bandwidhts in the conditional density estimators. Second, we propose a novel method for initializing the structure learning algorithm originally employed for hybrid random fields, which was meant instead for discrete variables. In order to test the accuracy of the proposed technique, we use a number of synthetic pattern classification benchmarks, generated from random distributions featuring nonlinear correlations between the variables. As compared to state-of-the-art nonparametric and semiparametric learning techniques for probabilistic graphical models, kernel-based hybrid random fields regularly outperform each considered alternative in terms of recognition accuracy, while preserving the scalability properties (with respect to the number of variables) that originally motivated their introduction. Antonino Freno, Edmondo Trentin, Marco Gori |
ECAI | 3 |
| 2010 | Learning with Convex Constraints
Marco Gori, Stefano Melacci |
ICANN (3) | 1 |
| 2010 | A template-based approach to automatic face enhancement
Stefano Melacci, Lorenzo Sarti, Marco Maggini, Marco Gori |
Pattern Anal. Appl. | 4 |
| 2009 | Semi-supervised Learning with Constraints for Multi-view Object Recognition
Stefano Melacci, Marco Maggini, Marco Gori |
ICANN (2) | 3 |
| 2009 | Scalable statistical learning: A modular bayesian/markov network approachabstractIn this paper we propose a hybrid probabilistic graphical model for pseudo-likelihood estimation in high-dimensional domains. The model is based on Bayesian networks and Markov random fields. On the one hand, we prove that the proposed model is more expressive than Bayesian networks in terms of the representable distributions. On the other hand, we develop a computationally efficient structure learning algorithm, and we provide theoretical and experimental evidence showing how the modular nature of our model allows structure learning to scale up very well to high-dimensional datasets. The capability of the hybrid model to accurately learn complex networks of conditional independencies is illustrated by promising results in pattern recognition applications. Antonino Freno, Edmondo Trentin, Marco Gori |
IJCNN | 3 |
| 2009 | Scalable pseudo-likelihood estimation in hybrid random fieldsabstractLearning probabilistic graphical models from high-dimensional datasets is a computationally challenging task. In many interesting applications, the domain dimensionality is such as to prevent state-of-the-art statistical learning techniques from delivering accurate models in reasonable time. This paper presents a hybrid random field model for pseudo-likelihood estimation in high-dimensional domains. A theoretical analysis proves that the class of pseudo-likelihood distributions representable by hybrid random fields strictly includes the class of joint probability distributions representable by Bayesian networks. In order to learn hybrid random fields from data, we develop the Markov Blanket Merging algorithm. Theoretical and experimental evidence shows that Markov Blanket Merging scales up very well to high-dimensional datasets. As compared to other widely used statistical learning techniques, Markov Blanket Merging delivers accurate results in a number of link prediction tasks, while achieving also significant improvements in terms of computational efficiency. Antonino Freno, Edmondo Trentin, Marco Gori |
KDD | 3 |
| 2009 | Users, Queries and Documents: A Unified Representation for Web MiningabstractThe collective feedback of the users of an Information Retrieval system has been proved to be useful in many tasks. A popular approach in the literature is to process the logs stored by Internet Service Providers (ISP), Intranet proxies or Web search engines to extract a query-document bi-partite graph. In this paper, we propose to use a richer data structure which is able to preserve most of the information available in the logs including query refinements, page visits and search activity. In particular, we represent the query refinements as separate transitions between the corresponding query nodes in the graph and we augment the graph by associating one node to each single user. Users are linked to the queries which they have issued and to the documents they have visited. The resulting data structure is a complete representation of the collective search activity performed by the users of a search engine or of an Intranet. The experimental results show that this more powerful representation can be successfully used to improve the quality of query clustering and to discover query suggestions. Michelangelo Diligenti, Marco Gori, Marco Maggini |
Web Intelligence | 2 |
| 2009 | A hybrid random field model for scalable statistical learning
Antonino Freno, Edmondo Trentin, Marco Gori |
Neural Networks | 3 |
| 2009 | Semantic-based regularization and Piaget's cognitive stages
Marco Gori |
Neural Networks | 1 |
| 2009 | The Graph Neural Network ModelabstractMany underlying relationships among data in several areas of science and engineering, e.g., computer vision, molecular chemistry, molecular biology, pattern recognition, and data mining, can be represented in terms of graphs. In this paper, we propose a new neural network model, called graph neural network (GNN) model, that extends existing neural network methods for processing the data represented in graph domains. This GNN model, which can directly process most of the practically useful types of graphs, e.g., acyclic, cyclic, directed, and undirected, implements a function tau(G,n) is an element of IR(m) that maps a graph G and one of its nodes n into an m-dimensional Euclidean space. A supervised learning algorithm is derived to estimate the parameters of the proposed GNN model. The computational cost of the proposed algorithm is also considered. Some experimental results are shown to validate the proposed learning algorithm, and to demonstrate its generalization capabilities. Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, Gabriele Monfardini |
IEEE Trans. Neural Networks | 2 |
| 2009 | Computational Capabilities of Graph Neural NetworksabstractIn this paper, we will consider the approximation properties of a recently introduced neural network model called graph neural network (GNN), which can be used to process-structured data inputs, e.g., acyclic graphs, cyclic graphs, and directed or undirected graphs. This class of neural networks implements a function tau(G,n) is an element of IR(m) that maps a graph G and one of its nodes n onto an m-dimensional Euclidean space. We characterize the functions that can be approximated by GNNs, in probability, up to any prescribed degree of precision. This set contains the maps that satisfy a property called preservation of the unfolding equivalence, and includes most of the practically useful functions on graphs; the only known exception is when the input graph contains particular patterns of symmetries when unfolding equivalence may not be preserved. The result can be considered an extension of the universal approximation property established for the classic feedforward neural networks (FNNs). Some experimental examples are used to show the computational capabilities of the proposed model. Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, Gabriele Monfardini |
IEEE Trans. Neural Networks | 2 |
| 2008 | A Fully Automatic Crossword GeneratorabstractThis paper presents a software system that is able to generate crosswords with no human intervention including definition generation and crossword compilation. In particular, the proposed system crawls relevant sources of the Web, extracts definitions from the downloaded pages using state-of-the-art natural language processing (NLP) techniques and, finally, attempts at compiling a crossword schema with the extracted definitions using a constrain satisfaction programming (CSP) solver. The crossword generator has relevant applications in entertainment, educational and rehabilitation contexts. Leonardo Rigutini, Michelangelo Diligenti, Marco Maggini, Marco Gori |
ICMLA | 4 |
| 2007 | Some Aspects of a Complexity Theory for Continuous Time Systems
Marco Gori, Klaus Meer |
CiE | 1 |
| 2007 | Learning in Hyperlinked Environments
Marco Gori |
ECIR | 1 |
| 2007 | An Adaptive Context-Based Algorithm for Term Weighting: Application to Single-Word Question Answering
Marco Ernandes, Giovanni Angelini, Marco Gori, Leonardo Rigutini, Franco Scarselli |
IJCAI | 3 |
| 2007 | ItemRank: A Random-Walk Based Scoring Algorithm for Recommender Engines
Marco Gori, Augusto Pucci |
IJCAI | 1 |
| 2006 | Adaptive Context-Based Term (Re)Weighting: An Experiment on Single-Word Question Answering
Marco Ernandes, Giovanni Angelini, Marco Gori, Leonardo Rigutini, Franco Scarselli |
ECAI | 3 |
| 2006 | Graph Neural Networks for Object Localization
Gabriele Monfardini, Vincenzo Di Massa, Franco Scarselli, Marco Gori |
ECAI | 4 |
| 2006 | A Comparison between Recursive Neural Networks and Graph Neural NetworksabstractRecursive neural networks (RNNs) and graph neural networks (GNNs) are two connectionist models that can directly process graphs. RNNs and GNNs exploit a similar processing framework, but they can be applied to different input domains. RNNs require the input graphs to be directed and acyclic, whereas GNNs can process any kind of graphs. The aim of this paper consists in understanding whether such a difference affects the behaviour of the models on a real application. An experimental comparison on an image classification problem is presented, showing that GNNs outperforms RNNs. Moreover the main differences between the models are also discussed w.r.t. their input domains, their approximation capabilities and their learning algorithms. Vincenzo Di Massa, Gabriele Monfardini, Lorenzo Sarti, Franco Scarselli, Marco Maggini, Marco Gori |
IJCNN | 6 |
| 2006 | Research Paper Recommender Systems: A Random-Walk Based ApproachabstractEvery day researchers from all over the world have to filter the huge mass of existing research papers with the crucial aim of finding out useful publications related to their current work. In this paper we propose a research paper recommending algorithm based on the citation graph and random-walker properties. The PaperRank algorithm is able to assign a preference score to a set of documents contained in a digital library and linked one each other by bibliographic references. A data set of papers extracted by ACM portal has been used for testing and very promising performances have been measured Marco Gori, Augusto Pucci |
Web Intelligence | 1 |
| 2006 | Inversion-based nonlinear adaptation of noisy acoustic parameters for a neural/HMM speech recognizer
Edmondo Trentin, Marco Gori |
Neurocomputing | 2 |
| 2006 | Recursive processing of cyclic graphsabstractRecursive neural networks are a powerful tool for processing structured data. According to the recursive learning paradigm, the input information consists of directed positional acyclic graphs (DPAGs). In fact, recursive networks are fed following the partial order defined by the links of the graph. Unfortunately, the hypothesis of processing DPAGs is sometimes too restrictive, being the nature of some real-world problems intrinsically cyclic. In this paper, a methodology is proposed, which allows us to process any cyclic directed graph. Therefore, the computational power of recursive networks is definitely established, also clarifying the underlying limitations of the model. Monica Bianchini, Marco Gori, Lorenzo Sarti, Franco Scarselli |
IEEE Trans. Neural Networks | 2 |
| 2005 | WebCrow: A Web-Based System for Crossword Solving
Marco Ernandes, Giovanni Angelini, Marco Gori |
AAAI | 3 |
| 2005 | A Neural Network Approach to Web Graph Processing
Ah Chung Tsoi, Franco Scarselli, Marco Gori, Markus Hagenbuchner, Sweah Liang Yong |
APWeb | 3 |
| 2005 | Learning Web Page Scores by Error Back-Propagation
Michelangelo Diligenti, Marco Gori, Marco Maggini |
IJCAI | 2 |
| 2005 | A new model for learning in graph domainsabstractIn several applications the information is naturally represented by graphs. Traditional approaches cope with graphical data structures using a preprocessing phase which transforms the graphs into a set of flat vectors. However, in this way, important topological information may be lost and the achieved results may heavily depend on the preprocessing stage. This paper presents a new neural model, called graph neural network (GNN), capable of directly processing graphs. GNNs extends recursive neural networks and can be applied on most of the practically useful kinds of graphs, including directed, undirected, labelled and cyclic graphs. A learning algorithm for GNNs is proposed and some experiments are discussed which assess the properties of the model. Marco Gori, Gabriele Monfardini, Franco Scarselli |
IJCNN | 1 |
| 2005 | Feature Normalization via ANN/HMM Inversion for Speech Recognition Under Noisy ConditionsabstractSpoken human-machine interaction in real-world environments requires acoustic models that are robust to changes in acoustic conditions, e.g. presence of noise. Unfortunately, the popular hidden Markov models (HMM) are not noise tolerant. One way to increase recognition performance rely on the acquisition of a small adaptation set of noisy utterances, that is used to estimate a normalization mapping between noisy and clean features to be fed into the acoustic model. In this research we develop a maximum-likelihood gradient-ascent training algorithm (instead of the usual least squares regression) for a neural normalization module to be combined with a hybrid connectionist/HMM recognizer. The algorithm is inspired by the so-called "inversion principle". Simulation results on a real-world speaker-independent continuous speech corpus of connected Italian digits, corrupted by additive noise, validate the approach: a small neural net (13 hidden neurons) trained over a single adaptation utterance for just one iteration yields 18.79% relative word error rate (WER) reduction over the bare hybrid, and a 65.10% relative WER reduction over the Gaussian-based HMM Edmondo Trentin, Marco Gori |
MMSP | 2 |
| 2005 | Graph Neural Networks for Ranking Web PagesabstractAn artificial neural network model, capable of processing general types of graph structured data, has recently been proposed. This paper applies the new model to the computation of customised page ranks problem in the World Wide Web. The class of customised page ranks that can be implemented in this way is very general and easy because the neural network model is learned by examples. Some preliminary experimental findings show that the model generalizes well over unseen Web pages, and hence, may be suitable for the task of page rank computation on a large Web graph. Franco Scarselli, Sweah Liang Yong, Marco Gori, Markus Hagenbuchner, Ah Chung Tsoi, Marco Maggini |
Web Intelligence | 3 |
| 2005 | The loading problem for recursive neural networks
Marco Gori, Alessandro Sperduti |
Neural Networks | 1 |
| 2005 | Exact and Approximate Graph Matching Using Random WalksabstractIn this paper, we propose a general framework for graph matching which is suitable for different problems of pattern recognition. The pattern representation we assume is at the same time highly structured, like for classic syntactic and structural approaches, and of subsymbolic nature with real-valued features, like for connectionist and statistic approaches. We show that random walk based models, inspired by Google's PageRank, give rise to a spectral theory that nicely enhances the graph topological features at node level. As a straightforward consequence, we derive a polynomial algorithm for the classic graph isomorphism problem, under the restriction of dealing with Markovian spectrally distinguishable graphs (MSD), a class of graphs that does not seem to be easily reducible to others proposed in the literature. The experimental results that we found on different test-beds of the TC-15 graph database show that the defined MSD class "almost always" covers the database, and that the proposed algorithm is significantly more efficient than top scoring VF algorithm on the same data. Most interestingly, the proposed approach is very well-suited for dealing with partial and approximate graph matching problems, derived for instance from image retrieval tasks. We consider the objects of the COIL-100 visual collection and provide a graph-based representation, whose node's labels contain appropriate visual features. We show that the adoption of classic bipartite graph matching algorithms offers a straightforward generalization of the algorithm given for graph isomorphism and, finally, we report very promising experimental results on the COIL-100 visual collection. Marco Gori, Marco Maggini, Lorenzo Sarti |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Artificial Neural Networks for Document Analysis and RecognitionabstractArtificial neural networks have been extensively applied to document analysis and recognition. Most efforts have been devoted to the recognition of isolated handwritten and printed characters with widely recognized successful results. However, many other document processing tasks, like preprocessing, layout analysis, character segmentation, word recognition, and signature verification, have been effectively faced with very promising results. This paper surveys the most significant problems in the area of offline document image processing, where connectionist-based approaches have been applied. Similarities and differences between approaches belonging to different categories are discussed. A particular emphasis is given on the crucial role of prior knowledge for the conception of both appropriate architectures and learning algorithms. Finally, the paper provides a critical analysis on the reviewed approaches and depicts the most promising research guidelines in the field. In particular, a second generation of connectionist-based models are foreseen which are based on appropriate graphical representations of the learning environment. Simone Marinai, Marco Gori, Giovanni Soda |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Preface
Simone Marinai, Marco Gori |
Pattern Recognit. Lett. | 2 |
| 2005 | Inside PageRankabstractAlthough the interest of a Web page is strictly related to its content and to the subjective readers' cultural background, a measure of the page authority can be provided that only depends on the topological structure of the Web. PageRank is a noticeable way to attach a score to Web pages on the basis of the Web connectivity. In this article, we look inside PageRank to disclose its fundamental properties concerning stability, complexity of computational scheme, and critical role of parameters involved in the computation. Moreover, we introduce a circuit analysis that allows us to understand the distribution of the page score, the way different Web communities interact each other, the role of dangling pages (pages with no outlinks), and the secrets for promotion of Web pages. Monica Bianchini, Marco Gori, Franco Scarselli |
ACM Trans. Internet Techn. | 2 |
| 2004 | Likely-Admissible and Sub-Symbolic Heuristics
Marco Ernandes, Marco Gori |
ECAI | 2 |
| 2004 | MumbleSearch Extraction of High Quality Web information for SMEabstractAlthough search engines are playing a crucial role for the retrieval of information from the Web, they cannot guarantee the quality required for most relevant business activities as well as for many top-level research projects. In this paper we present MumbleSearch, a Web Content Monitor which is especially conceived to extract and organize topic-based information with emphasis on quality requirements. We present the architecture of the software platform and its deployment for a real-world application, involving Italian Small and Medium Enterprises (SME). Nicola Baldini, Marco Gori, Marco Maggini |
Web Intelligence | 2 |
| 2004 | Neural computation, social networks, and topological spectra
Michelangelo Diligenti, Marco Gori, Marco Maggini |
Theor. Comput. Sci. | 2 |
| 2004 | A Unified Probabilistic Framework for Web Page Scoring SystemsabstractThe definition of efficient page ranking algorithms is becoming an important issue in the design of the query interface of Web search engines. Information flooding is a common experience especially when broad topic queries are issued. Queries containing only one or two keywords usually match a huge number of documents, while users can only afford to visit the first positions of the returned list, which do not necessarily refer to the most appropriate answers. Some successful approaches to page ranking in a hyperlinked environment, like the Web, are based on link analysis. We propose a general probabilistic framework for Web page scoring systems (WPSS), which incorporates and extends many of the relevant models proposed in the literature. In particular, we introduce scoring systems for both generic (horizontal) and focused (vertical) search engines. Whereas horizontal scoring algorithms are only based on the topology of the Web graph, vertical ranking also takes the page contents into account and are the base for focused and user adapted search interfaces. Experimental results are reported to show the properties of some of the proposed scoring systems with special emphasis on vertical search. Michelangelo Diligenti, Marco Gori, Marco Maggini |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2004 | Encoding nondeterministic fuzzy tree automata into recursive neural networksabstractFuzzy neural systems have been a subject of great interest in the last few years, due to their abilities to facilitate the exchange of information between symbolic and subsymbolic domains. However, the models in the literature are not able to deal with structured organization of information, that is typically required by symbolic processing. In many application domains, the patterns are not only structured, but a fuzziness degree is attached to each subsymbolic pattern primitive. The purpose of this paper is to show how recursive neural networks, properly conceived for dealing with structured information, can represent nondeterministic fuzzy frontier-to-root tree automata. Whereas available prior knowledge expressed in terms of fuzzy state transition rules are injected into a recursive network, unknown rules are supposed to be filled in by data-driven learning. We also prove the stability of the encoding algorithm, extending previous results on the injection of fuzzy finite-state dynamics in high-order recurrent networks. Marco Gori, Alfredo Petrosino |
IEEE Trans. Neural Networks | 1 |
| 2003 | An introduction to learning in web domains
Michelangelo Diligenti, Marco Gori, Marco Maggini, Franco Scarselli, Ah Chung Tsoi |
ESANN | 2 |
| 2003 | A Learning Algorithm for Web Page Scoring Systems
Michelangelo Diligenti, Marco Gori, Marco Maggini |
IJCAI | 2 |
| 2003 | A recursive neural network model for processing directed acyclic graphs with labeled edgesabstractThe recursive paradigm extends the neural network processing and learning algorithms to deal with structured inputs. In particular, recursive neural network (RNN) models have been proposed to process information coded as directed positional acyclic graphs (DPAGs) whose maximum node outdegree is known a priori. Unfortunately, the hypothesis of processing DPAGs having a given maximum node outdegree is sometimes too restrictive, being the nature of some real-world problems intrinsically disordered. In many applications the node outdegrees can vary considerably among the nodes in the graph, it may be unnatural to define a position for each child of a given node, and it may be necessary to prune some edges to reduce the number of the network parameters, which is proportional to the maximum node outdegree. In this paper, we proposed a new recursive neural network model which allows us to process directed acyclic graphs (DAGs) with labeled edges, relaxing the positional constraint and the correlated maximum outdegree limit. The effectiveness of the new scheme is experimentally tested on an image classification task. The results show that the new RNN model outperforms the standard RNN architecture, also allowing us to use a smaller number of free parameters. Marco Gori, Marco Maggini, Lorenzo Sarti |
IJCNN | 1 |
| 2003 | Evaluation on the Aurora 2 database of acoustic models that are less noise-sensitiveabstractThe Aurora 2 database may be used as a benchmark for evaluation of algorithms under noisy conditions. In particular, the clean training/noisy test mode is aimed at evaluating models that are trained on clean data only without further adjustments on the noisy data, i.e. under severe mismatch between the training and test conditions. While several researchers proposed techniques at the front-end level to improve recognition performance over the reference hideen Markov model (HMM) baseline, investigations at the back-end level are sought. In this respect, the goal is to develop acoustic models that are intrinsically less noise sensitive. This paper presents the word accuracy yielded by a non-parametric HMM with connectionist estimates of the emission probabilities, i.e. a neural network is applied instead of the usual parametric (Gaussian mixture) probability densities. A regularization technique, relying on a maximum-likelihood parameter grouping algorithm, is explicitly introduced to increase the generalization capability of the model and, in turn, its noise-robustness. Results show that a 15,43% relative word error rate reduction w.r.t. the Gaussianmixture HMM is obtained by averaging over the different noises and SNRs of Aurora 2 test set A. Edmondo Trentin, Marco Matassoni, Marco Gori |
INTERSPEECH | 3 |
| 2003 | PageRank and Web CommunitiesabstractThe definition of the ordering of the Web pages, returned on a given query, is a crucial topic, which gives rise to the notion of Web visibility. A fundamental contribution towards the conception of appropriate ordering criteria has been given by means of the introduction of PageRank, which takes into account only the hyper-linked structure of the Web, regardless of the content of the pages. We introduce a circuit analysis which allows us to understand the distribution of PageRank, and show some basic results for understanding the way it migrates amongst communities. In particular, we highlight some topological properties which suggest methods for the promotion of Web communities. These results confirm the importance and the effectiveness of PageRank for discovering relevant information but, at the same time, point out its vulnerability to spamming. Monica Bianchini, Marco Gori, Franco Scarselli |
Web Intelligence | 2 |
| 2003 | Detecting Near-Replicas on the Web by Content and Hyperlink AnalysisabstractThe presence of near-replicas of documents is very common on the Web. Documents may be replicated completely or partially for different reasons (versions, mirrors, etc.), or the same resource can be associated to different URLs (dynamically generated pages, etc.). Whilst replication can improve information accessibility by the users, the presence of near-replicated documents can hinder the effectiveness of search engines (for example, decreasing the coverage). We propose a method to detect similar pages, in particular replicas and near-replicas, which is based on a pair of signatures. The first signature is obtained by a random projection of the bag-of-words vector representing the page contents. The second signature is computed by a recursive equation which exploits the connectivity among the Web pages to code the context of each page. The accuracy of the proposed approach is analyzed and validated by experimental results which show that on the given dataset near-replicas can be detected with a precision-recall of 93%. Ernesto Di Iorio, Michelangelo Diligenti, Marco Gori, Marco Maggini, Augusto Pucci |
Web Intelligence | 3 |
| 2003 | Analysis and understanding of multi-class invoices
Francesca Cesarini, Enrico Francesconi, Marco Gori, Giovanni Soda |
Int. J. Document Anal. Recognit. | 3 |
| 2003 | Hidden Tree Markov Models for Document Image ClassificationabstractClassification is an important problem in image document processing and is often a preliminary step toward recognition, understanding, and information extraction. In this paper, the problem is formulated in the framework of concept learning and each category corresponds to the set of image documents with similar physical structure. We propose a solution based on two algorithmic ideas. First, we obtain a structured representation of images based on labeled XY-trees (this representation informs the learner about important relationships between image subconstituents). Second, we propose a probabilistic architecture that extends hidden Markov models for learning probability distributions defined on spaces of labeled trees. Finally, a successful application of this method to the categorization of commercial invoices is presented. Michelangelo Diligenti, Paolo Frasconi, Marco Gori |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2003 | Edge-backpropagation for noisy logo recognition
Marco Gori, Marco Maggini, Simone Marinai, Jianqing Sheng, Giovanni Soda |
Pattern Recognit. | 1 |
| 2003 | Using attributed plex grammars for the generation of image and graph databases
Markus Hagenbuchner, Marco Gori, Horst Bunke, Ah Chung Tsoi, Christophe Irniger |
Pattern Recognit. Lett. | 2 |
| 2003 | Similarity learning for graph-based image representations
Ciro de Mauro, Michelangelo Diligenti, Marco Gori, Marco Maggini |
Pattern Recognit. Lett. | 3 |
| 2003 | A Hybrid Model for the Prediction of the Linguistic Origin of SurnamesabstractThe prediction of the linguistic origin of surnames is a basic functionality required in the design of high-quality multilanguage speech synthesizers. The assignment of a given string representing a surname to a specific language is typically based on a set of rules which can hardly be written in an explicit form. The approach we propose faces this problem combining a rule-based system with a module based on evidential reasoning and a module based on neural networks. The resulting hybrid system combines the different sources of information, merging both knowledge from experts on linguistics and knowledge automatically acquired using learning from examples. The system has been validated on a large database containing surnames belonging to four different languages, showing its effectiveness for real-world applications. Patrizia Bonaventura, Marco Gori, Marco Maggini, Franco Scarselli, Jianqing Sheng |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2003 | Robust combination of neural networks and hidden Markov models for speech recognitionabstractAcoustic modeling in state-of-the-art speech recognition systems usually relies on hidden Markov models (HMMs) with Gaussian emission densities. HMMs suffer from intrinsic limitations, mainly due to their arbitrary parametric assumption. Artificial neural networks (ANNs) appear to be a promising alternative in this respect, but they historically failed as a general solution to the acoustic modeling problem. This paper introduces algorithms based on a gradient-ascent technique for global training of a hybrid ANN/HMM system, in which the ANN is trained for estimating the emission probabilities of the states of the HMM. The approach is related to the major hybrid systems proposed by Bourlard and Morgan and by Bengio, with the aim of combining their benefits within a unified framework and to overcome their limitations. Several viable solutions to the "divergence problem"-that may arise when training is accomplished over the maximum-likelihood (ML) criterion-are proposed. Experimental results in speaker-independent, continuous speech recognition over Italian digit-strings validate the novel hybrid framework, allowing for improved recognition performance over HMMs with mixtures of Gaussian components, as well as over Bourlard and Morgan's paradigm. In particular, it is shown that the maximum a posteriori (MAP) version of the algorithm yields a 46.34% relative word error rate reduction with respect to standard HMMs. Edmondo Trentin, Marco Gori |
IEEE Trans. Neural Networks | 2 |
| 2002 | Recursive Neural Networks Applied to Discourse Representation Theory
Antonella Bua, Marco Gori, Fabrizio Santini |
ICANN | 2 |
| 2002 | Recognition of Common Areas in a Web Page Using Visual Information: a possible application in a page classificationabstractExtracting and processing information from Web pages is an important task in many areas like constructing search engines, information retrieval, and data mining from the Web. A common approach in the extraction process is to represent a page as a "bag of words" and then to perform additional processing on such a flat representation. We propose a new, hierarchical representation that includes browser screen coordinates for every HTML object in a page. Using visual information one is able to define heuristics for the recognition of common page areas such as header, left and right menu, footer and center of a page. We show in initial experiments that using our heuristics defined objects are recognized properly in 73% of cases. Finally, we show that a Naive Bayes classifier, taking into account the proposed representation, clearly outperforms the same classifier using only information about the content of documents. Milos Kovacevic, Michelangelo Diligenti, Marco Gori, Veljko M. Milutinovic |
ICDM | 3 |
| 2002 | Web page scoring systems for horizontal and vertical searchabstractPage ranking is a fundamental step towards the construction of effective search engines for both generic (horizontal) and focused (vertical) search. Ranking schemes for horizontal search like the PageRank algorithm used by Google operate on the topology of the graph, regardless of the page content. On the other hand, the recent development of vertical portals (vortals) makes it useful to adopt scoring systems focussed on the topic and taking the page content into account.In this paper, we propose a general framework for Web Page Scoring Systems (WPSS) which incorporates and extends many of the relevant models proposed in the literature. Finally, experimental results are given to assess the features of the proposed scoring systems with special emphasis on vertical search. Michelangelo Diligenti, Marco Gori, Marco Maggini |
WWW | 2 |
| 2001 | Searching the Web: learning based techniques
Michelangelo Diligenti, Marco Gori, Marco Maggini, Franco Scarselli |
ESANN | 2 |
| 2001 | Continuous Speech Recognition with a Robust Connectionist/Markovian Hybrid Model
Edmondo Trentin, Marco Gori |
ICANN | 2 |
| 2001 | Classification of HTML Documents by Hidden Tree-Markov ModelsabstractContent-based search and organization of Web documents poses new issues in information retrieval. We propose a novel approach for the classification of HTML documents based on a structured representation of their contents which are split into logical contexts (paragraphs, sections, anchors, etc.). The classification is performed using Hidden Tree-Markov Models (HTMMs), an extension of Hidden Markov Models for processing structured objects. We report some promising experimental results showing that the use of the structured representation improves the classification accuracy in most of the cases. Michelangelo Diligenti, Marco Gori, Marco Maggini, Franco Scarselli |
ICDAR | 2 |
| 2001 | APEX An Adaptive Visual Information Retrieval SystemabstractGiven a user's visual query, most visual information retrieval (VIR) systems rank the images in the database according to a predefined measure of similarity and return the most similar ones. We propose an adaptive VIR system that uses a retrieval process based on relevance feedback in order to learn the similarity criterion from the user. Our system is based on a structured representation of the image which is then processed by a recursive neural network. The search algorithm refines its response trying to minimize the number of steps required to find the target image. Ciro de Mauro, Marco Gori, Marco Maggini |
ICDAR | 2 |
| 2001 | Toward noise-tolerant acoustic models
Edmondo Trentin, Marco Gori |
INTERSPEECH | 2 |
| 2001 | Automatic document classification and indexing in high-volume applications
Enrico Appiani, Francesca Cesarini, Anna Maria Colla, Michelangelo Diligenti, Marco Gori, Simone Marinai, Giovanni Soda |
Int. J. Document Anal. Recognit. | 5 |
| 2001 | A serial combination of connectionist-based classifiers for OCR
Enrico Francesconi, Marco Gori, Simone Marinai, Giovanni Soda |
Int. J. Document Anal. Recognit. | 2 |
| 2001 | A survey of hybrid ANN/HMM models for automatic speech recognition
Edmondo Trentin, Marco Gori |
Neurocomputing | 2 |
| 2001 | Adaptive graphical pattern recognition for the classification of company logos
Michelangelo Diligenti, Marco Gori, Marco Maggini, Enrico Martinelli |
Pattern Recognit. | 2 |
| 2001 | Optimal Algorithms for Well-Conditioned Nonlinear Systems of EquationsabstractWe propose solving nonlinear systems of equations by function optimization and we give an optimal algorithm which relies on a special canonical form of gradient descent. The algorithm can be applied under certain assumptions on the function to be optimized, that is, an upper bound must exist for the norm of the Hessian, whereas the norm of the gradient must be lower bounded. Due to its intrinsic structure, the algorithm looks particularly appealing for a parallel implementation. As a particular case, more specific results are given for linear systems. We prove that reaching a solution with a degree of precision /spl epsiv/ takes /spl Theta/(n/sup 2/k/sup 2/ log /sup k///sub /spl epsiv//), k being the condition number of A and n the problem dimension. Related results hold for systems of quadratic equations for which an estimation for the requested bounds can be devised. Finally, we report numerical results in order to establish the actual computational burden of the proposed method and to assess its performances with respect to classical algorithms for solving linear and quadratic equations. Monica Bianchini, Stefano Fanelli, Marco Gori |
IEEE Trans. Computers | 3 |
| 2001 | Guest Editors' Introduction: Special Section on Connectionist Models for Learning in Structured DomainsabstractGuest Editors' Introduction to the Special Section on Connectionist Models for Learning in Structured Domains Paolo Frasconi, Marco Gori, Alessandro Sperduti |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2001 | Theoretical properties of recursive neural networks with linear neuronsabstractRecursive neural networks are a powerful tool for processing structured data, thus filling the gap between connectionism, which is usually related to poorly organized data, and a great variety of real-world problems, where the information is naturally encoded in the relationships among the basic entities. In this paper, some theoretical results about linear recursive neural networks are presented that allow one to establish conditions on their dynamical properties and their capability to encode and classify structured information. A lot of the limitations of the linear model, intrinsically related to recursive processing, are inherited by the general model, thus establishing their computational capabilities and range of applicability. As a byproduct of our study some connections with the classical linear system theory are given where the processing is extended from sequences to graphs. Monica Bianchini, Marco Gori |
IEEE Trans. Neural Networks | 2 |
| 2001 | Processing directed acyclic graphs with recursive neural networksabstractRecursive neural networks are conceived for processing graphs and extend the well-known recurrent model for processing sequences. In Frasconi et al. (1998), recursive neural networks can deal only with directed ordered acyclic graphs (DOAGs), in which the children of any given node are ordered. While this assumption is reasonable in some applications, it introduces unnecessary constraints in others. In this paper, it is shown that the constraint on the ordering can be relaxed by using an appropriate weight sharing, that guarantees the independence of the network output with respect to the permutations of the arcs leaving from each node. The method can be used with graphs having low connectivity and, in particular, few outcoming arcs. Some theoretical properties of the proposed architecture are given. They guarantee that the approximation capabilities are maintained, despite the weight sharing. Monica Bianchini, Marco Gori, Franco Scarselli |
IEEE Trans. Neural Networks | 2 |
| 2000 | Learning Efficiently with Neural Networks: A Theoretical Comparison between Structured and Flat Representations
Marco Gori, Paolo Frasconi, Alessandro Sperduti |
ECAI | 1 |
| 2000 | Computational capabilities of linear recursive networksabstractRecursive neural networks are a new connectionist model introduced for processing graphs. Linear recursive networks are a special subclass where the neurons have linear activation functions. The approximation properties of recursive networks are tightly connected to the possibility of distinguishing the patterns by generating a different internal encoding for each input of the domain. In this paper, it is shown that, even if linear recursive networks can distinguish the patterns of any finite set of trees, such a result requires a prohibitive memory consumption. However, it is also proved that the problem disappears when the domain is restricted to set of trees belonging to special sub-classes. Monica Bianchini, Marco Gori, Franco Scarselli |
KES | 2 |
| 2000 | Focused Crawling Using Context Graphs
Michelangelo Diligenti, Frans Coetzee, Steve Lawrence, C. Lee Giles, Marco Gori |
VLDB | 5 |
| 2000 | Competitive radial basis functions training for phone classification
Piero Cosi, Paolo Frasconi, Marco Gori, Luca Lastrucci, Giovanni Soda |
Neurocomputing | 3 |
| 2000 | Using Physical and Logical Constraints for Invoice Understanding
Francesca Cesarini, Enrico Francesconi, Marco Gori |
Pattern Anal. Appl. | 3 |
| 1999 | Neural learning of approximate simple regular languages
Mikel L. Forcada, Antonio M. Corbí-Bellot, Marco Gori, Marco Maggini |
ESANN | 3 |
| 1999 | Learning in structured domains
Marco Gori |
ESANN | 1 |
| 1999 | A Two Level Knowledge Approach for Understanding Documents of a Multi-Class DomainabstractIn this paper an architecture for understanding documents of a domain that can be grouped into classes is shown. Documents are grouped with respect to the physical structure. The architecture is based on two knowledge descriptions of the domain: one is independent from the classes and one related to the classes. Such knowledge levels are used to understand the documents of the domain. The understanding phase is described in relation with the phases of analysis and classification of such documents. Francesca Cesarini, Enrico Francesconi, Marco Gori, Giovanni Soda |
ICDAR | 3 |
| 1999 | Structured Document Segmentation and Representation by the Modified X-Y treeabstractWe describe a top-down approach to the segmentation and representation of documents containing tabular structures. Examples of these documents are invoices and technical papers with tables. The segmentation is based on an extension of X-Y trees, where the regions are split by means of cuts along separators (e.g. lines), in addition to cuts along white spaces. The leaves describe regions containing homogeneous information and cutting separators. Adjacency links among leaves of the tree describe local relationships between corresponding regions. Francesca Cesarini, Marco Gori, Simone Marinai, Giovanni Soda |
ICDAR | 2 |
| 1999 | A connectionist-based model for predicting the linguistic origin of surnamesabstractThis paper describes an application to the prediction of the linguistic origin of surnames. The problem can be stated as follows: "Given an input string representing a surname, decide which language the surname belongs to". We present an approach that integrates methodologies from rule-based systems, evidential reasoning, and neural networks. Our hybrid solution maximizes the used information and allows one to deal with aspects of the problem that could have not been solved otherwise. In fact, our predictor exploits both knowledge from experts on languages (by means of a rule-based system) and knowledge automatically acquired by examples (through statistical analysis and a neural network). Patrizia Bonaventura, Marco Gori, Franco Scarselli, Pierluigi Salza, Jianqing Sheng |
IJCNN | 2 |
| 1999 | Recurrent neural networks can learn simple, approximate regular languagesabstractA number of researchers have shown that discrete-time recurrent neural networks (DTRNN) are capable of inferring deterministic finite automata from sets of example and counterexample strings; however, discrete algorithmic methods are much better at this task and clearly outperform DTRNN in terms of space and time complexity. We show how DTRNN may be used to learn not the exact language that explains the whole learning set but an approximate and much simpler language that explains a great majority of the examples by using simpler rules. This is accomplished by gradually varying the error function in such a way that the DTRNN is eventually allowed to classify clearly but incorrectly those strings that it has found to be difficult to learn, which are treated as exceptions. The results show that in this way, the DTRNN usually manages to learn a simplified approximate language. Mikel L. Forcada, Antonio M. Corbí-Bellot, Marco Gori, Marco Maggini |
IJCNN | 3 |
| 1999 | Feature extraction from data structures with unsupervised recursive neural networksabstractIn the case of static data of high dimension it is often useful to reduce the dimensionality before performing pattern recognition and learning tasks. One of the main reasons for this is that models for lower-dimensional data usually have fewer parameters to be determined. The problem of finding fixed-length vector representations for labelled directed ordered acyclic graphs (DOAGs) can be regarded as a feature extraction problem in which the dimensionality of the input space is infinite. We address the fundamental problem of finding fixed-length vector representations for DOAGs in an unsupervised way using a maximum entropy approach. Some preliminary experiments on image retrieval are reported. Christoph Goller, Marco Gori, Marco Maggini |
IJCNN | 2 |
| 1999 | Data Categorization Using Decision TrellisesabstractWe introduce a probabilistic graphical model for supervised learning on databases with categorical attributes. The proposed belief network contains hidden variables that play a role similar to nodes in decision trees and each of their states either corresponds to a class label or to a single attribute test. As a major difference with respect to decision trees, the selection of the attribute to be tested is probabilistic. Thus, the model can be used to assess the probability that a tuple belongs to some class, given the predictive attributes. Unfolding the network along the hidden states dimension yields a trellis structure having a signal flow similar to second order connectionist networks. The network encodes context specific probabilistic independencies to reduce parametric complexity. We present a custom tailored inference algorithm and derive a learning procedure based on the expectation-maximization algorithm. We propose decision trellises as an alternative to decision trees in the context of tuple categorization in databases, which is an important step for building data mining systems. Preliminary experiments on standard machine learning databases are reported, comparing the classification accuracy of decision trellises and decision trees induced by C4.5. In particular, we show that the proposed model can offer significant advantages for sparse databases in which many predictive attributes are missing. Paolo Frasconi, Marco Gori, Giovanni Soda |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1999 | On the implementation of frontier-to-root tree automata in recursive neural networksabstractIn this paper we explore the node complexity of recursive neural network implementations of frontier-to-root tree automata (FRA). Specifically, we show that an FRAO (Mealy version) with m states, l input-output labels, and maximum rank N can be implemented by a recursive neural network with O(radical(log l+log m)lm(N)/log l+N log m) units and four computational layers, i.e., without counting the input layer. A lower bound is derived which is tight when no restrictions are placed on the number of layers. Moreover, we present a construction with three computational layers having node complexity of O((log l + log m)radical lmN) and O((log l + log m) lmN) connections. A construction with two computational layers is given that implements any given FRAO with a node complexity of O(lmN) and O((log l+N log m)lmN) connections. As a corollary we also get a new upper bound for the implementation of finite-state automata (FSA) into recurrent neural networks with three computational layers. Marco Gori, Andreas Küchler, Alessandro Sperduti |
IEEE Trans. Neural Networks | 1 |
| 1998 | An Application of ELISA to Perfect Hashing with Deterministic Ordering
Monica Bianchini, Stefano Fanelli, Marco Gori |
ICONIP | 3 |
| 1998 | INFORMys: A Flexible Invoice-Like Form-Reader SystemabstractWe describe a flexible form-reader system capable of extracting textual information from accounting documents, like invoices and bills of service companies. In this kind of document, the extraction of some information fields cannot take place without having detected the corresponding instruction fields, which are only constrained to range in given domains. We propose modeling the document's layout by means of attributed relational graphs, which turn out to be very effective for form registration, as well as for performing a focused search for instruction fields. This search is carried out by means of a hybrid model, where proper algorithms, based on morphological operations and connected components, are integrated with connectionist models. Experimental results are given in order to assess the actual performance of the system. Francesca Cesarini, Marco Gori, Simone Marinai, Giovanni Soda |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Are Multilayer Perceptrons Adequate for Pattern Recognition and Verification?abstractDiscusses the ability of multilayer perceptrons (MLPs) to model the probability distribution of data in typical pattern recognition and verification problems. It is proven that multilayer perceptrons with sigmoidal units and a number of hidden units less or equal than the number of inputs are unable to model patterns distributed in typical clusters, since these networks draw open separation surfaces in the pattern space. When using more hidden units than inputs, the separation surfaces can be closed but, unfortunately it is proven that determining whether or not a MLP draws closed separation surfaces in the pattern space is NP-hard. The major conclusion of the paper is somewhat opposite to what is believed and reported in many application papers: MLPs are definitely not adequate for applications of pattern recognition requiring a reliable rejection and, especially, they are not adequate for pattern verification tasks. Marco Gori, Franco Scarselli |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | A general framework for adaptive processing of data structuresabstractA structured organization of information is typically required by symbolic processing. On the other hand, most connectionist models assume that data are organized according to relatively poor structures, like arrays or sequences. The framework described in this paper is an attempt to unify adaptive models like artificial neural nets and belief nets for the problem of processing structured information. In particular, relations between data variables are expressed by directed acyclic graphs, where both numerical and categorical values coexist. The general framework proposed in this paper can be regarded as an extension of both recurrent neural networks and hidden Markov models to the case of acyclic graphs. In particular we study the supervised learning problem as the problem of learning transductions from an input structured space to an output structured space, where transductions are assumed to admit a recursive hidden statespace representation. We introduce a graphical formalism for representing this class of adaptive transductions by means of recursive networks, i.e., cyclic graphs where nodes are labeled by variables and edges are labeled by generalized delay elements. This representation makes it possible to incorporate the symbolic and subsymbolic nature of data. Structures are processed by unfolding the recursive network into an acyclic graph called encoding network. In so doing, inference and learning algorithms can be easily inherited from the corresponding algorithms for artificial neural networks or probabilistic graphical model. Paolo Frasconi, Marco Gori, Alessandro Sperduti |
IEEE Trans. Neural Networks | 2 |
| 1998 | Inductive inference from noisy examples using the hybrid finite state filterabstractRecurrent neural networks processing symbolic strings can be regarded as adaptive neural parsers. Given a set of positive and negative examples, picked up from a given language, adaptive neural parsers can effectively be trained to infer the language grammar. In this paper we use adaptive neural parsers to face the problem of inferring grammars from examples that are corrupted by a kind of noise that simply changes their membership.We propose a training algorithm, referred to as hybrid finite state filter (HFF), which is based on a parsimony principle that penalizes the development of complex rules.We report very promising experimental results showing that the proposed inductive inference scheme is indeed capable of capturing rules, while removing noise. Marco Gori, Marco Maggini, Enrico Martinelli, Giovanni Soda |
IEEE Trans. Neural Networks | 1 |
| 1998 | On the closure of the set of functions that can be realized by a given multilayer perceptronabstractGiven a multilayer perceptron (MLP) with a fixed architecture, there are functions that can be approximated up to any degree of accuracy, without having to increase the number of the hidden nodes. Those functions belong to the closure F of the set F of the maps realizable by the MLP. In this paper, we give a list of maps with this property. In particular, it is proven that 1) rational functions belongs to F for networks with inverse tangent activation function and 2) products of polynomials and exponentials belongs to F for networks with sigmoid activation function. Moreover, for a restricted class of MLP's, we prove that the list is complete and give an analytic definition of F. Marco Gori, Franco Scarselli, Ah Chung Tsoi |
IEEE Trans. Neural Networks | 1 |
| 1998 | Comments on local minima free conditions in multilayer perceptronsabstractIn this letter we point out that multilayer neural networks (MLP's) with either sigmoidal units or radial basis functions can be given a canonical form with positive interunits weights, which does not restrict the well-known MLP universal computational capabilities. We give some results on the local minima of the error function using this canonical form. In particular, we prove that the local minima free conditions established in previous works can be relaxed significantly. Marco Gori, Ah Chung Tsoi |
IEEE Trans. Neural Networks | 1 |
| 1997 | A Neural-Based Architecture for Spot-Noisy Logo RecognitionabstractMuch attention has recently been paid to the recognition of graphical objects, such as company logos and trademarks. Recognizing these objects facilitates the recognition of document classes. Some promising results have been achieved by using autoassociator-based artificial neural networks (AANN) in the presence of homogeneously distributed noise. However, the performance drops significantly when dealing with spot-noisy logos, where strips or blobs produce a partial obstruction of the pictures. We propose a new approach for training AANNs especially conceived for dealing with spot noise. The basic idea is to introduce new metrics for assessing the reproduction error in AANNs. The proposed algorithm, referred to as spot-backpropagation (S-BP), is significantly more robust with respect to spot-noise than classical Euclidean norm-based backpropagation (BP). Our experimental results are based on a database of 88 real logos that are artificially corrupted by spot-noise. Francesca Cesarini, Enrico Francesconi, Marco Gori, Simone Marinai, Jianqing Sheng, Giovanni Soda |
ICDAR | 3 |
| 1997 | Rectangle Labelling for an Invoice Understanding SystemabstractWe present a method for the logical labelling of physical rectangles, extracted from invoices, based on a conceptual model which describes, as generally as possible, the invoice universe. This general knowledge is used in the semi automatic construction of a model for each class of invoices. Once the model is constructed, it can be applied to understand an invoice instance, whose class is univocally identified by its logo. This approach is used to design a flexible system which is able to learn, from a nucleus of general knowledge, a monotonic set of specific knowledge for each class of invoices (document models), in terms of physical coordinates for each rectangle and related semantic label. Francesca Cesarini, Enrico Francesconi, Marco Gori, Simone Marinai, Jianqing Sheng, Giovanni Soda |
ICDAR | 3 |
| 1997 | Solving Linear Systems by a Neural Network Canonical Form of Efficient Gradient Descent
Monica Bianchini, Stefano Fanelli, Marco Gori, Marco Protasi |
ICONIP (1) | 3 |
| 1997 | On the Efficient Classification of Data Structures by Neural Networks
Paolo Frasconi, Marco Gori, Alessandro Sperduti |
IJCAI | 2 |
| 1997 | Terminal attractor algorithms: A critical analysis
Monica Bianchini, Stefano Fanelli, Marco Gori, Marco Maggini |
Neurocomputing | 3 |
| 1997 | Links between LVQ and Backpropagation
Paolo Frasconi, Marco Gori, Giovanni Soda |
Pattern Recognit. Lett. | 2 |
| 1996 | Optimal learning in artificial neural networks: A review of theoretical results
Monica Bianchini, Marco Gori |
Neurocomputing | 2 |
| 1996 | Representation of Finite State Automata in Recurrent Radial Basis Function Networks
Paolo Frasconi, Marco Gori, Marco Maggini, Giovanni Soda |
Mach. Learn. | 2 |
| 1996 | Autoassociator-based models for speaker verification
Marco Gori, Luca Lastrucci, Giovanni Soda |
Pattern Recognit. Lett. | 1 |
| 1996 | Computational capabilities of local-feedback recurrent networks acting as finite-state machinesabstractIn this paper we explore the expressive power of recurrent networks with local feedback connections for symbolic data streams. We rely on the analysis of the maximal set of strings that can be shattered by the concept class associated to these networks (i.e. strings that can be arbitrarily classified as positive or negative), and find that their expressive power is inherently limited, since there are sets of strings that cannot be shattered, regardless of the number of hidden units. Although the analysis holds for networks with hard threshold units, we claim that the incremental computational capabilities gained when using sigmoidal units are severely paid in terms of robustness of the corresponding representation. Paolo Frasconi, Marco Gori |
IEEE Trans. Neural Networks | 2 |
| 1996 | A neural network-based model for paper currency recognition and verificationabstractThis paper describes the neural-based recognition and verification techniques used in a banknote machine, recently implemented for accepting paper currency of different countries. The perception mechanism is based on low-cost optoelectronic devices which produce a signal associated with the light refracted by the banknotes. The classification and verification steps are carried out by a society of multilayer perceptrons whose operation is properly scheduled by an external controlling algorithm, which guarantees real-time implementation on a standard microcontroller-based platform. The verification relies mainly on the property of autoassociators to generate closed separation surfaces in the pattern space. The experimental results are very interesting, particularly when considering that the recognition and verification steps are based on low-cost sensors. Angelo Frosini, Marco Gori, Paolo Priami |
IEEE Trans. Neural Networks | 2 |
| 1996 | Optimal convergence of on-line backpropagationabstractMany researchers are quite skeptical about the actual behavior of neural network learning algorithms like backpropagation. One of the major problems is with the lack of clear theoretical results on optimal convergence, particularly for pattern mode algorithms. In this paper, we prove the companion of Rosenblatt's PC (perceptron convergence) theorem for feedforward networks (1960), stating that pattern mode backpropagation converges to an optimal solution for linearly separable patterns. Marco Gori, Marco Maggini |
IEEE Trans. Neural Networks | 1 |
| 1995 | Data Extraction from Form Images
Francesca Cesarini, Marco Gori, Simone Marinai, Giovanni Soda |
DEXA | 2 |
| 1995 | A system for data extraction from forms of known classabstractIn this paper, we describe a flexible and efficient system for processing forms of a known class. The model is based on attributed relational graphs and the system performs form registration and location of information fields using algorithms based on the hypothesize-and-verify paradigm. A special emphasis has been placed at the low level, where an autoassociator-based connectionist model has exhibited successful results in finding the instruction fields in very noisy forms. Francesca Cesarini, Marco Gori, Simone Marinai, Giovanni Soda |
ICDAR | 2 |
| 1995 | Recurrent neural networks and prior knowledge for sequence processing: a constrained nondeterministic approach
Paolo Frasconi, Marco Gori, Giovanni Soda |
Knowl. Based Syst. | 2 |
| 1995 | Unified Integration of Explicit Knowledge and Learning by Example in Recurrent NetworksabstractProposes a novel unified approach for integrating explicit knowledge and learning by example in recurrent networks. The explicit knowledge is represented by automaton rules, which are directly injected into the connections of a network. This can be accomplished by using a technique based on linear programming, instead of learning from random initial weights. Learning is conceived as a refinement process and is mainly responsible for uncertain information management. We present preliminary results for problems of automatic speech recognition.> Paolo Frasconi, Marco Gori, Marco Maggini, Giovanni Soda |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1995 | Learning in multilayered networks used as autoassociatorsabstractGradient descent learning algorithms may get stuck in local minima, thus making the learning suboptimal. In this paper, we focus attention on multilayered networks used as autoassociators and show some relationships with classical linear autoassociators. In addition, by using the theoretical framework of our previous research, we derive a condition which is met at the end of the learning process and show that this condition has a very intriguing geometrical meaning in the pattern space. Monica Bianchini, Paolo Frasconi, Marco Gori |
IEEE Trans. Neural Networks | 3 |
| 1995 | Learning without local minima in radial basis function networksabstractLearning from examples plays a central role in artificial neural networks. The success of many learning schemes is not guaranteed, however, since algorithms like backpropagation may get stuck in local minima, thus providing suboptimal solutions. For feedforward networks, optimal learning can be achieved provided that certain conditions on the network and the learning environment are met. This principle is investigated for the case of networks using radial basis functions (RBF). It is assumed that the patterns of the learning environment are separable by hyperspheres. In that case, we prove that the attached cost function is local minima free with respect to all the weights. This provides us with some theoretical foundations for a massive application of RBF in pattern recognition. Monica Bianchini, Paolo Frasconi, Marco Gori |
IEEE Trans. Neural Networks | 3 |
| 1994 | On the problem of local minima in recurrent neural networksabstractMany researchers have recently focused their efforts on devising efficient algorithms, mainly based on optimization schemes, for learning the weights of recurrent neural networks. As in the case of feedforward networks, however, these learning algorithms may get stuck in local minima during gradient descent, thus discovering sub-optimal solutions. This paper analyses the problem of optimal learning in recurrent networks by proposing conditions that guarantee local minima free error surfaces. An example is given that also shows the constructive role of the proposed theory in designing networks suitable for solving a given task. Moreover, a formal relationship between recurrent and static feedforward networks is established such that the examples of local minima for feedforward networks already known in the literature can be associated with analogous ones in recurrent networks. Monica Bianchini, Marco Gori, Marco Maggini |
IEEE Trans. Neural Networks | 2 |
| 1992 | Phonetic recognition experiments with recurrent neural networks
Piero Cosi, Paolo Frasconi, Marco Gori, N. Griggio |
ICSLP | 3 |
| 1992 | Local Feedback Multilayered NetworksabstractIn this paper, we investigate the capabilities of local feedback multilayered networks, a particular class of recurrent networks, in which feedback connections are only allowed from neurons to themselves. In this class, learning can be accomplished by an algorithm that is local in both space and time. We describe the limits and properties of these networks and give some insights on their use for solving practical problems. Paolo Frasconi, Marco Gori, Giovanni Soda |
Neural Comput. | 2 |
| 1992 | On the Problem of Local Minima in BackpropagationabstractThe authors propose a theoretical framework for backpropagation (BP) in order to identify some of its limitations as a general learning procedure and the reasons for its success in several experiments on pattern recognition. The first important conclusion is that examples can be found in which BP gets stuck in local minima. A simple example in which BP can get stuck during gradient descent without having learned the entire training set is presented. This example guarantees the existence of a solution with null cost. Some conditions on the network architecture and the learning environment that ensure the convergence of the BP algorithm are proposed. It is proven in particular that the convergence holds if the classes are linearly separable. In this case, the experience gained in several experiments shows that multilayered neural networks (MLNs) exceed perceptrons in generalization to new examples.> Marco Gori, Alberto Tesi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1992 | Learning the dynamic nature of speech with back-propagation for sequences
Yoshua Bengio, Renato De Mori, Marco Gori |
Pattern Recognit. Lett. | 3 |
| 1990 | Performance analysis of tightly-coupled multiprocessor minicomputers
Luca Pistolesi, Marco Spadoni, Marco Gori |
Microprocessing and Microprogramming | 3 |