Artur S. d'Avila Garcez

dblp:26/4820 · also Artur d'Avila Garcez · DBLP profile ↗
← Back
69ranked-venue papers
15as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 60 · 10 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 1 first-authorTheory of computation · 5 · 4 first-author
YearPublicationVenuePosition
2025 Inducing Grokking with Distribution Shifts
abstract
Grokking, or delayed generalization, is an intriguing learning phenomenon where a model's test set performance improves sharply but only long after its training performance has converged. This challenges conventional understanding of training dynamics in deep learning. In this paper, we formalize and investigate grokking, highlighting that a distribution shift between training and test data is a key factor in its emergence. We introduce two novel synthetic datasets specifically designed to systematically induce and analyze the phenomenon. By controlling the imbalance in the sampling of data sub-categories, we can reliably reproduce grokking, demonstrating that while small sample sizes are associated with grokking, they are primarily only a mechanism for achieving the necessary distribution shift. We show that when classes possess a relational structure, such as in an equivariant mapping, the model can leverage this information to generalize to entirely unseen subclasses. To explore the limits of this phenomenon, we extend our analysis to the MNIST dataset. Our findings indicate that while distribution shifts can systematically cause grokking in structured synthetic settings, grokking may not emerge as readily in more complex real-world scenarios, suggesting that grokking may depend on the interplay between data distribution and the dataset's intrinsic structure.
Breno W. Carvalho, Artur S. d'Avila Garcez, Luís C. Lamb, Emilio Vital Brazil
ICTAI2
2025 A Semantic Framework for Neurosymbolic Computation (Abstract Reprint)
abstract
The field of neurosymbolic AI aims to benefit from the combination of neural networks and symbolic systems. A cornerstone of the field is the translation or encoding of symbolic knowledge into neural networks. Although many neurosymbolic methods and approaches have been proposed, and with a large increase in recent years, no common definition of encoding exists that can enable a precise, theoretical comparison of neurosymbolic methods. This paper addresses this problem by introducing a semantic framework for neurosymbolic AI. We start by providing a formal definition of semantic encoding, specifying the components and conditions under which a knowledge-base can be encoded correctly by a neural network. We then show that many neurosymbolic approaches are accounted for by this definition. We provide a number of examples and correspondence proofs applying the proposed framework to the neural encoding of various forms of knowledge representation. Many, at first sight disparate, neurosymbolic methods, are shown to fall within the proposed formalization. This is expected to provide guidance to future neurosymbolic encodings by placing them in the broader context of semantic encodings of entire families of existing neurosymbolic systems. The paper hopes to help initiate a discussion around the provision of a theory for neurosymbolic AI and a semantics for deep learning.
Simon Odense, Artur S. d'Avila Garcez
IJCAI2
2025 Towards Robust Neurosymbolic Relational Learning
abstract
Traditional neural networks (NNs) learn primarily from data, which limits their capacity to represent relational knowledge or handle symbolic relational data effectively. Although graph neural networks (GNNs) address this limitation at the level of relational data, they continue to struggle at learning relational knowledge. Neural-symbolic learning offers a solution by combining machine learning with knowledge representation, enabling the development of interpretable logic-based models learned from neural networks. Bottom clause propositionalization (BCP) is a prominent approach that transforms relational knowledge into attribute-value examples. A bottom clause is a logical representation created from each example as a starting point for the search process. BCP can be used with symbolic learners or neural networks to tackle relational domains. However, BCP often faces significant memory storage problems when handling larger datasets due to the volume of logical literals that it generates. Semi-propositionalization can alleviate these storage problems by grouping logical literals. However, it does not eliminate the substantial time requirements to create a bottom clause for each example. This paper investigates the application of sampling to the examples used for bottom clause generation. The hypothesis is that the number of examples needed to generate bottom clauses can be reduced significantly. Finding representative bottom clauses from data should enable relational learning to take place at an adequate level of abstract relational knowledge rather than simply at the level of the relations between any two data points. We evaluate this hypothesis by training a classifier with different sampling from five relational datasets. We experimentally validate the size of each sampling for each dataset. Experimental results show that training classifiers with fewer relational examples produces competitive results compared to using the entire dataset. The best results are obtained with up to 50% reduction in the set of examples.
Thais Luca, Aline Paes, Gerson Zaverucha, Artur S. d'Avila Garcez
IJCNN4
2025 A semantic framework for neurosymbolic computation
abstract
The field of neurosymbolic AI aims to benefit from the combination of neural networks and symbolic systems. A cornerstone of the field is the translation or encoding of symbolic knowledge into neural networks . Although many neurosymbolic methods and approaches have been proposed, and with a large increase in recent years, no common definition of encoding exists that can enable a precise, theoretical comparison of neurosymbolic methods. This paper addresses this problem by introducing a semantic framework for neurosymbolic AI. We start by providing a formal definition of semantic encoding , specifying the components and conditions under which a knowledge-base can be encoded correctly by a neural network. We then show that many neurosymbolic approaches are accounted for by this definition. We provide a number of examples and correspondence proofs applying the proposed framework to the neural encoding of various forms of knowledge representation. Many, at first sight disparate, neurosymbolic methods, are shown to fall within the proposed formalization. This is expected to provide guidance to future neurosymbolic encodings by placing them in the broader context of semantic encodings of entire families of existing neurosymbolic systems. The paper hopes to help initiate a discussion around the provision of a theory for neurosymbolic AI and a semantics for deep learning .
Simon Odense, Artur S. d'Avila Garcez
Artif. Intell.2
2024 Symbolic Knowledge Extraction and Distillation into Convolutional Neural Networks to Improve Medical Image Classification
abstract
Convolutional Neural Networks (CNNs) have achieved outstanding performance in radiology tasks. However, CNNs lack the transparency and explainability necessary to enable their practical clinical adoption. This paper introduces a neural-symbolic approach allowing domain experts to intervene in the training of CNNs. Following extraction and expert validation of meaningful symbolic knowledge from a trained CNN, such knowledge is distilled back into a streamlined CNN. The approach is shown to enhance user control over conventional CNN training by combining interpretable symbolic representations into a simplified CNN, allowing domain experts to control the decision making process. The kernels in a given CNN layer are mapped to symbolic knowledge representations in the form of logic programming rules. Extracted knowledge is evaluated against known radiomics features, allowing doctors to decide based on best practice which kernels to keep or reject. Expert intervention takes place through relevant knowledge distillation back into a more compact CNN. Our results show that a student CNN can learn successfully even from multiple teachers (different knowledge-bases) to replicate the selected relevant kernels and corresponding classification results. The proposed approach delivers a trainable parameter reduction of at least 56.3% while achieving high cosine similarity for kernel replication and a fidelity score of 99.2%. Expert validation highlights the importance of this approach at fostering greater trust in AI-driven medical decision making.
Kwun Ho Ngan, James Phelan, Joe Townsend, Artur S. d'Avila Garcez
IJCNN4
2024 Hypergraph Neural Networks with Logic Clauses
abstract
The analysis of structure in complex datasets has become essential to solving difficult Machine Learning problems. Relational aspects of data, capturing relationships between objects, play a crucial role in understanding the underlying data structure. While traditional graph algorithms have been widely used for binary relations, recent evidence suggests that hypergraphs can provide a more effective approach for modeling complex, non-binary relations. Hypergraph Neural Networks (HGNN) have been shown to offer a small improvement in performance when compared to Graph Neural Networks (GNN). In this paper, a new approach is proposed for inserting relational domain knowledge into HGNNs using a logic clause expressing non-binary relations. We evaluate the performance of this new hypergraph model, called Bottom-clause HGNN (BHGNN), in comparison with well-known approaches. Results show that BHGNN can achieve statistically significant improvement of performance, based on the Wilcoxon signed-ranks test, in comparison with HGNN and GNNs.
João Pedro Gandarela de Souza, Gerson Zaverucha, Artur S. d'Avila Garcez
IJCNN3
2024 FocusLearn: Fully-Interpretable, High-Performance Modular Neural Networks for Time Series
abstract
Multivariate time series have had many applications in areas from healthcare and finance to meteorology and life sciences. Although deep neural networks have shown excellent predictive performance for time series, they have been criticised for being non-interpretable. Neural Additive Models, however, are known to be fully-interpretable by construction, but may achieve far lower predictive performance than deep networks when applied to time series. This paper introduces FocusLearn, a fully-interpretable modular neural network capable of matching or surpassing the predictive performance of deep networks trained on multivariate time series. In FocusLearn, a recurrent neural network learns the temporal dependencies in the data, while a multi-headed attention layer learns to weight selected features while also suppressing redundant features. Modular neural networks are then trained in parallel and independently, one for each selected feature. This modular approach allows the user to inspect how features influence outcomes in the exact same way as with additive models. Experimental results show that this new approach outperforms additive models in both regression and classification of time series tasks, achieving predictive performance that is comparable to state-of-the-art, non-interpretable deep networks applied to time series.
Qiqi Su, Christos Kloukinas, Artur S. d'Avila Garcez
IJCNN3
2023 Neurosymbolic Reasoning and Learning with Restricted Boltzmann Machines
abstract
Knowledge representation and reasoning in neural networks has been a long-standing endeavour which has attracted much attention recently. The principled integration of reasoning and learning in neural networks is a main objective of the area of neurosymbolic Artificial Intelligence. In this paper, a neurosymbolic system is introduced that can represent any propositional logic formula. A proof of equivalence is presented showing that energy minimization in restricted Boltzmann machines corresponds to logical reasoning. We demonstrate the application of our approach empirically on logical reasoning and learning from data and knowledge. Experimental results show that reasoning can be performed effectively for a class of logical formulae. Learning from data and knowledge is also evaluated in comparison with learning of logic programs using neural networks. The results show that our approach can improve on state-of-the-art neurosymbolic systems. The theorems and empirical results presented in this paper are expected to reignite the research on the use of neural networks as massively-parallel models for logical reasoning and promote the principled integration of reasoning and learning in deep networks.
Son N. Tran, Artur S. d'Avila Garcez
AAAI2
2023 Contrastive counterfactual visual explanations with overdetermination
abstract
Abstract A novel explainable AI method called CLEAR Image is introduced in this paper. CLEAR Image is based on the view that a satisfactory explanation should be contrastive, counterfactual and measurable. CLEAR Image seeks to explain an image’s classification probability by contrasting the image with a representative contrast image, such as an auto-generated image obtained via adversarial learning. This produces a salient segmentation and a way of using image perturbations to calculate each segment’s importance. CLEAR Image then uses regression to determine a causal equation describing a classifier’s local input–output behaviour. Counterfactuals are also identified that are supported by the causal equation. Finally, CLEAR Image measures the fidelity of its explanation against the classifier. CLEAR Image was successfully applied to a medical imaging case study where it outperformed methods such as Grad-CAM and LIME by an average of 27% using a novel pointing game metric. CLEAR Image also identifies cases of causal overdetermination, where there are multiple segments in an image that are sufficient individually to cause the classification probability to be close to one.
Adam White 0002, Kwun Ho Ngan, James Phelan, Kevin Ryan, Saman Sadeghi Afgeh, Constantino Carlos Reyes-Aldasoro, Artur S. d'Avila Garcez
Mach. Learn.7
2022 Extracting Meaningful High-Fidelity Knowledge from Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) have been widely used for complex image recognition tasks. Due to the highly entangled correlations learned by the latent features in the convolutional kernels of CNNs, deriving human-comprehensible knowledge from CNNs has been proven difficult. As such, reasoning from relationships between kernels has been limited, resulting in little knowledge transfer from one task to another related task learned by CNNs. This paper introduces a neural-symbolic approach for providing semantically meaningful explanations to CNNs using logical rules and a shared conceptual representation space to capture the meaning of the knowledge learned. The validity of the proposed approach is demonstrated using benchmark chest x-rays of two respiratory conditions: pleural effusion and COVID-19. Our results show empirically that symbolic rules can be associated with semantically meaningful explanations obtained from different but related CNN models, even in domains requiring specialised knowledge such as medical imaging. This work is expected to aid the analysis of black-box CNNs by associating the predictions obtained from the CNNs with clinical research findings.
Kwun Ho Ngan, Artur S. d'Avila Garcez, Joseph Townsend
IJCNN2
2022 Formalizing Consistency and Coherence of Representation Learning
abstract
In the study of reasoning in neural networks, recent efforts have sought to improve consistency and coherence of sequence models, leading to important developments in the area of neuro-symbolic AI. In symbolic AI, the concepts of consistency and coherence can be defined and verified formally, but for neural networks these definitions are lacking. The provision of such formal definitions is crucial to offer a common basis for the quantitative evaluation and systematic comparison of connectionist, neuro-symbolic and transfer learning approaches. In this paper, we introduce formal definitions of consistency and coherence for neural systems. To illustrate the usefulness of our definitions, we propose a new dynamic relation-decoder model built around the principles of consistency and coherence. We compare our results with several existing relation-decoders using a partial transfer learning task based on a novel data set introduced in this paper. Our experiments show that relation-decoders that maintain consistency over unobserved regions of representation space retaincoherence across domains, whilst achieving better transfer learning performance.
Harald Strömfelt, Luke Dickens, Artur S. d'Avila Garcez, Alessandra Russo
NeurIPS3
2022 Logic Tensor Networks
Samy Badreddine, Artur S. d'Avila Garcez, Luciano Serafini, Michael Spranger
Artif. Intell.2
2020 Measurable Counterfactual Local Explanations for Any Classifier
abstract
We propose a novel method for explaining the predictions of any classifier.In our approach, local explanations are expected to explain both the outcome of a prediction and how that prediction would change if 'things had been different'.Furthermore, we argue that satisfactory explanations cannot be dissociated from a notion and measure of fidelity, as advocated in the early days of neural networks' knowledge extraction.We introduce a definition of fidelity to the underlying classifier for local explanation models which is based on distances to a target decision boundary.A system called CLEAR: Counterfactual Local Explanations via Regression, is introduced and evaluated.CLEAR generates b-counterfactual explanations that state minimum changes necessary to flip a prediction's classification.CLEAR then builds local regression models, using the b-counterfactuals to measure and improve the fidelity of its regressions.By contrast, the popular LIME method [17], which also uses regression to generate local explanations, neither measures its own fidelity nor generates counterfactuals.CLEAR's regressions are found to have significantly higher fidelity than LIME's, averaging over 40% higher in this paper's five case studies.
Adam White 0002, Artur S. d'Avila Garcez
ECAI2
2020 Neural-Symbolic Relational Reasoning on Graph Models: Effective Link Inference and Computation from Knowledge Bases
Henrique Lemos dos Santos, Pedro H. C. Avelar, Marcelo O. R. Prates, Artur S. d'Avila Garcez, Luís C. Lamb
ICANN (1)4
2020 Graph Neural Networks Meet Neural-Symbolic Computing: A Survey and Perspective
abstract
Neural-symbolic computing has now become the subject of interest of both academic and industry research laboratories. Graph Neural Networks (GNNs) have been widely used in relational and symbolic domains, with widespread application of GNNs in combinatorial optimization, constraint satisfaction, relational reasoning and other scientific domains. The need for improved explainability, interpretability and trust of AI systems in general demands principled methodologies, as suggested by neural-symbolic computing. In this paper, we review the state-of-the-art on the use of GNNs as a model of neural-symbolic computing. This includes the application of GNNs in several domains as well as their relationship to current developments in neural-symbolic computing.
Luís C. Lamb, Artur S. d'Avila Garcez, Marco Gori, Marcelo O. R. Prates, Pedro H. C. Avelar, Moshe Y. Vardi
IJCAI2
2020 Semi-supervised GANs for Fraud Detection*
abstract
Over the years the online gambling industry has evolved into one of the most profitable industries on the Internet. At the same time, new stringent regulations have required the online industry to become a lot more vigilant. Although standards have improved, the methods used to process finance from illicit activities also evolved and became more sophisticated. Detecting these fraudulent activities in real life with high accuracy requires a learning system to be trained with balanced data sets of fraudulent and normal transactions. However, in the real-world, the number of fraudulent cases is significantly lower than normal cases. In this paper, to deal with data imbalance, we propose a novel generative adversarial framework based on semi-supervised learning of sparse auto-encoders for the detection of fraud in online gambling. Experimental results show that the proposed framework outperforms mainstream discriminative techniques such as logistic regression, random forest and multi-layer perceptron. We validate further the approach by applying it to other domains that suffer from the problem of class imbalance obtaining promising results.
Charitos Charitou, Artur S. d'Avila Garcez, Simo Dragicevic
IJCNN2
2020 Neuro-Symbolic Probabilistic Argumentation Machines
abstract
Neural-symbolic systems combine the strengths of neural networks and symbolic formalisms. In this paper, we introduce a neural-symbolic system which combines restricted Boltzmann machines and probabilistic semi-abstract argumentation. We propose to train networks on argument labellings explaining the data, so that any sampled data outcome is associated with an argument labelling. Argument labellings are integrated as constraints within restricted Boltzmann machines, so that the neural networks are used to learn probabilistic dependencies amongst argument labels. Given a dataset and an argumentation graph as prior knowledge, for every example/case K in the dataset, we use a so-called K-maxconsistent labelling of the graph, and an explanation of case K refers to a K-maxconsistent labelling of the given argumentation graph. The abilities of the proposed system to predict correct labellings were evaluated and compared with standard machine learning techniques. Experiments revealed that such argumentation Boltzmann machines can outperform other classification models, especially in noisy settings.
Régis Riveret, Son N. Tran, Artur S. d'Avila Garcez
KR3
2020 Probabilistic approaches for music similarity using restricted Boltzmann machines
Son N. Tran, Ngo Tung Son, Artur S. d'Avila Garcez
Neural Comput. Appl.3
2020 Sequence Classification Restricted Boltzmann Machines With Gated Units
abstract
For the classification of sequential data, dynamic Bayesian networks and recurrent neural networks (RNNs) are the preferred models. While the former can explicitly model the temporal dependences between the variables, and the latter have the capability of learning representations. The recurrent temporal restricted Boltzmann machine (RTRBM) is a model that combines these two features. However, learning and inference in RTRBMs can be difficult because of the exponential nature of its gradient computations when maximizing log likelihoods. In this article, first, we address this intractability by optimizing a conditional rather than a joint probability distribution when performing sequence classification. This results in the "sequence classification restricted Boltzmann machine" (SCRBM). Second, we introduce gated SCRBMs (gSCRBMs), which use an information processing gate, as an integration of SCRBMs with long short-term memory (LSTM) models. In the experiments reported in this article, we evaluate the proposed models on optical character recognition, chunking, and multiresident activity recognition in smart homes. The experimental results show that gSCRBMs achieve the performance comparable to that of the state of the art in all three tasks. gSCRBMs require far fewer parameters in comparison with other recurrent networks with memory gates, in particular, LSTMs and gated recurrent units (GRUs).
Son N. Tran, Artur S. d'Avila Garcez, Tillman Weyde, Jie Yin 0001, Qing Zhang 0001, Mohan Karunanithi
IEEE Trans. Neural Networks Learn. Syst.2
2019 Editorial: Booming of Neural Networks and Learning Systems
abstract
As you open this January issue of the IEEE Transactions on Neural Networks and Learning Systems (TNNLS), I hope everyone enjoyed a great holiday season and is excited for the new year of 2019. I am very delighted and honored to report several key metrics of IEEE TNNLS to the community.
Akira Hirose 0001, Alessio Micheli, Artur S. d'Avila Garcez, Choon Ki Ahn, Gang Pan 0001, Hamid Reza Karimi, Jianbing Shen, José de Jesús Rubio, Lei Zhang 0005, Lingjia Liu 0001, Lorenzo Livi, Nishchal K. Verma, Pedro Antonio Gutiérrez, Qi Tian 0001, Qinglai Wei, Seiichi Ozawa, Stuart Harvey Rubin, Weineng Chen, Xi Li 0001, Xiaofeng Liao 0001, Youmin Zhang 0001, Zhen Ni, Haibo He
IEEE Trans. Neural Networks Learn. Syst.3
2018 Speaker recognition with hybrid features from a deep belief network
Hazrat Ali, Son Ngoc Tran, Emmanouil Benetos, Artur S. d'Avila Garcez
Neural Comput. Appl.4
2018 Deep Logic Networks: Inserting and Extracting Knowledge From Deep Belief Networks
abstract
Developments in deep learning have seen the use of layerwise unsupervised learning combined with supervised learning for fine-tuning. With this layerwise approach, a deep network can be seen as a more modular system that lends itself well to learning representations. In this paper, we investigate whether such modularity can be useful to the insertion of background knowledge into deep networks, whether it can improve learning performance when it is available, and to the extraction of knowledge from trained deep networks, and whether it can offer a better understanding of the representations learned by such networks. To this end, we use a simple symbolic language-a set of logical rules that we call confidence rules-and show that it is suitable for the representation of quantitative reasoning in deep networks. We show by knowledge extraction that confidence rules can offer a low-cost representation for layerwise networks (or restricted Boltzmann machines). We also show that layerwise extraction can produce an improvement in the accuracy of deep belief networks. Furthermore, the proposed symbolic characterization of deep networks provides a novel method for the insertion of prior knowledge and training of deep networks. With the use of this method, a deep neural-symbolic system is proposed and evaluated, with the experimental results indicating that modularity through the use of confidence rules and knowledge insertion can be beneficial to network performance.
Son Ngoc Tran, Artur S. d'Avila Garcez
IEEE Trans. Neural Networks Learn. Syst.2
2017 Generalising the Discriminative Restricted Boltzmann Machines
Srikanth Cherla, Son Ngoc Tran, Artur S. d'Avila Garcez, Tillman Weyde
ICANN (2)3
2017 Extracting M of N Rules from Restricted Boltzmann Machines
Simon Odense, Artur S. d'Avila Garcez
ICANN (2)2
2017 Logic Tensor Networks for Semantic Image Interpretation
abstract
Semantic Image Interpretation (SII) is the task of extracting structured semantic descriptions from images. It is widely agreed that the combined use of visual data and background knowledge is of great importance for SII. Recently, Statistical Relational Learning (SRL) approaches have been developed for reasoning under uncertainty and learning in the presence of data and rich knowledge. Logic Tensor Networks (LTNs) are a SRL framework which integrates neural networks with first-order fuzzy logic to allow (i) efficient learning from noisy data in the presence of logical constraints, and (ii) reasoning with logical formulas describing general properties of the data. In this paper, we develop and apply LTNs to two of the main tasks of SII, namely, the classification of an image's bounding boxes and the detection of the relevant part-of relations between objects. To the best of our knowledge, this is the first successful application of SRL to such SII tasks. The proposed approach is evaluated on a standard image processing benchmark. Experiments show that background knowledge in the form of logical constraints can improve the performance of purely data-driven approaches, including the state-of-the-art Fast Region-based Convolutional Neural Networks (Fast R-CNN). Moreover, we show that the use of logical background knowledge adds robustness to the learning system when errors are present in the labels of the training data.
Ivan Donadello, Luciano Serafini, Artur S. d'Avila Garcez
IJCAI3
2017 On the memory properties of recurrent neural models
abstract
In this paper, we investigate the memory properties of two popular gated units: long short term memory (LSTM) and gated recurrent units (GRU), which have been used in recurrent neural networks (RNN) to achieve state-of-the-art performance on several machine learning tasks. We propose five basic tasks for isolating and examining specific capabilities relating to the implementation of memory. Results show that (i) both types of gated unit perform less reliably than standard RNN units on tasks testing fixed delay recall, (ii) the reliability of stochastic gradient descent decreases as network complexity increases, and (iii) gated units are found to perform better than standard RNNs on tasks that require values to be stored in memory and updated conditionally upon input to the network. Task performance is found to be surprisingly independent of network depth (number of layers) and connection architecture. Finally, visualisations of the solutions found by these networks are presented and explored, exposing for the first time how logic operations are implemented by individual gated cells and small groups of these cells.
Arthur Jack Russell, Emmanouil Benetos, Artur S. d'Avila Garcez
IJCNN3
2016 The Need for Knowledge Extraction: Understanding Harmful Gambling Behavior with Neural Networks
abstract
Responsible gambling is a field of study that involves supporting gamblers so as to reduce the harm that their gambling activity might cause. Recently in the literature, machine learning algorithms have been introduced as a way to predict potentially harmful gambling based on patterns of gambling behavior, such as trends in amounts wagered and the time spent gambling. In this paper, neural network models are analyzed to help predict the outcome of a partial proxy for harmful gambling behavior: when a gambler “self-excludes”, requesting a gambling operator to prevent them from accessing gambling opportunities. Drawing on survey and interview insights from industry and public officials as to the importance of interpretability, a variant of the knowledge extraction algorithm TREPAN is proposed which can produce compact, human-readable logic rules efficiently, given a neural network trained on gambling data. To the best of our knowledge, this paper reports the first industrial-strength application of knowledge extraction from neural networks, which otherwise are black-boxes unable to provide the explanatory insights which are crucially required in this area of application. We show that through knowledge extraction one can explore and validate the kinds of behavioral and demographic profiles that best predict self-exclusion, while developing a machine learning approach with greater potential for adoption by industry and treatment providers. Experimental results reported in this paper indicate that the rules extracted can achieve high fidelity to the trained neural network while maintaining competitive accuracy and providing useful insight to domain experts in responsible gambling.
Christian Percy, Artur S. d'Avila Garcez, Simo Dragicevic, Manoel V. M. França, Gregory Slabaugh, Tillman Weyde
ECAI2
2016 Adaptive Transferred-profile Likelihood Learning
abstract
The recent success of representation learning is built upon the learning of relevant features, in particular from unlabelled data available in different domains. This raises the question of how to transfer and reuse such knowledge effectively so that the learning of a new task can be made easier or be improved. This poses a difficult challenge for the area of transfer learning where there is no label in the source data, and no source data is ever transferred to the target domain. In previous work, the most capable approach has been self-taught learning which, however, relies heavily upon the compatibility across the domains. In this paper, we propose a novel transfer learning framework called Adaptive Transferred-profile Likelihood Learning (aTPL), which performs adaptation on the representations to be transferred, so that they become more compatible with the target domain. At the same time, it learns supplementary knowledge about the target domain. Experiments on five images datasets and a sentiment dataset demonstrate the effectiveness of the approach in comparison with self-taught learning and other common feature extraction methods. The results also indicate that the new transfer method is less sensitive to negative transfer.
Son Ngoc Tran, Artur S. d'Avila Garcez
IJCNN2
2016 Fat-Fast VG-RAM WNN: A high performance approach
Avelino Forechi, Alberto Ferreira de Souza, Jorcy de Oliveira Neto, Edilson de Aguiar, Claudine Badue, Artur S. d'Avila Garcez, Thiago Oliveira-Santos
Neurocomputing6
2015 A hybrid recurrent neural network for music transcription
abstract
We investigate the problem of incorporating higher-level symbolic score-like information into Automatic Music Transcription (AMT) systems to improve their performance. We use recurrent neural networks (RNNs) and their variants as music language models (MLMs) and present a generative architecture for combining these models with predictions from a frame level acoustic classifier. We also compare different neural network architectures for acoustic modeling. The proposed model computes a distribution over possible output sequences given the acoustic input signal and we present an algorithm for performing a global search for good candidate transcriptions. The performance of the proposed model is evaluated on piano music from the MAPS dataset and we observe that the proposed model consistently outperforms existing transcription methods.
Siddharth Sigtia, Emmanouil Benetos, Nicolas Boulanger-Lewandowski, Tillman Weyde, Artur S. d'Avila Garcez, Simon Dixon
ICASSP5
2015 Discriminative learning and inference in the Recurrent Temporal RBM for melody modelling
abstract
We are interested in modelling musical pitch sequences in melodies in the symbolic form. The task here is to learn a model to predict the probability distribution over the various possible values of pitch of the next note in a melody, given those leading up to it. For this task, we propose the Recurrent Temporal Discriminative Restricted Boltzmann Machine (RTDRBM). It is obtained by carrying out discriminative learning and inference as put forward in the Discriminative RBM (DRBM), in a temporal setting by incorporating the recurrent structure of the Recurrent Temporal RBM (RTRBM). The model is evaluated on the cross entropy of its predictions using a corpus containing 8 datasets of folk and chorale melodies, and compared with n-grams and other standard connectionist models. Results show that the RTDRBM has a better predictive performance than the rest of the models, and that the improvement is statistically significant.
Srikanth Cherla, Son Ngoc Tran, Artur S. d'Avila Garcez, Tillman Weyde
IJCNN3
2015 Neural-symbolic monitoring and adaptation
abstract
Runtime monitors check the execution of a system under scrutiny against a set of formal specifications describing a prescribed behaviour. The two core properties for monitoring systems are scalability and adaptability. In this paper we show how RuleRunner, our previous neural-symbolic monitoring system, can exploit learning strategies in order to integrate desired deviations with the initial set of specification. The resulting system allows for fast conformance checking and can suggest possible enhanced models when the initial set of specifications has to be adapted in order to include new patterns.
Alan Perotti, Artur S. d'Avila Garcez, Guido Boella
IJCNN2
2015 Efficient representation ranking for transfer learning
abstract
Representation learning has emerged recently as a useful tool in the extraction of features from data. In a range of applications, features learned from data have been shown superior to their hand-crafted counterpart. Many deep learning approaches have taken advantage of such feature extraction. However, further research is needed on how such features can be evaluated for re-use in related applications, hopefully then improving performance on such applications. In this paper, we present a new method for ranking the representations learned by a Restricted Boltzmann Machine, which has been used regularly as a feature learner by deep networks. We show that high-ranking features, according to our method, should capture more information than low-ranking ones. We then apply representation ranking for pruning the network, and propose a new transfer learning algorithm, which uses such features extracted from a trained network to improve learning performance in another network trained on an analogous domain. We show that by transferring a small number of highest scored representations from source domain our method encourages the learning of new knowledge in target domain while preserving most of the information of the source domain during the transfer. This transfer learning is similar to self-taught learning in that it does not use the source domain data during the transfer process.
Son Ngoc Tran, Artur S. d'Avila Garcez
IJCNN2
2015 Runtime Verification Through Forward Chaining
Alan Perotti, Guido Boella, Artur S. d'Avila Garcez
RV3
2014 Low-Cost Representation for Restricted Boltzmann Machines
Son Ngoc Tran, Artur S. d'Avila Garcez
ICONIP (1)2
2014 Applying Neural-Symbolic Cognitive Agents in Intelligent Transport Systems to reduce CO2 emissions
abstract
Providing personalized feedback in Intelligent Transport Systems is a powerful tool for instigating a change in driving behaviour and the reduction of CO2emissions. This requires a system that is capable of detecting driver characteristics from real-time vehicle data. In this paper, we apply the architecture and theory of a Neural-Symbolic Cognitive Agent (NSCA) to effectively learn and reason about observed driving behaviour and related driver characteristics. The NSCA architecture combines neural learning and reasoning with symbolic temporal knowledge representation and is capable of encoding background knowledge, learning new hypotheses from observed data, and inferring new beliefs based on these hypotheses. Furthermore, it deals with uncertainty and errors in the data using a Bayesian inference model, and it scales well to hundreds of thousands of data samples as in the application reported in this paper. We have applied the NSCA in an Intelligent Transport System to reduce CO2emissions as part of an European Union project, called EcoDriver. Results reported in this paper show that the NSCA outperforms the state-of-the-art in this application area, and is applicable to very large data.
Leo de Penning, Artur S. d'Avila Garcez, Luís C. Lamb, Arjan Stuiver, John-Jules Ch. Meyer
IJCNN2
2014 Neural Networks for Runtime Verification
abstract
A recent trend in High-Performance Computation is parallel computing, and the field of Neural Networks is showing impressive improvements in performance, especially with the use of GPU accelerators. In this paper, we use neural networks to improve the performance of Runtime Verification. Runtime verification is used in a variety of domains -from policy enforcement to electronic fraud detection-to automatically check whether a system meets a temporal specification, by observing the output of the system. In this paper, we present a novel run-time monitoring system, RuleRunner, and we exploit results from the Neural-Symbolic Integration area to encode it in a recurrent neural network. The results show that neural networks can perform real-time online runtime verification. Performance was improved by the parallel architecture and the matrix-based implementation with GPU.
Alan Perotti, Artur S. d'Avila Garcez, Guido Boella
IJCNN2
2014 Learning motion-difference features using Gaussian restricted Boltzmann machines for efficient human action recognition
abstract
Learning visual words from video frames is challenging because deciding which word to assign to each subset of frames is a difficult task. For example, two similar frames may have different meanings in describing human actions such as starting to run and starting to walk. In order to associate richer information to vector-quantization and generate visual words, several approaches have been proposed recently that use complex algorithms to extract or learn spatio-temporal features from 3-D volumes of video frames. In this paper, we propose an efficient method to use Gaussian RBMs for learning motion-difference features from actions in videos. The difference between two video frames is defined by a subtraction function of one frame by another that preserves positive and negative changes, thus creating a simple spatio-temporal saliency map for an action. This subtraction function removes, by construction, the common shapes and background images that should not be relevant for action learning and recognition, and highlights the movement patterns in space, making it easier to learn the actions from such saliency maps using shallow feature learning models such as RBMs. In the experiments reported in this paper, we used a Gaussian restricted Boltzmann machine to learn the actions from saliency maps of different motion images. Despite its simplicity, the motion-difference method achieved very good performance in benchmark datasets, specifically the Weizmann dataset (98.81%) and the KTH dataset (88.89%). A comparative analysis with hand-crafted and learned features using similar classifiers indicates that motion-difference can be competitive and very efficient.
Son Ngoc Tran, Emmanouil Benetos, Artur S. d'Avila Garcez
IJCNN3
2014 Fast relational learning using bottom clause propositionalization with artificial neural networks
Manoel V. M. França, Gerson Zaverucha, Artur S. d'Avila Garcez
Mach. Learn.3
2012 Multi-instance learning using recurrent neural networks
abstract
Multiple instance learning is an increasingly important area in machine learning. In multi-instance learning, the training set is structured into subsets (or bags) of instances. The bags are labelled, but the label of each instance is unknown or irrelevant. In this paper, we revisit the connectionist approach to multi-instance learning. We propose a recurrent neural network model for multi-instance learning. We have applied the new model to a benchmark multi-instance dataset. The results provide evidence that connectionist multi-instance learning is more promising than previously anticipated. We argue that a principled connectionist approach should provide robust and efficient multi-instance learning, yet comparative results should be taken with caution as a result of varying methodologies.
Artur S. d'Avila Garcez, Gerson Zaverucha
IJCNN1
2011 Learning to adapt requirements specifications of evolving systems
abstract
We propose a novel framework for adapting and evolving software requirements models. The framework uses model checking and machine learning techniques for verifying properties and evolving model descriptions. The paper offers two novel contributions and a preliminary evaluation and application of the ideas presented. First, the framework is capable of coping with errors in the specification process so that performance degrades gracefully. Second, the framework can also be used to re-engineer a model from examples only, when an initial model is not available. We provide a preliminary evaluation of our framework by applying it to a Pump System case study, and integrate our prototype tool with the NuSMV model checker. We show how the tool integrates verification and evolution of abstract models, and also how it is capable of re-engineering partial models given examples from an existing system.
Rafael V. Borges, Artur S. d'Avila Garcez, Luís C. Lamb, Bashar Nuseibeh
ICSE2
2011 A Neural-Symbolic Cognitive Agent for Online Learning and Reasoning
abstract
In real-world applications, the effective integration of learning and reasoning in a cognitive agent model is a difficult task. However, such integration may lead to a better understanding, use and construction of more realistic models. Unfortunately, existing models are either oversimplified or require much processing time, which is unsuitable for online learning and reasoning. Currently, controlled environments like training simulators do not effectively integrate learning and reasoning. In particular, higher-order concepts and cognitive abilities have many unknown temporal relations with the data, making it impossible to represent such relationships by hand. We introduce a novel cognitive agent model and architecture for online learning and reasoning that seeks to effectively represent, learn and reason in complex training environments. The agent architecture of the model combines neural learning with symbolic knowledge representation. It is capable of learning new hypotheses from observed data, and infer new beliefs based on these hypotheses. Furthermore, it deals with uncertainty and errors in the data using a Bayesian inference model. The validation of the model on real-time simulations and the results presented here indicate the promise of the approach when performing online learning and reasoning in real-world scenarios, with possible applications in a range of areas.
Leo de Penning, Artur S. d'Avila Garcez, Luís C. Lamb, John-Jules Ch. Meyer
IJCAI2
2011 Learning and Representing Temporal Knowledge in Recurrent Networks
abstract
The effective integration of knowledge representation, reasoning, and learning in a robust computational model is one of the key challenges of computer science and artificial intelligence. In particular, temporal knowledge and models have been fundamental in describing the behavior of computational systems. However, knowledge acquisition of correct descriptions of a system's desired behavior is a complex task. In this paper, we present a novel neural-computation model capable of representing and learning temporal knowledge in recurrent networks. The model works in an integrated fashion. It enables the effective representation of temporal knowledge, the adaptation of temporal models given a set of desirable system properties, and effective learning from examples, which in turn can lead to temporal knowledge extraction from the corresponding trained networks. The model is sound from a theoretical standpoint, but it has also been tested on a case study in the area of model verification and adaptation. The results contained in this paper indicate that model verification and learning can be integrated within the neural computation paradigm, contributing to the development of predictive temporal knowledge-based systems and offering interpretable results that allow system researchers and engineers to improve their models and specifications. The model has been implemented and is available as part of a neural-symbolic computational toolkit.
Rafael V. Borges, Artur S. d'Avila Garcez, Luís C. Lamb
IEEE Trans. Neural Networks2
2010 Representing, Learning and Extracting Temporal Knowledge from Neural Networks: A Case Study
Rafael V. Borges, Artur S. d'Avila Garcez, Luís C. Lamb
ICANN (2)2
2010 Neuro-symbolic Representation of Logic Programs Defining Infinite Sets
Ekaterina Komendantskaya, Krysia Broda, Artur S. d'Avila Garcez
ICANN (1)3
2010 First-order logic learning in Artificial Neural Networks
abstract
Artificial Neural Networks have previously been applied in neuro-symbolic learning to learn ground logic program rules. However, there are few results of learning relations using neuro-symbolic learning. This paper presents the system PAN, which can learn relations. The inputs to PAN are one or more atoms, representing the conditions of a logic rule, and the output is the conclusion of the rule. The symbolic inputs may include functional terms of arbitrary depth and arity, and the output may include terms constructed from the input functors. Symbolic inputs are encoded as an integer using an invertible encoding function, which is used in reverse to extract the output terms. The main advance of this system is a convention to allow construction of Artificial Neural Networks able to learn rules with the same power of expression as first order definite clauses. The system is tested on three examples and the results are discussed.
Mathieu Guillame-Bert, Krysia Broda, Artur S. d'Avila Garcez
IJCNN3
2010 SOAR - Sparse Oracle-based Adaptive Rule extraction: Knowledge extraction from large-scale datasets to detect credit card fraud
abstract
This paper presents a novel approach to knowledge extraction from large-scale datasets using a neural network when applied to the real-world problem of payment card fraud detection. Fraud is a serious and long term threat to a peaceful and democratic society. We present SOAR (Sparse Oracle-based Adaptive Rule) extraction, a practical approach to process large datasets and extract key generalizing rules that are comprehensible using a trained neural network as an oracle to locate key decision boundaries. Experimental results indicate a high level of rule comprehensibility with an acceptable level of accuracy can be achieved. The SOAR extraction outperformed the best decision tree induction method and produced over 10 times fewer rules aiding comprehensibility. Moreover, the extracted rules discovered fraud facts of key interest to industry fraud analysts.
Nick F. Ryman-Tubb, Artur S. d'Avila Garcez
IJCNN2
2010 Integrating model verification and self-adaptation
abstract
In software development, formal verification plays an important role in improving the quality and safety of products and processes. Model checking is a successful approach to verification, used both in academic research and industrial applications. One important improvement regarding utilization of model checking is the development of automated processes to evolve models according to information obtained from verification. In this paper, we propose a new framework that make use of artificial intelligence and machine learning to generate and evolve models from partial descriptions and examples created by the model checking process. This was implemented as a tool that is integrated with a model checker. Our work extends model checking to be applicable when initial description of a system is not available, through observation of actual behaviour of this system. The framework is capable of integrated verification and evolution of abstract models, but also of reengineering partial models of a system.
Rafael V. Borges, Artur S. d'Avila Garcez, Luís C. Lamb
ASE2
2008 Symbolic Knowledge Extraction from Support Vector Machines: A Geometric Approach
Artur S. d'Avila Garcez
ICONIP (2)2
2007 A Connectionist Cognitive Model for Temporal Synchronisation and Learning
Luís C. Lamb, Rafael V. Borges, Artur S. d'Avila Garcez
AAAI3
2007 Reasoning and Learning About Past Temporal Knowledge in Connectionist Models
abstract
The integration of logic-based inference systems and connectionist learning architectures may lead to the construction of semantically sound cognitive models in artificial intelligence. The use of hybrid systems has shown promising results as regards the computation and learning of classical reasoning within neural networks. However, there still remains a number of open research issues on the integration of non-classical logics and neural networks. We present a new model for integrating symbolic reasoning about past temporal information and neural learning systems. We propose algorithms that translate background knowledge into a neural network and analyse the effectiveness of learning algorithms when subject to symbolic temporal knowledge. This opens several interesting research paths with possible applications to agents' decision making, cognitive modelling and knowledge-based systems.
Rafael V. Borges, Luís C. Lamb, Artur S. d'Avila Garcez
IJCNN3
2007 Connectionist modal logic: Representing modalities in neural networks
Artur S. d'Avila Garcez, Luís C. Lamb, Dov M. Gabbay
Theor. Comput. Sci.1
2006 Combining Architectures for Temporal Learning in Neural-Symbolic Systems
Rafael V. Borges, Luís C. Lamb, Artur S. d'Avila Garcez
HIS3
2006 Improving VG-RAM Neural Networks Performance Using Knowledge Correlation
Raphael V. Carneiro, Stiven S. Dias, Dijalma Fardin, Hallysson Oliveira, Artur S. d'Avila Garcez, Alberto Ferreira de Souza
ICONIP (1)5
2006 A Connectionist Computational Model for Epistemic and Temporal Reasoning
abstract
The importance of the efforts to bridge the gap between the connectionist and symbolic paradigms of artificial intelligence has been widely recognized. The merging of theory (background knowledge) and data learning (learning from examples) into neural-symbolic systems has indicated that such a learning system is more effective than purely symbolic or purely connectionist systems. Until recently, however, neural-symbolic systems were not able to fully represent, reason, and learn expressive languages other than classical propositional and fragments of first-order logic. In this article, we show that nonclassical logics, in particular propositional temporal logic and combinations of temporal and epistemic (modal) reasoning, can be effectively computed by artificial neural networks. We present the language of a connectionist temporal logic of knowledge (CTLK). We then present a temporal algorithm that translates CTLK theories into ensembles of neural networks and prove that the translation is correct. Finally, we apply CTLK to the muddy children puzzle, which has been widely used as a test-bed for distributed knowledge representation. We provide a complete solution to the puzzle with the use of simple neural networks, capable of reasoning about knowledge evolution in time and of knowledge acquisition through learning.
Artur S. d'Avila Garcez, Luís C. Lamb
Neural Comput.1
2006 Connectionist computations of intuitionistic reasoning
Artur S. d'Avila Garcez, Luís C. Lamb, Dov M. Gabbay
Theor. Comput. Sci.1
2005 Fewer Epistemological Challenges for Connectionism
Artur S. d'Avila Garcez
CiE1
2005 A Connectionist Model for Constructive Modal Reasoning
abstract
We present a new connectionist model for constructive, intuitionistic modal reasoning. We use ensembles of neural networks to represent in- tuitionistic modal theories, and show that for each intuitionistic modal program there exists a corresponding neural network ensemble that com- putes the program. This provides a massively parallel model for intu- itionistic modal reasoning, and sets the scene for integrated reasoning, knowledge representation, and learning of intuitionistic theories in neural networks, since the networks in the ensemble can be trained by examples using standard neural learning algorithms.
Artur S. d'Avila Garcez, Luís C. Lamb, Dov M. Gabbay
NIPS1
2005 Value-based Argumentation Frameworks as Neural-symbolic Learning Systems
abstract
While neural networks have been successfully used in a number of machine learning applications, logical languages have been the standard for the representation of argumentative reasoning. In this paper, we establish a relationship between neural networks and argumentation networks, combining reasoning and learning in the same argumentation framework. We do so by presenting a new neural argumentation algorithm, responsible for translating argumentation networks into standard neural networks. We then show a correspondence between the two networks. The algorithm works not only for acyclic argumentation networks, but also for circular networks, and it enables the accrual of arguments through learning as well as the parallel computation of arguments
Artur S. d'Avila Garcez, Dov M. Gabbay, Luís C. Lamb
J. Log. Comput.1
2004 Fibring Neural Networks
Artur S. d'Avila Garcez, Dov M. Gabbay
AAAI1
2004 Towards a Connectionist Argumentation Framework
Artur S. d'Avila Garcez, Dov M. Gabbay, Luís C. Lamb
ECAI1
2004 Argumentation Neural Networks
Artur S. d'Avila Garcez, Dov M. Gabbay, Luís C. Lamb
ICONIP1
2003 Neural-Symbolic Intuitionistic Reasoning
Artur S. d'Avila Garcez, Luís C. Lamb, Dov M. Gabbay
HIS1
2003 Reasoning about Time and Knowledge in Neural Symbolic Learning Systems
abstract
We show that temporal logic and combinations of temporal logics and modal logics of knowledge can be effectively represented in ar(cid:173) tificial neural networks. We present a Translation Algorithm from temporal rules to neural networks, and show that the networks compute a fixed-point semantics of the rules. We also apply the translation to the muddy children puzzle, which has been used as a testbed for distributed multi-agent systems. We provide a complete solution to the puzzle with the use of simple neural networks, capa(cid:173) ble of reasoning about time and of knowledge acquisition through inductive learning.
Artur S. d'Avila Garcez, Luís C. Lamb
NIPS1
2003 Revising Rules to Capture Requirements Traceability Relations: A Machine Learning Approach
George Spanoudakis, Artur S. d'Avila Garcez, Andrea Zisman
SEKE2
2001 An Analysis-Revision Cycle to Evolve Requirements Specifications
abstract
We argue that the evolution of requirements specifications can be supported by a cycle composed of two phases: analysis and revision. We investigate an instance of such a cycle, which combines two techniques of logical abduction and inductive learning to analyze and revise specifications respectively.
Artur S. d'Avila Garcez, Alessandra Russo, Bashar Nuseibeh, Jeff Kramer
ASE1
2001 Symbolic knowledge extraction from trained neural networks: A sound approach
Artur S. d'Avila Garcez, Krysia Broda, Dov M. Gabbay
Artif. Intell.1
1999 The Connectionist Inductive Learning and Logic Programming System
Artur S. d'Avila Garcez, Gerson Zaverucha
Appl. Intell.1
1998 Inducing Relational Concepts with Neural Networks via the LINUS System
Rodrigo Basilio, Gerson Zaverucha, Artur S. d'Avila Garcez
ICONIP3