EDBT 2026 Demo / reviewers in the wild / expert
Thomas Lukasiewicz
dblp:l/ThomasLukasiewicz
· DBLP profile ↗
211ranked-venue papers
51as first author
85since 2021 · last 2026
0000-0002-7644-1668ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 165 · 41 first-author · 65 since 2021Graphics, computer vision, multimedia, augmented reality and games · 61 · 11 first-author · 25 since 2021Theory of computation · 24 · 10 first-author · 1 since 2021Databases, data management, data science and information retrieval · 22 · 9 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 12 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GraphSynth: Resolving the Diversity-Reliability Trade-off with Probabilistic Factor GraphsabstractThe large language models offer a scaleable solution for the generation of synthetic data faced with a trade-off between maintaining the diversity of generation and achieving factually accurate results.This paper introduces Graphsynth, a framework which leverages a probabilistic factor graph modeling the universe of attributes.The framework leverages a high-level schema mapping compiled into efficient hard masks during the decoding phase for maintaining the syntactic truth and a span-synchronized verifier for dismissing logical contradictions at the decode time.The experiments conducted on biomedical, legal, and generic domains show that the method outperforms the state-of-the-art baselines with a structural integrity approaching perfection, a coverage of around 94% attributes on the factor graph solution, and a boost in performance on downstream tasks such as +17.9% on TruthfulQA. Zehua Cheng, Wei Dai 0015, Thomas Lukasiewicz |
ACL (1) | 4 |
| 2026 | AttCL-GAN: Attentional contrastive learning-based generative adversarial network for modality completion of medical images
Zhenghua Xu 0001, Jiaqi Tang 0013, Thomas Lukasiewicz |
Knowl. Based Syst. | 5 |
| 2026 | Semi-supervised medical image lesion detection based on multi-head feature fusion
Zhenghua Xu 0001, Hexiang Zhang, Runhe Yang, Weipeng Liu, Thomas Lukasiewicz |
Knowl. Based Syst. | 5 |
| 2026 | Advancing federated semi-supervised medical image segmentation: A duo of interactive denoising pseudo-labels and convolutional contrastive learning
Zhenghua Xu 0001, Bo Li 0099, Gaoxi Zhou, Xianglin Lu, Thomas Lukasiewicz |
Medical Image Anal. | 6 |
| 2026 | A survey on neuro-mimetic deep learning via predictive codingabstractArtificial intelligence (AI) is rapidly becoming one of the key technologies of this century. The majority of results in AI thus far have been achieved using deep neural networks trained with a learning algorithm called error backpropagation, always considered biologically implausible. To this end, recent works have studied learning algorithms for deep neural networks inspired by the neurosciences. One such theory, called predictive coding (PC), has shown promising properties that make it potentially valuable for the machine learning community: it can model information processing in different areas of the brain, can be used in control and robotics, has a solid mathematical foundation in variational inference, and performs its computations asynchronously. Inspired by such properties, works that propose novel PC-like algorithms are starting to be present in multiple sub-fields of machine learning and AI at large. Here, we survey such efforts by first providing a broad overview of the history of PC to provide common ground for the understanding of the recent developments, then by describing current efforts and results, and concluding with a large discussion of possible implications and ways forward. Tommaso Salvatori, Ankur Mali, Christopher L. Buckley, Thomas Lukasiewicz, Rajesh P. N. Rao, Karl J. Friston, Alexander Ororbia |
Neural Networks | 4 |
| 2026 | You Need Glimpse Before Segmentation: Stochastic Detector-Actor-Critic for Medical Image SegmentationabstractMedical images often contain more redundant background areas than natural images, potentially introducing noise and degrading image segmentation performance. Inspired by doctors' diagnostic processes, where they identify the lesion area before conducting a detailed analysis, we introduce a novel Stochastic Detector-Actor-Critic (SDAC) framework to tackle this challenge. SDAC initially glimpses the entire image using a detector network and policy gradient algorithms to filter out irrelevant background regions and focus on crucial, smaller areas for segmentation. The Actor-Critic algorithm then dynamically creates segmentation masks pixel by pixel without user intervention or coarse masks, forming a robust segmentation module. Both processes are trained jointly to reduce error propagation and ensure stability and ease of implementation. Our experiments on two commonly used medical image segmentation datasets demonstrate that SDAC achieves competitive results comparable to state-of-the-art methods while using 10x fewer parameters than the best-performing baseline in terms of DICE and IoU metrics. We also conduct detailed ablation studies to enhance understanding and facilitate practical use. Furthermore, SDAC performs well in low-resource settings (i.e., 50-shot or 100-shot), making it ideal for real-world scenarios. Its lightweight design make SDAC an excellent baseline for medical image segmentation tasks. Zhenghua Xu 0001, Bo Li 0099, Weipeng Liu, Thomas Lukasiewicz |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | AMLP: Adjustable Masking Lesion Patches for Self-Supervised Medical Image SegmentationabstractSelf-supervised masked image modeling (MIM) methods have shown promising performances on analyzing natural images. However, directly applying such methods to medical image segmentation tasks still cannot achieve satisfactory results. The challenges arise from the facts that (i) medical images are inherently more complex compared to natural images, and the subjects in medical images often exhibit more distinct contour features; (ii) moreover, the conventional high and fixed masking ratio in MIM is likely to mask the background, limiting the scope of learnable information. To address these problems, we propose a new self-supervised medical image segmentation framework, called Adjustable Masking Lesion Patches (AMLP), which employs Masked Patch Selection (MPS) strategy to identify patches with high probabilities of containing lesions to help model achieve precise lesion reconstruction. To improve the categorization of patches in MPS, we further introduce Relative Reconstruction Loss (RRL) to better learn hard-to-reconstruct lesion patches. Then, Category Consistency Loss (CCL) is proposed to refine patch categorization based on reconstruction difficulty, enhancing difference between lesions and backgrounds. Moreover, an Adjustable Masking Ratio (AMR) strategy is proposed to gradually increase the masking ratio over training to expand the scope of learnable mutual information. Extensive experiments on two medical segmentation datasets demonstrate the superior performances of the proposed AMLP w.r.t. the SOTA self-supervised methods; the results prove that AMLP effectively addresses the challenges of applying masked modeling to medical images and capturing accurate lesion details that are crucial for segmentation tasks. Xiangtao Wang, Thomas Lukasiewicz, Zhenghua Xu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Shh, don't say that! Domain Certification in LLMsabstractLarge language models (LLMs) are often deployed to do constrained tasks, with narrow domains. For example, customer support bots can be built on top of LLMs, relying on their broad language understanding and capabilities to enhance performance. However, these LLMs are adversarially susceptible, potentially generating outputs outside the intended domain. To formalize, assess and mitigate this risk, we introduce domain certification; a guarantee that accurately characterizes the out-of-domain behavior of language models. We then propose a simple yet effective approach dubbed VALID that provides adversarial bounds as a certificate. Finally, we evaluate our method across a diverse set of datasets, demonstrating that it yields meaningful certificates. Cornelius Emde, Alasdair Paren, Preetham Arvind, Maxime Kayser, Tom Rainforth, Thomas Lukasiewicz, Philip Torr 0001, Adel Bibi |
ICLR | 6 |
| 2025 | Towards Certification of Uncertainty Calibration under Adversarial AttacksabstractSince neural classifiers are known to be sensitive to adversarial perturbations that alter their accuracy, certification methods have been developed to provide provable guarantees on the insensitivity of their predictions to such perturbations. On the other hand, in safety-critical applications, the frequentist interpretation of the confidence of a classifier (also known as model calibration) can be of utmost importance. This property can be measured via the Brier score or the expected calibration error. We show that attacks can significantly harm calibration, and thus propose certified calibration providing worst-case bounds on calibration under adversarial perturbations. Specifically, we produce analytic bounds for the Brier score and approximate bounds via the solution of a mixed-integer program on the expected calibration error. Finally, we propose novel calibration attacks and demonstrate how they can improve model calibration through adversarial calibration training. The code will be publicly released upon acceptance. Cornelius Emde, Francesco Pinto, Thomas Lukasiewicz, Philip Torr 0001, Adel Bibi |
ICLR | 3 |
| 2025 | Benchmarking Predictive Coding Networks - Made SimpleabstractIn this work, we tackle the problems of efficiency and scalability for predictive coding networks (PCNs) in machine learning. To do so, we propose a library that focuses on performance and simplicity, and use it to implement a large set of standard benchmarks for the community to use for their experiments. As most works in the field propose their own tasks and architectures, do not compare one against each other, and focus on small-scale tasks, a simple and fast open-source library, and a comprehensive set of benchmarks, would address all of these concerns. Then, we perform extensive tests on such benchmarks using both existing algorithms for PCNs, as well as adaptations of other methods popular in the bio-plausible deep learning community. All of this has allowed us to (i) test architectures much larger than commonly used in the literature, on more complex datasets; (ii) reach new state-of-the-art results in all of the tasks and dataset provided; (iii) clearly highlight what the current limitations of PCNs are, allowing us to state important future research directions. With the hope of galvanizing community efforts towards one of the main open problems in the field, scalability, we will release the code, tests, and benchmarks. Luca Pinchetti, Chang Qi, Oleh Lokshyn, Cornelius Emde, Amine M'Charrak, Mufeng Tang, Simon Frieder, Bayar Menzat, Gaspard Oliviers, Rafal Bogacz, Thomas Lukasiewicz, Tommaso Salvatori |
ICLR | 11 |
| 2025 | Effective and Efficient Medical Image Segmentation with Hierarchical Context InteractionabstractThe U-Net models have become the predominant architecture within the domain of medical image segmentation. Recent advancements have showcased the potential of incorporating attention-based techniques into U-Net structures. Nevertheless, the inclusion of attention mechanisms often leads to a substantial increase in both computational demands and the number of parameters, with only a marginal improvement in the performance. This observation raises a critical evaluation of the efficiency associated with the integration of attention modules. In this paper, we propose a novel methodology termed Hierarchical Context Interaction (HCI), a parameter-efficient, attention-free enhancement that can be seamlessly incorporated into U-Net-based models. Experimental results demonstrate that our proposed HCI module attains state-of-the-art performance on two widely used benchmarks, i.e. Medical Segmentation Decathlon Datasets and Synapse Datasets, while concurrently sustaining a computationally efficient profile comparable to conventional U-Net configurations. Zehua Cheng, Wenhu Zhang, Thomas Lukasiewicz |
WACV | 4 |
| 2025 | Explanations for query answers under existential rulesabstractOntology-based data access is an extensively studied paradigm aiming at improving query answers with the use of an "ontology". An ontology is a specification of a domain of interest, which, in this context, is described via a logical theory. As a form of logical entailment, ontology-mediated query answering is fully interpretable, which makes it possible to derive explanations for ontological query answers. This is a quite important aspect, as the fact that many recent AI systems mostly operating as black boxes has led to some serious concerns. In the literature, various works on explanations in the context of description logics (DLs) have appeared, mostly focusing on explaining concept subsumption and concept unsatisfiability in the ontologies. Some works on explaining query entailment in DLs have appeared as well, however, mainly dealing with inconsistency-tolerant semantics and, actually, non-entailment of the queries. Surprisingly, explaining ontological query entailment has received little attention for ontology languages based on existential rules. In fact, although DLs are popular formalisms to model ontologies, it is generally agreed that rule-based ontologies are well-suited for data-intensive applications, as they allow us to conveniently deal with higher-arity relations, which naturally occur in standard relational databases. The goal of this work is to close this gap, and study the problem of explaining query entailment in the context of existential rules ontologies in terms of minimal subsets of database facts. We provide a thorough complexity analysis for several decision problems associated with minimal explanations for various classes of existential rules, and for different complexity measures. Ismail Ilkan Ceylan, Thomas Lukasiewicz, Enrico Malizia, Andrius Vaicenavicius |
Artif. Intell. | 2 |
| 2025 | Aggregated Mutual Learning between CNN and Transformer for semi-supervised medical image segmentation
Zhenghua Xu 0001, Hening Wang, Runhe Yang, Weipeng Liu, Thomas Lukasiewicz |
Knowl. Based Syst. | 6 |
| 2025 | Self-Supervised Medical Image Segmentation Using Deep Reinforced Adaptive MaskingabstractSelf-supervised learning aims to learn transferable representations from unlabeled data for downstream tasks. Inspired by masked language modeling in natural language processing, masked image modeling (MIM) has achieved certain success in the field of computer vision, but its effectiveness in medical images remains unsatisfactory. This is mainly due to the high redundancy and small discriminative regions in medical images compared to natural images. Therefore, this paper proposes an adaptive hard masking (AHM) approach based on deep reinforcement learning to expand the application of MIM in medical images. Unlike predefined random masks, AHM uses an asynchronous advantage actor-critic (A3C) model to predict reconstruction loss for each patch, enabling the model to learn where masking is valuable. By optimizing the non-differentiable sampling process using reinforcement learning, AHM enhances the understanding of key regions, thereby improving downstream task performance. Experimental results on two medical image datasets demonstrate that AHM outperforms state-of-the-art methods. Additional experiments under various settings validate the effectiveness of AHM in constructing masked images. Zhenghua Xu 0001, Thomas Lukasiewicz |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Hybrid Reinforced Medical Report Generation With M-Linear Attention and Repetition PenaltyabstractTo reduce doctors' workload, deep-learning-based automatic medical report generation has recently attracted more and more research efforts, where deep convolutional neural networks (CNNs) are employed to encode the input images, and recurrent neural networks (RNNs) are used to decode the visual features into medical reports automatically. However, these state-of-the-art methods mainly suffer from three shortcomings: 1) incomprehensive optimization; 2) low-order and unidimensional attention; and 3) repeated generation. In this article, we propose a hybrid reinforced medical report generation method with m-linear attention and repetition penalty mechanism (HReMRG-MR) to overcome these problems. Specifically, a hybrid reward with different weights is employed to remedy the limitations of single-metric-based rewards, and a local optimal weight search algorithm is proposed to significantly reduce the complexity of searching the weights of the rewards from exponential to linear. Furthermore, we use m-linear attention modules to learn multidimensional high-order feature interactions and to achieve multimodal reasoning, while a new repetition penalty is proposed to apply penalties to repeated terms adaptively during the model's training process. Extensive experimental studies on two public benchmark datasets show that HReMRG-MR greatly outperforms the state-of-the-art baselines in terms of all metrics. The effectiveness and necessity of all components in HReMRG-MR are also proved by ablation studies. Additional experiments are further conducted and the results demonstrate that our proposed local optimal weight search algorithm can significantly reduce the search time while maintaining superior medical report generation performances. Zhenghua Xu 0001, Wenting Xu, Junyang Chen 0001, Chang Qi, Thomas Lukasiewicz |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | The Defeat of the Winograd Schema Challenge (Abstract Reprint)abstractThe Winograd Schema Challenge—a set of twin sentences involving pronoun reference disambiguation that seem to require the use of commonsense knowledge—was proposed by Hector Levesque in 2011. By 2019, a number of AI systems, based on large pre-trained transformer-based language models and fine-tuned on these kinds of problems, achieved better than 90% accuracy. In this paper, we review the history of the Winograd Schema Challenge and discuss the lasting contributions of the flurry of research that has taken place on the WSC in the last decade. We discuss the significance of various datasets developed for WSC, and the research community's deeper understanding of the role of surrogate tasks in assessing the intelligence of an AI system. Vid Kocijan, Ernest Davis, Thomas Lukasiewicz, Gary Marcus 0001, Leora Morgenstern |
AAAI | 3 |
| 2024 | Hard Regularization to Prevent Deep Online Clustering Collapse without Data AugmentationabstractOnline deep clustering refers to the joint use of a feature extraction network and a clustering model to assign cluster labels to each new data point or batch as it is processed. While faster and more versatile than offline methods, online clustering can easily reach the collapsed solution where the encoder maps all inputs to the same point and all are put into a single cluster. Successful existing models have employed various techniques to avoid this problem, most of which require data augmentation or which aim to make the average soft assignment across the dataset the same for each cluster. We propose a method that does not require data augmentation, and that, differently from existing methods, regularizes the hard assignments. Using a Bayesian framework, we derive an intuitive optimization objective that can be straightforwardly included in the training of the encoder network. Tested on four image datasets, it consistently avoids collapse more robustly than other methods and leads to more accurate clustering. We also conduct further experiments and analyses justifying our choice to regularize the hard cluster assignments. Code is available at https://github.com/Lou1sM/online_hard_clustering. Louis Mahon, Thomas Lukasiewicz |
AAAI | 2 |
| 2024 | Affinity-Graph-Guided Contractive Learning for Pretext-Free Medical Image Segmentation with Minimal AnnotationabstractThe combination of semi-supervised learning (SemiSL) and contrastive learning (CL) has been successful in medical image segmentation with limited annotations. However, these works often rely on pretext tasks that lack the specificity required for pixel-level segmentation, and still face overfitting issues due to insufficient supervision signals resulting from too few annotations. Therefore, this paper proposes an affinitygraph-guided semi-supervised contrastive learning framework (Semi-AGCL) by establishing additional affinity-graph-based supervision signals between the student and teacher network, to achieve medical image segmentation with minimal annotations without pretext. The framework first designs an average-patchentropy-driven inter-patch sampling method, which can provide a robust initial feature space without relying on pretext tasks. Furthermore, the framework designs an affinity-graph-guided loss function, which can improve the quality of the learned representation and the model’s generalization ability by exploiting the inherent structure of the data, thus mitigating overfitting. Our experiments indicate that with merely 10% of the complete annotation set, our model approaches the accuracy of the fully annotated baseline, manifesting a marginal deviation of only 2.52%. Under the stringent conditions where only 5% of the annotations are employed, our model exhibits a significant enhancement in performance—surpassing the second-best baseline by 23.09% on the dice metric and achieving an improvement of 26.57% on the notably arduous CRAG and ACDC datasets. Zehua Cheng, Thomas Lukasiewicz |
BIBM | 3 |
| 2024 | Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support SettingabstractMaxime Kayser, Bayar Menzat, Cornelius Emde, Bogdan Bercean, Alex Novak, Abdala Espinosa, Bartlomiej W. Papiez, Susanne Gaube, Thomas Lukasiewicz, Oana-Maria Camburu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Maxime Kayser, Bayar Menzat, Cornelius Emde, Bogdan Bercean, Alex Novak, Abdalá Morgado, Bartlomiej Wladyslaw Papiez, Susanne Gaube, Thomas Lukasiewicz, Oana-Maria Camburu |
EMNLP | 9 |
| 2024 | A Stable, Fast, and Fully Automatic Learning Algorithm for Predictive Coding NetworksabstractPredictive coding networks are neuroscience-inspired models with roots in both Bayesian statistics and neuroscience. Training such models, however, is quite inefficient and unstable. In this work, we show how by simply changing the temporal scheduling of the update rule for the synaptic weights leads to an algorithm that is much more efficient and stable than the original one, and has theoretical guarantees in terms of convergence. The proposed algorithm, that we call incremental predictive coding (iPC) is also more biologically plausible than the original one, as it it fully automatic. In an extensive set of experiments, we show that iPC constantly performs better than the original formulation on a large number of benchmarks for image classification, as well as for the training of both conditional and masked language models, in terms of test accuracy, efficiency, and convergence with respect to a large set of hyperparameters. Tommaso Salvatori, Yuhang Song 0001, Yordan Yordanov, Beren Millidge, Lei Sha, Cornelius Emde, Zhenghua Xu 0001, Rafal Bogacz, Thomas Lukasiewicz |
ICLR | 9 |
| 2024 | How Realistic Is Your Synthetic Data? Constraining Deep Generative Models for Tabular DataabstractDeep Generative Models (DGMs) have been shown to be powerful tools for generating tabular data, as they have been increasingly able to capture the complex distributions that characterize them. However, to generate realistic synthetic data, it is often not enough to have a good approximation of their distribution, as it also requires compliance with constraints that encode essential background knowledge on the problem at hand. In this paper, we address this limitation and show how DGMs for tabular data can be transformed into Constrained Deep Generative Models (C-DGMs), whose generated samples are guaranteed to be compliant with the given constraints. This is achieved by automatically parsing the constraints and transforming them into a Constraint Layer (CL) seamlessly integrated with the DGM. Our extensive experimental analysis with various DGMs and tasks reveals that standard DGMs often violate constraints, some exceeding 95% non-compliance, while their corresponding C-DGMs are never non-compliant. Then, we quantitatively demonstrate that, at training time, C-DGMs are able to exploit the background knowledge expressed by the constraints to outperform their standard counterparts with up to 4.5% improvement in utility and detection. Further, we show how our CL does not necessarily need to be integrated at training time, as it can be also used as a guardrail at inference time, still producing some improvements in the overall performance of the models. Finally, we show that our CL does not hinder the sample generation time of the models. Mihaela Catalina Stoian, Salijona Dyrmishi, Maxime Cordy, Thomas Lukasiewicz, Eleonora Giunchiglia |
ICLR | 4 |
| 2024 | Predictive Coding beyond CorrelationsabstractBiologically plausible learning algorithms offer a promising alternative to traditional deep learning techniques, especially in overcoming the limitations of backpropagation in fast and low-energy neuromorphic implementations. To this end, there has been extensive research in understanding what their capabilities are. In this work, we show how one of such algorithms, called predictive coding, is able to perform causal inference tasks. First, we show how a simple change in the inference process of predictive coding enables to compute interventions without the need to mutilate or redefine a causal graph. Then, we explore applications in cases where the graph is unknown, and has to be inferred from observational data. Empirically, we show how such findings can be used to improve the performance of predictive coding in image classification tasks, and conclude that such models are naturally able to perform causal inference tasks using a biologically plausible kind of message passing. Tommaso Salvatori, Luca Pinchetti, Amine M'Charrak, Beren Millidge, Thomas Lukasiewicz |
ICML | 5 |
| 2024 | PiShield: A PyTorch Package for Learning with Requirements
Mihaela Catalina Stoian, Alex Tatomir, Thomas Lukasiewicz, Eleonora Giunchiglia |
IJCAI | 3 |
| 2024 | Pre-training and diagnosing knowledge base completion models
Vid Kocijan, Myeongjun Jang 0001, Thomas Lukasiewicz |
Artif. Intell. | 3 |
| 2024 | CCN+: A neuro-symbolic framework for deep learning with requirementsabstractFor their outstanding ability of finding hidden patterns in data, deep learning models have been extensively applied in many different domains. However, recent works have shown that, if a set of requirements expressing inherent knowledge about the problem at hand is given, then neural networks often fail to comply with them. This represents a major drawback for deep learning models, as requirements compliance is normally considered a necessary condition for standard software deployment. In this paper, we propose a novel neuro-symbolic framework able to make any neural network compliant by design to a given set of requirements over the output space expressed in full propositional logic. This framework, called CCN+, integrates the requirements into the output layer of the neural network by applying multiple inference rules that ensure compliance with the requirements and adapts the standard binary cross-entropy loss function to the requirement output layer. As a result, not only the outputted predictions are guaranteed to be compliant with the requirements, but the neural network itself learns how to exploit the domain knowledge expressed by the requirements to get better performance. We conduct an extensive experimental evaluation of CCN+ on 19 real-world multi-label classification datasets with propositional logic requirements, including a challenging dataset for autonomous driving. Our experimental analysis confirms that CCN+ is able to outperform both its neural counterparts and the state-of-the-art models. Eleonora Giunchiglia, Alex Tatomir, Mihaela Catalina Stoian, Thomas Lukasiewicz |
Int. J. Approx. Reason. | 4 |
| 2024 | Minimum description length clustering to measure meaningful image complexityabstractWe present a new image complexity metric. Existing complexity metrics cannot distinguish meaningful content from noise, and give a high score to white noise images, which contain no meaningful information. We use the minimum description length principle to determine the number of clusters and designate certain points as outliers and, hence, correctly assign white noise a low score. The presented method is a step towards humans’ ability to detect when data contain a meaningful pattern. It also has similarities to theoretical ideas for measuring meaningful complexity. We conduct experiments on seven different sets of images, which show that our method assigns the most accurate scores to all images considered. Additionally, comparing the different levels of the hierarchy of clusters can reveal how complexity manifests at different scales, from local detail to global structure. We then present ablation studies showing the contribution of the components of our method, and that it continues to assign reasonable scores when the inputs are modified in certain ways, including the addition of Gaussian noise and the lowering of the resolution. Louis Mahon, Thomas Lukasiewicz |
Pattern Recognit. | 2 |
| 2024 | Text Attribute Control via Closed-Loop DisentanglementabstractAbstract Changing an attribute of a text without changing the content usually requires first disentangling the text into irrelevant attributes and content representations. After that, in the inference phase, the representation of one attribute is tuned to a different value, expecting that the corresponding attribute of the text can also be changed accordingly. The usual way of disentanglement is to add some constraints on the latent space of an encoder-decoder architecture, including adversarial-based constraints and mutual-information-based constraints. However, previous semi-supervised processes of attribute change are usually not enough to guarantee the success of attribute change and content preservation. In this paper, we propose a novel approach to achieve a robust control of attributes while enhancing content preservation. In this approach, we use a semi-supervised contrastive learning method to encourage the disentanglement of attributes in latent spaces. Differently from previous works, we re-disentangle the reconstructed sentence and compare the re-disentangled latent space with the original latent space, which makes a closed-loop disentanglement process. This also helps content preservation. In addition, the contrastive learning method is also able to replace the role of minimizing mutual information and adversarial training in the disentanglement process, which alleviates the computation cost. We conducted experiments on three text datasets, including the Yelp Service review dataset, the Amazon Product review dataset, and the GoEmotions dataset. The experimental results show the effectiveness of our model. Lei Sha, Thomas Lukasiewicz |
Trans. Assoc. Comput. Linguistics | 2 |
| 2024 | Multi-ConDoS: Multimodal Contrastive Domain Sharing Generative Adversarial Networks for Self-Supervised Medical Image SegmentationabstractExisting self-supervised medical image segmentation usually encounters the domain shift problem (i.e., the input distribution of pre-training is different from that of fine-tuning) and/or the multimodality problem (i.e., it is based on single-modal data only and cannot utilize the fruitful multimodal information of medical images). To solve these problems, in this work, we propose multimodal contrastive domain sharing (Multi-ConDoS) generative adversarial networks to achieve effective multimodal contrastive self-supervised medical image segmentation. Compared to the existing self-supervised approaches, Multi-ConDoS has the following three advantages: (i) it utilizes multimodal medical images to learn more comprehensive object features via multimodal contrastive learning; (ii) domain translation is achieved by integrating the cyclic learning strategy of CycleGAN and the cross-domain translation loss of Pix2Pix; (iii) novel domain sharing layers are introduced to learn not only domain-specific but also domain-sharing information from the multimodal medical images. Extensive experiments on two publicly multimodal medical image segmentation datasets show that, with only 5% (resp., 10%) of labeled data, Multi-ConDoS not only greatly outperforms the state-of-the-art self-supervised and semi-supervised medical image segmentation baselines with the same ratio of labeled data, but also achieves similar (sometimes even better) performances as fully supervised segmentation methods with 50% (resp., 100%) of labeled data, which thus proves that our work can achieve superior segmentation performances with very low labeling workload. Furthermore, ablation studies prove that the above three improvements are all effective and essential for Multi-ConDoS to achieve this very superior performance. Shuo Zhang 0017, Xiaoqian Shen, Thomas Lukasiewicz, Zhenghua Xu 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2023 | An Empirical Analysis of Parameter-Efficient Methods for Debiasing Pre-Trained Language ModelsabstractThe increasingly large size of modern pretrained language models not only makes them inherit more human-like biases from the training corpora, but also makes it computationally expensive to mitigate such biases.In this paper, we investigate recent parameter-efficient methods in combination with counterfactual data augmentation (CDA) for bias mitigation.We conduct extensive experiments with prefix tuning, prompt tuning, and adapter tuning on different language models and bias types to evaluate their debiasing performance and abilities to preserve the internal knowledge of a pre-trained model.We find that the parameter-efficient methods (i) are effective in mitigating gender bias, where adapter tuning is consistently the most effective one and prompt tuning is more suitable for GPT-2 than BERT, (ii) are less effective when it comes to racial and religious bias, which may be attributed to the limitations of CDA, and (iii) can perform similarly to or sometimes better than full fine-tuning with improved time and memory efficiency, as well as maintain the internal knowledge in BERT and GPT-2, evaluated via fact retrieval and downstream fine-tuning. Zhongbin Xie, Thomas Lukasiewicz |
ACL (1) | 2 |
| 2023 | Counter-GAP: Counterfactual Bias Evaluation through Gendered Ambiguous PronounsabstractBias-measuring datasets play a critical role in detecting biased behavior of language models and in evaluating progress of bias mitigation methods.In this work, we focus on evaluating gender bias through coreference resolution, where previous datasets are either hand-crafted or fail to reliably measure an explicitly defined bias.To overcome these shortcomings, we propose a novel method to collect diverse, natural, and minimally distant text pairs via counterfactual generation, and construct Counter-GAP, an annotated dataset consisting of 4008 instances grouped into 1002 quadruples.We further identify a bias cancellation problem in previous group-level metrics on Counter-GAP, and propose to use the difference between inconsistency across genders and within genders to measure bias at a quadruple level.Our results show that four pre-trained language models are significantly more inconsistent across different gender groups than within each group, and that a name-based counterfactual data augmentation method is more effective to mitigate such bias than an anonymization-based method. Zhongbin Xie, Vid Kocijan, Thomas Lukasiewicz, Oana-Maria Camburu |
EACL | 3 |
| 2023 | Associative Memories in the Feature SpaceabstractAn autoassociative memory model is a function that, given a set of data points, takes as input an arbitrary vector and outputs the most similar data point from the memorized set. However, popular memory models fail to retrieve images even when the corruption is mild and easy to detect for a human evaluator. This is because similarities are evaluated in the raw pixel space, which does not contain any semantic information about the images. This problem can be easily solved by computing similarities in an embedding space instead of the pixel space. We show that an effective way of computing such embeddings is via a network pretrained with a contrastive loss. As the dimension of embedding spaces is often significantly smaller than the pixel space, we also have a faster computation of similarity scores. We test this method on complex datasets such as CIFAR10 and STL10. An additional drawback of current models is the need of storing the whole dataset in the pixel space, which is often extremely large. We relax this condition and propose a class of memory models that only stores low-dimensional semantic embeddings, and uses them to retrieve similar, but not identical, memories. We demonstrate a proof of concept of this method on a simple task on the MNIST dataset. Tommaso Salvatori, Beren Millidge, Yuhang Song 0001, Rafal Bogacz, Thomas Lukasiewicz |
ECAI | 5 |
| 2023 | Improving Language Models' Meaning Understanding and Consistency by Learning Conceptual Roles from DictionaryabstractThe non-humanlike behaviour of contemporary pre-trained language models (PLMs) is a leading cause undermining their trustworthiness.A striking phenomenon of such faulty behaviours is the generation of inconsistent predictions, which produces logically contradictory results, such as generating different predictions for texts delivering the same meaning or violating logical properties.Previous studies exploited data augmentation or implemented specialised loss functions to alleviate the issue.However, their usage is limited, because they consume expensive training resources for largesized PLMs and can only handle a certain consistency type.To this end, we propose a practical approach that alleviates the inconsistent behaviour issue by fundamentally improving PLMs' meaning awareness.Based on the conceptual role theory, our method allows PLMs to capture accurate meaning by learning precise interrelationships between concepts from word-definition pairs in a dictionary.Next, we propose an efficient parameter integration technique that updates only a few additional parameters to combine the learned interrelationship with PLMs' pre-trained knowledge.Our experimental results reveal that the approach can concurrently improve multiple types of consistency, enables efficient knowledge integration, and easily applies to other languages. Myeongjun Jang 0001, Thomas Lukasiewicz |
EMNLP | 2 |
| 2023 | Consistency Analysis of ChatGPTabstractChatGPT has gained a huge popularity since its introduction.Its positive aspects have been reported through many media platforms, and some analyses even showed that ChatGPT achieved a decent grade in professional exams, adding extra support to the claim that AI can now assist and even replace humans in industrial fields.Others, however, doubt its reliability and trustworthiness.This paper investigates the trustworthiness of ChatGPT and GPT-4 regarding logically consistent behaviour, focusing specifically on semantic consistency and the properties of negation, symmetric, and transitive consistency.Our findings suggest that while both models appear to show an enhanced language understanding and reasoning ability, they still frequently fall short of generating logically consistent predictions.We also ascertain via experiments that prompt designing, fewshot learning and employing larger large language models (LLMs) are unlikely to be the ultimate solution to resolve the inconsistency issue of LLMs. Myeongjun Jang 0001, Thomas Lukasiewicz |
EMNLP | 2 |
| 2023 | MPS-AMS: Masked Patches Selection and Adaptive Masking Strategy Based Self-Supervised Medical Image SegmentationabstractExisting self-supervised learning methods based on contrastive learning and masked image modeling have demonstrated impressive performances. However, current masked image modeling methods are mainly utilized in natural images, and their applications in medical images are relatively lacking. Besides, their fixed high masking strategy limits the upper bound of conditional mutual information, and the gradient noise is considerable, making less the learned representation information. Motivated by these limitations, in this paper, we propose masked patches selection and adaptive masking strategy based self-supervised medical image segmentation method, named MPS-AMS. We leverage the masked patches selection strategy to choose masked patches with lesions to obtain more lesion representation information, and the adaptive masking strategy is utilized to help learn more mutual information and improve performance further. Extensive experiments on three public medical image segmentation datasets (BUSI, Hecktor, and Brats2018) show that our proposed method greatly outperforms the state-of-the-art self-supervised baselines. Xiangtao Wang, Shuo Zhang 0017, Junyang Chen 0001, Thomas Lukasiewicz, Zhenghua Xu 0001 |
ICASSP | 7 |
| 2023 | MvCo-DoT: Multi-View Contrastive Domain Transfer Network for Medical Report GenerationabstractIn clinical scenarios, multiple medical images with different views are usually generated at the same time, and they have high semantic consistency. However, the existing medical report generation methods cannot exploit the rich multi-view mutual information of medical images. Therefore, in this work, we propose the first multi-view medical report generation model, called MvCo-DoT. Specifically, MvCo-DoT first propose a multi-view contrastive learning (MvCo) strategy to help the deep reinforcement learning based model utilize the consistency of multi-view inputs for better model learning. Then, to close the performance gaps of using multi-view and single-view inputs, a domain transfer network is further proposed to ensure MvCo-DoT achieve almost the same performance as multi-view inputs using only single-view inputs. Extensive experiments on the IU X-Ray public dataset show that MvCo-DoT outperforms the SOTA medical report generation baselines in all metrics. Xiangtao Wang, Zhenghua Xu 0001, Wenting Xu, Junyang Chen 0001, Thomas Lukasiewicz |
ICASSP | 6 |
| 2023 | Multi-Head Feature Pyramid Networks for Breast Mass DetectionabstractAnalysis of X-ray images is one of the main tools to diagnose breast cancer. The ability to quickly and accurately detect the location of masses from the huge amount of image data is the key to reducing the morbidity and mortality of breast cancer. Currently, the main factor limiting the accuracy of breast mass detection is the unequal focus on the mass boxes, leading the network to focus too much on larger masses at the expense of smaller ones. In the paper, we propose the multi-head feature pyramid module (MHFPN) to solve the problem of unbalanced focus of target boxes during feature map fusion and design a multi-head breast mass detection network (MBMDnet). Experimental studies show that, comparing to the SOTA detection baselines, our method improves by 6.58% (in AP@50) and 5.4% (in TPR@50) on the commonly used IN-breast dataset, while about 6-8% improvements (in AP@20) are also observed on the public MIAS and BCS-DBT datasets. Hexiang Zhang, Zhenghua Xu 0001, Shuo Zhang 0017, Junyang Chen 0001, Thomas Lukasiewicz |
ICASSP | 6 |
| 2023 | Backpropagation at the Infinitesimal Inference Limit of Energy-Based Models: Unifying Predictive Coding, Equilibrium Propagation, and Contrastive Hebbian Learning
Beren Millidge, Yuhang Song 0001, Tommaso Salvatori, Thomas Lukasiewicz, Rafal Bogacz |
ICLR | 4 |
| 2023 | A Theoretical Framework for Inference and Learning in Predictive Coding Networks
Beren Millidge, Yuhang Song 0001, Tommaso Salvatori, Thomas Lukasiewicz, Rafal Bogacz |
ICLR | 4 |
| 2023 | Adaptive-Masking Policy with Deep Reinforcement Learning for Self-Supervised Medical Image SegmentationabstractAlthough self-supervised learning methods based on masked image modeling have achieved some success in improving the performance of deep learning models, these methods have difficulty in ensuring that the masked region is the most appropriate for each image, resulting in segmentation networks that do not get the best weights in pre-training. Therefore, we propose a new adaptive-masking policy self-supervised learning method. Specifically, we model the process of masking images as a reinforcement learning problem and use the results of the reconstruction model as a feedback signal to guide the agent to learn the masking policy to select a more appropriate mask position and size for each image, helping the reconstruction network to learn more fine-grained image representation information and thus improve the downstream segmentation model performance. We conduct extensive experiments on two datasets, Cardiac and TCIA, and the results show that our approach outperforms current state-of-the-art self-supervised learning methods. Shengxin Wang, Thomas Lukasiewicz, Zhenghua Xu 0001 |
ICME | 3 |
| 2023 | NP-SemiSeg: When Neural Processes meet Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation involves assigning pixel-wise labels to unlabeled images at training time. This is useful in a wide range of real-world applications where collecting pixel-wise labels is not feasible in time or cost. Current approaches to semi-supervised semantic segmentation work by predicting pseudo-labels for each pixel from a class-wise probability distribution output by a model. If this predicted probability distribution is incorrect, however, it leads to poor segmentation results which can have knock-on consequences in safety critical systems, like medical images or self-driving cars. It is, therefore, important to understand what a model does not know, which is mainly achieved by uncertainty quantification. Recently, neural processes (NPs) have been explored in semi-supervised image classification, and they have been a computationally efficient and effective method for uncertainty quantification. In this work, we move one step forward by adapting NPs to semi-supervised semantic segmentation, resulting in a new model called NP-SemiSeg. We experimentally evaluated NP-SemiSeg on the public benchmarks PASCAL VOC 2012 and Cityscapes, with different training settings, and the results verify its effectiveness. Daniela Massiceti, Xiaolin Hu 0001, Vladimir Pavlovic 0001, Thomas Lukasiewicz |
ICML | 5 |
| 2023 | Complexity of Inconsistency-Tolerant Query Answering in Datalog+/- under Preferred RepairsabstractInconsistency-tolerant semantics have been proposed to provide meaningful ontological query answers even in the presence of inconsistencies. Several such semantics rely on the notion of a repair, which is a "maximal" consistent subset of the database, where different maximality criteria might be adopted depending on the application at hand. Previous work in the context of Datalog+/- has considered only the subset and cardinality maximality criteria. We take here a step further and study inconsistency-tolerant semantics under maximality criteria based on weights and priority levels. We provide a thorough complexity analysis for a wide range of existential rule languages and for several complexity measures. Thomas Lukasiewicz, Enrico Malizia, Cristian Molinaro |
KR | 1 |
| 2023 | Mathematical Capabilities of ChatGPTabstractWe investigate the mathematical capabilities of two iterations of ChatGPT (released 9-January-2023 and 30-January-2023) and of GPT-4 by testing them on publicly available datasets, as well as hand-crafted ones, using a novel methodology. In contrast to formal mathematics, where large databases of formal proofs are available (e.g., mathlib, the Lean Mathematical Library), current datasets of natural-language mathematics used to benchmark language models either cover only elementary mathematics or are very small. We address this by publicly releasing two new datasets: GHOSTS and miniGHOSTS. These are the first natural-language datasets curated by working researchers in mathematics that (1) aim to cover graduate-level mathematics, (2) provide a holistic overview of the mathematical capabilities of language models, and (3) distinguish multiple dimensions of mathematical reasoning. These datasets test on 1636 human expert evaluations whether ChatGPT and GPT-4 can be helpful assistants to professional mathematicians by emulating use cases that arise in the daily professional activities of mathematicians. We benchmark the models on a range of fine-grained performance metrics. For advanced mathematics, this is the most detailed evaluation effort to date. We find that ChatGPT and GPT-4 can be used most successfully as mathematical assistants for querying facts, acting as mathematical search engines and knowledge base interfaces. GPT-4 can additionally be used for undergraduate-level mathematics but fails on graduate-level difficulty. Contrary to many positive reports in the media about GPT-4 and ChatGPT's exam-solving abilities (a potential case of selection bias), their overall mathematical performance is well below the level of a graduate student. Hence, if you aim to use ChatGPT to pass a graduate-level math exam, you would be better off copying from your average peer! Simon Frieder, Luca Pinchetti, Alexis Chevalier, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Petersen, Julius Berner |
NeurIPS | 6 |
| 2023 | The defeat of the Winograd Schema Challenge
Vid Kocijan, Ernest Davis, Thomas Lukasiewicz, Gary Marcus 0001, Leora Morgenstern |
Artif. Intell. | 3 |
| 2023 | Rationalizing predictions by adversarial information calibration
Lei Sha, Oana-Maria Camburu, Thomas Lukasiewicz |
Artif. Intell. | 3 |
| 2023 | Multi-modal contrastive mutual learning and pseudo-label re-learning for semi-supervised medical image segmentation
Shuo Zhang 0017, Thomas Lukasiewicz, Zhenghua Xu 0001 |
Medical Image Anal. | 4 |
| 2023 | ROAD-R: the autonomous driving dataset with logical requirementsabstractAbstract Neural networks have proven to be very powerful at computer vision tasks. However, they often exhibit unexpected behaviors, acting against background knowledge about the problem at hand. This calls for models (i) able to learn from requirements expressing such background knowledge, and (ii) guaranteed to be compliant with the requirements themselves. Unfortunately, the development of such models is hampered by the lack of real-world datasets equipped with formally specified requirements. In this paper, we introduce the ROad event Awareness Dataset with logical Requirements (ROAD-R), the first publicly available dataset for autonomous driving with requirements expressed as logical constraints. Given ROAD-R, we show that current state-of-the-art models often violate its logical constraints, and that it is possible to exploit them to create models that (i) have a better performance, and (ii) are guaranteed to be compliant with the requirements themselves. Eleonora Giunchiglia, Mihaela Catalina Stoian, Salman Khan 0004, Fabio Cuzzolin, Thomas Lukasiewicz |
Mach. Learn. | 5 |
| 2023 | Recurrent predictive coding models for associative memory employing covariance learningabstractThe computational principles adopted by the hippocampus in associative memory (AM) tasks have been one of the most studied topics in computational and theoretical neuroscience. Recent theories suggested that AM and the predictive activities of the hippocampus could be described within a unitary account, and that predictive coding underlies the computations supporting AM in the hippocampus. Following this theory, a computational model based on classical hierarchical predictive networks was proposed and was shown to perform well in various AM tasks. However, this fully hierarchical model did not incorporate recurrent connections, an architectural component of the CA3 region of the hippocampus that is crucial for AM. This makes the structure of the model inconsistent with the known connectivity of CA3 and classical recurrent models such as Hopfield Networks, which learn the covariance of inputs through their recurrent connections to perform AM. Earlier PC models that learn the covariance information of inputs explicitly via recurrent connections seem to be a solution to these issues. Here, we show that although these models can perform AM, they do it in an implausible and numerically unstable way. Instead, we propose alternatives to these earlier covariance-learning predictive coding networks, which learn the covariance information implicitly and plausibly, and can use dendritic structures to encode prediction errors. We show analytically that our proposed models are perfectly equivalent to the earlier predictive coding model learning covariance explicitly, and encounter no numerical issues when performing AM tasks in practice. We further show that our models can be combined with hierarchical predictive coding networks to model the hippocampo-neocortical interactions. Our models provide a biologically plausible approach to modelling the hippocampal network, pointing to a potential computational mechanism during hippocampal memory formation and recall, which employs both predictive coding and covariance learning based on the recurrent network structure of the hippocampus. Mufeng Tang, Tommaso Salvatori, Beren Millidge, Yuhang Song 0001, Thomas Lukasiewicz, Rafal Bogacz |
PLoS Comput. Biol. | 5 |
| 2023 | Hi-BEHRT: Hierarchical Transformer-Based Model for Accurate Prediction of Clinical Events Using Multimodal Longitudinal Electronic Health RecordsabstractElectronic health records (EHR) represent a holistic overview of patients' trajectories. Their increasing availability has fueled new hopes to leverage them and develop accurate risk prediction models for a wide range of diseases. Given the complex interrelationships of medical records and patient outcomes, deep learning models have shown clear merits in achieving this goal. However, a key limitation of current study remains their capacity in processing long sequences, and long sequence modelling and its application in the context of healthcare and EHR remains unexplored. Capturing the whole history of medical encounters is expected to lead to more accurate predictions, but the inclusion of records collected for decades and from multiple resources can inevitably exceed the receptive field of the most existing deep learning architectures. This can result in missing crucial, long-term dependencies. To address this gap, we present Hi-BEHRT, a hierarchical Transformer-based model that can significantly expand the receptive field of Transformers and extract associations from much longer sequences. Using a multimodal large-scale linked longitudinal EHR, the Hi-BEHRT exceeds the state-of-the-art deep learning models 1% to 5% for area under the receiver operating characteristic (AUROC) curve and 1% to 8% for area under the precision recall (AUPRC) curve on average, and 2% to 8% (AUROC) and 2% to 11% (AUPRC) for patients with long medical history for 5-year heart failure, diabetes, chronic kidney disease, and stroke risk prediction. Additionally, because pretraining for hierarchical Transformer is not well-established, we provide an effective end-to-end contrastive pre-training strategy for Hi-BEHRT using EHR, improving its transferability on predicting clinical events with relatively small training dataset. Yikuan Li, Mohammad Mamouei, Shishir Rao, Abdelaali Hassaïne, Dexter Canoy, Thomas Lukasiewicz, Kazem Rahimi |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | Toward Knowledge as a Service (KaaS): Predicting Popularity of Knowledge Services Leveraging Graph Neural NetworksabstractKnowledge services are becoming a rising star in the family of XaaS (Everything as a Service). In recent years, people are more willing to search for answers and share their knowledge directly over the Internet, which makes the knowledge service ecosystem prosperous. In this paper, we aim to predict the popularity of knowledge services, which will benefit the downstream industries. Toward such a task, the spatial interactions (e.g., hyperlinks in Wikipedia) and temporal observations (e.g., page views) provide crucial information. However, it is difficult to utilize this information due to: (i) complicated and different usage observations, (ii) intricate and evolutionary spatial interactions, and (iii) small world trait of the network. To tackle such issues, we propose evolutionary graph convolutional recurrent neural networks (E-GCRNNs) to simultaneously model both temporal and spatial dependencies of knowledge services from their evolving networks. Additionally, a localized mini-batch training scheme is developed, which allows the E-GCRNNs to work on large-scale knowledge services network and reduce the prediction bias caused by the small world trait. Extensive experiments on real-world datasets have demonstrated that the proposed E-GCRNNs outperform baselines in terms of prediction accuracy, especially with the prediction range being longer, while remaining computationally efficient. Haozhe Lin, Yushun Fan, Jia Zhang 0001, Zhenghua Xu 0001, Thomas Lukasiewicz |
IEEE Trans. Serv. Comput. | 6 |
| 2022 | Reverse Differentiation via Predictive CodingabstractDeep learning has redefined AI thanks to the rise of artificial neural networks, which are inspired by neuronal networks in the brain. Through the years, these interactions between AI and neuroscience have brought immense benefits to both fields, allowing neural networks to be used in a plethora of applications. Neural networks use an efficient implementation of reverse differentiation, called backpropagation (BP). This algorithm, however, is often criticized for its biological implausibility (e.g., lack of local update rules for the parameters). Therefore, biologically plausible learning methods that rely on predictive coding (PC), a framework for describing information processing in the brain, are increasingly studied. Recent works prove that these methods can approximate BP up to a certain margin on multilayer perceptrons (MLPs), and asymptotically on any other complex model, and that zerodivergence inference learning (Z-IL), a variant of PC, is able to exactly implement BP on MLPs. However, the recent literature shows also that there is no biologically plausible method yet that can exactly replicate the weight update of BP on complex models. To fill this gap, in this paper, we generalize (PC and) Z-IL by directly defining it on computational graphs, and show that it can perform exact reverse differentiation. What results is the first PC (and so biologically plausible) algorithm that is equivalent to BP in the way of updating parameters on any neural network, providing a bridge between the interdisciplinary research of neuroscience and deep learning. Furthermore, the above results in particular also immediately provide a novel local and parallel implementation of BP. Tommaso Salvatori, Yuhang Song 0001, Zhenghua Xu 0001, Thomas Lukasiewicz, Rafal Bogacz |
AAAI | 4 |
| 2022 | Efficient Deep Clustering of Human Activities and How to Improve Evaluation
Louis Mahon, Thomas Lukasiewicz |
ACML | 2 |
| 2022 | Image-to-Image Translation with Text Guidance
Bowen Li 0001, Philip Torr 0001, Thomas Lukasiewicz |
BMVC | 3 |
| 2022 | Memory-Driven Text-to-Image Generation
Bowen Li 0001, Philip Torr 0001, Thomas Lukasiewicz |
BMVC | 3 |
| 2022 | BECEL: Benchmark for Consistency Evaluation of Language ModelsabstractBehavioural consistency is a critical condition for a language model (LM) to become trustworthy like humans. Despite its importance, however, there is little consensus on the definition of LM consistency, resulting in different definitions across many studies. In this paper, we first propose the idea of LM consistency based on behavioural consistency and establish a taxonomy that classifies previously studied consistencies into several sub-categories. Next, we create a new benchmark that allows us to evaluate a model on 19 test cases, distinguished by multiple types of consistency and diverse downstream tasks. Through extensive experiments on the new benchmark, we ascertain that none of the modern pre-trained language models (PLMs) performs well in every test case, while exhibiting high inconsistency in many cases. Our experimental results suggest that a unified benchmark that covers broad aspects (i.e., multiple consistency types and tasks) is essential for a more precise evaluation. Myeongjun Jang 0001, Deuk Sin Kwon, Thomas Lukasiewicz |
COLING | 3 |
| 2022 | Rethinking Bayesian Deep Learning Methods for Semi-Supervised Volumetric Medical Image SegmentationabstractRecently, several Bayesian deep learning methods have been proposed for semi-supervised medical image segmentation. Although they have achieved promising results on medical benchmarks, some problems are still existing. Firstly, their overall architectures belong to the discriminative models, and hence, in the early stage of training, they only use labeled data for training, which might make them overfit to the labeled data. Secondly, in fact, they are only partially based on Bayesian deep learning, as their overall architectures are not designed under the Bayesian framework. However, unifying the overall architecture under the Bayesian perspective can make the architecture have a rigorous theoretical basis, so that each part of the architecture can have a clear probabilistic interpretation. Therefore, to solve the problems, we propose a new generative Bayesian deep learning (GBDL) architecture. GBDL belongs to the generative models, whose target is to estimate the joint distribution of input medical volumes and their corresponding labels. Estimating the joint distribution implicitly involves the distribution of data, so both labeled and unlabeled data can be utilized in the early stage of training, which alleviates the potential overfitting problem. Besides, GBDL is completely designed under the Bayesian framework, and thus we give its full Bayesian formulation, which lays a theoretical probabilistic foundation for our architecture. Extensive experiments show that our GBDL outperforms previous state-of-the-art methods in terms of four commonly used evaluation indicators on three public medical datasets. Thomas Lukasiewicz |
CVPR | 2 |
| 2022 | Syntactically Rich Discriminative Training: An Effective Method for Open Information ExtractionabstractOpen information extraction (OIE) is the task of extracting facts "(Subject, Relation, Object)" from natural language text.We propose several new methods for training neural OIE models in this paper.First, we propose a novel method for computing syntactically rich text embeddings using the structure of dependency trees.Second, we propose a new discriminative training approach to OIE in which tokens in the generated fact are classified as "real" or "fake", i.e., those tokens that are in both the generated and gold tuples, and those that are only in the generated tuple but not in the gold tuple.We also address the issue of repetitive tokens in generated facts and improve the models' ability to generate implicit facts.Our approach reduces repetitive tokens by a factor of 23%.Finally, we present paraphrased versions of the CaRB, OIE2016, and LSOIE datasets, and show that the models' performance substantially improves when trained on datasets augmented by such data.Our best model beats the SOTA of IMoJIE on the recent CaRB dataset, with an improvement of 39.63% in F 1 score. Frank Mtumbuka, Thomas Lukasiewicz |
EMNLP | 2 |
| 2022 | (Non-)Convergence Results for Predictive Coding NetworksabstractPredictive coding networks (PCNs) are (un)supervised learning models, coming from neuroscience, that approximate how the brain works. One major open problem around PCNs is their convergence behavior. In this paper, we use dynamical systems theory to formally investigate the convergence of PCNs as they are used in machine learning. Doing so, we put their theory on a firm, rigorous basis, by developing a precise mathematical framework for PCN and show that for sufficiently small weights and initializations, PCNs converge for any input. Thereby, we provide the theoretical assurance that previous implementations, whose convergence was assessed solely by numerical experiments, can indeed capture the correct behavior of PCNs. Outside of the identified regime of small weights and small initializations, we show via a counterexample that PCNs can diverge, countering common beliefs held in the community. This is achieved by identifying a Neimark-Sacker bifurcation in a PCN of small size, which gives rise to an unstable fixed point and an invariant curve around it. Simon Frieder, Thomas Lukasiewicz |
ICML | 2 |
| 2022 | Knowledge-Grounded Self-Rationalization via Extractive and Natural Language ExplanationsabstractModels that generate extractive rationales (i.e., subsets of features) or natural language explanations (NLEs) for their predictions are important for explainable AI. While an extractive rationale provides a quick view of the features most responsible for a prediction, an NLE allows for a comprehensive description of the decision-making process behind a prediction. However, current models that generate the best extractive rationales or NLEs often fall behind the state-of-the-art (SOTA) in terms of task performance. In this work, we bridge this gap by introducing RExC, a self-rationalizing framework that grounds its predictions and two complementary types of explanations (NLEs and extractive rationales) in background knowledge. Our framework improves over previous methods by: (i) reaching SOTA task performance while also providing explanations, (ii) providing two types of explanations, while existing models usually provide only one type, and (iii) beating by a large margin the previous SOTA in terms of quality of both types of explanations. Furthermore, a perturbation analysis in RExC shows a high degree of association between explanations and predictions, a necessary property of faithful explanations. Bodhisattwa Prasad Majumder, Oana-Maria Camburu, Thomas Lukasiewicz, Julian J. McAuley |
ICML | 3 |
| 2022 | Universal Hopfield Networks: A General Framework for Single-Shot Associative Memory ModelsabstractA large number of neural network models of associative memory have been proposed in the literature. These include the classical Hopfield networks (HNs), sparse distributed memories (SDMs), and more recently the modern continuous Hopfield networks (MCHNs), which possess close links with self-attention in machine learning. In this paper, we propose a general framework for understanding the operation of such memory networks as a sequence of three operations: similarity, separation, and projection. We derive all these memory models as instances of our general framework with differing similarity and separation functions. We extend the mathematical framework of Krotov et al (2020) to express general associative memory models using neural network dynamics with local computation, and derive a general energy function that is a Lyapunov function of the dynamics. Finally, using our framework, we empirically investigate the capacity of using different similarity functions for these associative memory models, beyond the dot product similarity measure, and demonstrate empirically that Euclidean or Manhattan distance similarity metrics perform substantially better in practice on many tasks, enabling a more robust retrieval and higher memory capacity than existing models. Beren Millidge, Tommaso Salvatori, Yuhang Song 0001, Thomas Lukasiewicz, Rafal Bogacz |
ICML | 4 |
| 2022 | NP-Match: When Neural Processes meet Semi-Supervised LearningabstractSemi-supervised learning (SSL) has been widely explored in recent years, and it is an effective way of leveraging unlabeled data to reduce the reliance on labeled data. In this work, we adjust neural processes (NPs) to the semi-supervised image classification task, resulting in a new method named NP-Match. NP-Match is suited to this task for two reasons. Firstly, NP-Match implicitly compares data points when making predictions, and as a result, the prediction of each unlabeled data point is affected by the labeled data points that are similar to it, which improves the quality of pseudolabels. Secondly, NP-Match is able to estimate uncertainty that can be used as a tool for selecting unlabeled samples with reliable pseudo-labels. Compared with uncertainty-based SSL methods implemented with Monte Carlo (MC) dropout, NP-Match estimates uncertainty with much less computational overhead, which can save time at both the training and the testing phases. We conducted extensive experiments on four public datasets, and NP-Match outperforms state-of-theart (SOTA) results or achieves competitive results on them, which shows the effectiveness of NPMatch and its potential for SSL. Thomas Lukasiewicz, Daniela Massiceti, Xiaolin Hu 0001, Vladimir Pavlovic 0001, Alexandros Neophytou |
ICML | 2 |
| 2022 | Deep Learning with Logical ConstraintsabstractIn recent years, there has been an increasing interest in exploiting logically specified background knowledge in order to obtain neural models (i) with a better performance, (ii) able to learn from less data, and/or (iii) guaranteed to be compliant with the background knowledge itself, e.g., for safety-critical applications. In this survey, we retrace such works and categorize them based on (i) the logical language that they use to express the background knowledge and (ii) the goals that they achieve. Eleonora Giunchiglia, Mihaela Catalina Stoian, Thomas Lukasiewicz |
IJCAI | 3 |
| 2022 | Explanations for Negative Query Answers under Inconsistency-Tolerant SemanticsabstractInconsistency-tolerant semantics have been proposed to provide meaningful query answers even in the presence of inconsistent knowledge. Recently, explainability has also become a prominent problem in different areas of AI. While the complexity of inconsistency-tolerant semantics is rather well-understood, not much attention has been paid yet to the problem of explaining query answers when inconsistencies may exist. Recent work on existential rules in the inconsistent setting has focused only on understanding why a query is entailed. In this paper, we address another important problem, which is explaining why a query is not entailed under an inconsistency-tolerant semantics. In particular, we consider three popular semantics, namely, the ABox repair, the intersection of repairs, and the intersection of closed repairs. We provide a thorough complexity analysis for a wide range of existential rule languages and for several complexity measures. Thomas Lukasiewicz, Enrico Malizia, Cristian Molinaro |
IJCAI | 1 |
| 2022 | Predictive Coding: Towards a Future of Deep Learning beyond Backpropagation?abstractThe backpropagation of error algorithm (BP) used to train deep neural networks has been fundamental to the successes of deep learning. However, it requires sequential backwards updates and non-local computations which make it challenging to parallelize at scale and is unlike how learning works in the brain. Neuroscience-inspired learning algorithms, however, such as \emph{predictive coding} which utilize local learning have the potential to overcome these limitations and advance beyond deep learning technologies in the future. While predictive coding originated in theoretical neuroscience as a model of information processing in the cortex, recent work has developed the idea into a general-purpose algorithm able to train neural networks using only local computations. In this survey, we review works that have contributed to this perspective and demonstrate the close connection between predictive coding and backpropagation in terms of generalization quality, as well as works that highlight the multiple advantages of using predictive coding models over backprop-trained neural networks. Specifically, we show the substantially greater flexibility of predictive coding networks against equivalent deep neural networks, which can function as classifiers, generators, and associative memories simultaneously, and can be defined on arbitrary graph topologies. Finally, we review direct benchmarks of predictive coding networks on machine learning classification tasks, as well as its close connections to control theory and applications in robotics. Beren Millidge, Tommaso Salvatori, Yuhang Song 0001, Rafal Bogacz, Thomas Lukasiewicz |
IJCAI | 5 |
| 2022 | Explaining Chest X-Ray Pathologies in Natural Language
Maxime Kayser, Cornelius Emde, Oana-Maria Camburu, Guy Parsons, Bartlomiej Wladyslaw Papiez, Thomas Lukasiewicz |
MICCAI (5) | 6 |
| 2022 | Clustering Generative Adversarial Networks for Story VisualizationabstractStory visualization aims to generate a series of images, semantically matching a given sequence of sentences, one for each, and different output images within a story should be consistent with each other. Current methods generate story images by using a heavy architecture with two generative adversarial networks (GANs), one for image quality, and one for story consistency, and also rely on additional segmentation masks or auxiliary captioning networks. In this paper, we aim to build a concise and single-GAN-based network, neither depending on additional semantic information nor captioning networks. To achieve this, we propose a contrastive-learning- and clustering-learning-based approach for story visualization. Our network utilizes contrastive losses between language and visual information to maximize the mutual information between them, and further extends it with clustering learning in the training process to capture semantic similarity across modalities. So, the discriminator in our approach provides comprehensive feedback to the generator, regarding both image quality and story consistency at the same time, allowing to have a single-GAN-based network to produce high-quality synthetic results. Extensive experiments on two datasets demonstrate that our single-GAN-based network has a smaller number of total parameters in the network, but achieves a major step up from previous methods, which improves FID from 78.64 to 39.17, and FSD from 94.53 to 41.18 on Pororo-SV, and establishes a strong benchmark FID of 76.51 and FSD of 19.74 on Abstract Scenes. Bowen Li 0001, Philip Torr 0001, Thomas Lukasiewicz |
ACM Multimedia | 3 |
| 2022 | Predictive Coding beyond Gaussian DistributionsabstractA large amount of recent research has the far-reaching goal of finding training methods for deep neural networks that can serve as alternatives to backpropagation~(BP). A prominent example is predictive coding (PC), which is a neuroscience-inspired method that performs inference on hierarchical Gaussian generative models. These methods, however, fail to keep up with modern neural networks, as they are unable to replicate the dynamics of complex layers and activation functions. In this work, we solve this problem by generalizing PC to arbitrary probability distributions, enabling the training of architectures, such as transformers, that are hard to approximate with only Gaussian assumptions. We perform three experimental analyses. First, we study the gap between our method and the standard formulation of PC on multiple toy examples. Second, we test the reconstruction quality on variational autoencoders, where our method reaches the same reconstruction quality as BP. Third, we show that our method allows us to train transformer networks and achieve performance comparable with BP on conditional language models. More broadly, this method allows neuroscience-inspired learning to be applied to multiple domains, since the internal distributions can be flexibly adapted to the data, tasks, and architectures used. Luca Pinchetti, Tommaso Salvatori, Yordan Yordanov, Beren Millidge, Yuhang Song 0001, Thomas Lukasiewicz |
NeurIPS | 6 |
| 2022 | Learning on Arbitrary Graph Topologies via Predictive CodingabstractTraining with backpropagation (BP) in standard deep learning consists of two main steps: a forward pass that maps a data point to its prediction, and a backward pass that propagates the error of this prediction back through the network. This process is highly effective when the goal is to minimize a specific objective function. However, it does not allow training on networks with cyclic or backward connections. This is an obstacle to reaching brain-like capabilities, as the highly complex heterarchical structure of the neural connections in the neocortex are potentially fundamental for its effectiveness. In this paper, we show how predictive coding (PC), a theory of information processing in the cortex, can be used to perform inference and learning on arbitrary graph topologies. We experimentally show how this formulation, called PC graphs, can be used to flexibly perform different tasks with the same network by simply stimulating specific neurons. This enables the model to be queried on stimuli with different structures, such as partial images, images with labels, or images without labels. We conclude by investigating how the topology of the graph influences the final performance, and comparing against simple baselines trained with BP. Tommaso Salvatori, Luca Pinchetti, Beren Millidge, Yuhang Song 0001, Tianyi Bao, Rafal Bogacz, Thomas Lukasiewicz |
NeurIPS | 7 |
| 2022 | Complexity results for preference aggregation over (m)CP-nets: Max and rank voting
Thomas Lukasiewicz, Enrico Malizia |
Artif. Intell. | 1 |
| 2022 | Inconsistency-tolerant query answering for existential rules
Thomas Lukasiewicz, Enrico Malizia, Maria Vanina Martinez, Cristian Molinaro, Andreas Pieris, Gerardo I. Simari |
Artif. Intell. | 1 |
| 2022 | ω-net: Dual supervised medical image segmentation with multi-dimensional self-attention and diversely-connected multi-scale convolution
Zhenghua Xu 0001, Junyang Chen 0001, Thomas Lukasiewicz, Zhigang Fu |
Neurocomputing | 6 |
| 2022 | NoiER: An Approach for Training More Reliable Fine-Tuned Downstream Task ModelsabstractThe recent development in pretrained language models that are trained in a self-supervised fashion, such as BERT, is driving rapid progress in natural language processing. However, their brilliant performance is based on leveraging syntactic artefacts of the training data rather than fully understanding the intrinsic meaning of language. The excessive exploitation of spurious artefacts is a problematic issue: the distribution collapse problem, which is the phenomenon that the model fine-tuned on downstream tasks is unable to distinguish out-of-distribution sentences while producing a high-confidence score. In this paper, we argue that the distribution collapse is a prevalent issue in pretrained language models and proposenoise entropy regularisation (NoiER)as an efficient learning paradigm that solves the problem without auxiliary models and additional data. The proposed approach improved traditional out-of-distribution detection evaluation metrics by 55% on average compared to the original fine-tuned models. Myeongjun Jang 0001, Thomas Lukasiewicz |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | An Explainable Transformer-Based Deep Learning Model for the Prediction of Incident Heart FailureabstractPredicting the incidence of complex chronic conditions such as heart failure is challenging. Deep learning models applied to rich electronic health records may improve prediction but remain unexplainable hampering their wider use in medical practice. We aimed to develop a deep-learning framework for accurate and yet explainable prediction of 6-month incident heart failure (HF). Using 100,071 patients from longitudinal linked electronic health records across the U.K., we applied a novel Transformer-based risk model using all community and hospital diagnoses and medications contextualized within the age and calendar year for each patient's clinical encounter. Feature importance was investigated with an ablation analysis to compare model performance when alternatively removing features and by comparing the variability of temporal representations. A post-hoc perturbation technique was conducted to propagate the changes in the input to the outcome for feature contribution analyses. Our model achieved 0.93 area under the receiver operator curve and 0.69 area under the precision-recall curve on internal 5-fold cross validation and outperformed existing deep learning models. Ablation analysis indicated medication is important for predicting HF risk, calendar year is more important than chronological age, which was further reinforced by temporal variability analysis. Contribution analyses identified risk factors that are closely related to HF. Many of them were consistent with existing knowledge from clinical and epidemiological research but several new associations were revealed which had not been considered in expert-driven risk prediction models. In conclusion, the results highlight that our deep learning model, in addition high predictive performance, can inform data-driven risk factor identification. Shishir Rao, Yikuan Li, Rema Ramakrishnan, Abdelaali Hassaïne, Dexter Canoy, John G. F. Cleland, Thomas Lukasiewicz, Kazem Rahimi |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | Preferred Explanations for Ontology-Mediated Queries under Existential RulesabstractRecently, explanations for query answers under existential rules have been investigated, where an explanation is an inclusion-minimal subset of a given database that, together with the ontology, entails the query. In this paper, we take a step further and study explanations under different minimality criteria. In particular, we first study cardinality-minimal explanations and hence focus on deriving explanations of minimum size. We then study a more general preference order induced by a weight distribution. We assume that every database fact is annotated with a (penalization) weight, and we are interested in explanations with minimum overall weight. For both preference orders, we study a variety of explanation problems, such as recognizing a preferred explanation, all preferred explanations, a relevant or necessary fact, and the existence of a preferred explanation not containing forbidden sets of facts. We provide a detailed complexity analysis for all the aforementioned problems, thereby providing a more complete picture for explaining query answers under existential rules. Ismail Ilkan Ceylan, Thomas Lukasiewicz, Enrico Malizia, Cristian Molinaro, Andrius Vaicenavicius |
AAAI | 2 |
| 2021 | The Gap on Gap: Tackling the Problem of Differing Data Distributions in Bias-Measuring DatasetsabstractDiagnostic datasets that can detect biased models are an important prerequisite for bias reduction within natural language processing. However, undesired patterns in the collected data can make such tests incorrect. For example, if the feminine subset of a gender-bias-measuring coreference resolution dataset contains sentences with a longer average distance between the pronoun and the correct candidate, an RNN-based model may perform worse on this subset due to long-term dependencies. In this work, we introduce a theoretically grounded method for weighting test samples to cope with such patterns in the test data. We demonstrate the method on the GAP dataset for coreference resolution. We annotate GAP with spans of all personal names and show that examples in the female subset contain more personal names and a longer distance between pronouns and their referents, potentially affecting the bias score in an undesired way. Using our weighting method, we find the set of weights on the test instances that should be used for coping with these correlations, and we re-evaluate 16 recently released coreference models. Vid Kocijan, Oana-Maria Camburu, Thomas Lukasiewicz |
AAAI | 3 |
| 2021 | Learning from the Best: Rationalizing Predictions by Adversarial Information CalibrationabstractExplaining the predictions of AI models is paramount in safety-critical applications, such as in legal or medical domains. One form of explanation for a prediction is an extractive rationale, i.e., a subset of features of an instance that lead the model to give its prediction on the instance. Previous works on generating extractive rationales usually employ a two-phase model: a selector that selects the most important features (i.e., the rationale) followed by a predictor that makes the prediction based exclusively on the selected features. One disadvantage of these works is that the main signal for learning to select features comes from the comparison of the final answers given by the predictor and the ground-truth answers. In this work, we propose to squeeze more information from the predictor via an information calibration method. More precisely, we train two models jointly: one is a typical neural model that solves the task at hand in an accurate but black-box manner, and the other is a selector-predictor model that additionally produces a rationale for its prediction. The first model is used as a guide to the second model. We use an adversarial-based technique to calibrate the information extracted by the two models such that the difference between them is an indicator of the missed or over-selected features. In addition, for natural language tasks, we propose to use a language-model-based regularizer to encourage the extraction of fluent rationales. Experimental results on a sentiment analysis task as well as on three tasks from the legal domain show the effectiveness of our approach to rationale extraction. Lei Sha, Oana-Maria Camburu, Thomas Lukasiewicz |
AAAI | 3 |
| 2021 | Multi-type Disentanglement without Adversarial TrainingabstractControlling the style of natural language by disentangling the latent space is an important step towards interpretable machine learning. After the latent space is disentangled, the style of a sentence can be transformed by tuning the style representation without affecting other features of the sentence. Previous works usually use adversarial training to guarantee that disentangled vectors do not affect each other. However, adversarial methods are difficult to train. Especially when there are multiple features (e.g., sentiment, or tense, which we call style types in this paper), each feature requires a separate discriminator for extracting a disentangled style vector corresponding to that feature. In this paper, we propose a unified distribution-controlling method, which provides each specific style value (the value of style types, e.g., positive sentiment, or past tense) with a unique representation. This method contributes a solid theoretical basis to avoid adversarial training in multi-type disentanglement. We also propose multiple loss functions to achieve a style-content disentanglement as well as a disentanglement among multiple style types. In addition, we observe that if two different style types always have some specific style values that occur together in the dataset, they will affect each other when transferring the style values. We call this phenomenon training bias , and we propose a loss function to alleviate such training bias while disentangling multiple types. We conduct experiments on two datasets (Yelp service reviews and Amazon product reviews) to evaluate the style-disentangling effect and the unsupervised style-transfer performance on two style types: sentiment and tense. The experimental results show the effectiveness of our model. Lei Sha, Thomas Lukasiewicz |
AAAI | 2 |
| 2021 | Lightweight Visual Question Answering using Scene GraphsabstractVisual question answering (VQA) is a challenging problem in machine perception, which requires a deep joint understanding of both visual and textual data. Recent research has advanced the automatic generation of high-quality scene graphs from images, while powerful yet elegant models like graph neural networks (GNNs) have shown great power in reasoning over graph-structured data. In this work, we propose to bridge the gap between scene graph generation and VQA by leveraging GNNs. In particular, we design a new model called Conditional Enhanced Graph ATtention network (CE-GAT) to encode pairs of visual and semantic scene graphs with both node and edge features, which is seamlessly integrated with a textual question encoder to generate answers through question-graph conditioning. Moreover, to alleviate the training difficulties of CE-GAT towards VQA, we enforce more useful inductive biases in the scene graphs through novel question-guided graph enriching and pruning. Finally, we evaluate the framework on one of the largest available VQA datasets (namely, GQA) with ground-truth scene graphs, achieving the accuracy of 77.87%, compared with the state of the art (namely, the neural state machine (NSM)), which gives 63.17%. Notably, by leveraging existing scene graphs, our framework is much lighter compared with end-to-end VQA methods (e.g., about 95.3% less parameters than a typical NSM). Sai Vidyaranya Nuthalapati, Ramraj Chandradevan, Eleonora Giunchiglia, Bowen Li 0001, Maxime Kayser, Thomas Lukasiewicz, Carl Yang 0001 |
CIKM | 6 |
| 2021 | RSG: A Simple but Effective Module for Learning Imbalanced DatasetsabstractImbalanced datasets widely exist in practice and are a great challenge for training deep neural models with a good generalization on infrequent classes. In this work, we propose a new rare-class sample generator (RSG) to solve this problem. RSG aims to generate some new samples for rare classes during training, and it has in particular the following advantages: (1) it is convenient to use and highly versatile, because it can be easily integrated into any kind of convolutional neural network, and it works well when combined with different loss functions, and (2) it is only used during the training phase, and therefore, no additional burden is imposed on deep neural networks during the testing phase. In extensive experimental evaluations, we verify the effectiveness of RSG. Furthermore, by leveraging RSG, we obtain competitive results on Imbalanced CIFAR and new state-of-the-art results on Places-LT, ImageNet-LT, and iNaturalist 2018. The source code is available at https://github.com/Jianf-Wang/RSG. Thomas Lukasiewicz, Xiaolin Hu 0001, Jianfei Cai 0001, Zhenghua Xu 0001 |
CVPR | 2 |
| 2021 | Knowledge Base Completion Meets Transfer LearningabstractThe aim of knowledge base completion is to predict unseen facts from existing facts in knowledge bases.In this work, we introduce the first approach for transfer of knowledge from one collection of facts to another without the need for entity or relation matching.The method works for both canonicalized knowledge bases and uncanonicalized or open knowledge bases, i.e., knowledge bases where more than one copy of a real-world entity or relation may exist.Such knowledge bases are a natural output of automated information extraction tools that extract structured data from unstructured text.Our main contribution is a method that can make use of a large-scale pre-training on facts, collected from unstructured text, to improve predictions on structured data from a specific domain.The introduced method is the most impactful on small datasets such as ReVerb20K, where we obtained 6% absolute increase of mean reciprocal rank and 65% relative decrease of mean rank over the previously best method, despite not relying on large pre-trained models like BERT. Vid Kocijan, Thomas Lukasiewicz |
EMNLP (1) | 2 |
| 2021 | e-ViL: A Dataset and Benchmark for Natural Language Explanations in Vision-Language TasksabstractRecently, there has been an increasing number of efforts to introduce models capable of generating natural language explanations (NLEs) for their predictions on vision-language (VL) tasks. Such models are appealing, because they can provide human-friendly and comprehensive explanations. However, there is a lack of comparison between existing methods, which is due to a lack of re-usable evaluation frameworks and a scarcity of datasets. In this work, we introduce e-ViL and e-SNLI-VE. e-ViL is a benchmark for explainable vision-language tasks that establishes a unified evaluation framework and provides the first comprehensive comparison of existing approaches that generate NLEs for VL tasks. It spans four models and three datasets and both automatic metrics and human evaluation are used to assess model-generated explanations. e-SNLI-VE is currently the largest existing VL dataset with NLEs (over 430k instances). We also propose a new model that combines UNITER [15], which learns joint embeddings of images and text, and GPT-2 [38], a pre-trained language model that is well-suited for text generation. It surpasses the previous state of the art by a large margin across all datasets. Code and data are available here: https://github.com/maximek3/e-ViL. Maxime Kayser, Oana-Maria Camburu, Leonard Salewski, Cornelius Emde, Virginie Do, Zeynep Akata, Thomas Lukasiewicz |
ICCV | 7 |
| 2021 | The Surprising Power of Graph Neural Networks with Random Node InitializationabstractGraph neural networks (GNNs) are effective models for representation learning on relational data. However, standard GNNs are limited in their expressive power, as they cannot distinguish graphs beyond the capability of the Weisfeiler-Leman graph isomorphism heuristic. In order to break this expressiveness barrier, GNNs have been enhanced with random node initialization (RNI), where the idea is to train and run the models with randomized initial node features. In this work, we analyze the expressive power of GNNs with RNI, and prove that these models are universal, a first such result for GNNs not relying on computationally demanding higher-order properties. This universality result holds even with partially randomized initial node features, and preserves the invariance properties of GNNs in expectation. We then empirically analyze the effect of RNI on GNNs, based on carefully constructed datasets. Our empirical findings support the superior performance of GNNs with RNI over standard GNNs. Ralph Abboud, Ismail Ilkan Ceylan, Martin Grohe, Thomas Lukasiewicz |
IJCAI | 4 |
| 2021 | Associative Memories via Predictive CodingabstractAssociative memories in the brain receive and store patterns of activity registered by the sensory neurons, and are able to retrieve them when necessary. Due to their importance in human intelligence, computational models of associative memories have been developed for several decades now. In this paper, we present a novel neural model for realizing associative memories, which is based on a hierarchical generative network that receives external stimuli via sensory neurons. It is trained using predictive coding, an error-based learning algorithm inspired by information processing in the cortex. To test the model's capabilities, we perform multiple retrieval experiments from both corrupted and incomplete data points. In an extensive comparison, we show that this new model outperforms in retrieval accuracy and robustness popular associative memory models, such as autoencoders trained via backpropagation, and modern Hopfield networks. In particular, in completing partial data points, our model achieves remarkable results on natural image datasets, such as ImageNet, with a surprisingly high accuracy, even when only a tiny fraction of pixels of the original images is presented. Our model provides a plausible framework to study learning and retrieval of memories in the brain, as it closely mimics the behavior of the hippocampus as a memory index and generative model. Tommaso Salvatori, Yuhang Song 0001, Yujian Hong, Lei Sha, Simon Frieder, Zhenghua Xu 0001, Rafal Bogacz, Thomas Lukasiewicz |
NeurIPS | 8 |
| 2021 | An ontology-based deep learning approach for triple classification with out-of-knowledge-base entities
Elvira Amador-Domínguez, Emilio Serrano, Daniel Manrique, Patrick Hohenecker, Thomas Lukasiewicz |
Inf. Sci. | 5 |
| 2021 | Stable Model Semantics for Guarded Existential Rules and Description Logics: Decidability and ComplexityabstractThis work investigates the decidability and complexity of database query answering under guarded existential rules with nonmonotonic negation according to the classical stable model semantics. In this setting, existential quantification is interpreted via Skolem functions, and the unique name assumption is adopted. As a first result, we show the decidability of answering first-order queries based on such rules by a translation into the satisfiability problem for guarded second-order formulas having the tree-model property. To obtain precise complexity results for unions of conjunctive queries, we transform the original problem in polynomial time into an intermediate problem that is easier to analyze: query answering for guarded disjunctive existential rules with stratified negation. We obtain precise bounds for the general setting and for various restricted settings. We also consider extensions of the original formalism with negative constraints, keys, and the possibility of negated atoms in queries. Finally, we show how the above results can be used to provide decidability and complexity results for a natural adaptation of the stable model semantics to description logics such as ELHI and the DL-Lite family. Georg Gottlob, André Hernich, Clemens Kupke, Thomas Lukasiewicz |
J. ACM | 4 |
| 2021 | Multi-Label Classification Neural Networks with Hard Logical ConstraintsabstractMulti-label classification (MC) is a standard machine learning problem in which a data point can be associated with a set of classes. A more challenging scenario is given by hierarchical multi-label classification (HMC) problems, in which every prediction must satisfy a given set of hard constraints expressing subclass relationships between classes. In this article, we propose C-HMCNN(h), a novel approach for solving HMC problems, which, given a network h for the underlying MC problem, exploits the hierarchy information in order to produce predictions coherent with the constraints and to improve performance. Furthermore, we extend the logic used to express HMC constraints in order to be able to specify more complex relations among the classes and propose a new model CCN(h), which extends C-HMCNN(h) and is again able to satisfy and exploit the constraints to improve performance. We conduct an extensive experimental analysis showing the superior performance of both C-HMCNN(h) and CCN(h) when compared to state-of-the-art models in both the HMC and the general MC setting with hard logical constraints. Eleonora Giunchiglia, Thomas Lukasiewicz |
J. Artif. Intell. Res. | 2 |
| 2020 | Arena: A General Evaluation Platform and Building Toolkit for Multi-Agent IntelligenceabstractLearning agents that are not only capable of taking tests, but also innovating is becoming a hot topic in AI. One of the most promising paths towards this vision is multi-agent learning, where agents act as the environment for each other, and improving each agent means proposing new problems for others. However, existing evaluation platforms are either not compatible with multi-agent settings, or limited to a specific game. That is, there is not yet a general evaluation platform for research on multi-agent intelligence. To this end, we introduce Arena, a general evaluation platform for multi-agent intelligence with 35 games of diverse logics and representations. Furthermore, multi-agent intelligence is still at the stage where many problems remain unexplored. Therefore, we provide a building toolkit for researchers to easily invent and build novel multi-agent problems from the provided game set based on a GUI-configurable social tree and five basic multi-agent reward schemes. Finally, we provide Python implementations of five state-of-the-art deep multi-agent reinforcement learning baselines. Along with the baseline implementations, we release a set of 100 best agents/teams that we can train with different training schemes for each game, as the base for evaluating agents with population performance. As such, the research community can perform comparisons under a stable and uniform standard. All the implementations and accompanied tutorials have been open-sourced for the community at https://sites.google.com/view/arena-unity/. Yuhang Song 0001, Andrzej Wojcicki, Thomas Lukasiewicz, Jianyi Wang, Abi Aryan, Zhenghua Xu 0001, Mai Xu, Lianlong Wu |
AAAI | 3 |
| 2020 | Mega-Reward: Achieving Human-Level Play without Extrinsic RewardsabstractIntrinsic rewards were introduced to simulate how human intelligence works; they are usually evaluated by intrinsically-motivated play, i.e., playing games without extrinsic rewards but evaluated with extrinsic rewards. However, none of the existing intrinsic reward approaches can achieve human-level performance under this very challenging setting of intrinsically-motivated play. In this work, we propose a novel megalomania-driven intrinsic reward (called mega-reward), which, to our knowledge, is the first approach that achieves human-level performance in intrinsically-motivated play. Intuitively, mega-reward comes from the observation that infants' intelligence develops when they try to gain more control on entities in an environment; therefore, mega-reward aims to maximize the control capabilities of agents on given entities in a given environment. To formalize mega-reward, a relational transition model is proposed to bridge the gaps between direct and latent control. Experimental studies show that mega-reward (i) can greatly outperform all state-of-the-art intrinsic reward approaches, (ii) generally achieves the same level of performance as Ex-PPO and professional human-level scores, and (iii) has also a superior performance when it is incorporated with extrinsic rewards. Yuhang Song 0001, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu 0001, Shangtong Zhang, Andrzej Wojcicki, Mai Xu |
AAAI | 3 |
| 2020 | Learning to Reason: Leveraging Neural Networks for Approximate DNF CountingabstractWeighted model counting (WMC) has emerged as a prevalent approach for probabilistic inference. In its most general form, WMC is #P-hard. Weighted DNF counting (weighted #DNF) is a special case, where approximations with probabilistic guarantees are obtained in O(nm), where n denotes the number of variables, and m the number of clauses of the input DNF, but this is not scalable in practice. In this paper, we propose a neural model counting approach for weighted #DNF that combines approximate model counting with deep learning, and accurately approximates model counts in linear time when width is bounded. We conduct experiments to validate our method, and show that our model learns and generalizes very well to large-scale #DNF instances. Ralph Abboud, Ismail Ilkan Ceylan, Thomas Lukasiewicz |
AAAI | 3 |
| 2020 | Explanations for Inconsistency-Tolerant Query Answering under Existential RulesabstractQuerying inconsistent knowledge bases is a problem that has attracted a great deal of interest over the last decades. While several semantics of query answering have been proposed, and their complexity is rather well-understood, little attention has been paid to the problem of explaining query answers. Explainability has recently become a prominent problem in different areas of AI. In particular, explaining query answers allows users to understand not only what is entailed by an inconsistent knowledge base, but also why. In this paper, we address the problem of explaining query answers for existential rules under three popular inconsistency-tolerant semantics, namely, the ABox repair, the intersection of repairs, and the intersection of closed repairs semantics. We provide a thorough complexity analysis for a wide range of existential rule languages and for different complexity measures. Thomas Lukasiewicz, Enrico Malizia, Cristian Molinaro |
AAAI | 1 |
| 2020 | Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language ExplanationsabstractTo increase trust in artificial intelligence systems, a promising research direction consists of designing neural models capable of generating natural language explanations for their predictions.In this work, we show that such models are nonetheless prone to generating mutually inconsistent explanations, such as "Because there is a dog in the image."and "Because there is no dog in the [same] image.",exposing flaws in either the decision-making process of the model or in the generation of the explanations.We introduce a simple yet effective adversarial framework for sanity checking models against the generation of inconsistent natural language explanations.Moreover, as part of the framework, we address the problem of adversarial attacks with full target sequences, a scenario that was not previously addressed in sequence-to-sequence attacks.Finally, we apply our framework on a state-of-the-art neural natural language inference model that provides natural language explanations for its predictions.Our framework shows that this model is capable of generating a significant number of inconsistent explanations.PREMISE: A guy in a red jacket is snowboarding in midair. Oana-Maria Camburu, Brendan Shillingford, Pasquale Minervini, Thomas Lukasiewicz, Phil Blunsom |
ACL | 4 |
| 2020 | ManiGAN: Text-Guided Image ManipulationabstractThe goal of our paper is to semantically edit parts of an image matching a given text that describes desired attributes (e.g., texture, colour, and background), while preserving other contents that are irrelevant to the text. To achieve this, we propose a novel generative adversarial network (ManiGAN), which contains two key components: text-image affine combination module (ACM) and detail correction module (DCM). The ACM selects image regions relevant to the given text and then correlates the regions with corresponding semantic words for effective manipulation. Meanwhile, it encodes original image features to help reconstruct text-irrelevant contents. The DCM rectifies mismatched attributes and completes missing contents of the synthetic image. Finally, we suggest a new metric for evaluating image manipulation results, in terms of both the generation of new attributes and the reconstruction of text-irrelevant contents. Extensive experiments on the CUB and COCO datasets demonstrate the superior performance of the proposed method. Bowen Li 0001, Xiaojuan Qi 0001, Thomas Lukasiewicz, Philip Torr 0001 |
CVPR | 3 |
| 2020 | Explanations for Ontology-Mediated Query Answering in Description LogicsabstractOntology-mediated query answering is a paradigm that seeks to exploit the semantic knowledge expressed in terms of ontologies to improve query answers over incomplete data sources. In this paper, we focus on description logic ontologies, and study the problem of explaining why an ontology-mediated query is entailed from a given data source. Specifically, we view explanations as minimal sets of assertions from an ABox, which satisfy the ontologymediated query. Based on such explanations, we study a variety of problems taken from the recent literature on explanations (studied for existential rules), such as recognizing all minimal explanations. Our results establish tight connections between intractable explanation problems and variants of propositional satisfiability problems. We provide insights on the inherent computational difficulty of deriving explanations for ontology-mediated queries Ismail Ilkan Ceylan, Thomas Lukasiewicz, Enrico Malizia, Andrius Vaicenavicius |
ECAI | 2 |
| 2020 | Systematic Comparison of Neural Architectures and Training Approaches for Open Information ExtractionabstractThe goal of open information extraction (OIE) is to extract facts from natural language text, and to represent them as structured triples of the form subject, predicate, object .For example, given the sentence »Beethoven composed the Ode to Joy.«, we are expected to extract the triple Beethoven, composed, Ode to Joy .In this work, we systematically compare different neural network architectures and training approaches, and improve the performance of the currently best models on the OIE16 benchmark (Stanovsky and Dagan, 2016) by 0.421 F 1 score and 0.420 AUC-PR, respectively, in our experiments (i.e., by more than 200% in both cases).Furthermore, we show that appropriate problem and loss formulations often affect the performance more than the network architecture.Ludwig van Beethoven was a world -famous composer of classical music . Patrick Hohenecker, Frank Mtumbuka, Vid Kocijan, Thomas Lukasiewicz |
EMNLP (1) | 4 |
| 2020 | Does the Objective Matter? Comparing Training Objectives for Pronoun ResolutionabstractHard cases of pronoun resolution have been used as a long-standing benchmark for commonsense reasoning.In the recent literature, pre-trained language models have been used to obtain state-of-the-art results on pronoun resolution.Overall, four categories of training and evaluation objectives have been introduced.The variety of training datasets and pretrained language models used in these works makes it unclear whether the choice of training objective is critical.In this work, we make a fair comparison of the performance and seedwise stability of four models that represent the four categories of objectives.Our experiments show that the objective of sequence ranking performs the best in-domain, while the objective of semantic similarity between candidates and pronoun performs the best out-of-domain.We also observe a seed-wise instability of the model using sequence ranking, which is not the case when the other objectives are used. Yordan Yordanov, Oana-Maria Camburu, Vid Kocijan, Thomas Lukasiewicz |
EMNLP (1) | 4 |
| 2020 | Hybrid Deep-Semantic Matrix Factorization for Tag-Aware Personalized RecommendationabstractMatrix factorization has now become a dominant solution for personalized recommendation on the Social Web. To alleviate the cold start problem, previous approaches have incorporated various additional sources of information into traditional matrix factorization models. These upgraded models, however, achieve only "marginal" enhancements on the performance of personalized recommendation. Therefore, inspired by the recent development of deep-semantic modeling, we propose a hybrid deep-semantic matrix factorization (HDMF) model to further improve the performance of tag-aware personalized recommendation by integrating the techniques of deep-semantic modeling, hybrid learning, and matrix factorization. Experimental results show that HDMF significantly outperforms the state-of-the-art baselines in tag-aware personalized recommendation, in terms of all evaluation metrics. Zhenghua Xu 0001, Thomas Lukasiewicz, Cheng Chen 0002, Yishu Miao, Guizhi Xu |
ICASSP | 3 |
| 2020 | Knowledge Graph Extraction from VideosabstractNearly all existing techniques for automated video annotation (or captioning) describe videos using natural language sentences. However, this has several shortcomings: (i) it is very hard to then further use the generated natural language annotations in automated data processing, (ii) generating natural language annotations requires solving the hard subtask of generating semantically precise and syntactically correct natural language sentences, which is actually unrelated to the task of video annotation, (iii) it is difficult to quantitatively measure performance, as standard metrics (e.g., accuracy and F1-score) are inapplicable, and (iv) annotations are language-specific. In this paper, we propose the new task of knowledge graph extraction from videos, i.e., producing a description in the form of a knowledge graph of the contents of a given video. Since no datasets exist for this task, we also include a method to automatically generate them, starting from datasets where videos are annotated with natural language. We then describe an initial deep-learning model for knowledge graph extraction from videos, and report results on MSVD* and MSR-VTT*, two datasets obtained from MSVD and MSR-VTT using our method. Louis Mahon, Eleonora Giunchiglia, Bowen Li 0001, Thomas Lukasiewicz |
ICMLA | 4 |
| 2020 | Ontology Reasoning with Deep Neural Networks (Extended Abstract)abstractThe ability to conduct logical reasoning is a fundamental aspect of intelligent human behavior, and thus an important problem along the way to human-level artificial intelligence. Traditionally, logic-based symbolic methods from the field of knowledge representation and reasoning have been used to equip agents with capabilities that resemble human logical reasoning qualities. More recently, however, there has been an increasing interest in using machine learning rather than logic-based symbolic formalisms to tackle these tasks. In this paper, we employ state-of-the-art methods for training deep neural networks to devise a novel model that is able to learn how to effectively perform logical reasoning in the form of basic ontology reasoning. Patrick Hohenecker, Thomas Lukasiewicz |
IJCAI | 2 |
| 2020 | Extracting Outcomes from Appellate Decisions in US State CourtsabstractPredicting the outcome of a legal process has recently gained considerable research attention. Numerous attempts have been made to predict the exact outcome, judgment, charge, and fines of a case given the textual description of its facts and metadata. However, most of the effort has been focused on Chinese and European law, for which there exist annotated datasets. In this paper, we introduce CASELAW4 — a new dataset of 350k common law judicial decisions from the U.S. Caselaw Access Project, of which 250k have been automatically annotated with binary outcome labels of AFFIRM or REVERSE by our hybrid learning system. To our knowledge, it is the first attempt to perform outcome extraction (a) on such a large volume of English-language judicial opinions, (b) on the Caselaw Access Project data, and (c) on US State Courts of Appeal cases, and it paves the way to large-scale outcome prediction and advanced legal analytics using U.S. Case Law. We set up baseline results for the outcome extraction task on the new dataset, achieving an F-measure of 82.32%. Alina Petrova, John Armour, Thomas Lukasiewicz |
JURIX | 3 |
| 2020 | Explanations for Negative Query Answers under Existential RulesabstractOntology-mediated query answering is an extensively studied paradigm, where the conceptual knowledge provided by an ontology is leveraged towards more enhanced querying of data sources. A major advantage of ontological reasoning is its interpretability, which allows one to derive explanations for query answers. Indeed, explanations have a long history in knowledge representation, and have also been investigated for ontology languages based on description logics and existential rules. Existing works on existential rules, however, merely focus on understanding why a query is entailed, i.e., explaining positive query answers. In this paper, we continue this line of research and address another important problem, namely, explaining why a query is not entailed under existential rules, i.e., explaining negative query answers. We consider various problems related to explaining non-entailments from the abduction literature, and also introduce new problems. For all considered problems, we give a detailed complexity analysis for a wide range of existential rule languages and complexity measures. Ismail Ilkan Ceylan, Thomas Lukasiewicz, Enrico Malizia, Cristian Molinaro, Andrius Vaicenavicius |
KR | 2 |
| 2020 | Can the Brain Do Backpropagation? - Exact Implementation of Backpropagation in Predictive Coding NetworksabstractBackpropagation (BP) has been the most successful algorithm used to train artificial neural networks. However, there are several gaps between BP and learning in biologically plausible neuronal networks of the brain (learning in the brain, or simply BL, for short), in particular, (1) it has been unclear to date, if BP can be implemented exactly via BL, (2) there is a lack of local plasticity in BP, i.e., weight updates require information that is not locally available, while BL utilizes only locally available information, and (3)~there is a lack of autonomy in BP, i.e., some external control over the neural network is required (e.g., switching between prediction and learning stages requires changes to dynamics and synaptic plasticity rules), while BL works fully autonomously. Bridging such gaps, i.e., understanding how BP can be approximated by BL, has been of major interest in both neuroscience and machine learning. Despite tremendous efforts, however, no previous model has bridged the gaps at a degree of demonstrating an equivalence to BP, instead, only approximations to BP have been shown. Here, we present for the first time a framework within BL that bridges the above crucial gaps. We propose a BL model that (1) produces \emph{exactly the same} updates of the neural weights as~BP, while (2)~employing local plasticity, i.e., all neurons perform only local computations, done simultaneously. We then modify it to an alternative BL model that (3) also works fully autonomously. Overall, our work provides important evidence for the debate on the long-disputed question whether the brain can perform~BP. Yuhang Song 0001, Thomas Lukasiewicz, Zhenghua Xu 0001, Rafal Bogacz |
NeurIPS | 2 |
| 2020 | BoxE: A Box Embedding Model for Knowledge Base CompletionabstractKnowledge base completion (KBC) aims to automatically infer missing facts by exploiting information already present in a knowledge base (KB). A promising approach for KBC is to embed knowledge into latent spaces and make predictions from learned embeddings. However, existing embedding models are subject to at least one of the following limitations: (1) theoretical inexpressivity, (2) lack of support for prominent inference patterns (e.g., hierarchies), (3) lack of support for KBC over higher-arity relations, and (4) lack of support for incorporating logical rules. Here, we propose a spatio-translational embedding model, called BoxE, that simultaneously addresses all these limitations. BoxE embeds entities as points, and relations as a set of hyper-rectangles (or boxes), which spatially characterize basic logical properties. This seemingly simple abstraction yields a fully expressive model offering a natural encoding for many desired logical properties. BoxE can both capture and inject rules from rich classes of rule languages, going well beyond individual inference patterns. By design, BoxE naturally applies to higher-arity KBs. We conduct a detailed experimental analysis, and show that BoxE achieves state-of-the-art performance, both on benchmark knowledge graphs and on more general KBs, and we empirically show the power of integrating logical rules. Ralph Abboud, Ismail Ilkan Ceylan, Thomas Lukasiewicz, Tommaso Salvatori |
NeurIPS | 3 |
| 2020 | Coherent Hierarchical Multi-Label Classification NetworksabstractHierarchical multi-label classification (HMC) is a challenging classification task extending standard multi-label classification problems by imposing a hierarchy constraint on the classes. In this paper, we propose C-HMCNN(h), a novel approach for HMC problems, which, given a network h for the underlying multi-label classification problem, exploits the hierarchy information in order to produce predictions coherent with the constraint and improve performance. We conduct an extensive experimental analysis showing the superior performance of C-HMCNN(h) when compared to state-of-the-art models. Eleonora Giunchiglia, Thomas Lukasiewicz |
NeurIPS | 2 |
| 2020 | Lightweight Generative Adversarial Networks for Text-Guided Image ManipulationabstractWe propose a novel lightweight generative adversarial network for efficient image manipulation using natural language descriptions. To achieve this, a new word-level discriminator is proposed, which provides the generator with fine-grained training feedback at word-level, to facilitate training a lightweight generator that has a small number of parameters, but can still correctly focus on specific visual attributes of an image, and then edit them without affecting other contents that are not described in the text. Furthermore, thanks to the explicit training signal related to each word, the discriminator can also be simplified to have a lightweight structure. Compared with the state of the art, our method has a much smaller number of parameters, but still achieves a competitive manipulation performance. Extensive experimental results demonstrate that our method can better disentangle different visual attributes, then correctly map them to corresponding semantic words, and thus achieve a more accurate image modification using natural language descriptions. Bowen Li 0001, Xiaojuan Qi 0001, Philip Torr 0001, Thomas Lukasiewicz |
NeurIPS | 4 |
| 2020 | Partially observable game-theoretic agent programming in Golog
Alberto Finzi, Thomas Lukasiewicz |
Int. J. Approx. Reason. | 2 |
| 2020 | Ontology Reasoning with Deep Neural NetworksabstractThe ability to conduct logical reasoning is a fundamental aspect of intelligent human behavior, and thus an important problem along the way to human-level artificial intelligence. Traditionally, logic-based symbolic methods from the field of knowledge representation and reasoning have been used to equip agents with capabilities that resemble human logical reasoning qualities. More recently, however, there has been an increasing interest in using machine learning rather than logic-based symbolic formalisms to tackle these tasks. In this paper, we employ state-of-the-art methods for training deep neural networks to devise a novel model that is able to learn how to effectively perform logical reasoning in the form of basic ontology reasoning. This is an important and at the same time very natural logical reasoning task, which is why the presented approach is applicable to a plethora of important real-world problems. We present the outcomes of several experiments, which show that our model is able to learn to perform highly accurate ontology reasoning on very large, diverse, and challenging benchmarks. Furthermore, it turned out that the suggested approach suffers much less from different obstacles that prohibit logic-based symbolic reasoning, and, at the same time, is surprisingly plausible from a biological point of view. Patrick Hohenecker, Thomas Lukasiewicz |
J. Artif. Intell. Res. | 2 |
| 2019 | Ontology-Mediated Query Answering over Log-Linear Probabilistic DataabstractLarge-scale knowledge bases are at the heart of modern information systems. Their knowledge is inherently uncertain, and hence they are often materialized as probabilistic databases. However, probabilistic database management systems typically lack the capability to incorporate implicit background knowledge and, consequently, fail to capture some intuitive query answers. Ontology-mediated query answering is a popular paradigm for encoding commonsense knowledge, which can provide more complete answers to user queries. We propose a new data model that integrates the paradigm of ontology-mediated query answering with probabilistic databases, employing a log-linear probability model. We compare our approach to existing proposals, and provide supporting computational results. Stefan Borgwardt, Ismail Ilkan Ceylan, Thomas Lukasiewicz |
AAAI | 3 |
| 2019 | Complexity of Inconsistency-Tolerant Query Answering in Datalog+/- under Cardinality-Based RepairsabstractQuerying inconsistent ontological knowledge bases is an important problem in practice, for which several inconsistencytolerant query answering semantics have been proposed, including query answering relative to all repairs, relative to the intersection of repairs, and relative to the intersection of closed repairs. In these semantics, one assumes that the input database is erroneous, and the notion of repair describes a maximally consistent subset of the input database, where different notions of maximality (such as subset and cardinality maximality) are considered. In this paper, we give a precise picture of the computational complexity of inconsistencytolerant (Boolean conjunctive) query answering in a wide range of Datalog± languages under the cardinality-based versions of the above three repair semantics. Thomas Lukasiewicz, Enrico Malizia, Andrius Vaicenavicius |
AAAI | 1 |
| 2019 | Diversity-Driven Extensible Hierarchical Reinforcement LearningabstractHierarchical reinforcement learning (HRL) has recently shown promising advances on speeding up learning, improving the exploration, and discovering intertask transferable skills. Most recent works focus on HRL with two levels, i.e., a master policy manipulates subpolicies, which in turn manipulate primitive actions. However, HRL with multiple levels is usually needed in many real-world scenarios, whose ultimate goals are highly abstract, while their actions are very primitive. Therefore, in this paper, we propose a diversitydriven extensible HRL (DEHRL), where an extensible and scalable framework is built and learned levelwise to realize HRL with multiple levels. DEHRL follows a popular assumption: diverse subpolicies are useful, i.e., subpolicies are believed to be more useful if they are more diverse. However, existing implementations of this diversity assumption usually have their own drawbacks, which makes them inapplicable to HRL with multiple levels. Consequently, we further propose a novel diversity-driven solution to achieve this assumption in DEHRL. Experimental studies evaluate DEHRL with nine baselines from four perspectives in two domains; the results show that DEHRL outperforms the state-of-the-art baselines in all four aspects. Yuhang Song 0001, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu 0001, Mai Xu |
AAAI | 3 |
| 2019 | A Surprisingly Robust Trick for the Winograd Schema ChallengeabstractThe Winograd Schema Challenge (WSC) dataset WSC273 and its inference counterpart WNLI are popular benchmarks for natural language understanding and commonsense reasoning.In this paper, we show that the performance of three language models on WSC273 consistently and robustly improves when finetuned on a similar pronoun disambiguation problem dataset (denoted WSCR).We additionally generate a large unsupervised WSClike dataset.By fine-tuning the BERT language model both on the introduced and on the WSCR dataset, we achieve overall accuracies of 72.5% and 74.7% on WSC273 and WNLI, improving the previous state-of-theart solutions by 8.8% and 9.6%, respectively.Furthermore, our fine-tuned models are also consistently more accurate on the "complex" subsets of WSC273, introduced by Trichelair et al. (2018). Vid Kocijan, Ana-Maria Cretu 0002, Oana-Maria Camburu, Yordan Yordanov, Thomas Lukasiewicz |
ACL (1) | 5 |
| 2019 | WikiCREM: A Large Unsupervised Corpus for Coreference ResolutionabstractVid Kocijan, Oana-Maria Camburu, Ana-Maria Cretu, Yordan Yordanov, Phil Blunsom, Thomas Lukasiewicz. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Vid Kocijan, Oana-Maria Camburu, Ana-Maria Cretu 0002, Yordan Yordanov, Phil Blunsom, Thomas Lukasiewicz |
EMNLP/IJCNLP (1) | 6 |
| 2019 | Long Text Analysis Using Sliced Recurrent Neural Networks with Breaking Point Information EnrichmentabstractSliced recurrent neural networks (SRNNs) are the state-of-the-art efficient solution for long text analysis tasks; however, their slicing operations inevitably result in long-term dependency loss in lower-level networks and thus limit their accuracy. Therefore, we propose a breaking point information enrichment mechanism to strengthen dependencies between sliced subsequences without hindering parallelization. Then, the resulting BPIE-SRNN model is further extended to a bidirectional model, BPIE-BiSRNN, to utilize the dependency information in not only the previous but also the following contexts. Experiments on four large public real-world datasets demonstrate that the BPIE-SRNN and BPIE-BiSRNN models always achieve a much better accuracy than SRNNs and BiSRNNs, while maintaining a superior training efficiency. Bo Li 0099, Zehua Cheng, Zhenghua Xu 0001, Wei Ye 0004, Thomas Lukasiewicz, Shikun Zhang |
ICASSP | 5 |
| 2019 | Explanations for Query Answers under Existential RulesabstractOntology-mediated query answering is an extensively studied paradigm, which aims at improving query answers with the use of a logical theory. As a form of logical entailment, ontology-mediated query answering is fully interpretable, which makes it possible to derive explanations for query answers. Surprisingly, however, explaining answers for ontology-mediated queries has received little attention for ontology languages based on existential rules. In this paper, we close this gap, and study the problem of explaining query answers in terms of minimal subsets of database facts. We provide a thorough complexity analysis for several decision problems associated with minimal explanations under existential rules. Ismail Ilkan Ceylan, Thomas Lukasiewicz, Enrico Malizia, Andrius Vaicenavicius |
IJCAI | 2 |
| 2019 | Controllable Text-to-Image GenerationabstractIn this paper, we propose a novel controllable text-to-image generative adversarial network (ControlGAN), which can effectively synthesise high-quality images and also control parts of the image generation according to natural language descriptions. To achieve this, we introduce a word-level spatial and channel-wise attention-driven generator that can disentangle different visual attributes, and allow the model to focus on generating and manipulating subregions corresponding to the most relevant words. Also, a word-level discriminator is proposed to provide fine-grained supervisory feedback by correlating words with image regions, facilitating training an effective generator which is able to manipulate specific visual attributes without affecting the generation of other content. Furthermore, perceptual loss is adopted to reduce the randomness involved in the image generation, and to encourage the generator to manipulate specific attributes required in the modified text. Extensive experiments on benchmark datasets demonstrate that our method outperforms existing state of the art, and is able to effectively manipulate synthetic images using natural language descriptions. Bowen Li 0001, Xiaojuan Qi 0001, Thomas Lukasiewicz, Philip Torr 0001 |
NeurIPS | 3 |
| 2019 | Complexity results for preference aggregation over (m)CP-nets: Pareto and majority votingabstractAggregating preferences over combinatorial domains has many applications in artificial intelligence (AI). Given the inherent exponential nature of preferences over combinatorial domains, compact representation languages are needed to represent them, and ( m )CP-nets are among the most studied ones. Sequential and global voting are two different ways of aggregating preferences represented via CP-nets. In sequential voting, agents' preferences are aggregated feature-by-feature. For this reason, sequential voting may exhibit voting paradoxes, i.e., the possibility to select sub-optimal outcomes when preferences have specific feature dependencies. To avoid paradoxes in sequential voting, one has often assumed the (quite) restrictive constraint of O -legality, which imposes a shared common topological order among all the agents' CP-nets. On the contrary, in global voting, CP-nets are considered as a whole during the preference aggregation process. For this reason, global voting is immune from the voting paradoxes of sequential voting, and hence there is no need to impose restrictions over the CP-nets' structure when preferences are aggregated via global voting. Sequential voting over O -legal CP-nets received much attention, and O -legality of CP-nets has often been required in other studies. On the other hand, global voting over non- O -legal CP-nets has not carefully been analyzed, despite it was explicitly stated in the literature that a theoretical comparison between global and sequential voting was highly promising and a precise complexity analysis for global voting has been asked for multiple times. In quite a few works, only very partial results on the complexity of global voting over CP-nets have been given. In this paper, we start to fill this gap by carrying out a thorough computational complexity analysis of global voting tasks, for Pareto and majority voting, over not necessarily O -legal acyclic binary polynomially connected ( m )CP-nets. We show that all these problems belong to various levels of the polynomial hierarchy, and some of them are even in P or LOGSPACE. Our results are a notable achievement, given that the previously known upper bound for most of these problems was the complexity class EXPTIME. We provide various exact complexity results showing tight lower bounds and matching upper bounds for problems that (up to now) did not have any explicit non-obvious lower bound. Thomas Lukasiewicz, Enrico Malizia |
Artif. Intell. | 1 |
| 2018 | Recent Advances in Querying Probabilistic Knowledge BasesabstractWe give a survey on recent advances at the forefront of research on probabilistic knowledge bases for representing and querying large-scale automatically extracted data. We concentrate especially on increasing the semantic expressivity of formalisms for representing and querying probabilistic knowledge (i) by giving up the closed-world assumption, (ii) by allowing for commonsense knowledge (and in parallel giving up the tuple-independence assumption), and (iii) by giving up the closed-domain assumption, while preserving some computational properties of query answering in such formalisms. Stefan Borgwardt, Ismail Ilkan Ceylan, Thomas Lukasiewicz |
IJCAI | 3 |
| 2018 | Complexity of Approximate Query Answering under Inconsistency in Datalog+/-abstractSeveral semantics have been proposed to query inconsistent ontological knowledge bases, including the intersection of repairs and the intersection of closed repairs as two approximate inconsistency-tolerant semantics. In this paper, we analyze the complexity of conjunctive query answering under these two semantics for a wide range of Datalog+/- languages. We consider both the standard setting, where errors may only be in the database, and the generalized setting, where also the rules of a Datalog+/- knowledge base may be erroneous. Thomas Lukasiewicz, Enrico Malizia, Cristian Molinaro |
IJCAI | 1 |
| 2018 | e-SNLI: Natural Language Inference with Natural Language ExplanationsabstractIn order for machine learning to garner widespread public adoption, models must be able to provide interpretable and robust explanations for their decisions, as well as learn from human-provided explanations at train time. In this work, we extend the Stanford Natural Language Inference dataset with an additional layer of human-annotated natural language explanations of the entailment relations. We further implement models that incorporate these explanations into their training process and output them at test time. We show how our corpus of explanations, which we call e-SNLI, can be used for various goals, such as obtaining full sentence justifications of a model’s decisions, improving universal sentence representations and transferring to out-of-domain NLI datasets. Our dataset thus opens up a range of research directions for using natural language explanations, both for improving models and for asserting their trust Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, Phil Blunsom |
NeurIPS | 3 |
| 2018 | Ontological query answering under many-valued group preferences in Datalog+/-
Bettina Fazzinga, Thomas Lukasiewicz, Maria Vanina Martinez, Gerardo I. Simari, Oana Tifrea-Marciuska |
Int. J. Approx. Reason. | 2 |
| 2017 | Ontology-Mediated Queries for Probabilistic DatabasesabstractProbabilistic databases (PDBs) are usually incomplete, e.g., containing only the facts that have been extracted from the Web with high confidence. However, missing facts are often treated as being false, which leads to unintuitive results when querying PDBs. Recently, open-world probabilistic databases (OpenPDBs) were proposed to address this issue by allowing probabilities of unknown facts to take any value from a fixed probability interval. In this paper, we extend OpenPDBs by Datalog+/- ontologies, under which both upper and lower probabilities of queries become even more informative, enabling us to distinguish queries that were indistinguishable before. We show that the dichotomy between P and PP in (Open)PDBs can be lifted to the case of first-order rewritable positive programs (without negative constraints); and that the problem can become NP^PP-complete, once negative constraints are allowed. We also propose an approximating semantics that circumvents the increase in complexity caused by negative constraints. Stefan Borgwardt, Ismail Ilkan Ceylan, Thomas Lukasiewicz |
AAAI | 3 |
| 2017 | Location-Aware News Recommendation Using Deep Localized Semantic Analysis
Cheng Chen 0002, Thomas Lukasiewicz, Xiangwu Meng, Zhenghua Xu 0001 |
DASFAA (1) | 2 |
| 2017 | Most Probable Explanations for Probabilistic Database QueriesabstractForming the foundations of large-scale knowledge bases, probabilistic databases have been widely studied in the literature. In particular, probabilistic query evaluation has been investigated intensively as a central inference mechanism. However, despite its power, query evaluation alone cannot extract all the relevant information encompassed in large-scale knowledge bases. To exploit this potential, we study two inference tasks; namely finding the most probable database and the most probable hypothesis for a given query. As natural counterparts of most probable explanations (MPE) and maximum a posteriori hypotheses (MAP) in probabilistic graphical models, they can be used in a variety of applications that involve prediction or diagnosis tasks. We investigate these problems relative to a variety of query languages, ranging from conjunctive queries to ontology-mediated queries, and provide a detailed complexity analysis. Ismail Ilkan Ceylan, Stefan Borgwardt, Thomas Lukasiewicz |
IJCAI | 3 |
| 2017 | Query Answering in Ontologies under Preference RankingsabstractWe present an ontological framework, based on preference rankings, that allows users to express their preferences between the knowledge explicitly available in the ontology. Using this formalism, the answers for a given query to an ontology can be ranked by preference, allowing users to retrieve the most preferred answers only. We provide a host of complexity results for the main computational tasks in this framework, for the general case, and for EL and DL-Lite_core as underlying ontology languages. Ismail Ilkan Ceylan, Thomas Lukasiewicz, Rafael Peñaloza, Oana Tifrea-Marciuska |
IJCAI | 2 |
| 2017 | Tag-Aware Personalized Recommendation Using a Hybrid Deep ModelabstractRecently, many efforts have been put into tag-aware personalized recommendation. However, due to uncontrolled vocabularies, social tags are usually redundant, sparse, and ambiguous. In this paper, we propose a deep neural network approach to solve this problem by mapping the tag-based user and item profiles to an abstract deep feature space, where the deep-semantic similarities between users and their target items (resp., irrelevant items) are maximized (resp., minimized). To ensure the scalability in practice, we further propose to improve this model's training efficiency by using hybrid deep learning and negative sampling. Experimental results show that our approach can significantly outperform the state-of-the-art baselines in tag-aware personalized recommendation (3.8 times better than the best baseline), and that using hybrid deep learning and negative sampling can dramatically enhance the model's training efficiency (hundreds of times quicker), while maintaining similar (and sometimes even better) training quality and recommendation performance. Zhenghua Xu 0001, Thomas Lukasiewicz, Cheng Chen 0002, Yishu Miao, Xiangwu Meng |
IJCAI | 2 |
| 2017 | A novel characterization of the complexity class ϴPk based on counting and comparisonabstractThe complexity class Θ2P, which is the class of languages recognizable by deterministic Turing machines in polynomial time with at most logarithmic many calls to an NP oracle, received extensive attention in the literature. Its complete problems can be characterized by different specific tasks, such as deciding whether the optimum solution of an NP problem is unique, or whether it is in some sense “odd” (e.g., whether its size is an odd number). In this paper, we introduce a new characterization of this class and its generalization ΘkP to the k-th level of the polynomial hierarchy. We show that problems in ΘkP are also those whose solution involves deciding, for two given sets A and B of instances of two Σk−1P-complete (or Πk−1P-complete) problems, whether the number of “yes”-instances in A is greater than those in B. Moreover, based on this new characterization, we provide a novel sufficient condition for ΘkP-hardness. We also define the general problem Comp-Validk, which is proven here Θk+1P-complete. Comp-Validk is the problem of deciding, given two sets A and B of quantified Boolean formulas with at most k alternating quantifiers, whether the number of valid formulas in A is greater than those in B. Notably, the problem Comp-Sat of deciding whether a set contains more satisfiable Boolean formulas than another set, which is a particular case of Comp-Valid1, demonstrates itself as a very intuitive Θ2P-complete problem. Nonetheless, to our knowledge, it eluded its formal definition to date. In fact, given its strict adherence to the count-and-compare semantics here introduced, Comp-Validk is among the most suitable tools to prove ΘkP-hardness of problems involving the counting and comparison of the number of “yes”-instances in two sets. We support this by showing that the Θ2P-hardness of the Max voting scheme over mCP-nets is easily obtained via the new characterization of ΘkP introduced in this paper. Thomas Lukasiewicz, Enrico Malizia |
Theor. Comput. Sci. | 1 |
| 2016 | On the Complexity of mCP-netsabstractmCP-nets are an expressive and intuitive formalism based on CP-nets to reason about preferences of groups of agents. The dominance semantics of mCP-nets is based on the concept of voting, and different voting schemes give rise to different dominance semantics for the group. Unlike CP-nets, which received an extensive complexity analysis, mCP-nets, as reported multiple times in the literature, lack a precise study of the voting tasks' complexity. Prior to this work, only a complexity analysis of brute-force algorithms for these tasks was available, and this analysis only gave EXPTIME upper bounds for most of those problems. In this paper, we start to fill this gap by carrying out a precise computational complexity analysis of voting tasks on acyclic binary polynomially connected mCP-nets whose constituents are standard CP-nets. Interestingly, all these problems actually belong to various levels of the polynomial hierarchy, and some of them even belong to PTIME or LOGSPACE. Furthermore, for most of these problems, we provide completeness results, which show tight lower bounds for problems that (up to date) did not have any explicit non-obvious lower bound. Thomas Lukasiewicz, Enrico Malizia |
AAAI | 1 |
| 2016 | Basic Probabilistic Ontological Data Exchange with Existential RulesabstractWe study the complexity of exchanging probabilistic data between ontology-based probabilistic databases. We consider the Datalog+/- family of languages as ontology and ontology mapping languages, and we assume different compact encodings of the probabilities of the probabilistic source databases via Boolean events. We provide an extensive complexity analysis of the problem of deciding the existence of a probabilistic (universal) solution for a given probabilistic source database relative to a (probabilistic) data exchange problem for the different languages considered. Thomas Lukasiewicz, Maria Vanina Martinez, Livia Predoiu, Gerardo I. Simari |
AAAI | 1 |
| 2016 | Tag-Aware Personalized Recommendation Using a Deep-Semantic Similarity Model with Negative SamplingabstractWith the rapid growth of social tagging systems, many efforts have been put on tag-aware personalized recommendation. However, due to uncontrolled vocabularies, social tags are usually redundant, sparse, and ambiguous. In this paper, we propose a deep neural network approach to solve this problem by mapping both the tag-based user and item profiles to an abstract deep feature space, where the deep-semantic similarities between users and their target items (resp., irrelevant items) are maximized (resp., minimized). Due to huge numbers of online items, the training of this model is usually computationally expensive in the real-world context. Therefore, we introduce negative sampling, which significantly increases the model's training efficiency (109.6 times quicker) and ensures the scalability in practice. Experimental results show that our model can significantly outperform the state-of-the-art baselines in tag-aware personalized recommendation: e.g., its mean reciprocal rank is between 5.7 and 16.5 times better than the baselines. Zhenghua Xu 0001, Cheng Chen 0002, Thomas Lukasiewicz, Yishu Miao, Xiangwu Meng |
CIKM | 3 |
| 2016 | Complexity Results for Probabilistic Datalog±abstractWe study the query evaluation problem in probabilistic databases in the presence of probabilistic existential rules. Our focus is on the Datalog±family of languages for which we define the probabilistic counterpart using a flexible and compact encoding of probabilities. This formalism can be viewed as a generalization of probabilistic databases, as it allows to generate new facts from the given ones, using so-called tuple-generating dependencies, or existential rules. We study the computational cost of this additional expressiveness under two different semantics. First, we use a conventional approach and assume that the probabilistic knowledge base is consistent and employ the standard possible world semantics. Thereafter, we introduce a probabilistic inconsistency-tolerant semantics, which we call inconsistency-tolerant possible world semantics. For both of these cases, we provide a thorough complexity analysis relative to different languages, drawing a complete picture of the complexity of probabilistic query answering in this family. Ismail Ilkan Ceylan, Thomas Lukasiewicz, Rafael Peñaloza |
ECAI | 2 |
| 2016 | Complexity of Threshold Query Answering in Probabilistic Ontological Data ExchangeabstractWe study the complexity of threshold query answering in the logical framework for probabilistic ontological data exchange, which is an extension of the classical probabilistic data exchange framework with (1) probabilistic databases compactly encoded with several different annotations according to three different probability models used and (2) existential rules of different expressiveness. The ontological data exchange framework provides a logical formalization of exchanging probabilistic data and knowledge from one ontology to another via either deterministic or probabilistic mappings. We define the threshold query answering task in this framework and provide a thorough analysis of its computational complexity for different classes of existential rules and types of complexity. We also delineate several classes of existential rules and a probability model along with a compact encoding in which the threshold query answering problem can be solved in polynomial time in the data complexity. Thomas Lukasiewicz, Livia Predoiu |
ECAI | 1 |
| 2016 | Preferential Query Answering over the Semantic Web with Possibilistic Networks
Stefan Borgwardt, Bettina Fazzinga, Thomas Lukasiewicz, Akanksha Shrivastava, Oana Tifrea-Marciuska |
IJCAI | 3 |
| 2016 | Generalized Consistent Query Answering under Existential Rules
Thomas Eiter, Thomas Lukasiewicz, Livia Predoiu |
KR | 2 |
| 2016 | Probabilistic Models over Weighted Orderings: Fixed-Parameter Tractable Variable Elimination
Thomas Lukasiewicz, Maria Vanina Martinez, David Poole 0001, Gerardo I. Simari |
KR | 1 |
| 2015 | From Classical to Consistent Query Answering under Existential RulesabstractQuerying inconsistent ontologies is an intriguing new problem that gave rise to a flourishing research activity in the description logic (DL) community. The computational complexity of consistent query answering under the main DLs is rather well understood; however, little is known about existential rules. The goal of the current work is to perform an in-depth analysis of the complexity of consistent query answering under the main decidable classes of existential rules enriched with negative constraints. Our investigation focuses on one of the most prominent inconsistency-tolerant semantics, namely, the AR semantics. We establish a generic complexity result, which demonstrates the tight connection between classical and consistent query answering. This result allows us to obtain in a uniform way a relatively complete picture of the complexity of our problem. Thomas Lukasiewicz, Maria Vanina Martinez, Andreas Pieris, Gerardo I. Simari |
AAAI | 1 |
| 2015 | Combining Existential Rules with the Power of CP-Theories
Tommaso Di Noia, Thomas Lukasiewicz, Maria Vanina Martinez, Gerardo I. Simari, Oana Tifrea-Marciuska |
IJCAI | 2 |
| 2014 | Probabilistic Preference Logic NetworksabstractReasoning about an entity's preferences (be it a user of an application, an individual targeted for marketing, or a group of people whose choices are of interest) has a long history in different areas of study. In this paper, we adopt the point of view that grows out of the intersection of databases and knowledge representation, where preferences are usually represented as strict partial orders over the set of tuples in a database or the consequences of a knowledge base. We introduce probabilistic preference logic networks (PPLNs), which flexibly combine such preferences with probabilistic uncertainty. Their applications are clear in domains such as the Social Semantic Web, where users often express preferences in an incomplete manner and through different means, many times in contradiction with each other. We show that the basic problems associated with reasoning with PPLNs (computing the probability of a world or a given query) are #P-hard, and then explore ways to make these computations tractable by: (i) leveraging results from order theory to obtain a polynomial-time randomized approximation scheme (FPRAS) under fixed-parameter assumptions; and (ii) studying a fragment of the language of PPLNs for which exact computations can be performed in fixed-parameter polynomial time. Thomas Lukasiewicz, Maria Vanina Martinez, Gerardo I. Simari |
ECAI | 1 |
| 2014 | Stable Model Semantics for Guarded Existential Rules and Description Logics
Georg Gottlob, André Hernich, Clemens Kupke, Thomas Lukasiewicz |
KR | 4 |
| 2014 | Datalog+/-: Questions and Answers
Georg Gottlob, Thomas Lukasiewicz, Andreas Pieris |
KR | 2 |
| 2014 | Ontology-Based Query Answering with Group PreferencesabstractThe Web has recently been evolving into a system that is in many ways centered on social interactions and is now more and more becoming what is called the Social Semantic Web. One of the many implications of such an evolution is that the ranking of search results no longer depends solely on the structure of the interconnections among Web pages—instead, the social components must also come into play. In this article, we argue that such rankings can be based on ontological background knowledge and on user preferences. Another aspect that has become increasingly important in recent times is that of uncertainty management, since uncertainty can arise due to many uncontrollable factors. To combine these two aspects, we propose extensions of the Datalog+/-- family of ontology languages that both allow for the management of partially ordered preferences of groups of users as well as uncertainty, which is represented via a probabilistic model. We focus on answering k -rank queries in this context, presenting different strategies to compute group preferences as an aggregation of the preferences of a collection of single users. We also study merging operators that are useful for combining the preferences of the users with those induced by the values obtained from the probabilistic model. We then provide algorithms to answer k -rank queries for DAQs (disjunctions of atomic queries) under these group preferences and uncertainty that generalizes top- k queries based on the iterative computation of classical skyline answers. We show that such DAQ answering in Datalog+/-- can be done in polynomial time in the data complexity, under certain reasonable conditions, as long as query answering can also be done in polynomial time (in the data complexity) in the underlying classical ontology. Finally, we present a prototype implementation of the query answering system, as well as experimental results (on the running time of our algorithms and the quality of their results) obtained from real-world ontological data and preference models, derived from information gathered from real users, showing in particular that our approach is feasible in practice. Thomas Lukasiewicz, Maria Vanina Martinez, Gerardo I. Simari, Oana Tifrea-Marciuska |
ACM Trans. Internet Techn. | 1 |
| 2013 | Preference-Based Query Answering in Datalog+/- Ontologies
Thomas Lukasiewicz, Maria Vanina Martinez, Gerardo I. Simari |
IJCAI | 1 |
| 2013 | Well-founded semantics for extended datalog and ontological reasoningabstractThe Datalog± family of expressive extensions of Datalog has recently been introduced as a new paradigm for query answering over ontologies, which captures and extends several common description logics. It extends plain Datalog by features such as existentially quantified rule heads and, at the same time, restricts the rule syntax so as to achieve decidability and tractability. In this paper, we continue the research on Datalog±. More precisely, we generalize the well-founded semantics (WFS), as the standard semantics for nonmonotonic normal programs in the database context, to Datalog± programs with negation under the unique name assumption (UNA). We prove that for guarded Datalog± with negation under the standard WFS, answering normal Boolean conjunctive queries is decidable, and we provide precise complexity results for this problem, namely, in particular, completeness for PTIME (resp., 2-EXPTIME) in the data (resp., combined) complexity. André Hernich, Clemens Kupke, Thomas Lukasiewicz, Georg Gottlob |
PODS | 3 |
| 2013 | Workshop on recommender systems meet big data & semantic technologies: SeRSy 2013abstractThe primary goal of the workshop is to showcase cutting edge research on the intersection of Recommender Systems and Semantic Technologies, by taking the best of the two worlds. This combination may provide the RecSys community with important scenarios where the potential of Semantic Technologies can be effectively exploited into systems performing complex tasks, such as recommendation engines processing Big Data. Marco de Gemmis, Tommaso Di Noia, Ora Lassila, Pasquale Lops, Thomas Lukasiewicz, Giovanni Semeraro |
RecSys | 5 |
| 2013 | Query Answering in Probabilistic Datalog+/- Ontologies under Group PreferencesabstractIn the recent years, the Web has been changing more and more towards the so-called Social Semantic Web. Rather than being based on the link structure between Web pages, the ranking of search results in the Social Semantic Web needs to be based on something new - we believe that it can be based on user preferences and underlying ontological knowledge. Modeling uncertainty is also playing an increasingly important role in these domains, since uncertainty can arise due to many uncontrollable factors. In this paper, we propose an extension of the Data log+/- ontology language with a model for representing preferences of groups of users and a model for representing the (probabilistic) uncertainty in the domain. Assuming that more probable answers are more preferable, this raises the question of how to rank query results, since the preferences of single users may be in conflict both with the probability-based preferences as well as with each other. To this end, we propose preference merging and aggregation operators, respectively, and study their semantic and computational properties. Based on these operators, we provide algorithms for answering k-rank queries for DAQs (disjunctions of atomic queries), which generalize top-k queries based on the iterative computation of classical skyline answers, and show that, under certain reasonable conditions, they run in polynomial time in the data complexity. Thomas Lukasiewicz, Maria Vanina Martinez, Gerardo I. Simari, Oana Tifrea-Marciuska |
Web Intelligence | 1 |
| 2012 | Equality-Friendly Well-Founded Semantics and Applications to Description LogicsabstractWe tackle the problem of defining a well-founded semantics for Datalog rules with existentially quantified variables in their heads and negations in their bodies. In particular, we provide a well-founded semantics (WFS) for the recent Datalog+/- family of ontology languages, which covers several important description logics (DLs). To do so, we generalize Datalog+/- by non-stratified nonmonotonic negation in rule bodies, and we define a WFS for this generalization via guarded fixed-point logic. We refer to this approach as equality-friendly WFS, since it has the advantage that it does not make the unique name assumption (UNA); this brings it close to OWL and its profiles as well as typical DLs, which also do not make the UNA. We prove that for guarded Datalog+/- with negation under the equality-friendly WFS, conjunctive query answering is decidable, and we provide precise complexity results for this problem. From these results, we obtain precise definitions of the standard WFS extensions of EL and of members of the DL-Lite family, as well as corresponding complexity results for query answering. Georg Gottlob, André Hernich, Clemens Kupke, Thomas Lukasiewicz |
AAAI | 4 |
| 2012 | Heuristic Ranking in Tightly Coupled Probabilistic Description Logics
Thomas Lukasiewicz, Maria Vanina Martinez, Giorgio Orsi 0001, Gerardo I. Simari |
UAI | 1 |
| 2012 | A general Datalog-based framework for tractable query answering over ontologies
Andrea Calì, Georg Gottlob, Thomas Lukasiewicz |
J. Web Semant. | 3 |
| 2011 | Well-founded semantics for description logic programs in the semantic webabstractThe realization of the Semantic Web vision, in which computational logic has a prominent role, has stimulated a lot of research on combining rules and ontologies, which are formulated in different formalisms. In particular, combining logic programming with the Web Ontology Language (OWL), which is a standard based on description logics, emerged as an important issue for linking the Rules and Ontology Layers of the Semantic Web. Nonmonotonic description logic programs (dl-programs) were introduced for such a combination, in which a pair(L,P)of a description logic knowledge baseLand a set of rulesPwith negation as failure is given a model-based semantics that generalizes the answer set semantics of logic programs. In this article, we reconsider dl-programs and present a well-founded semantics for them as an analog for the other main semantics of logic programs. It generalizes the canonical definition of the well-founded semantics based on unfounded sets, and, as we show, lifts many of the well-known properties from ordinary logic programs to dl-programs. Among these properties, our semantics amounts to a partial model approximating the answer set semantics, which yields for positive and stratified dl-programs, a total model coinciding with the answer set semantics; it has polynomial data complexity provided the access to the description logic knowledge base is polynomial; under suitable restrictions, it has lower complexity and even first-order rewritability is achievable. The results add to previous evidence that dl-programs are a versatile and robust combination approach, which moreover is implementable using legacy engines. Thomas Eiter, Giovambattista Ianni, Thomas Lukasiewicz, Roman Schindlauer |
ACM Trans. Comput. Log. | 3 |
| 2011 | Semantic Web search based on ontological conjunctive queries
Bettina Fazzinga, Giorgio Gianforme, Georg Gottlob, Thomas Lukasiewicz |
J. Web Semant. | 4 |
| 2010 | Ontological Reasoning with F-logic Lite and its ExtensionsabstractAnswering queries posed over knowledge bases is a central problem in knowledge representation and database theory. In the database area, checking query containment is an important query optimization and schema integration technique. In knowledge representation it has been used for object classification, schema integration, service discovery, and more. In the presence of a knowledge base, the problem of query containment is strictly related to that of query answering; indeed, the two are reducible to each other; we focus on the latter, and our results immediately extend to the former. Andrea Calì, Georg Gottlob, Michael Kifer, Thomas Lukasiewicz, Andreas Pieris |
AAAI | 4 |
| 2010 | Datalog+/-: A Family of Logical Knowledge Representation and Query Languages for New ApplicationsabstractThis paper summarizes results on a recently introduced family of Datalog-based languages, called Datalog+/-, which is a new framework for tractable ontology querying, and for a variety of other applications. Datalog+/- extends plain Datalog by features such as existentially quantified rule heads and, at the same time, restricts the rule syntax so as to achieve decidability and tractability. In particular, we discuss three paradigms ensuring decidability: chase termination, guardedness, and stickiness. Andrea Calì, Georg Gottlob, Thomas Lukasiewicz, Bruno Marnette, Andreas Pieris |
LICS | 3 |
| 2010 | A Novel Combination of Answer Set Programming with Description Logics for the Semantic WebabstractWe present a novel combination of disjunctive programs under the answer set semantics with description logics for the Semantic Web. The combination is based on a well-balanced interface between disjunctive programs and description logics, which guarantees the decidability of the resulting formalism without assuming syntactic restrictions. We show that the new formalism has very nice semantic properties. In particular, it faithfully extends both disjunctive programs and description logics. Furthermore, we describe algorithms for reasoning in the new formalism, and we give a precise picture of its computational complexity. We also define the well-founded semantics for the normal case, where normal programs are combined with tractable description logics, and we explore its semantic and computational properties. In particular, we show that the well-founded semantics approximates the answer set semantics. We also describe algorithms for the problems of consistency checking and literal entailment under the well-founded semantics, and we give a precise picture of their computational complexity. As a crucial property, in the normal case, consistency checking and literal entailment under the well-founded semantics are both tractable in the data complexity, and even first-order rewritable (and thus can be done in LogSpace in the data complexity) in a special case that is especially useful for representing mappings between ontologies. Thomas Lukasiewicz |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2009 | Datalog±: a unified approach to ontologies and integrity constraintsabstractWe report on a recently introduced family of expressive extensions of Datalog, called Datalog±, which is a new framework for representing ontological axioms in form of integrity constraints, and for query answering under such constraints. Datalog± is derived from Datalog by allowing existentially quantified variables in rule heads, and by enforcing suitable properties in rule bodies, to ensure decidable and efficient query answering. We first present different languages in the Datalog± family, providing tight complexity bounds for all cases but one (where we have a low complexity AC0 upper bound). We then show that such languages are general enough to capture the most common tractable ontology languages. In particular, we show that the DL-Lite family of description logics and F-Logic Lite are expressible in Datalog±. We finally show how stratified negation can be added to Datalog± while keeping ontology querying tractable in the data complexity. Datalog± is a natural and very general framework that can be successfully employed in different contexts such as data integration and exchange. This survey mainly summarizes two recent papers. Andrea Calì, Georg Gottlob, Thomas Lukasiewicz |
ICDT | 3 |
| 2009 | Inductive Query Answering and Concept Retrieval Exploiting Local ModelsabstractWe present a classification method, founded in the instance-based learning and the disjunctive version space approach, for performing approximate retrieval from knowledge bases expressed in Description Logics. It is able to supply answers, even though they are not logically entailed by the knowledge base (e.g. because of its incompleteness or when there are inconsistent assertions). Moreover, the method may also induce new knowledge that can be employed to make the ontology population task semiautomatic. The method has been experimentally tested showing that it is sound and effective. Claudia d'Amato, Nicola Fanizzi, Floriana Esposito, Thomas Lukasiewicz |
ISDA | 4 |
| 2009 | A general datalog-based framework for tractable query answering over ontologiesabstractIn this paper, we introduce a family of expressive extensions of Datalog, called Datalog+/-, as a new paradigm for query answering over ontologies. The Datalog+/- family admits existentially quantified variables in rule heads, and has suitable restrictions to ensure highly efficient ontology querying. We show in particular that Datalog+/- generalizes the DL-Lite family of tractable description logics, which are the most common tractable ontology languages in the context of the Semantic Web and databases. We also show how stratified negation can be added to Datalog+/- while keeping ontology querying tractable. Furthermore, the Datalog+/- family is of interest in its own right and can, moreover, be used in various contexts such as data integration and data exchange. Andrea Calì, Georg Gottlob, Thomas Lukasiewicz |
PODS | 3 |
| 2009 | Approximate Classification of Semantically Annotated Web Resources Exploiting Pseudo-metrics Induced by Local ModelsabstractWe present a classification method, founded in the instance-based learning and the disjunctive version space approach, for performing approximate retrieval from knowledge bases expressed in Description Logics. The method supplies answers even if the knowledge base of reference is inconsistent or incomplete. Moreover, the method may also induce new knowledge that can be suggested to the knowledge engineer, thus making the ontology population task semi-automatic. Claudia d'Amato, Nicola Fanizzi, Floriana Esposito, Thomas Lukasiewicz |
Web Intelligence | 4 |
| 2009 | Description logic programs under probabilistic uncertainty and fuzzy vagueness
Thomas Lukasiewicz, Umberto Straccia |
Int. J. Approx. Reason. | 1 |
| 2009 | Reasoning about actions with sensing under qualitative and probabilistic uncertaintyabstractWe focus on the aspect of sensing in reasoning about actions under qualitative and probabilistic uncertainty. We first define the action language E for reasoning about actions with sensing, which has a semantics based on the autoepistemic description logic ALCK NF , and which is given a formal semantics via a system of deterministic transitions between epistemic states. As an important feature, the main computational tasks in E can be done in linear and quadratic time. We then introduce the action language E + for reasoning about actions with sensing under qualitative and probabilistic uncertainty, which is an extension of E by actions with nondeterministic and probabilistic effects, and which is given a formal semantics in a system of deterministic, nondeterministic, and probabilistic transitions between epistemic states. We also define the notion of a belief graph, which represents the belief state of an agent after a sequence of deterministic, nondeterministic, and probabilistic actions, and which compactly represents a set of unnormalized probability distributions. Using belief graphs, we then introduce the notion of a conditional plan and its goodness for reasoning about actions under qualitative and probabilistic uncertainty. We formulate the problems of optimal and threshold conditional planning under qualitative and probabilistic uncertainty, and show that they are both uncomputable in general. We then give two algorithms for conditional planning in our framework. The first one is always sound, and it is also complete for the special case in which the relevant transitions between epistemic states are cycle-free. The second algorithm is a sound and complete solution to the problem of finite-horizon conditional planning in our framework. Under suitable assumptions, it computes every optimal finite-horizon conditional plan in polynomial time. We also describe an application of our formalism in a robotic-soccer scenario, which underlines its usefulness in realistic applications. Luca Iocchi, Thomas Lukasiewicz, Daniele Nardi, Riccardo Rosati 0001 |
ACM Trans. Comput. Log. | 2 |
| 2008 | Combining answer set programming with description logics for the Semantic Web
Thomas Eiter, Giovambattista Ianni, Thomas Lukasiewicz, Roman Schindlauer, Hans Tompits |
Artif. Intell. | 3 |
| 2008 | Expressive probabilistic description logics
Thomas Lukasiewicz |
Artif. Intell. | 1 |
| 2008 | Fuzzy Description Logic Programs under the Answer Set Semantics for the Semantic Web
Thomas Lukasiewicz |
Fundam. Informaticae | 1 |
| 2008 | Logical approaches to imprecise probabilities
Thomas Lukasiewicz |
Int. J. Approx. Reason. | 1 |
| 2008 | Probabilistic description logic programs under inheritance with overriding for the Semantic Web
Thomas Lukasiewicz |
Int. J. Approx. Reason. | 1 |
| 2008 | Tightly Coupled Fuzzy Description Logic Programs under the Answer Set Semantics for the Semantic WebabstractWe present a novel approach to fuzzy description logic programs (or simply fuzzy dl-programs) under the answer set semantics, which is a tight integration of fuzzy disjunctive logic programs under the answer set semantics with fuzzy description logics. From a different perspective, it is a generalization of tightly coupled disjunctive dl-programs by fuzzy vagueness in both the description logic and the logic program component. We show that the new formalism faithfully extends both fuzzy disjunctive logic programs and fuzzy description logics, and that under suitable assumptions, reasoning in the new formalism is decidable. We present a polynomial reduction of certain fuzzy dl-programs to tightly coupled disjunctive dl-programs, and we analyze the complexity of consistency checking and query processing for certain fuzzy dl-programs. Furthermore, we provide a special case of fuzzy dl-programs for which deciding consistency and query processing can both be done in polynomial time in the data complexity. Thomas Lukasiewicz, Umberto Straccia |
Int. J. Semantic Web Inf. Syst. | 1 |
| 2008 | Managing uncertainty and vagueness in description logics for the Semantic Web
Thomas Lukasiewicz, Umberto Straccia |
J. Web Semant. | 1 |
| 2007 | Description Logic Programs Under Probabilistic Uncertainty and Fuzzy Vagueness
Thomas Lukasiewicz, Umberto Straccia |
ECSQARU | 1 |
| 2007 | A Novel Combination of Answer Set Programming with Description Logics for the Semantic Web
Thomas Lukasiewicz |
ESWC | 1 |
| 2007 | Tightly Integrated Probabilistic Description Logic Programs for the Semantic Web
Andrea Calì, Thomas Lukasiewicz |
ICLP | 2 |
| 2007 | Team Programming in Golog under Partial Observability
Alessandro Farinelli, Alberto Finzi, Thomas Lukasiewicz |
IJCAI | 3 |
| 2007 | Reasoning with imprecise probabilities
Andrés Cano, Fábio G. Cozman, Thomas Lukasiewicz |
Int. J. Approx. Reason. | 3 |
| 2007 | Nonmonotonic probabilistic logics under variable-strength inheritance with overriding: Complexity, algorithms, and implementation
Thomas Lukasiewicz |
Int. J. Approx. Reason. | 1 |
| 2007 | Probabilistic description logic programs
Thomas Lukasiewicz |
Int. J. Approx. Reason. | 1 |
| 2007 | Variable-strength conditional preferences for ranking objects in ontologies
Thomas Lukasiewicz, Jörg Schellhase |
J. Web Semant. | 1 |
| 2006 | Adaptive Multi-Agent Programming in GTGolog
Alberto Finzi, Thomas Lukasiewicz |
ECAI | 2 |
| 2006 | Variable-Strength Conditional Preferences for Ranking Objects in Ontologies
Thomas Lukasiewicz, Jörg Schellhase |
ESWC | 1 |
| 2006 | Variable-Strength Conditional Preferences for Matchmaking in Description Logics
Thomas Lukasiewicz, Jörg Schellhase |
KR | 1 |
| 2006 | Causes and explanations in the structural-model approach: Tractable cases
Thomas Eiter, Thomas Lukasiewicz |
Artif. Intell. | 2 |
| 2005 | Probabilistic Description Logic Programs
Thomas Lukasiewicz |
ECSQARU | 1 |
| 2005 | Game-Theoretic Reasoning About Actions in Nonmonotonic Causal Theories
Alberto Finzi, Thomas Lukasiewicz |
LPNMR | 2 |
| 2005 | Weak nonmonotonic probabilistic logics
Thomas Lukasiewicz |
Artif. Intell. | 1 |
| 2004 | Game-Theoretic Agent Programming in Golog
Alberto Finzi, Thomas Lukasiewicz |
ECAI | 2 |
| 2004 | Reasoning about Actions with Sensing under Qualitative and Probabilistic Uncertainty
Luca Iocchi, Thomas Lukasiewicz, Daniele Nardi, Riccardo Rosati 0001 |
ECAI | 2 |
| 2004 | Relational Markov Games
Alberto Finzi, Thomas Lukasiewicz |
JELIA | 2 |
| 2004 | Combining Answer Set Programming with Description Logics for the Semantic Web
Thomas Eiter, Thomas Lukasiewicz, Roman Schindlauer, Hans Tompits |
KR | 2 |
| 2004 | Weak Nonmonotonic Probabilistic Logics
Thomas Lukasiewicz |
KR | 1 |
| 2004 | Complexity results for explanations in the structural-model approach
Thomas Eiter, Thomas Lukasiewicz |
Artif. Intell. | 2 |
| 2004 | Combining probabilistic logic programming with the power of maximum entropyabstractThis paper is on the combination of two powerful approaches to uncertain reasoning: logic programming in a probabilistic setting, on the one hand, and the information-theoretical principle of maximum entropy, on the other hand. More precisely, we present two approaches to probabilistic logic programming under maximum entropy. The first one is based on the usual notion of entailment under maximum entropy, and is defined for the very general case of probabilistic logic programs over Boolean events. The second one is based on a new notion of entailment under maximum entropy, where the principle of maximum entropy is coupled with the closed world assumption (CWA) from classical logic programming. It is only defined for the more restricted case of probabilistic logic programs over conjunctive events. We then analyze the nonmonotonic behavior of both approaches along benchmark examples and along general properties for default reasoning from conditional knowledge bases. It turns out that both approaches have very nice nonmonotonic features. In particular, they realize some inheritance of probabilistic knowledge along subclass relationships, without suffering from the problem of inheritance blocking and from the drowning problem. They both also satisfy the property of rational monotonicity and several irrelevance properties. We finally present algorithms for both approaches, which are based on generalizations of recent techniques for probabilistic logic programming under logical entailment. The algorithm for the first approach still produces quite large weighted entropy maximization problems, while the one for the second approach generates optimization problems of the same size as the ones produced in probabilistic logic programming under logical entailment. Gabriele Kern-Isberner, Thomas Lukasiewicz |
Artif. Intell. | 2 |
| 2003 | Probabilistic Lexicographic Entailment under Variable-Strength Inheritance with Overriding
Thomas Lukasiewicz |
ECSQARU | 1 |
| 2003 | Probabilistic Reasoning about Actions in Nonmonotonic Causal Theories
Thomas Eiter, Thomas Lukasiewicz |
UAI | 2 |
| 2003 | Structure-Based Causes and Explanations in the Independent Choice Logic
Alberto Finzi, Thomas Lukasiewicz |
UAI | 2 |
| 2003 | Temporal Probabilistic Object BasesabstractThere are numerous applications where we have to deal with temporal uncertainty associated with objects. The ability to automatically store and manipulate time, probabilities, and objects is important. We propose a data model and algebra for temporal probabilistic object bases (TPOBs), which allows us to specify the probability with which an event occurs at a given time point. In explicit TPOB-instances, the sets of time points along with their probability intervals are explicitly enumerated. In implicit TPOB-instances, sets of time points are expressed by constraints and their probability intervals by probability distribution functions. Thus, implicit object base instances are succinct representations of explicit ones; they allow for an efficient implementation of algebraic operations, while their explicit counterparts make defining algebraic operations easy. We extend the relational algebra to both explicit and implicit instances and prove that the operations on implicit instances correctly implement their counterpart on explicit instances. Veronica Biazzo, Rosalba Giugno, Thomas Lukasiewicz, V. S. Subrahmanian |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2002 | P-SHOQ(D): A Probabilistic Extension of SHOQ(D) for Probabilistic Ontologies in the Semantic Web
Rosalba Giugno, Thomas Lukasiewicz |
JELIA | 2 |
| 2002 | Complexity Results for Explanations in the Structural-Model Approach
Thomas Eiter, Thomas Lukasiewicz |
KR | 2 |
| 2002 | Causes and Explanations in the Structural-Model Approach : Tractable Cases
Thomas Eiter, Thomas Lukasiewicz |
UAI | 2 |
| 2002 | Complexity results for structure-based causality
Thomas Eiter, Thomas Lukasiewicz |
Artif. Intell. | 2 |
| 2001 | Probabilistic Logic under Coherence, Model-Theoretic Probabilistic Logic, and Default Reasoning
Veronica Biazzo, Angelo Gilio, Thomas Lukasiewicz, Giuseppe Sanfilippo |
ECSQARU | 3 |
| 2001 | Complexity Results for Structure-Based Causality
Thomas Eiter, Thomas Lukasiewicz |
IJCAI | 2 |
| 2001 | Fixpoint Characterizations for Many-Valued Disjunctive Logic Programs with Probabilistic Semantics
Thomas Lukasiewicz |
LPNMR | 1 |
| 2001 | Probabilistic Logic Programming under Inheritance with Overriding
Thomas Lukasiewicz |
UAI | 1 |
| 2001 | Probabilistic logic programming with conditional constraintsabstractWe introduce a new approach to probabilistic logic programming in which probabilities are defined over a set of possible worlds. More precisely, classical program clauses are extended by a subinterval of [0,1] that describes a range for the conditional probability of the head of a clause given its body. We then analyze the complexity of selected probabilistic logic programming tasks. It turns out that probabilistic logic programming is computationally more complex than classical logic programming, More precisely, the tractability of special cases of classical logic programming generally does not carry over to the corresponding special cases of probabilistic logic programming. Moreover, we also draw a precise picture of the complexity of deciding and computing tight logical consequences in probabilistic reasoning with conditional constraints in general. We then present linear optimization techniques for deciding satisfiability and computing tight logical consequencesof probabilistic logic programs. These techniques are efficient in the special case in which we have little relevant purely probabilistic knowledge. We finally show that probabilistic logic programming under certain syntactic and semantic restrictions is closely related to van Emden's quantitative deduction, and thus has computational properties similar to calssical logic programming. Based on this result, we present an efficient approximation technique for probabilistic logic programming. Thomas Lukasiewicz |
ACM Trans. Comput. Log. | 1 |
| 2001 | Probabilistic object basesabstractAlthough there are many applications where an object-oriented data model is a good way of representing and querying data, current object database systems are unable to handle objects whose attributes are uncertain. In this article, we extend previous work by Kornatzky and Shimony to develop an algebra to handle object bases with uncertainty. We propose concepts of consistency for such object bases, together with an NP-completeness result, and classes of probabilistic object bases for which consistency is polynomially checkable. In addition, as certain operations involve conjunctions and disjunctions of events, and as the probability of conjunctive and disjunctive events depends both on the probabilities of the primitive events involved as well as on what is known (if anything) about the relationship between the events, we show how all our algebraic operations may be performed under arbitrary probabilistic conjunction and disjunction strategies. We also develop a host of equivalence results in our algebra, which may be used as rewrite rules for query optimization. Last but not least, we have developed a prototype probabilistic object base server on top of ObjectStore. We describe experiments to assess the efficiency of different possible rewrite rules. Thomas Eiter, James J. Lu, Thomas Lukasiewicz, V. S. Subrahmanian |
ACM Trans. Database Syst. | 3 |
| 2000 | Complexity Results for Default Reasoning from Conditional Knowledge Bases
Thomas Eiter, Thomas Lukasiewicz |
KR | 2 |
| 2000 | Credal Networks under Maximum Entropy
Thomas Lukasiewicz |
UAI | 1 |
| 2000 | Default reasoning from conditional knowledge bases: Complexity and tractable cases
Thomas Eiter, Thomas Lukasiewicz |
Artif. Intell. | 2 |
| 1999 | Many-Valued Disjunctive Logic Programs with Probabilistic Semantics
Thomas Lukasiewicz |
LPNMR | 1 |
| 1999 | Local probabilistic deduction from taxonomic and probabilistic knowledge-bases over conjunctive eventsabstractWe elaborate locally complete inference rules for probabilistic deduction from taxonomic and probabilistic knowledge-bases over conjunctive events. We integrate the presented inference rules into a local probabilistic deduction technique, which exploits taxonomic knowledge for an efficient representation of conjunctive events. This local probabilistic deduction technique is less incomplete and more efficient than already existing local approaches to probabilistic deduction. However, we show that it cannot compete with global probabilistic deduction by linear programming. Surprisingly, we can provide examples of globally very incomplete probabilistic deductions in the presented local approach. More generally, we even show that all systems of inference rules for probabilistic deduction in taxonomic and probabilistic knowledge-bases over conjunctive events that have a limited number of probabilistic formulas in the premises of their inference patterns are globally incomplete. Furthermore, we show that the presented local approach is not more efficient than the linear programming approach for that framework. We conclude that probabilistic deduction by the iterative application of inference rules on interval restrictions for conditional probabilities, even though considered very promising in the literature so far, is very limited in its field of application. Thomas Lukasiewicz |
Int. J. Approx. Reason. | 1 |
| 1999 | Probabilistic Deduction with Conditional Constraints over Basic EventsabstractWe study the problem of probabilistic deduction with conditional constraints over basic events. We show that globally complete probabilistic deduction with conditional constraints over basic events is NP-hard. We then concentrate on the special case of probabilistic deduction in conditional constraint trees. We elaborate very efficient techniques for globally complete probabilistic deduction. In detail, for conditional constraint trees with point probabilities, we present a local approach to globally complete probabilistic deduction, which runs in linear time in the size of the conditional constraint trees. For conditional constraint trees with interval probabilities, we show that globally complete probabilistic deduction can be done in a global approach by solving nonlinear programs. We show how these nonlinear programs can be transformed into equivalent linear programs, which are solvable in polynomial time in the size of the conditional constraint trees. Thomas Lukasiewicz |
J. Artif. Intell. Res. | 1 |
| 1998 | Probabilistic Logic Programming
Thomas Lukasiewicz |
ECAI | 1 |
| 1998 | Probabilistic Deduction with Conditional Constraints over Basic Events
Thomas Lukasiewicz |
KR | 1 |
| 1998 | Magic Inference Rules for Probabilistic Deduction under Taxonomic Knowledge
Thomas Lukasiewicz |
UAI | 1 |
| 1997 | Efficient Global Probabilistic Deduction from Taxonomic and Probabilistic Knowledge-Bases over Conjunctive EventsabstractWe present a new, efficient linear programming approach to probabilistic deduction from probabilistic knowledge-bases over conjunctive events. We show that this approach enables us to solve the classical problem of probabilistic deduction along a chain of basic events in polynomial time in the length of the chain. We then elaborate how taxonomic knowledge can be exploited in our new approach for an increased efficiency. We also present important new results for the classical linear programming approach to probabilistic deduction under taxonomic knowledge. 1 Introduction There are many approaches to non-Bayesian probabilistic deduction in the literature. They can be classified in global techniques based on linear programming and in local methods founded on the iterative application of inference rules. Non-Bayesian probabilistic deduction by solving linear programs is discussed e.g. in [23], [13], [24], [17], [14], [2], [15], and [22]. It can be performed within rich probabilistic lang... Thomas Lukasiewicz |
CIKM | 1 |
| 1995 | Taxonomic and Uncertain Integrity Constraints in Object-Oriented Databases - the TOP ApproachabstractWe present a coherent modeling and reasoning methodology to extend object-oriented databases towards taxonomic and uncertain integrity constraints. Our so-called TOP database model enriches current ISA-hierarchies by more general tclasses to improve conceptual modeling. The t-classes themselves are then integrated with probabilistic constraints to express uncertainty. We give an efficient algorithm for checking the modeling-consistency of a probabilistic knowledgebase. As a typical application domain for TOP, we exemplify how various aspects of a portfolio management system can be modeled. We also demonstrate that recent probabilistic inference methods, relying on a careful interaction between taxonomic and uncertain knowledge, can be applied in this context. 1 Introduction The evolution of database technology from the relational model into object-oriented databases (OODBs) and deductive databases has been pushed forward to a state where stable and usable systems are becoming widely ... Thomas Lukasiewicz, Werner Kießling, Gerhard Köstler, Ulrich Güntzer |
CIKM | 1 |
| 1995 | Uncertain Reasoning in Concept Lattices
Thomas Lukasiewicz |
ECSQARU | 1 |