Sebastijan Dumancic

dblp:182/1967 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0003-0915-8034ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 T4G: Trace-based P4 Program Generation
Chenxing Ji, Timo Jugariu, Sebastijan Dumancic, Fernando A. Kuipers
INFOCOM3
2025 A Divide, Align and Conquer Strategy For Program Synthesis
abstract
A major bottleneck in search-based program synthesis, which learns programs from input/output examples, is the synthesis of large programs. As the size of the target program increases, so does the search depth, which leads to an exponentially growing number of candidate programs. Humans mitigate the combinatorial explosion that arises from deep program search: they build complex programs from smaller parts. We introduce a new strategy for program synthesis called Divide, Align & Conquer (DA&C) that exploits the compositionality of real-world domains to guide the synthesis towards useful subprograms. Divide decomposes each example using a segmentation procedure that is synthesized as part of the learning problem. Align matches the components in the decomposed input/output examples to steer the search toward combinations that lead to the synthesis of useful subprograms, and Conquer then solves a standalone synthesis problem on each pair of aligned input/output components. We show how replacing a deep program search with a linear number of much smaller synthesis tasks leads us to efficiently discover useful subprograms that are then combined into a solution program. Our agent outperforms current Inductive Logic Programming (ILP) methods on string transformation tasks even with minimal knowledge priors. Unlike existing methods, the predictive accuracy of our agent monotonically increases for additional examples. It approximates an average time complexity of O(m) in the size m of subprograms for highly structured and, hence, decomposable domains such as strings. Finally, we demonstrate the scalability of our technique on highdimensional abstract visual reasoning tasks from the Abstract Reasoning Corpus (ARC) for which ILP methods were previously infeasible. We are competitive with state-of-the-art agents outside of ILP, despite generating only 0.2% as many candidate programs from a knowledge prior of only 11 generic geometric primitives.
Jonas Witt, Sebastijan Dumancic, Tias Guns, Claus-Christian Carbon
J. Artif. Intell. Res.2
2024 DeepSaDe: Learning Neural Networks That Guarantee Domain Constraint Satisfaction
abstract
As machine learning models, specifically neural networks, are becoming increasingly popular, there are concerns regarding their trustworthiness, specially in safety-critical applications, e.g. actions of an autonomous vehicle must be safe. There are approaches that can train neural networks where such domain requirements are enforced as constraints, but they either cannot guarantee that the constraint will be satisfied by all possible predictions (even on unseen data) or they are limited in the type of constraints that can be enforced. In this paper, we present an approach to train neural networks which can enforce a wide variety of constraints and guarantee that the constraint is satisfied by all possible predictions. The approach builds on earlier work where learning linear models is formulated as a constraint satisfaction problem (CSP). To make this idea applicable to neural networks, two crucial new elements are added: constraint propagation over the network layers, and weight updates based on a mix of gradient descent and CSP solving. Evaluation on various machine learning tasks demonstrates that our approach is flexible enough to enforce a wide variety of domain constraints and is able to guarantee them in neural networks.
Kshitij Goyal, Sebastijan Dumancic, Hendrik Blockeel
AAAI2
2024 Learning Logic Programs by Discovering Higher-Order Abstractions
Céline Hocquette, Sebastijan Dumancic, Andrew Cropper
IJCAI2
2024 From statistical relational to neurosymbolic artificial intelligence: A survey
Giuseppe Marra, Sebastijan Dumancic, Robin Manhaeve, Luc De Raedt
Artif. Intell.2
2023 Embedding a Long Short-Term Memory Network in a Constraint Programming Framework for Tomato Greenhouse Optimisation
abstract
Increasing global food demand, accompanied by the limited number of expert growers, brings the need for more sustainable and efficient horticulture. The controlled environment of greenhouses enable data collection and precise control. For optimally controlling the greenhouse climate, a grower not only looks at crop production, but rather aims at maximising the profit. However this is a complex, long term optimisation task. In this paper, Constraint Programming (CP) is applied to task of optimal greenhouse economic control, by leveraging a learned greenhouse climate model through a CP embedding. In collaboration with an industrial partner, we demonstrate how to model the greenhouse climate with an LSTM model, embed this LSTM into a CP optimisation framework, and optimise the expected profit of the grower. This data-to-decision pipeline is being integrated into a decision support system for multiple greenhouses in the Netherlands.
Dirk van Bokkem, Max van den Hemel, Sebastijan Dumancic, Neil Yorke-Smith
AAAI3
2022 SaDe: Learning Models that Provably Satisfy Domain Constraints
Kshitij Goyal, Sebastijan Dumancic, Hendrik Blockeel
ECML/PKDD (5)2
2022 Inductive Logic Programming At 30: A New Introduction
abstract
Inductive logic programming (ILP) is a form of machine learning. The goal of ILP is to induce a hypothesis (a set of logical rules) that generalises training examples. As ILP turns 30, we provide a new introduction to the field. We introduce the necessary logical notation and the main learning settings; describe the building blocks of an ILP system; compare several systems on several dimensions; describe four systems (Aleph, TILDE, ASPAL, and Metagol); highlight key application areas; and, finally, summarise current limitations and directions for future research.
Andrew Cropper, Sebastijan Dumancic
J. Artif. Intell. Res.2
2022 Inductive logic programming at 30
abstract
Abstract Inductive logic programming (ILP) is a form of logic-based machine learning. The goal is to induce a hypothesis (a logic program) that generalises given training examples and background knowledge. As ILP turns 30, we review the last decade of research. We focus on (i) new meta-level search methods, (ii) techniques for learning recursive programs, (iii) new approaches for predicate invention, and (iv) the use of different technologies. We conclude by discussing current limitations of ILP and directions for future research.
Andrew Cropper, Sebastijan Dumancic, Richard Evans 0001, Stephen H. Muggleton
Mach. Learn.2
2021 Knowledge Refactoring for Inductive Program Synthesis
abstract
Humans constantly restructure knowledge to use it more efficiently. Our goal is to give a machine learning system similar abilities so that it can learn more efficiently. We introduce the knowledge refactoring problem, where the goal is to restructure a learner's knowledge base to reduce its size and to minimise redundancy in it. We focus on inductive logic programming, where the knowledge base is a logic program. We introduce Knorf, a system which solves the refactoring problem using constraint optimisation. A key feature of Knorf is that, rather than simply removing knowledge, it also introduces new knowledge through predicate invention. We evaluate our approach on two domains: building Lego structures and real-world string transformations. Our experiments show that learning from refactored knowledge can improve predictive accuracies fourfold and reduce learning times by half.
Sebastijan Dumancic, Tias Guns, Andrew Cropper
AAAI1
2021 Automated Reasoning and Learning for Automated Payroll Management
abstract
While payroll management is a crucial aspect of any business venture, anticipating the future financial impact of changes to the payroll policy is a challenging task due to the complexity of tax legislature. The goal of this work is to automatically explore potential payroll policies and find the optimal set of policies that satisfies the user's needs. To achieve this goal, we overcome two major challenges. First, we translate the tax legislative knowledge into a formal representation flexible enough to support a variety of scenarios in payroll calculations. Second, the legal knowledge is further compiled into a set of constraints from which a constraint solver can find the optimal policy. Furthermore, payroll computation is performed on the individual basis which might be expensive for companies with a large number of employees. To make the optimisation more efficient, we integrate it with a machine learning model that learns from the previous optimisation runs and speeds up the optimisation engine. The results of this work have been deployed by a social insurance fund.
Sebastijan Dumancic, Wannes Meert, Stijn Goethals, Tim Stuyckens, Jelle Huygen, Koen Denies
AAAI1
2021 avatar - Automated Feature Wrangling for Machine Learning
Gust Verbruggen, Elia Van Wolputte, Sebastijan Dumancic, Luc De Raedt
IDA3
2021 Neural probabilistic logic programming in DeepProbLog
Robin Manhaeve, Sebastijan Dumancic, Angelika Kimmig, Thomas Demeester, Luc De Raedt
Artif. Intell.2
2020 Learning Large Logic Programs By Going Beyond Entailment
abstract
A major challenge in inductive logic programming (ILP) is learning large programs. We argue that a key limitation of existing systems is that they use entailment to guide the hypothesis search. This approach is limited because entailment is a binary decision: a hypothesis either entails an example or does not, and there is no intermediate position. To address this limitation, we go beyond entailment and use 'example-dependent' loss functions to guide the search, where a hypothesis can partially cover an example. We implement our idea in Brute, a new ILP system which uses best-first search, guided by an example-dependent loss function, to incrementally build programs. Our experiments on three diverse program synthesis domains (robot planning, string transformations, and ASCII art), show that Brute can substantially outperform existing ILP systems, both in terms of predictive accuracies and learning times, and can learn programs 20 times larger than state-of-the-art systems.
Andrew Cropper, Sebastijan Dumancic
IJCAI2
2020 Turning 30: New Ideas in Inductive Logic Programming
abstract
Common criticisms of state-of-the-art machine learning include poor generalisation, a lack of interpretability, and a need for large amounts of training data. We survey recent work in inductive logic programming (ILP), a form of machine learning that induces logic programs from data, which has shown promise at addressing these limitations. We focus on new methods for learning recursive programs that generalise from few examples, a shift from using hand-crafted background knowledge to learning background knowledge, and the use of different technologies, notably answer set programming and neural networks. As ILP approaches 30, we also discuss directions for future research.
Andrew Cropper, Sebastijan Dumancic, Stephen H. Muggleton
IJCAI2
2020 From Statistical Relational to Neuro-Symbolic Artificial Intelligence
abstract
Neuro-symbolic and statistical relational artificial intelligence both integrate frameworks for learning with logical reasoning. This survey identifies several parallels across seven different dimensions between these two fields. These cannot only be used to characterize and position neuro-symbolic artificial intelligence approaches but also to identify a number of directions for further research.
Luc De Raedt, Sebastijan Dumancic, Robin Manhaeve, Giuseppe Marra
IJCAI2
2020 Tackling Noise in Active Semi-supervised Clustering
Jonas Soenen, Sebastijan Dumancic, Toon van Craenendonck, Hendrik Blockeel
ECML/PKDD (2)2
2019 Learning Relational Representations with Auto-encoding Logic Programs
abstract
Deep learning methods capable of handling relational data have proliferated over the past years. In contrast to traditional relational learning methods that leverage first-order logic for representing such data, these methods aim at re-representing symbolic relational data in Euclidean space. They offer better scalability, but can only approximate rich relational structures and are less flexible in terms of reasoning. This paper introduces a novel framework for relational representation learning that combines the best of both worlds. This framework, inspired by the auto-encoding principle, uses first-order logic as a data representation language, and the mapping between the the original and latent representation is done by means of logic programs instead of neural networks. We show how learning can be cast as a constraint optimisation problem for which existing solvers can be used. The use of logic as a representation language makes the proposed framework more accurate (as the representation is exact, rather than approximate), more flexible, and more interpretable than deep learning methods. We experimentally show that these latent representations are indeed beneficial in relational learning tasks.
Sebastijan Dumancic, Tias Guns, Wannes Meert, Hendrik Blockeel
IJCAI1
2019 A Comparative Study of Distributional and Symbolic Paradigms for Relational Learning
abstract
Many real-world domains can be expressed as graphs and, more generally, as multi-relational knowledge graphs. Though reasoning and learning with knowledge graphs has traditionally been addressed by symbolic approaches such as Statistical relational learning, recent methods in (deep) representation learning have shown promising results for specialised tasks such as knowledge base completion. These approaches, also known as distributional, abandon the traditional symbolic paradigm by replacing symbols with vectors in Euclidean space. With few exceptions, symbolic and distributional approaches are explored in different communities and little is known about their respective strengths and weaknesses. In this work, we compare distributional and symbolic relational learning approaches on various standard relational classification and knowledge base completion tasks. Furthermore, we analyse the properties of the datasets and relate them to the performance of the methods in the comparison. The results reveal possible indicators that could help in choosing one approach over the other for particular knowledge graphs.
Sebastijan Dumancic, Alberto García-Durán, Mathias Niepert
IJCAI1
2018 COBRASTS: A New Approach to Semi-supervised Clustering of Time Series
Toon van Craenendonck, Wannes Meert, Sebastijan Dumancic, Hendrik Blockeel
DS3
2018 Learning Sequence Encoders for Temporal Knowledge Graph Completion
abstract
Research on link prediction in knowledge graphs has mainly focused on static multirelational data.In this work we consider temporal knowledge graphs where relations between entities may only hold for a time interval or a specific point in time.In line with previous work on static knowledge graphs, we propose to address this problem by learning latent entity and relation type representations.To incorporate temporal information, we utilize recurrent neural networks to learn timeaware representations of relation types which can be used in conjunction with existing latent factorization methods.The proposed approach is shown to be robust to common challenges in real-world KGs: the sparsity and heterogeneity of temporal expressions.Experiments show the benefits of our approach on four temporal KGs.The data sets are available under a permissive BSD-3 license 1 .
Alberto García-Durán, Sebastijan Dumancic, Mathias Niepert
EMNLP2
2018 COBRAS: Interactive Clustering with Pairwise Queries
Toon van Craenendonck, Sebastijan Dumancic, Elia Van Wolputte, Hendrik Blockeel
IDA2
2018 DeepProbLog: Neural Probabilistic Logic Programming
abstract
We introduce DeepProbLog, a probabilistic logic programming language that incorporates deep learning by means of neural predicates. We show how existing inference and learning techniques can be adapted for the new language. Our experiments demonstrate that DeepProbLog supports (i) both symbolic and subsymbolic representations and inference, (ii) program induction, (iii) probabilistic (logic) programming, and (iv) (deep) learning from examples. To the best of our knowledge, this work is the first to propose a framework where general-purpose neural networks and expressive probabilistic-logical modeling and reasoning are integrated in a way that exploits the full expressiveness and strengths of both worlds and can be trained end-to-end based on examples.
Robin Manhaeve, Sebastijan Dumancic, Angelika Kimmig, Thomas Demeester, Luc De Raedt
NeurIPS2
2018 Interactive Time Series Clustering with COBRASTS
Toon van Craenendonck, Wannes Meert, Sebastijan Dumancic, Hendrik Blockeel
ECML/PKDD (3)3
2017 COBRA: A Fast and Simple Method for Active Clustering with Pairwise Constraints
abstract
Clustering is inherently ill-posed: there often exist multiple valid clusterings of a single dataset, and without any additional information a clustering system has no way of knowing which clustering it should produce. This motivates the use of constraints in clustering, as they allow users to communicate their interests to the clustering system. Active constraint-based clustering algorithms select the most useful constraints to query, aiming to produce a good clustering using as few constraints as possible. We propose COBRA, an active method that first over-clusters the data by running K-means with a $K$ that is intended to be too large, and subsequently merges the resulting small clusters into larger ones based on pairwise constraints. In its merging step, COBRA is able to keep the number of pairwise queries low by maximally exploiting constraint transitivity and entailment. We experimentally show that COBRA outperforms the state of the art in terms of clustering quality and runtime, without requiring the number of clusters in advance.
Toon van Craenendonck, Sebastijan Dumancic, Hendrik Blockeel
IJCAI2
2017 Clustering-Based Relational Unsupervised Representation Learning with an Explicit Distributed Representation
abstract
The goal of unsupervised representation learning is to extract a new representation of data, such that solving many different tasks becomes easier. Existing methods typically focus on vectorized data and offer little support for relational data, which additionally describes relationships among instances. In this work we introduce an approach for relational unsupervised representation learning. Viewing a relational dataset as a hypergraph, new features are obtained by clustering vertices and hyperedges. To find a representation suited for many relational learning tasks, a wide range of similarities between relational objects is considered, e.g. feature and structural similarities. We experimentally evaluate the proposed approach and show that models learned on such latent representations perform better, have lower complexity, and outperform the existing approaches on classification tasks.
Sebastijan Dumancic, Hendrik Blockeel
IJCAI1
2017 Demystifying Relational Latent Representations
Sebastijan Dumancic, Hendrik Blockeel
ILP1
2017 An expressive dissimilarity measure for relational clustering using neighbourhood trees
Sebastijan Dumancic, Hendrik Blockeel
Mach. Learn.1
2016 An Efficient and Expressive Similarity Measure for Relational Clustering Using Neighbourhood Trees
abstract
Clustering is an underspecified task: there are no universal criteria for what makes a good clustering. This is especially true for relational data, where similarity can be based on the features of individuals, the relationships between them, or a mix of both. Existing methods for relational clustering have strong and often implicit biases in this respect. In this paper, we introduce a novel similarity measure for relational data. It is the first measure to incorporate a wide variety of types of similarity, including similarity of attributes, similarity of relational context, and proximity in a hypergraph. We experimentally evaluate how using this similarity affects the quality of clustering on very different types of datasets. The experiments demonstrate that (a) using this similarity in standard clustering methods consistently gives good results, whereas other measures work well only on datasets that match their bias; and (b) on most datasets, the novel similarity outperforms even the best among the existing ones.
Sebastijan Dumancic, Hendrik Blockeel
ECAI1