VLDB 2026 Research / reviewers in the wild / expert
Sebastijan Dumancic
dblp:182/1967
· DBLP profile ↗
29ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0003-0915-8034ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | T4G: Trace-based P4 Program Generation
Chenxing Ji, Timo Jugariu, Sebastijan Dumancic, Fernando A. Kuipers |
INFOCOM | 3 |
| 2025 | A Divide, Align and Conquer Strategy For Program SynthesisabstractA major bottleneck in search-based program synthesis, which learns programs from input/output examples, is the synthesis of large programs. As the size of the target program increases, so does the search depth, which leads to an exponentially growing number of candidate programs. Humans mitigate the combinatorial explosion that arises from deep program search: they build complex programs from smaller parts. We introduce a new strategy for program synthesis called Divide, Align & Conquer (DA&C) that exploits the compositionality of real-world domains to guide the synthesis towards useful subprograms. Divide decomposes each example using a segmentation procedure that is synthesized as part of the learning problem. Align matches the components in the decomposed input/output examples to steer the search toward combinations that lead to the synthesis of useful subprograms, and Conquer then solves a standalone synthesis problem on each pair of aligned input/output components. We show how replacing a deep program search with a linear number of much smaller synthesis tasks leads us to efficiently discover useful subprograms that are then combined into a solution program. Our agent outperforms current Inductive Logic Programming (ILP) methods on string transformation tasks even with minimal knowledge priors. Unlike existing methods, the predictive accuracy of our agent monotonically increases for additional examples. It approximates an average time complexity of O(m) in the size m of subprograms for highly structured and, hence, decomposable domains such as strings. Finally, we demonstrate the scalability of our technique on highdimensional abstract visual reasoning tasks from the Abstract Reasoning Corpus (ARC) for which ILP methods were previously infeasible. We are competitive with state-of-the-art agents outside of ILP, despite generating only 0.2% as many candidate programs from a knowledge prior of only 11 generic geometric primitives. Jonas Witt, Sebastijan Dumancic, Tias Guns, Claus-Christian Carbon |
J. Artif. Intell. Res. | 2 |
| 2024 | DeepSaDe: Learning Neural Networks That Guarantee Domain Constraint SatisfactionabstractAs machine learning models, specifically neural networks, are becoming increasingly popular, there are concerns regarding their trustworthiness, specially in safety-critical applications, e.g. actions of an autonomous vehicle must be safe. There are approaches that can train neural networks where such domain requirements are enforced as constraints, but they either cannot guarantee that the constraint will be satisfied by all possible predictions (even on unseen data) or they are limited in the type of constraints that can be enforced. In this paper, we present an approach to train neural networks which can enforce a wide variety of constraints and guarantee that the constraint is satisfied by all possible predictions. The approach builds on earlier work where learning linear models is formulated as a constraint satisfaction problem (CSP). To make this idea applicable to neural networks, two crucial new elements are added: constraint propagation over the network layers, and weight updates based on a mix of gradient descent and CSP solving. Evaluation on various machine learning tasks demonstrates that our approach is flexible enough to enforce a wide variety of domain constraints and is able to guarantee them in neural networks. Kshitij Goyal, Sebastijan Dumancic, Hendrik Blockeel |
AAAI | 2 |
| 2024 | Learning Logic Programs by Discovering Higher-Order Abstractions
Céline Hocquette, Sebastijan Dumancic, Andrew Cropper |
IJCAI | 2 |
| 2024 | From statistical relational to neurosymbolic artificial intelligence: A survey
Giuseppe Marra, Sebastijan Dumancic, Robin Manhaeve, Luc De Raedt |
Artif. Intell. | 2 |
| 2023 | Embedding a Long Short-Term Memory Network in a Constraint Programming Framework for Tomato Greenhouse OptimisationabstractIncreasing global food demand, accompanied by the limited number of expert growers, brings the need for more sustainable and efficient horticulture. The controlled environment of greenhouses enable data collection and precise control. For optimally controlling the greenhouse climate, a grower not only looks at crop production, but rather aims at maximising the profit. However this is a complex, long term optimisation task. In this paper, Constraint Programming (CP) is applied to task of optimal greenhouse economic control, by leveraging a learned greenhouse climate model through a CP embedding. In collaboration with an industrial partner, we demonstrate how to model the greenhouse climate with an LSTM model, embed this LSTM into a CP optimisation framework, and optimise the expected profit of the grower. This data-to-decision pipeline is being integrated into a decision support system for multiple greenhouses in the Netherlands. Dirk van Bokkem, Max van den Hemel, Sebastijan Dumancic, Neil Yorke-Smith |
AAAI | 3 |
| 2022 | SaDe: Learning Models that Provably Satisfy Domain Constraints
Kshitij Goyal, Sebastijan Dumancic, Hendrik Blockeel |
ECML/PKDD (5) | 2 |
| 2022 | Inductive Logic Programming At 30: A New IntroductionabstractInductive logic programming (ILP) is a form of machine learning. The goal of ILP is to induce a hypothesis (a set of logical rules) that generalises training examples. As ILP turns 30, we provide a new introduction to the field. We introduce the necessary logical notation and the main learning settings; describe the building blocks of an ILP system; compare several systems on several dimensions; describe four systems (Aleph, TILDE, ASPAL, and Metagol); highlight key application areas; and, finally, summarise current limitations and directions for future research. Andrew Cropper, Sebastijan Dumancic |
J. Artif. Intell. Res. | 2 |
| 2022 | Inductive logic programming at 30abstractAbstract Inductive logic programming (ILP) is a form of logic-based machine learning. The goal is to induce a hypothesis (a logic program) that generalises given training examples and background knowledge. As ILP turns 30, we review the last decade of research. We focus on (i) new meta-level search methods, (ii) techniques for learning recursive programs, (iii) new approaches for predicate invention, and (iv) the use of different technologies. We conclude by discussing current limitations of ILP and directions for future research. Andrew Cropper, Sebastijan Dumancic, Richard Evans 0001, Stephen H. Muggleton |
Mach. Learn. | 2 |
| 2021 | Knowledge Refactoring for Inductive Program SynthesisabstractHumans constantly restructure knowledge to use it more efficiently. Our goal is to give a machine learning system similar abilities so that it can learn more efficiently. We introduce the knowledge refactoring problem, where the goal is to restructure a learner's knowledge base to reduce its size and to minimise redundancy in it. We focus on inductive logic programming, where the knowledge base is a logic program. We introduce Knorf, a system which solves the refactoring problem using constraint optimisation. A key feature of Knorf is that, rather than simply removing knowledge, it also introduces new knowledge through predicate invention. We evaluate our approach on two domains: building Lego structures and real-world string transformations. Our experiments show that learning from refactored knowledge can improve predictive accuracies fourfold and reduce learning times by half. Sebastijan Dumancic, Tias Guns, Andrew Cropper |
AAAI | 1 |
| 2021 | Automated Reasoning and Learning for Automated Payroll ManagementabstractWhile payroll management is a crucial aspect of any business venture, anticipating the future financial impact of changes to the payroll policy is a challenging task due to the complexity of tax legislature. The goal of this work is to automatically explore potential payroll policies and find the optimal set of policies that satisfies the user's needs. To achieve this goal, we overcome two major challenges. First, we translate the tax legislative knowledge into a formal representation flexible enough to support a variety of scenarios in payroll calculations. Second, the legal knowledge is further compiled into a set of constraints from which a constraint solver can find the optimal policy. Furthermore, payroll computation is performed on the individual basis which might be expensive for companies with a large number of employees. To make the optimisation more efficient, we integrate it with a machine learning model that learns from the previous optimisation runs and speeds up the optimisation engine. The results of this work have been deployed by a social insurance fund. Sebastijan Dumancic, Wannes Meert, Stijn Goethals, Tim Stuyckens, Jelle Huygen, Koen Denies |
AAAI | 1 |
| 2021 | avatar - Automated Feature Wrangling for Machine Learning
Gust Verbruggen, Elia Van Wolputte, Sebastijan Dumancic, Luc De Raedt |
IDA | 3 |
| 2021 | Neural probabilistic logic programming in DeepProbLog
Robin Manhaeve, Sebastijan Dumancic, Angelika Kimmig, Thomas Demeester, Luc De Raedt |
Artif. Intell. | 2 |
| 2020 | Learning Large Logic Programs By Going Beyond EntailmentabstractA major challenge in inductive logic programming (ILP) is learning large programs. We argue that a key limitation of existing systems is that they use entailment to guide the hypothesis search. This approach is limited because entailment is a binary decision: a hypothesis either entails an example or does not, and there is no intermediate position. To address this limitation, we go beyond entailment and use 'example-dependent' loss functions to guide the search, where a hypothesis can partially cover an example. We implement our idea in Brute, a new ILP system which uses best-first search, guided by an example-dependent loss function, to incrementally build programs. Our experiments on three diverse program synthesis domains (robot planning, string transformations, and ASCII art), show that Brute can substantially outperform existing ILP systems, both in terms of predictive accuracies and learning times, and can learn programs 20 times larger than state-of-the-art systems. Andrew Cropper, Sebastijan Dumancic |
IJCAI | 2 |
| 2020 | Turning 30: New Ideas in Inductive Logic ProgrammingabstractCommon criticisms of state-of-the-art machine learning include poor generalisation, a lack of interpretability, and a need for large amounts of training data. We survey recent work in inductive logic programming (ILP), a form of machine learning that induces logic programs from data, which has shown promise at addressing these limitations. We focus on new methods for learning recursive programs that generalise from few examples, a shift from using hand-crafted background knowledge to learning background knowledge, and the use of different technologies, notably answer set programming and neural networks. As ILP approaches 30, we also discuss directions for future research. Andrew Cropper, Sebastijan Dumancic, Stephen H. Muggleton |
IJCAI | 2 |
| 2020 | From Statistical Relational to Neuro-Symbolic Artificial IntelligenceabstractNeuro-symbolic and statistical relational artificial intelligence both integrate frameworks for learning with logical reasoning. This survey identifies several parallels across seven different dimensions between these two fields. These cannot only be used to characterize and position neuro-symbolic artificial intelligence approaches but also to identify a number of directions for further research. Luc De Raedt, Sebastijan Dumancic, Robin Manhaeve, Giuseppe Marra |
IJCAI | 2 |
| 2020 | Tackling Noise in Active Semi-supervised Clustering
Jonas Soenen, Sebastijan Dumancic, Toon van Craenendonck, Hendrik Blockeel |
ECML/PKDD (2) | 2 |
| 2019 | Learning Relational Representations with Auto-encoding Logic ProgramsabstractDeep learning methods capable of handling relational data have proliferated over the past years. In contrast to traditional relational learning methods that leverage first-order logic for representing such data, these methods aim at re-representing symbolic relational data in Euclidean space. They offer better scalability, but can only approximate rich relational structures and are less flexible in terms of reasoning. This paper introduces a novel framework for relational representation learning that combines the best of both worlds. This framework, inspired by the auto-encoding principle, uses first-order logic as a data representation language, and the mapping between the the original and latent representation is done by means of logic programs instead of neural networks. We show how learning can be cast as a constraint optimisation problem for which existing solvers can be used. The use of logic as a representation language makes the proposed framework more accurate (as the representation is exact, rather than approximate), more flexible, and more interpretable than deep learning methods. We experimentally show that these latent representations are indeed beneficial in relational learning tasks. Sebastijan Dumancic, Tias Guns, Wannes Meert, Hendrik Blockeel |
IJCAI | 1 |
| 2019 | A Comparative Study of Distributional and Symbolic Paradigms for Relational LearningabstractMany real-world domains can be expressed as graphs and, more generally, as multi-relational knowledge graphs. Though reasoning and learning with knowledge graphs has traditionally been addressed by symbolic approaches such as Statistical relational learning, recent methods in (deep) representation learning have shown promising results for specialised tasks such as knowledge base completion. These approaches, also known as distributional, abandon the traditional symbolic paradigm by replacing symbols with vectors in Euclidean space. With few exceptions, symbolic and distributional approaches are explored in different communities and little is known about their respective strengths and weaknesses. In this work, we compare distributional and symbolic relational learning approaches on various standard relational classification and knowledge base completion tasks. Furthermore, we analyse the properties of the datasets and relate them to the performance of the methods in the comparison. The results reveal possible indicators that could help in choosing one approach over the other for particular knowledge graphs. Sebastijan Dumancic, Alberto García-Durán, Mathias Niepert |
IJCAI | 1 |
| 2018 | COBRASTS: A New Approach to Semi-supervised Clustering of Time Series
Toon van Craenendonck, Wannes Meert, Sebastijan Dumancic, Hendrik Blockeel |
DS | 3 |
| 2018 | Learning Sequence Encoders for Temporal Knowledge Graph CompletionabstractResearch on link prediction in knowledge graphs has mainly focused on static multirelational data.In this work we consider temporal knowledge graphs where relations between entities may only hold for a time interval or a specific point in time.In line with previous work on static knowledge graphs, we propose to address this problem by learning latent entity and relation type representations.To incorporate temporal information, we utilize recurrent neural networks to learn timeaware representations of relation types which can be used in conjunction with existing latent factorization methods.The proposed approach is shown to be robust to common challenges in real-world KGs: the sparsity and heterogeneity of temporal expressions.Experiments show the benefits of our approach on four temporal KGs.The data sets are available under a permissive BSD-3 license 1 . Alberto García-Durán, Sebastijan Dumancic, Mathias Niepert |
EMNLP | 2 |
| 2018 | COBRAS: Interactive Clustering with Pairwise Queries
Toon van Craenendonck, Sebastijan Dumancic, Elia Van Wolputte, Hendrik Blockeel |
IDA | 2 |
| 2018 | DeepProbLog: Neural Probabilistic Logic ProgrammingabstractWe introduce DeepProbLog, a probabilistic logic programming language that incorporates deep learning by means of neural predicates. We show how existing inference and learning techniques can be adapted for the new language. Our experiments demonstrate that DeepProbLog supports (i) both symbolic and subsymbolic representations and inference, (ii) program induction, (iii) probabilistic (logic) programming, and (iv) (deep) learning from examples. To the best of our knowledge, this work is the first to propose a framework where general-purpose neural networks and expressive probabilistic-logical modeling and reasoning are integrated in a way that exploits the full expressiveness and strengths of both worlds and can be trained end-to-end based on examples. Robin Manhaeve, Sebastijan Dumancic, Angelika Kimmig, Thomas Demeester, Luc De Raedt |
NeurIPS | 2 |
| 2018 | Interactive Time Series Clustering with COBRASTS
Toon van Craenendonck, Wannes Meert, Sebastijan Dumancic, Hendrik Blockeel |
ECML/PKDD (3) | 3 |
| 2017 | COBRA: A Fast and Simple Method for Active Clustering with Pairwise ConstraintsabstractClustering is inherently ill-posed: there often exist multiple valid clusterings of a single dataset, and without any additional information a clustering system has no way of knowing which clustering it should produce. This motivates the use of constraints in clustering, as they allow users to communicate their interests to the clustering system. Active constraint-based clustering algorithms select the most useful constraints to query, aiming to produce a good clustering using as few constraints as possible. We propose COBRA, an active method that first over-clusters the data by running K-means with a $K$ that is intended to be too large, and subsequently merges the resulting small clusters into larger ones based on pairwise constraints. In its merging step, COBRA is able to keep the number of pairwise queries low by maximally exploiting constraint transitivity and entailment. We experimentally show that COBRA outperforms the state of the art in terms of clustering quality and runtime, without requiring the number of clusters in advance. Toon van Craenendonck, Sebastijan Dumancic, Hendrik Blockeel |
IJCAI | 2 |
| 2017 | Clustering-Based Relational Unsupervised Representation Learning with an Explicit Distributed RepresentationabstractThe goal of unsupervised representation learning is to extract a new representation of data, such that solving many different tasks becomes easier. Existing methods typically focus on vectorized data and offer little support for relational data, which additionally describes relationships among instances. In this work we introduce an approach for relational unsupervised representation learning. Viewing a relational dataset as a hypergraph, new features are obtained by clustering vertices and hyperedges. To find a representation suited for many relational learning tasks, a wide range of similarities between relational objects is considered, e.g. feature and structural similarities. We experimentally evaluate the proposed approach and show that models learned on such latent representations perform better, have lower complexity, and outperform the existing approaches on classification tasks. Sebastijan Dumancic, Hendrik Blockeel |
IJCAI | 1 |
| 2017 | Demystifying Relational Latent Representations
Sebastijan Dumancic, Hendrik Blockeel |
ILP | 1 |
| 2017 | An expressive dissimilarity measure for relational clustering using neighbourhood trees
Sebastijan Dumancic, Hendrik Blockeel |
Mach. Learn. | 1 |
| 2016 | An Efficient and Expressive Similarity Measure for Relational Clustering Using Neighbourhood TreesabstractClustering is an underspecified task: there are no universal criteria for what makes a good clustering. This is especially true for relational data, where similarity can be based on the features of individuals, the relationships between them, or a mix of both. Existing methods for relational clustering have strong and often implicit biases in this respect. In this paper, we introduce a novel similarity measure for relational data. It is the first measure to incorporate a wide variety of types of similarity, including similarity of attributes, similarity of relational context, and proximity in a hypergraph. We experimentally evaluate how using this similarity affects the quality of clustering on very different types of datasets. The experiments demonstrate that (a) using this similarity in standard clustering methods consistently gives good results, whereas other measures work well only on datasets that match their bias; and (b) on most datasets, the novel similarity outperforms even the best among the existing ones. Sebastijan Dumancic, Hendrik Blockeel |
ECAI | 1 |