VLDB 2026 Research / reviewers in the wild / expert
Jacopo Urbani
dblp:15/7454
· DBLP profile ↗
42ranked-venue papers
15as first author
6since 2021 · last 2025
0000-0002-0717-3559ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 30 · 9 first-author · 3 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-authorSystems, architecture and hardware · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorTheory of computation · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The ART of Link Prediction with KGEsabstractLink Prediction (LP) in Knowledge Graphs (KGs) is typically framed as ranking candidate entities for a query of the form $(entity, relation,?)$, with models evaluated on their ability to rank the correct entities for each query. At the same time, Knowledge Graph Embedding (KGE) models used for this task produce unnormalised scores, making it unclear how to interpret their belief in the truthfulness of triples across different queries. Together, these two factors create a blind spot: models can achieve perfect rankings while assigning scores that are not comparable across queries, limiting their utility in downstream tasks or even in identifying the most plausible triples overall. Indeed, this issue becomes clear when test triples are ranked globally and evaluated with IR metrics, revealing that models with unnormalized scores often perform poorly due to inconsistent scoring across queries. To address this problem, we propose a new KGE model, called ART, which exploits probabilistic Auto-Regressive modelling and hence is normalised by design. Despite its conceptual simplicity, we show that ART outperforms prior art for discriminative and generative LP as well as other post-hoc calibration techniques. Yannick Brunink, Michael Cochez, Jacopo Urbani |
NeSy | 3 |
| 2023 | Probabilistic Reasoning at Scale: Trigger Graphs to the RescueabstractThe role of uncertainty in data management has become more prominent than ever before, especially because of the growing importance of machine learning-driven applications that produce large uncertain databases. A well-known approach to querying such databases is to blend rule-based reasoning with uncertainty. However, techniques proposed so far struggle with large databases. In this paper, we address this problem by presenting a new technique for probabilistic reasoning that exploits Trigger Graphs (TGs) -- a notion recently introduced for the non-probabilistic setting. The intuition is that TGs can effectively store a probabilistic model by avoiding an explicit materialization of the lineage and by grouping together similar derivations of the same fact. Firstly, we show how TGs can be adapted to support the possible world semantics. Then, we describe techniques for efficiently computing a probabilistic model and formally establish the correctness of our approach. We also present an extensive empirical evaluation using a prototype called LTGs. Our comparison against other leading engines shows that LTGs is not only faster, even against approximate reasoning techniques, but can also reason over probabilistic databases that existing engines cannot scale to. Efthymia Tsamoura, Jacopo Urbani |
Proc. ACM Manag. Data | 3 |
| 2022 | Ensemble-Based Fact Classification with Knowledge Graph Embeddings
Unmesh Joshi, Jacopo Urbani |
ESWC | 2 |
| 2022 | Chasing Streams with Existential Rules
Jacopo Urbani, Markus Krötzsch, Thomas Eiter |
KR | 1 |
| 2021 | Tribrid: Stance Classification with Neural Inconsistency DetectionabstractWe study the problem of performing automatic stance classification on social media with neural architectures such as BERT.Although these architectures deliver impressive results, their level is not yet comparable to the one of humans and they might produce errors that have a significant impact on the downstream task (e.g., fact-checking).To improve the performance, we present a new neural architecture where the input also includes automatically generated negated perspectives over a given claim.The model is jointly learned to make simultaneously multiple predictions, which can be used either to improve the classification of the original perspective or to filter out doubtful predictions.In the first case, we propose a weakly supervised method for combining the predictions into a final one.In the second case, we show that using the confidence scores to remove doubtful predictions allows our method to achieve human-like performance over the retained information, which is still a sizable part of the original input. Jacopo Urbani |
EMNLP (1) | 2 |
| 2021 | Materializing Knowledge Bases via Trigger GraphsabstractThe chase is a well-established family of algorithms used to materialize Knowledge Bases (KBs) for tasks like query answering under dependencies or data cleaning. A general problem of chase algorithms is that they might perform redundant computations. To counter this problem, we introduce the notion of Trigger Graphs (TGs), which guide the execution of the rules avoiding redundant computations. We present the results of an extensive theoretical and empirical study that seeks to answer when and how TGs can be computed and what are the benefits of TGs when applied over real-world KBs. Our results include introducing algorithms that compute (minimal) TGs. We implemented our approach in a new engine, called GLog, and our experiments show that it can be significantly more efficient than the chase enabling us to materialize Knowledge Graphs with 17B facts in less than 40 min using a single machine with commodity hardware. Efthymia Tsamoura, David Carral, Enrico Malizia, Jacopo Urbani |
Proc. VLDB Endow. | 4 |
| 2020 | Checking Chase Termination over Ontologies of Existential Rules with Equality
David Carral, Jacopo Urbani |
AAAI | 2 |
| 2020 | Extracting N-ary Facts from Wikipedia Table ClustersabstractTables in Wikipedia articles contain a wealth of knowledge that would be useful for many applications if it were structured in a more coherent, queryable form. An important problem is that many of such tables contain the same type of knowledge, but have different layouts and/or schemata. Moreover, some tables refer to entities that we can link to Knowledge Bases (KBs), while others do not. Finally, some tables express entity-attribute relations, while others contain more complex n-ary relations. We propose a novel knowledge extraction technique that tackles these problems. Our method first transforms and clusters similar tables into fewer unified ones to overcome the problem of table diversity. Then, the unified tables are linked to the KB so that knowledge about popular entities propagates to the unpopular ones. Finally, our method applies a technique that relies on functional dependencies to judiciously interpret the table and extract n-ary relations. Our experiments over 1.5M Wikipedia tables show that our clustering can group many semantically similar tables. This leads to the extraction of many novel n-ary relations. Benno Kruit, Peter Boncz, Jacopo Urbani |
CIKM | 3 |
| 2020 | Rewrite or Not Rewrite? ML-Based Algorithm Selection for Datalog Query Answering on Knowledge GraphsabstractQuery-driven reasoning techniques with Datalog rules, like Magic Sets (MS), are ideal for implementing query answering on Knowledge Graphs (KGs). For some queries, executing a rewriting procedure like MS is the best choice, but for others a non-rewriting procedure like Query-subquery (QSQ) can be faster. Choosing beforehand which procedure should be used is not trivial and mistakes can be costly. To address this problem, we describe a first-of-its-kind method that builds a Machine Learning (ML) model to predict whether a query should be answered with MS or with QSQ. Experiments on several well-known KGs show that our method can return accurate predictions, and this leads to a significant reduction of the response time of query answering. Unmesh Joshi, Ceriel J. H. Jacobs, Jacopo Urbani |
ECAI | 3 |
| 2020 | Handling Impossible Derivations During Stream Reasoning
Hamid R. Bazoobandi, Henri E. Bal, Frank van Harmelen, Jacopo Urbani |
ESWC | 4 |
| 2020 | Tab2Know: Building a Knowledge Base from Tables in Scientific Papers
Benno Kruit, Jacopo Urbani |
ISWC (1) | 3 |
| 2020 | Searching for Embeddings in a Haystack: Link Prediction on Knowledge Graphs with Subgraph PruningabstractEmbedding-based models of Knowledge Graphs (KGs) can be used to predict the existence of missing links by ranking the entities according to some likelihood scores. An exhaustive computation of all likelihood scores is very expensive if the KG is large. To counter this problem, we propose a technique to reduce the search space by identifying smaller subsets of promising entities. Our technique first creates embeddings of subgraphs using the embeddings from the model. Then, it ranks the subgraphs with some proposed ranking functions and considers only the entities in the top k subgraphs. Our experiments show that our technique is able to reduce the search space significantly while maintaining a good recall. Unmesh Joshi, Jacopo Urbani |
WWW | 2 |
| 2020 | Adaptive Low-level Storage of Very Large Knowledge GraphsabstractThe increasing availability and usage of Knowledge Graphs (KGs) on the Web calls for scalable and general-purpose solutions to store this type of data structures. We propose Trident, a novel storage architecture for very large KGs on centralized systems. Trident uses several interlinked data structures to provide fast access to nodes and edges, with the physical storage changing depending on the topology of the graph to reduce the memory footprint. In contrast to single architectures designed for single tasks, our approach offers an interface with few low-level and general-purpose primitives that can be used to implement tasks like SPARQL query answering, reasoning, or graph analytics. Our experiments show that Trident can handle graphs with 1011 edges using inexpensive hardware, delivering competitive performance on multiple workloads. Jacopo Urbani, Ceriel J. H. Jacobs |
WWW | 1 |
| 2019 | Datalog Reasoning over Compressed RDF Knowledge BasesabstractMaterialisation is often used in RDF systems as a preprocessing step to derive all facts implied by given RDF triples and rules. Although widely used, materialisation considers all possible rule applications and can use a lot of memory for storing the derived facts, which can hinder performance. We present a novel materialisation technique that compresses the RDF triples so that the rules can sometimes be applied to multiple facts at once, and the derived facts can be represented using structure sharing. Our technique can thus require less space, as well as skip certain rule applications. Our experiments show that our technique can be very effective: when the rules are relatively simple, our system is both faster and requires less memory than prominent state-of-the-art RDF systems. Pan Hu 0001, Jacopo Urbani, Boris Motik, Ian Horrocks 0001 |
CIKM | 2 |
| 2019 | Predicting Entity Mentions in Scientific LiteratureabstractPredicting which entities are likely to be mentioned in scientific articles is a task with significant academic and commercial value. For instance, it can lead to monetary savings if the articles are behind paywalls, or be used to recommend articles that are not yet available. Despite extensive prior work on entity prediction in Web documents, the peculiarities of scientific literature make it a unique scenario for this task. In this paper, we present an approach that uses a neural network to predict whether the (unseen) body of an article contains entities defined in domain-specific knowledge bases (KBs). The network uses features from the abstracts and the KB, and it is trained using open-access articles and authors’ prior works. Our experiments on biomedical literature show that our method is able to predict subsets of entities with high accuracy. As far as we know, our method is the first of its kind and is currently used in several commercial settings. Yalung Zheng, Jon Ezeiza, Mehdi Farzanehpour, Jacopo Urbani |
ESWC | 4 |
| 2019 | VLog: A Rule Engine for Knowledge Graphs
David Carral, Irina Dragoste, Larry González, Ceriel J. H. Jacobs, Markus Krötzsch, Jacopo Urbani |
ISWC (2) | 6 |
| 2019 | Extracting Novel Facts from Tables for Knowledge Graph Completion
Benno Kruit, Peter Boncz, Jacopo Urbani |
ISWC (1) | 3 |
| 2019 | ExFaKT: A Framework for Explaining Facts over Knowledge Graphs and TextabstractFact-checking is a crucial task for accurately populating, updating and curating knowledge graphs. Manually validating candidate facts is time-consuming. Prior work on automating this task focuses on estimating truthfulness using numerical scores which are not human-interpretable. Others extract explicit mentions of the candidate fact in the text as an evidence for the candidate fact, which can be hard to directly spot. In our work, we introduce ExFaKT, a framework focused on generating human-comprehensible explanations for candidate facts. ExFaKT uses background knowledge encoded in the form of Horn clauses to rewrite the fact in question into a set of other easier-to-spot facts. The final output of our framework is a set of semantic traces for the candidate fact from both text and knowledge graphs. The experiments demonstrate that our rewritings significantly increase the recall of fact-spotting while preserving high precision. Moreover, we show that the explanations effectively help humans to perform fact-checking and can also be exploited for automating this task. Mohamed H. Gad-Elrab, Daria Stepanova 0001, Jacopo Urbani, Gerhard Weikum |
WSDM | 3 |
| 2019 | Tracy: Tracing Facts over Knowledge Graphs and TextabstractIn order to accurately populate and curate Knowledge Graphs (KGs), it is important to distinguish ?s?p?o? facts that can be traced back to sources from facts that cannot be verified. Manually validating each fact is time-consuming. Prior work on automating this task relied on numerical confidence scores which might not be easily interpreted. To overcome this limitation, we present Tracy, a novel tool that generates human-comprehensible explanations for candidate facts. Our tool relies on background knowledge in the form of rules to rewrite the fact in question into other easier-to-spot facts. These rewritings are then used to reason over the candidate fact creating semantic traces that can aid KG curators. The goal of our demonstration is to illustrate the main features of our system and to show how the semantic traces can be computed over both text and knowledge graphs with a simple and intuitive user interface. Mohamed H. Gad-Elrab, Daria Stepanova 0001, Jacopo Urbani, Gerhard Weikum |
WWW | 3 |
| 2018 | A Deep Dive into Word Sense Disambiguation with LSTMabstractLSTM-based language models have been shown effective in Word Sense Disambiguation (WSD). In particular, the technique proposed by Yuan et al. (2016) returned state-of-the-art performance in several benchmarks, but neither the training data nor the source code was released. This paper presents the results of a reproduction study and analysis of this technique using only openly available datasets (GigaWord, SemCor, OMSTI) and software (TensorFlow). Our study showed that similar results can be obtained with much less data than hinted at by Yuan et al. (2016). Detailed analyses shed light on the strengths and weaknesses of this method. First, adding more unannotated training data is useful, but is subject to diminishing returns. Second, the model can correctly identify both popular and unpopular meanings. Finally, the limited sense coverage in the annotated datasets is a major limitation. All code and trained models are made freely available. Minh Le, Marten Postma, Jacopo Urbani, Piek Vossen |
COLING | 3 |
| 2017 | Enhancing Knowledge Graph Completion By Embedding CorrelationsabstractDespite their large sizes, modern Knowledge Graphs (KGs) are still highly incomplete. Statistical relational learning methods can detect missing links by "embedding" the nodes and relations into latent feature tensors. Unfortunately, these methods are unable to learn good embeddings if the nodes are not well-connected. Our proposal is to learn embeddings for correlations between subgraphs and add a post-prediction phase to counter the lack of training data. This technique, applied on top of methods like TransE or HolE, can significantly increase the predictions on realistic KGs. Soumajit Pal, Jacopo Urbani |
CIKM | 2 |
| 2017 | Expressive Stream Reasoning with Laser
Hamid R. Bazoobandi, Harald Beck, Jacopo Urbani |
ISWC (1) | 3 |
| 2017 | An Empirical Study on How the Distribution of Ontologies Affects Reasoning on the Web
Hamid R. Bazoobandi, Jacopo Urbani, Frank van Harmelen, Henri E. Bal |
ISWC (1) | 2 |
| 2016 | Commonsense in Parts: Mining Part-Whole Relations from the Web and Image TagsabstractCommonsense knowledge about part-whole relations (e.g., screen partOf notebook) is important for interpreting user input in web search and question answering, or for object detection in images. Prior work on knowledge base construction has compiled part-whole assertions, but with substantial limitations: i) semantically different kinds of part-whole relations are conflated into a single generic relation, ii) the arguments of a part-whole assertion are merely words with ambiguous meaning, iii) the assertions lack additional attributes like visibility (e.g., a nose is visible but a kidney is not) and cardinality information (e.g., a bird has two legs while a spider eight), iv) limited coverage of only tens of thousands of assertions. This paper presents a new method for automatically acquiring part-whole commonsense from Web contents and image tags at an unprecedented scale, yielding many millions of assertions, while specifically addressing the four shortcomings of prior work. Our method combines pattern-based information extraction methods with logical reasoning. We carefully distinguish different relations: physicalPartOf, memberOf, substanceOf. We consistently map the arguments of all assertions onto WordNet senses, eliminating the ambiguity of word-level assertions. We identify whether the parts can be visually perceived, and infer cardinalities for the assertions. The resulting commonsense knowledge base has very high quality and high coverage, with an accuracy of 89% determined by extensive sampling, and is publicly available. Niket Tandon, Charles Hariman, Jacopo Urbani, Anna Rohrbach, Marcus Rohrbach, Gerhard Weikum |
AAAI | 3 |
| 2016 | Column-Oriented Datalog Materialization for Large Knowledge GraphsabstractThe evaluation of Datalog rules over large Knowledge Graphs (KGs) is essential for many applications. In this paper, we present a new method of materializing Datalog inferences, which combines a column-based memory layout with novel optimization methods that avoid redundant inferences at runtime. The pro-active caching of certain subqueries further increases efficiency. Our empirical evaluation shows that this approach can often match or even surpass the performance of state-of-the-art systems, especially under restricted resources. Jacopo Urbani, Ceriel J. H. Jacobs, Markus Krötzsch |
AAAI | 1 |
| 2016 | KOGNAC: Efficient Encoding of Large Knowledge Graphs
Jacopo Urbani, Sourav Dutta 0001, Sairam Gurajada, Gerhard Weikum |
IJCAI | 1 |
| 2016 | Exception-Enriched Rule Learning from Knowledge Graphs
Mohamed H. Gad-Elrab, Daria Stepanova 0001, Jacopo Urbani, Gerhard Weikum |
ISWC (1) | 3 |
| 2015 | A Compact In-Memory Dictionary for RDF Data
Hamid R. Bazoobandi, Steven de Rooij, Jacopo Urbani, Annette ten Teije, Frank van Harmelen, Henri E. Bal |
ESWC | 3 |
| 2014 | AJIRA: A Lightweight Distributed Middleware for MapReduce and Stream ProcessingabstractCurrently, MapReduce is the most popular programming model for large-scale data processing and this motivated the research community to improve its efficiency either with new extensions, algorithmic optimizations, or hardware. In this paper we address two main limitations of MapReduce: one relates to the model's limited expressiveness, which prevents the implementation of complex programs that require multiple steps or iterations. The other relates to the efficiency of its most popular implementations (e.g., Hadoop), which provide good resource utilization only for massive volumes of input, operating sub optimally for smaller or rapidly changing input. To address these limitations, we present AJIRA, a new middleware designed for efficient and generic data processing. At a conceptual level, AJIRA replaces the traditional map/reduce primitives by generic operators that can be dynamically allocated, allowing the execution of more complex batch and stream processing jobs. At a more technical level, AJIRA adopts a distributed, multi-threaded architecture that strives at minimizing overhead for non-critical functionality. These characteristics allow AJIRA to be used as a single programming model for both batch and stream processing. To this end, we evaluated its performance against Hadoop, Spark, Esper, and Storm, which are state of the art systems for both batch and stream processing. Our evaluation shows that AJIRA is competitive in a wide range of scenarios both in terms of processing time and scalability, making it an ideal choice where flexibility, extensibility, and the processing of both large and dynamic data with a single programming model are either desirable or even mandatory requirements. Jacopo Urbani, Alessandro Margara, Ceriel J. H. Jacobs, Spyros Voulgaris, Henri E. Bal |
ICDCS | 1 |
| 2014 | Streaming the Web: Reasoning over dynamic data
Alessandro Margara, Jacopo Urbani, Frank van Harmelen, Henri E. Bal |
J. Web Semant. | 2 |
| 2013 | Seven Commandments for Benchmarking Semantic Flow Processing Systems
Thomas Scharrenbach, Jacopo Urbani, Alessandro Margara, Emanuele Della Valle, Abraham Bernstein |
ESWC | 2 |
| 2013 | DynamiTE: Parallel Materialization of Dynamic RDF Data
Jacopo Urbani, Alessandro Margara, Ceriel J. H. Jacobs, Frank van Harmelen, Henri E. Bal |
ISWC (1) | 1 |
| 2013 | Scalable RDF data compression with MapReduceabstractSUMMARY The Semantic Web contains many billions of statements, which are released using the resource description framework (RDF) data model. To better handle these large amounts of data, high performance RDF applications must apply a compression technique. Unfortunately, because of the large input size, even this compression is challenging. In this paper, we propose a set of distributed MapReduce algorithms to efficiently compress and decompress a large amount of RDF data. Our approach uses a dictionary encoding technique that maintains the structure of the data. We highlight the problems of distributed data compression and describe the solutions that we propose. We have implemented a prototype using the Hadoop framework, and evaluate its performance. We show that our approach is able to efficiently compress a large amount of data and scales linearly on both input size and number of nodes. Copyright © 2012 John Wiley & Sons, Ltd. Jacopo Urbani, Jason Maassen, Niels Drost, Frank J. Seinstra, Henri E. Bal |
Concurr. Comput. Pract. Exp. | 1 |
| 2012 | Robust Runtime Optimization and Skew-Resistant Execution of Analytical SPARQL Queries on Pig
Spyros Kotoulas, Jacopo Urbani, Peter Boncz, Peter Mika |
ISWC (1) | 2 |
| 2012 | WebPIE: A Web-scale Parallel Inference Engine using MapReduce
Jacopo Urbani, Spyros Kotoulas, Jason Maassen, Frank van Harmelen, Henri E. Bal |
J. Web Semant. | 1 |
| 2012 | Reply to comment on "WebPIE: A Web-scale parallel inference engine using MapReduce"
Jacopo Urbani, Spyros Kotoulas, Jason Maassen, Frank van Harmelen, Henri E. Bal |
J. Web Semant. | 1 |
| 2012 | Corrigendum to "WebPIE: A Web-scale Parallel Inference Engine using MapReduce" [Web Semant. Sci. Serv. Agents World Wide Web 10 (2012) 59-75]
Jacopo Urbani, Spyros Kotoulas, Jason Maassen, Frank van Harmelen, Henri E. Bal |
J. Web Semant. | 1 |
| 2011 | QueryPIE: Backward Reasoning for OWL Horst over Very Large Knowledge Bases
Jacopo Urbani, Frank van Harmelen, Stefan Schlobach, Henri E. Bal |
ISWC (1) | 1 |
| 2010 | Scalable and Parallel Reasoning in the Semantic Web
Jacopo Urbani |
ESWC (2) | 1 |
| 2010 | OWL Reasoning with WebPIE: Calculating the Closure of 100 Billion Triples
Jacopo Urbani, Spyros Kotoulas, Jason Maassen, Frank van Harmelen, Henri E. Bal |
ESWC (1) | 1 |
| 2010 | Massive Semantic Web data compression with MapReduceabstractThe Semantic Web consists of many billions of statements made of terms that are either URIs or literals. Since these terms usually consist of long sequences of characters, an effective compression technique must be used to reduce the data size and increase the application performance. One of the best known techniques for data compression is dictionary encoding. In this paper we propose a MapReduce algorithm that efficiently compresses and decompresses a large amount of Semantic Web data. We have implemented a prototype using the Hadoop framework and we report an evaluation of the performance. The evaluation shows that our approach is able to efficiently compress a large amount of data and that it scales linearly regarding the input size and number of nodes. Jacopo Urbani, Jason Maassen, Henri E. Bal |
HPDC | 1 |
| 2009 | Scalable Distributed Reasoning Using MapReduce
Jacopo Urbani, Spyros Kotoulas, Eyal Oren, Frank van Harmelen |
ISWC | 1 |