Fabian M. Suchanek

dblp:13/1396 · DBLP profile ↗
← Back
77ranked-venue papers
17as first author
19since 2021 · last 2026
0000-0001-7189-2796ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 59 · 16 first-author · 6 since 2021Artificial intelligence and machine learning · 29 · 5 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Theory of computation · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Qiana: A First-Order Formalism to Quantify over Contexts and Formulas with Temporality
abstract
We introduce Qiana, a logic framework for reasoning on formulas that are true only in specific contexts. In Qiana, it is possible to quantify over both formulas and contexts to express, e.g., that “everyone knows everything Alice says”. Qiana also permits paraconsistent logics within contexts, so that contexts can contain contradictions. Furthermore, Qiana is based on first-order logic, and is finitely axiomatizable, so that Qiana theories are compatible with pre-existing first-order logic theorem provers. We show how Qiana can be used to represent temporality, event calculus, and modal logic. We also discuss different design alternatives of Qiana.
Simon Coumes, Pierre-Henri Paris, François Schwarzentruber, Fabian M. Suchanek
J. Artif. Intell. Res.4
2025 FLORA: Unsupervised Knowledge Graph Alignment by Fuzzy Logic
Yiwen Peng, Thomas Bonald, Fabian M. Suchanek
ISWC (1)3
2024 The Factuality of Large Language Models in the Legal Domain
abstract
This paper investigates the factuality of large language models (LLMs) as knowledge bases in the legal domain, in a realistic usage scenario: we allow for acceptable variations in the answer, and let the model abstain from answering when uncertain. First, we design a dataset of diverse factual questions about case law and legislation. We then use the dataset to evaluate several LLMs under different evaluation methods, including exact, alias, and fuzzy matching. Our results show that the performance improves significantly under the alias and fuzzy matching methods. Further, we explore the impact of abstaining and in-context examples, finding that both strategies enhance precision. Finally, we demonstrate that additional pre-training on legal documents, as seen with SaulLM, further improves factual precision from 63% to 81%.
Rajaa El Hamdani, Thomas Bonald, Fragkiskos D. Malliaros, Nils Holzenberger, Fabian M. Suchanek
CIKM5
2024 PyClause - Simple and Efficient Rule Handling for Knowledge Graphs
Patrick Betz, Luis Galárraga, Simon Ott, Christian Meilicke, Fabian M. Suchanek, Heiner Stuckenschmidt
IJCAI5
2024 Qiana: A First-Order Formalism to Quantify over Contexts and Formulas
abstract
We introduce Qiana, a logic framework for reasoning on formulas that are true only in specific contexts. In Qiana, it is possible to quantify over both formulas and contexts to express, e.g., that ``everyone knows everything Alice says''. Qiana also permits paraconsistent logics within contexts, so that contexts can contain contradictions. Furthermore, Qiana is based on first-order logic, and is finitely axiomatizable, so that Qiana theories are compatible with pre-existing first-order logic theorem provers.
Simon Coumes, Pierre-Henri Paris, François Schwarzentruber, Fabian M. Suchanek
KR4
2024 MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and Classification
abstract
Chadi Helwe, Tom Calamai, Pierre-Henri Paris, Chloé Clavel, Fabian Suchanek. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Chadi Helwe, Tom Calamai, Pierre-Henri Paris, Chloé Clavel, Fabian M. Suchanek
NAACL-HLT5
2024 A Survey of Meaning Representations - From Theory to Practical Utility
abstract
Zacchary Sadeddine, Juri Opitz, Fabian Suchanek. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zacchary Sadeddine, Juri Opitz, Fabian M. Suchanek
NAACL-HLT3
2024 YAGO 4.5: A Large and Clean Knowledge Base with a Rich Taxonomy
abstract
International audience
Fabian M. Suchanek, Mehwish Alam, Thomas Bonald, Lihu Chen, Pierre-Henri Paris, Jules Soria
SIGIR1
2024 Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story Evaluation
abstract
Abstract Storytelling is an integral part of human experience and plays a crucial role in social interactions. Thus, Automatic Story Evaluation (ASE) and Generation (ASG) could benefit society in multiple ways, but they are challenging tasks which require high-level human abilities such as creativity, reasoning, and deep understanding. Meanwhile, Large Language Models (LLMs) now achieve state-of-the-art performance on many NLP tasks. In this paper, we study whether LLMs can be used as substitutes for human annotators for ASE. We perform an extensive analysis of the correlations between LLM ratings, other automatic measures, and human annotations, and we explore the influence of prompting on the results and the explainability of LLM behaviour. Most notably, we find that LLMs outperform current automatic measures for system-level evaluation but still struggle at providing satisfactory explanations for their answers.
Cyril Chhun, Fabian M. Suchanek, Chloé Clavel
Trans. Assoc. Comput. Linguistics2
2023 GLADIS: A General and Large Acronym Disambiguation Benchmark
abstract
Acronym Disambiguation (AD) is crucial for natural language understanding on various sources, including biomedical reports, scientific papers, and search engine queries.However, existing acronym disambiguation benchmarks and tools are limited to specific domains, and the size of prior benchmarks is rather small.To accelerate the research on acronym disambiguation, we construct a new benchmark named GLADIS with three components: (1) a much larger acronym dictionary with 1.5M acronyms and 6.4M long forms;(2) a pre-training corpus with 160 million sentences; (3) three datasets that cover the general, scientific, and biomedical domains.We then pre-train a language model, AcroBERT, on our constructed corpus for general acronym disambiguation, and show the challenges and values of our new benchmark.
Lihu Chen, Gaël Varoquaux, Fabian M. Suchanek
EACL3
2023 Knowledge Bases and Language Models: Complementing Forces
Fabian M. Suchanek, Anh Tuan Luu
RuleML+RR1
2022 Imputing Out-of-Vocabulary Embeddings with LOVE Makes LanguageModels Robust with Little Cost
abstract
State-of-the-art NLP systems represent inputs with word embeddings, but these are brittle when faced with Out-of-Vocabulary (OOV) words.To address this issue, we follow the principle of mimick-like models to generate vectors for unseen words, by learning the behavior of pre-trained embeddings using only the surface form of words.We present a simple contrastive learning framework, LOVE, which extends the word representation of an existing pre-trained language model (such as BERT), and makes it robust to OOV with few additional parameters.Extensive evaluations demonstrate that our lightweight model achieves similar or even better performances than prior competitors, both on original datasets and on corrupted variants.Moreover, it can be used in a plug-and-play fashion with FastText and BERT, where it significantly improves their robustness.
Lihu Chen, Gaël Varoquaux, Fabian M. Suchanek
ACL (1)3
2022 Of Human Criteria and Automatic Metrics: A Benchmark of the Evaluation of Story Generation
abstract
Research on Automatic Story Generation (ASG) relies heavily on human and automatic evaluation. However, there is no consensus on which human evaluation criteria to use, and no analysis of how well automatic criteria correlate with them. In this paper, we propose to re-evaluate ASG evaluation. We introduce a set of 6 orthogonal and comprehensive human criteria, carefully motivated by the social sciences literature. We also present HANNA, an annotated dataset of 1,056 stories produced by 10 different ASG systems. HANNA allows us to quantitatively evaluate the correlations of 72 automatic metrics with human criteria. Our analysis highlights the weaknesses of current metrics for ASG and allows us to formulate practical recommendations for ASG evaluation.
Cyril Chhun, Pierre Colombo, Fabian M. Suchanek, Chloé Clavel
COLING3
2022 Using a Knowledge Base to Automatically Annotate Speech Corpora and to Identify Sociolinguistic Variation
abstract
Speech characteristics vary from speaker to speaker. While some variation phenomena are due to the overall communication setting, others are due to diastratic factors such as gender, provenance, age, and social background. The analysis of these factors, although relevant for both linguistic and speech technology communities, is hampered by the need to annotate existing corpora or to recruit, categorise, and record volunteers as a function of targeted profiles. This paper presents a methodology that uses a knowledge base to provide speaker-specific information. This can facilitate the enrichment of existing corpora with new annotations extracted from the knowledge base. The method also helps the large scale analysis by automatically extracting instances of speech variation to correlate with diastratic features. We apply our method to an over 120-hour corpus of broadcast speech in French and investigate variation patterns linked to reduction phenomena and/or specific to connected speech such as disfluencies. We find significant differences in speech rate, the use of filler words, and the rate of non-canonical realisations of frequent segments as a function of different professional categories and age groups.
Yaru Wu, Fabian M. Suchanek, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker
LREC2
2022 An Experimental Study of State-of-the-Art Entity Alignment Approaches
abstract
Entity alignment (EA) finds equivalent entities that are located in different knowledge graphs (KGs), which is an essential step to enhance the quality of KGs, and hence of significance to downstream applications (e.g., question answering and recommendation). Recent years have witnessed a rapid increase of EA approaches, yet the relative performance of them remains unclear, partly due to the incomplete empirical evaluations, as well as the fact that comparisons were carried out under different settings (i.e., datasets, information used as input, etc.). In this paper, we fill in the gap by conducting a comprehensive evaluation and detailed analysis of state-of-the-art EA approaches. We first propose a general EA framework that encompasses all the current methods, and then group existing methods into three major categories. Next, we judiciously evaluate these solutions on a wide range of use cases, based on their effectiveness, efficiency and robustness. Finally, we construct a new EA dataset to mirror the real-life challenges of alignment, which were largely overlooked by existing literature. This study strives to provide a clear picture of the strengths and weaknesses of current EA approaches, so as to inspire quality follow-up research.
Xiang Zhao 0002, Weixin Zeng, Jiuyang Tang, Wei Wang 0011, Fabian M. Suchanek
IEEE Trans. Knowl. Data Eng.5
2021 A Lightweight Neural Model for Biomedical Entity Linking
abstract
Biomedical entity linking aims to map biomedical mentions, such as diseases and drugs, to standard entities in a given knowledge base. The specific challenge in this context is that the same biomedical entity can have a wide range of names, including synonyms, morphological variations, and names with different word orderings. Recently, BERT-based methods have advanced the state-of-the-art by allowing for rich representations of word sequences. However, they often have hundreds of millions of parameters and require heavy computing resources, which limits their applications in resource-limited scenarios. Here, we propose a lightweight neural method for biomedical entity linking, which needs just a fraction of the parameters of a BERT model and much less computing resources. Our method uses a simple alignment layer with attention mechanisms to capture the variations between mention and entity names. Yet, we show that our model is competitive with previous work on standard evaluation benchmarks.
Lihu Chen, Gaël Varoquaux, Fabian M. Suchanek
AAAI3
2021 Neural Knowledge Base Repairs
Thomas Pellissier Tanon, Fabian M. Suchanek
ESWC2
2021 Confident Interpretations of Black Box Classifiers
abstract
Deep Learning models provide state of the art classification results, but are not human-interpretable. We propose a novel method to interpret the classification results of a black box model a posteriori. We emulate the complex classifier by surrogate decision trees. Each tree mimics the behavior of the complex classifier by overestimating one of the classes. This yields a global, interpretable approximation of the black box classifier. Our method provides interpretations that are at the same time general (applying to many data points), confident (generalizing well to other data points), faithful to the original model (making the same predictions), and simple (easy to understand). Our experiments show that our method beats competing methods in these desiderata, and our user study shows that users prefer this type of interpretations over others.
Nedeljko Radulovic, Albert Bifet, Fabian M. Suchanek
IJCNN3
2021 On the Limits of Machine Knowledge: Completeness, Recall and Negation in Web-scale Knowledge Bases
abstract
General-purpose knowledge bases (KBs) are an important component of several data-driven applications. Pragmatically constructed from available web sources, these KBs are far from complete, which poses a set of challenges in curation as well as consumption. In this tutorial we discuss how completeness, recall and negation in DBs and KBs can be represented, extracted, and inferred. We proceed in 5 parts: (i) We introduce the logical foundations of knowledge representation and querying under partial closed-world semantics. (ii) We show how information about recall can be identified in KBs and in text, and (iii) how it can be estimated via statistical patterns. (iv) We show how interesting negative statements can be identified, and (v) how recall can be targeted in a comparative notion.
Simon Razniewski, Hiba Arnaout, Shrestha Ghosh, Fabian M. Suchanek
Proc. VLDB Endow.4
2020 Computing and Illustrating Query Rewritings on Path Views with Binding Patterns
abstract
In this system demonstration, we study views with binding patterns, which are a formalization of REST Web services. Such views are database queries that can be evaluated using the service, but only if values for the input variables are provided. We investigate how to use such views to answer a complex user query, by rewriting it as an execution plan, i.e., an orchestration of calls to the views. In general, it is undecidable to determine whether a given user query can be answered with the available views. In this demo, we illustrate a particular scenario studied in our earlier work [11], where the problem is not only decidable but has a particularly intuitive graphical solution. Our demo allows users to play with views defined by real Web services, and to animate the construction of execution plans visually.
Julien Romero, Nicoleta Preda, Antoine Amarilli, Fabian M. Suchanek
CIKM4
2020 Fast and Exact Rule Mining with AMIE 3
Jonathan Lajus, Luis Galárraga, Fabian M. Suchanek
ESWC3
2020 Equivalent Rewritings on Path Views with Binding Patterns
Julien Romero, Nicoleta Preda, Antoine Amarilli, Fabian M. Suchanek
ESWC4
2020 YAGO 4: A Reason-able Knowledge Base
abstract
YAGO is one of the large knowledge bases in the Linked Open Data cloud. In this resource paper, we present its latest version, YAGO 4, which reconciles the rigorous typing and constraints of schema.org with the rich instance data of Wikidata. The resulting resource contains 2 billion type-consistent triples for 64 Million entities, and has a consistent ontology that allows semantic reasoning with OWL 2 description logics.
Thomas Pellissier Tanon, Gerhard Weikum, Fabian M. Suchanek
ESWC3
2019 Anytime Large-Scale Analytics of Linked Open Data
Arnaud Soulet, Fabian M. Suchanek
ISWC (1)2
2019 Learning How to Correct a Knowledge Base from the Edit History
abstract
The curation of a knowledge base is a crucial but costly task. In this work, we propose to take advantage of the edit history of the knowledge base in order to learn how to correct constraint violations. Our method is based on rule mining, and uses the edits that solved some violations in the past to infer how to solve similar violations in the present. The experimental evaluation of our method on Wikidata shows significant improvements over baselines.
Thomas Pellissier Tanon, Camille Bourgaux, Fabian M. Suchanek
WWW3
2018 Text to Brain: Predicting the Spatial Distribution of Neuroimaging Observations from Text Reports
Jérôme Dockès, Demian Wassermann, Russell A. Poldrack, Fabian M. Suchanek, Bertrand Thirion, Gaël Varoquaux
MICCAI (3)4
2018 Adding Missing Words to Regular Expressions
Thomas Rebele, Katerina Tzompanaki, Fabian M. Suchanek
PAKDD (2)3
2018 Bash Datalog: Answering Datalog Queries with Unix Shell Commands
Thomas Rebele, Thomas Pellissier Tanon, Fabian M. Suchanek
ISWC (1)3
2018 Representativeness of Knowledge Bases with the Generalized Benford's Law
Arnaud Soulet, Arnaud Giacometti, Béatrice Bouchou-Markhoff, Fabian M. Suchanek
ISWC (1)4
2018 Are All People Married?: Determining Obligatory Attributes in Knowledge Bases
abstract
An attribute is obligatory for a class in a Knowledge Base (KB), if all instances of the class have the attribute in the real world. For example, has­Birth­Date is an obligatory attribute for the class Person, while has ­ Spouse is not. In this paper, we propose a new way to model incompleteness in KBs. From this model, we derive a method to automatically determine obligatory attributes -- using only the data from the KB. Our algorithm can detect such attributes with a precision of up to 90%.
Jonathan Lajus, Fabian M. Suchanek
WWW2
2017 VICKEY: Mining Conditional Keys on Knowledge Bases
Danai Symeonidou, Luis Galárraga, Nathalie Pernelle, Fatiha Saïs, Fabian M. Suchanek
ISWC (1)5
2017 Predicting Completeness in Knowledge Bases
abstract
Knowledge bases such as Wikidata, DBpedia, or YAGO contain millions of entities and facts. In some knowledge bases, the correctness of these facts has been evaluated. However, much less is known about their completeness, i.e., the proportion of real facts that the knowledge bases cover. In this work, we investigate different signals to identify the areas where a knowledge base is complete. We show that we can combine these signals in a rule mining approach, which allows us to predict where facts may be missing. We also show that completeness predictions can help other applications such as fact prediction.
Luis Galárraga, Simon Razniewski, Antoine Amarilli, Fabian M. Suchanek
WSDM4
2016 Thymeflow, A Personal Knowledge Base with Spatio-temporal Data
abstract
The typical Internet user has data spread over several devices and across several online systems. We demonstrate an open-source system for integrating user's data from different sources into a single Knowledge Base. Our system integrates data of different kinds into a coherent whole, starting with email messages, calendar, contacts, and location history. It is able to detect event periods in the user's location data and align them with calendar events. We will demonstrate how to query the system within and across different dimensions, and perform analytics over emails, events, and locations.
David Montoya, Thomas Pellissier Tanon, Serge Abiteboul, Fabian M. Suchanek
CIKM4
2016 Open Digital Forms
Hiep Le, Thomas Rebele, Fabian M. Suchanek
TPDL3
2016 YAGO: A Multilingual Knowledge Base from Wikipedia, Wordnet, and Geonames
abstract
YAGO is a large knowledge base that is built automatically from Wikipedia, WordNet and GeoNames. The project combines information from Wikipedias in 10 different languages into a coherent whole, thus giving the knowledge a multilingual dimension. It also attaches spatial and temporal information to many facts, and thus allows the user to query the data over space and time. YAGO focuses on extraction quality and achieves a manually evaluated precision of 95 %. In this paper, we explain how YAGO is built from its sources, how its quality is evaluated, how a user can access it, and how other projects utilize it.
Thomas Rebele, Fabian M. Suchanek, Johannes Hoffart, Asia J. Biega, Erdal Kuzey, Gerhard Weikum
ISWC (2)2
2016 Can You Imagine... A Language for Combinatorial Creativity?
Fabian M. Suchanek, Colette Menard, Meghyn Bienvenu, Cyril Chapellier
ISWC (1)1
2015 YAGO3: A Knowledge Base from Multilingual Wikipedias
Farzaneh Mahdisoltani, Asia J. Biega, Fabian M. Suchanek
CIDR3
2015 The elephant in the room: getting value from Big Data
abstract
International audience
Serge Abiteboul, Xin Dong 0001, Oren Etzioni, Divesh Srivastava, Gerhard Weikum, Julia Stoyanovich, Fabian M. Suchanek
WebDB7
2015 IBEX: Harvesting Entities from the Web Using Unique Identifiers
abstract
In this paper we study the prevalence of unique entity identifiers on the Web. These are, e.g., ISBNs (for books), GTINs (for commercial products), DOIs (for documents), email addresses, and others. We show how these identifiers can be harvested systematically from Web pages, and how they can be associated with humanreadable names for the entities at large scale.
Aliaksandr Talaika, Asia J. Biega, Antoine Amarilli, Fabian M. Suchanek
WebDB4
2015 Fast rule mining in ontological knowledge bases with AMIE+
Luis Galárraga, Christina Teflioudi, Katja Hose, Fabian M. Suchanek
VLDB J.4
2015 Editorial
Roberto Navigli, Fabian M. Suchanek
J. Web Semant.2
2014 Recent Topics of Research around the YAGO Knowledge Base
Antoine Amarilli, Luis Galárraga, Nicoleta Preda, Fabian M. Suchanek
APWeb4
2014 Canonicalizing Open Knowledge Bases
abstract
Open information extraction approaches have led to the creation of large knowledge bases from the Web. The problem with such methods is that their entities and relations are not canonicalized, leading to redundant and ambiguous facts. For example, they may store {Barack Obama, was born, Honolulu and {Obama, place of birth, Honolulu}. In this paper, we present an approach based on machine learning methods that can canonicalize such Open IE triples, by clustering synonymous names and phrases.
Luis Galárraga, Geremy Heitz, Kevin Murphy 0002, Fabian M. Suchanek
CIKM4
2014 WebChild: harvesting and organizing commonsense knowledge from the web
abstract
This paper presents a method for automatically constructing a large commonsense knowledge base, called WebChild, from Web contents. WebChild contains triples that connect nouns with adjectives via fine-grained relations like hasShape, hasTaste, evokesEmotion, etc. The arguments of these assertions, nouns and adjectives, are disambiguated by mapping them onto their proper WordNet senses. Our method is based on semi-supervised Label Propagation over graphs of noisy candidate assertions. We automatically derive seeds from WordNet and by pattern matching from Web text collections. The Label Propagation algorithm provides us with domain sets and range sets for 19 different relations, and with confidence-ranked assertions between WordNet senses. Large-scale experiments demonstrate the high accuracy (more than 80 percent) and coverage (more than four million fine grained disambiguated assertions) of WebChild.
Niket Tandon, Gerard de Melo, Fabian M. Suchanek, Gerhard Weikum
WSDM3
2014 Semantic Culturomics (vision paper)
abstract
Newspapers are testimonials of history. The same is increasingly true of social media such as online forums, online communities, and blogs. By looking at the sequence of articles over time, one can discover the birth and the development of trends that marked society and history -- a field known as "Culturomics". But Culturomics has so far been limited to statistics on keywords. In this vision paper, we argue that the advent of large knowledge bases (such as YAGO [37], NELL [5], DBpedia [3], and Freebase) will revolutionize the field. If their knowledge is combined with the news articles, it can breathe life into what is otherwise just a sequence of words for a machine. This will allow discovering trends in history and culture, explaining them through explicit logical rules, and making predictions about the events of the future. We predict that this could open up a new field of research, "Semantic Culturomics", in which no longer human text helps machines build up knowledge bases, but knowledge bases help humans understand their society.
Fabian M. Suchanek, Nicoleta Preda
Proc. VLDB Endow.1
2014 Knowledge Bases in the Age of Big Data Analytics
abstract
This tutorial gives an overview on state-of-the-art methods for the automatic construction of large knowledge bases and harnessing them for data and text analytics. It covers both big-data methods for building knowledge bases and knowledge bases being assets for big-data applications. The tutorial also points out challenges and research opportunities.
Fabian M. Suchanek, Gerhard Weikum
Proc. VLDB Endow.1
2013 PIKM 2013: the 6th ACM workshop for ph.d. students in information and knowledge management
abstract
The PIKM workshop gives Ph.D. students an opportunity to present their dissertation proposals at a global stage. Similarly to the CIKM, the PIKM workshop covers a wide range of topics in the areas of databases, information retrieval and knowledge management. Interdisciplinary work across these tracks is particularly encouraged.
Fabian M. Suchanek, Anisoara Nica
CIKM1
2013 AKBC 2013: third workshop on automated knowledge base construction
abstract
The AKBC 2013 workshop aims to be a venue of excellence and vision in the area of knowledge base construction. This year's workshop will feature keynotes by ten leading researchers in the field, including from Google, Microsoft, Stanford, and CMU. The submissions focus on visionary ideas instead of on experimental evaluation. Nineteen accepted papers will be presented as posters, with nine exceptional papers also highlighted as spotlight talks. Thereby, the workshop aims provides a vivid forum of discussion about the field of automated knowledge base construction.
Fabian M. Suchanek, Sebastian Riedel 0001, Sameer Singh 0001, Partha P. Talukdar
CIKM1
2013 SUSIE: Search using services and information extraction
abstract
The API of a Web service restricts the types of queries that the service can answer. For example, a Web service might provide a method that returns the songs of a given singer, but it might not provide a method that returns the singers of a given song. If the user asks for the singer of some specific song, then the Web service cannot be called - even though the underlying database might have the desired piece of information. This asymmetry is particularly problematic if the service is used in a Web service orchestration system. In this paper, we propose to use on-the-fly information extraction to collect values that can be used as parameter bindings for the Web service. We show how this idea can be integrated into a Web service orchestration system. Our approach is fully implemented in a prototype called SUSIE. We present experiments with real-life data and services to demonstrate the practical viability and good performance of our approach.
Nicoleta Preda, Fabian M. Suchanek, Wenjun Yuan, Gerhard Weikum
ICDE2
2013 Knowledge harvesting from text and Web sources
abstract
The proliferation of knowledge-sharing communities such as Wikipedia and the progress in scalable information extraction from Web and text sources has enabled the automatic construction of very large knowledge bases. Recent endeavors of this kind include academic research projects such as DBpedia, KnowItAll, Probase, ReadTheWeb, and YAGO, as well as industrial ones such as Freebase and Trueknowledge. These projects provide automatically constructed knowledge bases of facts about named entities, their semantic classes, and their mutual relationships. Such world knowledge in turn enables cognitive applications and knowledge-centric services like disambiguating natural-language text, deep question answering, and semantic search for entities and relations in Web and enterprise data. Prominent examples of how knowledge bases can be harnessed include the Google Knowledge Graph and the IBM Watson question answering system. This tutorial presents state-of-the-art methods, recent advances, research opportunities, and open challenges along this avenue of knowledge harvesting and its applications.
Fabian M. Suchanek, Gerhard Weikum
ICDE1
2013 YAGO2: A Spatially and Temporally Enhanced Knowledge Base from Wikipedia: Extended Abstract
Johannes Hoffart, Fabian M. Suchanek, Klaus Berberich, Gerhard Weikum
IJCAI2
2013 Knowledge harvesting in the big-data era
abstract
The proliferation of knowledge-sharing communities such as Wikipedia and the progress in scalable information extraction from Web and text sources have enabled the automatic construction of very large knowledge bases. Endeavors of this kind include projects such as DBpedia, Freebase, KnowItAll, ReadTheWeb, and YAGO. These projects provide automatically constructed knowledge bases of facts about named entities, their semantic classes, and their mutual relationships. They contain millions of entities and hundreds of millions of facts about them. Such world knowledge in turn enables cognitive applications and knowledge-centric services like disambiguating natural-language text, semantic search for entities and relations in Web and enterprise data, and entity-oriented analytics over unstructured contents. Prominent examples of how knowledge bases can be harnessed include the Google Knowledge Graph and the IBM Watson question answering system. This tutorial presents state-of-the-art methods, recent advances, research opportunities, and open challenges along this avenue of knowledge harvesting and its applications. Particular emphasis will be on the twofold role of knowledge bases for big-data analytics: using scalable distributed algorithms for harvesting knowledge from Web and text sources, and leveraging entity-centric knowledge for deeper interpretation of and better intelligence with Big Data.
Fabian M. Suchanek, Gerhard Weikum
SIGMOD Conference1
2013 AMIE: association rule mining under incomplete evidence in ontological knowledge bases
abstract
Recent advances in information extraction have led to huge knowledge bases (KBs), which capture knowledge in a machine-readable format. Inductive Logic Programming (ILP) can be used to mine logical rules from the KB. These rules can help deduce and add missing knowledge to the KB. While ILP is a mature field, mining logical rules from KBs is different in two aspects: First, current rule mining systems are easily overwhelmed by the amount of data (state-of-the art systems cannot even run on today's KBs). Second, ILP usually requires counterexamples. KBs, however, implement the open world assumption (OWA), meaning that absent data cannot be used as counterexamples. In this paper, we develop a rule mining model that is explicitly tailored to support the OWA scenario. It is inspired by association rule mining and introduces a novel measure for confidence. Our extensive experiments show that our approach outperforms state-of-the-art approaches in terms of precision and coverage. Furthermore, our system, AMIE, mines rules orders of magnitude faster than state-of-the-art approaches.
Luis Galárraga, Christina Teflioudi, Katja Hose, Fabian M. Suchanek
WWW4
2013 YAGO2: A spatially and temporally enhanced knowledge base from Wikipedia
Johannes Hoffart, Fabian M. Suchanek, Klaus Berberich, Gerhard Weikum
Artif. Intell.2
2012 PIKM 2012: 5th ACM workshop for PhD students in information and knowledge management
abstract
The PIKM 2012 workshop is the 5th of its kind after 4 successful PhD workshops at ACM CIKM. This PhD workshop invites papers that describe the Ph.D. dissertation proposals of doctoral students in any of the CIKM areas: databases, information retrieval, data mining and knowledge management. Interdisciplinary work across these tracks is particularly encouraged. This year PIKM has received around 25 submissions from over 12 countries across the globe, among which 10 have been accepted as full papers for oral presentation while 4 have been accepted as short ones for poster presentation. The selection has been conducted based on reviews submitted by an expert team comprising 21 PC members spanning 12 countries and 6 continents with a good balance of industry and academia.
Aparna S. Varde, Fabian M. Suchanek
CIKM2
2012 PATTY: A Taxonomy of Relational Patterns with Semantic Types
Ndapandula Nakashole, Gerhard Weikum, Fabian M. Suchanek
EMNLP-CoNLL3
2012 Discovering and Exploring Relations on the Web
abstract
We propose a demonstration of PATTY, a system for learning semantic relationships from the Web. PATTY is a collection of relations learned automatically from text. It aims to be to patterns what WordNet is to words. The semantic types of PATTY relations enable advanced search over subject-predicate-object data. With the ongoing trends of enriching Web data (both text and tables) with entity-relationship-oriented semantic annotations, we believe a demo of the PATTY system will be of interest to the database community.
Ndapandula Nakashole, Gerhard Weikum, Fabian M. Suchanek
Proc. VLDB Endow.3
2011 PIKM 2011: the 4th ACM workshop for Ph.D. students in information and knowledge management
abstract
The PIKM workshop gives Ph.D. students an opportunity to present their dissertation proposals at a global stage. Similarly to the CIKM, the PIKM workshop covers a wide range of topics in the areas of databases, information retrieval and knowledge management. Interdisciplinary work across these tracks is particularly encouraged.
Anisoara Nica, Fabian M. Suchanek
CIKM2
2011 The hidden web, XML and the Semantic Web: scientific data management perspectives
abstract
The World Wide Web no longer consists just of HTML pages. Our work sheds light on a number of trends on the Internet that go beyond simple Web pages. The hidden Web provides a wealth of data in semi-structured form, accessible through Web forms and Web services. These services, as well as numerous other applications on the Web, commonly use XML, the eXtensible Markup Language. XML has become the lingua franca of the Internet that allows customized markups to be defined for specific domains. On top of XML, the Semantic Web grows as a common structured data source. In this work, we first explain each of these developments in detail. Using real-world examples from scientific domains of great interest today, we then demonstrate how these new developments can assist the managing, harvesting, and organization of data on the Web. On the way, we also illustrate the current research avenues in these domains. We believe that this effort would help bridge multiple database tracks, thereby attracting researchers with a view to extend database technology.
Fabian M. Suchanek, Aparna S. Varde, Richi Nayak, Pierre Senellart
EDBT1
2011 Watermarking for Ontologies
abstract
In this paper, we study watermarking methods to prove the ownership of an ontology. Different from existing approaches, we propose to watermark not by altering existing statements, but by removing them. Thereby, our approach does not introduce false statements into the ontology. We show how ownership of ontologies can be established with provably tight probability bounds, even if only parts of the ontology are being re-used. We finally demonstrate the viability of our approach on real-world ontologies.
Fabian M. Suchanek, David Gross-Amblard, Serge Abiteboul
ISWC (1)1
2011 PARIS: Probabilistic Alignment of Relations, Instances, and Schema
abstract
One of the main challenges that the Semantic Web faces is the integration of a growing number of independently designed ontologies. In this work, we present paris, an approach for the automatic alignment of ontologies. paris aligns not only instances, but also relations and classes. Alignments at the instance level cross-fertilize with alignments at the schema level. Thereby, our system provides a truly holistic solution to the problem of ontology alignment. The heart of the approach is probabilistic, i.e., we measure degrees of matchings based on probability estimates. This allows paris to run without any parameter tuning. We demonstrate the efficiency of the algorithm and its precision through extensive experiments. In particular, we obtain a precision of around 90% in experiments with some of the world's largest ontologies.
Fabian M. Suchanek, Serge Abiteboul, Pierre Senellart
Proc. VLDB Endow.1
2010 Active knowledge: dynamically enriching RDF knowledge bases by web services
abstract
The proliferation of knowledge-sharing communities and the advances in information extraction have enabled the construction of large knowledge bases using the RDF data model to represent entities and relationships. However, as the Web and its latently embedded facts evolve, a knowledge base can never be complete and up-to-date. On the other hand, a rapidly increasing suite of Web services provide access to timely and high-quality information, but this is encapsulated by the service interface. We propose to leverage the information that could be dynamically obtained from Web services in order to enrich RDF knowledge bases on the fly whenever the knowledge base does not suffice to answer a user query.
Nicoleta Preda, Gjergji Kasneci, Fabian M. Suchanek, Thomas Neumann 0001, Wenjun Yuan, Gerhard Weikum
SIGMOD Conference3
2009 Knowledge Discovery over the Deep Web, Semantic Web and XML
Aparna S. Varde, Fabian M. Suchanek, Richi Nayak, Pierre Senellart
DASFAA2
2009 STAR: Steiner-Tree Approximation in Relationship Graphs
abstract
Large graphs and networks are abundant in modern information systems: entity-relationship graphs over relational data or Web-extracted entities, biological networks, social online communities, knowledge bases, and many more. Often such data comes with expressive node and edge labels that allow an interpretation as a semantic graph, and edge weights that reflect the strengths of semantic relations between entities. Finding close relationships between a given set of two, three, or more entities is an important building block for many search, ranking, and analysis tasks. From an algorithmic point of view, this translates into computing the best Steiner trees between the given nodes, a classical NP-hard problem. In this paper, we present a new approximation algorithm, coined STAR, for relationship queries over large relationship graphs. We prove that for n query entities, STAR yields an O(log(n))-approximation of the optimal Steiner tree in pseudopolynomial run-time, and show that in practical cases the results returned by STAR are qualitatively comparable to or even better than the results returned by a classical 2-approximation algorithm. We then describe an extension to our algorithm to return the top-k Steiner trees. Finally, we evaluate our algorithm over both main-memory as well as completely diskresident graphs containing millions of nodes. Our experiments show that in terms of efficiency STAR outperforms the best state-of-the-art database methods by a large margin, and also returns qualitatively better results.
Gjergji Kasneci, Maya Ramanath, Mauro Sozio, Fabian M. Suchanek, Gerhard Weikum
ICDE4
2009 Graffiti: node labeling in heterogeneous networks
abstract
We introduce a multi-label classification model and algorithm for labeling heterogeneous networks, where nodes belong to different types and different types have different sets of classification labels. We present a graph-based approach which models the mutual influence between nodes in the network as a random walk. When viewing class labels as "colors", the random surfer is "spraying" different node types with different color palettes; hence the name Graffiti. We demonstrate the performance gains of our method by comparing it to three state-of-the-art techniques for graph-based classification.
Ralitsa Angelova, Gjergji Kasneci, Fabian M. Suchanek, Gerhard Weikum
WWW3
2009 SOFIE: a self-organizing framework for information extraction
abstract
This paper presents SOFIE, a system for automated ontology extension. SOFIE can parse natural language documents, extract ontological facts from them and link the facts into an ontology. SOFIE uses logical reasoning on the existing knowledge and on the new knowledge in order to disambiguate words to their most probable meaning, to reason on the meaning of text patterns and to take into account world knowledge axioms. This allows SOFIE to check the plausibility of hypotheses and to avoid inconsistencies with the ontology. The framework of SOFIE unites the paradigms of pattern matching, word sense disambiguation and ontological reasoning in one unified model. Our experiments show that SOFIE delivers high-quality output, even from unstructured Internet documents.
Fabian M. Suchanek, Mauro Sozio, Gerhard Weikum
WWW1
2009 ANGIE: Active Knowledge for Interactive Exploration
abstract
We present ANGIE, a system that can answer user queries by combining knowledge from a local database with knowledge retrieved from Web services. If a user poses a query that cannot be answered by the local database alone, ANGIE calls the appropriate Web services to retrieve the missing information. This information is integrated seamlessly and transparently into the local database, so that the user can query and browse the knowledge base while appropriate Web services are called automatically in the background.
Nicoleta Preda, Fabian M. Suchanek, Gjergji Kasneci, Thomas Neumann 0001, Maya Ramanath, Gerhard Weikum
Proc. VLDB Endow.2
2008 Social tags: meaning and suggestions
abstract
This paper aims to quantify two common assumptions about social tagging: (1) that tags are "meaningful" and (2) that the tagging process is influenced by tag suggestions. For (1), we analyze the semantic properties of tags and the relationship between the tags and the content of the tagged page. Our analysis is based on a corpus of search keywords, contents, titles, and tags applied to several thousand popular Web pages. Among other results, we find that the more popular tags of a page tend to be the more meaningful ones. For (2), we develop a model of how the influence of tag suggestions can be measured. From a user study with over 4,000 participants, we conclude that roughly one third of the tag applications may be induced by the suggestions. Our results would be of interest for designers of social tagging systems and are a step towards understanding how to best leverage social tags for applications such as search and information extraction.
Fabian M. Suchanek, Milan Vojnovic, Dinan Gunawardena
CIKM1
2008 NAGA: Searching and Ranking Knowledge
abstract
The Web has the potential to become the world's largest knowledge base. In order to unleash this potential, the wealth of information available on the Web needs to be extracted and organized. There is a need for new querying techniques that are simple and yet more expressive than those provided by standard keyword-based search engines. Searching for knowledge rather than Web pages needs to consider inherent semantic structures like entities (person, organization, etc.) and relationships (isA, located In, etc.). In this paper, we propose NAGA, a new semantic search engine. NAGA builds on a knowledge base, which is organized as a graph with typed edges, and consists of millions of entities and relationships extracted from Web-based corpora. A graph-based query language enables the formulation of queries with additional semantic information. We introduce a novel scoring model, based on the principles of generative language models, which formalizes several notions such as confidence, informativeness and compactness and uses them to rank query results. We demonstrate NAGA's superior result quality over state-of-the-art search engines and question answering systems.
Gjergji Kasneci, Fabian M. Suchanek, Georgiana Ifrim, Maya Ramanath, Gerhard Weikum
ICDE2
2008 Integrating YAGO into the Suggested Upper Merged Ontology
abstract
Ontologies are becoming more and more popular as background knowledge for intelligent applications. Up to now, there has been a schism between manually assembled, highly axiomatic ontologies and large, automatically constructed knowledge bases. This paper discusses how the two worlds can be brought together by combining the high-level axiomatizations from the standard upper merged ontology (SUMO) with the extensive world knowledge of the YAGO ontology. The result is a new large-scale formal ontology, which provides information about millions of entities such as people, cities, organizations, and companies.
Gerard de Melo, Fabian M. Suchanek, Adam Pease
ICTAI (1)2
2008 NAGA: harvesting, searching and ranking knowledge
abstract
The presence of encyclopedic Web sources, such as Wikipedia, the Internet Movie \nDatabase (IMDB), World Factbook, etc. calls for new querying techniques that \nare simple and yet more expressive than those provided by standard \nkeyword-based search engines. Searching for explicit knowledge needs to \nconsider inherent semantic structures involving entities and relationships.\n\nIn this demonstration proposal, we describe a semantic search system named \nNAGA. NAGA operates on a knowledge graph, which contains millions of entities \nand relationships derived from various encyclopedic Web sources, such as the \nones above. NAGA's graph-based query language is geared towards expressing \nqueries with additional semantic information. Its scoring model is based on the \nprinciples of generative language models, and formalizes several desiderata \nsuch as confidence, informativeness and compactness of answers.\n\nWe propose a demonstration of NAGA which will allow users to browse the \nknowledge base through a user interface, enter queries in NAGA's query language \nand tune the ranking parameters to test various ranking aspects.
Gjergji Kasneci, Fabian M. Suchanek, Georgiana Ifrim, Shady Elbassuoni, Maya Ramanath, Gerhard Weikum
SIGMOD Conference2
2008 TOB: Timely Ontologies for Business Relations
Fabian M. Suchanek, Lihua Yue, Gerhard Weikum
WebDB2
2008 YAGO: A Large Ontology from Wikipedia and WordNet
Fabian M. Suchanek, Gjergji Kasneci, Gerhard Weikum
J. Web Semant.1
2007 ESTER: efficient search on text, entities, and relations
abstract
We present ESTER, a modular and highly efficient system for combined full-text and ontology search. ESTER builds on a query engine that supports two basic operations: prefix search and join. Both of these can be implemented very efficiently with a compact index, yet in combination provide powerful querying capabilities. We show how ESTER can answer basic SPARQL graph-pattern queries on the ontology by reducing them to a small number of these two basic operations. ESTER further supports a natural blend of such semantic queries with ordinary full-text queries. Moreover, the prefix search operation allows for a fully interactive and proactive user interface, which after every keystroke suggests to the user possible semantic interpretations of his or her query, and speculatively executes the most likely of these interpretations. As a proof of concept, we applied ESTER to the English Wikipedia, which contains about 3 million documents, combined with the recent YAGO ontology, which contains about 2.5 million facts. For a variety of complex queries, ESTER achieves worst-case query processing times of a fraction of a second, on a single machine, with an index size of about 4 GB.
Hannah Bast, Alexandru Chitea, Fabian M. Suchanek, Ingmar Weber
SIGIR3
2007 How NAGA uncoils: searching with entities and relations
abstract
Current keyword-oriented search engines for theWorld WideWeb do not allow specifying the semantics of queries. We address this limitation with NAGA1, a new semantic search engine. NAGA builds on a large semantic knowledge base of binary relationships (facts) derived from the Web. NAGA provides a simple, yet expressive query language to query this knowledge base. The results are then ranked with an intuitive scoring mechanism. We show the effectiveness and utility of NAGA by comparing its output with that of Googleon some interesting queries.
Gjergji Kasneci, Fabian M. Suchanek, Maya Ramanath, Gerhard Weikum
WWW2
2007 Yago: a core of semantic knowledge
abstract
We present YAGO, a light-weight and extensible ontology with high coverage and quality. YAGO builds on entities and relations and currently contains more than 1 million entities and 5 million facts. This includes the Is-A hierarchy as well as non-taxonomic relations between entities (such as HASONEPRIZE). The facts have been automatically extracted from Wikipedia and unified with WordNet, using a carefully designed combination of rule-based and heuristic methods described in this paper. The resulting knowledge base is a major step beyond WordNet: in quality by adding knowledge about individuals like persons, organizations, products, etc. with their semantic relationships - and in quantity by increasing the number of facts by more than an order of magnitude. Our empirical evaluation of fact correctness shows an accuracy of about 95%. YAGO is based on a logically clean model, which is decidable, extensible, and compatible with RDFS. Finally, we show how YAGO can be further extended by state-of-the-art information extraction techniques.
Fabian M. Suchanek, Gjergji Kasneci, Gerhard Weikum
WWW1
2006 Combining linguistic and statistical analysis to extract relations from web documents
abstract
The World Wide Web provides a nearly endless source of knowledge, which is mostly given in natural language. A first step towards exploiting this data automatically could be to extract pairs of a given semantic relation from text documents - for example all pairs of a person and her birthdate. One strategy for this task is to find text patterns that express the semantic relation, to generalize these patterns, and to apply them to a corpus to find new pairs. In this paper, we show that this approach profits significantly when deep linguistic structures are used instead of surface text patterns. We demonstrate how linguistic structures can be represented for machine learning, and we provide a theoretical analysis of the pattern matching approach. We show the benefits of our approach by extensive experiments with our prototype system LEILA.
Fabian M. Suchanek, Georgiana Ifrim, Gerhard Weikum
KDD1