EDBT 2026 Demo / reviewers in the wild / expert
Johannes Hoffart
dblp:99/7948
· DBLP profile ↗
18ranked-venue papers
7as first author
5since 2021 · last 2023
0000-0001-5426-785XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 13 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Can Persistent Homology provide an efficient alternative for Evaluation of Knowledge Graph Completion Methods?abstractIn this paper we present a novel method, Knowledge Persistence (), for faster evaluation of Knowledge Graph (KG) completion approaches. Current ranking-based evaluation is quadratic in the size of the KG, leading to long evaluation times and consequently a high carbon footprint. addresses this by representing the topology of the KG completion methods through the lens of topological data analysis, concretely using persistent homology. The characteristics of persistent homology allow to evaluate the quality of the KG completion looking only at a fraction of the data. Experimental results on standard datasets show that the proposed metric is highly correlated with ranking metrics (Hits@N, MR, MRR). Performance evaluation shows that is computationally efficient: In some cases, the evaluation time (validation+test) of a KG completion method has been reduced from 18 hours (using Hits@10) to 27 seconds (using ), and on average (across methods & data) reduces the evaluation time (validation+test) by ≈ 99.96%. Anson Bastos, Kuldeep Singh 0001, Abhishek Nadgeri, Johannes Hoffart, Manish Singh 0002, Toyotaro Suzumura |
WWW | 4 |
| 2021 | HopfE: Knowledge Graph Representation Learning using Inverse Hopf FibrationsabstractRecently, several Knowledge Graph Embedding (KGE) approaches have been devised to represent entities and relations in a dense vector space and employed in downstream tasks such as link prediction. A few KGE techniques address interpretability, i.e., mapping the connectivity patterns of the relations (symmetric/asymmetric, inverse, and composition) to a geometric interpretation such as rotation. Other approaches model the representations in higher dimensional space such as four-dimensional space (4D) to enhance the ability to infer the connectivity patterns (i.e., expressiveness). However, modeling relation and entity in a 4D space often comes at the cost of interpretability. We propose HopfE, a novel KGE approach aiming to achieve the interpretability of inferred relations in the four-dimensional space. HopfE models the structural embeddings in 3D Euclidean space. Next, we map the entity embedding vector from a 3D Euclidean space to a 4D hypersphere using the inverse Hopf Fibration, in which we embed the semantic information from the KG ontology. Thus, HopfE considers the structural and semantic properties of the entities without losing expressivity and interpretability. Our empirical results on four well-known benchmarks achieve state-of-the-art performance for KG completion. Anson Bastos, Kuldeep Singh 0001, Abhishek Nadgeri, Saeedeh Shekarpour, Isaiah Onando Mulang', Johannes Hoffart |
CIKM | 6 |
| 2021 | CHOLAN: A Modular Approach for Neural Entity Linking on Wikipedia and WikidataabstractManoj Prabhakar Kannan Ravi, Kuldeep Singh, Isaiah Onando Mulang’, Saeedeh Shekarpour, Johannes Hoffart, Jens Lehmann. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Manoj Prabhakar Kannan Ravi, Kuldeep Singh 0001, Isaiah Onando Mulang', Saeedeh Shekarpour, Johannes Hoffart, Jens Lehmann 0001 |
EACL | 5 |
| 2021 | Unsupervised Multi-View Post-OCR Error Correction With Language ModelsabstractWe investigate post-OCR correction in a setting where we have access to different OCR views of the same document.The goal of this study is to understand if a pretrained language model (LM) can be used in an unsupervised way to reconcile the different OCR views such that their combination contains fewer errors than each individual view.This approach is motivated by scenarios in which unconstrained text generation for error correction is too risky.We evaluated different pretrained LMs on two datasets and found significant gains in realistic scenarios with up to 15% WER improvement over the best OCR view.We also show the importance of domain adaptation for post-OCR correction on out-of-domain documents. Luciano Del Corro, Samuel Broscheit, Johannes Hoffart, Eliot Brenner |
EMNLP (1) | 4 |
| 2021 | RECON: Relation Extraction using Knowledge Graph Context in a Graph Neural NetworkabstractIn this paper, we present a novel method named RECON, that automatically identifies relations in a sentence (sentential relation extraction) and aligns to a knowledge graph (KG). RECON uses a graph neural network to learn representations of both the sentence as well as facts stored in a KG, improving the overall extraction quality. These facts, including entity attributes (label, alias, description, instance-of) and factual triples, have not been collectively used in the state of the art methods. We evaluate the effect of various forms of representing the KG context on the performance of RECON. The empirical evaluation on two standard relation extraction datasets shows that RECON significantly outperforms all state of the art methods on NYT Freebase and Wikidata datasets. Anson Bastos, Abhishek Nadgeri, Kuldeep Singh 0001, Isaiah Onando Mulang', Saeedeh Shekarpour, Johannes Hoffart, Manohar Kaul |
WWW | 6 |
| 2020 | Evaluating the Impact of Knowledge Graph Context on Entity Disambiguation ModelsabstractPretrained Transformer models have emerged as state-of-the-art approaches that learn contextual information from the text to improve the performance of several NLP tasks. These models, albeit powerful, still require specialized knowledge in specific scenarios. In this paper, we argue that context derived from a knowledge graph (in our case: Wikidata) provides enough signals to inform pretrained transformer models and improve their performance for named entity disambiguation (NED) on Wikidata KG. We further hypothesize that our proposed KG context can be standardized for Wikipedia, and we evaluate the impact of KG context on the state of the art NED model for the Wikipedia knowledge base. Our empirical results validate that the proposed KG context can be generalized (for Wikipedia), and providing KG context in transformer architectures considerably outperforms the existing baselines, including the vanilla transformer models. Isaiah Onando Mulang', Kuldeep Singh 0001, Chaitali Prabhu, Abhishek Nadgeri, Johannes Hoffart, Jens Lehmann 0001 |
CIKM | 5 |
| 2016 | Discovering Entities with Just a Little Help from YouabstractLinking entities like people, organizations, books, music groups and their songs in text to knowledge bases (KBs) is a fundamental task for many downstream search and mining applications. Achieving high disambiguation accuracy crucially depends on a rich and holistic representation of the entities in the KB. For popular entities, such a representation can be easily mined from Wikipedia, and many current entity disambiguation and linking methods make use of this fact. However, Wikipedia does not contain long-tail entities that only few people are interested in, and also at times lags behind until newly emerging entities are added. For such entities, mining a suitable representation in a fully automated fashion is very difficult, resulting in poor linking accuracy. Johannes Hoffart, Avishek Anand |
CIKM | 2 |
| 2016 | YAGO: A Multilingual Knowledge Base from Wikipedia, Wordnet, and GeonamesabstractYAGO is a large knowledge base that is built automatically from Wikipedia, WordNet and GeoNames. The project combines information from Wikipedias in 10 different languages into a coherent whole, thus giving the knowledge a multilingual dimension. It also attaches spatial and temporal information to many facts, and thus allows the user to query the data over space and time. YAGO focuses on extraction quality and achieves a manually evaluated precision of 95 %. In this paper, we explain how YAGO is built from its sources, how its quality is evaluated, how a user can access it, and how other projects utilize it. Thomas Rebele, Fabian M. Suchanek, Johannes Hoffart, Asia J. Biega, Erdal Kuzey, Gerhard Weikum |
ISWC (2) | 3 |
| 2016 | Context-Sensitive Auto-Completion for Searching with Entities and CategoriesabstractWhen searching in a document collection by keywords, good auto-completion suggestions can be derived from query logs and corpus statistics. On the other hand, when querying documents which have automatically been linked to entities and semantic categories, auto-completion has not been investigated much. We have developed a semantic auto-completion system, where suggestions for entities and categories are computed in real-time from the context of already entered entities or categories and from entity-level co-occurrence statistics for the underlying corpus. Given the huge size of the knowledge bases that underlie this setting, a challenge is to compute the best suggestions fast enough for interactive user experience. Our demonstration shows the effectiveness of our method, and its interactive usability. Andreas Schmidt 0002, Johannes Hoffart, Dragan Milchevski, Gerhard Weikum |
SIGIR | 2 |
| 2014 | AESTHETICS: Analytics with Strings, Things, and CatsabstractThis paper describes an advanced news analytics and exploration system that allows users to visualize trends of entities like politicians, countries, and organizations in continuously updated news articles. Our system improves state-of-the-art text analytics by linking ambiguous names in news articles to entities in knowledge bases like Freebase, DBpedia or YAGO. This step enables indexing entities and interpreting the contents in terms of entities. This way, the analysis of trends and co-occurrences of entities gains accuracy, and by leveraging the taxonomic type hierarchy of knowledge bases, also in expressiveness and usability. In particular, we can analyze not only individual entities, but also categories of entities and their combinations, including co-occurrences with informative text phrases. Our Web-based system demonstrates the power of this approach by insightful anecdotic analysis of recent events in the news. Johannes Hoffart, Dragan Milchevski, Gerhard Weikum |
CIKM | 1 |
| 2014 | STICS: searching with strings, things, and catsabstractThis paper describes an advanced search engine that supports users in querying documents by means of keywords, entities, and categories. Users simply type words, which are automatically mapped onto appropriate suggestions for entities and categories. Based on named-entity disambiguation, the search engine returns documents containing the query's entities and prominent entities from the query's categories. Johannes Hoffart, Dragan Milchevski, Gerhard Weikum |
SIGIR | 1 |
| 2014 | Discovering emerging entities with ambiguous namesabstractKnowledge bases (KB's) contain data about a large number of people, organizations, and other entities. However, this knowledge can never be complete due to the dynamics of the ever-changing world: new companies are formed every day, new songs are composed every minute and become of interest for addition to a KB. To keep up with the real world's entities, the KB maintenance process needs to continuously discover newly emerging entities in news and other Web streams. In this paper we focus on the most difficult case where the names of new entities are ambiguous. This raises the technical problem to decide whether an observed name refers to a known entity or represents a new entity. This paper presents a method to solve this problem with high accuracy. It is based on a new model of measuring the confidence of mapping an ambiguous mention to an existing entity, and a new model of representing a new entity with the same ambiguous name as a set of weighted keyphrases. The method can handle both Wikipedia-derived entities that typically constitute the bulk of large KB's as well as entities that exist only in other Web sources such as online communities about music or movies. Experiments show that our entity discovery method outperforms previous methods for coping with out-of-KB entities (called unlinkable in entity linking). Johannes Hoffart, Yasemin Altun, Gerhard Weikum |
WWW | 1 |
| 2013 | YAGO2: A Spatially and Temporally Enhanced Knowledge Base from Wikipedia: Extended Abstract
Johannes Hoffart, Fabian M. Suchanek, Klaus Berberich, Gerhard Weikum |
IJCAI | 1 |
| 2013 | YaLi: a crowdsourcing plug-in for NERDabstractWe demonstrate the YaLi browser plug-in which discovers named entities in Web pages and provides background knowledge about them. The plug-in is implemented with two purposes. From a user perspective, it enriches the browsing experience with entities, helping users with their information needs. From the research perspective, we aim to improve the methods that are used for named entity recognition and disambiguation (NERD) by leveraging the plug-in as an implicit crowdsourcing platform. YaLi tracks the system's errors and the users' corrections, and also gathers implicit training data for improving NERD accuracy. Yafang Wang, Lili Jiang 0002, Johannes Hoffart, Gerhard Weikum |
SIGIR | 3 |
| 2013 | YAGO2: A spatially and temporally enhanced knowledge base from Wikipedia
Johannes Hoffart, Fabian M. Suchanek, Klaus Berberich, Gerhard Weikum |
Artif. Intell. | 1 |
| 2012 | KORE: keyphrase overlap relatedness for entity disambiguationabstractMeasuring the semantic relatedness between two entities is the basis for numerous tasks in IR, NLP, and Web-based knowledge extraction. This paper focuses on disambiguating names in a Web or text document by jointly mapping all names onto semantically related entities registered in a knowledge base. To this end, we have developed a novel notion of semantic relatedness between two entities represented as sets of weighted (multi-word) keyphrases, with consideration of partially overlapping phrases. This measure improves the quality of prior link-based models, and also eliminates the need for (usually Wikipedia-centric) explicit interlinkage between entities. Thus, our method is more versatile and can cope with long-tail and newly emerging entities that have few or no links associated with them. For efficiency, we have developed approximation techniques based on min-hash sketches and locality-sensitive hashing. Our experiments on semantic relatedness and on named entity disambiguation demonstrate the superiority of our method compared to state-of-the-art baselines. Johannes Hoffart, Stephan Seufert, Dat Ba Nguyen, Martin Theobald, Gerhard Weikum |
CIKM | 1 |
| 2011 | Robust Disambiguation of Named Entities in Text
Johannes Hoffart, Mohamed Amir Yosef, Ilaria Bordino, Hagen Fürstenau, Manfred Pinkal, Marc Spaniol, Bilyana Taneva, Stefan Thater, Gerhard Weikum |
EMNLP | 1 |
| 2011 | AIDA: An Online Tool for Accurate Disambiguation of Named Entities in Text and Tables
Mohamed Amir Yosef, Johannes Hoffart, Ilaria Bordino, Marc Spaniol, Gerhard Weikum |
Proc. VLDB Endow. | 2 |