Harald Sack

dblp:s/HaraldSack · DBLP profile ↗
← Back
21ranked-venue papers in the field
0as first author
7since 2021 · last 2026
0000-0001-7069-9804ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 11Information Retrieval & Web Search · 6Other / Interdisciplinary · 2Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 NERdME: A Named Entity Recognition Dataset for Indexing Research Artifacts in Code Repositories
abstract
Existing scholarly information extraction (SIE) datasets focus on scientific papers and overlook implementation-level details in code repositories. README files describe datasets, source code, and other implementation-level artifacts, however, their free-form Markdown offers little semantic structure, making automatic information extraction difficult. To address this gap, NERdME is introduced: 200 manually annotated README files with over νm10000 labeled spans and 10 entity types. Baseline results using large language models and fine-tuned transformers show clear differences between paper-level and implementation-level entities, indicating the value of extending SIE benchmarks with entity types available in README files. A downstream entity-linking experiment was conducted to demonstrate that entities derived from READMEs can support artifact discovery and metadata integration.
Genet Asefa Gesese, Zongxiong Chen, Shufan Jiang 0001, Mary Ann Tan, Zhaotai Liu, Sonja Schimmler, Harald Sack
WWW7
2025 End-to-end Information Extraction from Archival Records with Multimodal Large Language Models
abstract
Semi-structured Document Understanding presents a challenging research task due to the significant variations in layout, style, font, and content of documents.This complexity is further amplified when dealing with born-analogue historical documents, such as digitised archival records, which contain degraded print, handwritten annotations, stamps, marginalia and inconsistent formatting resulting from historical production and digitisation processes.Traditional approaches for extracting information from semi-structured documents rely on manual labour, making them costly and inefficient.This is partly due to the fact that within document collections, there are various layout types, each requiring customised optimisation to account for structural differences, which substantially increases the effort needed to achieve consistent quality.The emergence of Multimodal Large Language Models (MLLMs) has significantly advanced Document Understanding by enabling flexible, prompt-based understanding of document images, needless of OCR outputs or layout encodings.Moreover, the encoder-decoder architectures have overcome the limitations of encoder-only models, such as reliance on annotated datasets and fixed input lengths.However, there still remains a gap in effectively applying these models in real-world scenarios.To address this gap, we first introduce BZKOpen, a new annotated dataset designed for key information extraction from historical German index cards.Furthermore, we systematically assess the capabilities of several state-of-the-art MLLMs-including the open-source InternVL2.0and InternVL2.5 series, and the commercial GPT-4o-mini-on the task of extracting
Mahsa Vafaie, Sven Hertling, Inger Banse-Strobel, Kevin Dubout, Harald Sack
CIKM5
2023 CourtDocs Ontology: Towards a Data Model for Representation of Historical Court Proceedings
abstract
For several decades researchers have studied legal documents for insights into the evolution of legal norms and strategies, in their social and cultural context. Analysing these documents and the associated legislative sessions, trials and court cases helps uncover hidden narratives and patterns, as well as showcase the lessons learnt. The field of knowledge engineering has contributed to the growing interest in the development and use of legal ontologies that aim at providing machine-readable foundations to model legal concepts, relations and processes. Legal ontologies have been used for legal knowledge management and as knowledge bases in legal knowledge systems. With a focus on the Wiedergutmachung project as a use case, this paper presents an overview of the existing legal ontologies, demonstrates the gap to align them with the essential conceptual framework required to model historical court proceedings with respect to provenance information, and presents the ongoing work towards developing the CourtDocs Ontology by utilising existing standards and ontologies on the intersection of the legal domain, history and archival sciences. The Wiedergutmachung project centres around constructing a knowledge graph as a backbone for information systems, based on historical archival records from the compensation procedure in post-World War II Germany.
Mahsa Vafaie, Oleksandra Bruns, Nastasja Pilz, Jörg Waitelonis, Harald Sack
K-CAP5
2022 Entity Type Prediction Leveraging Graph Walks and Entity Descriptions
Russa Biswas, Jan Portisch, Heiko Paulheim, Harald Sack, Mehwish Alam
ISWC4
2021 Cat2Type: Wikipedia Category Embeddings for Entity Typing in Knowledge Graphs
abstract
The entity type information in Knowledge Graphs (KGs) such as DBpedia, Freebase, etc. is often incomplete due to automated generation. Entity Typing is the task of assigning or inferring the semantic type of an entity in a KG. This paper introduces an approach named Cat2Type which exploits the Wikipedia Categories to predict the missing entity types in a KG. This work extracts information from Wikipedia Category names and the Wikipedia Category graph which are the sources of rich semantic information about the entities. In Cat2Type, the characteristic features of the entities encapsulated in Wikipedia Category names are exploited using Neural Language Models. On the other hand, a Wikipedia Category graph is constructed to capture the connection between the categories. The Node level representations are learned by optimizing the neighbourhood information on the Wikipedia category graph. These representations are then used for entity type prediction via classification. The performance of Cat2Type is assessed on two real-world benchmark datasets DBpedia630k and FIGER. The experiments depict that Cat2Type obtained a significant improvement over state-of-the-art approaches.
Russa Biswas, Radina Sofronova, Harald Sack, Mehwish Alam
K-CAP3
2021 LiterallyWikidata - A Benchmark for Knowledge Graph Completion Using Literals
Genet Asefa Gesese, Mehwish Alam, Harald Sack
ISWC3
2021 L2D 2021: First International Workshop on Enabling Data-Driven Decisions from Learning on the Web
abstract
By offering courses and resources, learning platforms on the Web have been attracting lots of participants, and the interactions with these systems have generated a vast amount of learning-related data. Their collection, processing and analysis have promoted a significant growth of learning analytics and have opened up new opportunities for supporting and assessing educational experiences. To provide all the stakeholders involved in the educational process with a timely guidance, being able to understand student's behavior and enable models which provide data-driven decisions pertaining to the learning domain is a primary property of online platforms, aiming at maximizing learning outcomes. In this workshop, we focus on collecting new contributions in this emerging area and on providing a common ground for researchers and practitioners (Website: https://mirkomarras.github.io/l2d-wsdm2021).
Danilo Dessì, Tanja Käser, Mirko Marras, Elvira Popescu, Harald Sack
WSDM5
2020 CSSA'20: Workshop on Combining Symbolic and Sub-Symbolic Methods and their Applications
abstract
There has been a rapid growth in the use of symbolic representations along with their applications in many important tasks. Symbolic representations, in the form of Knowledge Graphs (KGs), constitute large networks of real-world entities and their relationships. On the other hand, sub-symbolic artificial intelligence has also become a mainstream area of research. This workshop brought together researchers to discuss and foster collaborations on the intersection of these two areas.
Mehwish Alam, Paul Groth, Pascal Hitzler, Heiko Paulheim, Harald Sack, Volker Tresp
CIKM5
2020 Entity-Based Short Text Classification Using Convolutional Neural Networks
Mehwish Alam, Qingyuan Bie, Rima Dessi, Harald Sack
EKAW4
2020 AI-KG: An Automatically Generated Knowledge Graph of Artificial Intelligence
abstract
Scientific knowledge has been traditionally disseminated and preserved through research articles published in journals, conference proceedings, and online archives. However, this article-centric paradigm has been often criticized for not allowing to automatically process, categorize, and reason on this knowledge. An alternative vision is to generate a semantically rich and interlinked description of the content of research publications. In this paper, we present the Artificial Intelligence Knowledge Graph (AI-KG), a large-scale automatically generated knowledge graph that describes 820K research entities. AI-KG includes about 14M RDF triples and 1.2M reified statements extracted from 333K research publications in the field of AI, and describes 5 types of entities (tasks, methods, metrics, materials, others) linked by 27 relations. AI-KG has been designed to support a variety of intelligent services for analyzing and making sense of research dynamics, supporting researchers in their daily job, and helping to inform decision-making in funding bodies and research policymakers. AI-KG has been generated by applying an automatic pipeline that extracts entities and relationships using three tools: DyGIE++, Stanford CoreNLP, and the CSO Classifier. It then integrates and filters the resulting triples using a combination of deep learning and semantic technologies in order to produce a high-quality knowledge graph. This pipeline was evaluated on a manually crafted gold standard, yielding competitive results. AI-KG is available under CC BY 4.0 and can be downloaded as a dump or queried via a SPARQL endpoint.
Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta, Harald Sack
ISWC (2)6
2020 Weakly Supervised Short Text Categorization Using World Knowledge
Rima Dessi, Lei Zhang 0034, Mehwish Alam, Harald Sack
ISWC (1)4
2019 Knowledge-Based Short Text Categorization Using Entity and Category Embedding
abstract
Short text categorization is an important task due to the rapid growth of online available short texts in various domains such as web search snippets, etc. Most of the traditional methods suffer from sparsity and shortness of the text. Moreover, supervised learning methods require a significant amount of training data and manually labeling such data can be very time-consuming and costly. In this study, we propose a novel probabilistic model for Knowledge-Based Short Text Categorization (KBSTC), which does not require any labeled training data to classify a short text. This is achieved by leveraging entities and categories from large knowledge bases, which are further embedded into a common vector space, for which we propose a new entity and category embedding model. Given a short text, its category (e.g. Business , Sports , etc.) can then be derived based on the entities mentioned in the text by exploiting semantic similarity between entities and categories. To validate the effectiveness of the proposed method, we conducted experiments on two real-world datasets, i.e., AG News and Google Snippets. The experimental results show that our approach significantly outperforms the classification approaches which do not require any labeled data, while it comes close to the results of the supervised approaches.
Rima Dessi, Lei Zhang 0034, Maria Koutraki, Harald Sack
ESWC4
2016 TIB|AV-Portal: Integrating Automatically Generated Video Annotations into the Web of Data
Jörg Waitelonis, Margret Plank, Harald Sack
TPDL3
2016 I am a Machine, Let Me Understand Web Media!
Magnus Knuth, Jörg Waitelonis, Harald Sack
ICWE3
2015 GERBIL: General Entity Annotator Benchmarking Framework
abstract
We present GERBIL, an evaluation framework for semantic entity annotation. The rationale behind our framework is to provide developers, end users and researchers with easy-to-use interfaces that allow for the agile, fine-grained and uniform evaluation of annotation tools on multiple datasets. By these means, we aim to ensure that both tool developers and end users can derive meaningful insights pertaining to the extension, integration and use of annotation applications. In particular, GERBIL provides comparable results to tool developers so as to allow them to easily discover the strengths and weaknesses of their implementations with respect to the state of the art. With the permanent experiment URIs provided by our framework, we ensure the reproducibility and archiving of evaluation results. Moreover, the framework generates data in machine-processable format, allowing for the efficient querying and post-processing of evaluation results. Finally, the tool diagnostics provided by GERBIL allows deriving insights pertaining to the areas in which tools should be further refined, thus allowing developers to create an informed agenda for extensions and end users to detect the right tools for their purposes. GERBIL aims to become a focal point for the state of the art, driving the research agenda of the community by presenting comparable objective evaluation results.
Ricardo Usbeck, Michael Röder, Axel-Cyrille Ngonga Ngomo, Ciro Baron, Andreas Both 0001, Martin Brümmer, Diego Ceccarelli, Marco Cornolti, Didier Cherix, Bernd Eickmann, Paolo Ferragina, Christiane Lemke, Andrea Moro 0001, Roberto Navigli, Francesco Piccinno, Giuseppe Rizzo 0002, Harald Sack, René Speck, Raphaël Troncy, Jörg Waitelonis, Lars Wesemann
WWW17
2015 PatchR: A Framework for Linked Data Change Requests
abstract
Incorrect or outdated data is a common problem when working with Linked Data in real world applications. Linked Data is distributed over the web and under control of various dataset publishers. It is difficult for data publishers to ensure the quality and timeliness of the data all by themselves, though they might receive individual complaints by data users, who identified incorrect or missing data. Indeed, the authors see Linked Data consumers equally responsible for the quality of the datasets they use. PatchR provides a vocabulary to report incorrect data and to propose changes to correct them. Based on the PatchR ontology a framework is suggested that allows users to efficiently report and data publishers to handle change requests for their datasets.
Magnus Knuth, Harald Sack
Int. J. Semantic Web Inf. Syst.2
2013 Semantic Multimedia Information Retrieval Based on Contextual Descriptions
Nadine Steinmetz, Harald Sack
ESWC2
2012 Evaluating Entity Summarization Using a Game-Based Ground Truth
Andreas Thalhammer 0001, Magnus Knuth, Harald Sack
ISWC (2)3
2008 Who Reads and Writes the Social Web? A Security Architecture for Web 2.0 Applications
abstract
The World Wide Web has changed during the last decade. The so-called Web 2.0 enables inexperienced users to become worldwide publishers. Most often these users also don't have any idea about how to protect their own user-generated content or how to trust in content provided by aggregated and syndicated services. Public key infrastructures, digital signatures, and reputation services are well established but hard to understand and to handle for the layperson. We propose an efficient and user-friendly security architecture based on the popular tagging paradigm that connects user-defined tags with security policies, rules, and social network information to ensure access control, data integrity, and confidence also in derived and syndicated data.
Matthias Quasthoff, Harald Sack, Christoph Meinel
ICIW2
2007 Why HTTPS Is Not Enough - A Signature-Based Architecture for Trusted Content on the Social Web
abstract
Easy to use, interactive web applications accumulating data from heterogeneous sources represent a recent trend on the World Wide Web, referred to as the Social Web. There however, security standards are often disregarded in favor of interface design or brand new features. This prevents the new services from gaining ground in the enterprise, in medical or e-government environments. We propose the deployment of XML Digital Signatures on web content and demonstrate how an architecture enabling for various security properties would look like. The solution proposed will benefit from the research on security engineering in Service-Oriented Architectures and thus allows for an in-depth analysis on the results.
Matthias Quasthoff, Harald Sack, Christoph Meinel
Web Intelligence2
2006 SOGOS - A Distributed Meta Level Architecture for the Self-Organizing Grid of Services
abstract
Handling highly dynamic scenarios as they arise in emergency situations requires lots of semantic information about the situation and an extremely flexible, selforganizing IT infrastructure that provides services that can be used to manage the situation. We show that a distributed meta level architecture is particularly suited for the implementation of such a self-organizing grid of services. This architecture (SOGOS) distinguishes between an object level and a meta level. The middleware processes of the grid are running on the object level. The meta level defines an explicitly and declaratively represented dynamic meta model that provides the semantics for the object level processes. Additionally, this level runs processes that plan, supervise and control mobile agents on the object level. The levels are linked together by reflection processes that ensure that relevant changes on the object level are reflected in the meta model and vice versa. The corresponding reflection principles provide the basis for the implementation of the selforganizing mechanisms that govern the overall system.
Clemens Beckstein, Peter Dittrich, Christian Erfurth, Dietmar Fey, Birgitta König-Ries, Martin Mundhenk, Harald Sack
MDM7