EDBT 2026 Demo / reviewers in the wild / expert
Jacco van Ossenbruggen
dblp:v/JaccovanOssenbruggen
· DBLP profile ↗
34ranked-venue papers in the field
3as first author
10since 2021 · last 2025
0000-0002-7748-4715ORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 18Information Retrieval & Web Search · 16 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bridging Data Gaps: Harnessing Semantic Associations for Knowledge Discovery in Colonial HeritageabstractCultural heritage data, particularly from colonial contexts, frequently presents an incomplete and biased view, reflecting historical institutional priorities more than contemporary knowledge requirements. Consequently, knowledge graphs derived from these records often contain incomplete, fragmented, and skewed data, including absent attributes or values, missing semantic links, and under-represented perspectives. This work addresses the challenge of knowledge discovery under such limitations, presenting a real-world case study on the provenance research of colonial cultural heritage. We present a task-aware design method for building a tool to facilitate this process. The design approach of this application is rooted in a Knowledge Discovery in Database (KDD) framework and is particularly novel due to its formalisation and operationalisation of three distinct types of semantic association: explicit, abstract, and implicit. These semantic associations, grounded in domain interpretation, are crucial for bridging data gaps where user information needs cannot be directly met by existing data. We further demonstrate how these associations can be effectively communicated through user interface components, enabling users to infer new knowledge. We evaluated the resultant application through a user study among domain experts to assess its efficacy. The evaluation confirms the effectiveness of the tool in enabling new knowledge discovery and reveals opportunities to improve the representation of the underlying data, as users could successfully infer insights even when information was missing or poorly captured in the original data sets. Sarah Binta Alam Shoilee, Victor de Boer, Annastiina Ahola, Heikki Rantala, Eero Hyvönen, Jacco van Ossenbruggen, Susan Legêne |
K-CAP | 6 |
| 2024 | A Framework for Evaluating Entity Alignment Impact on Downstream Knowledge Discovery
Sarah Binta Alam Shoilee, Victor de Boer, Jacco van Ossenbruggen |
EKAW | 3 |
| 2024 | How Contentious Terms About People and Cultures are Used in Linked Open DataabstractWeb resources in linked open data (LOD) are comprehensible to humans through literal textual values attached to them, such as labels, notes, or comments. Word choices in literals may not always be neutral. When culturally stereotyping terminology is used in literals, they may appear as offensive to users in interfaces and propagate stereotypes to algorithms trained on them. We study how frequently and in which literals contentious terms about people and cultures occur in LOD and whether there are attempts to mark the usage of such terms. For our analysis, we reuse English and Dutch terms from a knowledge graph that provides opinions of experts from the cultural heritage domain about terms' contentiousness. We inspect occurrences of these terms in four widely used datasets: Wikidata, The Getty Art & Architecture Thesaurus, Princeton WordNet, and Open Dutch WordNet. Some terms are ambiguous and contentious only in particular senses. Applying word sense disambiguation, we generate a set of literals relevant to our analysis. We found that contentious terms frequently appear in descriptive and labelling literals, such as preferred labels that are usually displayed in interfaces and used for indexing. In some cases, LOD contributors mark contentious terms with words and phrases in literals (implicit markers) or properties linked to resources (explicit markers). However, such marking is rare and non-consistent in all datasets. Our quantitative and qualitative insights could be helpful in developing more systematic approaches to address the propagation of stereotypes via LOD. Andrei Nesterov, Laura Hollink, Jacco van Ossenbruggen |
WWW | 3 |
| 2024 | Reproducing Popularity Bias in Recommendation: The Effect of Evaluation StrategiesabstractThe extent to which popularity bias is propagated by media recommender systems is a current topic within the community, as is the uneven propagation among users with varying interests for niche items. Recent work focused on exactly this topic, with movies being the domain of interest. Later on, two different research teams reproduced the methodology in the domains of music and books, respectively. The results across the different domains diverge. In this paper, we reproduce the three studies and identify four aspects that are relevant in investigating the differences in results: data, algorithms, division of users in groups and evaluation strategy. We run a set of experiments in which we measure general popularity bias propagation and unfair treatment of certain users with various combinations of these aspects. We conclude that all aspects account to some degree for the divergence in results, and should be carefully considered in future studies. Further, we find that the divergence in findings can be in large part attributed to the choice of evaluation strategy. Savvina Daniil, Mirjam Cuper, Cynthia C. S. Liem, Jacco van Ossenbruggen, Laura Hollink |
Trans. Recomm. Syst. | 4 |
| 2023 | A Knowledge Graph of Contentious Terminology for Inclusive Representation of Cultural Heritage
Andrei Nesterov, Laura Hollink, Marieke van Erp, Jacco van Ossenbruggen |
ESWC | 4 |
| 2023 | The Role of Serendipity in User-Curated Music PlaylistsabstractIn this paper, we study the role of serendipity in music playlists. Serendipity is an important construct in recommendations, and finding an indicator of serendipity in a user-created playlist can facilitate the recommendation task. In particular, we want to know how the serendipity level of playlists is affected by the creator’s ability and by the context they are created. To do so, we (1) measure the serendipity level of music playlists using a previously established Linked Open Data-based approach, (2) assess whether the ability of the creator of the playlists has an effect on the serendipity level, and (3) assess whether different contexts facilitate a higher or lower serendipity level of playlists. The serendipity level of playlists is calculated with the cosine distance between Linked Open Data Paths that connect the songs contained in the playlist. The ability of the creator to generate serendipitous recommendations is estimated by measuring his/her coping potential and assessing the genre diversity of listening history. We instrument a study using a Spotify playlists dataset. Previous results in different contexts suggest that the coping potential is a good proxy for the curiosity level of a person, and, in turn, for the diversified knowledge this person has. Our analyses confirm these findings also in the music context: we find that playlist creators with higher coping potential have a more diversified knowledge. They create a higher number of playlists that span across multiple contexts and genres. Conversely, a lower copying potential implies a lower number of less coherent playlists. Valentina Maccatrozzo, Tobias Kuhn, Davide Ceolin, Jacco van Ossenbruggen |
K-CAP | 4 |
| 2023 | Advancing data sharing and reusability for restricted access data on the Web: introducing the DataSet-Variable OntologyabstractIn response to the increasing volume of research data being generated, more and more data portals have been designed to facilitate data findability and accessibility. However, a significant portion of this data remains confidential or restricted due to its sensitive nature, such as patient data or census microdata. While maintaining confidentiality prohibits its public release, the emergence of portals supporting rich metadata can help enable researchers to at least discover the existence of restricted access data, empowering them to assess the suitability of the data before requesting access. Margherita Martorana, Tobias Kuhn, Ronny Siebes, Jacco van Ossenbruggen |
K-CAP | 4 |
| 2021 | Comparing Methods for Finding Search Sessions on a Specified Topic: A Double Case Study
Tessel Bogaard, Aysenur Bilgin, Jan Wielemaker, Laura Hollink, Kees Ribbens, Jacco van Ossenbruggen |
TPDL | 6 |
| 2021 | Capturing Contentiousness: Constructing the Contentious Terms in Context CorpusabstractRecent initiatives by cultural heritage institutions in addressing outdated and offensive language used in their collections demonstrate the need for further understanding into when terms are problematic or contentious. This paper presents an annotated dataset of 2,715 unique samples of terms in context, drawn from a historical newspaper archive, collating 21,800 annotations of contentiousness from expert and crowd workers. We describe the contents of the corpus by analysing inter-rater agreement and differences between experts and crowd workers. In addition, we demonstrate the potential of the corpus for automated detection of contentiousness. We show that a simple classifier applied to the embedding representation of a target word provides a better than baseline performance in predicting contentiousness. We find that the term itself and the context play a role in whether a term is considered contentious. Ryan Brate, Andrei Nesterov, Valentin Vogelmann, Jacco van Ossenbruggen, Laura Hollink, Marieke van Erp |
K-CAP | 4 |
| 2021 | Expressing High-Level Scientific Claims with Formal SemanticsabstractThe use of semantic technologies is gaining significant traction in science communication with a wide array of applications in disciplines including the life sciences, computer science, and the social sciences. Languages like RDF, OWL, and other formalisms based on formal logic are applied to make scientific knowledge accessible not only to human readers but also to automated systems. These approaches have mostly focused on the structure of scientific publications themselves, on the used scientific methods and equipment, or on the structure of the used datasets. The core claims or hypotheses of scientific work have only been covered in a shallow manner, such as by linking mentioned entities to established identifiers. In this research, we therefore want to find out whether we can use existing semantic formalisms to fully express the content of high-level scientific claims using formal semantics in a systematic way. Analyzing the main claims from a sample of scientific articles from all disciplines, we find that their semantics are more complex than what a straight-forward application of formalisms like RDF or OWL account for, but we managed to elicit a clear semantic pattern which we call the "super-pattern''. We show here how the instantiation of the five slots of this super-pattern leads to a strictly defined statement in higher-order logic. We successfully applied this super-pattern to an enlarged sample of scientific claims. We show that knowledge representation experts, when instructed to independently instantiate the super-pattern with given scientific claims, show a high degree of consistency and convergence given the complexity of the task and the subject. These results therefore open the door on the longer run for allowing researchers to express their high-level scientific findings in a manner they can be automatically interpreted. This in turn will allow for automated consistency checking, question answering, aggregation, and much more. Cristina-Iulia Bucur, Tobias Kuhn, Davide Ceolin, Jacco van Ossenbruggen |
K-CAP | 4 |
| 2020 | Understanding User Behavior in Digital Libraries Using the MAGUS Session Visualization Tool
Tessel Bogaard, Jan Wielemaker, Laura Hollink, Lynda Hardman, Jacco van Ossenbruggen |
TPDL | 5 |
| 2019 | Searching for Old News: User Interests and Behavior within a National CollectionabstractModeling user interests helps to improve system support or refine recommendations in Interactive Information Retrieval. The aim of this study is to identify user interests in different parts of an online collection and investigate the related search behavior. To do this, we propose to use the metadata of selected facets and clicked documents as features for clustering sessions identified in user logs. We evaluate the session clusters by measuring their stability over a six-month period. Tessel Bogaard, Laura Hollink, Jan Wielemaker, Lynda Hardman, Jacco van Ossenbruggen |
CHIIR | 5 |
| 2017 | Time-based tags for fiction movies: comparing experts to novices using a video labeling gameabstractThe cultural heritage sector has embraced social tagging as a way to increase both access to online content and to engage users with their digital collections. In this article, we build on two current lines of research. (a) We use Waisda?, an existing labeling game, to add time‐based annotations to content. (b) In this context, we investigate the role of experts in human‐based computation (nichesourcing). We report on a small‐scale experiment in which we applied Waisda? to content from film archives. We study the differences in the type of time‐based tags between experts and novices for film clips in a crowdsourcing setting. The findings show high similarity in the number and type of tags (mostly factual). In the less frequent tags, however, experts used more domain‐specific terms. We conclude that competitive games are not suited to elicit real expert‐level descriptions. We also confirm that providing guidelines, based on conceptual frameworks that are more suited to moving images in a time‐based fashion, could result in increasing the quality of the tags, thus allowing for creating more tag‐based innovative services for online audiovisual heritage. Liliana Melgar, Michiel Hildebrand, Victor de Boer, Jacco van Ossenbruggen |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2016 | Comparing Topic Coverage in Breadth-First and Depth-First Crawls Using Anchor Texts
Thaer Samar, Myriam C. Traub, Jacco van Ossenbruggen, Arjen P. de Vries |
TPDL | 3 |
| 2015 | Supporting Exploration of Historical Perspectives Across Collections
Daan Odijk, Cristina Garbacea, Thomas Schoegje, Laura Hollink, Victor de Boer, Kees Ribbens, Jacco van Ossenbruggen |
TPDL | 7 |
| 2015 | Impact Analysis of OCR Quality on Research Tasks in Digital Archives
Myriam C. Traub, Jacco van Ossenbruggen, Lynda Hardman |
TPDL | 2 |
| 2014 | Measuring the Effectiveness of Gamesourcing Expert Oil Painting Annotations
Myriam C. Traub, Jacco van Ossenbruggen, Jiyin He, Lynda Hardman |
ECIR | 2 |
| 2013 | An Evaluation of Labelling-Game Data for Video Retrieval
Riste Gligorov, Michiel Hildebrand, Jacco van Ossenbruggen, Lora Aroyo, Guus Schreiber |
ECIR | 3 |
| 2012 | Supporting Linked Data Production for Cultural Heritage Institutes: The Amsterdam Museum Case Study
Victor de Boer, Jan Wielemaker, Judith van Gent, Michiel Hildebrand, Antoine Isaac, Jacco van Ossenbruggen, Guus Schreiber |
ESWC | 6 |
| 2011 | Interactive Vocabulary Alignment
Jacco van Ossenbruggen, Michiel Hildebrand, Victor de Boer |
TPDL | 1 |
| 2011 | On the role of user-generated metadata in audio visual collectionsabstractRecently, various crowdsourcing initiatives showed that targeted efforts of \nuser communities result in massive amounts of tags. For example, the \nNetherlands Institute for Sound and Vision collected a large number of tags \nwith the video labeling game \\emph{Waisda?}. To successfully utilize these \ntags, a better understanding of their characteristics is required. \nThe goal of this paper is twofold: (i) to investigate the vocabulary that \nusers employ when describing videos and compare it to the vocabularies used by \nprofessionals; and (ii) to establish which aspects of the video are typically \ndescribed and what type of tags are used for this. We report on an analysis of \nthe tags collected with \\emph{Waisda?}. With respect to the first goal, we \ncompared the the tags with a typical domain thesaurus used by professionals, \nas well as with a more general vocabulary. With respect to the second goal, we \ncompare the tags to the video subtitles to determine how many tags are derived \nfrom the audio signal. In addition, we perform a qualitative study in which a \ntag sample is interpreted in terms of an existing annotation classification \nframework. The results suggest that the tags complement the metadata provided \nby professional cataloguers, the tags describe both the audio and the visual \naspects of the video, and the users primarily describe objects in the video \nusing general descriptions. Riste Gligorov, Michiel Hildebrand, Jacco van Ossenbruggen, Guus Schreiber, Lora Aroyo |
K-CAP | 3 |
| 2011 | Hacking history via event extractionabstractWithin cultural heritage collections, objects are often grounded in a particular historical setting. This setting can currently not be made explicit, as structured descriptions of events are either missing or not marked up explicitly. This paper reports a study on automatic extraction of an historical event thesaurus from unstructured texts. We show how this preliminary thesaurus accommodates event- and object-driven search and browsing of two cultural heritage collections. Roxane Segers, Marieke van Erp, Lourens van der Meij, Lora Aroyo, Jacco van Ossenbruggen, Guus Schreiber, Bob J. Wielinga, Johan Oomen, Geertje Jacobs |
K-CAP | 5 |
| 2011 | Let's agree to disagree: on the evaluation of vocabulary alignmentabstractGold standard mappings created by experts are at the core of alignment evaluation. At the same time, the process of manual evaluation is rarely discussed. While the practice of having multiple raters evaluate results is accepted, their level of agreement is often not measured. In this paper we describe three experiments in manual evaluation and study the way different raters evaluate mappings. We used alignments generated using different techniques and between vocabularies of different type. In each experiment, five raters evaluated alignments and talked through their decisions using the think aloud method. In all three experiments we found that inter-rater agreement was low and analyzed our data to find the reasons for it. Our analysis shows which variables can be controlled to affect the level of agreement including the mapping relations, the evaluation guidelines and the background of the raters. On the other hand, differences in the perception of raters, and the complexity of the relations between often ill-defined natural language concepts remain inherent sources of disagreement. Our results indicate that the manual evaluation of ontology alignments is by no means an easy task and that the ontology alignment community should be careful in the construction and use of reference alignments. Anna Tordai, Jacco van Ossenbruggen, Guus Schreiber, Bob J. Wielinga |
K-CAP | 2 |
| 2010 | Aligning Large SKOS-Like Vocabularies: Two Case Studies
Anna Tordai, Jacco van Ossenbruggen, Guus Schreiber, Bob J. Wielinga |
ESWC (1) | 2 |
| 2009 | Organizing Suggestions in Autocompletion Interfaces
Alia Amin, Michiel Hildebrand, Jacco van Ossenbruggen, Vanessa Evers, Lynda Hardman |
ECIR | 3 |
| 2009 | Combining vocabulary alignment techniquesabstractIdentifying alignments between vocabularies has become a central knowledge engineering activity. A plethora of alignment techniques has been developed over the past years. In this paper we present a case study in which we examine and evaluate the practical use of three typical alignment techniques. The study involves the alignment of two vocabularies used in a semantic-search engine for cultural-heritage objects. We show that a sequence can be beneficial. The case study gives insight into evaluation issues, such as techniques for identification of false positives. We see this work as a step to a badly-needed methodology for alignment. Anna Tordai, Jacco van Ossenbruggen, Guus Schreiber |
K-CAP | 2 |
| 2008 | Thesaurus-Based Search in Large Heterogeneous Collections
Jan Wielemaker, Michiel Hildebrand, Jacco van Ossenbruggen, Guus Schreiber |
ISWC | 3 |
| 2008 | Semantic annotation and search of cultural-heritage collections: The MultimediaN E-Culture demonstrator
Guus Schreiber, Alia Amin, Lora Aroyo, Mark van Assem, Victor de Boer, Lynda Hardman, Michiel Hildebrand, Borys Omelayenko, Jacco van Ossenbruggen, Anna Tordai, Jan Wielemaker, Bob J. Wielinga |
J. Web Semant. | 9 |
| 2006 | /facet: A Browser for Heterogeneous Semantic Web Repositories
Michiel Hildebrand, Jacco van Ossenbruggen, Lynda Hardman |
ISWC | 2 |
| 2006 | MultimediaN E-Culture Demonstrator
Guus Schreiber, Alia Amin, Mark van Assem, Victor de Boer, Lynda Hardman, Michiel Hildebrand, Laura Hollink, Zhisheng Huang, Janneke van Kersen, Marco de Niet, Borys Omelayenko, Jacco van Ossenbruggen, Ronny Siebes, Jos Taekema, Jan Wielemaker, Bob J. Wielinga |
ISWC | 12 |
| 2005 | Making RDF presentable: integrated global and local semantic Web browsingabstractThis paper discusses generating document structure from annotated media repositories in a domain-independent manner. This approaches the vision of a universal RDF browser. We start by applying the search-and-browse paradigm established for the WWW to RDF presentation. Furthermore, this paper adds to this paradigm the clustering-based derivation of document structure from search returns, providing simple but domain-independent hypermedia generation from RDF stores. While such generated presentations hardly meet the standards of those written by humans, they provide quick access to media repositories when the required document has not yet been written. The resulting system allows a user to specify a topic for which it generates a hypermedia document providing guided navigation through virtually any RDF repository. The impact for content providers is that as soon as one adds new media items and their annotations to a repository, they become immediately available for automatic integration into subsequently requested presentations. Lloyd Rutledge, Jacco van Ossenbruggen, Lynda Hardman |
WWW | 2 |
| 2003 | Towards Ontology-Driven Discourse: From Semantic Graphs to Multimedia Presentations
Joost Geurts, Stefano Bocconi, Jacco van Ossenbruggen, Lynda Hardman |
ISWC | 3 |
| 2003 | Towards a multimedia formatting vocabularyabstractTime-based, media-centric Web presentations can be described declaratively in the XML world through the development of languages such as SMIL. It is difficult, however, to fully integrate them in a complete document transformation processing chain. In order to achieve the desired processing of data-driven, time-based, media-centric presentations, the text-flow based formatting vocabularies used by style languages such as XSL, CSS and DSSSL need to be extended. The paper presents a selection of use cases which are used to derive a list of requirements for a multimedia style and transformation formatting vocabulary. The boundaries of applicability of existing text-based formatting models for media-centric transformations are analyzed. The paper then discusses the advantages and disadvantages of a fully-fledged time-based multimedia formatting model. Finally, the discussion is illustrated by describing the key properties of the example multimedia formatting vocabulary currently implemented in the back-end of our Cuypers multimedia transformation engine. Jacco van Ossenbruggen, Lynda Hardman, Joost Geurts, Lloyd Rutledge |
WWW | 1 |
| 2001 | Towards second and third generation web-based multimediaabstractFirst generation Web-content encodes information in handwritten (HTML) Web pages. Second generation Web content generates HTML pages on demand, e.g. by filling in templates with content retrieved dynamically from a database or transformation of structured documents using style sheets (e.g. XSLT). Third generation Web pages will make use of rich markup (e.g. XML) along with metadata (e.g. RDF) schemes to make the content not only machine readable but also machine processable — a necessary pre-requisite to the Semantic Web. While text-based content on the Web is already rapidly approaching the third generation, multimedia content is still trying to catch up with second generation techniques. Multimedia document processing has a number of fundamentally different requirements from text which make it more difficult to incorporate within the document processing chain. In particular, multimedia transformation uses different document and presentation abstractions, its formatting rules cannot be based on text-flow, it requires feedback from the formatting back-end and is hard to describe in the functional style of current style languages. We state the requirements for second generation processing of multimedia and describe how these have been incorporated in our prototype multimedia document transformation environment, Cuypers. The system overcomes a number of the restrictions of the text-flow based toolsets by integrating a number of conceptually distinct processing steps in a single runtime execution environment. We describe the need for these different processing steps and describe them in turn (semantic structure, communicative device, qualitative constraints, quantitative constraints, finalform presentation), and illustrate our approach by means of an example. We conclude by discussing the models and techniques required for the creation of third generation multimedia content. Jacco van Ossenbruggen, Joost Geurts, Frank Cornelissen, Lynda Hardman, Lloyd Rutledge |
WWW | 1 |