Marieke van Erp

dblp:81/5832 · also Maria Godefrida Jacoba van Erp · DBLP profile ↗
← Back
21ranked-venue papers in the field
6as first author
11since 2021 · last 2025
0000-0001-9195-8203ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 20 (6 first)Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 Detecting Changing Culinary Trends Through Historical Recipes
abstract
Culinary trends evolve in response to social, economic, and cultural influences, reflecting broader historical transformations. We present an exploration into Dutch culinary trends from 1910 to 1995 by analysing recipes from housekeeping school cookbooks and newspaper recipe collections. Using computational techniques, we extract and examine ingredient frequency, recipe complexity, and shifts in recipe categories to identify trends in Dutch cuisine from a quantitative point of view. Additionally, we experimented with Large Language Models (LLMs) to structure and extract recipes’ features, demonstrating their potential for historical recipe parsing.
Gauri Bhagwat, Marieke van Erp, Teresa Paccosi, Rik Hoekstra
LDK2
2025 Tracing Organisation Evolution in Wikidata
abstract
Entities change over time, and while information about entity change is contained in knowledge graphs (KGs), it is often not stated explicitly. This makes KGs less useful for investigating entities over time, or downstream tasks such as historical entity linking. In this paper, we present an approach and experiments that make explicit entity change in Wikidata. Our contributions are a mapping between an existing change ontology and Wikidata properties to identify types of change, and a dataset of entities with explicit evolution information and analytics on this dataset.
Marieke van Erp, Vera Provatorova
LDK1
2025 Old Reviews, New Aspects: Aspect Based Sentiment Analysis and Entity Typing for Book Reviews with LLMs
abstract
This paper faces the problem of the limited availability of datasets for Aspect-Based Sentiment Analysis (ABSA) in the Cultural Heritage domain. Currently, the main datasets for ABSA are product or restaurant reviews. We expand this to book reviews. Our methodology employs an LLM to maintain domain relevance while preserving the linguistic authenticity and natural variations found in genuine reviews. Entity types are annotated through the tool Text2AMR2FRED and evaluated manually. Additionally, we finetuned Llama 3.1 8B as a baseline model that not only performs ABSA, but also performs Entity Typing (ET) with a set of classes from DOLCE foundational ontology, enabling precise categorization of target aspects within book reviews. We present three key contributions as a step forward expanding ABSA: 1) a semi-synthetic set of book reviews, 2) an evaluation of Llama-3-1-Instruct 8B on the ABSA task, and 3) a fine-tuned version of Llama-3-1-Instruct 8B for ABSA.
Andrea Schimmenti, Stefano De Giorgis, Fabio Vitali, Marieke van Erp
LDK4
2023 A Knowledge Graph of Contentious Terminology for Inclusive Representation of Cultural Heritage
Andrei Nesterov, Laura Hollink, Marieke van Erp, Jacco van Ossenbruggen
ESWC3
2023 Unflattening Knowledge Graphs
abstract
Large general-purpose knowledge graphs (KGs) are a critical component for knowledge-driven applications. However, most KGs represent only a limited view of the entities and concepts they describe. The concept coffee can, for example, refer to the plant that yields coffee seeds, the beverage ‘coffee’, and the activity of drinking the beverage. Moreover, it has a long history that is deeply connected to colonialism and status. All of these notions are an intricate part of national identities, have changed dramatically over time, and connect to many different narratives with different opinions on them. This complexity is not captured in current KGs. In this vision paper, I present the three crucial challenges for unflattening knowledge graphs and directions for future work.
Marieke van Erp
K-CAP1
2023 Contextual Profiling of Charged Terms in Historical Newspapers
Ryan Brate, Marieke van Erp, Antal van den Bosch
LDK2
2022 Capturing the Semantics of Smell: The Odeuropa Data Model for Olfactory Heritage Information
Pasquale Lisena, Daniel Schwabe 0001, Marieke van Erp, Raphaël Troncy, William Tullett, Inger Leemans, Lizzie Marx, Sofia Colette Ehrich
ESWC3
2021 A Polyvocal and Contextualised Semantic Web
Marieke van Erp, Victor de Boer
ESWC1
2021 Capturing Contentiousness: Constructing the Contentious Terms in Context Corpus
abstract
Recent initiatives by cultural heritage institutions in addressing outdated and offensive language used in their collections demonstrate the need for further understanding into when terms are problematic or contentious. This paper presents an annotated dataset of 2,715 unique samples of terms in context, drawn from a historical newspaper archive, collating 21,800 annotations of contentiousness from expert and crowd workers. We describe the contents of the corpus by analysing inter-rater agreement and differences between experts and crowd workers. In addition, we demonstrate the potential of the corpus for automated detection of contentiousness. We show that a simple classifier applied to the embedding representation of a target word provides a better than baseline performance in predicting contentiousness. We find that the term itself and the context play a role in whether a term is considered contentious.
Ryan Brate, Andrei Nesterov, Valentin Vogelmann, Jacco van Ossenbruggen, Laura Hollink, Marieke van Erp
K-CAP6
2021 Marriage is a Peach and a Chalice: Modelling Cultural Symbolism on the Semantic Web
abstract
In this work, we fill the gap in the Semantic Web in the context of Cultural Symbolism. Building upon earlier work in \citesartini_towards_2021, we introduce the Simulation Ontology, an ontology that models the background knowledge of symbolic meanings, developed by combining the concepts taken from the authoritative theory of Simulacra and Simulations of Jean Baudrillard with symbolic structures and content taken from "Symbolism: a Comprehensive Dictionary'' by Steven Olderr. We re-engineered the symbolic knowledge already present in heterogeneous resources by converting it into our ontology schema to create HyperReal, the first knowledge graph completely dedicated to cultural symbolism. A first experiment run on the knowledge graph is presented to show the potential of quantitative research on symbolism.
Bruno Sartini, Marieke van Erp, Aldo Gangemi
K-CAP2
2021 The Wind in Our Sails: Developing a Reusable and Maintainable Dutch Maritime History Knowledge Graph
abstract
Digital sources are more prevalent than ever but effectively using them can be challenging. One core challenge is that digitized sources are often distributed, thus forcing researchers to spend time collecting, interpreting, and aligning different sources. A knowledge graph can accelerate research by providing a single connected source of truth that humans and machines can query. During two design-test cycles, we convert four data sets from the historical maritime domain into a knowledge graph. The focus during these cycles is on creating a sustainable and usable approach that can be adopted in other linked data conversion efforts. Furthermore, our knowledge graph is available for maritime historians and other interested users to investigate the daily business of the Dutch East India Company through a unified portal.
Stijn Schouten, Victor de Boer, Lodewijk Petram, Marieke van Erp
K-CAP4
2019 A Proposal for a Two-Way Journey on Validating Locations in Unstructured and Structured Data
abstract
The Web of Data has grown explosively over the past few years, and as with any dataset, there are bound to be invalid statements in the data, as well as gaps. Natural Language Processing (NLP) is gaining interest to fill gaps in data by transforming (unstructured) text into structured data. However, there is currently a fundamental mismatch in approaches between Linked Data and NLP as the latter is often based on statistical methods, and the former on explicitly modelling knowledge. However, these fields can strengthen each other by joining forces. In this position paper, we argue that using linked data to validate the output of an NLP system, and using textual data to validate Linked Open Data (LOD) cloud statements is a promising research avenue. We illustrate our proposal with a proof of concept on a corpus of historical travel stories.
Ilkcan Keles, Omar Qawasmeh, Tabea Tietz, Ludovica Marinucci, Roberto Reda, Marieke van Erp
LDK6
2018 Slicing and Dicing a Newspaper Corpus for Historical Ecology Research
Marieke van Erp, Jesse de Does, Katrien Depuydt, Rob Lenders, Thomas van Goethem
EKAW1
2018 Constructing a Recipe Web from Historical Newspapers
Marieke van Erp, Melvin Wevers, Hugo C. Huurdeman
ISWC (1)1
2017 Multilingual Fine-Grained Entity Typing
Marieke van Erp, Piek Vossen
LDK1
2017 Hunger for Contextual Knowledge and a Road Map to Intelligent Entity Linking
Filip Ilievski, Piek Vossen, Marieke van Erp
LDK3
2016 LOTUS: Adaptive Text Search for Big Linked Data
Filip Ilievski, Wouter Beek, Marieke van Erp, Laurens Rietveld, Stefan Schlobach
ESWC3
2016 Building event-centric knowledge graphs from news
Marco Rospocher, Marieke van Erp, Piek Vossen, Antske Fokkens, Itziar Aldabe, German Rigau, Aitor Soroa, Thomas Ploeger, Tessel Bogaard
J. Web Semant.2
2015 Analysis of named entity recognition and linking for tweets
Leon Derczynski, Diana Maynard, Giuseppe Rizzo 0002, Marieke van Erp, Genevieve Gorrell, Raphaël Troncy, Johann Petrak, Kalina Bontcheva
Inf. Process. Manag.4
2011 Linked open piracy
abstract
There is an abundance of semi-structured reports on events being written and made available on the World Wide Web on a daily basis. These reports are primarily meant for human use. A recent movement is the addition of RDF metadata to make automatic processing by computers easier. A fine example of this movement is the Open Government Data initiative which, by adding RDF meta-data to spreadsheets and textual reports, strives to speed up the creation of geographical mashups and visual analytics applications. In this paper, we present a new Open Linked Data RDF dataset and a method for automatically adding such RDF metadata to semi-structured reports. We showcase our method on piracy attack reports issued on the web by the International Chamber of Commerce's International Maritime Bureau (ICC-CCS IMB). We create a Semantic Web representation with the Simple Event Model (SEM) from screen scrapes of the ICC-CCS website. We show how the event layer makes it possible to easily analyze and visualize the aggregated reports to answer domain questions. Our pipeline includes conversion of the reports to RDF, linking their parts to external resources from the Linked Open Data cloud and exposing them to the Web through a ClioPatria web server that hosts the RDF.
Willem Robert van Hage, Véronique Malaisé, Marieke van Erp, Guus Schreiber
K-CAP3
2011 Hacking history via event extraction
abstract
Within cultural heritage collections, objects are often grounded in a particular historical setting. This setting can currently not be made explicit, as structured descriptions of events are either missing or not marked up explicitly. This paper reports a study on automatic extraction of an historical event thesaurus from unstructured texts. We show how this preliminary thesaurus accommodates event- and object-driven search and browsing of two cultural heritage collections.
Roxane Segers, Marieke van Erp, Lourens van der Meij, Lora Aroyo, Jacco van Ossenbruggen, Guus Schreiber, Bob J. Wielinga, Johan Oomen, Geertje Jacobs
K-CAP2