Elena Simperl

dblp:p/ElenaPaslaruBontasSimperl · also Elena Paslaru Bontas, Elena Paslaru Bontas Simperl · DBLP profile ↗
← Back
57ranked-venue papers in the field
5as first author
13since 2021 · last 2026
0000-0003-1722-947XORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 38 (4 first)Information Retrieval & Web Search · 10Data Mining & Knowledge Discovery · 5Database Systems & Data Management · 2 (1 first)Other / Interdisciplinary · 2
YearPublicationVenuePosition
2026 HERTy-Wiki: A Benchmark for Hierarchical Entity Reasoning and Typing in Wikidata
abstract
Entity typing is central to knowledge engineering, and as large language models (LLMs) are increasingly used to support knowledge graph construction, it becomes essential to assess how reliably they can perform this task. We present HERTy-Wiki (Hierarchical Entity Reasoning and Typing in Wikidata) - a human-verified benchmark for evaluating hierarchical reasoning in LLMs for knowledge graph entity typing (KGET). Unlike existing probing benchmarks that test factual recall, HERTy-Wiki evaluates whether models can select the most specific valid type from limited contextual evidence, reflecting realistic Wikidata scenarios, where editors often work with incomplete or noisy information. HERTy-Wiki comprises 8,767 multiple-choice questions derived from 3,776 Wikidata entities across five domains, with an optional multimodal extension. Each question requires models to discriminate between semantically related types within the same subclass hierarchy. We evaluate several state-of-the-art, reasoning-oriented LLMs under zero-shot, few-shot, and chain-of-thought prompting settings and compare them against a no-context baseline to quantify model reliance on memorised priors. Across models, macro-F1 scores remain around 0.50–0.60, with only modest improvements over the no-context baseline. Our analysis indicates that the emerging hierarchical reasoning abilities of models are overshadowed by strong dependence on memorised priors. These findings highlight fundamental challenges in using LLMs for reasoning in knowledge engineering pipelines, particularly in rapidly evolving or specialised domains with limited pre-training coverage. HERTy-Wiki provides a human-verified benchmark for developing and evaluating genuine reasoning-driven approaches to KGET.
Nicole Obretincheva, Helen Yannakoudakis, Elena Simperl
ESWC (2)3
2026 OntoChat Assistant for User Story Generation in Ontology Engineering
abstract
An ontology is a formal, explicit specification of a shared conceptualisation, which can be combined with problem-solving methods and reasoning functionality to develop high-quality technology and application systems efficiently. Ontology engineering typically involves extensive manual effort to elicit intended use cases (user stories) from users for the target ontology-based systems. Recent studies have demonstrated the positive potential of large language model-based conversational agents in supporting user story generation in OE. However, we argue that we are not leveraging LLM to its fullest potential by not supporting users in formulating effective prompts. To address this, we identify the prompt guidance users need during user story generation workflows by conducting a formative study (N = 10) using participatory prompting. We demonstrate its usefulness through the design and development of the OntoChat LLM-based system for OE, as well as a user evaluation with knowledge engineers (N = 24). To our knowledge, this is the first work to design and validate a prompt guidance framework that helps users leverage LLM to its fullest potential to generate effective requirements for ontology development. This advances how we interact with LLM for requirements elicitation.
Yihang Zhao 0004, Anelia Kurteva, Albert Meroño-Peñuela, Elena Simperl
ACM Trans. Intell. Syst. Technol.4
2026 Knowledge prompting: How knowledge engineers use generative AI
abstract
Despite many advances in knowledge engineering (KE), challenges remain in areas such as engineering knowledge graphs (KGs) at scale, automating tasks, and keeping pace with evolving domain knowledge. KE has used NLP demonstrating notable advantages in knowledge-intensive tasks, but the most effective use of generative AI to support knowledge engineers across the KE activities is still in its infancy. To explore how generative AI may enhance KE and change existing KE practices, we conducted a multi-method study during a KE hackathon. We investigated participants’ views on the use of generative AI, the challenges they face, the skills they may need to integrate generative AI into their practices, and how they use generative AI responsibly. We found participants felt LLMs could indeed contribute to improving efficiency when engineering KGs, but presented increased challenges around the already complex issues of evaluating KE task success. We discovered prompting to be a useful but undervalued skill for knowledge engineers working with LLMs, and note that NLP skills may become more relevant across more roles in KE workflows. Integrating generative AI into KE tasks needs to be done with awareness of potential risks and harms. Given the limited ethical training most knowledge engineers receive, solutions such as our proposed ‘KG Cards’ based on Data Cards could be a useful guide for KG construction. Our findings can support designers of KE AI copilots, KE researchers, and practitioners using advanced AI to develop trustworthy applications, propose new methodologies for KE and operate new technologies responsibly.
Elisavet Koutsiana, Johanna Walker, Michelle Nwachukwu, Bohui Zhang, Albert Meroño-Peñuela, Elena Simperl
J. Web Semant.6
2026 Corrigendum to "Knowledge prompting: How knowledge engineers use generative AI" [Journal of Web Semantics 88 (2026) 100873]
Elisavet Koutsiana, Johanna Walker, Michelle Nwachukwu, Bohui Zhang, Albert Meroño-Peñuela, Elena Simperl
J. Web Semant.6
2025 Agreeing and disagreeing in collaborative knowledge graph construction: An analysis of Wikidata
abstract
In this work, we study disagreements in discussions around Wikidata, an online knowledge community that builds the data backend of Wikipedia. Discussions are essential in collaborative work as they can increase contributor performance and encourage the emergence of shared norms and practices. While disagreements can play a productive role in discussions, they can also lead to conflicts and controversies, which impact contributor’ well-being and their motivation to engage. We want to understand if and when such phenomena arise in Wikidata, using a mix of quantitative and qualitative analyses to identify the types of topics people disagree about, the most common patterns of interaction, and roles people play when arguing for or against an issue. We find that decisions to create Wikidata properties are much faster than those to delete properties and that more than half of controversial discussions do not lead to consensus. Our analysis suggests that Wikidata is an inclusive community, considering different opinions when making decisions, and that conflict and vandalism are rare in discussions. At the same time, while one-fourth of the editors participating in controversial discussions contribute legitimate and insightful opinions about Wikidata’s emerging issues, they respond with one or two posts and do not remain engaged in the discussions to reach consensus. Our work contributes to the analysis of collaborative KG construction with insights about communication and decision-making in projects, as well as with methodological directions and open datasets. We hope our findings will help managers and designers support community decision-making and improve discussion tools and practices.
Elisavet Koutsiana, Tushita Yadav, Nitisha Jain, Albert Meroño-Peñuela, Elena Simperl
J. Web Semant.5
2025 KG.GOV: Knowledge graphs as the backbone of data governance in AI
abstract
As (generative) Artificial Intelligence continues to evolve, so do the challenges associated with governing the data that powers it. Ensuring data quality, privacy, security, and ethical use become more and more challenging due to the increasing volume and variety of the data, the complexity of AI models, and the rapid pace of technological advancement. Knowledge graphs have the potential to play a significant role in enabling data governance in AI, as we move beyond their traditional use as data organisational systems. To address this, we present KG.GOV, a framework that positions KGs at a higher abstraction level within AI workflows, and enables them as a backbone of AI data governance. We illustrate the three dimensions of KG.GOV: modelling data, alternative representations, and describing behaviour; and describe the insights and challenges of three use cases implementing them: Croissant, a vocabulary to model and document ML datasets; WikiPrompts, a collaborative KG of prompts and prompt workflows to study their behaviour at scale; and Multimodal transformations, an approach for multimodal KGs harmonisation and completion aiming at broadening access to knowledge.
Albert Meroño-Peñuela, Elena Simperl, Anelia Kurteva, Ioannis Reklos
J. Web Semant.2
2024 RevOnt: Reverse engineering of competency questions from knowledge graphs via language models
abstract
The process of developing ontologies – a formal, explicit specification of a shared conceptualisation – is addressed by well-known methodologies. As for any engineering development, its fundamental basis is the collection of requirements, which includes the elicitation of competency questions. Competency questions are defined through interacting with domain and application experts or by investigating existing datasets that may be used to populate the ontology i.e. its knowledge graph. The rise in popularity and accessibility of knowledge graphs provides an opportunity to support this phase with automatic tools. In this work, we explore the possibility of extracting competency questions from a knowledge graph. This reverses the traditional workflow in which knowledge graphs are built from ontologies, which in turn are engineered from competency questions. We describe in detail RevOnt, an approach that extracts and abstracts triples from a knowledge graph, generates questions based on triple verbalisations, and filters the resulting questions to yield a meaningful set of competency questions; the WDV dataset. This approach is implemented utilising the Wikidata knowledge graph as a use case, and contributes a set of core competency questions from 20 domains present in the WDV dataset. To evaluate RevOnt, we contribute a new dataset of manually-annotated high-quality competency questions, and compare the extracted competency questions by calculating their BLEU score against the human references. The results for the abstraction and question generation components of the approach show good to high quality. Meanwhile, the accuracy of the filtering component is above 86%, which is comparable to the state-of-the-art classifications.
Fiorela Ciroku, Jacopo de Berardinis, Jongmo Kim, Albert Meroño-Peñuela, Valentina Presutti, Elena Simperl
J. Web Semant.6
2023 Qrowdsmith: Enhancing Paid Microtask Crowdsourcing with Gamification and Furtherance Incentives
abstract
Microtask crowdsourcing platforms are social intelligence systems in which volunteers, called crowdworkers, complete small, repetitive tasks in return for a small fee. Beyond payments, task requesters are considering non-monetary incentives such as points, badges, and other gamified elements to increase performance and improve crowdworker experience. In this article, we present Qrowdsmith, a platform for gamifying microtask crowdsourcing. To design the system, we explore empirically a range of gamified and financial incentives and analyse their impact on how efficient, effective, and reliable the results are. To maintain participation over time and save costs, we propose furtherance incentives, which are offered to crowdworkers to encourage additional contributions in addition to the fee agreed upfront. In a series of controlled experiments, we find that while gamification can work as furtherance incentives, it impacts negatively on crowdworkers’ performance, both in terms of the quantity and quality of work, as compared to a baseline where they can continue to contribute voluntarily. Gamified incentives are also less effective than paid bonus equivalents. Our results contribute to the understanding of how best to encourage engagement in microtask crowdsourcing activities and design better crowd intelligence systems.
Eddy Maddalena, Luis-Daniel Ibáñez, Neal Reeves, Elena Simperl
ACM Trans. Intell. Syst. Technol.4
2023 An analysis of discussions in collaborative knowledge engineering through the lens of Wikidata
Elisavet Koutsiana, Gabriel Maia Rocha Amaral, Neal Reeves, Albert Meroño-Peñuela, Elena Simperl
J. Web Semant.5
2022 A comparison of dataset search behaviour of internal versus search engine referred sessions
abstract
Dataset discovery is a first step for data-centric tasks, from data storytelling to labelling for supervised machine learning. Previous qualitative research suggests that people use two types of search affordances to find the data they need: they either go to a data portal that probably contains the data and search there; or they start on a regular web search engine, which sometimes returns results that are datasets. For the first type of search, prior works have analysed logs from different data portals to understand basic tenets of search behaviour such as query length or topics. In this paper, we advance the state of the art in dataset search behaviour with a comprehensive transaction log analysis study (n = 236441 sessions) of an international open data portal, in which we compare sessions straight on a data portal (internal searches) against sessions that land on a dataset or SERP (search engine result page) through a referral from a web search engine (external). Using dataset downloads as a proxy for successful searches, we find a statistically significant, though weak relationship between the use of keyword search and session type and between the use of search facets and session type (moderate). We also discover and discuss behavioural patterns and user profiles across session types.
Luis-Daniel Ibáñez, Elena Simperl
CHIIR2
2022 An Analysis of Content Gaps Versus User Needs in the Wikidata Knowledge Graph
David Abián, Albert Meroño-Peñuela, Elena Simperl
ISWC3
2022 WDV: A Broad Data Verbalisation Dataset Built from Wikidata
Gabriel Maia Rocha Amaral, Odinaldo Rodrigues, Elena Simperl
ISWC3
2021 Learning to Recommend Items to Wikidata Editors
Kholoud Alghamdi, Miaojing Shi, Elena Simperl
ISWC3
2020 Pie Chart or Pizza: Identifying Chart Types and Their Virality on Twitter
Pavlos Vougiouklis, Les Carr, Elena Simperl
ICWSM3
2020 Enhancing Public Procurement in the European Union Through Constructing and Exploiting an Integrated Knowledge Graph
Ahmet Soylu, Óscar Corcho, Brian Elvesæter, Carlos Badenes-Olmedo, Francisco Yedro Martínez, Matej Kovacic, Matej Posinkovic, Ian Makgill, Chris Taggart, Elena Simperl, Till C. Lech, Dumitru Roman
ISWC (2)10
2020 Mapping Points of Interest Through Street View Imagery and Paid Crowdsourcing
abstract
We present the Virtual City Explorer (VCE), an online crowdsourcing platform for the collection of rich geotagged information in urban environments. Compared to other volunteered geographic information approaches, which are constrained by the number and availability of mapping enthusiasts on the ground, the VCE uses digital street imagery to allow people to virtually explore a city from anywhere in the world, using a browser or a mobile phone. In addition, contributions in VCE are designed as paid microtasks—small jobs that can be carried out without any specific knowledge of the local area or previous mapping expertise in exchange for a fee. We tested the VCE in two cities to map points of interest (PoIs) in transport and mobility, using FigureEight to recruit participants. We were able to show that our platform enables crowdworkers to submit PoI location seamlessly, cover almost all of the tested areas, and discover several PoIs not reported by other approaches. This allows the VCE to complement existing approaches that leverage experts or grassroot communities.
Eddy Maddalena, Luis-Daniel Ibáñez, Elena Simperl
ACM Trans. Intell. Syst. Technol.3
2020 Dataset search: a survey
abstract
Generating value from data requires the ability to find, access and make sense of datasets. There are many efforts underway to encourage data sharing and reuse, from scientific publishers asking authors to submit data alongside manuscripts to data marketplaces, open data portals and data communities. Google recently beta-released a search service for datasets, which allows users to discover data stored in various online repositories via keyword queries. These developments foreshadow an emerging research field around dataset search or retrieval that broadly encompasses frameworks, methods and tools that help match a user data need against a collection of datasets. Here, we survey the state of the art of research and commercial systems and discuss what makes dataset search a field in its own right, with unique challenges and open questions. We look at approaches and implementations from related areas dataset search is drawing upon, including information retrieval, databases, entity-centric and tabular search in order to identify possible paths to tackle these questions as well as immediate next steps that will take the field forward.
Adriane Chapman, Elena Simperl, Laura Koesten, George Konstantinidis 0001, Luis-Daniel Ibáñez, Emilia Kacprzak, Paul Groth
VLDB J.2
2019 Ranking Knowledge Graphs By Capturing Knowledge about Languages and Labels
abstract
Capturing knowledge about the mulitilinguality of a knowledge graph is of supreme importance to understand its applicability across multiple languages. Several metrics have been proposed for describing mulitilinguality at the level of a whole knowledge graph. Albeit enabling the understanding of the ecosystem of knowledge graphs in terms of the utilized languages, they are unable to capture a fine-grained description of the languages in which the different entities and properties of the knowledge graph are represented. This lack of representation prevents the comparison of existing knowledge graphs in order to decide which are the most appropriate for a multilingual application.
Lucie-Aimée Kaffee, Kemele M. Endris, Elena Simperl, Maria-Esther Vidal
K-CAP3
2019 An Assessment of Adoption and Quality of Linked Data in European Open Government Data
Luis-Daniel Ibáñez, Ian Millard, Hugh Glaser, Elena Simperl
ISWC (2)4
2019 Characterising dataset search - An analysis of search logs and data requests
Emilia Kacprzak, Laura Koesten, Luis-Daniel Ibáñez, Tom Blount, Jeni Tennison, Elena Simperl
J. Web Semant.6
2018 Studying Topical Relevance with Evidence-based Crowdsourcing
abstract
Information Retrieval systems rely on large test collections to measure their effectiveness in retrieving relevant documents. While the demand is high, the task of creating such test collections is laborious due to the large amounts of data that need to be annotated, and due to the intrinsic subjectivity of the task itself. In this paper we study the topical relevance from a user perspective by addressing the problems of subjectivity and ambiguity. We compare our approach and results with the established TREC annotation guidelines and results. The comparison is based on a series of crowdsourcing pilots experimenting with variables, such as relevance scale, document granularity, annotation template and the number of workers. Our results show correlation between relevance assessment accuracy and smaller document granularity, i.e., aggregation of relevance on paragraph level results in a better relevance accuracy, compared to assessment done at the level of the full document. As expected, our results also show that collecting binary relevance judgments results in a higher accuracy compared to the ternary scale used in the TREC annotation guidelines. Finally, the crowdsourced annotation tasks provided a more accurate document relevance ranking than a single assessor relevance label. This work resulted is a reliable test collection around the TREC Common Core track.
Oana Inel, Giannis Haralabopoulos, Dan Li 0015, Christophe Van Gysel, Zoltán Szlávik, Elena Simperl, Evangelos Kanoulas, Lora Aroyo
CIKM6
2018 Making Sense of Numerical Data - Semantic Labelling of Web Tables
Emilia Kacprzak, José M. Giménez-García, Alessandro Piscopo, Laura Koesten, Luis-Daniel Ibáñez, Jeni Tennison, Elena Simperl
EKAW7
2018 Smart Papers: Dynamic Publications on the Blockchain
Michal R. Hoffman, Luis-Daniel Ibáñez, Huw Fryer, Elena Simperl
ESWC4
2018 Mind the (Language) Gap: Generation of Multilingual Wikipedia Summaries from Wikidata for ArticlePlaceholders
Lucie-Aimée Kaffee, Hady ElSahar, Pavlos Vougiouklis, Christophe Gravier, Frédérique Laforest, Jonathon S. Hare, Elena Simperl
ESWC7
2018 What Does an Ontology Engineering Community Look Like? A Systematic Analysis of the schema.org Community
Samantha Kanza, Alex Stolz, Martin Hepp, Elena Simperl
ESWC4
2018 "A Game Without Competition Is Hardly a Game": The Impact of Competitions on Player Activity in a Human Computation Game"
abstract
Virtual citizen science (VCS) projects enable new forms of scientific research using crowdsourcing and human computation to gather and analyse large-scale datasets. To attract and sustain the number of participants and levels of participation necessary to achieve research aims, some VCS projects have introduced game elements such as competitions to tasks. However, we still know very little about how some game elements, particularly competitions, influence participation rates. To investigate the impact of game elements on player engagement, we conducted a two-part mixed-methods study of EyeWire, a VCS game. First, we interviewed EyeWire designers to understand their rationale for introducing competitions. Guided by their answers, we analysed two datasets of EyeWire user task contributions and chat logs to assess the effectiveness of competitions in achieving designers' goals. Our findings contribute to the growing understanding of how competitions influence participant activity in human computation initiatives and socio-technical systems such as VCS.
Neal Reeves, Peter West, Elena Simperl
HCOMP3
2018 DATA: SEARCH'18 - Searching Data on the Web
abstract
This half day workshop explores challenges in data search, with a particular focus on data on the web. We want to stimulate an interdisciplinary discussion around how to improve the description, discovery, ranking and presentation of structured and semi-structured data, across data formats and domain applications. We welcome contributions describing algorithms and systems, as well as frameworks and studies in human data interaction. The workshop aims to bring together communities interested in making the web of data more discoverable, easier to search and more user friendly.
Paul Groth, Laura Koesten, Philipp Mayr 0001, Maarten de Rijke, Elena Simperl
SIGIR5
2018 Neural Wikipedian: Generating Textual Summaries from Knowledge Base Triples
abstract
Most people need textual or visual interfaces in order to make sense of Semantic Web data. In this paper, we investigate the problem of generating natural language summaries for Semantic Web data using neural networks. Our end-to-end trainable architecture encodes the information from a set of triples into a vector of fixed dimensionality and generates a textual summary by conditioning the output on the encoded vector. We explore a set of different approaches that enable our models to verbalise entities from the input set of triples in the generated text. Our systems are trained and evaluated on two corpora of loosely aligned Wikipedia snippets with triples from DBpedia and Wikidata, with promising results.
Pavlos Vougiouklis, Hady ElSahar, Lucie-Aimée Kaffee, Christophe Gravier, Frédérique Laforest, Jonathon S. Hare, Elena Simperl
J. Web Semant.7
2017 A Query Log Analysis of Dataset Search
Emilia Kacprzak, Laura Koesten, Luis-Daniel Ibáñez, Elena Simperl, Jeni Tennison
ICWE4
2017 To Help or Hinder: Real-Time Chat in Citizen Science
Ramine Tinati, Elena Simperl, Markus Luczak-Rösch
ICWSM2
2017 Dataset Reuse: An Analysis of References in Community Discussions, Publications and Data
abstract
Following the Linked Data principles means maximising the reusability of data over the Web. Reuse of datasets can become apparent when datasets are linked to from other datasets, and referred in scientific articles or community discussions. It can thus be measured, similarly to citations of papers. In this paper we propose dataset reuse metrics and use these metrics to analyse indications of dataset reuse in different communication channels within a scientific community. In particular we consider mailing lists and publications in the Semantic Web community and their correlation with data interlinking. Our results demonstrate that indications of dataset reuse across different communication channels and reuse in terms of data interlinking are positively correlated.
Kemele M. Endris, José M. Giménez-García, Harsh Thakkar, Elena Demidova, Antoine Zimmermann, Christoph Lange 0002, Elena Simperl
K-CAP7
2017 Provenance Information in a Collaborative Knowledge Graph: An Evaluation of Wikidata External References
Alessandro Piscopo, Lucie-Aimée Kaffee, Christopher Phethean, Elena Simperl
ISWC (1)4
2017 Social Incentives in Paid Collaborative Crowdsourcing
abstract
Paid microtask crowdsourcing has traditionally been approached as an individual activity, with units of work created and completed independently by the members of the crowd. Other forms of crowdsourcing have, however, embraced more varied models, which allow for a greater level of participant interaction and collaboration. This article studies the feasibility and uptake of such an approach in the context of paid microtasks. Specifically, we compare engagement, task output, and task accuracy in a paired-worker model with the traditional, single-worker version. Our experiments indicate that collaboration leads to better accuracy and more output, which, in turn, translates into lower costs. We then explore the role of the social flow and social pressure generated by collaborating partners as sources of incentives for improved performance. We utilise a Bayesian method in conjunction with interface interaction behaviours to detect when one of the workers in a pair tries to exit the task. Upon this realisation, the other worker is presented with the opportunity to contact the exiting partner to stay: either for personal financial reasons (i.e., they have not completed enough tasks to qualify for a payment) or for fun (i.e., they are enjoying the task). The findings reveal that: (1) these socially motivated incentives can act as furtherance mechanisms to help workers attain and exceed their task requirements and produce better results than baseline collaborations; (2) microtask crowd workers are empathic (as opposed to selfish) agents, willing to go the extra mile to help their partners get paid; and, (3) social furtherance incentives create a win-win scenario for the requester and for the workers by helping more workers get paid by re-engaging them before they drop out.
Oluwaseyi Feyisetan, Elena Simperl
ACM Trans. Intell. Syst. Technol.2
2017 Enhancing answer completeness of SPARQL queries via crowdsourcing
Maribel Acosta, Elena Simperl, Fabian Flöck, Maria-Esther Vidal
J. Web Semant.2
2016 Please Stay vs Let's Play: Social Pressure Incentives in Paid Collaborative Crowdsourcing
Oluwaseyi Feyisetan, Elena Simperl
ICWE2
2015 Towards Hybrid NER: A Study of Content and Crowdsourcing-Related Performance Factors
Oluwaseyi Feyisetan, Markus Luczak-Rösch, Elena Simperl, Ramine Tinati, Nigel Shadbolt
ESWC3
2015 HARE: A Hybrid SPARQL Engine to Enhance Query Answers via Crowdsourcing
abstract
Due to the semi-structured nature of RDF data, missing values affect answer completeness of queries that are posed against RDF. To overcome this limitation, we present HARE, a novel hybrid query processing engine that brings together machine and human computation to execute SPARQL queries. We propose a model that exploits the characteristics of RDF in order to estimate the completeness of portions of a data set. The completeness model complemented by crowd knowledge is used by the HARE query engine to on-the-fly decide which parts of a query should be executed against the data set or via crowd computing. To evaluate HARE, we created and executed a collection of 50 SPARQL queries against the DBpedia data set. Experimental results clearly show that our solution accurately enhances answer completeness.
Maribel Acosta, Elena Simperl, Fabian Flöck, Maria-Esther Vidal
K-CAP2
2015 Improving Paid Microtasks through Gamification and Adaptive Furtherance Incentives
abstract
Crowdsourcing via paid microtasks has been successfully applied in a plethora of domains and tasks. Previous efforts for making such crowdsourcing more effective have considered aspects as diverse as task and workflow design, spam detection, quality control, and pricing models. Our work expands upon such efforts by examining the potential of adding gamification to microtask interfaces as a means of improving both worker engagement and effectiveness. We run a series of experiments in image labeling, one of the most common use cases for microtask crowdsourcing, and analyse worker behavior in terms of number of images completed, quality of annotations compared against a gold standard, and response to financial and game-specific rewards. Each experiment studies these parameters in two settings: one based on a state-of-the-art, non-gamified task on CrowdFlower and another one using an alternative interface incorporating several game elements. Our findings show that gamification leads to better accuracy and lower costs than conventional approaches that use only monetary incentives. In addition, it seems to make paid microtask work more rewarding and engaging, especially when sociality features are introduced. Following these initial insights, we define a predictive model for estimating the most appropriate incentives for individual workers, based on their previous contributions. This allows us to build a personalised game experience, with gains seen on the volume and quality of work completed.
Oluwaseyi Feyisetan, Elena Simperl, Max Van Kleek, Nigel Shadbolt
WWW2
2014 SPARQL Query Verbalization for Explaining Semantic Search Engine Queries
Basil Ell, Andreas Harth, Elena Simperl
ESWC3
2014 Why Won't Aliens Talk to Us? Content and Community Dynamics in Online Citizen Science
Markus Luczak-Rösch, Ramine Tinati, Elena Simperl, Max Van Kleek, Nigel Shadbolt, Robert J. Simpson
ICWSM3
2014 The Role of Ontology Engineering in Linked Data Publishing and Management: An Empirical Study
abstract
In this article the authors evaluate the adoption and applicability of established ontology engineering results by the Linked Data providers' community. The evaluation relies on a combination of qualitative and quantitative methods; in particular, the authors conducted an analytical survey containing structured interviews with data publishers in order to give an account of the current ontology engineering practice in Linked Data provisioning, and compared and expanded our findings with statistics on ontology development and usage provided by the Billion Triple Challenges datasets from 2012 (using the vocab.cc platform) and from 2014 and other related tools. The findings of the evaluation allow data practitioners and ontologists to yield a better understanding of the conceptual part of the LOD Cloud; and form the basis for the definition of purposeful, empirically grounded guidelines and best practices for developing, managing and using ontologies in the new application scenarios that arise in the context of Linked Data.
Markus Luczak-Rösch, Elena Simperl, Steffen Stadtmüller, Tobias Käfer
Int. J. Semantic Web Inf. Syst.2
2013 Crowdsourcing Linked Data Quality Assessment
Maribel Acosta, Amrapali Zaveri, Elena Simperl, Dimitris Kontokostas, Sören Auer, Jens Lehmann 0001
ISWC (2)3
2012 CrowdMap: Crowdsourcing Ontology Alignment with Microtasks
Cristina Sarasua, Elena Simperl, Natasha F. Noy
ISWC (1)2
2012 ONTOCOM: A reliable cost estimation method for ontology development projects
Elena Simperl, Tobias Bürger, Simon Hangl, Stephan Wörgl, Igor O. Popov
J. Web Semant.1
2011 SeaFish: A Game for Collaborative and Visual Image Annotation and Interlinking
Stefan Thaler, Katharina Siorpaes, David Mear, Elena Simperl, Carl Goodman
ESWC (2)4
2011 Wikiing pro: semantic wiki-based process editor
abstract
Recently, a trend toward collaborative, user-centric, on-line process modeling can be observed. Unfortunately, current social software approaches mostly focus on the graphical development of processes and do not consider existing textual process description like HowTos or guidelines. We address this issue by combining graphical process modeling techniques with a wiki-based light-weight knowledge capturing approach and a background semantic knowledge base. Our approach enables the collaborative maturing of process descriptions with a graphical representation, formal semantic annotations, and natural language. By translating existing textual process descriptions into graphical descriptions and formal semantic annotations, we provide a holistic approach for collaborative process development that is designed to foster knowledge reuse and maturing within the system.
Frank Dengler, Denny Vrandecic, Elena Simperl
K-CAP3
2011 Labels in the Web of Data
Basil Ell, Denny Vrandecic, Elena Simperl
ISWC (1)3
2010 FOLCOM or the Costs of Tagging
Elena Simperl, Tobias Bürger, Christian Hofer
EKAW1
2009 ONTOCOM Revisited: Towards Accurate Cost Predictions for Ontology Development Projects
Elena Simperl, Igor O. Popov, Tobias Bürger
ESWC1
2009 Reusing ontologies on the Semantic Web: A feasibility study
Elena Simperl
Data Knowl. Eng.1
2008 OMEGA: An Automatic Ontology Metadata Generation Algorithm
Rachanee Ungrangsi, Elena Simperl
EKAW2
2007 An Ontology-Driven Approach To Reflective Middleware
abstract
Recent work in the field of middleware technology proposes semantic spaces as a tool for coping with the scalability, heterogeneity and dynamism issues arising in large scale distributed environments. Reflective middleware moreover offers answers to the needs for adaptivity and selfdetermination of systems where mobility and ubiquity add to such environments. Based on experiences with traditional middleware we argue that ontology-driven management is a major advancement for semantic spaces and provides the fundamental means for reflection. By means of ontologies, and ontology-based reasoning services we can implement automatic adaptation of the middleware's functionality to environmental changes and user desires.
Reto Krummenacher, Elena Simperl, Dieter Fensel
Web Intelligence2
2006 DEMO - Design Environment for Metadata Ontologies
Jens Hartmann 0001, Elena Simperl, Raúl Palma, Asunción Gómez-Pérez
ESWC2
2006 : A Cost Estimation Model for Ontology Engineering
Elena Simperl, Christoph Tempich, York Sure-Vetter
ISWC1
2005 Enabling Real World Semantic Web Applications Through a Coordination Middleware
Robert Tolksdorf, Lyndon J. B. Nixon, Elena Simperl, Franziska Liebsch
ESWC3
2005 Towards a Tuplespace-Based Middleware for the Semantic Web
abstract
The realization of the semantic Web needs a set of specialized middleware as its infrastructure. In this paper we describe the principles of tuplespace computing, explain why tuplespaces are a suitable middleware for the semantic Web, envision "semantic Web spaces", and outline how our tuplespace platform XMLSpaces can be extended to support semantic Web technologies, like RDF(S) and OWL.
Robert Tolksdorf, Elena Simperl, Lyndon J. B. Nixon
Web Intelligence2
2004 Engineering a Semantic Web for Pathology
Robert Tolksdorf, Elena Simperl
ICWE2