Rianne Kaptein

dblp:42/2877 · DBLP profile ↗
← Back
14ranked-venue papers
11as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 10 first-authorArtificial intelligence and machine learning · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Information retrieval · 85% Knowledge graphs · 15%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Human-robot interaction
social robot
0.212016
Child's Culture-related Experiences with a Social Robot at Diabetes Camps · HRI 2016
Information retrieval › search engines › semantic search › entity retrieval
entity ranking
0.222013
Exploiting the category structure of Wikipedia for entity ranking · Artif. Intell. 2013
Using wikipedia categories for ad hoc search · SIGIR 2009
Information retrieval
retrieval models
0.222010
Linking wikipedia to the web · SIGIR 2010
Using parsimonious language models on web data · SIGIR 2008
Knowledge graphs › taxonomy
wikipedia category structure
0.212013
Exploiting the category structure of Wikipedia for entity ranking · Artif. Intell. 2013
Information retrieval › retrieval models
language model
0.112010
Linking wikipedia to the web · SIGIR 2010
Information retrieval › retrieval models
ad-hoc retrieval
0.112009
Using wikipedia categories for ad hoc search · SIGIR 2009
Information retrieval › document retrieval › concept-based retrieval
category-based retrieval
0.112009
Using wikipedia categories for ad hoc search · SIGIR 2009
Information retrieval › document processing › document analysis
document layout analysis
0.112009
Who said what to whom?: capturing the structure of debates · SIGIR 2009
Information retrieval
text summarization
0.112009
Who said what to whom?: capturing the structure of debates · SIGIR 2009
Information retrieval › retrieval models › language model
parsimonious language model
0.112008
Using parsimonious language models on web data · SIGIR 2008
Human-robot interaction
child-robot interaction
0.112016
Child's Culture-related Experiences with a Social Robot at Diabetes Camps · HRI 2016
Information retrieval › retrieval models › probabilistic retrieval model
document prior
0.012010
Linking wikipedia to the web · SIGIR 2010
Knowledge graphs › semantic web
semantic annotation
0.012009
Who said what to whom?: capturing the structure of debates · SIGIR 2009
Information retrieval › web search
web information retrieval
0.012008
Using parsimonious language models on web data · SIGIR 2008

Methods — techniques the papers use, named apart from their topics

questionnaire · 0.2observation · 0.2language modeling · 0.2category structure exploitation · 0.2social bookmarking · 0.1parsimonious language model · 0.1graph visualization · 0.1
YearPublicationVenuePosition
2018 Rethinking Summarization and Storytelling for Modern Social Multimedia
Stevan Rudinac, Tat-Seng Chua, Nicolás E. Díaz Ferreyra, Gerald Friedland, Tatjana Gornostaja, Benoit Huet, Rianne Kaptein, Krister Lindén, Marie-Francine Moens, Jaakko Peltonen, Miriam Redi, Markus Schedl, David A. Shamma, Alan F. Smeaton, Lexing Xie
MMM (1)7
2016 Child's Culture-related Experiences with a Social Robot at Diabetes Camps
abstract
This paper investigates the experiences of Italian and Dutch children while interacting with a social robot that is designed to support their diabetes self-management. Observations of children's behaviors and analyses of questionnaires at diabetes camps, showed positive experiences with variation (e.g., Italian children seemed to be more open and expressive, and more close to the robot compared to the Dutch children). A culture-aware robot should be sensitive to such differences.
Anouk Neerincx, Francesca Sacchitelli, Rianne Kaptein, Sylvia van der Pal, Elettra Oleari, Mark A. Neerincx
HRI3
2016 Estimating Reputation Polarity on Microblog Posts
Maria-Hendrike Peetz, Maarten de Rijke, Rianne Kaptein
Inf. Process. Manag.3
2014 Analyzing Discussions on Twitter: Case Study on HPV Vaccinations
Rianne Kaptein, Erik M. Boertjes, David Langley
ECIR1
2014 Needle Custom Search - Recall-Oriented Search on the Web Using Semantic Annotations
Rianne Kaptein, Gijs Koot, Mirjam Huis in 't Veld, Egon L. van den Broek
ECIR1
2013 Exploiting the category structure of Wikipedia for entity ranking
Rianne Kaptein, Jaap Kamps
Artif. Intell.1
2011 Explicit extraction of topical context
abstract
Abstract This article studies one of the main bottlenecks in providing more effective information access: the poverty on the query end. We explore whether users can classify keyword queries into categories from the DMOZ directory on different levels and whether this topical context can help retrieval performance. We have conducted a user study to let participants classify queries into DMOZ categories, either by freely searching the directory or by selection from a list of suggestions. Results of the study show that DMOZ categories are suitable for topic categorization. Both free search and list selection can be used to elicit topical context. Free search leads to more specific categories than the list selections. Participants in our study show moderate agreement on the categories they select, but broad agreement on the higher levels of chosen categories. The free search categories significantly improve retrieval effectiveness. The more general list selection categories and the top‐level categories do not lead to significant improvements. Combining topical context with blind relevance feedback leads to better results than applying either of them separately. We conclude that DMOZ is a suitable resource for interacting with users on topical categories applicable to their query, and can lead to better search results.
Rianne Kaptein, Jaap Kamps
J. Assoc. Inf. Sci. Technol.1
2010 Entity ranking using Wikipedia as a pivot
abstract
In this paper we investigate the task of Entity Ranking on the Web. Searchers looking for entities are arguably better served by presenting a ranked list of entities directly, rather than a list of web pages with relevant but also potentially redundant information about these entities. Since entities are represented by their web homepages, a naive approach to entity ranking is to use standard text retrieval. Our experimental results clearly demonstrate that text retrieval is effective at finding relevant pages, but performs poorly at finding entities. Our proposal is to use Wikipedia as a pivot for finding entities on the Web, allowing us to reduce the hard web entity ranking problem to easier problem of Wikipedia entity ranking. Wikipedia allows us to properly identify entities and some of their characteristics, and Wikipedia's elaborate category structure allows us to get a handle on the entity's type.
Rianne Kaptein, Pavel Serdyukov, Arjen P. de Vries, Jaap Kamps
CIKM1
2010 How Different Are Language Models andWord Clouds?
Rianne Kaptein, Djoerd Hiemstra, Jaap Kamps
ECIR1
2010 Linking wikipedia to the web
abstract
We investigate the task of finding links from Wikipedia pages to external web pages. Such external links significantly extend the information in Wikipedia with information from the Web at large, while retaining the encyclopedic organization of Wikipedia. We use a language modeling approach to create a full-text and anchor text runs, and experiment with different document priors. In addition we explore whether social bookmarking site Delicious can be exploited to further improve our performance. We have constructed a test collection of 53 topics, which are Wikipedia pages on different entities. Our findings are that the anchor text index is a very effective method to retrieve home pages. Url class and anchor text length priors and their combination leads to the best results. Using Delicious on its own does not lead to very good results, but it does contain valuable information. Combining the best anchor text run and the Delicious run leads to further improvements.
Rianne Kaptein, Pavel Serdyukov, Jaap Kamps
SIGIR1
2010 Focused retrieval and result aggregation with political data
abstract
This paper presents a case-study in which we use a large semi-structured data set consisting of official transcripts of meetings of the Dutch parliament for focused retrieval and result aggregation. Transcripts of meetings are a document genre characterized by a complex narrative structure. The essence is not only what is said, but also by who and to whom. We have notes of more than 40 years of Dutch parliamentary debates where this structure is exploited to automatically make semantic annotations. These annotations yield numerous new ways of searching, browsing, mining and summarizing these documents. Concerning result aggregation, we summarise and visualise the structure of meetings into tables of content and interruption graphs. The contents of meetings or parts of meetings are condensed into word clouds that are created using a parsimonious language model. Furthermore, we have developed a search engine that exploits the structure and annotations of our data making it possible to provide entry points, to group search results, and to use faceted search techniques for data-exploration. Evaluation shows that our content and structure summarization tools provide a good first impression of a debate. Users reported that, compared to a standard document retrieval system, our search engine gives a better overview of the data. Search tasks are performed faster and the users felt more certain of their answers.
Rianne Kaptein, Maarten Marx
Inf. Retr.1
2009 Using wikipedia categories for ad hoc search
abstract
In this paper we explore the use of category information for ad hoc retrieval in Wikipedia. We show that techniques for entity ranking exploiting this category information can also be applied to ad hoc topics and lead to significant improvements. Automatically assigned target categories are good surrogates for manually assigned categories, which perform only slightly better.
Rianne Kaptein, Marijn Koolen, Jaap Kamps
SIGIR1
2009 Who said what to whom?: capturing the structure of debates
abstract
Transcripts of meetings are a document genre characterized by a complex narrative structure. The essence is not only what is said, but also by who and to whom. This paper investigates whether we can use semantic annotations like the speaker in order to capture this debate structure, as well as the related content of the debate. The structure is visualized in a graph, while the content is condensed into word clouds, that are created using a parsimonious language model. Evaluation shows that both tools adequately capture the structure and content of the debate at an aggregated level.
Rianne Kaptein, Maarten Marx, Jaap Kamps
SIGIR1
2008 Using parsimonious language models on web data
abstract
In this paper we explore the use of parsimonious language models for web retrieval. These models are smaller thus more efficient than the standard language models and are therefore well suited for large-scale web retrieval. We have conducted experiments on four TREC topic sets, and found that the parsimonious language model results in improvement of retrieval effectiveness over the standard language model for all data-sets and measures. In all cases the improvement is significant, and more substantial than in earlier experiments on newspaper/newswire data.
Rianne Kaptein, Rongmei Li, Djoerd Hiemstra, Jaap Kamps
SIGIR1