Dolf Trieschnigg

dblp:26/3082 · also Rudolf Berend Trieschnigg · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 22 · 7 first-authorArtificial intelligence and machine learning · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
10 papers
Information retrieval · 84% Data mining · 7% Data models and query languages · 4%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 25 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › evaluation › relevance judgment
graded relevance
0.212014
Exploiting user disagreement for web search evaluation: an experimental approach · WSDM 2014
Information retrieval
retrieval evaluation
0.212014
Exploiting user disagreement for web search evaluation: an experimental approach · WSDM 2014
Information retrieval › distributed information retrieval
federated search
0.212013
SearchResultFinder: federated search made easy · SIGIR 2013
Information retrieval
web search
0.212013
SearchResultFinder: federated search made easy · SIGIR 2013
Information retrieval
distributed information retrieval
0.112012
Peer-to-Peer Information Retrieval: An Overview · ACM Trans. Inf. Syst. 2012
Information retrieval › distributed information retrieval
peer-to-peer search
0.112012
Peer-to-Peer Information Retrieval: An Overview · ACM Trans. Inf. Syst. 2012
Information retrieval › document retrieval › domain-specific retrieval
biomedical information retrieval
0.112009
MeSH Up: effective MeSH text classification for improved document retrieval · Bioinform. 2009
Data mining › text mining
text classification
0.112009
Response to comment on 'MeSH-up: effective MeSH text classification for improved document retrieval' · Bioinform. 2009
Knowledge graphs
concept relatedness
0.112008
Measuring concept relatedness using language models · SIGIR 2008
Data models and query languages
conceptual modeling
0.112008
Parsimonious concept modeling · SIGIR 2008
Information retrieval
cross-language information retrieval
0.112008
Biomedical cross-language information retrieval · SIGIR 2008
Information retrieval › text analysis
semantic relatedness
0.112008
Measuring concept relatedness using language models · SIGIR 2008
Information retrieval
indexing
0.112007
The influence of basic tokenization on biomedical document retrieval · SIGIR 2007
Information retrieval › document processing
tokenization
0.112007
The influence of basic tokenization on biomedical document retrieval · SIGIR 2007
Bioinformatics and computational biology
biomedical text mining
0.132008
Measuring concept relatedness using language models · SIGIR 2008
Biomedical cross-language information retrieval · SIGIR 2008
The influence of basic tokenization on biomedical document retrieval · SIGIR 2007
Information retrieval › text analysis › topic analysis
topic detection and tracking
0.112005
Scalable hierarchical topic detection: exploring a sample based approach · SIGIR 2005
Information retrieval › distributed information retrieval
distributed search
0.012012
Peer-to-Peer Information Retrieval: An Overview · ACM Trans. Inf. Syst. 2012
Information retrieval › search engines
search engine architecture
0.012012
Peer-to-Peer Information Retrieval: An Overview · ACM Trans. Inf. Syst. 2012
Distributed systems › peer-to-peer systems
peer-to-peer search
0.012012
Peer-to-Peer Information Retrieval: An Overview · ACM Trans. Inf. Syst. 2012
Distributed systems
peer-to-peer systems
0.012012
Peer-to-Peer Information Retrieval: An Overview · ACM Trans. Inf. Syst. 2012
Bioinformatics and computational biology › biomedical text mining
biomedical literature retrieval
0.012009
Response to comment on 'MeSH-up: effective MeSH text classification for improved document retrieval' · Bioinform. 2009
Information retrieval
evaluation
0.012007
The influence of basic tokenization on biomedical document retrieval · SIGIR 2007
Information retrieval › evaluation
retrieval effectiveness
0.012007
The influence of basic tokenization on biomedical document retrieval · SIGIR 2007
Data mining › clustering › hierarchical clustering
agglomerative clustering
0.012005
Scalable hierarchical topic detection: exploring a sample based approach · SIGIR 2005
Data mining
clustering
0.012005
Scalable hierarchical topic detection: exploring a sample based approach · SIGIR 2005

Methods — techniques the papers use, named apart from their topics

language model · 0.4vector space model · 0.2k-nearest neighbor · 0.2XPath extraction · 0.2cross-entropy · 0.1cross entropy · 0.1agglomerative clustering · 0.1DAG optimization · 0.1
YearPublicationVenuePosition
2015 Audience and the Use of Minority Languages on Twitter
Dong Nguyen 0002, Dolf Trieschnigg, Leonie Cornips
ICWSM2
2014 Using Crowdsourcing to Investigate Perception of Narrative Similarity
abstract
For many applications measuring the similarity between documents is essential. However, little is known about how users perceive similarity between documents. This paper presents the first large-scale empirical study that investigates perception of narrative similarity using crowdsourcing. As a dataset we use a large collection of Dutch folk narratives. We study the perception of narrative similarity by both experts and non-experts by analyzing their similarity ratings and motivations for these ratings. While experts focus mostly on the plot, characters and themes of narratives, non-experts also pay attention to dimensions such as genre and style. Our results show that a more nuanced view is needed of narrative similarity than captured by story types, a concept used by scholars to group similar folk narratives. We also evaluate to what extent unsupervised and supervised models correspond with how humans perceive narrative similarity.
Dong Nguyen 0002, Dolf Trieschnigg, Mariët Theune
CIKM2
2014 Aligning Vertical Collection Relevance with User Intent
abstract
Selecting and aggregating different types of content from multiple vertical search engines is becoming popular in web search. The user vertical intent, the verticals the user expects to be relevant for a particular information need, might not correspond to the vertical collection relevance, the verticals containing the most relevant content. In this work we propose different approaches to define the set of relevant verticals based on document judgments. We correlate the collection-based relevant verticals obtained from these approaches to the real user vertical intent, and show that they can be aligned relatively well. The set of relevant verticals defined by those approaches could therefore serve as an approximate but reliable ground-truth for evaluating vertical selection, avoiding the need for collecting explicit user vertical intent, and vice versa.
Ke Zhou 0003, Thomas Demeester, Dong Nguyen 0002, Djoerd Hiemstra, Dolf Trieschnigg
CIKM5
2014 Why Gender and Age Prediction from Tweets is Hard: Lessons from a Crowdsourcing Experiment
Dong Nguyen 0002, Dolf Trieschnigg, A. Seza Dogruöz, Rilana Gravel, Mariët Theune, Theo Meder, Franciska de Jong
COLING2
2014 Average Precision: Good Guide or False Friend to Multimedia Search Effectiveness?
Robin Aly, Dolf Trieschnigg, Kevin McGuinness, Noel E. O'Connor, Franciska de Jong
MMM (2)2
2014 Exploiting user disagreement for web search evaluation: an experimental approach
abstract
To express a more nuanced notion of relevance as compared to binary judgments, graded relevance levels can be used for the evaluation of search results. Especially in Web search, users strongly prefer top results over less relevant results, and yet they often disagree on which are the top results for a given information need. Whereas previous works have generally considered disagreement as a negative effect, this paper proposes a method to exploit this user disagreement by integrating it into the evaluation procedure.
Thomas Demeester, Robin Aly, Djoerd Hiemstra, Dong Nguyen 0002, Dolf Trieschnigg, Chris Develder
WSDM5
2013 Improving Cyberbullying Detection with User Context
Maral Dadvar, Dolf Trieschnigg, Roeland Ordelman, Franciska de Jong
ECIR2
2013 Snippet-Based Relevance Predictions for Federated Web Search
Thomas Demeester, Dong Nguyen 0002, Dolf Trieschnigg, Chris Develder, Djoerd Hiemstra
ECIR3
2013 Folktale Classification Using Learning to Rank
Dong Nguyen 0002, Dolf Trieschnigg, Mariët Theune
ECIR2
2013 Using a Stack Decoder for Structured Search
Kien-Tsoi T. E. Tjin-Kam-Jet, Dolf Trieschnigg, Djoerd Hiemstra
FQAS2
2013 "How Old Do You Think I Am?" A Study of Language and Age in Twitter
Dong Nguyen 0002, Rilana Gravel, Dolf Trieschnigg, Theo Meder
ICWSM3
2013 SearchResultFinder: federated search made easy
abstract
Building a federated search engine based on a large number existing web search engines is a challenge: implementing the programming interface (API) for each search engine is an exacting and time-consuming job. In this demonstration we present SearchResultFinder, a browser plugin which speeds up determining reusable XPaths for extracting search result items from HTML search result pages. Based on a single search result page, the tool presents a ranked list of candidate extraction XPaths and allows highlighting to view the extraction result. An evaluation with 148 web search engines shows that in 90% of the cases a correct XPath is suggested.
Dolf Trieschnigg, Kien-Tsoi T. E. Tjin-Kam-Jet, Djoerd Hiemstra
SIGIR1
2012 Federated search in the wild: the combined power of over a hundred search engines
abstract
Federated search has the potential of improving web search: the user becomes less dependent on a single search provider and parts of the deep web become available through a unified interface, leading to a wider variety in the retrieved search results. However, a publicly available dataset for federated search reflecting an actual web environment has been absent. As a result, it has been difficult to assess whether proposed systems are suitable for the web setting. We introduce a new test collection containing the results from more than a hundred actual search engines, ranging from large general web search engines such as Google and Bing to small domain-specific engines. We discuss the design and analyze the effect of several sampling methods. For a set of test queries, we collected relevance judgements for the top 10 results of each search engine. The dataset is publicly available and is useful for researchers interested in resource selection for web search collections, result merging and size estimation of uncooperative resources.
Dong Nguyen 0002, Thomas Demeester, Dolf Trieschnigg, Djoerd Hiemstra
CIKM3
2012 Towards User Modelling in the Combat against Cyberbullying
Maral Dadvar, Roeland Ordelman, Franciska de Jong, Dolf Trieschnigg
NLDB4
2012 Peer-to-Peer Information Retrieval: An Overview
abstract
Peer-to-peer technology is widely used for file sharing. In the past decade a number of prototype peer-to-peer information retrieval systems have been developed. Unfortunately, none of these has seen widespread real-world adoption and thus, in contrast with file sharing, information retrieval is still dominated by centralized solutions. In this article we provide an overview of the key challenges for peer-to-peer information retrieval and the work done so far. We want to stimulate and inspire further research to overcome these challenges. This will open the door to the development and large-scale deployment of real-world peer-to-peer information retrieval systems that rival existing centralized client-server solutions in terms of scalability, performance, user satisfaction, and freedom.
Almer S. Tigelaar, Djoerd Hiemstra, Dolf Trieschnigg
ACM Trans. Inf. Syst.3
2011 Free-Text Search versus Complex Web Forms
Kien-Tsoi T. E. Tjin-Kam-Jet, Dolf Trieschnigg, Djoerd Hiemstra
ECIR2
2011 Classic Children's Literature - Difficult to Read?
Dolf Trieschnigg, Claudia Hauff
ECIR1
2010 A cross-lingual framework for monolingual biomedical information retrieval
abstract
An important challenge for biomedical information retrieval (IR) is dealing with the complex, inconsistent and ambiguous biomedical terminology. Frequently, a concept-based representation defined in terms of a domain-specific terminological resource is employed to deal with this challenge. In this paper, we approach the incorporation of a concept-based representation in monolingual biomedical IR from a cross-lingual perspective. In the proposed framework, this is realized by translating and matching between text and concept-based representations. The approach allows for deployment of a rich set of techniques proposed and evaluated in traditional cross-lingual IR. We compare six translation models and measure their effectiveness in the biomedical domain. We demonstrate that the approach can result in significant improvements in retrieval effectiveness over word-based retrieval. Moreover, we demonstrate increased effectiveness of a CLIR framework for monolingual biomedical IR if basic translations models are combined. © 2010 ACM.
Dolf Trieschnigg, Djoerd Hiemstra, Franciska de Jong, Wessel Kraaij
CIKM1
2010 Conceptual language models for domain-specific retrieval
Edgar Meij, Dolf Trieschnigg, Maarten de Rijke, Wessel Kraaij
Inf. Process. Manag.2
2009 MeSH Up: effective MeSH text classification for improved document retrieval
abstract
MOTIVATION: Controlled vocabularies such as the Medical Subject Headings (MeSH) thesaurus and the Gene Ontology (GO) provide an efficient way of accessing and organizing biomedical information by reducing the ambiguity inherent to free-text data. Different methods of automating the assignment of MeSH concepts have been proposed to replace manual annotation, but they are either limited to a small subset of MeSH or have only been compared with a limited number of other systems. RESULTS: We compare the performance of six MeSH classification systems [MetaMap, EAGL, a language and a vector space model-based approach, a K-Nearest Neighbor (KNN) approach and MTI] in terms of reproducing and complementing manual MeSH annotations. A KNN system clearly outperforms the other published approaches and scales well with large amounts of text using the full MeSH thesaurus. Our measurements demonstrate to what extent manual MeSH annotations can be reproduced and how they can be complemented by automatic annotations. We also show that a statistically significant improvement can be obtained in information retrieval (IR) when the text of a user's query is automatically annotated with MeSH concepts, compared to using the original textual query alone. CONCLUSIONS: The annotation of biomedical texts using controlled vocabularies such as MeSH can be automated to improve text-only IR. Furthermore, the automatic MeSH annotation system we propose is highly scalable and it generates improvements in IR comparable with those observed for manual annotations.
Dolf Trieschnigg, Piotr Pezik, Vivian Lee, Franciska de Jong, Wessel Kraaij, Dietrich Rebholz-Schuhmann
Bioinform.1
2009 Response to comment on 'MeSH-up: effective MeSH text classification for improved document retrieval'
abstract
Abstract Contact: [email protected]; [email protected] As developers and primary users of MTI and MetaMap, Névéol et al. made a number of interesting comments on our recent publication in Bioinformatics. However, some of the results and conclusions found in the reply seem premature and lack proper clarification.
Dolf Trieschnigg, Piotr Pezik, Vivian Lee, Franciska de Jong, Wessel Kraaij, Dietrich Rebholz-Schuhmann
Bioinform.1
2008 Parsimonious concept modeling
abstract
No abstract available.
Edgar Meij, Dolf Trieschnigg, Maarten de Rijke, Wessel Kraaij
SIGIR2
2008 Biomedical cross-language information retrieval
abstract
No abstract available.
Dolf Trieschnigg
SIGIR1
2008 Measuring concept relatedness using language models
abstract
Over the years, the notion of concept relatedness has attracted considerable attention. A variety of approaches, based on ontology structure, information content, association, or context have been proposed to indicate the relatedness of abstract ideas. We propose a method based on the cross entropy reduction between language models of concepts which are estimated based on document-concept assignments. The approach shows improved or competitive results compared to state-of-the-art methods on two test sets in the biomedical domain.
Dolf Trieschnigg, Edgar Meij, Maarten de Rijke, Wessel Kraaij
SIGIR1
2007 The influence of basic tokenization on biomedical document retrieval
abstract
Tokenization is a fundamental preprocessing step in Information Retrieval systems in which text is turned into index terms. This paper quantifies and compares the influence of various simple tokenization techniques on document retrieval effectiveness in two domains: biomedicine and news. As expected, biomedical retrieval is more sensitive to small changes in the tokenization method. The tokenization strategy can make the difference between a mediocre and well performing IR system, especially in the biomedical domain.
Dolf Trieschnigg, Wessel Kraaij, Franciska de Jong
SIGIR1
2005 Scalable hierarchical topic detection: exploring a sample based approach
abstract
Hierarchical topic detection is a new task in the TDT 2004 evaluation program, which aims to organize an unstructured news collection in a directed acyclic graph (DAG) structure, reflecting the topics discussed. We present a scalable architecture for HTD and compare several alternative choices for agglomerative clustering and DAG optimization in order to minimize the HTD cost metric.
Dolf Trieschnigg, Wessel Kraaij
SIGIR1