EDBT 2026 Demo / reviewers in the wild / expert
Wessel Kraaij
dblp:k/WesselKraaij
· DBLP profile ↗
47ranked-venue papers
5as first author
5since 2021 · last 2024
0000-0001-7797-619XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 25 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 1 first-authorHuman-computer interaction and ubiquitous computing · 6 · 1 first-authorComputer networks · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Better Together - Empowering Citizen Collectives with Community Learning
Wessel Kraaij, Geiske Bouma, Marloes van der Klauw, Pepijn van Empelen |
I4CS | 1 |
| 2024 | Injecting the score of the first-stage retriever as text improves BERT-based re-rankersabstractAbstract In this paper we propose a novel approach for combining first-stage lexical retrieval models and Transformer-based re-rankers: we inject the relevance score of the lexical model as a token into the input of the cross-encoder re-ranker. It was shown in prior work that interpolation between the relevance score of lexical and Bidirectional Encoder Representations from Transformers (BERT) based re-rankers may not consistently result in higher effectiveness. Our idea is motivated by the finding that BERT models can capture numeric information. We compare several representations of the Best Match 25 (BM25) and Dense Passage Retrieval (DPR) scores and inject them as text in the input of four different cross-encoders. Since knowledge distillation, i.e., teacher-student training, proved to be highly effective for cross-encoder re-rankers, we additionally analyze the effect of injecting the relevance score into the student model while training the model by three larger teacher models. Evaluation on the MSMARCO Passage collection and the TREC DL collections shows that the proposed method significantly improves over all cross-encoder re-rankers as well as the common interpolation methods. We show that the improvement is consistent for all query types. We also find an improvement in exact matching capabilities over both the first-stage rankers and the cross-encoders. Our findings indicate that cross-encoder re-rankers can efficiently be improved without additional computational burden or extra steps in the pipeline by adding the output of the first-stage ranker to the model input. This effect is robust for different models and query types. Arian Askari, Amin Abolghasemi, Gabriella Pasi, Wessel Kraaij, Suzan Verberne |
Discov. Comput. | 4 |
| 2024 | Retrieval for Extremely Long Queries and Documents with RPRS: A Highly Efficient and Effective Transformer-based Re-RankerabstractRetrieval with extremely long queries and documents is a well-known and challenging task in information retrieval and is commonly known as Query-by-Document (QBD) retrieval. Specifically designed Transformer models that can handle long input sequences have not shown high effectiveness in QBD tasks in previous work. We propose a Re-Ranker based on the novel Proportional Relevance Score (RPRS) to compute the relevance score between a query and the top- k candidate documents. Our extensive evaluation shows RPRS obtains significantly better results than the state-of-the-art models on five different datasets. Furthermore, RPRS is highly efficient, since all documents can be pre-processed, embedded, and indexed before query time that gives our re-ranker the advantage of having a complexity of O(N) , where N is the total number of sentences in the query and candidate documents. Furthermore, our method solves the problem of the low-resource training in QBD retrieval tasks as it does not need large amounts of training data and has only three parameters with a limited range that can be optimized with a grid search even if a small amount of labeled data is available. Our detailed analysis shows that RPRS benefits from covering the full length of candidate documents and queries. Arian Askari, Suzan Verberne, Amin Abolghasemi, Wessel Kraaij, Gabriella Pasi |
ACM Trans. Inf. Syst. | 4 |
| 2023 | Injecting the BM25 Score as Text Improves BERT-Based Re-rankers
Arian Askari, Amin Abolghasemi, Gabriella Pasi, Wessel Kraaij, Suzan Verberne |
ECIR (1) | 4 |
| 2023 | How do others cope? Extracting coping strategies for adverse drug events from social mediaabstractPatients advise their peers on how to cope with their illness in daily life on online support groups. To date, no efforts have been made to automatically extract recommended coping strategies from online patient discussion groups. We introduce this new task, which poses a number of challenges including complex, long entities, a large long-tailed label space, and cross-document relations. We present an initial ontology for coping strategies as a starting point for future research on coping strategies, and the first end-to-end pipeline for extracting coping strategies for side effects. We also compared two possible computational solutions for this novel and highly challenging task; multi-label classification and named entity recognition (NER) with entity linking (EL). We evaluated our methods on the discussion forum from the Facebook group of the worldwide patient support organization ‘GIST support international’ (GSI); GIST support international donated the data to us. We found that coping strategy extraction is difficult and both methods attain limited performance (measured with F1 score) on held out test sets; multi-label classification outperforms NER+EL (F1=0.220 vs F1=0.155). An inspection of the multi-label classification output revealed that for some of the incorrect predictions, the reference label is close to the predicted label in the ontology (e.g. the predicted label ‘juice’ instead of the more specific reference label ‘grapefruit juice’). Performance increased to F1=0.498 when we evaluated at a coarser level of the ontology. We conclude that our pipeline can be used in a semi-automatic setting, in interaction with domain experts to discover coping strategies for side effects from a patient forum. For example, we found that patients recommend ginger tea for nausea and magnesium and potassium supplements for cramps. This information can be used as input for patient surveys or clinical studies. Anne Dirkson, Suzan Verberne, Gerard Van Oortmerssen, Hans Gelderblom, Wessel Kraaij |
J. Biomed. Informatics | 5 |
| 2020 | Personalized support for well-being at work: an overview of the SWELL projectabstractRecent advances in wearable sensor technology and smartphones enable simple and affordable collection of personal analytics. This paper reflects on the lessons learned in the SWELL project that addressed the design of user-centered ICT applications for self-management of vitality in the domain of knowledge workers. These workers often have a sedentary lifestyle and are susceptible to mental health effects due to a high workload. We present the sense–reason–act framework that is the basis of the SWELL approach and we provide an overview of the individual studies carried out in SWELL. In this paper, we revisit our work on reasoning: interpreting raw heterogeneous sensor data, and acting: providing personalized feedback to support behavioural change. We conclude that simple affordable sensors can be used to classify user behaviour and heath status in a physically non-intrusive way. The interpreted data can be used to inform personalized feedback strategies. Further longitudinal studies can now be initiated to assess the effectiveness of m-Health interventions using the SWELL methods. Wessel Kraaij, Suzan Verberne, Saskia Koldijk, Elsbeth de Korte, Saskia van Dantzig, Maya Sappelli, Muhammad Shoaib 0001, Steven Bosems, Reinoud Achterkamp, Alberto Bonomi, John G. M. Schavemaker, R. J. Hulsebosch, Thymen Wabeke, Miriam M. R. Vollenbroek-Hutten, Mark A. Neerincx, Marten van Sinderen |
User Model. User Adapt. Interact. | 1 |
| 2018 | Open Knowledge Discovery and Data Mining from Patient ForumsabstractIn the biomedical realm, open knowledge discovery from text has been limited to semi -structured data, such as electronic health records, and biomedical literature [1]. Despite a wealth of research on adverse drug effect (ADE) extraction from patient forum data [2], open knowledge discovery from patient forums has only been conducted incidentally by patient communities themselves [3]-[5]. These patient forums, however, contain a wealth of knowledge: the experiences of the patients themselves. Indeed, patients can offer personal information and experiences that clinicians cannot provide [6], for example on effective daily coping mechanisms. Anne Dirkson, Suzan Verberne, Gerard Van Oortmerssen, Hans Gelderblom, Wessel Kraaij |
eScience | 5 |
| 2018 | Detecting Work Stress in Offices by Combining Unobtrusive SensorsabstractEmployees often report the experience of stress at work. In the SWELL project we investigate how new context aware pervasive systems can support knowledge workers to diminish stress. The focus of this paper is on developing automatic classifiers to infer working conditions and stress related mental states from a multimodal set of sensor data (computer logging, facial expressions, posture and physiology). We address two methodological and applied machine learning challenges: 1) Detecting work stress using several (physically) unobtrusive sensors, and 2) Taking into account individual differences. A comparison of several classification approaches showed that, for our SWELL-KW dataset, neutral and stressful working conditions can be distinguished with 90 percent accuracy by means of SVM. Posture yields most valuable information, followed by facial expressions. Furthermore, we found that the subjective variable`mental effort' can be better predicted from sensor data than, e.g., `perceived stress'. A comparison of several regression approaches showed that mental effort can be predicted best by a decision tree (correlation of 0.82). Facial expressions yield most valuable information, followed by posture. We find that especially for estimating mental states it makes sense to address individual differences. When we train models on particular subgroups of similar users, (in almost all cases) a specialized model performs equally well or better than a generic model. Saskia Koldijk, Mark A. Neerincx, Wessel Kraaij |
IEEE Trans. Affect. Comput. | 3 |
| 2017 | Evaluation of context-aware recommendation systems for information re-findingabstractIn this article we evaluate context‐aware recommendation systems for information re‐finding by knowledge workers. We identify 4 criteria that are relevant for evaluating the quality of knowledge worker support: context relevance, document relevance, prediction of user action, and diversity of the suggestions. We compare 3 different context‐aware recommendation methods for information re‐finding in a writing support task. The first method uses contextual prefiltering and content‐based recommendation (CBR), the second uses the just‐in‐time information retrieval paradigm (JITIR), and the third is a novel network‐based recommendation system where context is part of the recommendation model (CIA). We found that each method has its own strengths: CBR is strong at context relevance, JITIR captures document relevance well, and CIA achieves the best result at predicting user action. Weaknesses include that CBR depends on a manual source to determine the context and in JITIR the context query can fail when the textual content is not sufficient. We conclude that to truly support a knowledge worker, all 4 evaluation criteria are important. In light of that conclusion, we argue that the network‐based approach the CIA offers has the highest robustness and flexibility for context‐aware information recommendation. Maya Sappelli, Suzan Verberne, Wessel Kraaij |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2017 | Improving video event retrieval by user feedbackabstractIn content based video retrieval videos are often indexed with semantic labels ( concepts ) using pre-trained classifiers. These pre-trained classifiers ( concept detectors ), are not perfect, and thus the labels are noisy. Additionally, the amount of pre-trained classifiers is limited. Often automatic methods cannot represent the query adequately in terms of the concepts available. This problem is also apparent in the retrieval of events, such as bike trick or birthday party . Our solution is to obtain user feedback. This user feedback can be provided on two levels: concept level and video level . We introduce the method Adaptive Relevance Feedback ( ARF ) on video level feedback. ARF is based on the classical Rocchio relevance feedback method from Information Retrieval. Furthermore, we explore methods on concept level feedback, such as the re-weighting and Query Point Modification (QPM) methods as well as a method that changes the semantic space the concepts are represented in. Methods on both concept level and video level are evaluated on the international benchmark TRECVID Multimedia Event Detection (MED) and compared to state of the art methods. Results show that relevance feedback on both concept and video level improves performance compared to using no relevance feedback; relevance feedback on video level obtains higher performance compared to relevance feedback on concept level; our proposed ARF method on video level outperforms a state of the art k-NN method, all methods on concept level and even manually selected concepts. Maaike de Boer, Geert Pingen, Douwe Knook, Klamer Schutte, Wessel Kraaij |
Multim. Tools Appl. | 5 |
| 2017 | Semantic Reasoning in Zero Example Video Event RetrievalabstractSearching in digital video data for high-level events, such as a parade or a car accident, is challenging when the query is textual and lacks visual example images or videos. Current research in deep neural networks is highly beneficial for the retrieval of high-level events using visual examples, but without examples it is still hard to (1) determine which concepts are useful to pre-train ( Vocabulary challenge ) and (2) which pre-trained concept detectors are relevant for a certain unseen high-level event ( Concept Selection challenge ). In our article, we present our Semantic Event Retrieval System which (1) shows the importance of high-level concepts in a vocabulary for the retrieval of complex and generic high-level events and (2) uses a novel concept selection method ( i-w2v ) based on semantic embeddings. Our experiments on the international TRECVID Multimedia Event Detection benchmark show that a diverse vocabulary including high-level concepts improves performance on the retrieval of high-level events in videos and that our novel method outperforms a knowledge-based concept selection method. Maaike de Boer, Yi-Jie Lu, Hao Zhang 0047, Klamer Schutte, Chong-Wah Ngo, Wessel Kraaij |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2016 | Longitudinal Navigation Log Data on a Large Web DomainabstractWe have collected the access logs for our university's web domain over a time span of 4.5 years. We now release the pre-processed data of a 3-month period for research into user navigation behavior. We preprocessed the data so that only successful GET requests of web pages by non-bot users are kept. The resulting 3-month collection comprises 9.6M page visits (190K unique URLs) by 744K unique visitors. Suzan Verberne, Bram Arends, Wessel Kraaij, Arjen P. de Vries |
SIGIR | 3 |
| 2016 | Evaluation and analysis of term scoring methods for term extractionabstractWe evaluate five term scoring methods for automatic term extraction on four different types of text collections: personal document collections, news articles, scientific articles and medical discharge summaries. Each collection has its own use case: author profiling, boolean query term suggestion, personalized query suggestion and patient query expansion. The methods for term scoring that have been proposed in the literature were designed with a specific goal in mind. However, it is as yet unclear how these methods perform on collections with characteristics different than what they were designed for, and which method is the most suitable for a given (new) collection. In a series of experiments, we evaluate, compare and analyse the output of six term scoring methods for the collections at hand. We found that the most important factors in the success of a term scoring method are the size of the collection and the importance of multi-word terms in the domain. Larger collections lead to better terms; all methods are hindered by small collection sizes (below 1000 words). The most flexible method for the extraction of single-word and multi-word terms is pointwise Kullback–Leibler divergence for informativeness and phraseness. Overall, we have shown that extracting relevant terms using unsupervised term scoring methods is possible in diverse use cases, and that the methods are applicable in more contexts than their original design purpose. Suzan Verberne, Maya Sappelli, Djoerd Hiemstra, Wessel Kraaij |
Inf. Retr. J. | 4 |
| 2016 | Assessing e-mail intent and tasks in e-mail messages
Maya Sappelli, Gabriella Pasi, Suzan Verberne, Maaike de Boer, Wessel Kraaij |
Inf. Sci. | 5 |
| 2016 | Knowledge based query expansion in complex multimedia event detectionabstractA common approach in content based video information retrieval is to perform automatic shot annotation with semantic labels using pre-trained classifiers. The visual vocabulary of state-of-the-art automatic annotation systems is limited to a few thousand concepts, which creates a semantic gap between the semantic labels and the natural language query. One of the methods to bridge this semantic gap is to expand the original user query using knowledge bases. Both common knowledge bases such as Wikipedia and expert knowledge bases such as a manually created ontology can be used to bridge the semantic gap. Expert knowledge bases have highest performance, but are only available in closed domains. Only in closed domains all necessary information, including structure and disambiguation, can be made available in a knowledge base. Common knowledge bases are often used in open domain, because it covers a lot of general information. In this research, query expansion using common knowledge bases ConceptNet and Wikipedia is compared to an expert description of the topic applied to content-based information retrieval of complex events. We run experiments on the Test Set of TRECVID MED 2014. Results show that 1) Query Expansion can improve performance compared to using no query expansion in the case that the main noun of the query could not be matched to a concept detector; 2) Query expansion using expert knowledge is not necessarily better than query expansion using common knowledge; 3) ConceptNet performs slightly better than Wikipedia; 4) Late fusion can slightly improve performance. To conclude, query expansion has potential in complex event detection. Maaike de Boer, Klamer Schutte, Wessel Kraaij |
Multim. Tools Appl. | 3 |
| 2016 | Adapting the Interactive Activation Model for Context Recognition and Identification
Maya Sappelli, Suzan Verberne, Wessel Kraaij |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2015 | User Simulations for Interactive Search: Evaluating Personalized Query Suggestion
Suzan Verberne, Maya Sappelli, Kalervo Järvelin, Wessel Kraaij |
ECIR | 4 |
| 2014 | Query Term Suggestion in Academic Search
Suzan Verberne, Maya Sappelli, Wessel Kraaij |
ECIR | 3 |
| 2014 | The SWELL Knowledge Work Dataset for Stress and User Modeling ResearchabstractThis paper describes the new multimodal SWELL knowledge work (SWELL-KW) dataset for research on stress and user modeling. The dataset was collected in an experiment, in which 25 people performed typical knowledge work (writing reports, making presentations, reading e-mail, searching for information). We manipulated their working conditions with the stressors: email interruptions and time pressure. A varied set of data was recorded: computer logging, facial expression from camera recordings, body postures from a Kinect 3D sensor and heart rate (variability) and skin conductance from body sensors. The dataset made available not only contains raw data, but also preprocessed data and extracted features. The participants' subjective experience on task load, mental effort, emotion and perceived stress was assessed with validated questionnaires as a ground truth. The resulting dataset on working behavior and affect is a valuable contribution to several research fields, such as work psychology, user modeling and context aware systems. Saskia Koldijk, Maya Sappelli, Suzan Verberne, Mark A. Neerincx, Wessel Kraaij |
ICMI | 5 |
| 2014 | Privacy and User Trust in Context-Aware Systems
Saskia Koldijk, Gijs Koot, Mark A. Neerincx, Wessel Kraaij |
UMAP | 4 |
| 2014 | Requirements for multimedia metadata schemes in surveillance applications for security
Jeroen van Rest, F. A. Grootjen, Marc Grootjen, Remco Wijn, Olav Aarts, M. L. Roelofs, Gertjan J. Burghouts, Henri Bouma, Lejla Alic, Wessel Kraaij |
Multim. Tools Appl. | 10 |
| 2014 | Content-Based Video Copy Detection Benchmarking at TRECVIDabstractThis article presents an overview of the video copy detection benchmark which was run over a period of 4 years (2008--2011) as part of the TREC Video Retrieval (TRECVID) workshop series. The main contributions of the article include i) an examination of the evolving design of the evaluation framework and its components (system tasks, data, measures); ii) a high-level overview of results and best-performing approaches; and iii) a discussion of lessons learned over the four years. The content-based copy detection (CCD) benchmark worked with a large collection of synthetic queries, which is atypical for TRECVID, as was the use of a normalized detection cost framework. These particular evaluation design choices are motivated and appraised. George Awad, Paul Over, Wessel Kraaij |
ACM Trans. Inf. Syst. | 3 |
| 2013 | Recommending personalized touristic sights using google placesabstractThe purpose of the Contextual Suggestion track, an evaluation task at the TREC 2012 conference, is to suggest personalized tourist activities to an individual, given a certain location and time. In our content-based approach, we collected initial recommendations using the location context as search query in Google Places. We first ranked the recommendations based on their textual similarity to the user profiles. In order to improve the ranking of popular sights, we combined the initial ranking with rankings based on Google Search, popularity and categories. Finally, we performed filtering based on the temporal context. Overall, our system performed well above average and median, and outperformed the baseline - Google Places only -- run. Maya Sappelli, Suzan Verberne, Wessel Kraaij |
SIGIR | 3 |
| 2013 | Unobtrusive Monitoring of Knowledge Workers for Stress Self-regulation
Saskia Koldijk, Maya Sappelli, Mark A. Neerincx, Wessel Kraaij |
UMAP | 4 |
| 2013 | Reliability and validity of query intent assessmentsabstractIn most intent recognition studies, annotations of query intent are created post hoc by external assessors who are not the searchers themselves. It is important for the field to get a better understanding of the quality of this process as an approximation for determining the searcher's actual intent. Some studies have investigated the reliability of the query intent annotation process by measuring the interassessor agreement. However, these studies did not measure the validity of the judgments, that is, to what extent the annotations match the searcher's actual intent. In this study, we asked both the searchers themselves and external assessors to classify queries using the same intent classification scheme. We show that of the seven dimensions in our intent classification scheme, four can reliably be used for query annotation. Of these four, only the annotations on the topic and spatial sensitivity dimension are valid when compared with the searcher's annotations. The difference between the interassessor agreement and the assessor‐searcher agreement was significant on all dimensions, showing that the agreement between external assessors is not a good estimator of the validity of the intent classifications. Therefore, we encourage the research community to consider using query intent classifications by the searchers themselves as test data. Suzan Verberne, Maarten van der Heijden, Max Hinne, Maya Sappelli, Saskia Koldijk, Eduard Hoenkamp, Wessel Kraaij |
J. Assoc. Inf. Sci. Technol. | 7 |
| 2012 | Special issue on searching speechabstractNo abstract available. Martha A. Larson, Franciska de Jong, Wessel Kraaij, Steve Renals |
ACM Trans. Inf. Syst. | 3 |
| 2011 | Bringing Why-QA to Web Search
Suzan Verberne, Lou Boves, Wessel Kraaij |
ECIR | 3 |
| 2011 | Speech Transcript Evaluation for Information RetrievalabstractSpeech recognition transcripts are being used in various fields of research and practical applications, putting various demands on their accuracy. Traditionally ASR research has used intrinsic evaluation measures such as word error rate to determine transcript quality. In non-dictation-type applications such as speech retrieval, it is better to use extrinsic (or task specific) measures. Indexation and the associated processing may eliminate certain errors, whereas the search query may reveal others. In this work, we argue that the standard extrinsic speech retrieval measure average precision is unpractical for ASR evaluation. As an alternative we propose the use of ranked correlation measures on the output of the speech retrieval task, with the goal of predicting relative mean average precision. The measures we used showed a reasonably high correlation with average precision, but require much less human effort to calculate and can be more easily deployed in a variety of real-life settings. Laurens van der Werff, Wessel Kraaij, Franciska de Jong |
INTERSPEECH | 2 |
| 2010 | A cross-lingual framework for monolingual biomedical information retrievalabstractAn important challenge for biomedical information retrieval (IR) is dealing with the complex, inconsistent and ambiguous biomedical terminology. Frequently, a concept-based representation defined in terms of a domain-specific terminological resource is employed to deal with this challenge. In this paper, we approach the incorporation of a concept-based representation in monolingual biomedical IR from a cross-lingual perspective. In the proposed framework, this is realized by translating and matching between text and concept-based representations. The approach allows for deployment of a rich set of techniques proposed and evaluated in traditional cross-lingual IR. We compare six translation models and measure their effectiveness in the biomedical domain. We demonstrate that the approach can result in significant improvements in retrieval effectiveness over word-based retrieval. Moreover, we demonstrate increased effectiveness of a CLIR framework for monolingual biomedical IR if basic translations models are combined. © 2010 ACM. Dolf Trieschnigg, Djoerd Hiemstra, Franciska de Jong, Wessel Kraaij |
CIKM | 4 |
| 2010 | Classifier Calibration for Multi-Domain Sentiment Classification
Stephan Raaijmakers, Wessel Kraaij |
ICWSM | 2 |
| 2010 | Multimedia content with a speech track: ACM multimedia 2010 workshop on searching spontaneous conversational speechabstractNo abstract available. Martha A. Larson, Roeland Ordelman, Florian Metze, Wessel Kraaij, Franciska de Jong |
ACM Multimedia | 4 |
| 2010 | Conceptual language models for domain-specific retrieval
Edgar Meij, Dolf Trieschnigg, Maarten de Rijke, Wessel Kraaij |
Inf. Process. Manag. | 4 |
| 2009 | Searching multimedia content with a spontaneous conversational speech trackabstractNo abstract available. Martha A. Larson, Roeland Ordelman, Franciska de Jong, Wessel Kraaij, Joachim Köhler |
ACM Multimedia | 4 |
| 2009 | Annotation of URLs: more than the sum of partsabstractRecently a number of studies have demonstrated that search engine logfiles are an important resource to determine the relevance relation between URLs and query terms. We hypothesized that the queries associated with a URL could also be presented as useful URL metadata in a search engine result list, e.g. for helping to determine the semantic category of a URL. We evaluated this hypothesis by a classification experiment based on the DMOZ dataset. Our method can also annotate URLs that have no associated queries. Max Hinne, Wessel Kraaij, Stephan Raaijmakers, Suzan Verberne, Theo P. van der Weide, Maarten van der Heijden |
SIGIR | 2 |
| 2009 | MeSH Up: effective MeSH text classification for improved document retrievalabstractMOTIVATION: Controlled vocabularies such as the Medical Subject Headings (MeSH) thesaurus and the Gene Ontology (GO) provide an efficient way of accessing and organizing biomedical information by reducing the ambiguity inherent to free-text data. Different methods of automating the assignment of MeSH concepts have been proposed to replace manual annotation, but they are either limited to a small subset of MeSH or have only been compared with a limited number of other systems. RESULTS: We compare the performance of six MeSH classification systems [MetaMap, EAGL, a language and a vector space model-based approach, a K-Nearest Neighbor (KNN) approach and MTI] in terms of reproducing and complementing manual MeSH annotations. A KNN system clearly outperforms the other published approaches and scales well with large amounts of text using the full MeSH thesaurus. Our measurements demonstrate to what extent manual MeSH annotations can be reproduced and how they can be complemented by automatic annotations. We also show that a statistically significant improvement can be obtained in information retrieval (IR) when the text of a user's query is automatically annotated with MeSH concepts, compared to using the original textual query alone. CONCLUSIONS: The annotation of biomedical texts using controlled vocabularies such as MeSH can be automated to improve text-only IR. Furthermore, the automatic MeSH annotation system we propose is highly scalable and it generates improvements in IR comparable with those observed for manual annotations. Dolf Trieschnigg, Piotr Pezik, Vivian Lee, Franciska de Jong, Wessel Kraaij, Dietrich Rebholz-Schuhmann |
Bioinform. | 5 |
| 2009 | Response to comment on 'MeSH-up: effective MeSH text classification for improved document retrieval'abstractAbstract Contact: [email protected]; [email protected] As developers and primary users of MTI and MetaMap, Névéol et al. made a number of interesting comments on our recent publication in Bioinformatics. However, some of the results and conclusions found in the reply seem premature and lack proper clarification. Dolf Trieschnigg, Piotr Pezik, Vivian Lee, Franciska de Jong, Wessel Kraaij, Dietrich Rebholz-Schuhmann |
Bioinform. | 5 |
| 2008 | A Shallow Approach to Subjectivity Classification
Stephan Raaijmakers, Wessel Kraaij |
ICWSM | 2 |
| 2008 | Parsimonious concept modelingabstractNo abstract available. Edgar Meij, Dolf Trieschnigg, Maarten de Rijke, Wessel Kraaij |
SIGIR | 4 |
| 2008 | Measuring concept relatedness using language modelsabstractOver the years, the notion of concept relatedness has attracted considerable attention. A variety of approaches, based on ontology structure, information content, association, or context have been proposed to indicate the relatedness of abstract ideas. We propose a method based on the cross entropy reduction between language models of concepts which are estimated based on document-concept assignments. The approach shows improved or competitive results compared to state-of-the-art methods on two test sets in the biomedical domain. Dolf Trieschnigg, Edgar Meij, Maarten de Rijke, Wessel Kraaij |
SIGIR | 4 |
| 2007 | The influence of basic tokenization on biomedical document retrievalabstractTokenization is a fundamental preprocessing step in Information Retrieval systems in which text is turned into index terms. This paper quantifies and compares the influence of various simple tokenization techniques on document retrieval effectiveness in two domains: biomedicine and news. As expected, biomedical retrieval is more sensitive to small changes in the tokenization method. The tokenization strategy can make the difference between a mediocre and well performing IR system, especially in the biomedical domain. Dolf Trieschnigg, Wessel Kraaij, Franciska de Jong |
SIGIR | 2 |
| 2005 | Scalable hierarchical topic detection: exploring a sample based approachabstractHierarchical topic detection is a new task in the TDT 2004 evaluation program, which aims to organize an unstructured news collection in a directed acyclic graph (DAG) structure, reflecting the topics discussed. We present a scalable architecture for HTD and compare several alternative choices for agglomerative clustering and DAG optimization in order to minimize the HTD cost metric. Dolf Trieschnigg, Wessel Kraaij |
SIGIR | 2 |
| 2004 | TRECVID: evaluating the effectiveness of information retrieval tasks on digital videoabstractTRECVID is an annual exercise which encourages research in information retrieval from digital video by providing a large video test collection, uniform scoring procedures, and a forum for organizations interested in comparing their results. TRECVID benchmarking covers both interactive and manual searching by end users, as well as the benchmarking of some supporting technologies including shot boundary detection, extraction of some semantic features, and the automatic segmentation of TV news broadcasts into non-overlapping news stories. TRECVID has a broad range of over 40 participating groups from across the world and as it is now (2004) in its 4th annual cycle it is opportune to stand back and look at the lessons we have learned from the cumulative activity. In this paper we shall present a brief and high-level overview of the TRECVID activity covering the data, the benchmarked tasks, the overall results obtained by groups to date and an overview of the approaches taken by selective groups in some tasks. While progress from one year to the next cannot be measured directly because of the changing nature of the video data we have been using, we shall present a summary of the lessons we have learned from TRECVID and include some pointers on what we feel are the most important of these lessons. Alan F. Smeaton, Paul Over, Wessel Kraaij |
ACM Multimedia | 3 |
| 2003 | Embedding Web-Based Statistical Translation Models in Cross-Language Information RetrievalabstractAlthough more and more language pairs are covered by machine translation (MT) services, there are still many pairs that lack translation resources. Cross-language information retrieval (CLIR) is an application that needs translation functionality of a relatively low level of sophistication, since current models for information retrieval (IR) are still based on a bag of words. The Web provides a vast resource for the automatic construction of parallel corpora that can be used to train statistical translation models automatically. The resulting translation models can be embedded in several ways in a retrieval model. In this article, we will investigate the problem of automatically mining parallel texts from the Web and different ways of integrating the translation models within the retrieval process. Our experiments on standard test collections for CLIR show that the Web-based translation models can surpass commercial MT systems in CLIR tasks. These results open the perspective of constructing a fully automatic query translation device for CLIR at a very low cost. Wessel Kraaij, Jian-Yun Nie, Michel Simard |
Comput. Linguistics | 1 |
| 2002 | The Importance of Prior Probabilities for Entry Page SearchabstractAn important class of searches on the world-wide-web has the goal to find an entry page (homepage) of an organisation. Entry page search is quite different from Ad Hoc search. Indeed a plain Ad Hoc system performs disappointingly. We explored three non-content features of web pages: page length, number of incoming links and URL form. Especially the URL form proved to be a good predictor. Using URL form priors we found over 70% of all entry pages at rank 1, and up to 89% in the top 10. Non-content features can easily be embedded in a language model framework as a prior probability. Wessel Kraaij, Thijs Westerveld, Djoerd Hiemstra |
SIGIR | 1 |
| 1998 | Twenty-One: Cross-Language Disclosure and Retrieval of Multimedia Documents on Sustainable Development
Wilco G. ter Stal, J.-H. Beijert, G. de Bruin, J. van Gent, Franciska de Jong, Wessel Kraaij, Klaus Netter, G. Smart |
Comput. Networks | 6 |
| 1996 | Viewing Stemming as Recall EnhancementabstractPrevious research on stemming has shown both positive and negative effects on retrieval performance.This paper describes an experiment in which several linguistic and non-linguistic stemmers are evaluated on a Dutch test collection.Experiments especially focus on the measurement of Recall.Results show that linguistic stemming restricted to inflection yields a significant improvement over full linguistic and non-linguistic stemming, both in average Precision and R-Recall.Best results are obtained with a linguistic stemmer which is enhanced with compound analysis.This version has a significantly better Recall than a system without stemming, without a significant deterioration of Precision. Wessel Kraaij, Renée Pohlmann |
SIGIR | 1 |
| 1990 | Ambiguity resolution and the retrieval of idioms: two approaches
Erik-Jan van der Linden, Wessel Kraaij |
COLING | 2 |