Kevyn Collins-Thompson

dblp:42/4842 · DBLP profile ↗
← Back
49ranked-venue papers
15as first author
3since 2021 · last 2026
0000-0002-8178-5035ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 33 · 11 first-authorArtificial intelligence and machine learning · 22 · 7 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSystems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
21 papers
Information retrieval · 72% Recommender systems · 15% Data mining · 8%
Human-computer interaction and pervasive computing
3 papers
Learning and educational technologies · 82% Wearable and physiological sensing · 11% Usability and user experience research · 6%

Topics — the 30 heaviest of 58, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval models
0.742017
Retrieval Algorithms Optimized for Human Learning · SIGIR 2017
Copulas for information retrieval · SIGIR 2013
Characterizing web content, user interests, and search behavior by reading level and topic · WSDM 2012
Information retrieval
web search
0.532013
Toward whole-session relevance: exploring intrinsic diversity in web search · SIGIR 2013
Shame to be sham: addressing content-based grey hat search engine optimization · SIGIR 2013
Probabilistic models for personalizing web search · WSDM 2012
Learning and educational technologies › educational content generation
question generation
0.412020
Improving Learning Outcomes with Gaze Tracking and Automatic Question Generation · WWW 2020
Information retrieval
personalized search
0.422017
Retrieval Algorithms Optimized for Human Learning · SIGIR 2017
Probabilistic models for personalizing web search · WSDM 2012
Recommender systems
cross-domain recommendation
0.312018
Demographic Inference Via Knowledge Transfer in Cross-Domain Recommender Systems · ICDM 2018
Web and social media mining › social media analysis
demographic inference
0.312018
Demographic Inference Via Knowledge Transfer in Cross-Domain Recommender Systems · ICDM 2018
Recommender systems › cross-domain recommendation
knowledge transfer
0.312018
Demographic Inference Via Knowledge Transfer in Cross-Domain Recommender Systems · ICDM 2018
Information retrieval › user interaction
personalization
0.312017
Retrieval Algorithms Optimized for Human Learning · SIGIR 2017
Information retrieval › text analysis
readability assessment
0.212016
Predicting the Relative Difficulty of Single Sentences With and Without Surrounding Context · EMNLP 2016
Recommender systems › context-aware recommendation
contextual suggestion
0.212015
Query Suggestion and Data Fusion in Contextual Disambiguation · WWW 2015
Information retrieval › query understanding
query disambiguation
0.212015
Query Suggestion and Data Fusion in Contextual Disambiguation · WWW 2015
Information retrieval
query suggestion
0.212015
Query Suggestion and Data Fusion in Contextual Disambiguation · WWW 2015
Information retrieval › user behavior
search session analysis
0.222013
Toward whole-session relevance: exploring intrinsic diversity in web search · SIGIR 2013
Personalizing atypical web search sessions · WSDM 2013
Information retrieval
retrieval evaluation
0.222017
Visualizing differences in web search algorithms using the expected weighted hoeffding distance · WWW 2010
Retrieval Algorithms Optimized for Human Learning · SIGIR 2017
Information retrieval › query understanding
query intent
0.212014
Understanding Intrinsic Diversity in Web Search: Improving Whole-Session Relevance · ACM Trans. Inf. Syst. 2014
Information retrieval › user behavior
search logs
0.212014
Understanding Intrinsic Diversity in Web Search: Improving Whole-Session Relevance · ACM Trans. Inf. Syst. 2014
Information retrieval › interactive information retrieval
session search
0.212014
Understanding Intrinsic Diversity in Web Search: Improving Whole-Session Relevance · ACM Trans. Inf. Syst. 2014
Recommender systems
user modeling
0.222012
Characterizing web content, user interests, and search behavior by reading level and topic · WSDM 2012
Probabilistic models for personalizing web search · WSDM 2012
Data mining › crowdsourcing
crowdsourced ranking
0.212013
Pairwise ranking aggregation in a crowdsourced setting · WSDM 2013
Information retrieval › retrieval models
probabilistic retrieval model
0.212013
Copulas for information retrieval · SIGIR 2013
Information retrieval › ranking
rank aggregation
0.212013
Pairwise ranking aggregation in a crowdsourced setting · WSDM 2013
Information retrieval
ranking
0.212013
Pairwise ranking aggregation in a crowdsourced setting · WSDM 2013
Information retrieval › web search
search engine optimization
0.212013
Shame to be sham: addressing content-based grey hat search engine optimization · SIGIR 2013
Information retrieval › web search
search personalization
0.212013
Personalizing atypical web search sessions · WSDM 2013
Data mining › anomaly detection
spam detection
0.212013
Shame to be sham: addressing content-based grey hat search engine optimization · SIGIR 2013
Information retrieval › query reformulation
query expansion
0.222008
Estimating Robust Query Models with Convex Optimization · NIPS 2008
Estimation and use of uncertainty in pseudo-relevance feedback · SIGIR 2007
Information retrieval › ranking
learning to rank
0.112012
Robust ranking models via risk-sensitive optimization · SIGIR 2012
Information retrieval
query understanding
0.112012
Characterizing web content, user interests, and search behavior by reading level and topic · WSDM 2012
Information retrieval › ranking › learning to rank
robust ranking
0.112012
Robust ranking models via risk-sensitive optimization · SIGIR 2012
Data mining › text mining
topic modeling
0.112012
Characterizing web content, user interests, and search behavior by reading level and topic · WSDM 2012

Methods — techniques the papers use, named apart from their topics

user study · 0.5logistic regression · 0.5interaction log analysis · 0.5bayesian rating system · 0.5gaze tracking · 0.4machine learning · 0.4probabilistic matrix factorization · 0.3optimization · 0.3cognitive learning model · 0.3bayesian inference · 0.2active learning · 0.2risk-sensitive optimization · 0.1learning-to-rank · 0.1statistical estimation · 0.1metric labeling · 0.1convex optimization · 0.1pipeline evaluation · 0.0
YearPublicationVenuePosition
2026 Measuring Simulation Fidelity via Statistical Detectability: A Diagnostic Framework for AI-Generated Tutoring Conversations
Michael Ion, Kevyn Collins-Thompson
L@S2
2024 Finding Educationally Supportive Contexts for Vocabulary Learning with Attention-Based Models
abstract
When learning new vocabulary, both humans and machines acquire critical information about the meaning of an unfamiliar word through contextual information in a sentence or passage. However, not all contexts are equally helpful for learning an unfamiliar ‘target’ word. Some contexts provide a rich set of semantic clues to the target word’s meaning, while others are less supportive. We explore the task of finding educationally supportive contexts with respect to a given target word for vocabulary learning scenarios, particularly for improving student literacy skills. Because of their inherent context-based nature, attention-based deep learning methods provide an ideal starting point. We evaluate attention-based approaches for predicting the amount of educational support from contexts, ranging from a simple custom model using pre-trained embeddings with an additional attention layer, to a commercial Large Language Model (LLM). Using an existing major benchmark dataset for educational context support prediction, we found that a sophisticated but generic LLM had poor performance, while a simpler model using a custom attention-based approach achieved the best-known performance to date on this dataset.
SungJin Nam, Kevyn Collins-Thompson, David Jurgens
LREC/COLING2
2024 Generation and Assessment of Multiple-Choice Questions from Video Transcripts using Large Language Models
abstract
We present an empirical study evaluating the quality of multiple-choice questions (MCQs) generated by Large Language Models (LLMs) from a corpus of video transcripts of course lectures in an online data science degree program. With our database of thousands of generated questions, we conducted both human and automated judging of question quality on a representative sample using a broad set of criteria, including well-established Item Writing Flaw (IWF) categories. We found the number of average IWFs per MCQ ranged from 1.6 (rule-based verification) to 2.18 (LLM-based). Among the most frequently identified MCQ flaws were lack of enough context (17%) or answer choices with at least one implausible distractor (57%). Both human and automated assessment identified implausible distractors as one of the most frequent flaw categories. Results from our human annotation study were generally more positive (51--65% good items) compared to our automated assessment study results, which tended toward greater flaw identification (15--25% good items), depending on evaluation method.
Taimoor Arif, Sumit Asthana, Kevyn Collins-Thompson
L@S3
2020 Improving Learning Outcomes with Gaze Tracking and Automatic Question Generation
abstract
As AI technology advances, it offers promising opportunities to improve educational outcomes when integrated with an overall learning experience. We investigate forward-looking interactive reading experiences that leverage both automatic question generation and analysis of attention signals, such as gaze tracking, to improve short- and long-term learning outcomes. We aim to expand the known pedagogical benefits of adjunct questions to more general reading scenarios, by investigating the benefits of adjunct questions generated after participants attend to passages in an article, based on their gaze behavior. We also compare the effectiveness of manually-written questions with those produced by Automatic Question Generation (AQG). We further investigate gaze and reading patterns indicative of low vs. high learning in both short- and long-term scenarios (one-week followup). We show AQG-generated adjunct questions have promise as a way to scale to a wide variety of reading material where the cost of manually curating questions may be prohibitive.
Rohail Syed, Kevyn Collins-Thompson, Paul N. Bennett, Mengqiu Teng, Shane Williams, Wendy W. Tay, Shamsi T. Iqbal
WWW2
2018 Exploring Document Retrieval Features Associated with Improved Short- and Long-term Vocabulary Learning Outcomes
abstract
A growing body of information retrieval research has studied the potential of search engines as effective, scalable platforms for self-directed learning. Towards this goal, we explore document representations for retrieval that include features associated with effective learning outcomes. While prior studies have investigated different retrieval models designed for teaching, this study is the first to investigate how document-level features are associated with actual learning outcomes when users get results from a personalized learning-oriented retrieval algorithm. We also conduct what is, to our knowledge, the first crowdsourced longitudinal study of long-term learning retention, in which we gave a subset of users who participated in an initial learning and assessment study a delayed post-test approximately nine months later. With this data, we were able to analyze how the three retrieval conditions in the original study were associated with changes in long-term vocabulary knowledge. We found that while users who read the documents in the personalized retrieval condition had immediate learning gains comparable to the other two conditions, they had better long-term retention of more difficult vocabulary.
Rohail Syed, Kevyn Collins-Thompson
CHIIR2
2018 Demographic Inference Via Knowledge Transfer in Cross-Domain Recommender Systems
abstract
User demographics such as age and gender are very useful in recommender systems for applications such as personalization services and marketing, but may not always be available for individual users. Existing approaches can infer users' private demographics based on ratings, given labeled data from users who share demographics. However, such labeled information is not always available in many e-commerce services, particularly small online retailers and most media sites, for which no user registration is required. We introduce a novel probabilistic matrix factorization model for demographic transfer that enables knowledge transfer from the source domain, in which users' ratings and the corresponding demographics are available, to the target domain, in which we would like to infer unknown user demographics from ratings. Our proposed method is based on two observations: (1) Items from different but related domains may share the same latent factors such as genres and styles, and (2) Users who share similar demographics are likely to prefer similar genres across domains. This approach can align latent factors across domains that share neither common users nor common items, associating user demographics with latent factors in a unified framework. Experiments on cross-domain datasets demonstrate that the proposed method consistently improves demographic classification accuracy over existing methods.
Jin Shang 0001, Mingxuan Sun 0001, Kevyn Collins-Thompson
ICDM3
2017 Social work in the classroom? A tool to evaluate topical relevance in student writing
Heeryung Choi, Zijian Wang 0002, Christopher Brooks 0001, Kevyn Collins-Thompson, Beth Glover Reed, Dale Fitch
EDM4
2017 Predicting Short- and Long-Term Vocabulary Learning via Semantic Features of Partial Word Knowledge
SungJin Nam, Gwen A. Frishkoff, Kevyn Collins-Thompson
EDM3
2017 What does student writing tell us about their thinking on social justice?
abstract
In this work we investigate the use of deep learning for text analysis to measure elements of student thinking related to issues of privilege, oppression, diversity and social justice. We leverage historical expert annotations as well as a large lexical model to create a more generalizable vocabulary for identifying these characteristics in short student writing. We demonstrate the feasibility of this approach, and identify further areas for research.
Heeryung Choi, Christopher Brooks 0001, Kevyn Collins-Thompson
LAK3
2017 Retrieval Algorithms Optimized for Human Learning
abstract
While search technology is widely used for learning-oriented information needs, the results provided by popular services such as Web search engines are optimized primarily for generic relevance, not effective learning outcomes. As a result, the typical information trail that a user must follow while searching to achieve a learning goal may be an inefficient one involving unnecessarily easy or difficult content, or material that is irrelevant to actual learning progress relative to a user's existing knowledge. We address this problem by introducing a novel theoretical framework, algorithms, and empirical analysis of an information retrieval model that is optimized for learning outcomes instead of generic relevance. We do this by formulating an optimization problem that incorporates a cognitive learning model into a retrieval objective, and then give an algorithm for an efficient approximate solution to find the search results that represent the best 'training set' for a human learner. Our model can personalize results for an individual user's learning goals, as well as account for the effort required to achieve those goals for a given set of retrieval results. We investigate the effectiveness and efficiency of our retrieval framework relative to a commercial search engine baseline ('Google') through a crowdsourced user study involving a vocabulary learning task, and demonstrate the effectiveness of personalized results from our model on word learning outcomes.
Rohail Syed, Kevyn Collins-Thompson
SIGIR2
2017 Optimizing search results for human learning goals
Rohail Syed, Kevyn Collins-Thompson
Inf. Retr. J.2
2016 Assessing Learning Outcomes in Web Search: A Comparison of Tasks and Query Strategies
abstract
Users make frequent use of Web search for learning-related tasks, but little is known about how different Web search interaction strategies affect outcomes for learning-oriented tasks, or what implicit or explicit indicators could reliably be used to assess search-related learning on the Web. We describe a lab-based user study in which we investigated potential indicators of learning in web searching, effective query strategies for learning, and the relationship between search behavior and learning outcomes. Using questionnaires, analysis of written responses to knowledge prompts, and search log data, we found that searchers' perceived learning outcomes closely matched their actual learning outcomes; that the amount searchers wrote in post-search questionnaire responses was highly correlated with their cognitive learning scores; and that the time searchers spent per document while searching was also highly and consistently correlated with higher-level cognitive learning scores. We also found that of the three query interaction conditions we applied, an intrinsically diverse presentation of results was associated with the highest percentage of users achieving combined factual and conceptual knowledge gains. Our study provides deeper insight into which aspects of search interaction are most effective for supporting superior learning outcomes, and the difficult problem of how learning may be assessed effectively during Web search.
Kevyn Collins-Thompson, Soo Young Rieh, Carl C. Haynes, Rohail Syed
CHIIR1
2016 Predicting the Relative Difficulty of Single Sentences With and Without Surrounding Context
abstract
The problem of accurately predicting relative reading difficulty across a set of sentences arises in a number of important natural language applications, such as finding and curating effective usage examples for intelligent language tutoring systems. Yet while significant research has explored document- and passage-level reading difficulty, the special challenges involved in assessing aspects of readability for single sentences have received much less attention, particularly when considering the role of surrounding passages. We introduce and evaluate a novel approach for estimating the relative reading difficulty of a set of sentences, with and without surrounding context. Using different sets of lexical and grammatical features, we explore models for predicting pairwise relative difficulty using logistic regression, and examine rankings generated by aggregating pairwise difficulty labels using a Bayesian rating system to form a final ranking. We also compare rankings derived for sentences assessed with and without context, and find that contextual features can help predict differences in relative difficulty judgments across these two conditions.
Elliot Schumacher, Maxine Eskénazi, Gwen A. Frishkoff, Kevyn Collins-Thompson
EMNLP4
2016 User Behavior in Asynchronous Slow Search
abstract
Conventional Web search is predicated on returning results to users as quickly as possible. However, for some search tasks, users have reported a willingness to wait for the perfect set of results. In this work, we present the first study to analyze users' willingness to wait and their search success, when given a Web search system that embodies characteristics of slow search, where speed can be traded for an improvement in quality. We conducted a between-subjects user study involving tasks that required multiple queries to complete, providing a Web search system that gave users the option to additionally issue asynchronous queries for which results improve in relevance over time as users continued working. We analyze the resulting survey results and interaction log data to investigate how users spent their time while waiting, and how behavior and search outcomes changes when users are given the option of using a system with asynchronous slow search capabilities. We find that when given a slow search system, users are able to perceive the improvement in quality over time, and find tasks to be easier compared to a baseline conventional Web search system. Additionally, we find that users continue to issue their own queries and examine additional documents while the slow search queries are processed in the background, and use the slow search feature more effectively as they gain exposure to its behavior across tasks. Our study significantly advances our understanding of the benefits and tradeoffs involved in providing slow search scenarios for Web search.
Ryan Burton, Kevyn Collins-Thompson
SIGIR2
2016 Assessing the readability of ClinicalTrials.gov
abstract
OBJECTIVE: ClinicalTrials.gov serves critical functions of disseminating trial information to the public and helping the trials recruit participants. This study assessed the readability of trial descriptions at ClinicalTrials.gov using multiple quantitative measures. MATERIALS AND METHODS: The analysis included all 165,988 trials registered at ClinicalTrials.gov as of April 30, 2014. To obtain benchmarks, the authors also analyzed 2 other medical corpora: (1) all 955 Health Topics articles from MedlinePlus and (2) a random sample of 100,000 clinician notes retrieved from an electronic health records system intended for conveying internal communication among medical professionals. The authors characterized each of the corpora using 4 surface metrics, and then applied 5 different scoring algorithms to assess their readability. The authors hypothesized that clinician notes would be most difficult to read, followed by trial descriptions and MedlinePlus Health Topics articles. RESULTS: Trial descriptions have the longest average sentence length (26.1 words) across all corpora; 65% of their words used are not covered by a basic medical English dictionary. In comparison, average sentence length of MedlinePlus Health Topics articles is 61% shorter, vocabulary size is 95% smaller, and dictionary coverage is 46% higher. All 5 scoring algorithms consistently rated CliniclTrials.gov trial descriptions the most difficult corpus to read, even harder than clinician notes. On average, it requires 18 years of education to properly understand these trial descriptions according to the results generated by the readability assessment algorithms. DISCUSSION AND CONCLUSION: Trial descriptions at CliniclTrials.gov are extremely difficult to read. Significant work is warranted to improve their readability in order to achieve CliniclTrials.gov's goal of facilitating information dissemination and subject recruitment.
Danny T. Y. Wu, David A. Hanauer, Qiaozhu Mei, Patricia M. Clark, Lawrence C. An, Joshua Proulx, Qing T. Zeng, V. G. Vinod Vydiswaran, Kevyn Collins-Thompson, Kai Zheng 0002
J. Am. Medical Informatics Assoc.9
2016 Using the Crowd to Improve Search Result Ranking and the Search Experience
abstract
Despite technological advances, algorithmic search systems still have difficulty with complex or subtle information needs. For example, scenarios requiring deep semantic interpretation are a challenge for computers. People, on the other hand, are well suited to solving such problems. As a result, there is an opportunity for humans and computers to collaborate during the course of a search in a way that takes advantage of the unique abilities of each. While search tools that rely on human intervention will never be able to respond as quickly as current search engines do, recent research suggests that there are scenarios where a search engine could take more time if it resulted in a much better experience. This article explores how crowdsourcing can be used at query time to augment key stages of the search pipeline. We first explore the use of crowdsourcing to improve search result ranking. When the crowd is used to replace or augment traditional retrieval components such as query expansion and relevance scoring, we find that we can increase robustness against failure for query expansion and improve overall precision for results filtering. However, the gains that we observe are limited and unlikely to make up for the extra cost and time that the crowd requires. We then explore ways to incorporate the crowd into the search process that more drastically alter the overall experience. We find that using crowd workers to support rich query understanding and result processing appears to be a more worthwhile way to make use of the crowd during search. Our results confirm that crowdsourcing can positively impact the search experience but suggest that significant changes to the search process may be required for crowdsourcing to fulfill its potential in search systems.
Yubin Kim 0001, Kevyn Collins-Thompson, Jaime Teevan
ACM Trans. Intell. Syst. Technol.2
2015 Query Suggestion and Data Fusion in Contextual Disambiguation
abstract
Queries issued to a search engine are often under-specified or ambiguous. The user's search context or background may provide information that disambiguates their information need in order to automatically predict and issue a more effective query. The disambiguation can take place at different stages of the retrieval process. For instance, contextual query suggestions may be computed and recommended to users on the result page when appropriate, an approach that does not require modifying the original query's results. Alternatively, the search engine can attempt to provide efficient access to new relevant documents by injecting these documents directly into search results based on the user's context.
Milad Shokouhi, Marc Sloan, Paul N. Bennett, Kevyn Collins-Thompson, Siranush Sarkizova
WWW4
2015 Overview of the Special Issue on Contextual Search and Recommendation
abstract
editorial Free AccessOverview of the Special Issue on Contextual Search and Recommendation Editors: Paul N. Bennett View Profile , Kevyn Collins-Thompson View Profile , Diane Kelly View Profile , Ryen W. White View Profile , Yi Zhang View Profile Authors Info & Claims ACM Transactions on Information SystemsVolume 33Issue 1March 2015 Article No.: 1epp 1–7https://doi.org/10.1145/2691351Published:17 March 2015Publication History 12citation543DownloadsMetricsTotal Citations12Total Downloads543Last 12 Months32Last 6 weeks10 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Paul N. Bennett, Kevyn Collins-Thompson, Diane Kelly 0001, Ryen W. White, Yi Zhang 0001
ACM Trans. Inf. Syst.2
2014 Understanding Intrinsic Diversity in Web Search: Improving Whole-Session Relevance
abstract
Current research on Web search has focused on optimizing and evaluating single queries. However, a significant fraction of user queries are part of more complex tasks [Jones and Klinkner 2008] which span multiple queries across one or more search sessions [Liu and Belkin 2010; Kotov et al. 2011]. An ideal search engine would not only retrieve relevant results for a user's particular query but also be able to identify when the user is engaged in a more complex task and aid the user in completing that task [Morris et al. 2008; Agichtein et al. 2012]. Toward optimizing whole-session or task relevance, we characterize and address the problem of intrinsic diversity (ID) in retrieval [Radlinski et al. 2009], a type of complex task that requires multiple interactions with current search engines. Unlike existing work on extrinsic diversity [Carbonell and Goldstein 1998; Zhai et al. 2003; Chen and Karger 2006] that deals with ambiguity in intent across multiple users, ID queries often have little ambiguity in intent but seek content covering a variety of aspects on a shared theme. In such scenarios, the underlying needs are typically exploratory, comparative, or breadth-oriented in nature. We identify and address three key problems for ID retrieval: identifying authentic examples of ID tasks from post-hoc analysis of behavioral signals in search logs; learning to identify initiator queries that mark the start of an ID search task; and given an initiator query, predicting which content to prefetch and rank.
Karthik Raman 0001, Paul N. Bennett, Kevyn Collins-Thompson
ACM Trans. Inf. Syst.3
2013 Designing Human-Readable User Profiles for Search Evaluation
Carsten Eickhoff, Kevyn Collins-Thompson, Paul N. Bennett, Susan T. Dumais
ECIR2
2013 Copulas for information retrieval
abstract
In many domains of information retrieval, system estimates of document relevance are based on multidimensional quality criteria that have to be accommodated in a unidimensional result ranking. Current solutions to this challenge are often inconsistent with the formal probabilistic framework in which constituent scores were estimated, or use sophisticated learning methods that make it difficult for humans to understand the origin of the final ranking. To address these issues, we introduce the use of copulas, a powerful statistical framework for modeling complex multi-dimensional dependencies, to information retrieval tasks. We provide a formal background to copulas and demonstrate their effectiveness on standard IR tasks such as combining multidimensional relevance estimates and fusion of results from multiple search engines. We introduce copula-based versions of standard relevance estimators and fusion methods and show that these lead to significant performance improvements on several tasks, as evaluated on large-scale standard corpora, compared to their non-copula counterparts. We also investigate criteria for understanding the likely effect of using copula models in a given retrieval scenario.
Carsten Eickhoff, Arjen P. de Vries, Kevyn Collins-Thompson
SIGIR3
2013 Shame to be sham: addressing content-based grey hat search engine optimization
abstract
We present an initial study identifying a form of content-based grey hat search engine optimization, in which a Web page contains both potentially relevant content and manipulated content: we call such pages sham documents, because they lie in the grey area between 'ham' (clearly normal) and 'spam' (clearly fake). Sham documents are often ranked artificially high in response to certain queries, but also may contain some useful information and cannot be considered as absolute spam. We report a novel annotation effort performed with the ClueWeb09 benchmark where pages were labeled as being spam, sham, or legitimate content. Significant inter-annotator agreement rates support the claim that there are sham documents that are highly ranked by a very effective retrieval approach, yet are not spam. We also present an initial study of predictors that may indicate whether a query is the target of shamming.
Fiana Raiber, Kevyn Collins-Thompson, Oren Kurland
SIGIR2
2013 Toward whole-session relevance: exploring intrinsic diversity in web search
abstract
Current research on web search has focused on optimizing and evaluating single queries. However, a significant fraction of user queries are part of more complex tasks [20] which span multiple queries across one or more search sessions [26,24]. An ideal search engine would not only retrieve relevant results for a user's particular query but also be able to identify when the user is engaged in a more complex task and aid the user in completing that task [29,1]. Toward optimizing whole-session or task relevance, we characterize and address the problem of intrinsic diversity (ID) in retrieval [30], a type of complex task that requires multiple interactions with current search engines. Unlike existing work on extrinsic diversity [30] that deals with ambiguity in intent across multiple users, ID queries often have little ambiguity in intent but seek content covering a variety of aspects on a shared theme. In such scenarios, the underlying needs are typically exploratory, comparative, or breadth-oriented in nature. We identify and address three key problems for ID retrieval: identifying authentic examples of ID tasks from post-hoc analysis of behavioral signals in search logs; learning to identify initiator queries that mark the start of an ID search task; and given an initiator query, predicting which content to prefetch and rank.
Karthik Raman 0001, Paul N. Bennett, Kevyn Collins-Thompson
SIGIR3
2013 Pairwise ranking aggregation in a crowdsourced setting
abstract
Inferring rankings over elements of a set of objects, such as documents or images, is a key learning problem for such important applications as Web search and recommender systems. Crowdsourcing services provide an inexpensive and efficient means to acquire preferences over objects via labeling by sets of annotators. We propose a new model to predict a gold-standard ranking that hinges on combining pairwise comparisons via crowdsourcing. In contrast to traditional ranking aggregation methods, the approach learns about and folds into consideration the quality of contributions of each annotator. In addition, we minimize the cost of assessment by introducing a generalization of the traditional active learning scenario to jointly select the annotator and pair to assess while taking into account the annotator quality, the uncertainty over ordering of the pair, and the current model uncertainty. We formalize this as an active learning strategy that incorporates an exploration-exploitation tradeoff and implement it using an efficient online Bayesian updating scheme. Using simulated and real-world data, we demonstrate that the active learning strategy achieves significant reductions in labeling cost while maintaining accuracy.
Paul N. Bennett, Kevyn Collins-Thompson, Eric Horvitz
WSDM3
2013 Personalizing atypical web search sessions
abstract
Most research in Web search personalization models users as static or slowly evolving entities with a given set of preferences defined by their past behavior. However, recent publications as well as empirical evidence suggest that for a significant number of search sessions, users diverge from their regular search profiles in order to satisfy atypical, limited-duration information needs. In this work, we conduct a large-scale inspection of real-life search sessions to further understand this scenario. Subsequently, we design an automatic means of detecting and supporting such atypical sessions. We demonstrate significant improvements over state-of-the-art Web search personalization techniques by accounting for the typicality of search sessions. The proposed method is evaluated based on Web-scale search session data spanning several months of user activity.
Carsten Eickhoff, Kevyn Collins-Thompson, Paul N. Bennett, Susan T. Dumais
WSDM2
2012 Definition Response Scoring with Probabilistic Ordinal Regression
abstract
Word knowledge is often partial, rather than all-or-none. In this paper, we describe a method for estimating partial word knowledge on a trial-by-trial basis. Users generate a free-form synonym for a newly learned word. We then apply a probabilistic regression model that combines features based on Latent Semantic Analysis (LSA) with features derived from a large-scale, multi-relation word graph model to estimate the similarity of the user response to the actual meaning. This method allows us to predict multiple levels of accuracy, i.e., responses that precisely capture a word's meaning versus those that are partially correct or incorrect. We train and evaluate our approach using a new gold-standard corpus of expert responses, and find consistently superior performance compared to a state-of-the-art multi-class logistic regression baseline. These findings are a promising step toward a new kind of adaptive tutoring system that provides fine-grained, continuous feedback as learners acquire richer, more complete knowledge of words.
Kevyn Collins-Thompson, Gwen A. Frishkoff, Scott A. Crossley
ICCE1
2012 Robust ranking models via risk-sensitive optimization
abstract
Many techniques for improving search result quality have been proposed. Typically, these techniques increase average effectiveness by devising advanced ranking features and/or by developing sophisticated learning to rank algorithms. However, while these approaches typically improve average performance of search results relative to simple baselines, they often ignore the important issue of robustness. That is, although achieving an average gain overall, the new models often hurt performance on many queries. This limits their application in real-world retrieval scenarios. Given that robustness is an important measure that can negatively impact user satisfaction, we present a unified framework for jointly optimizing effectiveness and robustness. We propose an objective that captures the tradeoff between these two competing measures and demonstrate how we can jointly optimize for these two measures in a principled learning framework. Experiments indicate that ranking models learned this way significantly decreased the worst ranking failures while maintaining strong average effectiveness on par with current state-of-the-art models.
Paul N. Bennett, Kevyn Collins-Thompson
SIGIR3
2012 Characterizing web content, user interests, and search behavior by reading level and topic
abstract
A user's expertise or ability to understand a document on a given topic is an important aspect of that document's relevance. However, this aspect has not been well-explored in information retrieval systems, especially those at Web scale where the great diversity of content, users, and tasks presents an especially challenging search problem. To help improve our modeling and understanding of this diversity, we apply automatic text classifiers, based on reading difficulty and topic prediction, to estimate a novel type of profile for important entities in Web search -- users, websites, and queries. These profiles capture topic and reading level distributions, which we then use in conjunction with search log data to characterize and compare different entities.
Jin Young Kim 0005, Kevyn Collins-Thompson, Paul N. Bennett, Susan T. Dumais
WSDM2
2012 Probabilistic models for personalizing web search
abstract
We present a new approach for personalizing Web search results to a specific user. Ranking functions for Web search engines are typically trained by machine learning algorithms using either direct human relevance judgments or indirect judgments obtained from click-through data from millions of users. The rankings are thus optimized to this generic population of users, not to any specific user. We propose a generative model of relevance which can be used to infer the relevance of a document to a specific user for a search query. The user-specific parameters of this generative model constitute a compact user profile. We show how to learn these profiles from a user's long-term search history. Our algorithm for computing the personalized ranking is simple and has little computational overhead. We evaluate our personalization approach using historical search data from thousands of users of a major Web search engine. Our findings demonstrate gains in retrieval performance for queries with high ambiguity, with particularly large improvements for acronym queries.
David A. Sontag, Kevyn Collins-Thompson, Paul N. Bennett, Ryen W. White, Susan T. Dumais, Bodo Billerbeck
WSDM2
2011 Personalizing web search results by reading level
abstract
Traditionally, search engines have ignored the reading difficulty of documents and the reading proficiency of users in computing a document ranking. This is one reason why Web search engines do a poor job of serving an important segment of the population: children. While there are many important problems in interface design, content filtering, and results presentation related to addressing children's search needs, perhaps the most fundamental challenge is simply that of providing relevant results at the right level of reading difficulty. At the opposite end of the proficiency spectrum, it may also be valuable for technical users to find more advanced material or to filter out material at lower levels of difficulty, such as tutorials and introductory texts. We show how reading level can provide a valuable new relevance signal for both general and personalized Web search. We describe models and algorithms to address the three key problems in improving relevance for search using reading difficulty: estimating user proficiency, estimating result difficulty, and re-ranking based on the difference between user and result reading level profiles. We evaluate our methods on a large volume of Web query traffic and provide a large-scale log analysis that highlights the importance of finding results at an appropriate reading level for the user.
Kevyn Collins-Thompson, Paul N. Bennett, Ryen W. White, Sebastian de la Chica, David A. Sontag
CIKM1
2011 Statistical information retrieval modelling: from the probability ranking principle to recent advances in diversity, portfolio theory, and beyond
abstract
Statistical modelling of Information Retrieval (IR) systems is a key driving force in the development of the IR field. The goal of this tutorial is to provide a comprehensive and up-to-date introduction to statistical IR modelling. We take a fresh and systematic perspective from the viewpoint of portfolio theory of IR and risk management. A unified treatment and new insights will be given to reflect the recent developments of considering the ranked retrieval results as a whole. Recent research progress in diversification, risk management, and portfolio theory will be covered, in addition to classic methods such as Maron and Kuhns' Probabilistic Indexing, Robertson-Sparck Jones model (and the resulting BM25 formula) and language modelling approaches. The tutorial also reviews the resulting practical algorithms of risk-aware query expansion, diverse ranking, IR metric optimization as well as their performance evaluations. Practical IR applications such as web search, multimedia retrieval, and collaborative filtering are also introduced, as well as discussion of new opportunities for future research and applications that intersect among information retrieval, knowledge management, and databases.
Jun Wang 0012, Kevyn Collins-Thompson
CIKM2
2010 A unified optimization framework for robust pseudo-relevance feedback algorithms
abstract
We present a flexible new optimization framework for finding effective, reliable pseudo-relevance feedback models that unifies existing complementary approaches in a principled way. The result is an algorithmic approach that not only brings together different benefits of previous methods, such as parameter self-tuning and risk reduction from term dependency modeling, but also allows a rich new space of model search strategies to be investigated. We compare the effectiveness of a unified algorithm to existing methods by examining iterative performance and risk-reward tradeoffs. We also discuss extensions for generating new algorithms within our framework.
Joshua V. Dillon, Kevyn Collins-Thompson
CIKM2
2010 Predicting Query Performance via Classification
Kevyn Collins-Thompson, Paul N. Bennett
ECIR1
2010 Visualizing differences in web search algorithms using the expected weighted hoeffding distance
abstract
We introduce a new dissimilarity function for ranked lists, the expected weighted Hoeffding distance, that has several advantages over current dissimilarity measures for ranked search results. First, it is easily customized for users who pay varying degrees of attention to websites at different ranks. Second, unlike existing measures such as generalized Kendall's tau, it is based on a true metric, preserving meaningful embeddings when visualization techniques like multi-dimensional scaling are applied. Third, our measure can effectively handle partial or missing rank information while retaining a probabilistic interpretation. Finally, the measure can be made computationally tractable and we give a highly efficient algorithm for computing it. We then apply our new metric with multi-dimensional scaling to visualize and explore relationships between the result sets from different search engines, showing how the weighted Hoeffding distance can distinguish important differences in search engine behavior that are not apparent with other rank-distance metrics. Such visualizations are highly effective at summarizing and analyzing insights on which search engines to use, what search strategies users can employ, and how search results evolve over time. We demonstrate our techniques using a collection of popular search engines, a representative set of queries, and frequently used query manipulation methods.
Mingxuan Sun 0001, Guy Lebanon, Kevyn Collins-Thompson
WWW3
2009 Reducing the risk of query expansion via robust constrained optimization
abstract
We introduce a new theoretical derivation, evaluation methods, and extensive empirical analysis for an automatic query expansion framework in which model estimation is cast as a robust constrained optimization problem. This framework provides a powerful method for modeling and solving complex expansion problems, by allowing multiple sources of domain knowledge or evidence to be encoded as simultaneous optimization constraints. Our robust optimization approach provides a clean theoretical way to model not only expansion benefit, but also expansion risk, by optimizing over uncertainty sets for the data. In addition, we introduce risk-reward curves to visualize expansion algorithm performance and analyze parameter sensitivity. We show that a robust approach significantly reduces the number and magnitude of expansion failures for a strong baseline algorithm, with no loss in average gain. Our approach is implemented as a highly efficient post-processing step that assumes little about the baseline expansion method used as input, making it easy to apply to existing expansion methods. We provide analysis showing that this approach is a natural and effective way to do selective expansion, automatically reducing or avoiding expansion in risky scenarios, and successfully attenuating noise in poor baseline methods.
Kevyn Collins-Thompson
CIKM1
2009 Statistical Estimation of Word Acquisition with Application to Readability Prediction
Paul Kidwell, Guy Lebanon, Kevyn Collins-Thompson
EMNLP3
2009 Tutorial summary: Machine learning in IR: recent successes and new opportunities
abstract
No abstract available.
Paul N. Bennett, Mikhail Bilenko, Kevyn Collins-Thompson
ICML3
2009 Estimating query performance using class predictions
abstract
We investigate using topic prediction data, as a summary of document content, to compute measures of search result quality. Unlike existing quality measures such as query clarity that require the entire content of the top-ranked results, class-based statistics can be computed efficiently online, because class information is compact enough to precompute and store in the index. In an empirical study we compare the performance of class-based statistics to their language-model counterparts for predicting two measures: query difficulty and expansion risk. Our findings suggest that using class predictions can offer comparable performance to full language models while reducing computation overhead.
Kevyn Collins-Thompson, Paul N. Bennett
SIGIR1
2008 Estimating Robust Query Models with Convex Optimization
abstract
Query expansion is a long-studied approach for improving retrieval effectiveness by enhancing the user’s original query with additional related terms. Current algorithms for automatic query expansion have been shown to consistently improve retrieval accuracy on average, but are highly unstable and have bad worst-case performance for individual queries. We introduce a novel risk framework that formulates query model estimation as a constrained metric labeling problem on a graph of term relations. Themodel combines assignment costs based on a baseline feedback algorithm, edge weights based on term similarity, and simple constraints to enforce aspect balance, aspect coverage, and term centrality. Results across multiple standard test collections show consistent and dramatic reductions in the number and magnitude of expansion failures, while retaining the strong positive gains of the baseline algorithm.
Kevyn Collins-Thompson
NIPS1
2007 Automatic and Human Scoring of Word Definition Responses
Kevyn Collins-Thompson, Jamie Callan
HLT-NAACL1
2007 Combining Lexical and Grammatical Features to Improve Readability Measures for First and Second Language Texts
Michael Heilman, Kevyn Collins-Thompson, Jamie Callan, Maxine Eskénazi
HLT-NAACL2
2007 Estimation and use of uncertainty in pseudo-relevance feedback
abstract
Existing pseudo-relevance feedback methods typically perform averaging over the top-retrieved documents, but ignore an important statistical dimension: the risk or variance associated with either the individual document models, or their combination. Treating the baseline feedback method as a black box, and the output feedback model as a random variable, we estimate a posterior distribution for the feed-back model by resampling a given query's top-retrieved documents, using the posterior mean or mode as the enhanced feedback model. We then perform model combination over several enhanced models, each based on a slightly modified query sampled from the original query. We find that resampling documents helps increase individual feedback model precision by removing noise terms, while sampling from the query improves robustness (worst-case performance) by emphasizing terms related to multiple query aspects. The result is a meta-feedback algorithm that is both more robust and more precise than the original strong baseline method.
Kevyn Collins-Thompson, Jamie Callan
SIGIR1
2006 Classroom success of an intelligent tutoring system for lexical practice and reading comprehension
abstract
We present an intelligent tutoring system called REAP that provides reader-specific lexical practice for improved reading comprehension. REAP offers individualized practice to students by presenting authentic and appropriate reading materials selected automatically from the web. We encountered a number of challenges that must be met in order for the system to be effective in a classroom setting. These include general challenges for a system that uses authentic materials, as well as more specific challenges that arise from integrating the system with pre-existing classroom curricula. We discuss how these challenges were met, and present evidence that REAP has gained acceptance into the classroom at the English Language Institute at the University of Pittsburgh. 1. System Description
Michael Heilman, Kevyn Collins-Thompson, Jamie Callan, Maxine Eskénazi
INTERSPEECH2
2005 Query expansion using random walk models
abstract
It has long been recognized that capturing term relationships is an important aspect of information retrieval. Even with large amounts of data, we usually only have significant evidence for a fraction of all potential term pairs. It is therefore important to consider whether multiple sources of evidence may be combined to predict term relations more accurately. This is particularly important when trying to predict the probability of relevance of a set of terms given a query, which may involve both lexical and semantic relations between the terms.We describe a Markov chain framework that combines multiple sources of knowledge on term associations. The stationary distribution of the model is used to obtain probability estimates that a potential expansion term reflects aspects of the original query. We use this model for query expansion and evaluate the effectiveness of the model by examining the accuracy and robustness of the expansion methods, and investigate the relative effectiveness of various sources of term evidence. Statistically significant differences in accuracy were observed depending on the weighting of evidence in the random walk. For example, using co-occurrence data later in the walk was generally better than using it early, suggesting further improvements in effectiveness may be possible by learning walk behaviors.
Kevyn Collins-Thompson, Jamie Callan
CIKM1
2005 Predicting reading difficulty with statistical language models
abstract
Abstract A potentially useful feature of information retrieval systems for students is the ability to identify documents that not only are relevant to the query but also match the student's reading level. Manually obtaining an estimate of reading difficulty for each document is not feasible for very large collections, so we require an automated technique. Traditional readability measures, such as the widely used Flesch‐Kincaid measure, are simple to apply but perform poorly on Web pages and other nontraditional documents. This work focuses on building a broadly applicable statistical model of text for different reading levels that works for a wide range of documents. To do this, we recast the well‐studied problem of readability in terms of text categorization and use straightforward techniques from statistical language modeling. We show that with a modified form of text categorization, it is possible to build generally applicable classifiers with relatively little training data. We apply this method to the problem of classifying Web pages according to their reading difficulty level and show that by using a mixture model to interpolate evidence of a word's frequency across grades, it is possible to build a classifier that achieves an average root mean squared error of between one and two grade levels for 9 of 12 grades. Such classifiers have very efficient implementations and can be applied in many different scenarios. The models can be varied to focus on smaller or larger grade ranges or easily retrained for a variety of tasks or populations.
Kevyn Collins-Thompson, Jamie Callan
J. Assoc. Inf. Sci. Technol.1
2004 A Language Modeling Approach to Predicting Reading Difficulty
Kevyn Collins-Thompson, Jamie Callan
HLT-NAACL1
2004 Information retrieval for language tutoring: an overview of the REAP project
abstract
No abstract available.
Kevyn Collins-Thompson, Jamie Callan
SIGIR1
2004 The effect of document retrieval quality on factoid question answering performance
abstract
INTRODUCTION A widely-used architecture for factoid question answering (QA) involves the use of a multi-step pipeline consisting of: 1) initial question analysis, 2) document and/or passage retrieval, and 3) answer extraction. In this study, we examine the relationship between the quality of document retrieval and the overall accuracy of QA systems. We evaluate two QA systems using TREC 2002 test set questions [9]: Carnegie Mellon's JAVELIN system [7] and Waterloo's MultiText QA system [2]. We adapt the two QA systems in order to use di#erent sets of documents as input, and seven different document retrieval methods to create the list of documents including a combination of di#erent systems. The set of known relevant documents was used as a baseline to compare the di#erent retrieval methods. Documents with exact or inexact judgments are considered relevant. Our main hypothesis for this study is that there is a positive relationship between improved document retrieval and QA accuracy
Kevyn Collins-Thompson, Jamie Callan, Egidio L. Terra, Charles L. A. Clarke
SIGIR1
2001 Improved String Matching Under Noisy Channel Conditions
abstract
Many document-based applications, including popular Web browsers, email viewers, and word processors, have a 'Find on this Page' feature that allows a user to find every occurrence of a given string in the document. If the document text being searched is derived from a noisy process such as optical character recognition (OCR), the effectiveness of typical string matching can be greatly reduced. This paper describes an enhanced string-matching algorithm for degraded text that improves recall, while keeping precision at acceptable levels. The algorithm is more general than most approximate matching algorithms and allows string-to-string edits with arbitrary costs. We develop a method for evaluating our technique and use it to examine the relative effectiveness of each sub-component of the algorithm. Of the components we varied, we find that using confidence information from the recognition process lead to the largest improvements in matching accuracy.
Kevyn Collins-Thompson, Charles Schweizer, Susan T. Dumais
CIKM1