Peter Bailey

dblp:52/4837 · DBLP profile ↗
← Back
26ranked-venue papers
10as first author
2since 2021 · last 2022
0009-0005-9456-1865ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 22 · 9 first-authorArtificial intelligence and machine learning · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
15 papers
Information retrieval · 88% Recommender systems · 4% Data integration and cleaning · 4%
Human-computer interaction and pervasive computing
5 papers
Human-AI interaction · 34% Collaborative and social computing · 23% User interface design and tools · 14%
Artificial intelligence
1 paper
Language models and text generation · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
evaluation
1.262017
Incorporating User Expectations and Behavior into the Measurement of Search Effectiveness · ACM Trans. Inf. Syst. 2017
Retrieval Consistency in the Presence of Query Variations · SIGIR 2017
UQV100: A Test Collection with Query Variability · SIGIR 2016
Collaborative and social computing › computer-supported cooperative work
task management
0.612022
Imagining future digital assistants at work: A study of task management needs · Int. J. Hum. Comput. Stud. 2022
Information retrieval › evaluation
test collection
0.632016
UQV100: A Test Collection with Query Variability · SIGIR 2016
User Variability and IR System Evaluation · SIGIR 2015
Relevance assessment: are judges exchangeable and does it matter · SIGIR 2008
Information retrieval › query understanding
query variation
0.522017
Retrieval Consistency in the Presence of Query Variations · SIGIR 2017
UQV100: A Test Collection with Query Variability · SIGIR 2016
Information retrieval › evaluation
relevance judgment
0.432016
UQV100: A Test Collection with Query Variability · SIGIR 2016
Evaluating whole-page relevance · SIGIR 2010
Relevance assessment: are judges exchangeable and does it matter · SIGIR 2008
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.412020
Storytelling with Dialogue: A Critical Role Dungeons and Dragons Dataset · ACL 2020
Natural language and speech › Language models and text generation
text summarization
0.412020
Storytelling with Dialogue: A Critical Role Dungeons and Dragons Dataset · ACL 2020
Computational social science and digital humanities › computational linguistics
dialogue analysis
0.412020
Storytelling with Dialogue: A Critical Role Dungeons and Dragons Dataset · ACL 2020
Information retrieval
interactive information retrieval
0.422016
Ingrams: A Neuropsychological Explanation For Why People Search · SIGIR 2016
Explicit feedback in local search tasks · SIGIR 2013
Information retrieval › document retrieval › domain-specific retrieval
email search
0.412019
Evaluating User Actions as a Proxy for Email Significance · WWW 2019
Ubiquitous computing and smart environments
personal information management
0.412019
Evaluating User Actions as a Proxy for Email Significance · WWW 2019
Information retrieval
query formulation
0.322015
User Variability and IR System Evaluation · SIGIR 2015
Understanding the relationship of information need specificity to search query length · SIGIR 2007
Data integration and cleaning
data fusion
0.312017
Retrieval Consistency in the Presence of Query Variations · SIGIR 2017
Information retrieval › evaluation
effectiveness metrics
0.312017
Incorporating User Expectations and Behavior into the Measurement of Search Effectiveness · ACM Trans. Inf. Syst. 2017
Information retrieval
ranking
0.312017
Retrieval Consistency in the Presence of Query Variations · SIGIR 2017
Information retrieval › evaluation
user-oriented evaluation
0.312017
Incorporating User Expectations and Behavior into the Measurement of Search Effectiveness · ACM Trans. Inf. Syst. 2017
Information retrieval › evaluation › relevance judgment
crowdsourced relevance judgment
0.212016
UQV100: A Test Collection with Query Variability · SIGIR 2016
Information retrieval › interactive information retrieval
information needs
0.212016
Ingrams: A Neuropsychological Explanation For Why People Search · SIGIR 2016
Information retrieval › user behavior
search behavior modeling
0.212016
Ingrams: A Neuropsychological Explanation For Why People Search · SIGIR 2016
Information retrieval › web search
local search
0.212013
Explicit feedback in local search tasks · SIGIR 2013
Information retrieval
relevance feedback
0.212013
Explicit feedback in local search tasks · SIGIR 2013
Collaborative and social computing › computer-mediated communication › messaging
social messaging
0.112021
"I Can't Reply with That": Characterizing Problematic Email Reply Suggestions · CHI 2021
Information retrieval › web search
search personalization
0.112012
Modeling the impact of short- and long-term behavior on search personalization · SIGIR 2012
Web and social media mining › user behavior analysis
user behavior modeling
0.112012
Modeling the impact of short- and long-term behavior on search personalization · SIGIR 2012
Games and playful interaction › game genre
role-playing games
0.112020
Storytelling with Dialogue: A Critical Role Dungeons and Dragons Dataset · ACL 2020
Recommender systems
user interest modeling
0.112009
Predicting user interests from contextual information · SIGIR 2009
Recommender systems
user modeling
0.112009
Predicting user interests from contextual information · SIGIR 2009
Information retrieval › evaluation › test collection
test collection construction
0.112017
Incorporating User Expectations and Behavior into the Measurement of Search Effectiveness · ACM Trans. Inf. Syst. 2017
Information retrieval › query understanding
query specificity
0.112007
Understanding the relationship of information need specificity to search query length · SIGIR 2007
Data mining › text mining › text classification
genre classification
0.112006
Towards practical genre classification of web documents · WWW 2006

Methods — techniques the papers use, named apart from their topics

neural summarization · 1.3data augmentation · 1.3in-lab qualitative study · 0.9user action analysis · 0.8interview study · 0.6qualitative interviews · 0.5mixed methods · 0.5crowdsourced experiment · 0.5crowdsourced study · 0.4crowd-sourced study · 0.4relevance judgment · 0.3residual analysis · 0.3rank-biased overlap · 0.3data fusion · 0.3crowd-sourced assessment · 0.3user modeling · 0.2neuropsychological theory · 0.2
YearPublicationVenuePosition
2022 Imagining future digital assistants at work: A study of task management needs
Yonchanok Khaokaew, Indigo Holcombe-James, Mohammad Saiedur Rahaman, Jonathan Liono, Johanne R. Trippas, Damiano Spina, Peter Bailey, Nicholas J. Belkin, Paul N. Bennett, Yongli Ren, Mark Sanderson, Falk Scholer, Ryen W. White, Flora D. Salim
Int. J. Hum. Comput. Stud.7
2021 "I Can't Reply with That": Characterizing Problematic Email Reply Suggestions
abstract
In email interfaces, providing users with reply suggestions may simplify or accelerate correspondence. While the “success” of such systems is typically quantified using the number of suggestions selected by users, this ignores the impact of social context, which can change how suggestions are perceived. To address this, we developed a mixed-methods framework involving qualitative interviews and crowdsourced experiments to characterize problematic email reply suggestions. Our interviews revealed issues with over-positive, dissonant, cultural, and gender-assuming replies, as well as contextual politeness. In our experiments, crowdworkers assessed email scenarios that we generated and systematically controlled, showing that contextual factors like social ties and the presence of salutations impacts users’ perceptions of email correspondence. These assessments created a novel dataset of human-authored corrections for problematic email replies. Our study highlights the social complexity of providing suggestions for email correspondence, raising issues that may apply to all social messaging systems.
Ronald E. Robertson, Alexandra Olteanu, Fernando Diaz 0001, Milad Shokouhi, Peter Bailey
CHI5
2020 Storytelling with Dialogue: A Critical Role Dungeons and Dragons Dataset
abstract
This paper describes the Critical Role Dungeons and Dragons Dataset (CRD3) and related analyses.Critical Role is an unscripted, live-streamed show where a fixed group of people play Dungeons and Dragons, an openended role-playing game.The dataset is collected from 159 Critical Role episodes transcribed to text dialogues, consisting of 398,682 turns.It also includes corresponding abstractive summaries collected from the Fandom wiki.The dataset is linguistically unique in that the narratives are generated entirely through player collaboration and spoken interaction.For each dialogue, there are a large number of turns, multiple abstractive summaries with varying levels of detail, and semantic ties to the previous dialogues.In addition, we provide a data augmentation method that produces 34,243 summarydialogue chunk pairs to support current neural ML approaches, and we provide an abstractive summarization benchmark and evaluation.
Revanth Rameshkumar, Peter Bailey
ACL2
2020 The Impact of More Transparent Interfaces on Behavior in Personalized Recommendation
abstract
Many interactive online systems, such as social media platforms or news sites, provide personalized experiences through recommendations or news feed customization based on people's feedback and engagement on individual items (e.g., liking items). In this paper, we investigate how we can support a greater degree of user control in such systems by changing the way the system allows people to gauge the consequences of their feedback actions. To this end, we consider two important aspects of how the system responds to feedback actions: (i) immediacy, i.e., how quickly the system responds with an update, and (ii) visibility, i.e., whether or not changes will get highlighted. We used both an in-lab qualitative study and a large-scale crowd-sourced study to examine the impact of these factors on people's reported preferences and observed behavioral metrics. We demonstrate that UX design which enables people to preview the impact of their actions and highlights changes results in a higher reported transparency, an overall preference for this design, and a greater selectivity in which items are liked.
Tobias Schnabel, Saleema Amershi, Paul N. Bennett, Peter Bailey, Thorsten Joachims
SIGIR4
2019 Learning About Work Tasks to Inform Intelligent Assistant Design
abstract
Intelligent assistants can serve many purposes, including entertainment (e.g. playing music), home automation, and task management (e.g. timers, reminders). The role of these assistants is evolving to also support people engaged in work tasks, in workplaces and beyond. To design truly useful intelligent assistants for work, it is important to better understand the work tasks that people are performing. Based on a survey of 401 respondents' daily tasks and activities in a work setting, we present a classification of work-related tasks, and analyze their key characteristics, including the frequency of their self-reported tasks, the environment in which they undertake the tasks, and which, if any, electronic devices are used. We also investigate the cyber, physical, and social aspects of tasks. Finally, we reflect on how intelligent assistants could influence and help people in a work environment to complete their tasks, and synthesize our findings to provide insight on the future of intelligent assistants in support of amplifying personal productivity.
Johanne R. Trippas, Damiano Spina, Falk Scholer, Ahmed Awadallah 0001, Peter Bailey, Paul N. Bennett, Ryen W. White, Jonathan Liono, Yongli Ren, Flora D. Salim, Mark Sanderson
CHIIR5
2019 Evaluating User Actions as a Proxy for Email Significance
abstract
Email remains a critical channel for communicating information in both personal and work accounts. The number of emails people receive every day can be overwhelming, which in turn creates challenges for efficient information management and consumption. Having a good estimate of the significance of emails forms the foundation for many downstream tasks (e.g. email prioritization); but determining significance at scale is expensive and challenging.
Tarfah Alrashed, Peter Bailey, Christopher H. Lin, Milad Shokouhi, Susan T. Dumais
WWW3
2017 Retrieval Consistency in the Presence of Query Variations
abstract
A search engine that can return the ideal results for a person's information need, independent of the specific query that is used to express that need, would be preferable to one that is overly swayed by the individual terms used; search engines should be consistent in the presence of syntactic query variations responding to the same information need. In this paper we examine the retrieval consistency of a set of five systems responding to syntactic query variations over one hundred topics, working with the UQV100 test collection, and using Rank-Biased Overlap (RBO) relative to a centroid ranking over the query variations per topic as a measure of consistency. We also introduce a new data fusion algorithm, Rank-Biased Centroid (RBC), for constructing a centroid ranking over a set of rankings from query variations for a topic. RBC is compared with alternative data fusion algorithms.
Peter Bailey, Alistair Moffat, Falk Scholer, Paul Thomas 0001
SIGIR1
2017 Incorporating User Expectations and Behavior into the Measurement of Search Effectiveness
abstract
Information retrieval systems aim to help users satisfy information needs. We argue that the goal of the person using the system, and the pattern of behavior that they exhibit as they proceed to attain that goal, should be incorporated into the methods and techniques used to evaluate the effectiveness of IR systems, so that the resulting effectiveness scores have a useful interpretation that corresponds to the users’ search experience. In particular, we investigate the role of search task complexity, and show that it has a direct bearing on the number of relevant answer documents sought by users in response to an information need, suggesting that useful effectiveness metrics must be goal sensitive . We further suggest that user behavior while scanning results listings is affected by the rate at which their goal is being realized, and hence that appropriate effectiveness metrics must be adaptive to the presence (or not) of relevant documents in the ranking. In response to these two observations, we present a new effectiveness metric, INST, that has both of the desired properties: INST employs a parameter T , a direct measure of the user’s search goal that adjusts the top-weightedness of the evaluation score; moreover, as progress towards the target T is made, the modeled user behavior is adapted, to reflect the remaining expectations. INST is experimentally compared to previous effectiveness metrics, including Average Precision (AP), Normalized Discounted Cumulative Gain (NDCG), and Rank-Biased Precision (RBP), demonstrating our claims as to INST’s usefulness. Like RBP, INST is a weighted-precision metric, meaning that each score can be accompanied by a residual that quantifies the extent of the score uncertainty caused by unjudged documents. As part of our experimentation, we use crowd-sourced data and score residuals to demonstrate that a wide range of queries arise for even quite specific information needs, and that these variant queries introduce significant levels of residual uncertainty into typical experimental evaluations. These causes of variability have wide-reaching implications for experiment design, and for the construction of test collections.
Alistair Moffat, Peter Bailey, Falk Scholer, Paul Thomas 0001
ACM Trans. Inf. Syst.2
2016 Ingrams: A Neuropsychological Explanation For Why People Search
abstract
Why do people start a search? Why do they stop? Why do they do what they do in-between? Our goal in this paper is to provide a simple yet general explanation for these acts that has its basis in neuropsychology and observed user behavior. We coin the term "ingram", as an information counterpart to Richard Semon's ?engram? or "memory trace". People search to create ingrams. People stop searching because they have created sufficient ingrams, or given up. We describe these acts through a pair of user models and use it to explain various user behaviors in search activity. Understanding people?s search acts in terms of ingrams may help us predict or model the interaction of people?s information needs, the queries they issue, and the information they consume. If we could observe certain decision-making acts within these activities, we might also gain new insight into the relationships between textual information and knowledge representation.
Peter Bailey, Nick Craswell
SIGIR1
2016 UQV100: A Test Collection with Query Variability
abstract
We describe the UQV100 test collection, designed to incorporate variability from users. Information need ?backstories? were written for 100 topics (or sub-topics) from the TREC 2013 and 2014 Web Tracks. Crowd workers were asked to read the backstories, and provide the queries they would use; plus effort estimates of how many useful documents they would have to read to satisfy the need. A total of 10,835 queries were collected from 263 workers. After normalization and spell-correction, 5,764 unique variations remained; these were then used to construct a document pool via Indri-BM25 over the ClueWeb12-B corpus. Qualified crowd workers made relevance judgments relative to the backstories, using a relevance scale similar to the original TREC approach; first to a pool depth of ten per query, then deeper on a set of targeted documents. The backstories, query variations, normalized and spell-corrected queries, effort estimates, run outputs, and relevance judgments are made available collectively as the UQV100 test collection. We also make available the judging guidelines and the gold hits we used for crowd-worker qualification and spam detection. We believe this test collection will unlock new opportunities for novel investigations and analysis, including for problems such as task-intent retrieval performance and consistency (independent of query variation), query clustering, query difficulty prediction, and relevance feedback, among others.
Peter Bailey, Alistair Moffat, Falk Scholer, Paul Thomas 0001
SIGIR1
2015 Pooled Evaluation Over Query Variations: Users are as Diverse as Systems
abstract
Evaluation of information retrieval systems with test collections makes use of a suite of fixed resources: a document corpus; a set of topics; and associated judgments of the relevance of each document to each topic. With large modern collections, exhaustive judging is not feasible. Therefore an approach called pooling is typically used where, for example, the documents to be judged can be determined by taking the union of all documents returned in the top positions of the answer lists returned by a range of systems. Conventionally, pooling uses system variations to provide diverse documents to be judged for a topic; different user queries are not considered. We explore the ramifications of user query variability on pooling, and demonstrate that conventional test collections do not cover this source of variation. The effect of user query variation on the size of the judging pool is just as strong as the effect of retrieval system variation. We conclude that user query variation should be incorporated early in test collection construction, and cannot be considered effectively post hoc.
Alistair Moffat, Falk Scholer, Paul Thomas 0001, Peter Bailey
CIKM4
2015 User Variability and IR System Evaluation
abstract
Test collection design eliminates sources of user variability to make statistical comparisons among information retrieval (IR) systems more affordable. Does this choice unnecessarily limit generalizability of the outcomes to real usage scenarios? We explore two aspects of user variability with regard to evaluating the relative performance of IR systems, assessing effectiveness in the context of a subset of topics from three TREC collections, with the embodied information needs categorized against three levels of increasing task complexity. First, we explore the impact of widely differing queries that searchers construct for the same information need description. By executing those queries, we demonstrate that query formulation is critical to query effectiveness. The results also show that the range of scores characterizing effectiveness for a single system arising from these queries is comparable or greater than the range of scores arising from variation among systems using only a single query per topic. Second, our experiments reveal that searchers display substantial individual variation in the numbers of documents and queries they anticipate needing to issue, and there are underlying significant differences in these numbers in line with increasing task complexity levels. Our conclusion is that test collection design would be improved by the use of multiple query variations per topic, and could be further improved by the use of metrics which are sensitive to the expected numbers of useful documents.
Peter Bailey, Alistair Moffat, Falk Scholer, Paul Thomas 0001
SIGIR1
2014 Relevance and Effort: An Analysis of Document Utility
abstract
In this paper, we study one important source of the mis-match between user data and relevance judgments, those due to the high degree of effort required by users to identify and consume the information in a document. Information retrieval relevance judges are trained to search for evidence of relevance when assessing documents. For complex documents, this can lead to judges' spending substantial time considering each document. However, in practice, search users are often much more impatient: if they do not see evidence of relevance quickly, they tend to give up.
Emine Yilmaz, Manisha Verma, Nick Craswell, Filip Radlinski, Peter Bailey
CIKM5
2013 Explicit feedback in local search tasks
abstract
Modern search engines make extensive use of people's contextual information to finesse result rankings. Using a searcher's location provides an especially strong signal for adjusting results for certain classes of queries where people may have clear preference for local results, without explicitly specifying the location in the query direct-ly. However, if the location estimate is inaccurate or searchers want to obtain many results from a particular location, they have limited control on the location focus in the search results returned. In this paper we describe a user study that examines the effect of offering searchers more control over how local preferences are gathered and used. We studied providing users with functionality to offer explicit relevance feedback (ERF) adjacent to results automatically identi-fied as location-dependent (i.e., more from this location). They can use this functionality to indicate whether they are interested in a particular search result and desire more results from that result's location. We compared the ERF system against a baseline (NoERF) that used the same underlying mechanisms to retrieve and rank results, but did not offer ERF support. User performance was as-sessed across 12 experimental participants over 12 location-sensitive topics, in a fully counter-balanced design. We found that participants interacted with ERF frequently, and there were signs that ERF has the potential to improve success rates and lead to more efficient searching for location-sensitive search tasks than NoERF.
Dmitry Lagun, Avneesh Sud, Ryen W. White, Peter Bailey, Georg Buscher
SIGIR4
2012 Modeling the impact of short- and long-term behavior on search personalization
abstract
User behavior provides many cues to improve the relevance of search results through personalization. One aspect of user behavior that provides especially strong signals for delivering better relevance is an individual's history of queries and clicked documents. Previous studies have explored how short-term behavior or long-term behavior can be predictive of relevance. Ours is the first study to assess how short-term (session) behavior and long-term (historic) behavior interact, and how each may be used in isolation or in combination to optimally contribute to gains in relevance through search personalization. Our key findings include: historic behavior provides substantial benefits at the start of a search session; short-term session behavior contributes the majority of gains in an extended search session; and the combination of session and historic behavior out-performs using either alone. We also characterize how the relative contribution of each model changes throughout the duration of a session. Our findings have implications for the design of search systems that leverage user behavior to personalize the search experience.
Paul N. Bennett, Ryen W. White, Susan T. Dumais, Peter Bailey, Fedor Borisyuk, Xiaoyuan Cui
SIGIR5
2011 Recommending interesting activity-related local entities
abstract
When searching for entities with a strong local character (e.g., a museum), people may also be interested in discovering proximal activity-related entities (e.g., a café). Geographical proximity is a necessary, but not sufficient, qualifier for recommending other entities such that they are related in a useful manner (e.g., interest in a fish market does not imply interest in nearby bookshops, but interest in other produce stores is more likely). We describe and evaluate methods to identify such activity-related local entities.
Ryen W. White, Peter Bailey
SIGIR3
2010 Evaluating whole-page relevance
abstract
Whole page relevance defines how well the surface-level representation of all elements on a search result page and the corresponding holistic attributes of the presentation respond to users’ information needs. We introduce a method for evaluating the whole-page relevance of Web search engine results pages. Our key contribution is that the method allows us to investigate aspects of component relevance that are difficult or impossible to judge in isolation. Such aspects include component-level information redundancy and cross-component coherence. The method we describe complements traditional document relevance measurement, affords comparative relevance assessment across multiple search engines, and facilitates the study of important factors such as brand presentation effects and component-level quality.
Peter Bailey, Nick Craswell, Ryen W. White, Ashwin Satyanarayana, Seyed M. M. Tahaghoghi
SIGIR1
2010 Mining Historic Query Trails to Label Long and Rare Search Engine Queries
abstract
Web search engines can perform poorly for long queries (i.e., those containing four or more terms), in part because of their high level of query specificity. The automatic assignment of labels to long queries can capture aspects of a user’s search intent that may not be apparent from the terms in the query. This affords search result matching or reranking based on queries and labels rather than the query text alone. Query labels can be derived from interaction logs generated from many users’ search result clicks or from query trails comprising the chain of URLs visited following query submission. However, since long queries are typically rare, they are difficult to label in this way because little or no historic log data exists for them. A subset of these queries may be amenable to labeling by detecting similarities between parts of a long and rare query and the queries which appear in logs. In this article, we present the comparison of four similarity algorithms for the automatic assignment of Open Directory Project category labels to long and rare queries, based solely on matching against similar satisfied query trails extracted from log data. Our findings show that although the similarity-matching algorithms we investigated have tradeoffs in terms of coverage and accuracy, one algorithm that bases similarity on a popular search result ranking function (effectively regarding potentially-similar queries as “documents”) outperforms the others. We find that it is possible to correctly predict the top label better than one in five times, even when no past query trail exactly matches the long and rare query. We show that these labels can be used to reorder top-ranked search results leading to a significant improvement in retrieval performance over baselines that do not utilize query labeling, but instead rank results using content-matching or click-through logs. The outcomes of our research have implications for search providers attempting to provide users with highly-relevant search results for long queries.
Peter Bailey, Ryen W. White, Han Liu 0001, Giridhar Kumaran
ACM Trans. Web1
2009 Accelerating Lattice Boltzmann Fluid Flow Simulations Using Graphics Processors
abstract
Lattice Boltzmann methods (LBM) are used for the computational simulation of Newtonian fluid dynamics. LBM-based simulations are readily parallelizable; they have been implemented on general-purpose processors, field-programmable gate arrays (FPGAs), and graphics processing units (GPUs). Of the three methods, the GPU implementations achieved the highest simulation performance per chip. With memory bandwidth of up to 141 GB/s and a theoretical maximum floating point performance of over 600 GFLOPS, CUDA-ready GPUs from NVIDIA provide an attractive platform for a wide range of scientific simulations, including LBM. This paper improves upon prior single-precision GPU LBM results for the D3Q19 model by increasing GPU multiprocessor occupancy, resulting in an increase in maximum performance by 20%, and by introducing a space-efficient storage method which reduces GPU RAM requirements by 50% at a slight detriment to performance. Both GPU implementations are over 28 times faster than a single-precision quad-core CPU version utilizing OpenMP.
Peter Bailey, Joe Myre, Stuart D. C. Walsh, David J. Lilja, Martin O. Saar
ICPP1
2009 Predicting user interests from contextual information
abstract
Search and recommendation systems must include contextual information to effectively model users' interests. In this paper, we present a systematic study of the effectiveness of five variant sources of contextual information for user interest modeling. Post-query navigation and general browsing behaviors far outweigh direct search engine interaction as an information-gathering activity. Therefore we conducted this study with a focus on Website recommendations rather than search results. The five contextual information sources used are: social, historic, task, collection, and user interaction. We evaluate the utility of these sources, and overlaps between them, based on how effectively they predict users' future interests. Our findings demonstrate that the sources perform differently depending on the duration of the time window used for future prediction, and that context overlap outperforms any isolated source. Designers of Website suggestion systems can use our findings to provide improved support for post-query navigation and general browsing behaviors.
Ryen W. White, Peter Bailey
SIGIR2
2008 Relevance assessment: are judges exchangeable and does it matter
abstract
We investigate to what extent people making relevance judgements for a reusable IR test collection are exchangeable. We consider three classes of judge: "gold standard" judges, who are topic originators and are experts in a particular information seeking task; "silver standard" judges, who are task experts but did not create topics; and "bronze standard" judges, who are those who did not define topics and are not experts in the task.
Peter Bailey, Nick Craswell, Ian Soboroff, Paul Thomas 0001, Arjen P. de Vries, Emine Yilmaz
SIGIR1
2007 Understanding the relationship of information need specificity to search query length
abstract
When searching, people's information needs flowthrough to expressing an information retrieval request posed to asearch engine. We hypothesise that the degree of specificity of anIR request might correspond to the length of a search query. Ourresults show a strong correlation between decreasing query lengthand increasing broadness or generality of the IR request. We foundan average cross-over point of specificity from broad to narrow of 3words in the query. These results have implications for searchengines in responding to queries of differing lengths.
Nina Phan, Peter Bailey, Ross Wilkinson
SIGIR2
2006 Secure search in enterprise webs: tradeoffs in efficient implementation for document level security
abstract
Document level security (DLS) -- enforcing permissions prevailing at the time of search -- is specified as a mandatory requirement in many enterprise search applications. Unfortunately, depending upon implementation details and values of key parameters, DLS may come at a high price in increased query processing time, leading to an unacceptably slow search experience. In this paper we present a model and a method for carrying out secure search in the presence of DLS within enterprise webs. We report on two alternative commercial DLS search implementations. Using a 10,000 document experimental DLS environment, we graph the dependence of query processing time on result set size and visibility density for different classes of user. Scaled up to collections of tens of thousands of documents, our results suggest that query times will be unacceptable if exact counts of matching documents are required and also for users who can view only a small proportion of documents. We show that the time to conduct access checks is dramatically increased if requests must be sent off-server, even on a local network, and discuss methods for reducing the cost of security checks. We conclude that enterprises can effectively reduce DLS overheads by organizing documents in such a way that most access checking can be at collection rather than document level, by forgoing accurate match counts, by using caching, batching or hierarchical methods to cut costs of DLS checking and, if applicable, by using a single portal both to access and search documents.
Peter Bailey, David Hawking, Brett Matson
CIKM1
2006 Towards practical genre classification of web documents
abstract
Classification of documents by genre is typically done either using linguistic analysis or term frequency based techniques. The former provides better classification accuracy than the latter but at the cost of two orders of magnitude more computation time. While term frequency analysis requires much less computational resources than linguistic analysis,it returns poor classification accuracy when the genres are not sufficiently distinct. A method that removes or approximates the expensive portions of linguistic analysis is presented.The accuracy and computation time of this method then compared with both linguistic analysis and term frequency analysis. The results in this paper show that this method can significantly reduce the computation of both time of linguistic analysis and term frequency analysis, while retaining an accuracy that is higher than that of term frequency analysis.
George Ferizis, Peter Bailey
WWW2
2003 Engineering a multi-purpose test collection for Web retrieval experiments
Peter Bailey, Nick Craswell, David Hawking
Inf. Process. Manag.1
2001 Measuring Search Engine Quality
David Hawking, Nick Craswell, Peter Bailey, Kathleen Griffiths
Inf. Retr.3