EDBT 2026 Demo / reviewers in the wild / expert
Jaime Teevan
dblp:38/129
· DBLP profile ↗
47ranked-venue papers in the field
16as first author
6since 2021 · last 2023
0000-0002-2786-0209ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 32 (12 first)Data Mining & Knowledge Discovery · 15 (4 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Targeted Training for Multi-organization RecommendationabstractMaking recommendations for users in diverse organizations (orgs) is a challenging task for workplace social platforms such as Microsoft Teams and Slack. The current industry-standard model training approaches either use data from all organizations to maximize information or train organization-specific models to minimize noise. Our real-world experiments show that both approaches are poorly suited for the multi-org recommendation setting where different organizations’ interaction patterns vary in their generalizability. We introducetargeted training, which improves on standard practices by automatically selecting a subset of orgs for model development whose data are cleanest and best represent global trends. We demonstrate how and when targeted training improves over global training through theoretical analysis and simulation. Our experiments on large-scale datasets from Microsoft Teams, SharePoint, Stack Exchange, DBLP, and Reddit show that in many cases targeted training can improve mean average precision (MAP) across orgs by 10–15% over global training, is more robust to orgs with lower data quality, and generalizes better to unseen orgs. Our training framework is applicable to a wide range of inductive recommendation models, from simple regression models to graph neural networks (GNNs). Kiran Tomlinson, Mengting Wan, Cao Lu, Brent J. Hecht, Jaime Teevan, Longqi Yang 0001 |
Trans. Recomm. Syst. | 5 |
| 2022 | How Hybrid Work Will Make Work More IntelligentabstractWe are in the middle of the most significant change to work practices in generations. For hundreds of years, physical space was the most important technology people used to get things done. The coming Hybrid Work Era, however, will be shaped by digital technology. The recent rapid shift to remote work accelerated the digital transformation already underway at many organizations, and new types of work-related data are now being generated at an unprecedented rate. For example, the average Microsoft Teams user spends 252% more time in the application now than they did in February 2020. Jaime Teevan |
CIKM | 1 |
| 2022 | Learning Causal Effects on HypergraphsabstractHypergraphs provide an effective abstraction for modeling multi-way group interactions among nodes, where each hyperedge can connect any number of nodes. Different from most existing studies which leverage statistical dependencies, we study hypergraphs from the perspective of causality. Specifically, in this paper, we focus on the problem of individual treatment effect (ITE) estimation on hypergraphs, aiming to estimate how much an intervention (e.g., wearing face covering) would causally affect an outcome (e.g., COVID-19 infection) of each individual node. Existing works on ITE estimation either assume that the outcome on one individual should not be influenced by the treatment assignments on other individuals (i.e., no interference), or assume the interference only exists between pairs of connected individuals in an ordinary graph. We argue that these assumptions can be unrealistic on real-world hypergraphs, where higher-order interference can affect the ultimate ITE estimations due to the presence of group interactions. In this work, we investigate high-order interference modeling, and propose a new causality learning framework powered by hypergraph neural networks. Extensive experiments on real-world hypergraphs verify the superiority of our framework over existing baselines. Jing Ma 0002, Mengting Wan, Longqi Yang 0001, Jundong Li, Brent J. Hecht, Jaime Teevan |
KDD | 6 |
| 2022 | Searching for a New and Better Future of WorkabstractSearch engines were one of the first intelligent cloud-based applications that people used to get things done, and they have since become an extremely important productivity tool. This is in part because much of what a person is doing when they search is thinking. Search engines do not merely support that thinking, however, but can also actually shape it. For example, some search results are more likely to spur learning than others and influence a person's future queries [1]. This means that as information retrieval researchers the new approaches that we develop can actively shape the future of work. Jaime Teevan |
SIGIR | 1 |
| 2022 | How the Web Will Shape the Hybrid Work Era: Keynote TalkabstractJaime Teevan has given a Keynote Talk at The Web Conference on Friday 29th April 2022. This page provides a summary of the topics she addressed during her talk. Jaime Teevan |
WWW | 1 |
| 2021 | Learning to Represent Human Motives for Goal-directed Web BrowsingabstractMotives or goals are recognized in psychology literature as the most fundamental drive that explains and predicts why people do what they do, including when they browse the web. Although providing enormous value, these higher-ordered goals are often unobserved, and little is known about how to leverage such goals to assist people’s browsing activities. This paper proposes to take a new approach to address this problem, which is fulfilled through a novel neural framework, Goal-directed Web Browsing (GoWeB). We adopt a psychologically-sound taxonomy of higher-ordered goals and learn to build their representations in a structure-preserving manner. Then we incorporate the resulting representations for enhancing the experiences of common activities people perform on the web. Experiments on large-scale data from Microsoft Edge web browser show that GoWeB significantly outperforms competitive baselines for in-session web page recommendation, re-visitation classification, and goal-based web page grouping. A follow-up analysis further characterizes how the variety of human motives can affect the difference observed in human behavioral patterns. Jyun-Yu Jiang, Longqi Yang 0001, Bahareh Sarrafzadeh, Brent J. Hecht, Jaime Teevan |
RecSys | 6 |
| 2019 | Attending to What MattersabstractOnline services are increasingly intelligent. They evolve intelligently through A/B testing and experimentation, employ artificial intelligence in their core functionality using machine learning, and seamlessly engage human intelligence by connecting people in a low-friction manner. All of this has resulted in incredibly engaging experiences -- but not particularly productive ones. As more and more of people's most important tasks move online we need to think carefully about the underlying influence online services have on people's ability to attend to what matters to them. There is an opportunity to use intelligence for this to do more than just not distract people and actually start helping people attend to what matters even better than they would otherwise. This presentation explores the ways we might make it as compelling and easy to start an important task as it is to check social media. Jaime Teevan |
WSDM | 1 |
| 2017 | CrowdMask: Using Crowds to Preserve Privacy in Crowd-Powered Systems via Progressive FilteringabstractCrowd-powered systems leverage human intelligence to go beyond the capabilities of automated systems, but also introduce privacy and security concerns because unknown people must view the data that the system processes. While automated approaches cannot robustly filter private information from these datasets, people have the ability to do so if the risk from them viewing the data can be mitigated. We present a crowd-powered approach to masking private content in data by segmenting and distributing smaller segments to crowd workers so that individual workers can identify potentially private content without being able to fully view it themselves. We introduce a novel pyramid workflow for segmentation that uses segments at multiple levels of granularity to overcome problems with fixed-sized approaches. We implement our approach in CrowdMask, a system that allows images with potentially sensitive content to be masked by appearing in progressively larger, more identifiable segments, and masking portions of the image as soon as a risk is identified. Our experiments with 4134 Mechanical Turk workers show that CrowdMask can effectively mask private content from images without revealing sensitive content to constituent workers, while still enabling future systems to use the filtered result. Harmanpreet Kaur, Mitchell L. Gordon, Yiwei Yang 0004, Jeffrey P. Bigham, Jaime Teevan, Ece Kamar, Walter S. Lasecki |
HCOMP | 5 |
| 2017 | Design and in-situ evaluation of a mixed-initiative approach to information organizationabstractOrganizing personal information by folders or tags has proved to be effective for finding, remembering, and understanding information. However, past studies have shown that the cost of organization can be too high for some users to be worth the effort. Mixed‐initiative approaches attempt to reduce the burden of manual organization by automatically identifying and suggesting organizational units such as folders to users. However, little is known about how such mixed‐initiative approaches influence users' organizational experiences. In this paper, we explore a mixed‐initiative approach that suggests high‐level organizational units to users to facilitate e‐mail organization. In 2 in‐situ experiments with 34 knowledge workers, we study how our mixed‐initiative approach influenced users' experience with organization. We show that our approach made it easier to create organizational units without negatively affecting recall of those units, and led to the creation of units that otherwise would have not been created. Our findings suggest ways computers and people can most effectively work together to organize information. Mona Haraty, Zhongyuan Wang 0006, Helen J. Wang, Shamsi T. Iqbal, Jaime Teevan |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2016 | Explicit In Situ User Feedback for Web Search ResultsabstractGathering evidence about whether a search result is relevant is a core concern in the evaluation and improvement of information retrieval systems. Two common sources of evidence for establishing relevance are judgements from trained assessors and logs of online user behavior. However, both are limited; it is hard for a trained assessor to know exactly what users want to find, and user behavior only provides an implicit and ambiguous signal. In this paper, we aim to address these limitations by collecting explicit feedback on web search results from users in situ as they search. When users return to the search result page via the browser back button after having clicked on a result, we ask them to provide a binary thumbs up or thumbs down judgment and text feedback. We collect in situ feedback from a large commercial search engine, and compare this feedback with the judgments provided by trained assessors. We find that in situ feedback differs significantly from traditional relevance judgments, and that it suggests a different interpretation of behavior signals, with the dwell time threshold between negative and positive in situ feedback being 87 seconds, longer than the more common heuristic of 30 seconds. Using text feedback from users, we discuss why user feedback may differ from editorial judgments. Jin Young Kim 0005, Jaime Teevan, Nick Craswell |
SIGIR | 2 |
| 2016 | Using the Crowd to Improve Search Result Ranking and the Search ExperienceabstractDespite technological advances, algorithmic search systems still have difficulty with complex or subtle information needs. For example, scenarios requiring deep semantic interpretation are a challenge for computers. People, on the other hand, are well suited to solving such problems. As a result, there is an opportunity for humans and computers to collaborate during the course of a search in a way that takes advantage of the unique abilities of each. While search tools that rely on human intervention will never be able to respond as quickly as current search engines do, recent research suggests that there are scenarios where a search engine could take more time if it resulted in a much better experience. This article explores how crowdsourcing can be used at query time to augment key stages of the search pipeline. We first explore the use of crowdsourcing to improve search result ranking. When the crowd is used to replace or augment traditional retrieval components such as query expansion and relevance scoring, we find that we can increase robustness against failure for query expansion and improve overall precision for results filtering. However, the gains that we observe are limited and unlikely to make up for the extra cost and time that the crowd requires. We then explore ways to incorporate the crowd into the search process that more drastically alter the overall experience. We find that using crowd workers to support rich query understanding and result processing appears to be a more worthwhile way to make use of the crowd during search. Our results confirm that crowdsourcing can positively impact the search experience but suggest that significant changes to the search process may be required for crowdsourcing to fulfill its potential in search systems. Yubin Kim 0001, Kevyn Collins-Thompson, Jaime Teevan |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2015 | Slow Search: Improving Information Retrieval Using Human AssistanceabstractWe live in a world where the pace of everything from communication to transportation is getting faster. In recent years a number of "slow movements" have emerged that advocate for reducing speed in exchange for increasing quality, including the slow food movement, slow parenting, slow travel, and even slow science. We propose the concept of "slow search," where search engines use additional time to provide a higher quality search experience than is possible given conventional time constraints. While additional time can be used to identify relevant results within the existing search engine framework, it can also be used to create new search artifacts and enable previously unimaginable user experiences. This talk will focus on how search engines can make use of additional time to employ a resource that is inherently slow: other people. Using crowdsourcing and friendsourcing, it will highlight opportunities for search systems to support new search experiences with high quality result content that takes time to identify. Jaime Teevan |
CIKM | 1 |
| 2015 | Crowdsourcing in the Field: A Case Study Using Local Crowds for Event ReportingabstractWhile crowd work typically involves tasks that performed at any time and anywhere, some tasks inherently require the physical presence of workers at a specific time and location. This paper presents a case study of a hybrid crowdsourcing process that involves the collaborative production of event reports using a combination of local and remote workers. The process extends human computation into the physical world by using local workers to collect information in person at events and remote workers to curate the collected information and generate event reports. We deployed the process at 11 events, employing 84 workers, and identified the challenges local workers face as constraints in mobility, time available to perform tasks, unpredictability of events, and interaction with others. We discuss issues related to collaboration with remote workers and bias in field reporting, and conduct a qualitative analysis to make design recommendations for extending human computation into the physical environment. Elena Agapie, Jaime Teevan, Andrés Monroy-Hernández |
HCOMP | 2 |
| 2014 | A Crowd of Your Own: Crowdsourcing for On-Demand PersonalizationabstractPersonalization is a way for computers to support people’s diverse interests and needs by providing content tailored to the individual. While strides have been made in algorithmic approaches to personalization, most require access to a significant amount of data. However, even when data is limited online crowds can be used to infer an individual’s personal preferences. Aided by the diversity of tastes among online crowds and their ability to understand others, we show that crowdsourcing is an effective on-demand tool for personalization. Unlike typical crowdsourcing approaches that seek a ground truth, we present and evaluate two crowdsourcing approaches designed to capture personal preferences. The first, taste-matching, identifies workers with similar taste to the requester and uses their taste to infer the requester’s taste. The second, taste-grokking, asks workers to explicitly predict the requester’s taste based on training examples. These techniques are evaluated on two subjective tasks, personalized image recommendation and tailored textual summaries. Taste-matching and taste-grokking both show improvement over the use of generic workers, and have different benefits and drawbacks depending on the complexity of the task and the variability of the taste space. Peter Organisciak, Jaime Teevan, Susan T. Dumais, Rob Miller 0001, Adam Tauman Kalai |
HCOMP | 2 |
| 2014 | Choices and constraints: research goals and approaches in information retrieval (part 1)abstractAll research projects begin with a goal, for instance to describe search behavior, to predict when a person will enter a second query, or to discover which IR system performs the best. Different research goals suggest different research approaches, ranging from field studies to lab studies to online experimentation. This tutorial will provide an overview of the different types of research goals, common evaluation approaches used to address each type, and the constraints each approach entails. Participants will come away with a broad perspective of research goals and approaches in IR, and an understanding of the benefits and limitations of these research approaches. The tutorial will take place in two independent, but interrelated parts, each focusing on a unique set of research approaches but with the same intended tutorial outcomes. These outcomes will be accomplished by deconstructing and analyzing our own published research papers, with further illustrations of each technique using the broader literature. By using our own research as anchors, we will provide insight about the research process, revealing the difficult choices and trade-offs researchers make when designing and conducting IR studies. Diane Kelly 0001, Filip Radlinski, Jaime Teevan |
SIGIR | 3 |
| 2014 | Choices and constraints: research goals and approaches in information retrieval (part 2)abstractAll research projects begin with a goal, for instance to describe search behavior, to predict when a person will enter a second query, or to discover which IR system performs the best. Different research goals suggest different research approaches, ranging from field studies to lab studies to online experimentation. This tutorial will provide an overview of the different types of research goals, common evaluation approaches used to address each type, and the constraints each approach entails. Participants will come away with a broad perspective of research goals and approaches in IR, and an understanding of the benefits and limitations of these research approaches. The tutorial will take place in two independent, but interrelated parts, each focusing on a unique set of research approaches but with the same intended tutorial outcomes. These outcomes will be accomplished by deconstructing and analyzing our own published research papers, with further illustrations of each technique using the broader literature. By using our own research as anchors, we will provide insight about the research process, revealing the difficult choices and trade-offs researchers make when designing and conducting IR studies. Diane Kelly 0001, Filip Radlinski, Jaime Teevan |
SIGIR | 3 |
| 2014 | Characterizing multi-click search behavior and the risks and opportunities of changing results during useabstractAlthough searchers often click on more than one result following a query, little is known about how they interact with search results after their first click. Using large scale query log analysis, we characterize what people do when they return to a result page after having visited an initial result. We find that the initial click provides insight into the searcher's subsequent behavior, with short initial dwell times suggesting more future interaction and later clicks occurring close in rank to the first. Although users think of a search result list as static, when people return to a result list following a click there is the opportunity for the list to change, potentially providing additional relevant content. Such change, however, can be confusing, leading to increased abandonment and slower subsequent clicks. We explore the risks and opportunities of changing search results during use, observing, for example, that when results change above a user's initial click that user is less likely to find new content, whereas changes below correlate with increased subsequent interaction. Our results can be used to improve people's search experience during the course of a single query by seamlessly providing new, more relevant content as the user interacts with a search result page, helping them find what they are looking for without having to issue a new query. Jaime Teevan, Sebastian de la Chica |
SIGIR | 2 |
| 2014 | CiteSight: supporting contextual citation recommendation using differential searchabstractA person often uses a single search engine for very different tasks. For example, an author editing a manuscript may use the same academic search engine to find the latest work on a particular topic or to find the correct citation for a familiar article. The author's tolerance for latency and accuracy may vary according to task. However, search engines typically employ a consistent approach for processing all queries. In this paper we explore how a range of search needs and expectations can be supported within a single search system using differential search. We introduce CiteSight, a system that provides personalized citation recommendations to author groups that vary based on task. CiteSight presents cached recommendations instantaneously for online tasks (e.g., active paper writing), and refines these recommendations in the background for offline tasks (e.g., future literature review). We develop an active cache-warming process to enhance the system as the author works, and context-coupling, a technique for augment sparse citation networks. By evaluating the quality of the recommendations and collecting user feedback, we show that differential search can provide a high level of accuracy for different tasks on different time scales. We believe that differential search can be used in many situations where the user's tolerance for latency and desired response vary dramatically based on use. Avishay Livne, Vivek Gokuladas, Jaime Teevan, Susan T. Dumais, Eytan Adar |
SIGIR | 3 |
| 2014 | Lessons from the journey: a query log analysis of within-session learningabstractThe Internet is the largest source of information in the world. Search engines help people navigate the huge space of available data in order to acquire new skills and knowledge. In this paper, we present an in-depth analysis of sessions in which people explicitly search for new knowledge on the Web based on the log files of a popular search engine. We investigate within-session and cross-session developments of expertise, focusing on how the language and search behavior of a user on a topic evolves over time. In this way, we identify those sessions and page visits that appear to significantly boost the learning process. Our experiments demonstrate a strong connection between clicks and several metrics related to expertise. Based on models of the user and their specific context, we present a method capable of automatically predicting, with good accuracy, which clicks will lead to enhanced learning. Our findings provide insight into how search engines might better help users learn as they search. Carsten Eickhoff, Jaime Teevan, Ryen W. White, Susan T. Dumais |
WSDM | 2 |
| 2013 | Understanding how people interact with web search results that change in real-time using implicit feedbackabstractThe way a searcher interacts with query results can reveal a lot about what is being sought. Considerable research has gone into using implicit relevance feedback to identify relevant con-tent in real-time, but little is known about how to best present this newly identified relevant content to users. In this paper we compare a traditional search interface with one that dynamical-ly re-ranks and recommends search results as the user interacts with it in order to build a picture of how and when users should be offered dynamically identified relevant content. We present several studies that compare logged behavior for hun-dreds of thousands of users and millions of queries as well as self-reported measures of success across the two interaction models. Compared to traditional web search, users presented with dynamically ranked results exhibit higher engagement and find information faster, particularly during exploratory tasks. These findings have implications for how search engines might best exploit implicit feedback in real-time in order to help users identify the most relevant results as quickly as possible. Jin Young Kim 0005, Mark Cramer, Jaime Teevan, Dmitry Lagun |
CIKM | 3 |
| 2013 | A Crowd-Powered Socially Embedded Search Engine
Jin-Woo Jeong, Meredith Ringel Morris, Jaime Teevan, Daniel J. Liebling |
ICWSM | 3 |
| 2013 | Towards Supporting Search over Trending Events with Social Media
Sanjay Ram Kairam, Meredith Ringel Morris, Jaime Teevan, Daniel J. Liebling, Susan T. Dumais |
ICWSM | 3 |
| 2013 | 3rd workshop on context-awareness in retrieval and recommendationabstractContext-aware information is widely available in various ways and is becoming more and more important for enhancing retrieval performance and recommendation results. The current main issue to cope with is not only recommending or retrieving the most relevant items and content, but defining them ad hoc. Other relevant issues include personalizing and adapting the information and the way it is displayed to the user's current situation and interests. Ubiquitous computing further provides new means for capturing user feedback on items and providing information. Matthias Böhmer 0001, Ernesto William De Luca, Alan Said, Jaime Teevan |
WSDM | 4 |
| 2013 | Behavioral dynamics on the web: Learning, modeling, and predictionabstractThe queries people issue to a search engine and the results clicked following a query change over time. For example, after the earthquake in Japan in March 2011, the query japan spiked in popularity and people issuing the query were more likely to click government-related results than they would prior to the earthquake. We explore the modeling and prediction of such temporal patterns in Web search behavior. We develop a temporal modeling framework adapted from physics and signal processing and harness it to predict temporal patterns in search behavior using smoothing, trends, periodicities, and surprises. Using current and past behavioral data, we develop a learning procedure that can be used to construct models of users' Web search activities. We also develop a novel methodology that learns to select the best prediction model from a family of predictive models for a given query or a class of queries. Experimental results indicate that the predictive models significantly outperform baseline models that weight historical evidence the same for all queries. We present two applications where new methods introduced for the temporal modeling of user behavior significantly improve upon the state of the art. Finally, we discuss opportunities for using models of temporal dynamics to enhance other areas of Web search and information retrieval. Kira Radinsky, Krysta M. Svore, Susan T. Dumais, Milad Shokouhi, Jaime Teevan, Alex Bocharov, Eric Horvitz |
ACM Trans. Inf. Syst. | 5 |
| 2012 | SearchBuddies: Bringing Search Engines into the Conversation
Brent J. Hecht, Jaime Teevan, Meredith Ringel Morris, Daniel J. Liebling |
ICWSM | 2 |
| 2012 | Creating temporally dynamic web search snippetsabstractContent on the Internet is always changing. We explore the value of biasing search result snippets towards new webpage content. We present results from a user study comparing traditional query-focused snippets with snippets that emphasize new page content for two query types: general and trending. Our results indicate that searchers prefer the inclusion of temporal information for trending queries but not for general queries, and that this is particularly valuable for pages that have not been recently crawled. Krysta M. Svore, Jaime Teevan, Susan T. Dumais, Anagha Kulkarni 0001 |
SIGIR | 2 |
| 2012 | Modeling and predicting behavioral dynamics on the webabstractUser behavior on the Web changes over time. For example, the queries that people issue to search engines, and the underlying informational goals behind the queries vary over time. In this paper, we examine how to model and predict this temporal user behavior. We develop a temporal modeling framework adapted from physics and signal processing that can be used to predict time-varying user behavior using smoothing and trends. We also explore other dynamics of Web behaviors, such as the detection of periodicities and surprises. We develop a learning procedure that can be used to construct models of users' activities based on features of current and historical behaviors. The results of experiments indicate that by using our framework to predict user behavior, we can achieve significant improvements in prediction compared to baseline models that weight historical evidence the same for all queries. We also develop a novel learning algorithm that explicitly learns when to apply a given prediction model among a set of such models. Our improved temporal modeling of user behavior can be used to enhance query suggestions, crawling policies, and result ranking. Kira Radinsky, Krysta M. Svore, Susan T. Dumais, Jaime Teevan, Alex Bocharov, Eric Horvitz |
WWW | 4 |
| 2011 | Factors Affecting Response Quantity, Quality, and Speed for Questions Asked Via Social Network Status Messages
Jaime Teevan, Meredith Ringel Morris, Katrina Panovich |
ICWSM | 1 |
| 2011 | Culture Matters: A Survey Study of Social Q&A Behavior
Meredith Ringel Morris, Jaime Teevan, Lada A. Adamic, Mark S. Ackerman |
ICWSM | 3 |
| 2011 | Modeling and analysis of cross-session search tasksabstractThe information needs of search engine users vary in complexity, depending on the task they are trying to accomplish. Some simple needs can be satisfied with a single query, whereas others require a series of queries issued over a longer period of time. While search engines effectively satisfy many simple needs, searchers receive little support when their information needs span session boundaries. In this work, we propose methods for modeling and analyzing user search behavior that extends over multiple search sessions. We focus on two problems: (i) given a user query, identify all of the related queries from previous sessions that the same user has issued, and (ii) given a multi-query task for a user, predict whether the user will return to this task in the future. We model both problems within a classification framework that uses features of individual queries and long-term user search behavior at different granularity. Experimental evaluation of the proposed models for both tasks indicates that it is possible to effectively model and analyze cross-session search behavior. Our findings have implications for improving search for complex information needs and designing search engine features to support cross-session search tasks. Alexander Kotov 0001, Paul N. Bennett, Ryen W. White, Susan T. Dumais, Jaime Teevan |
SIGIR | 5 |
| 2011 | Understanding temporal query dynamicsabstractWeb search is strongly influenced by time. The queries people issue change over time, with some queries occasionally spiking in popularity (e.g., earthquake) and others remaining relatively constant (e.g., youtube). The documents indexed by the search engine also change, with some documents always being about a particular query (e.g., the Wikipedia page on earthquakes is about the query earthquake) and others being about the query only at a particular point in time (e.g., the New York Times is only about earthquakes following a major seismic activity). The relationship between documents and queries can also change as people's intent changes (e.g., people sought different content for the query earthquake before the Haitian earthquake than they did after). In this paper, we explore how queries, their associated documents, and the query intent change over the course of 10 weeks by analyzing query log data, a daily Web crawl, and periodic human relevance judgments. We identify several interesting features by which changes to query popularity can be classified, and show that presence of these features, when accompanied by changes in result content, can be a good indicator of change in query intent. Anagha Kulkarni 0001, Jaime Teevan, Krysta M. Svore, Susan T. Dumais |
WSDM | 2 |
| 2011 | Understanding and predicting personal navigationabstractThis paper presents an algorithm that predicts with very high accuracy which Web search result a user will click for one sixth of all Web queries. Prediction is done via a straightforward form of personalization that takes advantage of the fact that people often use search engines to re-find previously viewed resources. In our approach, an individual's past navigational behavior is identified via query log analysis and used to forecast identical future navigational behavior by the same individual. We compare the potential value of personal navigation with general navigation identified using aggregate user behavior. Although consistent navigational behavior across users can be useful for identifying a subset of navigational queries, different people often use the same queries to navigate to different resources. This is true even for queries comprised of unambiguous company names or URLs and typically thought of as navigational. We build an understanding of what personal navigation looks like, and identify ways to improve its coverage and accuracy by taking advantage of people's consistency over time and across groups of individuals. Jaime Teevan, Daniel J. Liebling, Gayathri Ravichandran Geetha |
WSDM | 1 |
| 2011 | #TwitterSearch: a comparison of microblog search and web searchabstractSocial networking Web sites are not just places to maintain relationships; they can also be valuable information sources. However, little is known about how and why people search socially-generated content. In this paper we explore search behavior on the popular microblogging/social networking site Twitter. Using analysis of large-scale query logs and supplemental qualitative data, we observe that people search Twitter to find temporally relevant information (e.g., breaking news, real-time content, and popular trends) and information related to people (e.g., content directed at the searcher, information about people of interest, and general sentiment and opinion). Twitter queries are shorter, more popular, and less likely to evolve as part of a session than Web queries. It appears people repeat Twitter queries to monitor the associated search results, while changing and developing Web queries to learn about a topic. The results returned from the different corpora support these different uses, with Twitter results including more social chatter and social events, and Web results containing more basic facts and navigational content. We discuss the implications of these findings for the design of next-generation Web search tools that incorporate social media. Jaime Teevan, Daniel Ramage, Meredith Ringel Morris |
WSDM | 1 |
| 2011 | Addressing people's information needs directly in a web search result pageabstractWeb search engines have historically focused on connecting people with information resources. For example, if a person wanted to know when their flight to Hyderabad was leaving, a search engine might connect them with the airline where they could find flight status information. However, search engines have recently begun to try to meet people's search needs directly, providing, for example, flight status information in response to queries that include an airline and a flight number. In this paper, we use large scale query log analysis to explore the challenges a search engine faces when trying to meet an information need directly in the search result page. We look at how people's interaction behavior changes when inline content is returned, finding that such content can cannibalize clicks from the algorithmic results. We see that in the absence of interaction behavior, an individual's repeat search behavior can be useful in understanding the content's value. We also discuss some of the ways user behavior can be used to provide insight into when inline answers might better trigger and what types of additional information might be included in the results. Lydia B. Chilton, Jaime Teevan |
WWW | 2 |
| 2010 | A Comparison of Information Seeking Using Search Engines and Social Networks
Meredith Ringel Morris, Jaime Teevan, Katrina Panovich |
ICWSM | 2 |
| 2010 | Large scale query log analysis of re-findingabstractAlthough Web search engines are targeted towards helping people find new information, people regularly use them to re-find Web pages they have seen before. Researchers have noted the existence of this phenomenon, but relatively little is understood about how re-finding behavior differs from the finding of new information. This paper dives deeply into the differences via analysis of three large-scale data sources: 1) query logs (queries, clicks, result impressions), 2) Web browsing logs (URL visits), and 3) a daily Web crawl (page content). It appears that people learn valuable information about the pages they find that helps them re-find what they are looking for later; compared to the initial finding query, re-finding queries are typically shorter, and rank the re-found URL higher. While many instances of re-finding probably serve as a type of bookmark for a known URL, others seem to represent the resumption of a previous task; results clicked at the end of a session are more likely than those at the beginning to be re-found during a later session, while re-finding is more likely to happen at the beginning of a session than at the end. Additionally, we observe differences in cross-session and intra-session re-finding that may indicate different types of re-finding tasks. Our findings suggest there is a rich opportunity for search engines to take advantage of re-finding behavior as a means to improve the search experience. Sarah K. Tyler, Jaime Teevan |
WSDM | 2 |
| 2009 | The web changes everything: understanding the dynamics of web contentabstractThe Web is a dynamic, ever changing collection of information. This paper explores changes in Web content by analyzing a crawl of 55,000 Web pages, selected to represent different user visitation patterns. Although change over long intervals has been explored on random (and potentially unvisited) samples of Web pages, little is known about the nature of finer grained changes to pages that are actively consumed by users, such as those in our sample. We describe algorithms, analyses, and models for characterizing changes in Web content, focusing on both time (by using hourly and sub-hourly crawls) and structure (by looking at page-, DOM-, and term-level changes). Change rates are higher in our behavior-based sample than found in previous work on randomly sampled pages, with a large portion of pages changing more than hourly. Detailed content and structure analyses identify stable and dynamic content within each page. The understanding of Web change we develop in this paper has implications for tools designed to help people interact with dynamic Web content, such as search engines, advertising, and Web browsers. Eytan Adar, Jaime Teevan, Susan T. Dumais, Jonathan L. Elsas |
WSDM | 2 |
| 2009 | Discovering and using groups to improve personalized searchabstractPersonalized Web search takes advantage of information about an individual to identify the most relevant results for that person. A challenge for personalization lies in collecting user profiles that are rich enough to do this successfully. One way an individual's profile can be augmented is by using data from other people. To better understand whether groups of people can be used to benefit personalized search, we explore the similarity of query selection, desktop information, and explicit relevance judgments across people grouped in different ways. The groupings we explore fall along two dimensions: the longevity of the group members' relationship, and how explicitly the group is formed. We find that some groupings provide valuable insight into what members consider relevant to queries related to the group focus, but that it can be difficult to identify valuable groups implicitly. Building on these findings, we explore an algorithm to "groupize" (versus "personalize") Web search results that leads to a significant improvement in result ranking on group-relevant queries. Jaime Teevan, Meredith Ringel Morris, Steve Bush |
WSDM | 1 |
| 2009 | Characterizing the influence of domain expertise on web search behaviorabstractDomain experts search differently than people with little or no domain knowledge. Previous research suggests that domain experts employ different search strategies and are more successful in finding what they are looking for than non-experts. In this paper we present a large-scale, longitudinal, log-based analysis of the effect of domain expertise on web search behavior in four different domains (medicine, finance, law, and computer science). We characterize the nature of the queries, search sessions, web sites visited, and search success for users identified as experts and non-experts within these domains. Large-scale analysis of real-world interactions allows us to understand how expertise relates to vocabulary, resource use, and search task under more realistic search conditions than has been possible in previous small-scale studies. Building upon our analysis we develop a model to predict expertise based on search behavior, and describe how knowledge about domain expertise can be used to present better results and query suggestions to users and to help non-experts gain expertise. Ryen W. White, Susan T. Dumais, Jaime Teevan |
WSDM | 3 |
| 2008 | To personalize or not to personalize: modeling queries with variation in user intentabstractIn most previous work on personalized search algorithms, the results for all queries are personalized in the same manner. However, as we show in this paper, there is a lot of variation across queries in the benefits that can be achieved through personalization. For some queries, everyone who issues the query is looking for the same thing. For other queries, different people want very different results even though they express their need in the same way. We examine variability in user intent using both explicit relevance judgments and large-scale log analysis of user behavior patterns. While variation in user behavior is correlated with variation in explicit relevance judgments the same query, there are many other factors, such as result entropy, result quality, and task that can also affect the variation in behavior. We characterize queries using a variety of features of the query, the results returned for the query, and people's interaction history with the query. Using these features we build predictive models to identify queries that can benefit from personalization. Jaime Teevan, Susan T. Dumais, Daniel J. Liebling |
SIGIR | 1 |
| 2008 | How medical expertise influences web search interactionabstractDomain expertise can have an important influence on how people search. In this poster we present findings from a log-based study into how medical domain experts search the Web for information related to their expertise, as compared with non-experts. We find differences in sites visited, query vocabulary, and search behavior. The findings have implications for the automatic identification of domain experts from interaction logs, and the use of domain knowledge in applications such as query suggestion or page recommendation to support non-experts. Ryen W. White, Susan T. Dumais, Jaime Teevan |
SIGIR | 3 |
| 2008 | How people recall, recognize, and reuse search resultsabstractWhen a person issues a query, that person has expectations about the search results that will be returned. These expectations can be based on the current information need, but are also influenced by how the searcher believes the search engine works, where relevant results are expected to be ranked, and any previous searches the individual has run on the topic. This paper looks in depth at how the expectations people develop about search result lists during an initial query affect their perceptions of and interactions with future repeat search result lists. Three studies are presented that give insight into how people recall, recognize, and reuse results. The first study (a study of recall ) explores what people recall about previously viewed search result lists. The second study (a study of recognition ) builds on the first to reveal that people often recognize a result list as one they have seen before even when it is quite different. As long as those aspects that the searcher remembers about the initial list remain the same, other aspects can change significantly. This is advantageous because, as the third study (a study of reuse ) shows, when a result list appears to have changed, people have trouble re-using the previously viewed content in the list. They are less likely to find what they are looking for, less happy with the result quality, more likely to find the task hard, and more likely to take a long time searching. Although apparent consistency is important for reuse, people's inability to recognize change makes consistency without stagnation possible. New relevant results can be presented where old results have been forgotten, making both old and new content easy to find. Jaime Teevan |
ACM Trans. Inf. Syst. | 1 |
| 2007 | Information re-retrieval: repeat queries in Yahoo's logsabstractPeople often repeat Web searches, both to find new information on topics they have previously explored and to re-find information they have seen in the past. The query associated with a repeat search may differ from the initial query but can nonetheless lead to clicks on the same results. This paper explores repeat search behavior through the analysis of a one-year Web query log of 114 anonymous users and a separate controlled survey of an additional 119 volunteers. Our study demonstrates that as many as 40% of all queries are re-finding queries. Re-finding appears to be an important behavior for search engines to explicitly support, and we explore how this can be done. We demonstrate that changes to search engine results can hinder re-finding, and provide a way to automatically detect repeat searches and predict repeat clicks. Jaime Teevan, Eytan Adar, Rosie Jones, Michael A. S. Potts |
SIGIR | 1 |
| 2007 | Characterizing the value of personalizing searchabstractWe investigate the diverse goals that people have when they issue the same query to a search engine, and the ability of current search engines to address such diversity. We quantify the potential value of personalizing search results based on this analysis. Great variance was found in the results that different individuals rated as relevant for the same query -- even when the same information goal was expressed. Our analysis suggests that while search engines do a good job of ranking results to maximize global happiness, they do not do a very good job for specific individuals. Jaime Teevan, Susan T. Dumais, Eric Horvitz |
SIGIR | 1 |
| 2006 | History repeats itself: repeat queries in Yahoo's logsabstractThanks to the ubiquity of the Internet search engine search box, users have come to depend on search engines both to find and re-find information. However, re-finding behavior has not been significantly addressed. Here we look at re-finding queries issued to the Yahoo! search engine by 114 users over a year. Jaime Teevan, Eytan Adar, Rosie Jones, Michael A. S. Potts |
SIGIR | 1 |
| 2005 | Personalizing search via automated analysis of interests and activitiesabstractWe formulate and study search algorithms that consider a user's prior interactions with a wide variety of content to personalize that user's current Web search. Rather than relying on the unrealistic assumption that people will precisely specify their intent when searching, we pursue techniques that leverage implicit information about the user's interests. This information is used to re-rank Web search results within a relevance feedback framework. We explore rich models of user interests, built from both search-related information, such as previously issued queries and previously visited Web pages, and other information about the user such as documents and email the user has read and created. Our research suggests that rich representations of the user and the corpus are important for personalization, but that it is possible to approximate these representations and provide efficient client-side algorithms for personalizing search. We show that such personalization algorithms can significantly improve on current Web search. Jaime Teevan, Susan T. Dumais, Eric Horvitz |
SIGIR | 1 |
| 2003 | Empirical development of an exponential probabilistic model for text retrieval: using textual analysis to build a better modelabstractMuch work in information retrieval focuses on using a model of documents and queries to derive retrieval algorithms. Model based development is a useful alternative to heuristic development because in a model the assumptions are explicit and can be examined and refined independent of the particular retrieval algorithm. We explore the explicit assumptions underlying the naïve framework by performing computational analysis of actual corpora and queries to devise a generative document model that closely matches text. Our thesis is that a model so developed will be more accurate than existing models, and thus more useful in retrieval, as well as other applications. We test this by learning from a corpus the best document model. We find the learned model better predicts the existence of text data and has improved performance on certain IR tasks. Jaime Teevan, David R. Karger |
SIGIR | 1 |