EDBT 2026 Demo / reviewers in the wild / expert
Rosie Jones
dblp:40/5446
· DBLP profile ↗
39ranked-venue papers in the field
12as first author
7since 2021 · last 2023
0009-0000-3821-1207ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 30 (9 first)Data Mining & Knowledge Discovery · 6 (2 first)Database Systems & Data Management · 2 (1 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Learning to Understand Audio and Multimodal Content
Rosie Jones |
WSDM | 1 |
| 2022 | CHIIR Workshop on Audio Collection Human Interaction (AudioCHI 2022): http: //speechretrievalworkshop.github.ioabstractThe AudioCHI 2022 workshop focusses on human engagement with spoken material in search settings, including live stream audio and collections. Spoken material comes in many forms, including for example: factual or entertaining (or both!), timely or of historical interest, local or global, single speaker or conversations. Users engage with spoken material for a variety of reasons, including entertainment, current affairs, education, and research. While there has been considerable previous work studying spoken document retrieval or more generally spoken content retrieval, AudioCHI 2022 is the first meeting to explore user engagement with audio content, including discussing: (i) how content analysis might establish verbal and non-verbal features for rich content representations, and (ii) and use cases and human factors in interaction with spoken audio content, and their interaction with more established topics relating to spoken content retrieval. The workshop brings together researchers in spoken content retrieval with expertise on human computer interaction in information access to examine opportunities and challenges for advancing technologies for search and interaction with spoken content. Gareth J. F. Jones, Maria Eskevich, Ben Carterette, Joana Correia, Rosie Jones, Jussi Karlgren, Ian Soboroff |
CHIIR | 5 |
| 2022 | The Contribution of Lyrics and Acoustics to Collaborative Understanding of Mood
Shahrzad Naseri, Sravana Reddy, Joana Correia, Jussi Karlgren, Rosie Jones |
ICWSM | 5 |
| 2022 | What Makes a Good Podcast Summary?abstractAbstractive summarization of podcasts is motivated by the growing popularity of podcasts and the needs of their listeners. Podcasting is a markedly different domain from news and other media that are commonly studied in the context of automatic summarization. As such, the qualities of a good podcast summary are yet unknown. Using a collection of podcast summaries produced by different algorithms alongside human judgments of summary quality obtained from the TREC 2020 Podcasts Track, we study the correlations between various automatic evaluation metrics and human judgments, as well as the linguistic aspects of summaries that result in strong evaluations. Rezvaneh Rezapour, Sravana Reddy, Rosie Jones, Ian Soboroff |
SIGIR | 3 |
| 2021 | PodRecs 2021: 2nd Workshop on Podcast RecommendationsabstractPodcasts have continued to experience rapid growth in both cultural relevance as well as research attention. Coming off the success of the first PodRecs Workshop for Podcast Recommendations at RecSys in 2020, as well as to build upon the research datasets and prior work released in the last year, the second PodRecs Workshop for Podcast Recommendations was held at RecSys 2021 to further develop the community of researchers and practitioners interested in the recommendation of podcasts. Ching-Wei Chen, Rosie Jones, Zahra Nazari, Longqi Yang 0001, Maria Eskevich, Gareth J. F. Jones, Sergio Oramas |
RecSys | 2 |
| 2021 | Podcast Metadata and Content: Episode Relevance and Attractiveness in Ad Hoc SearchabstractRapidly growing online podcast archives contain diverse content on a wide range of topics. These archives form an important resource for entertainment and professional use, but their value can only be realized if users can rapidly and reliably locate content of interest. Search for relevant content can be based on metadata provided by content creators, but also on transcripts of the spoken content itself. Excavating relevant content from deep within these audio streams for diverse types of information needs requires varying the approach to systems prototyping. We describe a set of diverse podcast information needs and different approaches to assessing retrieved content for relevance. We use these information needs in an investigation of the utility and effectiveness of these information sources. Based on our analysis, we recommend approaches for indexing and retrieving podcast content for ad hoc search. Ben Carterette, Rosie Jones, Gareth J. F. Jones, Maria Eskevich, Sravana Reddy, Ann Clifton, Jussi Karlgren, Ian Soboroff |
SIGIR | 2 |
| 2021 | Current Challenges and Future Directions in Podcast Information AccessabstractPodcasts are spoken documents across a wide-range of genres and styles, with growing listenership across the world, and a rapidly lowering barrier to entry for both listeners and creators. The great strides in search and recommendation in research and industry have yet to see impact in the podcast space, where recommendations are still largely driven by word of mouth. In this perspective paper, we highlight the many differences between podcasts and other media, and discuss our perspective on challenges and future research directions in the domain of podcast information access. Rosie Jones, Hamed Zamani, Markus Schedl, Ching-Wei Chen, Sravana Reddy, Ann Clifton, Jussi Karlgren, Helia Hashemi, Aasish Pappu, Zahra Nazari, Longqi Yang 0001, Oguz Semerci, Hugues Bouchard, Ben Carterette |
SIGIR | 1 |
| 2020 | PodRecs: Workshop on Podcast RecommendationsabstractThe last year has been a breakout year for podcasts. There are now over 1 million podcast shows and over 64 million podcast episodes available through public RSS feeds. In the United States, 32% of all people listened to a podcast every month, and forecasts point to global podcast listenership to reach 2.2 billion monthly listeners by 2024. The workshop on Podcast Recommendations (PodRecs), collocated with RecSys 2020, introduces researchers in other domains of recommender systems to the special characteristics and challenges of podcast recommendations: how podcasts are as created, consumed, and how we might see algorithms being designed specifically for podcast content. We hope this workshop will help grow a community of researchers and foster an active research and innovation in the field. Ching-Wei Chen, Longqi Yang 0001, Hongyi Wen, Rosie Jones, Vladan Radosavljevic, Hugues Bouchard |
RecSys | 4 |
| 2020 | The New TREC Track on Podcast Search and SummarizationabstractPodcasts are exploding in popularity. As this medium grows, it becomes increasingly important to understand the content of podcasts (e.g. what exactly is being covered, by whom, and how?), and how we can use this to connect users to shows that align with their interests. Given the explosion of new material, how do listeners find the needle in the haystack, and connect to those shows or episodes that speak to them? Furthermore, once they are presented with potential podcasts to listen to, how can they decide if this is what they want? Rosie Jones |
SIGIR | 1 |
| 2015 | Characterizing and Predicting Voice Query ReformulationabstractVoice interactions are becoming more prevalent as the usage of voice search and intelligent assistants gains more popularity. Users frequently reformulate their requests in hope of getting better results either because the system was unable to recognize what they said or because it was able to recognize it but was unable to return the desired response. Query reformulation has been extensively studied in the context of text input. Many of the characteristics studied in the context of text query reformulation are potentially useful for voice query reformulation. However, voice query reformulation has its unique characteristics in terms of the reasons that lead users to reformulating their queries and how they reformulate them. In this paper, we study the problem of voice query reformulation. We perform a large scale human annotation study to collect thousands of labeled instances of voice reformulation and non-reformulation query pairs. We use this data to compare and contrast characteristics of reformulation and non-reformulation queries over a large a number of dimensions. We then train classifiers to distinguish between reformulation and non-reformulation query pairs and to predict the rationale behind reformulation. We demonstrate through experiments with the human labeled data that our classifiers achieve good performance in both tasks. Ahmed Awadallah 0001, Ranjitha Gurunath Kulkarni, Umut Ozertem, Rosie Jones |
CIKM | 4 |
| 2015 | Automatic Online Evaluation of Intelligent AssistantsabstractVoice-activated intelligent assistants, such as Siri, Google Now, and Cortana, are prevalent on mobile devices. However, it is challenging to evaluate them due to the varied and evolving number of tasks supported, e.g., voice command, web search, and chat. Since each task may have its own procedure and a unique form of correct answers, it is expensive to evaluate each task individually. This paper is the first attempt to solve this challenge. We develop consistent and automatic approaches that can evaluate different tasks in voice-activated intelligent assistants. We use implicit feedback from users to predict whether users are satisfied with the intelligent assistant as well as its components, i.e., speech recognition and intent classification. Using this approach, we can potentially evaluate and compare different tasks within and across intelligent assistants ac-cording to the predicted user satisfaction rates. Our approach is characterized by an automatic scheme of categorizing user-system interaction into task-independent dialog actions, e.g., the user is commanding, selecting, or confirming an action. We use the action sequence in a session to predict user satisfaction and the quality of speech recognition and intent classification. We also incorporate other features to further improve our approach, including features derived from previous work on web search satisfaction prediction, and those utilizing acoustic characteristics of voice requests. We evaluate our approach using data collected from a user study. Results show our approach can accurately identify satisfactory and unsatisfactory sessions. Jiepu Jiang, Ahmed Awadallah 0001, Rosie Jones, Umut Ozertem, Imed Zitouni, Ranjitha Gurunath Kulkarni, Omar Zia Khan |
WWW | 3 |
| 2014 | Mobile query reformulationsabstractUsers frequently interact with web search systems on their mobile devices via multiple modalities, including touch and speech. These interaction modes are substantially different from the user experience on desktop search. As a result, system designers have new challenges and questions around understanding the intent on these platforms. In this paper, we study the query reformulation patterns in mobile logs. We group query reformulations based on their input method into four categories; text-text, text-voice, voice-text and voice-voice. We discuss the unique characteristics of each of these groups by comparing them against each other and desktop logs. We also compare the distribution of reformulation types (e.g. adding/dropping words) against desktop logs and show that there are new classes of reformulations that are caused by errors in speech recognition. Our results suggest that users do not tend to switch between different input types (e.g. voice and text). Voice to text switches are largely caused by speech recognition errors, and text to voice switches are unlikely to be about the same intent. Milad Shokouhi, Rosie Jones, Umut Ozertem, Karthik Raghunathan, Fernando Diaz 0001 |
SIGIR | 2 |
| 2011 | Classification of proxy labeled examples for marketing segment generationabstractMarketers often rely on a set of descriptive segments, or qualitative subsets of the population, to specify the audiences of targeted advertising campaigns. For example, the descriptive segment "Empty Nesters" might describe a desirable target audience for extended vacation package offers. While some segments may be easily described and generated using demographic data as ground truth, others such as "Soccer Moms" or "Urban Hipsters" reflect a combination of demographic and behavioral attributes. Ideally, these attributes would be available as the basis for ground truth labeling of a classifier training set or even direct member selection from the population. Unfortunately, ground truth attributes are often scarce or unavailable, in which case a proxy labeling scheme is needed. Dean Cerrato, Rosie Jones, Avinash Gupta |
KDD | 2 |
| 2011 | Evaluating new search engine configurations with pre-existing judgments and clicksabstractWe provide a novel method of evaluating search results, which allows us to combine existing editorial judgments with the relevance estimates generated by click-based user browsing models. There are evaluation methods in the literature that use clicks and editorial judgments together, but our approach is novel in the sense that it allows us to predict the impact of unseen search models without online tests to collect clicks and without requesting new editorial data, since we are only re-using existing editorial data, and clicks observed for previous result set configurations. Since the user browsing model and the pre-existing editorial data cannot provide relevance estimates for all documents for the selected set of queries, one important challenge is to obtain this performance estimation where there are a lot of ranked documents with missing relevance values. We introduce a query and rank based smoothing to overcome this problem. We show that a hybrid of these smoothing techniques performs better than both query and position based smoothing, and despite the high percentage of missing judgments, the resulting method is significantly correlated (0.74) with DCG values evaluated using fully judged datasets, and approaches inter-annotator agreement. We show that previously published techniques, applicable to frequent queries, degrade when applied to a random sample of queries, with a correlation of only 0.29. While our experiments focus on evaluation using DCG, our method is also applicable to other commonly used metrics. Umut Ozertem, Rosie Jones, Benoît Dumoulin |
WWW | 2 |
| 2010 | Predicting searcher frustrationabstractWhen search engine users have trouble finding information, they may become frustrated, possibly resulting in a bad experience (even if they are ultimately successful). In a user study in which participants were given difficult information seeking tasks, half of all queries submitted resulted in some degree of self-reported frustration. A third of all successful tasks involved at least one instance of frustration. By modeling searcher frustration, search engines can predict the current state of user frustration and decide when to intervene with alternative search strategies to prevent the user from becoming more frustrated, giving up, or switching to another search engine. We present several models to predict frustration using features extracted from query logs and physical sensors. We are able to predict frustration with a mean average precision of 65% from the physical sensors, and 87% from the query log features. Henry Allen Feild, James Allan 0001, Rosie Jones |
SIGIR | 3 |
| 2010 | Beyond DCG: user behavior as a predictor of a successful searchabstractWeb search engines are traditionally evaluated in terms of the relevance of web pages to individual queries. However, relevance of web pages does not tell the complete picture, since an individual query may represent only a piece of the user's information need and users may have different information needs underlying the same queries. In this work, we address the problem of predicting user search goal success by modeling user behavior. We show empirically that user behavior alone can give an accurate picture of the success of the user's web search goals, without considering the relevance of the documents displayed. In fact, our experiments show that models using user behavior are more predictive of goal success than those using document relevance. We build novel sequence models incorporating time distributions for this task and our experiments show that the sequence and time distribution models are more accurate than static models based on user behavior, or predictions based on document relevance. Ahmed Awadallah 0001, Rosie Jones, Kristina Lisa Klinkner |
WSDM | 2 |
| 2010 | Applications of open search toolsabstractIt costs about 300M dollars to just build a search system that scales to the web. Web services which open up web search results to the public allow academics, developers, and entrepreneurs to achieve instant web search parity for free and enable them to focus on building their additional secret sauce on top to create even grander relevant services. For example, a social network could leverage open search and their data about users to personalize web search. Additionally, one of the best ways to gather data about web search behavior is to build your own search system. Proto-type IR and web search systems based on open search can be used to gather user interaction data and test the applicability of research ideas. Open Source tools like Lucene and Nutch and open search services like from major search engines can greatly help developers implement these types of systems quickly, allowing for fast production and evaluation. We will give detailed overviews of the popular open search tools, showcasing examples of search and IR algorithms and systems implemented using them, as well as discussing how evaluation can be performed. Rosie Jones, Ted Drake |
WWW | 1 |
| 2009 | A case study of using geographic cues to predict query news intentabstractGeographic information retrieval encompasses important tasks including finding the location of a user, and locations relevant to their search queries. Web-based search engines receive queries from numerous users located in very different parts of the world. A typical way for people to find news is through a general web search engine, which makes it important for search engines to recognize queries with news intent. An important question for geographic information retrieval is how we can benefit from geographic cues to predict the intent of users. This work presents a case study of an application using geographic features to improve the quality of an important web search task, involving predicting which queries have news intent and hence are likely to receive clicks on news search results. Our case study suggests that information derived from geographic features can help the task. The information we consider includes cues derived from the location of the user, from the IP address, the location relevant to the query, automatically extracted from the query string, and the relation between the two locations. We build a classifier that uses geographical cues to predict whether a query will result in a news click or not. We compare our classifier to a strong baseline that use non-geographic click-based features and we show that our classifier outperforms the baseline for geographic queries. Ahmed Awadallah 0001, Rosie Jones, Fernando Diaz 0001 |
GIS | 2 |
| 2009 | Privacy in Web Search Query Log Mining
Rosie Jones |
ECML/PKDD (1) | 1 |
| 2009 | Improving search relevance for implicitly temporal queriesabstractNo abstract available. Donald Metzler, Rosie Jones, Fuchun Peng, Ruiqiang Zhang |
SIGIR | 2 |
| 2009 | Mining user web search activity with layered bayesian networks or how to capture a click in its contextabstractMining user web search activity potentially has a broad range of applications including web result pre-fetching, automatic search query reformulation, click spam detection, estimation of document relevance and prediction of user satisfaction. This analysis is difficult because the data recorded by search engines while users interact with them, although abundant, is very noisy. In this work, we explore the utility of mining search behavior of users, represented by observed variables including the time the user spends on the page, and whether the user reformulated his or her query. As a case study, we examine the contribution this data makes to predicting the relevance of a document in the absence of document content models. To this end, we first propose a method for grouping the interactions of a particular user according to the different tasks he or she undertakes. With each task corresponding to a distinct information need, we then propose a Bayesian Network to holistically model these interactions. The aim is to identify distinct patterns of search behaviors. Finally, we join these patterns to a list of custom features and we use gradient boosted decision trees to predict the relevance of a set of query document pairs for which we have relevance assessments. The experimental results confirm the potential of our model, with significant improvements in precision for predicting the relevance of documents based on a model of the user's search and click behavior, over a baseline model using only click and query features, with no Bayesian Network input. Benjamin Piwowarski, Georges Dupret, Rosie Jones |
WSDM | 3 |
| 2008 | Beyond the session timeout: automatic hierarchical segmentation of search topics in query logsabstractMost analysis of web search relevance and performance takes a single query as the unit of search engine interaction. When studies attempt to group queries together by task or session, a timeout is typically used to identify the boundary. However, users query search engines in order to accomplish tasks at a variety of granularities, issuing multiple queries as they attempt to accomplish tasks. In this work we study real sessions manually labeled into hierarchical tasks, and show that timeouts, whatever their length, are of limited utility in identifying task boundaries, achieving a maximum precision of only 70%. We report on properties of this search task hierarchy, as seen in a random sample of user interactions from a major web search engine's log, annotated by human editors, learning that 17% of tasks are interleaved, and 20% are hierarchically organized. No previous work has analyzed or addressed automatic identification of interleaved and hierarchically organized search tasks. We propose and evaluate a method for the automated segmentation of users' query streams into hierarchical units. Our classifiers can improve on timeout segmentation, as well as other previously published approaches, bringing the accuracy up to 92% for identifying fine-grained task boundaries, and 89-97% for identifying pairs of queries from the same task when tasks are interleaved hierarchically. This is the first work to identify, measure and automatically segment sequences of user queries into their hierarchical structure. The ability to perform this kind of segmentation paves the way for evaluating search engines in terms of user task completion. Rosie Jones, Kristina Lisa Klinkner |
CIKM | 1 |
| 2008 | Vanity fair: privacy in querylog bundlesabstractA recently proposed approach to address privacy concerns in storing web search querylogs is bundling logs of multiple users together. In this work we investigate privacy leaks that are possible even when querylogs from multiple users are bundled together, without any user or session identifiers. We begin by quantifying users' propensity to issue own-name vanity queries and geographically revealing queries. We show that these propensities interact badly with two forms of vulnerabilities in the bundling scheme. First, structural vulnerabilities arise due to properties of the heavy tail of the user search frequency distribution, or the distribution of locations that appear within a user's queries. These heavy tails may cause a user to appear visibly different from other users in the same bundle. Second, we demonstrate analytical vulnerabilities based on the ability to separate the queries in a bundle into threads corresponding to individual users. These vulnerabilities raise privacy issues suggesting that bundling must be handled with great care. Rosie Jones, Ravi Kumar 0001, Bo Pang 0001, Andrew Tomkins |
CIKM | 1 |
| 2008 | Geographic intention and modification in web searchabstractWeb searchers signal their geographic intent by using place‐names in search queries. They also indicate their flexibility about geographic specificity by reformulating their queries. By examining this data we can learn to understand web searcher flexibility with respect to geographic intent. We examine aggregated data of queries with locations, and locations identified from IP addresses, to identify overall distance preferences, as well as distance preferences by search topic. We also examine query rewriting: both deliberate query rewriting, conducted in web search sessions, and automated query rewriting, with manual relevance judgments of geo‐modified queries. We find geo‐specification in 12.7% of user query rewrites in search sessions, and show the breakdown into sub‐classes such as same‐city, same‐state, same‐country and different‐country. We also measure the dependence between US‐state‐name and distance‐of‐modified‐location‐from‐original‐location, finding that Vermont web searchers modify their locations greater distances than California web searchers. We find that automatically‐modified queries are perceived as much more relevant when the geographic component is unchanged. We look at the relationship between the non‐location part of a query and the distance from the user. We see that people search for child day‐care near their locations and maps far from where they are located. We also give distance profiles for the top topics which cooccur with place‐names in queries, which could be used to set document priors based on document proximity and query topic. Rosie Jones, Wei Vivian Zhang, Benjamin Rey, Pradhuman Jhala, Eugene Stipp |
Int. J. Geogr. Inf. Sci. | 1 |
| 2007 | "I know what you did last summer": query logs and user privacyabstractWe investigate the subtle cues to user identity that may be exploited in attacks on the privacy of users in web search query logs. We study the application of simple classifiers to map a sequence of queries into the gender, age, and location of the user issuing the queries. We then show how these classifiers may be carefully combined at multiple granularities to map a sequence of queries into a set of candidate users that is 300-600 times smaller than random chance would allow. We show that this approach remains accurate even after removing personally identifiable information such as names/numbers or limiting the size of the query log. Rosie Jones, Ravi Kumar 0001, Bo Pang 0001, Andrew Tomkins |
CIKM | 1 |
| 2007 | Information re-retrieval: repeat queries in Yahoo's logsabstractPeople often repeat Web searches, both to find new information on topics they have previously explored and to re-find information they have seen in the past. The query associated with a repeat search may differ from the initial query but can nonetheless lead to clicks on the same results. This paper explores repeat search behavior through the analysis of a one-year Web query log of 114 anonymous users and a separate controlled survey of an additional 119 volunteers. Our study demonstrates that as many as 40% of all queries are re-finding queries. Re-finding appears to be an important behavior for search engines to explicitly support, and we explore how this can be done. We demonstrate that changes to search engine results can hinder re-finding, and provide a way to automatically detect repeat searches and predict repeat clicks. Jaime Teevan, Eytan Adar, Rosie Jones, Michael A. S. Potts |
SIGIR | 3 |
| 2007 | Query rewriting using active learning for sponsored searchabstractSponsored search is a major revenue source for search companies. Web searchers can issue any queries, while advertisement keywords are limited. Query rewriting technique effectively matches user queries with relevant advertisement keywords, thus increases the amount of web advertisements available. The match relevance is critical for clicks. In this study, we aim to improve query rewriting relevance. For this purpose, we use an active learning algorithm called Transductive Experimental Design to select the most informative samples to train the query rewriting relevance model. Experiments show that this approach improves model accuracy and rewriting relevance. Wei Vivian Zhang, Xiaofei He 0001, Benjamin Rey, Rosie Jones |
SIGIR | 4 |
| 2007 | Temporal profiles of queriesabstractDocuments with timestamps, such as email and news, can be placed along a timeline. The timeline for a set of documents returned in response to a query gives an indication of how documents relevant to that query are distributed in time. Examining the timeline of a query result set allows us to characterize both how temporally dependent the topic is, as well as how relevant the results are likely to be. We outline characteristic patterns in query result set timelines, and show experimentally that we can automatically classify documents into these classes. We also show that properties of the query result set timeline can help predict the mean average precision of a query. These results show that meta-features associated with a query can be combined with text retrieval techniques to improve our understanding and treatment of text search on documents with timestamps. Rosie Jones, Fernando Diaz 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2006 | Measuring the meaning in time series clustering of text search queriesabstractWe use a combination of proven methods from time series analysis and machine learning to explore the relationship between temporal and semantic similarity in web query logs; we discover that the combination of correlation and cycles is a good, but not perfect, sign of semantic relationship. Bing Liu 0003, Rosie Jones, Kristina Lisa Klinkner |
CIKM | 2 |
| 2006 | History repeats itself: repeat queries in Yahoo's logsabstractThanks to the ubiquity of the Internet search engine search box, users have come to depend on search engines both to find and re-find information. However, re-finding behavior has not been significantly addressed. Here we look at re-finding queries issued to the Yahoo! search engine by 114 users over a year. Jaime Teevan, Eytan Adar, Rosie Jones, Michael A. S. Potts |
SIGIR | 3 |
| 2006 | Generating query substitutionsabstractWe introduce the notion of query substitution, that is, generating a new query to replace a user’s original search query. Our technique uses modifications based on typical substitutions web searchers make to their queries. In this way the new query is strongly related to the original query, containing terms closely related to all of the original terms. This contrasts with query expansion through pseudo-relevance feedback, which is costly and can lead to query drift. This also contrasts with query relaxation through boolean or TFIDF retrieval, which reduces the specificity of the query. We define a scale for evaluating query substitution, and show that our method performs well at generating new queries related to the original queries. We build a model for selecting between candidates, by using a number of features relating the query-candidate pair, and by fitting the model to human judgments of relevance of query suggestions. This further improves the quality of the candidates generated. Experiments show that our techniques significantly increase coverage and effectiveness in the setting of sponsored search. Rosie Jones, Benjamin Rey, Omid Madani, Wiley Greiner |
WWW | 1 |
| 2005 | Biasing web search results for topic familiarityabstractDepending on a web searcher's familiarity with a query's target topic, it may be more appropriate to show her introductory or advanced documents. The TREC HARD [1] track defined topic familiarity as meta-data associated with a user's query. We instead define a user-independent and query-independent model of topic-familiarity required to read a document, so it can be matched to a given user in response to a query. An introductory web page is defined as A web page that doesn't presuppose any background knowledge of the topic it is on, and to an extent introduces or defines the key terms in the topic. while an advanced web page is defined as A web page that assumes sufficient background knowledge of the topic it is on, and familiarity with the key technical/ important terms in the topic, and potentially builds on them. We develop a method for biasing the initial mix of documents returned by a search engine to increase the number of documents of desired familiarity level up to position 5, and up to position 10. Our method involves building a supervised text classifier, incorporating features based on reading level, the distribution of stop-words in the text, and non-text features such as average line-length. Using this familiarity classifier, we achieve statistically significant improvements at reranking the result set to show introductory documents higher up the ranked list. Our classifier can be seamlessly integrated into current search engine technology without involving any major modifications to existing architectures. Giridhar Kumaran, Rosie Jones, Omid Madani |
CIKM | 2 |
| 2005 | Building Minority Language Corpora by Learning to Generate Web Search Queries
Rayid Ghani, Rosie Jones, Dunja Mladenic |
Knowl. Inf. Syst. | 2 |
| 2004 | Using temporal profiles of queries for precision predictionabstractA key missing component in information retrieval systems is self-diagnostic tests to establish whether the system can provide reasonable results for a given query on a document collection. If we can measure properties of a retrieved set of documents which allow us to predict average precision, we can automate the decision of whether to elicit relevance feedback, or modify the retrieval system in other ways. We use meta-data attached to documents in the form of time stamps to measure the distribution of documents retrieved in response to a query, over the time domain, to create a temporal profile for a query. We define some useful features over this temporal profile. We find that using these temporal features, together with the content of the documents retrieved, we can improve the prediction of average precision for a query. Fernando Diaz 0001, Rosie Jones |
SIGIR | 2 |
| 2003 | Query word deletion predictionabstractWeb search query logs contain traces of users' search modifications. One strategy users employ is deleting terms, presumably to obtain greater coverage. It is useful to model and automate term deletion when arbitrary searches are conjunctively matched against a small hand constructed collection, such as a hand-built hierarchy, or collection of high-quality pages matched with key phrases. Queries with no matches can have words deleted till a match is obtained. We provide algorithms which perform substantially better than the baseline in predicting which word should be deleted from a reformulated query, for increasing query coverage in the context of web search on small high-quality collections. Rosie Jones, Daniel C. Fain |
SIGIR | 1 |
| 2001 | Mining the Web to Create Minority Language CorporaabstractThe Web is a valuable source of language specific resources but the process of collecting, organizing and utilizing these resources is difficult. We describe CorpusBuilder, an approach for automatically generating Web-search queries for collecting documents in a minority language. It differs from pseudo-relevance feedback in that retrieved documents are labeled by an automatic language classifier as relevant or irrelevant, and this feedback is used to generate new queries. We experiment with various query-generation methods and query-lengths to find inclusion/exclusion terms that are helpful for retrieving documents in the target language and find that using odds-ratio scores calculated over the documents acquired so far was one of the most consistently accurate query-generation methods. We also describe experiments using a handful of words elicited from a user instead of initial documents and show that the methods perform similarly. Experiments applying the same approach to multiple languages are also presented showing that our approach generalizes to a variety of languages. Rayid Ghani, Rosie Jones, Dunja Mladenic |
CIKM | 2 |
| 2001 | Automatic Web Search Query Generation to Create Minority Language CorporaabstractThe Web is a valuable source of language specific resources but collecting, organizing and utilizing this information is difficult. We describe CorpusBuilder, an approach for automatically generating Web-search queries to collect documents in a minority language. It differs from pseudo-relevance feedback in that retrieved documents are labeled by an automatic language classifier as relevant or irrelevant and a subset of documents is used to generate new queries. We experiment with various query-generation methods and query-lengths to find inclusion/exclusion terms that are helpful for finding documents in the target language and find that using odds-ratio scores calculated over the documents acquired so far was one of the most consistently accurate query-generation methods. We also describe experiments using a handful of words elicited from a user instead of initial documents and show that the methods perform similarly. Applying the same approach to multiple languages show that our system generalizes to a variety of languages. Rayid Ghani, Rosie Jones, Dunja Mladenic |
SIGIR | 2 |
| 2001 | Online Learning for Web Query Generation: Finding Documents Matching a Minority Concept on the Web
Rayid Ghani, Rosie Jones, Dunja Mladenic |
Web Intelligence | 2 |
| 2000 | Learning a Monolingual Language Model from a Multilingual Text Databaseabstract9999 9999 9999 9999 9999 9999 9999 9999 9999 + ,\t,+- , "#$\t!\t! %&\t' !($(*) 29039-47500 4* $%!3 29039-46450 &\t' !($(*) 29039-47500 '"6= 9 !$:\t! ,"<; 29659-45410 9039-47500 G"+\t\t!G"H '$/ !$"= 9039-47500 (/ ,"./ J+; 28939-43300 / !$"= 9039-47500 1\t'H+""Q%G1:\t'\t,H$%(DEH1:\t\t,R:\t' !$F\t@(-(/("/ $%\t !(+\t,+=(J""%\t!\t, ,R:\t' !$F\t@(-(/("/ 2 +\tV\t!F("$"#W\t!X\t/$(YV$%(+ (/("/ 2 $+\t,"$%/\t- '\tJ8\t!("""E(/(/ ("/ 28800 <\\?-+$"""/(:+(1/ ""E(/(/ ("/ 28800 1\t'] $%/\t>\t,6; 28969-37040 ""E(/(/ ("... Rayid Ghani, Rosie Jones |
CIKM | 2 |