EDBT 2026 Demo / reviewers in the wild / expert
Miles Efron
dblp:41/3635
· DBLP profile ↗
25ranked-venue papers
18as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 25 · 18 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
13 papers |
Information retrieval · 98% Web and social media mining · 1% Data mining · 1% | |
| Theoretical computer science
1 paper |
Information theory · 50% Coding theory · 50% |
Topics — the 30 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › document retrieval
temporal information retrieval |
0.8 | 5 | 2016 | What Makes a Query Temporally Sensitive? · SIGIR 2016 Temporal feedback for tweet search with non-parametric density estimation · SIGIR 2014 SIGIR 2013 workshop on time aware information access (#TAIA2013) · SIGIR 2013 |
Information retrieval
evaluation |
0.7 | 3 | 2016 | What Makes a Query Temporally Sensitive? · SIGIR 2016 Assessor Differences and User Preferences in Tweet Timeline Generation · SIGIR 2015 On run diversity in Evaluation as a Service · SIGIR 2014 |
Information retrieval › document processing › document analysis › document representation
document expansion |
0.4 | 2 | 2017 | Document Expansion Using External Collections · SIGIR 2017 Improving retrieval of short texts through document expansion · SIGIR 2012 |
Information retrieval › relevance feedback
pseudo-relevance feedback |
0.4 | 2 | 2017 | Document Expansion Using External Collections · SIGIR 2017 Estimation methods for ranking recent information · SIGIR 2011 |
Information retrieval › ranking › search relevance
relevance modeling |
0.3 | 1 | 2017 | Document Expansion Using External Collections · SIGIR 2017 |
Information retrieval › web search › web information retrieval › social media retrieval
microblog retrieval |
0.3 | 2 | 2012 | Improving retrieval of short texts through document expansion · SIGIR 2012 Hashtag retrieval in a microblogging environment · SIGIR 2010 |
Information retrieval › evaluation
test collection |
0.2 | 1 | 2016 | What Makes a Query Temporally Sensitive? · SIGIR 2016 |
Information retrieval › retrieval models
boolean retrieval |
0.2 | 1 | 2014 | Learning sufficient queries for entity filtering · SIGIR 2014 |
Information retrieval › information filtering
document filtering |
0.2 | 1 | 2014 | Learning sufficient queries for entity filtering · SIGIR 2014 |
Information retrieval › information filtering
entity filtering |
0.2 | 1 | 2014 | Learning sufficient queries for entity filtering · SIGIR 2014 |
Information retrieval › evaluation › test collection › test collection construction
pooling |
0.2 | 1 | 2014 | On run diversity in Evaluation as a Service · SIGIR 2014 |
Information retrieval › evaluation › test collection
test collection construction |
0.2 | 1 | 2014 | On run diversity in Evaluation as a Service · SIGIR 2014 |
Information retrieval › web search
tweet search |
0.2 | 1 | 2014 | Temporal feedback for tweet search with non-parametric density estimation · SIGIR 2014 |
Information retrieval
cross-language information retrieval |
0.2 | 1 | 2013 | Query representation for cross-temporal information retrieval · SIGIR 2013 |
Information retrieval › document retrieval › domain-specific retrieval
historical text retrieval |
0.2 | 1 | 2013 | Query representation for cross-temporal information retrieval · SIGIR 2013 |
Information retrieval
relevance feedback |
0.2 | 2 | 2013 | Hashtag retrieval in a microblogging environment · SIGIR 2010 Query representation for cross-temporal information retrieval · SIGIR 2013 |
Information retrieval
retrieval models |
0.1 | 2 | 2016 | What Makes a Query Temporally Sensitive? · SIGIR 2016 Model-averaged latent semantic indexing · SIGIR 2007 |
Information retrieval › document retrieval › text search
short text retrieval |
0.1 | 1 | 2012 | Improving retrieval of short texts through document expansion · SIGIR 2012 |
Information retrieval › query reformulation
query expansion |
0.1 | 1 | 2011 | Estimation methods for ranking recent information · SIGIR 2011 |
Information retrieval › retrieval models › language model
query likelihood model |
0.1 | 1 | 2011 | Estimation methods for ranking recent information · SIGIR 2011 |
Information retrieval › retrieval models › ad-hoc retrieval
time-aware retrieval |
0.1 | 2 | 2016 | What Makes a Query Temporally Sensitive? · SIGIR 2016 Improving retrieval of short texts through document expansion · SIGIR 2012 |
Information retrieval › retrieval models › latent semantic models
latent semantic indexing |
0.1 | 1 | 2007 | Model-averaged latent semantic indexing · SIGIR 2007 |
Information retrieval › evaluation › evaluation methodology
nugget-based evaluation |
0.1 | 1 | 2015 | Assessor Differences and User Preferences in Tweet Timeline Generation · SIGIR 2015 |
Web and social media mining
social media analysis |
0.1 | 1 | 2015 | SIGIR 2015 Workshop on Temporal, Social and Spatially-aware Information Access (#TAIA2015) · SIGIR 2015 |
Information retrieval › text summarization
summarization evaluation |
0.1 | 1 | 2015 | Assessor Differences and User Preferences in Tweet Timeline Generation · SIGIR 2015 |
Data mining
density estimation |
0.1 | 1 | 2014 | Temporal feedback for tweet search with non-parametric density estimation · SIGIR 2014 |
Information retrieval › machine learning for information retrieval
query learning |
0.1 | 1 | 2014 | Learning sufficient queries for entity filtering · SIGIR 2014 |
Information retrieval › reranking
search result re-ranking |
0.1 | 1 | 2014 | Temporal feedback for tweet search with non-parametric density estimation · SIGIR 2014 |
Information retrieval
evidence combination |
0.0 | 1 | 2013 | Query representation for cross-temporal information retrieval · SIGIR 2013 |
Information retrieval
web search |
0.0 | 1 | 2013 | SIGIR 2013 workshop on time aware information access (#TAIA2013) · SIGIR 2013 |
Methods — techniques the papers use, named apart from their topics
relevance feedback · 0.4regression analysis · 0.2quantitative analysis · 0.2qualitative analysis · 0.2spatial-temporal retrieval · 0.2information access · 0.2temporal feedback · 0.2kernel density estimation · 0.2deterministic filtering · 0.2boolean query learning · 0.2language modeling · 0.1latent semantic indexing · 0.1kullback-leibler divergence · 0.1AIC · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Document Expansion Using External CollectionsabstractDocument expansion has been shown to improve the effectiveness of information retrieval systems by augmenting documents' term probability estimates with those of similar documents, producing higher quality document representations. We propose a method to further improve document models by utilizing external collections as part of the document expansion process. Our approach is based on relevance modeling, a popular form of pseudo-relevance feedback; however, where relevance modeling is concerned with query expansion, we are concerned with document expansion. Our experiments demonstrate that the proposed model improves ad-hoc document retrieval effectiveness on a variety of corpus types, with a particular benefit on more heterogeneous collections of documents. Garrick Sherman, Miles Efron |
SIGIR | 2 |
| 2016 | What Makes a Query Temporally Sensitive?abstractThis work takes an in-depth look at the factors that affect manual classifications of 'temporally sensitive' information needs. We use qualitative and quantitative techniques to analyze 660 topics from the Text Retrieval Conference (TREC) previously used in the experimental evaluation of temporal retrieval models. Regression analysis is used to identify factors in previous manual classifications. We explore potential problems with the previous classifications, considering principles and guidelines for future work on temporal retrieval models. Craig Willis, Garrick Sherman, Miles Efron |
SIGIR | 3 |
| 2015 | Reproducible Experiments on Lexical and Temporal Feedback for Tweet Search
Jinfeng Rao, Jimmy Lin, Miles Efron |
ECIR | 3 |
| 2015 | SIGIR 2015 Workshop on Temporal, Social and Spatially-aware Information Access (#TAIA2015)abstractIn this workshop we aim to bring together practitioners and researchers to discuss their recent breakthroughs and the challenges with addressing spatial and temporal information access, both from the algorithmic and the architectural perspectives. Klaus Berberich, James Caverlee, Miles Efron, Claudia Hauff, Vanessa Murdock 0001, Milad Shokouhi, Bart Thomee |
SIGIR | 3 |
| 2015 | Assessor Differences and User Preferences in Tweet Timeline GenerationabstractIn information retrieval evaluation, when presented with an effectiveness difference between two systems, there are three relevant questions one might ask. First, are the differences statistically significant? Second, is the comparison stable with respect to assessor differences? Finally, is the difference actually meaningful to a user? This paper tackles the last two questions about assessor differences and user preferences in the context of the newly-introduced tweet timeline generation task in the TREC 2014 Microblog track, where the system's goal is to construct an informative summary of non-redundant tweets that addresses the user's information need. Central to the evaluation methodology is human-generated semantic clusters of tweets that contain substantively similar information. We show that the evaluation is stable with respect to assessor differences in clustering and that user preferences generally correlate with effectiveness metrics even though users are not explicitly aware of the semantic clustering being performed by the systems. Although our analyses are limited to this particular task, we believe that lessons learned could generalize to other evaluations based on establishing semantic equivalence between information units, such as nugget-based evaluations in question answering and temporal summarization. Yulu Wang, Garrick Sherman, Jimmy Lin, Miles Efron |
SIGIR | 4 |
| 2014 | Temporal feedback for tweet search with non-parametric density estimationabstractThis paper investigates the temporal cluster hypothesis: in search tasks where time plays an important role, do relevant documents tend to cluster together in time? We explore this question in the context of tweet search and temporal feedback: starting with an initial set of results from a baseline retrieval model, we estimate the temporal density of relevant documents, which is then used for result reranking. Our contributions lie in a method to characterize this temporal density function using kernel density estimation, with and without human relevance judgments, and an approach to integrating this information into a standard retrieval model. Experiments on TREC datasets confirm that our temporal feedback formulation improves search effectiveness, thus providing support for our hypothesis. Our approach out-performs both a standard baseline and previous temporal retrieval models. Temporal feedback improves over standard lexical feedback (with and without human judgments), illus- trating that temporal relevance signals exist independently of document content. Miles Efron, Jimmy Lin, Jiyin He, Arjen P. de Vries |
SIGIR | 1 |
| 2014 | Learning sufficient queries for entity filteringabstractEntity-centric document filtering is the task of analyzing a time-ordered stream of documents and emitting those that are relevant to a specified set of entities (e.g., people, places, organizations). This task is exemplified by the TREC Knowledge Base Acceleration (KBA) track and has broad applicability in other modern IR settings. In this paper, we present a simple yet effective approach based on learning high-quality Boolean queries that can be applied deterministically during filtering. We call these Boolean statements sufficient queries. We argue that using deterministic queries for entity-centric filtering can reduce confounding factors seen in more familiar "score-then-threshold" filtering methods. Experiments on two standard datasets show significant improvements over state-of-the-art baseline models. Miles Efron, Craig Willis, Garrick Sherman |
SIGIR | 1 |
| 2014 | On run diversity in Evaluation as a Serviceabstract"Evaluation as a service" (EaaS) is a new methodology that enables community-wide evaluations and the construction of test collections on documents that cannot be distributed. The basic idea is that evaluation organizers provide a service API through which the evaluation task can be completed. However, this concept violates some of the premises of traditional pool-based collection building and thus calls into question the quality of the resulting test collection. In particular, the service API might restrict the diversity of runs that contribute to the pool: this might hamper innovation by researchers and lead to incomplete judgment pools that affect the reusability of the collection. This paper shows that the distinctiveness of the retrieval runs used to construct the first test collection built using EaaS, the TREC 2013 Microblog collection, is not substantially different from that of the TREC-8 ad hoc collection, a high-quality collection built using traditional pooling. Further analysis using the `leave out uniques' test suggests that pools from the Microblog 2013 collection are less complete than those from TREC-8, although both collections benefit from the presence of distinctive and effective manual runs. Although we cannot yet generalize to all EaaS implementations, our analyses reveal no obvious flaws in the test collection built using the methodology in the TREC 2013 Microblog track. Ellen M. Voorhees, Jimmy Lin, Miles Efron |
SIGIR | 3 |
| 2013 | SIGIR 2013 workshop on time aware information access (#TAIA2013)abstractWeb content increasingly reflects the current state of the physical and social world, manifested both in traditional news media sources along with user-generated publishing sites such as Twitter, Foursquare, and Facebook. At the same time, web searching increasingly reflects problems grounded in the real world. As a result of this blending of the web with the real world, we observe that the web, both in its composition and use, has incorporated many of the dynamics of the real world. Few of the problems associated with searching dynamic collections are well understood, such as defining time-sensitive relevance, understanding user query behavior over time and understanding why certain web content changes. Fernando Diaz 0001, Susan T. Dumais, Miles Efron, Kira Radinsky, Maarten de Rijke, Milad Shokouhi |
SIGIR | 3 |
| 2013 | Query representation for cross-temporal information retrievalabstractThis paper addresses the problem of long-term language change in information retrieval (IR) systems. IR research has often ignored lexical drift. But in the emerging domain of massive digitized book collections, the risk of vocabulary mismatch due to language change is high. Collections such as Google Books and the Hathi Trust contain text written in the vernaculars of many centuries. With respect to IR, changes in vocabulary and orthography make 14th-Century English qualitatively different from 21st-Century English. This challenges retrieval models that rely on keyword matching. With this challenge in mind, we ask: given a query written in contemporary English, how can we retrieve relevant documents that were written in early English? We argue that search in historically diverse corpora is similar to cross-language retrieval (CLIR). By considering "modern" English and "archaic" English as distinct languages, CLIR techniques can improve what we call cross-temporal IR (CTIR). We focus on ways to combine evidence to improve CTIR effectiveness, proposing and testing several ways to handle language change during book search. We find that a principled combination of three sources of evidence during relevance feedback yields strong CTIR performance. Miles Efron |
SIGIR | 1 |
| 2012 | Improving retrieval of short texts through document expansionabstractCollections containing a large number of short documents are becoming increasingly common. As these collections grow in number and size, providing effective retrieval of brief texts presents a significant research problem. We propose a novel approach to improving information retrieval (IR) for short texts based on aggressive document expansion. Starting from the hypothesis that short documents tend to be about a single topic, we submit documents as pseudo-queries and analyze the results to learn about the documents themselves. Document expansion helps in this context because short documents yield little in the way of term frequency information. However, as we show, the proposed technique helps us model not only lexical properties, but also temporal properties of documents. We present experimental results using a corpus of microblog (Twitter) data and a corpus of metadata records from a federated digital library. With respect to established baselines, results of these experiments show that applying our proposed document expansion method yields significant improvements in effectiveness. Specifically, our method improves the lexical representation of documents and the ability to let time influence retrieval. Miles Efron, Peter Organisciak, Katrina Fenlon |
SIGIR | 1 |
| 2012 | Search User Interfaces. Marti A. Hearst. Cambridge, UK: Cambridge University Press, 2009. 404 pp. $55.00. (ISBN 978-0-521-11379-3) Interactive Information Seeking, Behaviour and Retrieval. Ian Ruthven and Diane Kelly (Eds.). London: Facet Publishing, 2011. 296 pp. $89.95. (ISBN 978-1-85604-707-4)
Miles Efron |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2011 | Estimation methods for ranking recent informationabstractTemporal aspects of documents can impact relevance for certain kinds of queries. In this paper, we build on earlier work of modeling temporal information. We propose an extension to the Query Likelihood Model that incorporates query-specific information to estimate rate parameters, and we introduce a temporal factor into language model smoothing and query expansion using pseudo-relevance feedback. We evaluate these extensions using a Twitter corpus and two newspaper article collections. Results suggest that, compared to prior approaches, our models are more effective at capturing the temporal variability of relevance associated with some topics. Miles Efron, Gene Golovchinsky |
SIGIR | 1 |
| 2011 | Ayşe Göker, John Davies (eds): Information retrieval: searching in the 21st century - John Wiley & Sons, 2009, 320 pp, Price £65.00/€74.80, ISBN: 978-0-470-02762-2
Miles Efron |
Inf. Retr. | 1 |
| 2011 | Information search and retrieval in microblogsabstractModern information retrieval (IR) has come to terms with numerous new media in efforts to help people find information in increasingly diverse settings. Among these new media are so-called microblogs. A microblog is a stream of text that is written by an author over time. It comprises many very brief updates that are presented to the microblog's readers in reverse-chronological order. Today, the service called Twitter is the most popular microblogging platform. Although microblogging is increasingly popular, methods for organizing and providing access to microblog data are still new. This review offers an introduction to the problems that face researchers and developers of IR systems in microblog settings. After an overview of microblogs and the behavior surrounding them, the review describes established problems in microblog retrieval, such as entity search and sentiment analysis, and modeling abstractions, such as authority and quality. The review also treats user-created metadata that often appear in microblogs. Because the problem of microblog search is so new, the review concludes with a discussion of particularly pressing research issues yet to be studied in the field. Miles Efron |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Hashtag retrieval in a microblogging environmentabstractMicroblog services let users broadcast brief textual messages to people who "follow" their activity. Often these posts contain terms called hashtags, markers of a post's meaning, audience, etc. This poster treats the following problem: given a user's stated topical interest, retrieve useful hashtags from microblog posts. Our premise is that a user interested in topic x might like to find hashtags that are often applied to posts about x. This poster proposes a language modeling approach to hashtag retrieval. The main contribution is a novel method of relevance feedback based on hashtags. The approach is tested on a corpus of data harvested from twitter.com. Miles Efron |
SIGIR | 1 |
| 2010 | Linear time series models for term weighting in information retrievalabstractAbstract Common measures of term importance in information retrieval (IR) rely on counts of term frequency; rare terms receive higher weight in document ranking than common terms receive. However, realistic scenarios yield additional information about terms in a collection. Of interest in this article is the temporal behavior of terms as a collection changes over time. We propose capturing each term's collection frequency at discrete time intervals over the lifespan of a corpus and analyzing the resulting time series. We hypothesize the collection frequency of a weakly discriminative term x at time t is predictable by a linear model of the term's prior observations. On the other hand, a linear time series model for a strong discriminators' collection frequency will yield a poor fit to the data. Operationalizing this hypothesis, we induce three time‐based measures of term importance and test these against state‐of‐the‐art term weighting models. Miles Efron |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Query polyrepresentation for ranking retrieval systems without relevance judgmentsabstractAbstract Ranking information retrieval (IR) systems with respect to their effectiveness is a crucial operation during IR evaluation, as well as during data fusion. This article offers a novel method of approaching the system‐ranking problem, based on the widely studied idea of polyrepresentation. The principle of polyrepresentation suggests that a single information need can be represented by many query articulations–what we call query aspects. By skimming the top k (where k is small) documents retrieved by a single system for multiple query aspects, we collect a set of documents that are likely to be relevant to a given test topic. Labeling these skimmed documents as putatively relevant lets us build pseudorelevance judgments without undue human intervention. We report experiments where using these pseudorelevance judgments delivers a rank ordering of IR systems that correlates highly with rankings based on human relevance judgments. Miles Efron, Megan A. Winget |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2009 | Using Multiple Query Aspects to Build Test Collections without Human Relevance Judgments
Miles Efron |
ECIR | 1 |
| 2008 | Query expansion and dimensionality reduction: Notions of optimality in Rocchio relevance feedback and latent semantic indexing
Miles Efron |
Inf. Process. Manag. | 1 |
| 2008 | Shannon Meets Shortz: A Probabilistic Model of Crossword Puzzle DifficultyabstractAbstract This article is concerned with the difficulty of crossword puzzles. A model is proposed that quantifies the difficulty of a Puzzle P with respect to its clues. Given a clue–answer pair (c,a), we model the difficulty of guessing a based on c using the conditional probability P(a|c); easier mappings should enjoy a higher conditional probability. The model is tested by two experiments, each of which involves estimating the difficulty of puzzles taken from The New York Times. Additionally, we discuss how the notion of information implicit in our model relates to more easily quantifiable types of information that figure into crossword puzzles. Miles Efron |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2007 | Model-averaged latent semantic indexingabstractThis poster introduces a novel approach to information retrieval that uses statistical model averaging to improve latent semantic indexing (LSI). Instead of choosing a single dimensionality $k$ for LSI , we propose using several models of differing dimensionality to inform retrieval. To manage this ensemble we weight each model's contribution to an extent inversely proportional to its AIC (Akaike information criterion). Thus each model contributes proportionally to its expected Kullback-Leibler divergence from the distribution that generated the data. We present results on three standard IR test collections, demonstrating significant improvement over both the traditional vector space model and single-model LSI. Miles Efron |
SIGIR | 1 |
| 2006 | Using cocitation information to estimate political orientation in web documents
Miles Efron |
Knowl. Inf. Syst. | 1 |
| 2005 | Eigenvalue-based model selection during latent semantic indexingabstractAbstract In this study amended parallel analysis (APA), a novel method for model selection in unsupervised learning problems such as information retrieval (IR), is described. At issue is the selection ofk, the number of dimensions retained under latent semantic indexing (LSI). Amended parallel analysis is an elaboration of Horn's parallel analysis, which advocates retaining eigenvalues larger than those that we would expect under term independence. Amended parallel analysis operates by deriving confidence intervals on these “null” eigenvalues. The technique amounts to a series of nonparametric hypothesis tests on the correlation matrix eigenvalues. In the study, APA is tested along with four established dimensionality estimators on six standard IR test collections. These estimates are evaluated with regard to two IR performance metrics. Additionally, results from simulated data are reported. In both rounds of experimentation APA performs well, predicting the best values ofkon 3 of 12 observations, with good predictions on several others, and never offering the worst estimate of optimal dimensionality. Miles Efron |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2004 | The liberal media and right-wing conspiracies: using cocitation information to estimate political orientation in web documentsabstractThis paper introduces a simple method for estimating cultural orientation, the affiliation of online entities in a polarized field of discourse. In particular, cocitation information is used to estimate the political orientation of hypertext documents. A type of cultural orientation, the political orientation of a document is the degree to which it participates in traditionally left- or right-wing beliefs. Estimating documents' political orientation is of interest for personalized information retrieval and recommender systems. In its application to politics, the method uses a simple probabilistic model to estimate the strength of association between a document and left- and right-wing communities. The model estimates the likelihood of cocitation between a document of interest and a small number of documents of known orientation. The model is tested on three sets of data, 695 partisan web documents, 162 political weblogs, and 72 non-partisan documents. Accuracy above 90% is obtained from the cocitation model, outperforming lexically based classifiers at statistically significant levels. Miles Efron |
CIKM | 1 |