EDBT 2026 Demo / reviewers in the wild / expert
Eric C. Jensen
dblp:69/6635
· DBLP profile ↗
19ranked-venue papers
4as first author
0since 2021 · last 2008
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 17 · 4 first-authorArtificial intelligence and machine learning · 3 · 1 first-authorSystems, architecture and hardware · 1Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
10 papers |
Information retrieval · 100% | |
| Artificial intelligence
1 paper |
Learning paradigms · 100% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › query understanding
query classification |
0.3 | 5 | 2007 | Automatic classification of Web queries using very large unlabeled query logs · ACM Trans. Inf. Syst. 2007 Varying approaches to topical web query classification · SIGIR 2007 Automatic web query classification using labeled and unlabeled training data · SIGIR 2005 |
Information retrieval
evaluation |
0.2 | 5 | 2007 | Repeatable evaluation of search services in dynamic environments · ACM Trans. Inf. Syst. 2007 Predicting query difficulty on the web by learning visual clues · SIGIR 2005 Using manually-built web directories for automatic evaluation of known-item retrieval · SIGIR 2003 |
Information retrieval
query log analysis |
0.1 | 2 | 2007 | Automatic classification of Web queries using very large unlabeled query logs · ACM Trans. Inf. Syst. 2007 Hourly analysis of a very large topically categorized web query log · SIGIR 2004 |
Information retrieval
web search |
0.1 | 3 | 2007 | Automatic classification of Web queries using very large unlabeled query logs · ACM Trans. Inf. Syst. 2007 Repeatable evaluation of search services in dynamic environments · ACM Trans. Inf. Syst. 2007 Predicting query difficulty on the web by learning visual clues · SIGIR 2005 |
Information retrieval › evaluation › test collection
test collection construction |
0.1 | 1 | 2007 | Repeatable evaluation of search services in dynamic environments · ACM Trans. Inf. Syst. 2007 |
Information retrieval › distributed information retrieval
metasearch |
0.1 | 2 | 2005 | Surrogate scoring for improved metasearch precision · SIGIR 2005 Predicting query difficulty on the web by learning visual clues · SIGIR 2005 |
Information retrieval › evaluation › query performance prediction
query difficulty estimation |
0.1 | 1 | 2005 | Predicting query difficulty on the web by learning visual clues · SIGIR 2005 |
Information retrieval › distributed information retrieval
result fusion |
0.1 | 1 | 2005 | Surrogate scoring for improved metasearch precision · SIGIR 2005 |
Information retrieval › information filtering
news filtering |
0.0 | 1 | 2004 | Evaluation of filtering current news search results · SIGIR 2004 |
Information retrieval › evaluation › evaluation methodology
automatic evaluation |
0.0 | 1 | 2003 | Using manually-built web directories for automatic evaluation of known-item retrieval · SIGIR 2003 |
Machine learning › Learning paradigms
semi-supervised learning |
0.0 | 1 | 2005 | Improving Automatic Query Classification via Semi-Supervised Learning · ICDM 2005 |
Information retrieval
precision-oriented retrieval |
0.0 | 1 | 2004 | Evaluation of filtering current news search results · SIGIR 2004 |
Information retrieval › query understanding
query disambiguation |
0.0 | 1 | 2004 | Hourly analysis of a very large topically categorized web query log · SIGIR 2004 |
Information retrieval › evaluation
relevance judgment |
0.0 | 1 | 2003 | Using manually-built web directories for automatic evaluation of known-item retrieval · SIGIR 2003 |
Methods — techniques the papers use, named apart from their topics
supervised learning · 0.2selectional preference mining · 0.1query log mining · 0.1computational linguistics · 0.1web taxonomies · 0.1text classification · 0.1pseudo-relevance judgments · 0.1linear model · 0.1ensemble classification · 0.1bootstrap · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2008 | Using a relational database for scalable XML search
Rebecca Cathey, Steven M. Beitzel, Eric C. Jensen, David A. Grossman, Ophir Frieder |
J. Supercomput. | 3 |
| 2007 | Relationally Mapping XML Queries For Scalable XML SearchabstractThe growing trend of using XML to share security data requires scalable technology to effectively manage the volume and variety of data. Although a wide variety of methods exist for storing and searching XML, the two most common techniques are conventional tree-based approaches and relational approaches. Tree-based approaches represent XML as a tree and use indexes and path join algorithms to process queries. In contrast, the relational approach seeks to utilize the power of a mature relational database to store and search XML. This method relationally maps XML queries to SQL and reconstructs the XML from the database results. We use the XBench benchmark to compare the scalability of the SQLGenerator, our relational approach, with eXist, a popular tree-based approach. Rebecca Cathey, Steven M. Beitzel, Eric C. Jensen, David A. Grossman, Ophir Frieder |
ISI | 3 |
| 2007 | Varying approaches to topical web query classificationabstractTopical classification of web queries has drawn recent interest because of the promise it offers in improving retrieval effectiveness and efficiency. However, much of this promise depends on whether classification is performed before or after the query is used to retrieve documents. We examine two previously unaddressed issues in query classification: pre versus post-retrieval classification effectiveness and the effect of training explicitly from classified queries versus bridging a classifier trained using a document taxonomy. Bridging classifiers map the categories of a document taxonomy onto those of a query classification problem to provide sufficient training data. We find that training classifiers explicitly from manually classified queries outperforms the bridged classifier by 48% in F1 score. Also, a pre-retrieval classifier using only the query terms performs merely 11% worse than the bridged classifier which requires snippets from retrieved documents. Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, Ophir Frieder |
SIGIR | 2 |
| 2007 | Temporal analysis of a very large topically categorized Web query logabstractAbstract The authors review a log of billions of Web queries that constituted the total query traffic for a 6‐month period of a general‐purpose commercial Web search service. Previously, query logs were studied from a single, cumulative view. In contrast, this study builds on the authors' previous work, which showed changes in popularity and uniqueness of topically categorized queries across the hours in a day. To further their analysis, they examine query traffic on a daily, weekly, and monthly basis by matching it against lists of queries that have been topically precategorized by human editors. These lists represent 13% of the query traffic. They show that query traffic from particular topical categories differs both from the query stream as a whole and from other categories. Additionally, they show that certain categories of queries trend differently over varying periods. The authors key contribution is twofold: They outline a method for studying both the static and topical properties of a very large query log over varying periods, and they identify and examine topical trends that may provide valuable insight for improving both retrieval effectiveness and efficiency. Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, Ophir Frieder, David A. Grossman |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2007 | Exploiting parallelism to support scalable hierarchical clusteringabstractAbstract A distributed memory parallel version of the group average hierarchical agglomerative clustering algorithm is proposed to enable scaling the document clustering problem to large collections. Using standard message passing operations reduces interprocess communication while maintaining efficient load balancing. In a series of experiments using a subset of a standard Text REtrieval Conference (TREC) test collection, our parallel hierarchical clustering algorithm is shown to be scalable in terms of processors efficiently used and the collection size. Results show that our algorithm performs close to the expectedO(n2/p) time onpprocessors rather than the worst‐caseO(n3/p) time. Furthermore, theO(n2/p) memory complexity per node allows larger collections to be clustered as the number of nodes increases. While partitioning algorithms such ask‐means are trivially parallelizable, our results confirm those of other studies which showed that hierarchical algorithms produce significantly tighter clusters in the document clustering task. Finally, we show how our parallel hierarchical agglomerative clustering algorithm can be used as the clustering subroutine for a parallel version of the buckshot algorithm to cluster the complete TREC collection at near theoretical runtime expectations. Rebecca Cathey, Eric C. Jensen, Steven M. Beitzel, Ophir Frieder, David A. Grossman |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2007 | Automatic classification of Web queries using very large unlabeled query logsabstractAccurate topical classification of user queries allows for increased effectiveness and efficiency in general-purpose Web search systems. Such classification becomes critical if the system must route queries to a subset of topic-specific and resource-constrained back-end databases. Successful query classification poses a challenging problem, as Web queries are short, thus providing few features. This feature sparseness, coupled with the constantly changing distribution and vocabulary of queries, hinders traditional text classification. We attack this problem by combining multiple classifiers, including exact lookup and partial matching in databases of manually classified frequent queries, linear models trained by supervised learning, and a novel approach based on mining selectional preferences from a large unlabeled query log. Our approach classifies queries without using external sources of information, such as online Web directories or the contents of retrieved pages, making it viable for use in demanding operational environments, such as large-scale Web search services. We evaluate our approach using a large sample of queries from an operational Web search engine and show that our combined method increases recall by nearly 40% over the best single method while maintaining adequate precision. Additionally, we compare our results to those from the 2005 KDD Cup and find that we perform competitively despite our operational restrictions. This suggests it is possible to topically classify a significant portion of the query stream without requiring external sources of information, allowing for deployment in operationally restricted environments. Steven M. Beitzel, Eric C. Jensen, David D. Lewis, Abdur Chowdhury, Ophir Frieder |
ACM Trans. Inf. Syst. | 2 |
| 2007 | Repeatable evaluation of search services in dynamic environmentsabstractIn dynamic environments, such as the World Wide Web, a changing document collection, query population, and set of search services demands frequent repetition of search effectiveness (relevance) evaluations. Reconstructing static test collections, such as in TREC, requires considerable human effort, as large collection sizes demand judgments deep into retrieved pools. In practice it is common to perform shallow evaluations over small numbers of live engines (often pairwise, engine A vs. engine B) without system pooling. Although these evaluations are not intended to construct reusable test collections, their utility depends on conclusions generalizing to the query population as a whole. We leverage the bootstrap estimate of the reproducibility probability of hypothesis tests in determining the query sample sizes required to ensure this, finding they are much larger than those required for static collections. We propose a semiautomatic evaluation framework to reduce this effort. We validate this framework against a manual evaluation of the top ten results of ten Web search engines across 896 queries in navigational and informational tasks. Augmenting manual judgments with pseudo-relevance judgments mined from Web taxonomies reduces both the chances of missing a correct pairwise conclusion, and those of finding an errant conclusion, by approximately 50%. Eric C. Jensen, Steven M. Beitzel, Abdur Chowdhury, Ophir Frieder |
ACM Trans. Inf. Syst. | 1 |
| 2006 | Query Phrase Suggestion from Topically Tagged Session Logs
Eric C. Jensen, Steven M. Beitzel, Abdur Chowdhury, Ophir Frieder |
FQAS | 1 |
| 2006 | On the development of name search techniques for ArabicabstractAbstract The need for effective identity matching systems has led to extensive research in the area of name search. For the most part, such work has been limited to English and other Latin‐based languages. Consequently, algorithms such as Soundex and n‐gram matching are of limited utility for languages such as Arabic, which has vastly different morphologic features that rely heavily on phonetic information. The dearth of work in this field is partly caused by the lack of standardized test data. Consequently, we have built a collection of 7,939 Arabic names, along with 50 training queries and 111 test queries. We use this collection to evaluate a variety of algorithms, including a derivative of Soundex tailored to Arabic (ASOUNDEX), measuring effectiveness by using standard information retrieval measures. Our results show an improvement of 70% over existing approaches. Syed Uzair Aqeel, Steven M. Beitzel, Eric C. Jensen, David A. Grossman, Ophir Frieder |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2005 | Improving Automatic Query Classification via Semi-Supervised LearningabstractAccurate topical classification of user queries allows for increased effectiveness and efficiency in general-purpose Web search systems. Such classification becomes critical if the system is to return results not just from a general Web collection but from topic-specific back-end databases as well. Maintaining sufficient classification recall is very difficult as Web queries are typically short, yielding few features per query. This feature sparseness coupled with the high query volumes typical for a large-scale search service makes manual and supervised learning approaches alone insufficient. We use an application of computational linguistics to develop an approach for mining the vast amount of unlabeled data in Web query logs to improve automatic topical Web query classification. We show that our approach in combination with manual matching and supervised learning allows us to classify a substantially larger proportion of queries than any single technique. We examine the performance of each approach on a real Web query stream and show that our combined method accurately classifies 46% of queries, outperforming the recall of best single approach by nearly 20%, with a 7% improvement in overall effectiveness. Steven M. Beitzel, Eric C. Jensen, Ophir Frieder, David D. Lewis, Abdur Chowdhury, Alek Kolcz |
ICDM | 2 |
| 2005 | Surrogate scoring for improved metasearch precisionabstractWe describe a method for improving the precision of metasearch results based upon scoring the visual features of documents' surrogate representations. These surrogate scores are used during fusion in place of the original scores or ranks provided by the underlying search engines. Visual features are extracted from typical search result surrogate information, such as title, snippet, URL, and rank. This approach specifically avoids the use of search engine-specific scores and collection statistics that are required by most traditional fusion strategies. This restriction correctly reflects the use of metasearch in practice, in which knowledge of the underlying search engines' strategies cannot be assumed. We evaluate our approach using a precision-oriented test collection of manually-constructed binary relevance judgments for the top ten results from ten web search engines over 896 queries. We show that our visual fusion approach significantly outperforms the rCombMNZ fusion algorithm by 5.71%, with 99% confidence, and the best individual web search engine by 10.9%, with 99% confidence. Steven M. Beitzel, Eric C. Jensen, Ophir Frieder, Abdur Chowdhury, Greg Pass |
SIGIR | 2 |
| 2005 | Automatic web query classification using labeled and unlabeled training dataabstractAccurate topical categorization of user queries allows for increased effectiveness, efficiency, and revenue potential in general-purpose web search systems. Such categorization becomes critical if the system is to return results not just from a general web collection but from topic-specific databases as well. Maintaining sufficient categorization recall is very difficult as web queries are typically short, yielding few features per query. We examine three approaches to topical categorization of general web queries: matching against a list of manually labeled queries, supervised learning of classifiers, and mining of selectional preference rules from large unlabeled query logs. Each approach has its advantages in tackling the web query classification recall problem, and combining the three techniques allows us to classify a substantially larger proportion of queries than any of the individual techniques. We examine the performance of each approach on a real web query stream and show that our combined method accurately classifies 46% of queries, outperforming the recall of the best single approach by nearly 20%, with a 7% improvement in overall effectiveness. Steven M. Beitzel, Eric C. Jensen, Ophir Frieder, David A. Grossman, David D. Lewis, Abdur Chowdhury, Alek Kolcz |
SIGIR | 2 |
| 2005 | Predicting query difficulty on the web by learning visual cluesabstractWe describe a method for predicting query difficulty in a precision-oriented web search task. Our approach uses visual features from retrieved surrogate document representations (titles, snippets, etc.) to predict retrieval effectiveness for a query. By training a supervised machine learning algorithm with manually evaluated queries, visual clues indicative of relevance are discovered. We show that this approach has a moderate correlation of 0.57 with precision at 10 scores from manual relevance judgments of the top ten documents retrieved by ten web search engines over 896 queries. Our findings indicate that difficulty predictors which have been successful in recall-oriented ad-hoc search, such as clarity metrics, are not nearly as correlated with engine performance in precision-oriented tasks such as this, yielding a maximum correlation of 0.3. Additionally, relying only on visual clues avoids the need for collection statistics that are required by these prior approaches. This enables our approach to be employed in environments where these statistics are unavailable or costly to retrieve, such as metasearch. Eric C. Jensen, Steven M. Beitzel, David A. Grossman, Ophir Frieder, Abdur Chowdhury |
SIGIR | 1 |
| 2004 | Hourly analysis of a very large topically categorized web query logabstractWe review a query log of hundreds of millions of queries that constitute the total query traffic for an entire week of a generalpurpose commercial web search service. Previously, query logs have been studied from a single, cumulative view. In contrast, our analysis shows changes in popularity and uniqueness of topically categorized queries across the hours of the day. We examine query traffic on an hourly basis by matching it against lists of queries that have been topically pre-categorized by human editors. This represents 13 % of the query traffic. We show that query traffic from particular topical categories differs both from the query stream as a whole and from other categories. This analysis provides valuable insight for improving retrieval effectiveness and efficiency. It is also relevant to the development of enhanced query disambiguation, routing, and caching algorithms. Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, David A. Grossman, Ophir Frieder |
SIGIR | 2 |
| 2004 | Evaluation of filtering current news search resultsabstractWe describe an evaluation of result set filtering techniques for providing ultra-high precision in the task of presenting related news for general web queries. In this task, the negative user experience generated by retrieving non-relevant documents has a much worse impact than not retrieving relevant ones. We adapt cost-based metrics from the document filtering domain to this result filtering problem in order to explicitly examine the tradeoff between missing relevant documents and retrieving non-relevant ones. A large manual evaluation of three simple threshold filters shows that the basic approach of counting matching title terms outperforms also incorporating selected abstract terms based on part-of-speech or higher-level linguistic structures. Simultaneously, leveraging these cost-based metrics allows us to explicitly determine what other tasks would benefit from these alternative techniques. Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, David A. Grossman, Ophir Frieder |
SIGIR | 2 |
| 2004 | Fusion of effective retrieval strategies in the same information retrieval systemabstractAbstract Prior efforts have shown that under certain situations retrieval effectiveness may be improved via the use of data fusion techniques. Although these improvements have been observed from the fusion of result sets from several distinct information retrieval systems, it has often been thought that fusing different document retrieval strategies in a single information retrieval system will lead to similar improvements. In this study, we show that this is not the case. We hold constant systemic differences such as parsing, stemming, phrase processing, and relevance feedback, and fuse result sets generated from highly effective retrieval strategies in the same information retrieval system. From this, we show that data fusion of highly effective retrieval strategies alone shows little or no improvement in retrieval effectiveness. Furthermore, we present a detailed analysis of the performance of modern data fusion approaches, and demonstrate the reasons why they do not perform well when applied to this problem. Detailed results and analyses are included to support our conclusions. Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, David A. Grossman, Ophir Frieder, Nazli Goharian |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2003 | Using titles and category names from editor-driven taxonomies for automatic evaluationabstractEvaluation of IR systems has always been difficult because of the need for manually assessed relevance judgments. The advent of large editor-driven taxonomies on the web opens the door to a new evaluation approach. We use the ODP (Open Directory Project) taxonomy to find sets of pseudo-relevant documents via one of two assumptions: 1) taxonomy entries are relevant to a given query if their editor-entered titles exactly match the query, or 2) all entries in a leaf-level taxonomy category are relevant to a given query if the category title exactly matches the query. We compare and contrast these two methodologies by evaluating six web search engines on a sample from an America Online log of ten million web queries, using MRR measures for the first method and precision-based measures for the second. We show that this technique is stable with respect to the query set selected and correlated with a reasonably large manual evaluation. Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, David A. Grossman |
CIKM | 2 |
| 2003 | Using manually-built web directories for automatic evaluation of known-item retrievalabstractInformation retrieval system evaluation is complicated by the need for manually assessed relevance judgments. Large manually-built directories on the web open the door to new evaluation procedures. By assuming that web pages are the known relevant items for queries that exactly match their title, we use the ODP (Open Directory Project) and Looksmart directories for system evaluation. We test our approach with a sample from a log of ten million web queries and show that such an evaluation is unbiased in terms of the directory used, stable with respect to the query set selected, and correlated with a reasonably large manual evaluation. Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, David A. Grossman, Ophir Frieder |
SIGIR | 2 |
| 2002 | Parallelizing the buckshot algorithm for efficient document clusteringabstractWe present a parallel implementation of the Buckshot document clustering algorithm. We demonstrate that this parallel approach is highly efficient both in terms of load balancing and minimization of communication. In a series of experiments using the 2GB of SGML data from TReC disks 4 and 5, our parallel approach was shown to be scalable in terms of processors efficiently used and the number of clusters created. Eric C. Jensen, Steven M. Beitzel, Angelo J. Pilotto, Nazli Goharian, Ophir Frieder |
CIKM | 1 |