Eric C. Jensen

dblp:69/6635 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
0since 2021 · last 2008
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 17 · 4 first-authorArtificial intelligence and machine learning · 3 · 1 first-authorSystems, architecture and hardware · 1Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
10 papers
Information retrieval · 100%
Artificial intelligence
1 paper
Learning paradigms · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › query understanding
query classification
0.352007
Automatic classification of Web queries using very large unlabeled query logs · ACM Trans. Inf. Syst. 2007
Varying approaches to topical web query classification · SIGIR 2007
Automatic web query classification using labeled and unlabeled training data · SIGIR 2005
Information retrieval
evaluation
0.252007
Repeatable evaluation of search services in dynamic environments · ACM Trans. Inf. Syst. 2007
Predicting query difficulty on the web by learning visual clues · SIGIR 2005
Using manually-built web directories for automatic evaluation of known-item retrieval · SIGIR 2003
Information retrieval
query log analysis
0.122007
Automatic classification of Web queries using very large unlabeled query logs · ACM Trans. Inf. Syst. 2007
Hourly analysis of a very large topically categorized web query log · SIGIR 2004
Information retrieval
web search
0.132007
Automatic classification of Web queries using very large unlabeled query logs · ACM Trans. Inf. Syst. 2007
Repeatable evaluation of search services in dynamic environments · ACM Trans. Inf. Syst. 2007
Predicting query difficulty on the web by learning visual clues · SIGIR 2005
Information retrieval › evaluation › test collection
test collection construction
0.112007
Repeatable evaluation of search services in dynamic environments · ACM Trans. Inf. Syst. 2007
Information retrieval › distributed information retrieval
metasearch
0.122005
Surrogate scoring for improved metasearch precision · SIGIR 2005
Predicting query difficulty on the web by learning visual clues · SIGIR 2005
Information retrieval › evaluation › query performance prediction
query difficulty estimation
0.112005
Predicting query difficulty on the web by learning visual clues · SIGIR 2005
Information retrieval › distributed information retrieval
result fusion
0.112005
Surrogate scoring for improved metasearch precision · SIGIR 2005
Information retrieval › information filtering
news filtering
0.012004
Evaluation of filtering current news search results · SIGIR 2004
Information retrieval › evaluation › evaluation methodology
automatic evaluation
0.012003
Using manually-built web directories for automatic evaluation of known-item retrieval · SIGIR 2003
Machine learning › Learning paradigms
semi-supervised learning
0.012005
Improving Automatic Query Classification via Semi-Supervised Learning · ICDM 2005
Information retrieval
precision-oriented retrieval
0.012004
Evaluation of filtering current news search results · SIGIR 2004
Information retrieval › query understanding
query disambiguation
0.012004
Hourly analysis of a very large topically categorized web query log · SIGIR 2004
Information retrieval › evaluation
relevance judgment
0.012003
Using manually-built web directories for automatic evaluation of known-item retrieval · SIGIR 2003

Methods — techniques the papers use, named apart from their topics

supervised learning · 0.2selectional preference mining · 0.1query log mining · 0.1computational linguistics · 0.1web taxonomies · 0.1text classification · 0.1pseudo-relevance judgments · 0.1linear model · 0.1ensemble classification · 0.1bootstrap · 0.1
YearPublicationVenuePosition
2008 Using a relational database for scalable XML search
Rebecca Cathey, Steven M. Beitzel, Eric C. Jensen, David A. Grossman, Ophir Frieder
J. Supercomput.3
2007 Relationally Mapping XML Queries For Scalable XML Search
abstract
The growing trend of using XML to share security data requires scalable technology to effectively manage the volume and variety of data. Although a wide variety of methods exist for storing and searching XML, the two most common techniques are conventional tree-based approaches and relational approaches. Tree-based approaches represent XML as a tree and use indexes and path join algorithms to process queries. In contrast, the relational approach seeks to utilize the power of a mature relational database to store and search XML. This method relationally maps XML queries to SQL and reconstructs the XML from the database results. We use the XBench benchmark to compare the scalability of the SQLGenerator, our relational approach, with eXist, a popular tree-based approach.
Rebecca Cathey, Steven M. Beitzel, Eric C. Jensen, David A. Grossman, Ophir Frieder
ISI3
2007 Varying approaches to topical web query classification
abstract
Topical classification of web queries has drawn recent interest because of the promise it offers in improving retrieval effectiveness and efficiency. However, much of this promise depends on whether classification is performed before or after the query is used to retrieve documents. We examine two previously unaddressed issues in query classification: pre versus post-retrieval classification effectiveness and the effect of training explicitly from classified queries versus bridging a classifier trained using a document taxonomy. Bridging classifiers map the categories of a document taxonomy onto those of a query classification problem to provide sufficient training data. We find that training classifiers explicitly from manually classified queries outperforms the bridged classifier by 48% in F1 score. Also, a pre-retrieval classifier using only the query terms performs merely 11% worse than the bridged classifier which requires snippets from retrieved documents.
Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, Ophir Frieder
SIGIR2
2007 Temporal analysis of a very large topically categorized Web query log
abstract
Abstract The authors review a log of billions of Web queries that constituted the total query traffic for a 6‐month period of a general‐purpose commercial Web search service. Previously, query logs were studied from a single, cumulative view. In contrast, this study builds on the authors' previous work, which showed changes in popularity and uniqueness of topically categorized queries across the hours in a day. To further their analysis, they examine query traffic on a daily, weekly, and monthly basis by matching it against lists of queries that have been topically precategorized by human editors. These lists represent 13% of the query traffic. They show that query traffic from particular topical categories differs both from the query stream as a whole and from other categories. Additionally, they show that certain categories of queries trend differently over varying periods. The authors key contribution is twofold: They outline a method for studying both the static and topical properties of a very large query log over varying periods, and they identify and examine topical trends that may provide valuable insight for improving both retrieval effectiveness and efficiency.
Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, Ophir Frieder, David A. Grossman
J. Assoc. Inf. Sci. Technol.2
2007 Exploiting parallelism to support scalable hierarchical clustering
abstract
Abstract A distributed memory parallel version of the group average hierarchical agglomerative clustering algorithm is proposed to enable scaling the document clustering problem to large collections. Using standard message passing operations reduces interprocess communication while maintaining efficient load balancing. In a series of experiments using a subset of a standard Text REtrieval Conference (TREC) test collection, our parallel hierarchical clustering algorithm is shown to be scalable in terms of processors efficiently used and the collection size. Results show that our algorithm performs close to the expectedO(n2/p) time onpprocessors rather than the worst‐caseO(n3/p) time. Furthermore, theO(n2/p) memory complexity per node allows larger collections to be clustered as the number of nodes increases. While partitioning algorithms such ask‐means are trivially parallelizable, our results confirm those of other studies which showed that hierarchical algorithms produce significantly tighter clusters in the document clustering task. Finally, we show how our parallel hierarchical agglomerative clustering algorithm can be used as the clustering subroutine for a parallel version of the buckshot algorithm to cluster the complete TREC collection at near theoretical runtime expectations.
Rebecca Cathey, Eric C. Jensen, Steven M. Beitzel, Ophir Frieder, David A. Grossman
J. Assoc. Inf. Sci. Technol.2
2007 Automatic classification of Web queries using very large unlabeled query logs
abstract
Accurate topical classification of user queries allows for increased effectiveness and efficiency in general-purpose Web search systems. Such classification becomes critical if the system must route queries to a subset of topic-specific and resource-constrained back-end databases. Successful query classification poses a challenging problem, as Web queries are short, thus providing few features. This feature sparseness, coupled with the constantly changing distribution and vocabulary of queries, hinders traditional text classification. We attack this problem by combining multiple classifiers, including exact lookup and partial matching in databases of manually classified frequent queries, linear models trained by supervised learning, and a novel approach based on mining selectional preferences from a large unlabeled query log. Our approach classifies queries without using external sources of information, such as online Web directories or the contents of retrieved pages, making it viable for use in demanding operational environments, such as large-scale Web search services. We evaluate our approach using a large sample of queries from an operational Web search engine and show that our combined method increases recall by nearly 40% over the best single method while maintaining adequate precision. Additionally, we compare our results to those from the 2005 KDD Cup and find that we perform competitively despite our operational restrictions. This suggests it is possible to topically classify a significant portion of the query stream without requiring external sources of information, allowing for deployment in operationally restricted environments.
Steven M. Beitzel, Eric C. Jensen, David D. Lewis, Abdur Chowdhury, Ophir Frieder
ACM Trans. Inf. Syst.2
2007 Repeatable evaluation of search services in dynamic environments
abstract
In dynamic environments, such as the World Wide Web, a changing document collection, query population, and set of search services demands frequent repetition of search effectiveness (relevance) evaluations. Reconstructing static test collections, such as in TREC, requires considerable human effort, as large collection sizes demand judgments deep into retrieved pools. In practice it is common to perform shallow evaluations over small numbers of live engines (often pairwise, engine A vs. engine B) without system pooling. Although these evaluations are not intended to construct reusable test collections, their utility depends on conclusions generalizing to the query population as a whole. We leverage the bootstrap estimate of the reproducibility probability of hypothesis tests in determining the query sample sizes required to ensure this, finding they are much larger than those required for static collections. We propose a semiautomatic evaluation framework to reduce this effort. We validate this framework against a manual evaluation of the top ten results of ten Web search engines across 896 queries in navigational and informational tasks. Augmenting manual judgments with pseudo-relevance judgments mined from Web taxonomies reduces both the chances of missing a correct pairwise conclusion, and those of finding an errant conclusion, by approximately 50%.
Eric C. Jensen, Steven M. Beitzel, Abdur Chowdhury, Ophir Frieder
ACM Trans. Inf. Syst.1
2006 Query Phrase Suggestion from Topically Tagged Session Logs
Eric C. Jensen, Steven M. Beitzel, Abdur Chowdhury, Ophir Frieder
FQAS1
2006 On the development of name search techniques for Arabic
abstract
Abstract The need for effective identity matching systems has led to extensive research in the area of name search. For the most part, such work has been limited to English and other Latin‐based languages. Consequently, algorithms such as Soundex and n‐gram matching are of limited utility for languages such as Arabic, which has vastly different morphologic features that rely heavily on phonetic information. The dearth of work in this field is partly caused by the lack of standardized test data. Consequently, we have built a collection of 7,939 Arabic names, along with 50 training queries and 111 test queries. We use this collection to evaluate a variety of algorithms, including a derivative of Soundex tailored to Arabic (ASOUNDEX), measuring effectiveness by using standard information retrieval measures. Our results show an improvement of 70% over existing approaches.
Syed Uzair Aqeel, Steven M. Beitzel, Eric C. Jensen, David A. Grossman, Ophir Frieder
J. Assoc. Inf. Sci. Technol.3
2005 Improving Automatic Query Classification via Semi-Supervised Learning
abstract
Accurate topical classification of user queries allows for increased effectiveness and efficiency in general-purpose Web search systems. Such classification becomes critical if the system is to return results not just from a general Web collection but from topic-specific back-end databases as well. Maintaining sufficient classification recall is very difficult as Web queries are typically short, yielding few features per query. This feature sparseness coupled with the high query volumes typical for a large-scale search service makes manual and supervised learning approaches alone insufficient. We use an application of computational linguistics to develop an approach for mining the vast amount of unlabeled data in Web query logs to improve automatic topical Web query classification. We show that our approach in combination with manual matching and supervised learning allows us to classify a substantially larger proportion of queries than any single technique. We examine the performance of each approach on a real Web query stream and show that our combined method accurately classifies 46% of queries, outperforming the recall of best single approach by nearly 20%, with a 7% improvement in overall effectiveness.
Steven M. Beitzel, Eric C. Jensen, Ophir Frieder, David D. Lewis, Abdur Chowdhury, Alek Kolcz
ICDM2
2005 Surrogate scoring for improved metasearch precision
abstract
We describe a method for improving the precision of metasearch results based upon scoring the visual features of documents' surrogate representations. These surrogate scores are used during fusion in place of the original scores or ranks provided by the underlying search engines. Visual features are extracted from typical search result surrogate information, such as title, snippet, URL, and rank. This approach specifically avoids the use of search engine-specific scores and collection statistics that are required by most traditional fusion strategies. This restriction correctly reflects the use of metasearch in practice, in which knowledge of the underlying search engines' strategies cannot be assumed. We evaluate our approach using a precision-oriented test collection of manually-constructed binary relevance judgments for the top ten results from ten web search engines over 896 queries. We show that our visual fusion approach significantly outperforms the rCombMNZ fusion algorithm by 5.71%, with 99% confidence, and the best individual web search engine by 10.9%, with 99% confidence.
Steven M. Beitzel, Eric C. Jensen, Ophir Frieder, Abdur Chowdhury, Greg Pass
SIGIR2
2005 Automatic web query classification using labeled and unlabeled training data
abstract
Accurate topical categorization of user queries allows for increased effectiveness, efficiency, and revenue potential in general-purpose web search systems. Such categorization becomes critical if the system is to return results not just from a general web collection but from topic-specific databases as well. Maintaining sufficient categorization recall is very difficult as web queries are typically short, yielding few features per query. We examine three approaches to topical categorization of general web queries: matching against a list of manually labeled queries, supervised learning of classifiers, and mining of selectional preference rules from large unlabeled query logs. Each approach has its advantages in tackling the web query classification recall problem, and combining the three techniques allows us to classify a substantially larger proportion of queries than any of the individual techniques. We examine the performance of each approach on a real web query stream and show that our combined method accurately classifies 46% of queries, outperforming the recall of the best single approach by nearly 20%, with a 7% improvement in overall effectiveness.
Steven M. Beitzel, Eric C. Jensen, Ophir Frieder, David A. Grossman, David D. Lewis, Abdur Chowdhury, Alek Kolcz
SIGIR2
2005 Predicting query difficulty on the web by learning visual clues
abstract
We describe a method for predicting query difficulty in a precision-oriented web search task. Our approach uses visual features from retrieved surrogate document representations (titles, snippets, etc.) to predict retrieval effectiveness for a query. By training a supervised machine learning algorithm with manually evaluated queries, visual clues indicative of relevance are discovered. We show that this approach has a moderate correlation of 0.57 with precision at 10 scores from manual relevance judgments of the top ten documents retrieved by ten web search engines over 896 queries. Our findings indicate that difficulty predictors which have been successful in recall-oriented ad-hoc search, such as clarity metrics, are not nearly as correlated with engine performance in precision-oriented tasks such as this, yielding a maximum correlation of 0.3. Additionally, relying only on visual clues avoids the need for collection statistics that are required by these prior approaches. This enables our approach to be employed in environments where these statistics are unavailable or costly to retrieve, such as metasearch.
Eric C. Jensen, Steven M. Beitzel, David A. Grossman, Ophir Frieder, Abdur Chowdhury
SIGIR1
2004 Hourly analysis of a very large topically categorized web query log
abstract
We review a query log of hundreds of millions of queries that constitute the total query traffic for an entire week of a generalpurpose commercial web search service. Previously, query logs have been studied from a single, cumulative view. In contrast, our analysis shows changes in popularity and uniqueness of topically categorized queries across the hours of the day. We examine query traffic on an hourly basis by matching it against lists of queries that have been topically pre-categorized by human editors. This represents 13 % of the query traffic. We show that query traffic from particular topical categories differs both from the query stream as a whole and from other categories. This analysis provides valuable insight for improving retrieval effectiveness and efficiency. It is also relevant to the development of enhanced query disambiguation, routing, and caching algorithms.
Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, David A. Grossman, Ophir Frieder
SIGIR2
2004 Evaluation of filtering current news search results
abstract
We describe an evaluation of result set filtering techniques for providing ultra-high precision in the task of presenting related news for general web queries. In this task, the negative user experience generated by retrieving non-relevant documents has a much worse impact than not retrieving relevant ones. We adapt cost-based metrics from the document filtering domain to this result filtering problem in order to explicitly examine the tradeoff between missing relevant documents and retrieving non-relevant ones. A large manual evaluation of three simple threshold filters shows that the basic approach of counting matching title terms outperforms also incorporating selected abstract terms based on part-of-speech or higher-level linguistic structures. Simultaneously, leveraging these cost-based metrics allows us to explicitly determine what other tasks would benefit from these alternative techniques.
Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, David A. Grossman, Ophir Frieder
SIGIR2
2004 Fusion of effective retrieval strategies in the same information retrieval system
abstract
Abstract Prior efforts have shown that under certain situations retrieval effectiveness may be improved via the use of data fusion techniques. Although these improvements have been observed from the fusion of result sets from several distinct information retrieval systems, it has often been thought that fusing different document retrieval strategies in a single information retrieval system will lead to similar improvements. In this study, we show that this is not the case. We hold constant systemic differences such as parsing, stemming, phrase processing, and relevance feedback, and fuse result sets generated from highly effective retrieval strategies in the same information retrieval system. From this, we show that data fusion of highly effective retrieval strategies alone shows little or no improvement in retrieval effectiveness. Furthermore, we present a detailed analysis of the performance of modern data fusion approaches, and demonstrate the reasons why they do not perform well when applied to this problem. Detailed results and analyses are included to support our conclusions.
Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, David A. Grossman, Ophir Frieder, Nazli Goharian
J. Assoc. Inf. Sci. Technol.2
2003 Using titles and category names from editor-driven taxonomies for automatic evaluation
abstract
Evaluation of IR systems has always been difficult because of the need for manually assessed relevance judgments. The advent of large editor-driven taxonomies on the web opens the door to a new evaluation approach. We use the ODP (Open Directory Project) taxonomy to find sets of pseudo-relevant documents via one of two assumptions: 1) taxonomy entries are relevant to a given query if their editor-entered titles exactly match the query, or 2) all entries in a leaf-level taxonomy category are relevant to a given query if the category title exactly matches the query. We compare and contrast these two methodologies by evaluating six web search engines on a sample from an America Online log of ten million web queries, using MRR measures for the first method and precision-based measures for the second. We show that this technique is stable with respect to the query set selected and correlated with a reasonably large manual evaluation.
Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, David A. Grossman
CIKM2
2003 Using manually-built web directories for automatic evaluation of known-item retrieval
abstract
Information retrieval system evaluation is complicated by the need for manually assessed relevance judgments. Large manually-built directories on the web open the door to new evaluation procedures. By assuming that web pages are the known relevant items for queries that exactly match their title, we use the ODP (Open Directory Project) and Looksmart directories for system evaluation. We test our approach with a sample from a log of ten million web queries and show that such an evaluation is unbiased in terms of the directory used, stable with respect to the query set selected, and correlated with a reasonably large manual evaluation.
Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, David A. Grossman, Ophir Frieder
SIGIR2
2002 Parallelizing the buckshot algorithm for efficient document clustering
abstract
We present a parallel implementation of the Buckshot document clustering algorithm. We demonstrate that this parallel approach is highly efficient both in terms of load balancing and minimization of communication. In a series of experiments using the 2GB of SGML data from TReC disks 4 and 5, our parallel approach was shown to be scalable in terms of processors efficiently used and the number of clusters created.
Eric C. Jensen, Steven M. Beitzel, Angelo J. Pilotto, Nazli Goharian, Ophir Frieder
CIKM1