VLDB 2026 Research / reviewers in the wild / expert
Oren Etzioni
dblp:e/OEtzioni
· DBLP profile ↗
110ranked-venue papers
27as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 84 · 23 first-author · 1 since 2021Databases, data management, data science and information retrieval · 26 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-authorTheory of computation · 7 · 5 first-authorHuman-computer interaction and ubiquitous computing · 5Computer networks · 4Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
61 papers |
Information extraction and text analysis · 41% Question answering and dialogue systems · 20% Knowledge representation and reasoning · 17% | |
| Databases, data mining, and information retrieval
22 papers |
Data mining · 50% Knowledge graphs · 27% Information retrieval · 12% | |
| Human-computer interaction and pervasive computing
5 papers |
Collaborative and social computing · 39% Games and playful interaction · 24% Interaction techniques and input · 23% |
Topics — the 30 heaviest of 117, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
multimodal communication |
0.5 | 1 | 2021 | Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and Text · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
open information extraction |
0.5 | 4 | 2012 | Open Language Learning for Information Extraction · EMNLP-CoNLL 2012 Open Information Extraction: The Second Generation · IJCAI 2011 Identifying Relations for Open Information Extraction · EMNLP 2011 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.4 | 2 | 2014 | Open question answering over curated and extracted knowledge bases · KDD 2014 Paraphrase-Driven Learning for Open Question Answering · ACL (1) 2013 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.3 | 3 | 2011 | Identifying Relations for Open Information Extraction · EMNLP 2011 Identifying Functional Relations in Web Text · EMNLP 2010 The Tradeoffs Between Open and Traditional Relation Extraction · ACL 2008 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.3 | 3 | 2011 | Named Entity Recognition in Tweets: An Experimental Study · EMNLP 2011 Locating Complex Named Entities in Web Text · IJCAI 2007 Unsupervised named-entity extraction from the Web: An experimental study · Artif. Intell. 2005 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › automated reasoning
knowledge base reasoning |
0.2 | 1 | 2016 | Combining Retrieval, Statistics, and Inference to Answer Elementary Science Questions · AAAI 2016 |
Collaborative and social computing
online communities |
0.2 | 1 | 2016 | Toward Automatic Bootstrapping of Online Communities Using Decision-theoretic Optimization · CSCW 2016 |
Data mining
text mining |
0.2 | 2 | 2014 | The battle for the future of data mining · KDD 2014 To buy or not to buy: that is the question · KDD 2013 |
Computer vision › 3D vision
geometric reasoning |
0.2 | 1 | 2015 | Solving Geometry Problems: Combining Text and Diagram Interpretation · EMNLP 2015 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
geometry problem solving |
0.2 | 1 | 2015 | Solving Geometry Problems: Combining Text and Diagram Interpretation · EMNLP 2015 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › probabilistic reasoning › probabilistic logic
markov logic networks |
0.2 | 1 | 2015 | Exploring Markov Logic Networks for Question Answering · EMNLP 2015 |
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
science question answering |
0.2 | 1 | 2015 | Exploring Markov Logic Networks for Question Answering · EMNLP 2015 |
Knowledge graphs
knowledge graph construction |
0.2 | 2 | 2014 | The battle for the future of data mining · KDD 2014 Machine reading at web scale · WSDM 2008 |
Data mining
predictive modeling |
0.2 | 2 | 2013 | To buy or not to buy: that is the question · KDD 2013 To buy or not to buy: mining airfare data to minimize ticket purchase price · KDD 2003 |
Data mining › predictive modeling › forecasting
price prediction |
0.2 | 2 | 2013 | To buy or not to buy: that is the question · KDD 2013 To buy or not to buy: mining airfare data to minimize ticket purchase price · KDD 2003 |
Data mining › text mining › sentiment analysis
review mining |
0.2 | 2 | 2013 | RevMiner: an extractive interface for navigating reviews on a smartphone · UIST 2012 To buy or not to buy: that is the question · KDD 2013 |
Computer vision › Vision and language › multimodal understanding
diagram understanding |
0.2 | 1 | 2014 | Diagram Understanding in Geometry Questions · AAAI 2014 |
Natural language and speech › Question answering and dialogue systems
knowledge base question answering |
0.2 | 1 | 2014 | Open question answering over curated and extracted knowledge bases · KDD 2014 |
Natural language and speech › Question answering and dialogue systems
math word problem solving |
0.2 | 1 | 2014 | Learning to Solve Arithmetic Word Problems with Verb Categorization · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis › lexical semantics
verb classification |
0.2 | 1 | 2014 | Learning to Solve Arithmetic Word Problems with Verb Categorization · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis › event analysis › event understanding
event schema induction |
0.2 | 1 | 2013 | Generating Coherent Event Schemas at Scale · EMNLP 2013 |
Natural language and speech › Information extraction and text analysis › relation extraction
joint entity and relation extraction |
0.2 | 1 | 2013 | Generating Coherent Event Schemas at Scale · EMNLP 2013 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.1 | 2 | 2008 | Machine reading at web scale · WSDM 2008 Machine Reading · AAAI 2006 |
Natural language and speech › Information extraction and text analysis
entity linking |
0.1 | 1 | 2012 | No Noun Phrase Left Behind: Detecting and Typing Unlinkable Entities · EMNLP-CoNLL 2012 |
Natural language and speech › Information extraction and text analysis › event extraction
event classification |
0.1 | 1 | 2012 | Open domain event extraction from twitter · KDD 2012 |
Natural language and speech › Information extraction and text analysis
event extraction |
0.1 | 1 | 2012 | Open domain event extraction from twitter · KDD 2012 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.1 | 1 | 2012 | Open domain event extraction from twitter · KDD 2012 |
Natural language and speech › Information extraction and text analysis › event extraction
open event extraction |
0.1 | 1 | 2012 | Open domain event extraction from twitter · KDD 2012 |
Knowledge graphs › knowledge graph construction › knowledge extraction
entity typing |
0.1 | 1 | 2012 | No Noun Phrase Left Behind: Detecting and Typing Unlinkable Entities · EMNLP-CoNLL 2012 |
Interaction techniques and input
mobile interaction |
0.1 | 1 | 2012 | RevMiner: an extractive interface for navigating reviews on a smartphone · UIST 2012 |
Methods — techniques the papers use, named apart from their topics
case study · 0.6simulation · 0.5decision-theoretic optimization · 0.5markov logic networks · 0.5submodular optimization · 0.4text mining · 0.4probabilistic inference · 0.3unsupervised learning · 0.3natural language processing · 0.3integer programming · 0.2information retrieval · 0.2corpus statistics · 0.2semantic parsing · 0.2first-order logic · 0.2dependency parsing · 0.2knowledge base integration · 0.2deep learning · 0.2data mining · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and TextabstractChristopher Clark, Jordi Salvador, Dustin Schwenk, Derrick Bonafilia, Mark Yatskar, Eric Kolve, Alvaro Herrasti, Jonghyun Choi, Sachin Mehta, Sam Skjonsberg, Carissa Schoenick, Aaron Sarnat, Hannaneh Hajishirzi, Aniruddha Kembhavi, Oren Etzioni, Ali Farhadi. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Jordi Salvador, Dustin Schwenk, Derrick Bonafilia, Mark Yatskar, Eric Kolve, Alvaro Herrasti, Sachin Mehta, Sam Skjonsberg, Carissa Schoenick, Aaron Sarnat, Hannaneh Hajishirzi, Aniruddha Kembhavi, Oren Etzioni, Ali Farhadi |
EMNLP (1) | 15 |
| 2016 | Combining Retrieval, Statistics, and Inference to Answer Elementary Science QuestionsabstractWhat capabilities are required for an AI system to pass standard 4th Grade Science Tests? Previous work has examined the use of Markov Logic Networks (MLNs) to represent the requisite background knowledge and interpret test questions, but did not improve upon an information retrieval (IR) baseline. In this paper, we describe an alternative approach that operates at three levels of representation and reasoning: information retrieval, corpus statistics, and simple inference over a semi-automatically constructed knowledge base, to achieve substantially improved results. We evaluate the methods on six years of unseen, unedited exam questions from the NY Regents Science Exam (using only non-diagram, multiple choice questions), and show that our overall system’s score is 71.3%, an improvement of 23.8% (absolute) over the MLN-based method described in previous work. We conclude with a detailed analysis, illustrating the complementary strengths of each method in the ensemble. Our datasets are being released to enable further research. Peter Clark, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter D. Turney, Daniel Khashabi |
AAAI | 2 |
| 2016 | Toward Automatic Bootstrapping of Online Communities Using Decision-theoretic OptimizationabstractSuccessful online communities (e.g., Wikipedia, Yelp, and StackOverflow) can produce valuable content. However, many communities fail in their initial stages. Starting an online community is challenging because there is not enough content to attract a critical mass of active members. This paper examines methods for addressing this cold-start problem in datamining-bootstrappable communities by attracting non-members to contribute to the community. We make four contributions: 1) we characterize a set of communities that are “datamining-bootstrappable” and define the bootstrapping problem in terms of decision-theoretic optimization, 2) we estimate the model parameters in a case study involving the Open AI Resources website, 3) we demonstrate that non-members' predicted interest levels and request design are important features that can significantly affect the contribution rate, and 4) we ran a simulation experiment using data generated with the learned parameters and show that our decision-theoretic optimization algorithm can generate as much community utility when bootstrapping the community as our strongest baseline while issuing only 55% as many contribution requests. Shih-Wen Huang, Jonathan Bragg, Isaac Cowhey, Oren Etzioni, Daniel S. Weld |
CSCW | 4 |
| 2016 | Question Answering via Integer Programming over Semi-Structured Knowledge
Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Peter Clark, Oren Etzioni, Dan Roth 0001 |
IJCAI | 5 |
| 2015 | Exploring Markov Logic Networks for Question AnsweringabstractElementary-level science exams pose sig-nificant knowledge acquisition and rea-soning challenges for automatic question answering. We develop a system that rea-sons with knowledge derived from text-books, represented in a subset of first-order logic. Automatic extraction, while scalable, often results in knowledge that is incomplete and noisy, motivating use of reasoning mechanisms that handle uncer-tainty. Markov Logic Networks (MLNs) seem a natural model for expressing such knowl-edge, but the exact way of leveraging MLNs is by no means obvious. We in-vestigate three ways of applying MLNs to our task. First, we simply use the extracted science rules directly as MLN clauses and exploit the structure present in hard con-straints to improve tractability. Second, we interpret science rules as describing prototypical entities, resulting in a drasti-cally simplified but brittle network. Our third approach, called Praline, uses MLNs to align lexical elements as well as define and control how inference should be per-formed in this task. Praline demonstrates a 15 % accuracy boost and a 10x reduction in runtime as compared to other MLN-based methods, and comparable accuracy to word-based baseline approaches. Tushar Khot, Niranjan Balasubramanian, Eric Gribkoff, Ashish Sabharwal, Peter Clark, Oren Etzioni |
EMNLP | 6 |
| 2015 | Solving Geometry Problems: Combining Text and Diagram InterpretationabstractThis paper introduces GEOS, the first automated system to solve unaltered SAT geometry questions by combining text understanding and diagram interpretation.We model the problem of understanding geometry questions as submodular optimization, and identify a formal problem description likely to be compatible with both the question text and diagram.GEOS then feeds the description to a geometric solver that attempts to determine the correct answer.In our experiments, GEOS achieves a 49% score on official SAT questions, and a score of 61% on practice questions. 1 Finally, we show that by integrating textual and visual information, GEOS boosts the accuracy of dependency and semantic parsing of the question text. Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi, Oren Etzioni, Clint Malcolm |
EMNLP | 4 |
| 2015 | The elephant in the room: getting value from Big DataabstractInternational audience Serge Abiteboul, Xin Dong 0001, Oren Etzioni, Divesh Srivastava, Gerhard Weikum, Julia Stoyanovich, Fabian M. Suchanek |
WebDB | 3 |
| 2015 | Parsing Algebraic Word Problems into EquationsabstractThis paper formalizes the problem of solving multi-sentence algebraic word problems as that of generating and scoring equation trees. We use integer linear programming to generate equation trees and score their likelihood by learning local and global discriminative models. These models are trained on a small set of word problems and their answers, without any manual annotation, in order to choose the equation that best matches the problem text. We refer to the overall system as Alges. We compare Alges with previous work and show that it covers the full gamut of arithmetic operations whereas Hosseini et al. (2014) only handle addition and subtraction. In addition, Alges overcomes the brittleness of the Kushman et al. (2014) approach on single-equation problems, yielding a 15% to 50% reduction in error. Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, Siena Dumas Ang |
Trans. Assoc. Comput. Linguistics | 4 |
| 2014 | Diagram Understanding in Geometry QuestionsabstractAutomatically solving geometry questions is a long-standing AI problem. A geometry question typically includes a textual description accompanied by a diagram. The first step in solving geometry questions is diagram understanding, which consists of identifying visual elements in the diagram, their locations, their geometric properties, and aligning them to corresponding textual descriptions. In this paper, we present a method for diagram understanding that identifies visual elements in a diagram while maximizing agreement between textual and visual data. We show that the method's objective function is submodular; thus we are able to introduce an efficient method for diagram understanding that is close to optimal. To empirically evaluate our method, we compile a new dataset of geometry questions (textual descriptions and diagrams) and compare with baselines that utilize standard vision techniques. Our experimental evaluation shows an F1 boost of more than 17% in identifying visual elements and 25% in aligning visual elements with their textual descriptions. Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi, Oren Etzioni |
AAAI | 4 |
| 2014 | Chinese Open Relation Extraction for Knowledge AcquisitionabstractYuen-Hsien Tseng, Lung-Hao Lee, Shu-Yen Lin, Bo-Shun Liao, Mei-Jun Liu, Hsin-Hsi Chen, Oren Etzioni, Anthony Fader. Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, volume 2: Short Papers. 2014. Yuen-Hsien Tseng, Lung-Hao Lee, Shu-Yen Lin, Bo-Shun Liao, Meijun Liu, Hsin-Hsi Chen, Oren Etzioni, Anthony Fader |
EACL | 7 |
| 2014 | Learning to Solve Arithmetic Word Problems with Verb CategorizationabstractThis paper presents a novel approach to learning to solve simple arithmetic word problems.Our system, ARIS, analyzes each of the sentences in the problem statement to identify the relevant variables and their values.ARIS then maps this information into an equation that represents the problem, and enables its (trivial) solution as shown in Figure 1.The paper analyzes the arithmetic-word problems "genre", identifying seven categories of verbs used in such problems.ARIS learns to categorize verbs with 81.2% accuracy, and is able to solve 77.7% of the problems in a corpus of standard primary school test questions.We report the first learning results on this task without reliance on predefined templates and make our data publicly available. Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, Nate Kushman |
EMNLP | 3 |
| 2014 | The battle for the future of data miningabstractDeep learning has catapulted to the front page of the New York Times, formed the core of the so-called 'Google brain', and achieved impressive results in vision, speech recognition, and elsewhere. Yet researchers have offered simple conundrums that deep learning doesn't address. For example, consider the sentence: 'The large ball crashed right through the table because it was made of Styrofoam.' What was made of Styrofoam? The large ball? Or the table? The answer is obviously 'the table', but if we change the word 'Styrofoam' to 'steel', the answer is clearly 'the large ball'. To automatically answer this type of question, our computers require an extensive body of knowledge. We believe that text mining can provide the requisite body of knowledge. My talk will describe work at the new Allen Institute for AI towards building the next-generation of text-mining systems. Oren Etzioni |
KDD | 1 |
| 2014 | Open question answering over curated and extracted knowledge basesabstractWe consider the problem of open-domain question answering (Open QA) over massive knowledge bases (KBs). Existing approaches use either manually curated KBs like Freebase or KBs automatically extracted from unstructured text. In this paper, we present OQA, the first approach to leverage both curated and extracted KBs. Anthony Fader, Luke Zettlemoyer, Oren Etzioni |
KDD | 3 |
| 2013 | Paraphrase-Driven Learning for Open Question Answering
Anthony Fader, Luke Zettlemoyer, Oren Etzioni |
ACL (1) | 3 |
| 2013 | Generating Coherent Event Schemas at ScaleabstractChambers and Jurafsky (2009) demonstrated that event schemas can be automatically induced from text corpora.However, our analysis of their schemas identifies several weaknesses, e.g., some schemas lack a common topic and distinct roles are incorrectly mixed into a single actor.It is due in part to their pair-wise representation that treats subjectverb independently from verb-object.This often leads to subject-verb-object triples that are not meaningful in the real-world.We present a novel approach to inducing open-domain event schemas that overcomes these limitations.Our approach uses cooccurrence statistics of semantically typed relational triples, which we call Rel-grams (relational n-grams).In a human evaluation, our schemas outperform Chambers's schemas by wide margins on several evaluation criteria.Both Rel-grams and event schemas are freely available to the research community. Niranjan Balasubramanian, Stephen Soderland, Mausam, Oren Etzioni |
EMNLP | 4 |
| 2013 | To buy or not to buy: that is the questionabstractShopping can be decomposed into three basic questions: what, where, and when to buy? In this talk, I'll describe how we utilize advanced data-mining and text-mining techniques at Decide.com (and earlier at Farecast) to solve these problems for on-line shoppers. Our algorithms have predicted prices utilizing billions of data points, and ranked products based on millions of reviews. Oren Etzioni |
KDD | 1 |
| 2013 | Towards Coherent Multi-Document Summarization
Janara Christensen, Mausam, Stephen Soderland, Oren Etzioni |
HLT-NAACL | 4 |
| 2013 | Modeling Missing Data in Distant Supervision for Information ExtractionabstractDistant supervision algorithms learn information extraction models given only large readily available databases and text collections. Most previous work has used heuristics for generating labeled data, for example assuming that facts not contained in the database are not mentioned in the text, and facts in the database must be mentioned at least once. In this paper, we propose a new latent-variable approach that models missing data. This provides a natural way to incorporate side information, for instance modeling the intuition that text will often mention rare entities which are likely to be missing in the database. Despite the added complexity introduced by reasoning about missing data, we demonstrate that a carefully designed local search approach to inference is very accurate and scales to large datasets. Experiments demonstrate improved performance for binary and unary relation extraction when compared to learning with heuristic labels, including on average a 27% increase in area under the precision recall curve in the binary case. Alan Ritter, Luke Zettlemoyer, Mausam, Oren Etzioni |
Trans. Assoc. Comput. Linguistics | 4 |
| 2012 | No Noun Phrase Left Behind: Detecting and Typing Unlinkable Entities
Thomas Lin, Mausam, Oren Etzioni |
EMNLP-CoNLL | 3 |
| 2012 | Open Language Learning for Information Extraction
Mausam, Michael Schmitz 0002, Stephen Soderland, Robert Bart, Oren Etzioni |
EMNLP-CoNLL | 5 |
| 2012 | Open domain event extraction from twitterabstractTweets are the most up-to-date and inclusive stream of in- formation and commentary on current events, but they are also fragmented and noisy, motivating the need for systems that can extract, aggregate and categorize important events. Previous work on extracting structured representations of events has focused largely on newswire text; Twitter's unique characteristics present new challenges and opportunities for open-domain event extraction. This paper describes TwiCal-- the first open-domain event-extraction and categorization system for Twitter. We demonstrate that accurately extracting an open-domain calendar of significant events from Twitter is indeed feasible. In addition, we present a novel approach for discovering important event categories and classifying extracted events based on latent variable models. By leveraging large volumes of unlabeled data, our approach achieves a 14% increase in maximum F1 over a supervised baseline. A continuously updating demonstration of our system can be viewed at http://statuscalendar.com; Our NLP tools are available at http://github.com/aritter/ twitter_nlp. Alan Ritter, Mausam, Oren Etzioni, Sam Clark |
KDD | 3 |
| 2012 | RevMiner: an extractive interface for navigating reviews on a smartphoneabstractSmartphones are convenient, but their small screens make searching, clicking, and reading awkward. Thus, perusing product reviews on a smartphone is difficult. In response, we introduce RevMiner - a novel smartphone interface that utilizes Natural Language Processing techniques to analyze and navigate reviews. RevMiner was run over 300K Yelp restaurant reviews extracting attribute-value pairs, where attributes represent restaurant attributes such as sushi and service, and values represent opinions about the attributes such as fresh or fast. These pairs were aggregated and used to: 1) answer queries such as "cheap Indian food", 2) concisely present information about each restaurant, and 3) identify similar restaurants. Our user studies demonstrate that on a smartphone, participants preferred RevMiner's interface to tag clouds and color bars, and that they preferred RevMiner's results to Yelp's, particularly for conjunctive queries (e.g., "great food and huge portions"). Demonstrations of RevMiner are available at revminer.com. Jeff Huang 0002, Oren Etzioni, Luke Zettlemoyer, Kevin Clark, Christian Lee |
UIST | 2 |
| 2011 | Identifying Relations for Open Information Extraction
Anthony Fader, Stephen Soderland, Oren Etzioni |
EMNLP | 3 |
| 2011 | Named Entity Recognition in Tweets: An Experimental Study
Alan Ritter, Sam Clark, Mausam, Oren Etzioni |
EMNLP | 4 |
| 2011 | Open Information Extraction: The Second GenerationabstractHow do we scale information extraction to the massive size and unprecedented heterogeneity of the Web corpus? Beginning in 2003, our KnowItAll project has sought to extract high-quality knowledge from the Web. In 2007, we introduced the Open Information Extraction (Open IE) paradigm which eschews handlabeled training examples, and avoids domainspecific verbs and nouns, to develop unlexicalized, domain-independent extractors that scale to the Web corpus. Open IE systems have extracted billions of assertions as the basis for both commonsense knowledge and novel question-answering systems. This paper describes the second generation of Open IE systems, which rely on a novel model of how relations and their arguments are expressed in English sentences to double precision/recall compared with previous systems such as TEXTRUNNER and WOE. 1 Oren Etzioni, Anthony Fader, Janara Christensen, Stephen Soderland, Mausam |
IJCAI | 1 |
| 2011 | An analysis of open information extraction based on semantic role labelingabstractOpen Information Extraction extracts relations from text without requiring a pre-specified domain or vocabulary. While existing techniques have used only shallow syntactic features, we investigate the use of semantic role labeling techniques for the task of Open IE. Semantic role labeling (SRL) and Open IE, although developed mostly in isolation, are quite related. We compare SRL-based open extractors, which perform computationally expensive, deep syntactic analysis, with TextRunner, an open extractor, which uses shallow syntactic analysis but is able to analyze many more sentences in a fixed amount of time and thus exploit corpus-level statistics. Our evaluation answers questions regarding these systems, including, can SRL extractors, which are trained on PropBank, cope with heterogeneous text found on the Web? Which extractor attains better precision, recall, f-measure, or running time? How does extractor performance vary for binary, n-ary and nested relations? How much do we gain by running multiple extractors? How do we select the optimal extractor given amount of data, available time, types of extractions desired? Janara Christensen, Mausam, Stephen Soderland, Oren Etzioni |
K-CAP | 4 |
| 2010 | Panlingual Lexical Translation via Probabilistic InferenceabstractThe bare minimum lexical resource required to translate between a pair of languages is a translation dictionary. Unfortunately, dictionaries exist only between a tiny fraction of the 49 million possible language-pairs making machine translation virtually impossible between most of the languages. This paper summarizes the last four years of our research motivated by the vision of panlingual communication. Our research comprises three key steps. First, we compile over 630 freely available dictionaries over the Web and convert this data into a single representation – the translation graph. Second, we build several inference algorithms that infer translations between word pairs even when no dictionary lists them as translations. Finally, we run our inference procedure offline to construct PANDICTIONARY– a sense-distinguished, massively multilingual dictionary that has translations in more than 1000 languages. Our experiments assess the quality of this dictionary and find that we have 4 times as many translations at a high precision of 0.9 compared to the English Wiktionary, which is the lexical resource closest to PANDICTIONARY. Mausam, Stephen Soderland, Oren Etzioni |
AAAI | 3 |
| 2010 | A Latent Dirichlet Allocation Method for Selectional Preferences
Alan Ritter, Mausam, Oren Etzioni |
ACL | 3 |
| 2010 | Identifying Functional Relations in Web Text
Thomas Lin, Mausam, Oren Etzioni |
EMNLP | 3 |
| 2010 | Learning First-Order Horn Clauses from Web Text
Stefan Schoenmackers, Jesse Davis, Oren Etzioni, Daniel S. Weld |
EMNLP | 3 |
| 2010 | Analysis of a probabilistic model of redundancy in unsupervised information extraction
Doug Downey, Oren Etzioni, Stephen Soderland |
Artif. Intell. | 2 |
| 2010 | Panlingual lexical translation via probabilistic inference
Mausam, Stephen Soderland, Oren Etzioni, Daniel S. Weld, Kobi Reiter, Michael Skinner, Marcus Sammer, Jeff A. Bilmes |
Artif. Intell. | 3 |
| 2009 | Compiling a Massive, Multilingual Dictionary via Probabilistic Inference
Mausam, Stephen Soderland, Oren Etzioni, Daniel S. Weld, Michael Skinner, Jeff A. Bilmes |
ACL/IJCNLP | 3 |
| 2009 | Identifying interesting assertions from the webabstractHow can we cull the facts we need from the overwhelming mass of information and misinformation that is the Web? The TextRunner extraction engine represents one approach, in which people pose keyword queries or simple questions and TextRunner returns concise answers based on tuples extracted from Web text. Unfortunately, the results returned by engines such as TextRunner include both informative facts (e.g., “the FDA banned ephedra”) and less useful statements (e.g., “the FDA banned products”). This paper therefore investigates filtering TextRunner results to enable people to better focus on interesting assertions. We first develop three distinct models of what assertions are likely to be interesting in response to a query. We then fully operationalize each of these models as a filter over TextRunner results. Finally, we develop a more sophisticated filter that combines the different models using relevance feedback. In a study of human ratings of the interestingness of TextRunner assertions, we show that our approach substantially enhances the quality of TextRunner results. Our best filter raises the fraction of interesting results in the top thirty from 41.6 % to 64.1%. Thomas Lin, Oren Etzioni, James Fogarty |
CIKM | 2 |
| 2009 | Lemmatic Machine Translation
Stephen Soderland, Christopher Lim, Mausam, Oren Etzioni, Jonathan Pool |
MTSummit | 5 |
| 2009 | Unsupervised Methods for Determining Object and Relation Synonyms on the WebabstractThe task of identifying synonymous relations and objects, or synonym resolution, is critical for high-quality information extraction. This paper investigates synonym resolution in the context of unsupervised information extraction, where neither hand-tagged training examples nor domain knowledge is available. The paper presents a scalable, fully-implemented system that runs in O(KN log N) time in the number of extractions, N, and the maximum number of synonyms per word, K. The system, called Resolver , introduces a probabilistic relational model for predicting whether two strings are co-referential based on the similarity of the assertions containing them. On a set of two million assertions extracted from the Web, Resolver resolves objects with 78% precision and 68% recall, and resolves relations with 90% precision and 35% recall. Several variations of resolver's probabilistic model are explored, and experiments demonstrate that under appropriate conditions these variations can improve F1 by 5%. An extension to the basic Resolver system allows it to handle polysemous names with 97% precision and 95% recall on a data set from the TREC corpus. Alexander Yates, Oren Etzioni |
J. Artif. Intell. Res. | 2 |
| 2008 | The Tradeoffs Between Open and Traditional Relation Extraction
Michele Banko, Oren Etzioni |
ACL | 2 |
| 2008 | It's a Contradiction - no, it's not: A Case Study using Functional Relations
Alan Ritter, Stephen Soderland, Doug Downey, Oren Etzioni |
EMNLP | 4 |
| 2008 | Scaling Textual Inference to the Web
Stefan Schoenmackers, Oren Etzioni, Daniel S. Weld |
EMNLP | 2 |
| 2008 | Look Ma, No Hands: Analyzing the Monotonic Feature Abstraction for Text ClassificationabstractIs accurate classification possible in the absence of hand-labeled data? This paper introduces the Monotonic Feature (MF) abstraction--where the probability of class membership increases monotonically with the MF's value. The paper proves that when an MF is given, PAC learning is possible with no hand-labeled data under certain assumptions. We argue that MFs arise naturally in a broad range of textual classification applications. On the classic "20 Newsgroups" data set, a learner given an MF and unlabeled data achieves classification accuracy equal to that of a state-of-the-art semi-supervised learner relying on 160 hand-labeled examples. Even when MFs are not given as input, their presence or absence can be determined from a small amount of hand-labeled data, which yields a new semi-supervised learning method that reduces error by 15% on the 20 Newsgroups data. Doug Downey, Oren Etzioni |
NIPS | 2 |
| 2008 | Machine reading at web scaleabstractNo abstract available. Oren Etzioni |
WSDM | 1 |
| 2007 | Sparse Information Extraction: Unsupervised Language Models to the Rescue
Doug Downey, Stefan Schoenmackers, Oren Etzioni |
ACL | 3 |
| 2007 | Structured Querying of Web Text Data: A Technical Challenge
Michael J. Cafarella, Christopher Ré, Dan Suciu, Oren Etzioni |
CIDR | 4 |
| 2007 | Open Information Extraction from the Web
Michele Banko, Michael J. Cafarella, Stephen Soderland, Matthew Broadhead, Oren Etzioni |
IJCAI | 5 |
| 2007 | Locating Complex Named Entities in Web Text
Doug Downey, Matthew Broadhead, Oren Etzioni |
IJCAI | 3 |
| 2007 | Strategies for lifelong knowledge extraction from the webabstractThe increasing availability of electronic text has made it possible to acquire information using a variety of techniques that leverage the expertise of both humans and machines. In particular, the field of Information Extraction (IE), in which knowledge is extracted automatically from text, has shown promise for large-scale knowledge acquisition. While IE systems can uncover assertions about individual entities with an increasing level of sophistication,alltext understanding -- the formation of a coherent theory from a textual corpus -- involves representation and learning abilities not currently achievable by today's IE systems. Compared to individual relational assertions outputted by IE systems, a theory includes coherent knowledge of abstract concepts and the relationships among them. We believe that the ability to fully discover the richness of knowledge present within large, unstructured and heterogeneous corpora will require a lifelong learning process in which earlier learned knowledge is used to guide subsequent learning. This paper introduces Alice, a lifelong learning agent whose goal is to automatically discovera collection of concepts, facts and generalizations that describe a particular topic of interest directly from a large volume of Web text. Building upon recent advances in unsupervised information extraction, we demonstrate that Alice can iteratively discover new concepts and compose general domain knowledge with a precision of 78%. Michele Banko, Oren Etzioni |
K-CAP | 2 |
| 2007 | Machine reading of web textabstractArticle Share on Machine reading of web text Author: Oren Etzioni University of Washington, Seattle, WA University of Washington, Seattle, WAView Profile Authors Info & Claims K-CAP '07: Proceedings of the 4th international conference on Knowledge captureOctober 2007 Pages 1–4https://doi.org/10.1145/1298406.1298407Published:28 October 2007Publication History 2citation320DownloadsMetricsTotal Citations2Total Downloads320Last 12 Months5Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Oren Etzioni |
K-CAP | 1 |
| 2007 | Lexical translation with application to image searching on the web
Oren Etzioni, Kobi Reiter, Stephen Soderland, Marcus Sammer |
MTSummit | 1 |
| 2007 | Unsupervised Resolution of Objects and Relations on the Web
Alexander Yates, Oren Etzioni |
HLT-NAACL | 2 |
| 2007 | Navigating Extracted Data with Schema Discovery
Michael J. Cafarella, Dan Suciu, Oren Etzioni |
WebDB | 3 |
| 2006 | Machine Reading
Oren Etzioni, Michele Banko, Michael J. Cafarella |
AAAI | 1 |
| 2006 | Detecting Parser Errors Using Web-based Semantic Filters
Alexander Yates, Stefan Schoenmackers, Oren Etzioni |
EMNLP | 3 |
| 2006 | Self-supervised Relation Extraction from the Web
Ronen Feldman, Binyamin Rosenfeld, Stephen Soderland, Oren Etzioni |
ISMIS | 4 |
| 2005 | A Probabilistic Model of Redundancy in Information Extraction
Doug Downey, Oren Etzioni, Stephen Soderland |
IJCAI | 2 |
| 2005 | A search engine for natural language applicationsabstractMany modern natural language-processing applications utilize search engines to locate large numbers of Web documents or to compute statistics over the Web corpus. Yet Web search engines are designed and optimized for simple human queries---they are not well suited to support such applications. As a result, these applications are forced to issue millions of successive queries resulting in unnecessary search engine load and in slow applications with limited scalability.In response, this paper introduces the Bindings Engine (BE), which supports queries containing typed variables and string-processing functions. For example, in response to the query "powerful ‹noun›" BE will return all the nouns in its index that immediately follow the word "powerful", sorted by frequency. In response to the query "Cities such as ProperNoun(Head(‹NounPhrase›))", BE will return a list of proper nouns likely to be city names.BE's novel neighborhood index enables it to do so with O(k) random disk seeks and O(k) serial disk reads, where k is the number of non-variable terms in its query. As a result, BE can yield several orders of magnitude speedup for large-scale language-processing applications. The main cost is a modest increase in space to store the index. We report on experiments validating these claims, and analyze how BE's space-time tradeoff scales with the size of its index and the number of variable types. Finally, we describe how a BE-based application extracts thousands of facts from the Web at interactive speeds in response to simple user queries. Michael J. Cafarella, Oren Etzioni |
WWW | 2 |
| 2005 | Unsupervised named-entity extraction from the Web: An experimental study
Oren Etzioni, Michael J. Cafarella, Doug Downey, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates |
Artif. Intell. | 1 |
| 2004 | Methods for Domain-Independent Information Extraction from the Web: An Experimental Comparison
Oren Etzioni, Michael J. Cafarella, Doug Downey, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates |
AAAI | 1 |
| 2004 | PRECISE on ATIS: Semantic Tractability and Experimental Results
Ana-Maria Popescu, Alex Armanasu, Oren Etzioni, David Ko, Alexander Yates |
AAAI | 3 |
| 2004 | Modern Natural Language Interfaces to Databases: Composing Statistical Parsing with Semantic Tractability
Ana-Maria Popescu, Alex Armanasu, Oren Etzioni, David Ko, Alexander Yates |
COLING | 3 |
| 2004 | The Specification of Agent Behavior by Ordinary People: A Case Study
Luke K. McDowell, Oren Etzioni, Alon Y. Halevy |
ISWC | 2 |
| 2004 | Web-scale information extraction in knowitall: (preliminary results)abstractManually querying search engines in order to accumulate a large bodyof factual information is a tedious, error-prone process of piecemealsearch. Search engines retrieve and rank potentially relevantdocuments for human perusal, but do not extract facts, assessconfidence, or fuse information from multiple documents. This paperintroduces KnowItAll, a system that aims to automate the tedious process ofextracting large collections of facts from the web in an autonomous,domain-independent, and scalable manner.The paper describes preliminary experiments in which an instance of KnowItAll, running for four days on a single machine, was able to automatically extract 54,753 facts. KnowItAll associates a probability with each fact enabling it to trade off precision and recall. The paper analyzes KnowItAll's architecture and reports on lessons learned for the design of large-scale information extraction systems. Oren Etzioni, Michael J. Cafarella, Doug Downey, Stanley Kok, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates |
WWW | 1 |
| 2004 | Semantic emailabstractThis paper investigates how the vision of the Semantic Web can be carried overto the realm of email. We introduce a general notion of semantice mail, in which an email message consists of an RDF query or update coupled with corresponding explanatory text. Semantic email opens the door to a wide range of automated, email-mediated applications with formally guaranteed properties. In particular, this paper introduces a broad class of semantic email processes. For example consider the process of sending an email to a program committee asking who will attend the PC dinner automatically collecting the responses and tallying them up. We define bothlogical and decision-theoretic models where an email process ismodeled as a set of updates to a data set on which we specify goals via certain constraints or utilities. We then describe a set ofinference problems that arise while trying to satisfy these goals and analyze their computational tractability. In particular weshow that for the logical model it is possible to automatically infer which email responses are acceptable w.r.t. a set ofconstraints in polynomial time and for the decision-theoreticmodel it is possible to compute the optimal message-handling policy in polynomial time. Finally we discuss our publicly available implementation of semantic email and outline research challenges inthis realm. Luke K. McDowell, Oren Etzioni, Alon Y. Halevy, Henry M. Levy |
WWW | 2 |
| 2004 | Semantic email: theory and applications
Luke K. McDowell, Oren Etzioni, Alon Y. Halevy |
J. Web Semant. | 2 |
| 2003 | Crossing the Structure Chasm
Alon Y. Halevy, Oren Etzioni, AnHai Doan, Zachary G. Ives, Jayant Madhavan, Luke K. McDowell, Igor Tatarinov |
CIDR | 2 |
| 2003 | Automatically Personalizing User Interfaces
Daniel S. Weld, Corin R. Anderson, Pedro M. Domingos, Oren Etzioni, Krzysztof Z. Gajos, Tessa A. Lau, Steven A. Wolfman |
IJCAI | 4 |
| 2003 | Towards a theory of natural language interfaces to databasesabstractThe need for Natural Language Interfaces to databases (NLIs) has become increasingly acute as more and more people access information through their web browsers, PDAs, and cell phones. Yet NLIs are only usable if they map natural language questions to SQL queries correctly. As Schneiderman and Norman have argued, people are unwilling to trade reliable and predictable user interfaces for intelligent but unreliable ones. In this paper, we introduce a theoretical framework for reliable NLIs, which is the foundation for the fully implemented Precise NLI. We prove that, for a broad class of semantically tractable natural language questions, Precise is guaranteed to map each question to the corresponding SQL query. We report on experiments testing Precise on several hundred questions drawn from user studies over three benchmark databases. We find that over 80% of the questions are semantically tractable questions, which Precise answers correctly. Precise automatically recognizes the 20% of questions that it cannot handle, and requests a paraphrase. Finally, we show that Precise compares favorably with Mooney's learning NLI and with Microsoft's English Query product Ana-Maria Popescu, Oren Etzioni, Henry A. Kautz |
IUI | 2 |
| 2003 | Towards a theory of natural language interfaces to databasesabstractThe need for Natural Language Interfaces (NLIs) to databases has become increasingly acute as more nontechnical people access information through their web browsers, PDAs and cell phones. Yet NLIs are only usable if they map natural language questions to SQL queries correctly. We introduce the Precise NLI [2], which reduces the semantic interpretation challenge in NLIs to a graph matching problem. Precise uses the max-flow algorithm to efficiently solve this problem. Each max-flow solution corresponds to a possible semantic interpretation of the sentence. precise collects max-flow solutions, discards the solutions that do not obey syntactic constraints and retains the rest as the basis for generating SQL queries corresponding to the question q. The syntactic information is extracted from the parse tree corresponding to the given question which is computed by a statistical parser [1]. For a broad, well-defined class of semantically tractable natural language questions, Precise is guaranteed to map each question to the corresponding SQL querySemantically tractable questions correspond to a natural, domain-independent subset of English that can be efficiently and accurately interpreted as nonrecursive Datalog clauses. Precise is transportable to arbitrary databases, such as the Restaurants,Jobs and Geography databases used in our implementation. Examples of semantically tractable questions include: "What Chinese restaurants with a 3.5 rating are in Seattle?", "What are the areas of US states with large populations?", "What jobs require 4 years of experience and desire a B.S.CS degree?".Given a question which is not semantically tractable, Precise recognizes it as such and informs the user that it cannot answer it.Given a semantically tractable question, Precise computes the set of non-equivalent SQL interpretations corresponding to the question. If a unique such SQL interpretation exists, Precise outputs it together with the corresponding result set obtained by querying the current database. If the set contains more than one SQL interpretation, the natural language question is ambiguous in the context of the current database. In this case, Precise asks for the user's help in determining which interpretation is the correct one.Our experiments have shown that Precise has high coverage and accuracy over common English questions. In future work, we plan to explore increasingly broad classes of questions and include Precise as a module in a full-fledged dialog system. An important direction for future work is helping users understand the types of questions Precise cannot handle via dialog, enabling them to build an accurate mental model of the system and its capabilities. Also, our own group's work on the EXACT natural language interface [3] builds on Precise and on the underlying theoretical framework. EXACT composes an extended version of Precise with a sound and complete planner to develop a powerful and provably reliable interface to household appliances Ana-Maria Popescu, Oren Etzioni, Henry A. Kautz |
IUI | 2 |
| 2003 | A reliable natural language interface to household appliancesabstractAs household appliances grow in complexity and sophistication, they become harder and harder to use, particularly because of their tiny display screens and limited keyboards. This paper describes a strategy for building natural language interfaces to appliances that circumvents these problems. Our approach leverages decades of research on planning and natural language interfaces to databases by reducing the appliance problem to the database problem; the reduction provably maintains desirable properties of the database interface. The paper goes on to describe the implementation and evaluation of the EXACT interface to appliances, which is based on this reduction. EXACT maps each English user request to an SQL query, which is transformed to create a PDDL goal, and uses the Blackbox planner [13] to map the planning problem to a sequence of appliance commands that satisfy the original request. Both theoretical arguments and experimental evaluation show that EXACT is highly reliable Alexander Yates, Oren Etzioni, Daniel S. Weld |
IUI | 2 |
| 2003 | To buy or not to buy: mining airfare data to minimize ticket purchase priceabstractAs product prices become increasingly available on the World Wide Web, consumers attempt to understand how corporations vary these prices over time. However, corporations change prices based on proprietary algorithms and hidden variables (e.g., the number of unsold seats on a flight). Is it possible to develop data mining techniques that will enable consumers to predict price changes under these conditions?This paper reports on a pilot study in the domain of airline ticket prices where we recorded over 12,000 price observations over a 41 day period. When trained on this data, Hamlet --- our multi-strategy data mining algorithm --- generated a predictive model that saved 341 simulated passengers $198,074 by advising them when to buy and when to postpone ticket purchases. Remarkably, a clairvoyant algorithm with complete knowledge of future prices could save at most $320,572 in our simulation, thus HAMLET's savings were 61.8% of optimal. The algorithm's savings of $198,074 represents an average savings of 23.8% for the 341 passengers for whom savings are possible. Overall, HAMLET saved 4.4% of the ticket price averaged over the entire set of 4,488 simulated passengers. Our pilot study suggests that mining of price data available over the web has the potential to save consumers substantial sums of money per annum. Oren Etzioni, Rattapoom Tuchinda, Craig A. Knoblock, Alexander Yates |
KDD | 1 |
| 2003 | Mangrove: Enticing Ordinary People onto the Semantic Web via Instant Gratification
Luke K. McDowell, Oren Etzioni, Steve D. Gribble, Alon Y. Halevy, Henry M. Levy, William Pentney, Stani Vlasseva |
ISWC | 2 |
| 2003 | Semantic Email: Adding Lightweight Data Manipulation Capabilities to the Email Habitat
Oren Etzioni, Alon Y. Halevy, Henry M. Levy, Luke K. McDowell |
WebDB | 1 |
| 2001 | Scaling question answering to the WebabstractThe wealth of information on the web makes it an attractive resource for seeking quick answers to simple, factual questions such as "who was the first American in space?" or "what is the second tallest mountain in the world?" Yet today's most advanced web search services (e.g., Google and AskJeeves) make it surprisingly tedious to locate answers to such questions. In this paper, we extend question-answering techniques, first studied in the information retrieval literature, to the web and experimentally evaluate their performance. First we introduce MULDER, which we believe to be the first general-purpose, fully-automated question-answering system available on the web. Second, we describe MULDER's architecture, which relies on multiple search-engine queries, natural-language parsing, and a novel voting procedure to yield reliable answers coupled with high recall. Finally, we compare MULDER's performance to that of Google and AskJeeves on questions drawn from the TREC-8 question track. We find that MULDER's recall is more than a factor of three higher than that of AskJeeves. In addition, we find that Google requires 6.6 times as much user effort to achieve the same level of recall as MULDER. 1. Cody C. T. Kwok, Oren Etzioni, Daniel S. Weld |
WWW | 2 |
| 2001 | Scaling question answering to the webabstractThe wealth of information on the web makes it an attractive resource for seeking quick answers to simple, factual questions such as “who was the first American in space?” or “what is the second tallest mountain in the world?” Yet today's most advanced web search services (e.g., Google and AskJeeves) make it surprisingly tedious to locate answers to such questions. In this paper, we extend question-answering techniques, first studied in the information retrieval literature, to the web and experimentally evaluate their performance.First we introduce Mulder, which we believe to be the first general-purpose, fully-automated question-answering system available on the web. Second, we describe Mulder's architecture, which relies on multiple search-engine queries, natural-language parsing, and a novel voting procedure to yield reliable answers coupled with high recall. Finally, we compare Mulder's performance to that of Google and AskJeeves on questions drawn from the TREC-8 question answering track. We find that Mulder's recall is more than a factor of three higher than that of AskJeeves. In addition, we find that Google requires 6.6 times as much user effort to achieve the same level of recall as Mulder. Cody C. T. Kwok, Oren Etzioni, Daniel S. Weld |
ACM Trans. Inf. Syst. | 2 |
| 2000 | Towards adaptive Web sites: Conceptual framework and case study
Mike Perkowitz, Oren Etzioni |
Artif. Intell. | 2 |
| 2000 | Query routing for Web search engines: architecture and experiments
Atsushi Sugiura, Oren Etzioni |
Comput. Networks | 2 |
| 2000 | Optimal Information Gathering on the Internet with Time and Cost ConstraintsabstractThe World Wide Web provides access to vast amounts of information, but content providers are considering charging for the information and services they supply. Thus the consumer may face the problem of balancing the benefit of asking for information against the cost (in terms of both money and time) of acquiring it. We study information-gathering strategies that maximize the expected value to the consumer. In our model there is a single information request, which has a known benefit to the consumer. To satisfy the request, queries can be sent simultaneously or in sequence to any of a finite set of independent information sources. For each source we know the monetary cost of making the query, the amount of time it will take, and the probability that the source will be able to provide the requested information. A policy specifies which sources to contact at which times, and the expected value of the policy can be defined as some function of the likelihood that the policy will yield an answer, the expected benefit, and the monetary cost and time delay associated with executing the policy. The problem is to find an expected-value-maximizing policy. We explore four variants of the objective function V: (i) V consists only of the benefit term subject to threshold constraints on both total cost and total elapsed time, (ii) V is linear in the expected total cost of the policy subject to the constraint that the total elapsed time never exceeds somedeadline, (iii) V is linear in the expected total elapsed time subject to the constraint that the total cost never exceeds some threshold, and (iv) V is linear in the expected total monetary cost and the expected time delay of the policy. The problems of devising an optimal querying policy for all four variants and approximating an optimal querying policy for variants (iii) and (iv) are shown to be NP-hard. For (i), and with a mild simplifying assumption for (iii), we give a fully polynomial time approximation scheme. For (ii), we consider batched querying policies, and design an O(n 2 ) time approximation algorithm with ratio $\frac{1}{2}$ and a polynomial time approximation scheme for optimal single-batch policies, and an O(kn 2 ) time approximation algorithm with ratio $\frac{1}{5}$ for optimal k-batch policies. Oren Etzioni, Steve Hanks, Tao Jiang 0001, Omid Madani |
SIAM J. Comput. | 1 |
| 1999 | Adaptive Web Sites: Conceptual Cluster Mining
Mike Perkowitz, Oren Etzioni |
IJCAI | 2 |
| 1999 | Towards Adaptive Web Sites: Conceptual Framework and Case Study
Mike Perkowitz, Oren Etzioni |
Comput. Networks | 2 |
| 1999 | Grouper: A Dynamic Clustering Interface to Web Search Results
Oren Zamir, Oren Etzioni |
Comput. Networks | 2 |
| 1998 | Web Document Clustering: A Feasibility DemonstrationabstractUsers of Web search engines are often forced to sift through the long ordered list of document returned by the engines. The IR community has explored document clustering as an alternative method of organizing retrieval results, but clustering has yet to be deployed on the major search engines. The paper articulates the unique requirements of Web document clustering and reports on the first evaluation of clustering methods in this domain. A key requirement is that the methods create their clusters based on the short snippets returned by Web search engines. Surprisingly, we find that clusters based on snippets are almost as good as clusters created using the full text of Web documents. To satisfy the stringent requirements of the Web domain, we introduce an incremental, linear time (in the document collection size) algorithm called Suffix Tree Clustering (STC). which creates clusters based on phrases shared between documents. We show that STC is faster than standard clustering methods in this domain, and argue that Web document clustering via STC is both feasible and potentially beneficial. Oren Zamir, Oren Etzioni |
SIGIR | 2 |
| 1997 | Adaptive Web Sites: an AI Challenge
Mike Perkowitz, Oren Etzioni |
IJCAI (1) | 2 |
| 1997 | Fast and Intuitive Clustering of Web Documents
Oren Zamir, Oren Etzioni, Omid Madani, Richard M. Karp |
KDD | 2 |
| 1997 | Sound and Efficient Closed-World Reasoning for Planning
Oren Etzioni, Keith Golden, Daniel S. Weld |
Artif. Intell. | 1 |
| 1997 | Dynamic Reference Sifting: A Case Study in the Homepage Domain
Jonathan Shakes, Marc Langheinrich, Oren Etzioni |
Comput. Networks | 3 |
| 1997 | Learning to Understand Information on the Internet: An Example-Based Approach
Mike Perkowitz, Robert B. Doorenbos, Oren Etzioni, Daniel S. Weld |
J. Intell. Inf. Syst. | 3 |
| 1996 | Efficient Information Gathering on the Internet (extended abstract)abstractThe Internet offers unprecedented access to information. At present most of this information is free, but information providers ore likely to start charging for their services in the near future. With that in mind this paper introduces the following information access problem: given a collection of n information sources, each of which has a known time delay, dollar cost and probability of providing the needed information, find an optimal schedule for querying the information sources. We study several variants of the problem which differ in the definition of an optimal schedule. We first consider a cost model in which the problem is to minimize the expected total cost (monetary and time) of the schedule, subject to the requirement that the schedule may terminate only when the query has been answered or all sources have been queried unsuccessfully. We develop an approximation algorithm for this problem and for an extension of the problem in which more than a single item of information is being sought. We then develop approximation algorithms for a reward model in which a constant reward is earned if the information is successfully provided, and we seek the schedule with the maximum expected difference between the reward and a measure of cost. The monetary and time costs may either appear in the cost measure or be constrained not to exceed a fixed upper bound; these options give rise to four different variants of the reward model. Oren Etzioni, Steve Hanks, Tao Jiang 0001, Richard M. Karp, Omid Madani, Orli Waarts |
FOCS | 1 |
| 1996 | Scaling Up Goal Recognition
Neal Lesh, Oren Etzioni |
KR | 2 |
| 1995 | A Sound and Fast Goal Recognizer
Neal Lesh, Oren Etzioni |
IJCAI | 2 |
| 1995 | Category Translation: Learning to Understand Information on the Internet
Mike Perkowitz, Oren Etzioni |
IJCAI (1) | 2 |
| 1994 | Learning About Software Errors Via Systematic Experimentation
Terrance Goan, Oren Etzioni |
AAAI | 2 |
| 1994 | Omnipotence Without Omniscience: Efficient Sensor Management for Planning
Keith Golden, Oren Etzioni, Daniel S. Weld |
AAAI | 2 |
| 1994 | Database Learning for Software Agents
Mike Perkowitz, Oren Etzioni |
AAAI | 2 |
| 1994 | Learning Decision Lists Using Homogeneous Rules
Richard B. Segal, Oren Etzioni |
AAAI | 2 |
| 1994 | The First Law of Robotics (A Call to Arms)
Daniel S. Weld, Oren Etzioni |
AAAI | 2 |
| 1994 | Tractable Closed World Reasoning with Updates
Oren Etzioni, Keith Golden, Daniel S. Weld |
KR | 1 |
| 1994 | Statistical Methods for Analyzing Speedup Learning Experiments
Oren Etzioni, Ruth Etzioni |
Mach. Learn. | 1 |
| 1993 | Acquiring Search-Control Knowledge Via Static Analysis
Oren Etzioni |
Artif. Intell. | 1 |
| 1993 | A Structural Theory of Explanation-Based Learning
Oren Etzioni |
Artif. Intell. | 1 |
| 1992 | An Asymptotic Analysis of Speedup Learning
Oren Etzioni |
ML | 1 |
| 1992 | Why EBL Produces Overly-Specific Knowledge: A Critique of the PRODIGY Approaches
Oren Etzioni, Steven Minton |
ML | 1 |
| 1992 | DYNAMIC: A New Role for Training Problems in EBL
M. Alicia Pérez, Oren Etzioni |
ML | 2 |
| 1992 | An Approach to Planning with Incomplete Information
Oren Etzioni, Steve Hanks, Daniel S. Weld, Denise Draper, Neal Lesh, Mike Williamson |
KR | 1 |
| 1991 | STATIC: A Problem-Space Compiler for PRODIGY
Oren Etzioni |
AAAI | 1 |
| 1991 | Integrating Abstraction and Explanation-Based Learning in PRODIGY
Craig A. Knoblock, Steven Minton, Oren Etzioni |
AAAI | 3 |
| 1991 | Integrating Efficient Model-Learning and Problem-Solving Algorithms in Permutation Environments
Prasad Chalasani, Oren Etzioni, John Mount |
KR | 2 |
| 1991 | Embedding Decision-Analytic Control in a Learning Architecture
Oren Etzioni |
Artif. Intell. | 1 |
| 1990 | Why PRODIGY/EBL Works
Oren Etzioni |
AAAI | 1 |
| 1989 | Tractable Decision-Analytic Control
Oren Etzioni |
KR | 1 |
| 1989 | Explanation-Based Learning: A Problem Solving Perspective
Steven Minton, Jaime G. Carbonell, Craig A. Knoblock, Daniel Kuokka, Oren Etzioni, Yolanda Gil |
Artif. Intell. | 5 |
| 1988 | Hypothesis Filtering: A Practical Approach to Reliable Learning
Oren Etzioni |
ML | 1 |