Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Oren Etzioni

dblp:e/OEtzioni · DBLP profile ↗
← Back
110ranked-venue papers
27as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 84 · 23 first-author · 1 since 2021Databases, data management, data science and information retrieval · 26 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-authorTheory of computation · 7 · 5 first-authorHuman-computer interaction and ubiquitous computing · 5Computer networks · 4Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
61 papers
Information extraction and text analysis · 41% Question answering and dialogue systems · 20% Knowledge representation and reasoning · 17%
Databases, data mining, and information retrieval
22 papers
Data mining · 50% Knowledge graphs · 27% Information retrieval · 12%
Human-computer interaction and pervasive computing
5 papers
Collaborative and social computing · 39% Games and playful interaction · 24% Interaction techniques and input · 23%

Topics — the 30 heaviest of 117, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
multimodal communication
0.512021
Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and Text · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis
open information extraction
0.542012
Open Language Learning for Information Extraction · EMNLP-CoNLL 2012
Open Information Extraction: The Second Generation · IJCAI 2011
Identifying Relations for Open Information Extraction · EMNLP 2011
Natural language and speech › Question answering and dialogue systems
open-domain question answering
0.422014
Open question answering over curated and extracted knowledge bases · KDD 2014
Paraphrase-Driven Learning for Open Question Answering · ACL (1) 2013
Natural language and speech › Information extraction and text analysis
relation extraction
0.332011
Identifying Relations for Open Information Extraction · EMNLP 2011
Identifying Functional Relations in Web Text · EMNLP 2010
The Tradeoffs Between Open and Traditional Relation Extraction · ACL 2008
Natural language and speech › Information extraction and text analysis
named entity recognition
0.332011
Named Entity Recognition in Tweets: An Experimental Study · EMNLP 2011
Locating Complex Named Entities in Web Text · IJCAI 2007
Unsupervised named-entity extraction from the Web: An experimental study · Artif. Intell. 2005
Knowledge, reasoning and agents › Knowledge representation and reasoning › automated reasoning
knowledge base reasoning
0.212016
Combining Retrieval, Statistics, and Inference to Answer Elementary Science Questions · AAAI 2016
Collaborative and social computing
online communities
0.212016
Toward Automatic Bootstrapping of Online Communities Using Decision-theoretic Optimization · CSCW 2016
Data mining
text mining
0.222014
The battle for the future of data mining · KDD 2014
To buy or not to buy: that is the question · KDD 2013
Computer vision › 3D vision
geometric reasoning
0.212015
Solving Geometry Problems: Combining Text and Diagram Interpretation · EMNLP 2015
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
geometry problem solving
0.212015
Solving Geometry Problems: Combining Text and Diagram Interpretation · EMNLP 2015
Knowledge, reasoning and agents › Knowledge representation and reasoning › probabilistic reasoning › probabilistic logic
markov logic networks
0.212015
Exploring Markov Logic Networks for Question Answering · EMNLP 2015
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
science question answering
0.212015
Exploring Markov Logic Networks for Question Answering · EMNLP 2015
Knowledge graphs
knowledge graph construction
0.222014
The battle for the future of data mining · KDD 2014
Machine reading at web scale · WSDM 2008
Data mining
predictive modeling
0.222013
To buy or not to buy: that is the question · KDD 2013
To buy or not to buy: mining airfare data to minimize ticket purchase price · KDD 2003
Data mining › predictive modeling › forecasting
price prediction
0.222013
To buy or not to buy: that is the question · KDD 2013
To buy or not to buy: mining airfare data to minimize ticket purchase price · KDD 2003
Data mining › text mining › sentiment analysis
review mining
0.222013
RevMiner: an extractive interface for navigating reviews on a smartphone · UIST 2012
To buy or not to buy: that is the question · KDD 2013
Computer vision › Vision and language › multimodal understanding
diagram understanding
0.212014
Diagram Understanding in Geometry Questions · AAAI 2014
Natural language and speech › Question answering and dialogue systems
knowledge base question answering
0.212014
Open question answering over curated and extracted knowledge bases · KDD 2014
Natural language and speech › Question answering and dialogue systems
math word problem solving
0.212014
Learning to Solve Arithmetic Word Problems with Verb Categorization · EMNLP 2014
Natural language and speech › Information extraction and text analysis › lexical semantics
verb classification
0.212014
Learning to Solve Arithmetic Word Problems with Verb Categorization · EMNLP 2014
Natural language and speech › Information extraction and text analysis › event analysis › event understanding
event schema induction
0.212013
Generating Coherent Event Schemas at Scale · EMNLP 2013
Natural language and speech › Information extraction and text analysis › relation extraction
joint entity and relation extraction
0.212013
Generating Coherent Event Schemas at Scale · EMNLP 2013
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.122008
Machine reading at web scale · WSDM 2008
Machine Reading · AAAI 2006
Natural language and speech › Information extraction and text analysis
entity linking
0.112012
No Noun Phrase Left Behind: Detecting and Typing Unlinkable Entities · EMNLP-CoNLL 2012
Natural language and speech › Information extraction and text analysis › event extraction
event classification
0.112012
Open domain event extraction from twitter · KDD 2012
Natural language and speech › Information extraction and text analysis
event extraction
0.112012
Open domain event extraction from twitter · KDD 2012
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.112012
Open domain event extraction from twitter · KDD 2012
Natural language and speech › Information extraction and text analysis › event extraction
open event extraction
0.112012
Open domain event extraction from twitter · KDD 2012
Knowledge graphs › knowledge graph construction › knowledge extraction
entity typing
0.112012
No Noun Phrase Left Behind: Detecting and Typing Unlinkable Entities · EMNLP-CoNLL 2012
Interaction techniques and input
mobile interaction
0.112012
RevMiner: an extractive interface for navigating reviews on a smartphone · UIST 2012

Methods — techniques the papers use, named apart from their topics

case study · 0.6simulation · 0.5decision-theoretic optimization · 0.5markov logic networks · 0.5submodular optimization · 0.4text mining · 0.4probabilistic inference · 0.3unsupervised learning · 0.3natural language processing · 0.3integer programming · 0.2information retrieval · 0.2corpus statistics · 0.2semantic parsing · 0.2first-order logic · 0.2dependency parsing · 0.2knowledge base integration · 0.2deep learning · 0.2data mining · 0.2
YearPublicationVenuePosition
2021 Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and Text
abstract
Christopher Clark, Jordi Salvador, Dustin Schwenk, Derrick Bonafilia, Mark Yatskar, Eric Kolve, Alvaro Herrasti, Jonghyun Choi, Sachin Mehta, Sam Skjonsberg, Carissa Schoenick, Aaron Sarnat, Hannaneh Hajishirzi, Aniruddha Kembhavi, Oren Etzioni, Ali Farhadi. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Jordi Salvador, Dustin Schwenk, Derrick Bonafilia, Mark Yatskar, Eric Kolve, Alvaro Herrasti, Sachin Mehta, Sam Skjonsberg, Carissa Schoenick, Aaron Sarnat, Hannaneh Hajishirzi, Aniruddha Kembhavi, Oren Etzioni, Ali Farhadi
EMNLP (1)15
2016 Combining Retrieval, Statistics, and Inference to Answer Elementary Science Questions
abstract
What capabilities are required for an AI system to pass standard 4th Grade Science Tests? Previous work has examined the use of Markov Logic Networks (MLNs) to represent the requisite background knowledge and interpret test questions, but did not improve upon an information retrieval (IR) baseline. In this paper, we describe an alternative approach that operates at three levels of representation and reasoning: information retrieval, corpus statistics, and simple inference over a semi-automatically constructed knowledge base, to achieve substantially improved results. We evaluate the methods on six years of unseen, unedited exam questions from the NY Regents Science Exam (using only non-diagram, multiple choice questions), and show that our overall system’s score is 71.3%, an improvement of 23.8% (absolute) over the MLN-based method described in previous work. We conclude with a detailed analysis, illustrating the complementary strengths of each method in the ensemble. Our datasets are being released to enable further research.
Peter Clark, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter D. Turney, Daniel Khashabi
AAAI2
2016 Toward Automatic Bootstrapping of Online Communities Using Decision-theoretic Optimization
abstract
Successful online communities (e.g., Wikipedia, Yelp, and StackOverflow) can produce valuable content. However, many communities fail in their initial stages. Starting an online community is challenging because there is not enough content to attract a critical mass of active members. This paper examines methods for addressing this cold-start problem in datamining-bootstrappable communities by attracting non-members to contribute to the community. We make four contributions: 1) we characterize a set of communities that are “datamining-bootstrappable” and define the bootstrapping problem in terms of decision-theoretic optimization, 2) we estimate the model parameters in a case study involving the Open AI Resources website, 3) we demonstrate that non-members' predicted interest levels and request design are important features that can significantly affect the contribution rate, and 4) we ran a simulation experiment using data generated with the learned parameters and show that our decision-theoretic optimization algorithm can generate as much community utility when bootstrapping the community as our strongest baseline while issuing only 55% as many contribution requests.
Shih-Wen Huang, Jonathan Bragg, Isaac Cowhey, Oren Etzioni, Daniel S. Weld
CSCW4
2016 Question Answering via Integer Programming over Semi-Structured Knowledge
Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Peter Clark, Oren Etzioni, Dan Roth 0001
IJCAI5
2015 Exploring Markov Logic Networks for Question Answering
abstract
Elementary-level science exams pose sig-nificant knowledge acquisition and rea-soning challenges for automatic question answering. We develop a system that rea-sons with knowledge derived from text-books, represented in a subset of first-order logic. Automatic extraction, while scalable, often results in knowledge that is incomplete and noisy, motivating use of reasoning mechanisms that handle uncer-tainty. Markov Logic Networks (MLNs) seem a natural model for expressing such knowl-edge, but the exact way of leveraging MLNs is by no means obvious. We in-vestigate three ways of applying MLNs to our task. First, we simply use the extracted science rules directly as MLN clauses and exploit the structure present in hard con-straints to improve tractability. Second, we interpret science rules as describing prototypical entities, resulting in a drasti-cally simplified but brittle network. Our third approach, called Praline, uses MLNs to align lexical elements as well as define and control how inference should be per-formed in this task. Praline demonstrates a 15 % accuracy boost and a 10x reduction in runtime as compared to other MLN-based methods, and comparable accuracy to word-based baseline approaches.
Tushar Khot, Niranjan Balasubramanian, Eric Gribkoff, Ashish Sabharwal, Peter Clark, Oren Etzioni
EMNLP6
2015 Solving Geometry Problems: Combining Text and Diagram Interpretation
abstract
This paper introduces GEOS, the first automated system to solve unaltered SAT geometry questions by combining text understanding and diagram interpretation.We model the problem of understanding geometry questions as submodular optimization, and identify a formal problem description likely to be compatible with both the question text and diagram.GEOS then feeds the description to a geometric solver that attempts to determine the correct answer.In our experiments, GEOS achieves a 49% score on official SAT questions, and a score of 61% on practice questions. 1 Finally, we show that by integrating textual and visual information, GEOS boosts the accuracy of dependency and semantic parsing of the question text.
Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi, Oren Etzioni, Clint Malcolm
EMNLP4
2015 The elephant in the room: getting value from Big Data
abstract
International audience
Serge Abiteboul, Xin Dong 0001, Oren Etzioni, Divesh Srivastava, Gerhard Weikum, Julia Stoyanovich, Fabian M. Suchanek
WebDB3
2015 Parsing Algebraic Word Problems into Equations
abstract
This paper formalizes the problem of solving multi-sentence algebraic word problems as that of generating and scoring equation trees. We use integer linear programming to generate equation trees and score their likelihood by learning local and global discriminative models. These models are trained on a small set of word problems and their answers, without any manual annotation, in order to choose the equation that best matches the problem text. We refer to the overall system as Alges. We compare Alges with previous work and show that it covers the full gamut of arithmetic operations whereas Hosseini et al. (2014) only handle addition and subtraction. In addition, Alges overcomes the brittleness of the Kushman et al. (2014) approach on single-equation problems, yielding a 15% to 50% reduction in error.
Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, Siena Dumas Ang
Trans. Assoc. Comput. Linguistics4
2014 Diagram Understanding in Geometry Questions
abstract
Automatically solving geometry questions is a long-standing AI problem. A geometry question typically includes a textual description accompanied by a diagram. The first step in solving geometry questions is diagram understanding, which consists of identifying visual elements in the diagram, their locations, their geometric properties, and aligning them to corresponding textual descriptions. In this paper, we present a method for diagram understanding that identifies visual elements in a diagram while maximizing agreement between textual and visual data. We show that the method's objective function is submodular; thus we are able to introduce an efficient method for diagram understanding that is close to optimal. To empirically evaluate our method, we compile a new dataset of geometry questions (textual descriptions and diagrams) and compare with baselines that utilize standard vision techniques. Our experimental evaluation shows an F1 boost of more than 17% in identifying visual elements and 25% in aligning visual elements with their textual descriptions.
Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi, Oren Etzioni
AAAI4
2014 Chinese Open Relation Extraction for Knowledge Acquisition
abstract
Yuen-Hsien Tseng, Lung-Hao Lee, Shu-Yen Lin, Bo-Shun Liao, Mei-Jun Liu, Hsin-Hsi Chen, Oren Etzioni, Anthony Fader. Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, volume 2: Short Papers. 2014.
Yuen-Hsien Tseng, Lung-Hao Lee, Shu-Yen Lin, Bo-Shun Liao, Meijun Liu, Hsin-Hsi Chen, Oren Etzioni, Anthony Fader
EACL7
2014 Learning to Solve Arithmetic Word Problems with Verb Categorization
abstract
This paper presents a novel approach to learning to solve simple arithmetic word problems.Our system, ARIS, analyzes each of the sentences in the problem statement to identify the relevant variables and their values.ARIS then maps this information into an equation that represents the problem, and enables its (trivial) solution as shown in Figure 1.The paper analyzes the arithmetic-word problems "genre", identifying seven categories of verbs used in such problems.ARIS learns to categorize verbs with 81.2% accuracy, and is able to solve 77.7% of the problems in a corpus of standard primary school test questions.We report the first learning results on this task without reliance on predefined templates and make our data publicly available.
Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, Nate Kushman
EMNLP3
2014 The battle for the future of data mining
abstract
Deep learning has catapulted to the front page of the New York Times, formed the core of the so-called 'Google brain', and achieved impressive results in vision, speech recognition, and elsewhere. Yet researchers have offered simple conundrums that deep learning doesn't address. For example, consider the sentence: 'The large ball crashed right through the table because it was made of Styrofoam.' What was made of Styrofoam? The large ball? Or the table? The answer is obviously 'the table', but if we change the word 'Styrofoam' to 'steel', the answer is clearly 'the large ball'. To automatically answer this type of question, our computers require an extensive body of knowledge. We believe that text mining can provide the requisite body of knowledge. My talk will describe work at the new Allen Institute for AI towards building the next-generation of text-mining systems.
Oren Etzioni
KDD1
2014 Open question answering over curated and extracted knowledge bases
abstract
We consider the problem of open-domain question answering (Open QA) over massive knowledge bases (KBs). Existing approaches use either manually curated KBs like Freebase or KBs automatically extracted from unstructured text. In this paper, we present OQA, the first approach to leverage both curated and extracted KBs.
Anthony Fader, Luke Zettlemoyer, Oren Etzioni
KDD3
2013 Paraphrase-Driven Learning for Open Question Answering
Anthony Fader, Luke Zettlemoyer, Oren Etzioni
ACL (1)3
2013 Generating Coherent Event Schemas at Scale
abstract
Chambers and Jurafsky (2009) demonstrated that event schemas can be automatically induced from text corpora.However, our analysis of their schemas identifies several weaknesses, e.g., some schemas lack a common topic and distinct roles are incorrectly mixed into a single actor.It is due in part to their pair-wise representation that treats subjectverb independently from verb-object.This often leads to subject-verb-object triples that are not meaningful in the real-world.We present a novel approach to inducing open-domain event schemas that overcomes these limitations.Our approach uses cooccurrence statistics of semantically typed relational triples, which we call Rel-grams (relational n-grams).In a human evaluation, our schemas outperform Chambers's schemas by wide margins on several evaluation criteria.Both Rel-grams and event schemas are freely available to the research community.
Niranjan Balasubramanian, Stephen Soderland, Mausam, Oren Etzioni
EMNLP4
2013 To buy or not to buy: that is the question
abstract
Shopping can be decomposed into three basic questions: what, where, and when to buy? In this talk, I'll describe how we utilize advanced data-mining and text-mining techniques at Decide.com (and earlier at Farecast) to solve these problems for on-line shoppers. Our algorithms have predicted prices utilizing billions of data points, and ranked products based on millions of reviews.
Oren Etzioni
KDD1
2013 Towards Coherent Multi-Document Summarization
Janara Christensen, Mausam, Stephen Soderland, Oren Etzioni
HLT-NAACL4
2013 Modeling Missing Data in Distant Supervision for Information Extraction
abstract
Distant supervision algorithms learn information extraction models given only large readily available databases and text collections. Most previous work has used heuristics for generating labeled data, for example assuming that facts not contained in the database are not mentioned in the text, and facts in the database must be mentioned at least once. In this paper, we propose a new latent-variable approach that models missing data. This provides a natural way to incorporate side information, for instance modeling the intuition that text will often mention rare entities which are likely to be missing in the database. Despite the added complexity introduced by reasoning about missing data, we demonstrate that a carefully designed local search approach to inference is very accurate and scales to large datasets. Experiments demonstrate improved performance for binary and unary relation extraction when compared to learning with heuristic labels, including on average a 27% increase in area under the precision recall curve in the binary case.
Alan Ritter, Luke Zettlemoyer, Mausam, Oren Etzioni
Trans. Assoc. Comput. Linguistics4
2012 No Noun Phrase Left Behind: Detecting and Typing Unlinkable Entities
Thomas Lin, Mausam, Oren Etzioni
EMNLP-CoNLL3
2012 Open Language Learning for Information Extraction
Mausam, Michael Schmitz 0002, Stephen Soderland, Robert Bart, Oren Etzioni
EMNLP-CoNLL5
2012 Open domain event extraction from twitter
abstract
Tweets are the most up-to-date and inclusive stream of in- formation and commentary on current events, but they are also fragmented and noisy, motivating the need for systems that can extract, aggregate and categorize important events. Previous work on extracting structured representations of events has focused largely on newswire text; Twitter's unique characteristics present new challenges and opportunities for open-domain event extraction. This paper describes TwiCal-- the first open-domain event-extraction and categorization system for Twitter. We demonstrate that accurately extracting an open-domain calendar of significant events from Twitter is indeed feasible. In addition, we present a novel approach for discovering important event categories and classifying extracted events based on latent variable models. By leveraging large volumes of unlabeled data, our approach achieves a 14% increase in maximum F1 over a supervised baseline. A continuously updating demonstration of our system can be viewed at http://statuscalendar.com; Our NLP tools are available at http://github.com/aritter/ twitter_nlp.
Alan Ritter, Mausam, Oren Etzioni, Sam Clark
KDD3
2012 RevMiner: an extractive interface for navigating reviews on a smartphone
abstract
Smartphones are convenient, but their small screens make searching, clicking, and reading awkward. Thus, perusing product reviews on a smartphone is difficult. In response, we introduce RevMiner - a novel smartphone interface that utilizes Natural Language Processing techniques to analyze and navigate reviews. RevMiner was run over 300K Yelp restaurant reviews extracting attribute-value pairs, where attributes represent restaurant attributes such as sushi and service, and values represent opinions about the attributes such as fresh or fast. These pairs were aggregated and used to: 1) answer queries such as "cheap Indian food", 2) concisely present information about each restaurant, and 3) identify similar restaurants. Our user studies demonstrate that on a smartphone, participants preferred RevMiner's interface to tag clouds and color bars, and that they preferred RevMiner's results to Yelp's, particularly for conjunctive queries (e.g., "great food and huge portions"). Demonstrations of RevMiner are available at revminer.com.
Jeff Huang 0002, Oren Etzioni, Luke Zettlemoyer, Kevin Clark, Christian Lee
UIST2
2011 Identifying Relations for Open Information Extraction
Anthony Fader, Stephen Soderland, Oren Etzioni
EMNLP3
2011 Named Entity Recognition in Tweets: An Experimental Study
Alan Ritter, Sam Clark, Mausam, Oren Etzioni
EMNLP4
2011 Open Information Extraction: The Second Generation
abstract
How do we scale information extraction to the massive size and unprecedented heterogeneity of the Web corpus? Beginning in 2003, our KnowItAll project has sought to extract high-quality knowledge from the Web. In 2007, we introduced the Open Information Extraction (Open IE) paradigm which eschews handlabeled training examples, and avoids domainspecific verbs and nouns, to develop unlexicalized, domain-independent extractors that scale to the Web corpus. Open IE systems have extracted billions of assertions as the basis for both commonsense knowledge and novel question-answering systems. This paper describes the second generation of Open IE systems, which rely on a novel model of how relations and their arguments are expressed in English sentences to double precision/recall compared with previous systems such as TEXTRUNNER and WOE. 1
Oren Etzioni, Anthony Fader, Janara Christensen, Stephen Soderland, Mausam
IJCAI1
2011 An analysis of open information extraction based on semantic role labeling
abstract
Open Information Extraction extracts relations from text without requiring a pre-specified domain or vocabulary. While existing techniques have used only shallow syntactic features, we investigate the use of semantic role labeling techniques for the task of Open IE. Semantic role labeling (SRL) and Open IE, although developed mostly in isolation, are quite related. We compare SRL-based open extractors, which perform computationally expensive, deep syntactic analysis, with TextRunner, an open extractor, which uses shallow syntactic analysis but is able to analyze many more sentences in a fixed amount of time and thus exploit corpus-level statistics. Our evaluation answers questions regarding these systems, including, can SRL extractors, which are trained on PropBank, cope with heterogeneous text found on the Web? Which extractor attains better precision, recall, f-measure, or running time? How does extractor performance vary for binary, n-ary and nested relations? How much do we gain by running multiple extractors? How do we select the optimal extractor given amount of data, available time, types of extractions desired?
Janara Christensen, Mausam, Stephen Soderland, Oren Etzioni
K-CAP4
2010 Panlingual Lexical Translation via Probabilistic Inference
abstract
The bare minimum lexical resource required to translate between a pair of languages is a translation dictionary. Unfortunately, dictionaries exist only between a tiny fraction of the 49 million possible language-pairs making machine translation virtually impossible between most of the languages. This paper summarizes the last four years of our research motivated by the vision of panlingual communication. Our research comprises three key steps. First, we compile over 630 freely available dictionaries over the Web and convert this data into a single representation – the translation graph. Second, we build several inference algorithms that infer translations between word pairs even when no dictionary lists them as translations. Finally, we run our inference procedure offline to construct PANDICTIONARY– a sense-distinguished, massively multilingual dictionary that has translations in more than 1000 languages. Our experiments assess the quality of this dictionary and find that we have 4 times as many translations at a high precision of 0.9 compared to the English Wiktionary, which is the lexical resource closest to PANDICTIONARY.
Mausam, Stephen Soderland, Oren Etzioni
AAAI3
2010 A Latent Dirichlet Allocation Method for Selectional Preferences
Alan Ritter, Mausam, Oren Etzioni
ACL3
2010 Identifying Functional Relations in Web Text
Thomas Lin, Mausam, Oren Etzioni
EMNLP3
2010 Learning First-Order Horn Clauses from Web Text
Stefan Schoenmackers, Jesse Davis, Oren Etzioni, Daniel S. Weld
EMNLP3
2010 Analysis of a probabilistic model of redundancy in unsupervised information extraction
Doug Downey, Oren Etzioni, Stephen Soderland
Artif. Intell.2
2010 Panlingual lexical translation via probabilistic inference
Mausam, Stephen Soderland, Oren Etzioni, Daniel S. Weld, Kobi Reiter, Michael Skinner, Marcus Sammer, Jeff A. Bilmes
Artif. Intell.3
2009 Compiling a Massive, Multilingual Dictionary via Probabilistic Inference
Mausam, Stephen Soderland, Oren Etzioni, Daniel S. Weld, Michael Skinner, Jeff A. Bilmes
ACL/IJCNLP3
2009 Identifying interesting assertions from the web
abstract
How can we cull the facts we need from the overwhelming mass of information and misinformation that is the Web? The TextRunner extraction engine represents one approach, in which people pose keyword queries or simple questions and TextRunner returns concise answers based on tuples extracted from Web text. Unfortunately, the results returned by engines such as TextRunner include both informative facts (e.g., “the FDA banned ephedra”) and less useful statements (e.g., “the FDA banned products”). This paper therefore investigates filtering TextRunner results to enable people to better focus on interesting assertions. We first develop three distinct models of what assertions are likely to be interesting in response to a query. We then fully operationalize each of these models as a filter over TextRunner results. Finally, we develop a more sophisticated filter that combines the different models using relevance feedback. In a study of human ratings of the interestingness of TextRunner assertions, we show that our approach substantially enhances the quality of TextRunner results. Our best filter raises the fraction of interesting results in the top thirty from 41.6 % to 64.1%.
Thomas Lin, Oren Etzioni, James Fogarty
CIKM2
2009 Lemmatic Machine Translation
Stephen Soderland, Christopher Lim, Mausam, Oren Etzioni, Jonathan Pool
MTSummit5
2009 Unsupervised Methods for Determining Object and Relation Synonyms on the Web
abstract
The task of identifying synonymous relations and objects, or synonym resolution, is critical for high-quality information extraction. This paper investigates synonym resolution in the context of unsupervised information extraction, where neither hand-tagged training examples nor domain knowledge is available. The paper presents a scalable, fully-implemented system that runs in O(KN log N) time in the number of extractions, N, and the maximum number of synonyms per word, K. The system, called Resolver , introduces a probabilistic relational model for predicting whether two strings are co-referential based on the similarity of the assertions containing them. On a set of two million assertions extracted from the Web, Resolver resolves objects with 78% precision and 68% recall, and resolves relations with 90% precision and 35% recall. Several variations of resolver's probabilistic model are explored, and experiments demonstrate that under appropriate conditions these variations can improve F1 by 5%. An extension to the basic Resolver system allows it to handle polysemous names with 97% precision and 95% recall on a data set from the TREC corpus.
Alexander Yates, Oren Etzioni
J. Artif. Intell. Res.2
2008 The Tradeoffs Between Open and Traditional Relation Extraction
Michele Banko, Oren Etzioni
ACL2
2008 It's a Contradiction - no, it's not: A Case Study using Functional Relations
Alan Ritter, Stephen Soderland, Doug Downey, Oren Etzioni
EMNLP4
2008 Scaling Textual Inference to the Web
Stefan Schoenmackers, Oren Etzioni, Daniel S. Weld
EMNLP2
2008 Look Ma, No Hands: Analyzing the Monotonic Feature Abstraction for Text Classification
abstract
Is accurate classification possible in the absence of hand-labeled data? This paper introduces the Monotonic Feature (MF) abstraction--where the probability of class membership increases monotonically with the MF's value. The paper proves that when an MF is given, PAC learning is possible with no hand-labeled data under certain assumptions. We argue that MFs arise naturally in a broad range of textual classification applications. On the classic "20 Newsgroups" data set, a learner given an MF and unlabeled data achieves classification accuracy equal to that of a state-of-the-art semi-supervised learner relying on 160 hand-labeled examples. Even when MFs are not given as input, their presence or absence can be determined from a small amount of hand-labeled data, which yields a new semi-supervised learning method that reduces error by 15% on the 20 Newsgroups data.
Doug Downey, Oren Etzioni
NIPS2
2008 Machine reading at web scale
abstract
No abstract available.
Oren Etzioni
WSDM1
2007 Sparse Information Extraction: Unsupervised Language Models to the Rescue
Doug Downey, Stefan Schoenmackers, Oren Etzioni
ACL3
2007 Structured Querying of Web Text Data: A Technical Challenge
Michael J. Cafarella, Christopher Ré, Dan Suciu, Oren Etzioni
CIDR4
2007 Open Information Extraction from the Web
Michele Banko, Michael J. Cafarella, Stephen Soderland, Matthew Broadhead, Oren Etzioni
IJCAI5
2007 Locating Complex Named Entities in Web Text
Doug Downey, Matthew Broadhead, Oren Etzioni
IJCAI3
2007 Strategies for lifelong knowledge extraction from the web
abstract
The increasing availability of electronic text has made it possible to acquire information using a variety of techniques that leverage the expertise of both humans and machines. In particular, the field of Information Extraction (IE), in which knowledge is extracted automatically from text, has shown promise for large-scale knowledge acquisition. While IE systems can uncover assertions about individual entities with an increasing level of sophistication,alltext understanding -- the formation of a coherent theory from a textual corpus -- involves representation and learning abilities not currently achievable by today's IE systems. Compared to individual relational assertions outputted by IE systems, a theory includes coherent knowledge of abstract concepts and the relationships among them. We believe that the ability to fully discover the richness of knowledge present within large, unstructured and heterogeneous corpora will require a lifelong learning process in which earlier learned knowledge is used to guide subsequent learning. This paper introduces Alice, a lifelong learning agent whose goal is to automatically discovera collection of concepts, facts and generalizations that describe a particular topic of interest directly from a large volume of Web text. Building upon recent advances in unsupervised information extraction, we demonstrate that Alice can iteratively discover new concepts and compose general domain knowledge with a precision of 78%.
Michele Banko, Oren Etzioni
K-CAP2
2007 Machine reading of web text
abstract
Article Share on Machine reading of web text Author: Oren Etzioni University of Washington, Seattle, WA University of Washington, Seattle, WAView Profile Authors Info & Claims K-CAP '07: Proceedings of the 4th international conference on Knowledge captureOctober 2007 Pages 1–4https://doi.org/10.1145/1298406.1298407Published:28 October 2007Publication History 2citation320DownloadsMetricsTotal Citations2Total Downloads320Last 12 Months5Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Oren Etzioni
K-CAP1
2007 Lexical translation with application to image searching on the web
Oren Etzioni, Kobi Reiter, Stephen Soderland, Marcus Sammer
MTSummit1
2007 Unsupervised Resolution of Objects and Relations on the Web
Alexander Yates, Oren Etzioni
HLT-NAACL2
2007 Navigating Extracted Data with Schema Discovery
Michael J. Cafarella, Dan Suciu, Oren Etzioni
WebDB3
2006 Machine Reading
Oren Etzioni, Michele Banko, Michael J. Cafarella
AAAI1
2006 Detecting Parser Errors Using Web-based Semantic Filters
Alexander Yates, Stefan Schoenmackers, Oren Etzioni
EMNLP3
2006 Self-supervised Relation Extraction from the Web
Ronen Feldman, Binyamin Rosenfeld, Stephen Soderland, Oren Etzioni
ISMIS4
2005 A Probabilistic Model of Redundancy in Information Extraction
Doug Downey, Oren Etzioni, Stephen Soderland
IJCAI2
2005 A search engine for natural language applications
abstract
Many modern natural language-processing applications utilize search engines to locate large numbers of Web documents or to compute statistics over the Web corpus. Yet Web search engines are designed and optimized for simple human queries---they are not well suited to support such applications. As a result, these applications are forced to issue millions of successive queries resulting in unnecessary search engine load and in slow applications with limited scalability.In response, this paper introduces the Bindings Engine (BE), which supports queries containing typed variables and string-processing functions. For example, in response to the query "powerful ‹noun›" BE will return all the nouns in its index that immediately follow the word "powerful", sorted by frequency. In response to the query "Cities such as ProperNoun(Head(‹NounPhrase›))", BE will return a list of proper nouns likely to be city names.BE's novel neighborhood index enables it to do so with O(k) random disk seeks and O(k) serial disk reads, where k is the number of non-variable terms in its query. As a result, BE can yield several orders of magnitude speedup for large-scale language-processing applications. The main cost is a modest increase in space to store the index. We report on experiments validating these claims, and analyze how BE's space-time tradeoff scales with the size of its index and the number of variable types. Finally, we describe how a BE-based application extracts thousands of facts from the Web at interactive speeds in response to simple user queries.
Michael J. Cafarella, Oren Etzioni
WWW2
2005 Unsupervised named-entity extraction from the Web: An experimental study
Oren Etzioni, Michael J. Cafarella, Doug Downey, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates
Artif. Intell.1
2004 Methods for Domain-Independent Information Extraction from the Web: An Experimental Comparison
Oren Etzioni, Michael J. Cafarella, Doug Downey, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates
AAAI1
2004 PRECISE on ATIS: Semantic Tractability and Experimental Results
Ana-Maria Popescu, Alex Armanasu, Oren Etzioni, David Ko, Alexander Yates
AAAI3
2004 Modern Natural Language Interfaces to Databases: Composing Statistical Parsing with Semantic Tractability
Ana-Maria Popescu, Alex Armanasu, Oren Etzioni, David Ko, Alexander Yates
COLING3
2004 The Specification of Agent Behavior by Ordinary People: A Case Study
Luke K. McDowell, Oren Etzioni, Alon Y. Halevy
ISWC2
2004 Web-scale information extraction in knowitall: (preliminary results)
abstract
Manually querying search engines in order to accumulate a large bodyof factual information is a tedious, error-prone process of piecemealsearch. Search engines retrieve and rank potentially relevantdocuments for human perusal, but do not extract facts, assessconfidence, or fuse information from multiple documents. This paperintroduces KnowItAll, a system that aims to automate the tedious process ofextracting large collections of facts from the web in an autonomous,domain-independent, and scalable manner.The paper describes preliminary experiments in which an instance of KnowItAll, running for four days on a single machine, was able to automatically extract 54,753 facts. KnowItAll associates a probability with each fact enabling it to trade off precision and recall. The paper analyzes KnowItAll's architecture and reports on lessons learned for the design of large-scale information extraction systems.
Oren Etzioni, Michael J. Cafarella, Doug Downey, Stanley Kok, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates
WWW1
2004 Semantic email
abstract
This paper investigates how the vision of the Semantic Web can be carried overto the realm of email. We introduce a general notion of semantice mail, in which an email message consists of an RDF query or update coupled with corresponding explanatory text. Semantic email opens the door to a wide range of automated, email-mediated applications with formally guaranteed properties. In particular, this paper introduces a broad class of semantic email processes. For example consider the process of sending an email to a program committee asking who will attend the PC dinner automatically collecting the responses and tallying them up. We define bothlogical and decision-theoretic models where an email process ismodeled as a set of updates to a data set on which we specify goals via certain constraints or utilities. We then describe a set ofinference problems that arise while trying to satisfy these goals and analyze their computational tractability. In particular weshow that for the logical model it is possible to automatically infer which email responses are acceptable w.r.t. a set ofconstraints in polynomial time and for the decision-theoreticmodel it is possible to compute the optimal message-handling policy in polynomial time. Finally we discuss our publicly available implementation of semantic email and outline research challenges inthis realm.
Luke K. McDowell, Oren Etzioni, Alon Y. Halevy, Henry M. Levy
WWW2
2004 Semantic email: theory and applications
Luke K. McDowell, Oren Etzioni, Alon Y. Halevy
J. Web Semant.2
2003 Crossing the Structure Chasm
Alon Y. Halevy, Oren Etzioni, AnHai Doan, Zachary G. Ives, Jayant Madhavan, Luke K. McDowell, Igor Tatarinov
CIDR2
2003 Automatically Personalizing User Interfaces
Daniel S. Weld, Corin R. Anderson, Pedro M. Domingos, Oren Etzioni, Krzysztof Z. Gajos, Tessa A. Lau, Steven A. Wolfman
IJCAI4
2003 Towards a theory of natural language interfaces to databases
abstract
The need for Natural Language Interfaces to databases (NLIs) has become increasingly acute as more and more people access information through their web browsers, PDAs, and cell phones. Yet NLIs are only usable if they map natural language questions to SQL queries correctly. As Schneiderman and Norman have argued, people are unwilling to trade reliable and predictable user interfaces for intelligent but unreliable ones. In this paper, we introduce a theoretical framework for reliable NLIs, which is the foundation for the fully implemented Precise NLI. We prove that, for a broad class of semantically tractable natural language questions, Precise is guaranteed to map each question to the corresponding SQL query. We report on experiments testing Precise on several hundred questions drawn from user studies over three benchmark databases. We find that over 80% of the questions are semantically tractable questions, which Precise answers correctly. Precise automatically recognizes the 20% of questions that it cannot handle, and requests a paraphrase. Finally, we show that Precise compares favorably with Mooney's learning NLI and with Microsoft's English Query product
Ana-Maria Popescu, Oren Etzioni, Henry A. Kautz
IUI2
2003 Towards a theory of natural language interfaces to databases
abstract
The need for Natural Language Interfaces (NLIs) to databases has become increasingly acute as more nontechnical people access information through their web browsers, PDAs and cell phones. Yet NLIs are only usable if they map natural language questions to SQL queries correctly. We introduce the Precise NLI [2], which reduces the semantic interpretation challenge in NLIs to a graph matching problem. Precise uses the max-flow algorithm to efficiently solve this problem. Each max-flow solution corresponds to a possible semantic interpretation of the sentence. precise collects max-flow solutions, discards the solutions that do not obey syntactic constraints and retains the rest as the basis for generating SQL queries corresponding to the question q. The syntactic information is extracted from the parse tree corresponding to the given question which is computed by a statistical parser [1]. For a broad, well-defined class of semantically tractable natural language questions, Precise is guaranteed to map each question to the corresponding SQL querySemantically tractable questions correspond to a natural, domain-independent subset of English that can be efficiently and accurately interpreted as nonrecursive Datalog clauses. Precise is transportable to arbitrary databases, such as the Restaurants,Jobs and Geography databases used in our implementation. Examples of semantically tractable questions include: "What Chinese restaurants with a 3.5 rating are in Seattle?", "What are the areas of US states with large populations?", "What jobs require 4 years of experience and desire a B.S.CS degree?".Given a question which is not semantically tractable, Precise recognizes it as such and informs the user that it cannot answer it.Given a semantically tractable question, Precise computes the set of non-equivalent SQL interpretations corresponding to the question. If a unique such SQL interpretation exists, Precise outputs it together with the corresponding result set obtained by querying the current database. If the set contains more than one SQL interpretation, the natural language question is ambiguous in the context of the current database. In this case, Precise asks for the user's help in determining which interpretation is the correct one.Our experiments have shown that Precise has high coverage and accuracy over common English questions. In future work, we plan to explore increasingly broad classes of questions and include Precise as a module in a full-fledged dialog system. An important direction for future work is helping users understand the types of questions Precise cannot handle via dialog, enabling them to build an accurate mental model of the system and its capabilities. Also, our own group's work on the EXACT natural language interface [3] builds on Precise and on the underlying theoretical framework. EXACT composes an extended version of Precise with a sound and complete planner to develop a powerful and provably reliable interface to household appliances
Ana-Maria Popescu, Oren Etzioni, Henry A. Kautz
IUI2
2003 A reliable natural language interface to household appliances
abstract
As household appliances grow in complexity and sophistication, they become harder and harder to use, particularly because of their tiny display screens and limited keyboards. This paper describes a strategy for building natural language interfaces to appliances that circumvents these problems. Our approach leverages decades of research on planning and natural language interfaces to databases by reducing the appliance problem to the database problem; the reduction provably maintains desirable properties of the database interface. The paper goes on to describe the implementation and evaluation of the EXACT interface to appliances, which is based on this reduction. EXACT maps each English user request to an SQL query, which is transformed to create a PDDL goal, and uses the Blackbox planner [13] to map the planning problem to a sequence of appliance commands that satisfy the original request. Both theoretical arguments and experimental evaluation show that EXACT is highly reliable
Alexander Yates, Oren Etzioni, Daniel S. Weld
IUI2
2003 To buy or not to buy: mining airfare data to minimize ticket purchase price
abstract
As product prices become increasingly available on the World Wide Web, consumers attempt to understand how corporations vary these prices over time. However, corporations change prices based on proprietary algorithms and hidden variables (e.g., the number of unsold seats on a flight). Is it possible to develop data mining techniques that will enable consumers to predict price changes under these conditions?This paper reports on a pilot study in the domain of airline ticket prices where we recorded over 12,000 price observations over a 41 day period. When trained on this data, Hamlet --- our multi-strategy data mining algorithm --- generated a predictive model that saved 341 simulated passengers $198,074 by advising them when to buy and when to postpone ticket purchases. Remarkably, a clairvoyant algorithm with complete knowledge of future prices could save at most $320,572 in our simulation, thus HAMLET's savings were 61.8% of optimal. The algorithm's savings of $198,074 represents an average savings of 23.8% for the 341 passengers for whom savings are possible. Overall, HAMLET saved 4.4% of the ticket price averaged over the entire set of 4,488 simulated passengers. Our pilot study suggests that mining of price data available over the web has the potential to save consumers substantial sums of money per annum.
Oren Etzioni, Rattapoom Tuchinda, Craig A. Knoblock, Alexander Yates
KDD1
2003 Mangrove: Enticing Ordinary People onto the Semantic Web via Instant Gratification
Luke K. McDowell, Oren Etzioni, Steve D. Gribble, Alon Y. Halevy, Henry M. Levy, William Pentney, Stani Vlasseva
ISWC2
2003 Semantic Email: Adding Lightweight Data Manipulation Capabilities to the Email Habitat
Oren Etzioni, Alon Y. Halevy, Henry M. Levy, Luke K. McDowell
WebDB1
2001 Scaling question answering to the Web
abstract
The wealth of information on the web makes it an attractive resource for seeking quick answers to simple, factual questions such as "who was the first American in space?" or "what is the second tallest mountain in the world?" Yet today's most advanced web search services (e.g., Google and AskJeeves) make it surprisingly tedious to locate answers to such questions. In this paper, we extend question-answering techniques, first studied in the information retrieval literature, to the web and experimentally evaluate their performance. First we introduce MULDER, which we believe to be the first general-purpose, fully-automated question-answering system available on the web. Second, we describe MULDER's architecture, which relies on multiple search-engine queries, natural-language parsing, and a novel voting procedure to yield reliable answers coupled with high recall. Finally, we compare MULDER's performance to that of Google and AskJeeves on questions drawn from the TREC-8 question track. We find that MULDER's recall is more than a factor of three higher than that of AskJeeves. In addition, we find that Google requires 6.6 times as much user effort to achieve the same level of recall as MULDER. 1.
Cody C. T. Kwok, Oren Etzioni, Daniel S. Weld
WWW2
2001 Scaling question answering to the web
abstract
The wealth of information on the web makes it an attractive resource for seeking quick answers to simple, factual questions such as “who was the first American in space?” or “what is the second tallest mountain in the world?” Yet today's most advanced web search services (e.g., Google and AskJeeves) make it surprisingly tedious to locate answers to such questions. In this paper, we extend question-answering techniques, first studied in the information retrieval literature, to the web and experimentally evaluate their performance.First we introduce Mulder, which we believe to be the first general-purpose, fully-automated question-answering system available on the web. Second, we describe Mulder's architecture, which relies on multiple search-engine queries, natural-language parsing, and a novel voting procedure to yield reliable answers coupled with high recall. Finally, we compare Mulder's performance to that of Google and AskJeeves on questions drawn from the TREC-8 question answering track. We find that Mulder's recall is more than a factor of three higher than that of AskJeeves. In addition, we find that Google requires 6.6 times as much user effort to achieve the same level of recall as Mulder.
Cody C. T. Kwok, Oren Etzioni, Daniel S. Weld
ACM Trans. Inf. Syst.2
2000 Towards adaptive Web sites: Conceptual framework and case study
Mike Perkowitz, Oren Etzioni
Artif. Intell.2
2000 Query routing for Web search engines: architecture and experiments
Atsushi Sugiura, Oren Etzioni
Comput. Networks2
2000 Optimal Information Gathering on the Internet with Time and Cost Constraints
abstract
The World Wide Web provides access to vast amounts of information, but content providers are considering charging for the information and services they supply. Thus the consumer may face the problem of balancing the benefit of asking for information against the cost (in terms of both money and time) of acquiring it. We study information-gathering strategies that maximize the expected value to the consumer. In our model there is a single information request, which has a known benefit to the consumer. To satisfy the request, queries can be sent simultaneously or in sequence to any of a finite set of independent information sources. For each source we know the monetary cost of making the query, the amount of time it will take, and the probability that the source will be able to provide the requested information. A policy specifies which sources to contact at which times, and the expected value of the policy can be defined as some function of the likelihood that the policy will yield an answer, the expected benefit, and the monetary cost and time delay associated with executing the policy. The problem is to find an expected-value-maximizing policy. We explore four variants of the objective function V: (i) V consists only of the benefit term subject to threshold constraints on both total cost and total elapsed time, (ii) V is linear in the expected total cost of the policy subject to the constraint that the total elapsed time never exceeds somedeadline, (iii) V is linear in the expected total elapsed time subject to the constraint that the total cost never exceeds some threshold, and (iv) V is linear in the expected total monetary cost and the expected time delay of the policy. The problems of devising an optimal querying policy for all four variants and approximating an optimal querying policy for variants (iii) and (iv) are shown to be NP-hard. For (i), and with a mild simplifying assumption for (iii), we give a fully polynomial time approximation scheme. For (ii), we consider batched querying policies, and design an O(n 2 ) time approximation algorithm with ratio $\frac{1}{2}$ and a polynomial time approximation scheme for optimal single-batch policies, and an O(kn 2 ) time approximation algorithm with ratio $\frac{1}{5}$ for optimal k-batch policies.
Oren Etzioni, Steve Hanks, Tao Jiang 0001, Omid Madani
SIAM J. Comput.1
1999 Adaptive Web Sites: Conceptual Cluster Mining
Mike Perkowitz, Oren Etzioni
IJCAI2
1999 Towards Adaptive Web Sites: Conceptual Framework and Case Study
Mike Perkowitz, Oren Etzioni
Comput. Networks2
1999 Grouper: A Dynamic Clustering Interface to Web Search Results
Oren Zamir, Oren Etzioni
Comput. Networks2
1998 Web Document Clustering: A Feasibility Demonstration
abstract
Users of Web search engines are often forced to sift through the long ordered list of document returned by the engines. The IR community has explored document clustering as an alternative method of organizing retrieval results, but clustering has yet to be deployed on the major search engines. The paper articulates the unique requirements of Web document clustering and reports on the first evaluation of clustering methods in this domain. A key requirement is that the methods create their clusters based on the short snippets returned by Web search engines. Surprisingly, we find that clusters based on snippets are almost as good as clusters created using the full text of Web documents. To satisfy the stringent requirements of the Web domain, we introduce an incremental, linear time (in the document collection size) algorithm called Suffix Tree Clustering (STC). which creates clusters based on phrases shared between documents. We show that STC is faster than standard clustering methods in this domain, and argue that Web document clustering via STC is both feasible and potentially beneficial.
Oren Zamir, Oren Etzioni
SIGIR2
1997 Adaptive Web Sites: an AI Challenge
Mike Perkowitz, Oren Etzioni
IJCAI (1)2
1997 Fast and Intuitive Clustering of Web Documents
Oren Zamir, Oren Etzioni, Omid Madani, Richard M. Karp
KDD2
1997 Sound and Efficient Closed-World Reasoning for Planning
Oren Etzioni, Keith Golden, Daniel S. Weld
Artif. Intell.1
1997 Dynamic Reference Sifting: A Case Study in the Homepage Domain
Jonathan Shakes, Marc Langheinrich, Oren Etzioni
Comput. Networks3
1997 Learning to Understand Information on the Internet: An Example-Based Approach
Mike Perkowitz, Robert B. Doorenbos, Oren Etzioni, Daniel S. Weld
J. Intell. Inf. Syst.3
1996 Efficient Information Gathering on the Internet (extended abstract)
abstract
The Internet offers unprecedented access to information. At present most of this information is free, but information providers ore likely to start charging for their services in the near future. With that in mind this paper introduces the following information access problem: given a collection of n information sources, each of which has a known time delay, dollar cost and probability of providing the needed information, find an optimal schedule for querying the information sources. We study several variants of the problem which differ in the definition of an optimal schedule. We first consider a cost model in which the problem is to minimize the expected total cost (monetary and time) of the schedule, subject to the requirement that the schedule may terminate only when the query has been answered or all sources have been queried unsuccessfully. We develop an approximation algorithm for this problem and for an extension of the problem in which more than a single item of information is being sought. We then develop approximation algorithms for a reward model in which a constant reward is earned if the information is successfully provided, and we seek the schedule with the maximum expected difference between the reward and a measure of cost. The monetary and time costs may either appear in the cost measure or be constrained not to exceed a fixed upper bound; these options give rise to four different variants of the reward model.
Oren Etzioni, Steve Hanks, Tao Jiang 0001, Richard M. Karp, Omid Madani, Orli Waarts
FOCS1
1996 Scaling Up Goal Recognition
Neal Lesh, Oren Etzioni
KR2
1995 A Sound and Fast Goal Recognizer
Neal Lesh, Oren Etzioni
IJCAI2
1995 Category Translation: Learning to Understand Information on the Internet
Mike Perkowitz, Oren Etzioni
IJCAI (1)2
1994 Learning About Software Errors Via Systematic Experimentation
Terrance Goan, Oren Etzioni
AAAI2
1994 Omnipotence Without Omniscience: Efficient Sensor Management for Planning
Keith Golden, Oren Etzioni, Daniel S. Weld
AAAI2
1994 Database Learning for Software Agents
Mike Perkowitz, Oren Etzioni
AAAI2
1994 Learning Decision Lists Using Homogeneous Rules
Richard B. Segal, Oren Etzioni
AAAI2
1994 The First Law of Robotics (A Call to Arms)
Daniel S. Weld, Oren Etzioni
AAAI2
1994 Tractable Closed World Reasoning with Updates
Oren Etzioni, Keith Golden, Daniel S. Weld
KR1
1994 Statistical Methods for Analyzing Speedup Learning Experiments
Oren Etzioni, Ruth Etzioni
Mach. Learn.1
1993 Acquiring Search-Control Knowledge Via Static Analysis
Oren Etzioni
Artif. Intell.1
1993 A Structural Theory of Explanation-Based Learning
Oren Etzioni
Artif. Intell.1
1992 An Asymptotic Analysis of Speedup Learning
Oren Etzioni
ML1
1992 Why EBL Produces Overly-Specific Knowledge: A Critique of the PRODIGY Approaches
Oren Etzioni, Steven Minton
ML1
1992 DYNAMIC: A New Role for Training Problems in EBL
M. Alicia Pérez, Oren Etzioni
ML2
1992 An Approach to Planning with Incomplete Information
Oren Etzioni, Steve Hanks, Daniel S. Weld, Denise Draper, Neal Lesh, Mike Williamson
KR1
1991 STATIC: A Problem-Space Compiler for PRODIGY
Oren Etzioni
AAAI1
1991 Integrating Abstraction and Explanation-Based Learning in PRODIGY
Craig A. Knoblock, Steven Minton, Oren Etzioni
AAAI3
1991 Integrating Efficient Model-Learning and Problem-Solving Algorithms in Permutation Environments
Prasad Chalasani, Oren Etzioni, John Mount
KR2
1991 Embedding Decision-Analytic Control in a Learning Architecture
Oren Etzioni
Artif. Intell.1
1990 Why PRODIGY/EBL Works
Oren Etzioni
AAAI1
1989 Tractable Decision-Analytic Control
Oren Etzioni
KR1
1989 Explanation-Based Learning: A Problem Solving Perspective
Steven Minton, Jaime G. Carbonell, Craig A. Knoblock, Daniel Kuokka, Oren Etzioni, Yolanda Gil
Artif. Intell.5
1988 Hypothesis Filtering: A Practical Approach to Reliable Learning
Oren Etzioni
ML1