VLDB 2026 Research / reviewers in the wild / expert
Eugene Charniak
dblp:c/ECharniak
· DBLP profile ↗
79ranked-venue papers
31as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 78 · 31 first-authorGraphics, computer vision, multimedia, augmented reality and games · 19 · 14 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
46 papers |
Information extraction and text analysis · 58% Language models and text generation · 13% Probabilistic and Bayesian machine learning · 8% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 92% Programming languages and type systems · 8% | |
| Theoretical computer science
6 papers |
Automata and formal languages · 85% Information theory · 13% Automated reasoning and model checking · 1% |
Topics — the 30 heaviest of 73, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
parsing |
0.5 | 2 | 2016 | Parsing as Language Modeling · EMNLP 2016 Syntactic Parse Fusion · EMNLP 2015 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.3 | 9 | 2006 | A Look at Parsing and Its Applications · AAAI 2006 Coarse-to-Fine n-Best Parsing and MaxEnt Discriminative Reranking · ACL 2005 A TAG-based noisy-channel model of speech repairs · ACL 2004 |
Natural language and speech › Information extraction and text analysis › sentiment analysis
irony detection |
0.2 | 1 | 2015 | Sparse, Contextually Informed Models for Irony Detection: Exploiting User Communities, Entities and Sentiment · ACL (1) 2015 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.2 | 1 | 2015 | Sparse, Contextually Informed Models for Irony Detection: Exploiting User Communities, Entities and Sentiment · ACL (1) 2015 |
Natural language and speech › Information extraction and text analysis › dialogue analysis
conversation analysis |
0.2 | 2 | 2011 | Disentangling Chat with Local Coherence Models · ACL 2011 You Talking to Me? A Corpus and Algorithm for Conversation Disentanglement · ACL 2008 |
Natural language and speech › Question answering and dialogue systems › multi-party dialogue
dialogue disentanglement |
0.2 | 2 | 2011 | Disentangling Chat with Local Coherence Models · ACL 2011 You Talking to Me? A Corpus and Algorithm for Conversation Disentanglement · ACL 2008 |
Machine learning › Generative modeling
generative model |
0.2 | 1 | 2013 | A Generative Joint, Additive, Sequential Model of Topics and Speech Acts in Patient-Doctor Communication · EMNLP 2013 |
Natural language and speech › Information extraction and text analysis › dialogue analysis
speech act recognition |
0.2 | 1 | 2013 | A Generative Joint, Additive, Sequential Model of Topics and Speech Acts in Patient-Doctor Communication · EMNLP 2013 |
Natural language and speech › Information extraction and text analysis
topic model |
0.2 | 1 | 2013 | A Generative Joint, Additive, Sequential Model of Topics and Speech Acts in Patient-Doctor Communication · EMNLP 2013 |
Natural language and speech › Information extraction and text analysis › lexical semantics
word sense induction |
0.2 | 1 | 2013 | Naive Bayes Word Sense Induction · EMNLP 2013 |
Automata and formal languages
tree adjoining grammar |
0.2 | 1 | 2013 | A Context Free TAG Variant · ACL (1) 2013 |
Automata and formal languages
parsing |
0.1 | 1 | 2010 | Top-Down Nearly-Context-Sensitive Parsing · EMNLP 2010 |
Automata and formal languages
parsing algorithms |
0.1 | 1 | 2010 | Top-Down Nearly-Context-Sensitive Parsing · EMNLP 2010 |
Automata and formal languages › parsing algorithms
top-down parsing |
0.1 | 1 | 2010 | Top-Down Nearly-Context-Sensitive Parsing · EMNLP 2010 |
Natural language and speech › Language models and text generation
language modeling |
0.1 | 1 | 2016 | Parsing as Language Modeling · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
statistical parsing |
0.1 | 2 | 2006 | Reranking and Self-Training for Parser Adaptation · ACL 2006 Parsing and Disfluency Placement · EMNLP 2002 |
Compilers and program optimization › parsing
constituency parsing |
0.1 | 1 | 2015 | Syntactic Parse Fusion · EMNLP 2015 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.1 | 1 | 2006 | Recognizing disfluencies in conversational speech · IEEE Trans. Speech Audio Process. 2006 |
Natural language and speech › Information extraction and text analysis › dialogue analysis
disfluency detection |
0.1 | 1 | 2006 | Recognizing disfluencies in conversational speech · IEEE Trans. Speech Audio Process. 2006 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
parser adaptation |
0.1 | 1 | 2006 | Reranking and Self-Training for Parser Adaptation · ACL 2006 |
Natural language and speech › Language models and text generation › reranking
discriminative reranking |
0.1 | 1 | 2005 | Coarse-to-Fine n-Best Parsing and MaxEnt Discriminative Reranking · ACL 2005 |
Natural language and speech › Language models and text generation › text summarization
sentence compression |
0.1 | 1 | 2005 | Supervised and Unsupervised Learning for Sentence Compression · ACL 2005 |
Natural language and speech › Language models and text generation
text summarization |
0.1 | 1 | 2005 | Supervised and Unsupervised Learning for Sentence Compression · ACL 2005 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › bayesian network › bayesian network classifiers
naive bayes |
0.0 | 1 | 2013 | Naive Bayes Word Sense Induction · EMNLP 2013 |
Programming languages and type systems
grammar formalisms |
0.0 | 1 | 2013 | A Context Free TAG Variant · ACL (1) 2013 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › grammar-based parsing
tree adjoining grammar parsing |
0.0 | 1 | 2004 | A TAG-based noisy-channel model of speech repairs · ACL 2004 |
Natural language and speech › Language models and text generation › language modeling
statistical language modeling |
0.0 | 2 | 2001 | Immediate-Head Parsing for Language Models · ACL 2001 Context-Sensitive Statistics For Improved Grammatical Language Models · AAAI 1994 |
Information theory › information measures
entropy |
0.0 | 1 | 2003 | Variation of Entropy and Parse Trees of Sentences as a Function of the Sentence Number · EMNLP 2003 |
Natural language and speech › Speech recognition and synthesis › spoken language understanding
speech parsing |
0.0 | 1 | 2002 | Parsing and Disfluency Placement · EMNLP 2002 |
Information theory › information measures › entropy › source entropy
entropy of natural language |
0.0 | 1 | 2002 | Entropy Rate Constancy in Text · ACL 2002 |
Methods — techniques the papers use, named apart from their topics
log-linear model · 0.5language modeling · 0.5clustering · 0.4formal language theory · 0.3sparse modeling · 0.2sentiment features · 0.2self-training · 0.2n-best list combination · 0.2entity features · 0.2discriminative reranking · 0.2naive bayes · 0.2joint modeling · 0.2extended naive bayes · 0.2noisy channel model · 0.2nearly-context-sensitive grammar · 0.1entropy analysis · 0.0entropy measurement · 0.0heuristic search · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Parsing as Language Modeling
Do Kook Choe, Eugene Charniak |
EMNLP | 2 |
| 2015 | Sparse, Contextually Informed Models for Irony Detection: Exploiting User Communities, Entities and SentimentabstractByron C. Wallace, Do Kook Choe, Eugene Charniak. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Byron C. Wallace, Do Kook Choe, Eugene Charniak |
ACL (1) | 3 |
| 2015 | Syntactic Parse FusionabstractModel combination techniques have consistently shown state-of-the-art performance across multiple tasks, including syntactic parsing.However, they dramatically increase runtime and can be difficult to employ in practice.We demonstrate that applying constituency model combination techniques to n-best lists instead of n different parsers results in significant parsing accuracy improvements.Parses are weighted by their probabilities and combined using an adapted version of Sagae and Lavie (2006).These accuracy gains come with marginal computational costs and are obtained on top of existing parsing techniques such as discriminative reranking and self-training, resulting in state-of-the-art accuracy: 92.6% on WSJ section 23.On out-of-domain corpora, accuracy is improved by 0.4% on average.We empirically confirm that six well-known n-best parsers benefit from the proposed methods across six domains. Do Kook Choe, David McClosky, Eugene Charniak |
EMNLP | 3 |
| 2015 | A Hybrid Generative/Discriminative Approach To Citation PredictionabstractText documents of varying nature (e.g., summary documents written by analysts or published, scientific papers) often cite others as a means of providing evidence to support a claim, attributing credit, or referring the reader to related work.We address the problem of predicting a document's cited sources by introducing a novel, discriminative approach which combines a content-based generative model (LDA) with author-based features.Further, our classifier is able to learn the importance and quality of each topic within our corpus -which can be useful beyond this task -and preliminary results suggest its metric is competitive with other standard metrics (Topic Coherence).Our flagship system, Logit-Expanded, provides state-of-the-art performance on the largest corpus ever used for this task. Chris Tanner, Eugene Charniak |
HLT-NAACL | 2 |
| 2014 | Identifying Differences in Physician Communication Styles with a Log-Linear Transition Component ModelabstractWe consider the task of grouping doctors with respect to communication patterns exhibited in outpatient visits. We propose a novel approach toward this end in which we model speech act transitions in conversations via a log-linear model incorporating physician specific components. We train this model over transcripts of outpatient visits annotated with speech act codes and then cluster physicians in (a transformation of) this parameter space. We find significant correlations between the induced groupings and patient survey response data comprising ratings of physician communication. Furthermore, the novel sequential component model we leverage to induce this clustering allows us to explore differences across these groups. This work demonstrates how statistical AI might be used to better understand (and ultimately improve) physician communication. Byron C. Wallace, Issa J. Dahabreh, Thomas A. Trikalinos, Michael Barton Laws, Ira B. Wilson, Eugene Charniak |
AAAI | 6 |
| 2014 | Domain-Specific Image CaptioningabstractWe present a data-driven framework for image caption generation which incorporates visual and textual features with varying degrees of spatial structure.We propose the task of domain-specific image captioning, where many relevant visual details cannot be captured by off-the-shelf general-domain entity detectors.We extract previously-written descriptions from a database and adapt them to new query images, using a joint visual and textual bag-of-words model to determine the correctness of individual words.We implement our model using a large, unlabeled dataset of women's shoes images and natural language descriptions (Berg et al., 2010).Using both automatic and human evaluations, we show that our captioning method effectively deletes inaccurate words from extracted captions while maintaining a high level of detail in the generated output. Rebecca Mason, Eugene Charniak |
CoNLL | 2 |
| 2014 | Data Driven Language Transfer HypothesesabstractLanguage transfer, the preferential second language behavior caused by similarities to the speaker's native language, requires considerable expertise to be detected by humans alone.Our goal in this work is to replace expert intervention by data-driven methods wherever possible.We define a computational methodology that produces a concise list of lexicalized syntactic patterns that are controlled for redundancy and ranked by relevancy to language transfer.We demonstrate the ability of our methodology to detect hundreds of such candidate patterns from currently available data sources, and validate the quality of the proposed patterns through classification experiments. Benjamin Swanson, Eugene Charniak |
EACL | 2 |
| 2013 | A Context Free TAG Variant
Benjamin Swanson, Elif Yamangil, Eugene Charniak, Stuart M. Shieber |
ACL (1) | 3 |
| 2013 | Naive Bayes Word Sense InductionabstractWe introduce an extended naive Bayes model for word sense induction (WSI) and apply it to a WSI task.The extended model incorporates the idea the words closer to the target word are more relevant in predicting its sense.The proposed model is very simple yet effective when evaluated on SemEval-2010 WSI data. Do Kook Choe, Eugene Charniak |
EMNLP | 2 |
| 2013 | A Generative Joint, Additive, Sequential Model of Topics and Speech Acts in Patient-Doctor CommunicationabstractWe develop a novel generative model of conversation that jointly captures both the topical content and the speech act type associated with each utterance.Our model expresses both token emission and state transition probabilities as log-linear functions of separate components corresponding to topics and speech acts (and their interactions).We apply this model to a dataset comprising annotated patient-physician visits and show that the proposed joint approach outperforms a baseline univariate model. Byron C. Wallace, Thomas A. Trikalinos, Michael Barton Laws, Ira B. Wilson, Eugene Charniak |
EMNLP | 5 |
| 2013 | Extracting the Native Language Signal for Second Language Acquisition
Benjamin Swanson, Eugene Charniak |
HLT-NAACL | 2 |
| 2012 | Apples to Oranges: Evaluating Image Annotations from Natural Language Processing Systems
Rebecca Mason, Eugene Charniak |
HLT-NAACL | 2 |
| 2011 | Disentangling Chat with Local Coherence Models
Micha Elsner, Eugene Charniak |
ACL | 2 |
| 2011 | $S^3$ - Statistical Sandhi Splitting
Abhiram Natarajan, Eugene Charniak |
IJCNLP | 2 |
| 2011 | The Brain as a Statistical Inference Engine - and You Can Tooabstractno difference in the spacing between syllables.(It is pretty boring, but the child is only subjected to two minutes of it.)After that the child is tested to see if he or she can distinguish real words from non-words.To test, either a word or a non-word is played from one of two speakers.This is not done until the child is already looking at that speaker and the word is replayed until the child looks away.The children are expected to gaze longer at the speaker that is playing a novel (non-) word than for words that they have already heard.Thus there are two testing conditions.In the first, the non-words are completely novel in the sense that the three-syllable combination did not occur in the two minutes of pretest training.On average, the children focus on the speaker 0.88 seconds longer for the novel words.The second condition is more interesting.Here the non-words are made up of sound combinations that have in fact occurred on the tape, but relatively infrequently because they consist of pieces of two different words.Here the question is not a categorical one (Have I heard this combination or not?) but a statistical one: Is this a frequent combination or is it rare?Now the focus differential is 0.83 seconds.The conclusion is that children are indeed sensitive to the statistical differences. What You See Where You Are Not LookingTry the following test.Keep your gaze on the plus sign in the following example and try to identify the letters to its left and its right. Eugene Charniak |
Comput. Linguistics | 1 |
| 2010 | Top-Down Nearly-Context-Sensitive Parsing
Eugene Charniak |
EMNLP | 1 |
| 2010 | Automatic Domain Adaptation for Parsing
David McClosky, Eugene Charniak, Mark Johnson 0001 |
HLT-NAACL | 2 |
| 2010 | Disentangling ChatabstractWhen multiple conversations occur simultaneously, a listener must decide which conversation each utterance is part of in order to interpret and respond to it appropriately. We refer to this task as disentanglement. We present a corpus of Internet Relay Chat dialogue in which the various conversations have been manually disentangled, and evaluate annotator reliability. We propose a graph-based clustering model for disentanglement, using lexical, timing, and discourse-based features. The model's predicted disentanglements are highly correlated with manual annotations. We conclude by discussing two extensions to the model, specificity tuning and conversation start detection, both of which are promising but do not currently yield practical improvements. Micha Elsner, Eugene Charniak |
Comput. Linguistics | 2 |
| 2009 | EM Works for Pronoun Anaphora Resolution
Eugene Charniak, Micha Elsner |
EACL | 1 |
| 2009 | Structured Generative Models for Unsupervised Named-Entity Clustering
Micha Elsner, Eugene Charniak, Mark Johnson 0001 |
HLT-NAACL | 2 |
| 2008 | You Talking to Me? A Corpus and Algorithm for Conversation Disentanglement
Micha Elsner, Eugene Charniak |
ACL | 2 |
| 2008 | Evaluating Unsupervised Part-of-Speech Tagging for Grammar Induction
William P. Headden III, David McClosky, Eugene Charniak |
COLING | 3 |
| 2008 | When is Self-Training Effective for Parsing?
David McClosky, Eugene Charniak, Mark Johnson 0001 |
COLING | 2 |
| 2007 | A Unified Local and Global Model for Discourse Coherence
Micha Elsner, Joseph L. Austerweil, Eugene Charniak |
HLT-NAACL | 3 |
| 2006 | A Look at Parsing and Its Applications
Matthew Lease, Eugene Charniak, Mark Johnson 0001, David McClosky |
AAAI | 2 |
| 2006 | Reranking and Self-Training for Parser AdaptationabstractStatistical parsers trained and tested on the Penn Wall Street Journal (WSJ) treebank have shown vast improvements over the last 10 years. Much of this improvement, however, is based upon an ever-increasing number of features to be trained on (typically) the WSJ treebank data. This has led to concern that such parsers may be too finely tuned to this corpus at the expense of portability to other genres. Such worries have merit. The standard "Charniak parser" checks in at a labeled precision-recall f-measure of 89.7% on the Penn WSJ test set, but only 82.9% on the test set from the Brown treebank corpus.This paper should allay these fears. In particular, we show that the reranking parser described in Charniak and Johnson (2005) improves performance of the parser on Brown to 85.2%. Furthermore, use of the self-training techniques described in (McClosky et al., 2006) raise this to 87.8% (an error reduction of 28%) again without any use of labeled Brown data. This is remarkable since training the parser and reranker on labeled Brown data achieves only 88.4%. David McClosky, Eugene Charniak, Mark Johnson 0001 |
ACL | 2 |
| 2006 | Learning Phrasal Categories
William P. Headden III, Eugene Charniak, Mark Johnson 0001 |
EMNLP | 2 |
| 2006 | SParseval: Evaluation Metrics for Parsing Speech
Brian Roark, Mary P. Harper, Eugene Charniak, Bonnie J. Dorr, Mark Johnson 0001, Jeremy G. Kahn, Yang Liu 0004, Mari Ostendorf, John Hale, Anna Krasnyanskaya, Matthew Lease, Izhak Shafran, Matthew G. Snover, Robin Stewart, Lisa Yung |
LREC | 3 |
| 2006 | Multilevel Coarse-to-Fine PCFG Parsing
Eugene Charniak, Mark Johnson 0001, Micha Elsner, Joseph L. Austerweil, David A. Ellis, Isaac Haxton, R. Shrivaths, Jeremy Moore, Michael Pozar, Theresa Vu |
HLT-NAACL | 1 |
| 2006 | Effective Self-Training for Parsing
David McClosky, Eugene Charniak, Mark Johnson 0001 |
HLT-NAACL | 2 |
| 2006 | Recognizing disfluencies in conversational speechabstractWe present a system for modeling disfluency in conversational speech: repairs, fillers, and self-interruption points (IPs). For each sentence, candidate repair analyses are generated by a stochastic tree adjoining grammar (TAG) noisy-channel model. A probabilistic syntactic language model scores the fluency of each analysis, and a maximum-entropy model selects the most likely analysis given the language model score and other features. Fillers are detected independently via a small set of deterministic rules, and IPs are detected by combining the output of repair and filler detection modules. In the recent Rich Transcription Fall 2004 (RT-04F) blind evaluation, systems competed to detect these three forms of disfluency under two input conditions: a best-case scenario of manually transcribed words and a fully automatic case of automatic speech recognition (ASR) output. For all three tasks and on both types of input, our system was the top performer in the evaluation Matthew Lease, Mark Johnson 0001, Eugene Charniak |
IEEE Trans. Speech Audio Process. | 3 |
| 2005 | Coarse-to-Fine n-Best Parsing and MaxEnt Discriminative RerankingabstractDiscriminative reranking is one method for constructing high-performance statistical parsers (Collins, 2000). A discriminative reranker requires a source of candidate parses for each sentence. This paper describes a simple yet novel method for constructing sets of 50-best parses based on a coarse-to-fine generative parser (Charniak, 2000). This method generates 50-best lists that are of substantially higher quality than previously obtainable. We used these parses as the input to a MaxEnt reranker (Johnson et al., 1999; Riezler et al., 2002) that selects the best parse from the set of parses for each sentence, obtaining an f-score of 91.0% on sentences of length 100 or less. Eugene Charniak, Mark Johnson 0001 |
ACL | 1 |
| 2005 | Supervised and Unsupervised Learning for Sentence CompressionabstractIn Statistics-Based Summarization - Step One: Sentence Compression, Knight and Marcu (Knight and Marcu, 2000) (K&M) present a noisy-channel model for sentence compression. The main difficulty in using this method is the lack of data; Knight and Marcu use a corpus of 1035 training sentences. More data is not easily available, so in addition to improving the original K&M noisy-channel model, we create unsupervised and semi-supervised models of the task. Finally, we point out problems with modeling the task in this way. They suggest areas for future research. Jenine Turner, Eugene Charniak |
ACL | 2 |
| 2005 | Parsing and its applications for conversational speechabstractThis paper provides an introduction to recent work in statistical parsing and its applications for conversational speech, with particular emphasis on the relationship between parsing and detecting speech repairs. While historically parsing and repair detection have been studied independently, we present a line of research which has spanned the boundary between the two and demonstrated the efficacy of this synergistic approach. Our presentation highlights successes to date, remaining challenges, and promising future work. Matthew Lease, Eugene Charniak, Mark Johnson 0001 |
ICASSP (5) | 2 |
| 2005 | Parsing Biomedical Literature
Matthew Lease, Eugene Charniak |
IJCNLP | 2 |
| 2004 | A TAG-based noisy-channel model of speech repairsabstractThis paper describes a noisy channel model of speech repairs, which can identify and correct repairs in speech transcripts. A syntactic parser is used as the source model, and a novel type of TAG-based transducer is the channel model. The use of TAG is motivated by the intuition that the reparandum is a "rough copy" of the repair. The model is trained and tested on the Switchboard disfluency-annotated corpus. Mark Johnson 0001, Eugene Charniak |
ACL | 2 |
| 2004 | Using the Penn Treebank to Evaluate Non-Treebank Parsers
Eric K. Ringger, Robert C. Moore, Eugene Charniak, Lucy Vanderwende, Hisami Suzuki |
LREC | 3 |
| 2004 | Sentence-Internal Prosody Does not Help Parsing the Way Punctuation Does
Michelle L. Gregory, Mark Johnson 0001, Eugene Charniak |
HLT-NAACL | 3 |
| 2003 | Variation of Entropy and Parse Trees of Sentences as a Function of the Sentence Number
Dmitriy Genzel, Eugene Charniak |
EMNLP | 2 |
| 2003 | Syntax-based language models for statistical machine translationabstractWe present a syntax-based language model for use in noisy-channel machine translation. In particular, a language model based upon that described in (Cha01) is combined with the syntax based translation-model described in (YK01). The resulting system was used to translate 347 sentences from Chinese to English and compared with the results of an IBM-model-4-based system, as well as that of (YK02), all trained on the same data. The translations were sorted into four groups: good/bad syntax crossed with good/bad meaning. While the total number of translations that preserved meaning were the same for (YK02) and the syntax-based system (and both higher than the IBM-model-4-based system), the syntax based system had 45% more translations that also had good syntax than did (YK02) (and approximately 70% more than IBM Model 4). The number of translations that did not preserve meaning, but at least had good grammar, also increased, though to less avail. Eugene Charniak, Kevin Knight, Kenji Yamada |
MTSummit | 1 |
| 2002 | Entropy Rate Constancy in TextabstractWe present a constancy rate principle governing language generation. We show that this principle implies that local measures of entropy (ignoring context) should increase with the sentence number. We demonstrate that this is indeed the case by measuring entropy in three different ways. We also show that this effect has both lexical (which words are used) and non-lexical (how the words are used) causes. Dmitriy Genzel, Eugene Charniak |
ACL | 2 |
| 2002 | Parsing and Disfluency PlacementabstractIt has been suggested that some forms of speech disfluencies, most notable interjections and parentheticals, tend to occur disproportionally at major clause boundaries [6] and thus might serve to aid parsers in establishing these boundaries. We have tested a current statistical parser [1] on Switchboard text with and without interjections and parentheticals and found that the parser performed better when not faced with these extra phenomena. This suggest that for current parsers, at least, interjection and parenthetical placement does not help in the parsing process. Don Engel, Eugene Charniak, Mark Johnson 0001 |
EMNLP | 2 |
| 2001 | Immediate-Head Parsing for Language ModelsabstractWe present two language models based upon an "immediate-head" parser --- our name for a parser that conditions all events below a constituent c upon the head of c. While all of the most accurate statistical parsers are of the immediate-head variety, no previous grammatical language model uses this technology. The perplexity for both of these models significantly improve upon the trigram model base-line as well as the best previous grammar-based language model. For the better of our two models these improvements are 24% and 14% respectively. We also suggest that improvement of the underlying parser should significantly improve the model's perplexity and that even in the near term there is a lot of potential for improvement in immediate-head language models. Eugene Charniak |
ACL | 1 |
| 2001 | Unsupervised Learning of Name Structure From Coreference Data
Eugene Charniak |
NAACL | 1 |
| 2001 | Edit Detection and Parsing for Transcribed Speech
Eugene Charniak, Mark Johnson 0001 |
NAACL | 1 |
| 1999 | Finding Parts in Very Large CorporaabstractWe present a method for extracting parts of objects from wholes (e.g. "speedometer" from "car"). Given a very large corpus our method finds part words with 55% accuracy for the top 50 words as ranked by the system. The part list could be scanned by an end-user and added to an existing ontology (such as WordNet), or used as a part of a rough semantic lexicon. Matthew Berland, Eugene Charniak |
ACL | 2 |
| 1999 | Automatic Compensation for Parser Figure-of-Merit FlawsabstractBest-first chart parsing utilises a figure of merit (FOM) to efficiently guide a parse by first attending to those edges judged better. In the past it has usually been static; this paper will show that with some extra information, a parser can compensate for FOM flaws which otherwise slow it down. Our results are faster than the prior best by a factor of 2.5; and the speedup is won with no significant decrease in parser accuracy. Don Blaheta, Eugene Charniak |
ACL | 2 |
| 1999 | Determining the specificity of nouns from text
Sharon A. Caraballo, Eugene Charniak |
EMNLP | 2 |
| 1998 | New Figures of Merit for Best-First Probabilistic Chart Parsing
Sharon A. Caraballo, Eugene Charniak |
Comput. Linguistics | 2 |
| 1996 | Figures of Merit for Best-First Probabilistic Chart Parsing
Sharon A. Caraballo, Eugene Charniak |
EMNLP | 2 |
| 1996 | Taggers for Parsers
Eugene Charniak, Glenn Carroll, John E. Adcock, Anthony R. Cassandra, Yoshihiko Gotoh, Jeremy Katz, Michael L. Littman, John McCann |
Artif. Intell. | 1 |
| 1994 | Context-Sensitive Statistics For Improved Grammatical Language Models
Eugene Charniak, Glenn Carroll |
AAAI | 1 |
| 1994 | Cost-Based Abduction and MAP Explanation
Eugene Charniak, Solomon Eyal Shimony |
Artif. Intell. | 1 |
| 1993 | Equations for Part-of-Speech Tagging
Eugene Charniak, Curtis Hendrickson, Neil Jacobson, Mike Perkowitz |
AAAI | 1 |
| 1993 | A Bayesian Model of Plan Recognition
Eugene Charniak, Robert P. Goldman |
Artif. Intell. | 1 |
| 1993 | A Language for Construction of Belief NetworksabstractA method for incrementally constructing belief networks, which are directed acyclic graph representations for probability distributions, is described. A network-construction language, FRAIL3, which is similar to a forward-chaining language using data dependencies but has additional features for specifying distributions, was developed. A particularly important feature of this language is that is allows the user to conveniently specify conditional probability matrices using stereotyped models of intercausal interaction. Using FRAIL3, one can define parmeterized classes of probabilistic models. These parameterized models make it possible to apply probabilistic reasoning to problems for which it is impractical to have a single large, static model.> Robert P. Goldman, Eugene Charniak |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1992 | Dynamic MAP Calculations for Abduction
Eugene Charniak, Eugene Santos Jr. |
AAAI | 1 |
| 1991 | A Probabilistic Model of Plan Recognition
Eugene Charniak, Robert P. Goldman |
AAAI | 1 |
| 1991 | A New Admissible Heuristic for Minimal-Cost Proofs
Eugene Charniak, Saadia Husain |
AAAI | 1 |
| 1991 | A Probabilistic Analysis of Marker-Passing Techniques for Plan-Recognition
Glenn Carroll, Eugene Charniak |
UAI | 2 |
| 1990 | Probabilistic Semantics for Cost Based Abduction
Eugene Charniak, Solomon Eyal Shimony |
AAAI | 1 |
| 1990 | Dynamic construction of belief networks
Robert P. Goldman, Eugene Charniak |
UAI | 2 |
| 1990 | A new algorithm for finding MAP assignments to belief networks
Solomon Eyal Shimony, Eugene Charniak |
UAI | 2 |
| 1989 | A Semantics for Probabilistic Quantifier-Free First-Order Languages, with Particular Application to Story Understanding
Eugene Charniak, Robert P. Goldman |
IJCAI | 1 |
| 1989 | Plan Recognition in Stories and in Life
Eugene Charniak, Robert P. Goldman |
UAI | 1 |
| 1988 | A Logic for Semantic InterpretationabstractWe propose that logic (enhanced to encode probability information) is a good way of characterizing semantic interpretation. In support of this we give a fragment of an axiomatization for word-sense disambiguation, nounphrase (and verb) reference, and case disambiguation. We describe an inference engine (Frail3) which actually takes this axiomatization and uses it to drive the semantic interpretation process. We claim three benefits from this scheme. First, the interface between semantic interpretation and pragmatics has always been problematic, since all of the above tasks in general require pragmatic inference. Now the interface is trival, since both semantic interpretation and pragmatics use the same vocabulary and inference engine. The second benefit, related to the first, is that semantic guidance of syntax is a side effect of the interpretation. The third benefit is the elegance of the semantic interpretation theory. A few simple rules capture a remarkable diversity of semantic phenomena. Eugene Charniak, Robert P. Goldman |
ACL | 1 |
| 1988 | Motivation Analysis, Abductive Unification, and Nonmonotonic Equality
Eugene Charniak |
Artif. Intell. | 1 |
| 1986 | A Neat Theory of Marker Passing
Eugene Charniak |
AAAI | 1 |
| 1986 | Time and Tense in EnglishabstractTense, temporal adverbs, and temporal connectives provide information about when events described in English sentences occur. To extract this temporal information from a sentence, it must be parsed into a semantic representation which captures the meaning of tense, temporal adverbs, and temporal connectives. Representations were developed for the basic tenses, some temporal adverbs, as well as some of the temporal connectives. Five criteria were suggested for judging these representations, and based on these criteria the representations were judged. Mary P. Harper, Eugene Charniak |
ACL | 2 |
| 1983 | The Bayesian Basis of Common Sense Medical Diagnosis
Eugene Charniak |
AAAI | 1 |
| 1982 | Word Sense and Case Slot Disambiguation
Graeme Hirst, Eugene Charniak |
AAAI | 2 |
| 1981 | Six Topics in Search of a Parser: An Overview of AI Language Research
Eugene Charniak |
IJCAI | 1 |
| 1981 | A Common Representation for Problem-Solving and Language-Comprehension Information
Eugene Charniak |
Artif. Intell. | 1 |
| 1978 | On the Use of Framed Knowledge in Language Comprehension
Eugene Charniak |
Artif. Intell. | 1 |
| 1977 | Ms. Maloprop, A Language Comprehension Program
Eugene Charniak |
IJCAI | 1 |
| 1977 | Natural Language Processing
Roger C. Schank, Eugene Charniak, Yorick Wilks, Terry Winograd, William A. Woods |
IJCAI | 2 |
| 1975 | A Partial Taxonomy of Knowledge about Actions
Eugene Charniak |
IJCAI | 1 |
| 1973 | Jack and Janet in Search of a Theory of Knowledge
Eugene Charniak |
IJCAI | 1 |
| 1969 | Computer Solution of Calculus Word Problems
Eugene Charniak |
IJCAI | 1 |