David McClosky

dblp:41/108 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
0since 2021 · last 2016
0009-0005-8897-0793ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Information extraction and text analysis · 84% Planning, search and constraint satisfaction · 14% Language models and text generation · 2%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.322015
Parsing Paraphrases with Joint Inference · ACL (1) 2015
A Look at Parsing and Its Applications · AAAI 2006
Compilers and program optimization
parsing
0.212015
Syntactic Parse Fusion · EMNLP 2015
Mathematical optimization
simultaneous inference
0.212015
Parsing Paraphrases with Joint Inference · ACL (1) 2015
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › constraint satisfaction
constraint learning
0.112012
Learning Constraints for Consistent Timeline Extraction · EMNLP-CoNLL 2012
Natural language and speech › Information extraction and text analysis
temporal information extraction
0.112012
Learning Constraints for Consistent Timeline Extraction · EMNLP-CoNLL 2012
Natural language and speech › Information extraction and text analysis › temporal information extraction
timeline extraction
0.112012
Learning Constraints for Consistent Timeline Extraction · EMNLP-CoNLL 2012
Natural language and speech › Information extraction and text analysis
event extraction
0.112011
Event Extraction as Dependency Parsing · ACL 2011
Compilers and program optimization › parsing
constituency parsing
0.112015
Syntactic Parse Fusion · EMNLP 2015
Natural language and speech › Information extraction and text analysis › syntactic parsing
parser adaptation
0.112006
Reranking and Self-Training for Parser Adaptation · ACL 2006
Natural language and speech › Information extraction and text analysis › syntactic parsing
statistical parsing
0.112006
Reranking and Self-Training for Parser Adaptation · ACL 2006
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
0.012011
Event Extraction as Dependency Parsing · ACL 2011

Methods — techniques the papers use, named apart from their topics

joint inference · 0.4self-training · 0.3n-best list combination · 0.2discriminative reranking · 0.2constraint learning · 0.1dependency parsing · 0.1reranking · 0.1parsing · 0.1
YearPublicationVenuePosition
2016 The Role of Context Types and Dimensionality in Learning Word Embeddings
abstract
Oren Melamud, David McClosky, Siddharth Patwardhan, Mohit Bansal. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Oren Melamud, David McClosky, Siddharth Patwardhan, Mohit Bansal
HLT-NAACL2
2015 Parsing Paraphrases with Joint Inference
abstract
Do Kook Choe, David McClosky. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Do Kook Choe, David McClosky
ACL (1)2
2015 Syntactic Parse Fusion
abstract
Model combination techniques have consistently shown state-of-the-art performance across multiple tasks, including syntactic parsing.However, they dramatically increase runtime and can be difficult to employ in practice.We demonstrate that applying constituency model combination techniques to n-best lists instead of n different parsers results in significant parsing accuracy improvements.Parses are weighted by their probabilities and combined using an adapted version of Sagae and Lavie (2006).These accuracy gains come with marginal computational costs and are obtained on top of existing parsing techniques such as discriminative reranking and self-training, resulting in state-of-the-art accuracy: 92.6% on WSJ section 23.On out-of-domain corpora, accuracy is improved by 0.4% on average.We empirically confirm that six well-known n-best parsers benefit from the proposed methods across six domains.
Do Kook Choe, David McClosky, Eugene Charniak
EMNLP2
2012 Learning Constraints for Consistent Timeline Extraction
David McClosky, Christopher D. Manning
EMNLP-CoNLL1
2012 Combining joint models for biomedical event extraction
abstract
BACKGROUND: We explore techniques for performing model combination between the UMass and Stanford biomedical event extraction systems. Both sub-components address event extraction as a structured prediction problem, and use dual decomposition (UMass) and parsing algorithms (Stanford) to find the best scoring event structure. Our primary focus is on stacking where the predictions from the Stanford system are used as features in the UMass system. For comparison, we look at simpler model combination techniques such as intersection and union which require only the outputs from each system and combine them directly. RESULTS: First, we find that stacking substantially improves performance while intersection and union provide no significant benefits. Second, we investigate the graph properties of event structures and their impact on the combination of our systems. Finally, we trace the origins of events proposed by the stacked model to determine the role each system plays in different components of the output. We learn that, while stacking can propose novel event structures not seen in either base model, these events have extremely low precision. Removing these novel events improves our already state-of-the-art F1 to 56.6% on the test set of Genia (Task 1). Overall, the combined system formed via stacking ("FAUST") performed well in the BioNLP 2011 shared task. The FAUST system obtained 1st place in three out of four tasks: 1st place in Genia Task 1 (56.0% F1) and Task 2 (53.9%), 2nd place in the Epigenetics and Post-translational Modifications track (35.0%), and 1st place in the Infectious Diseases track (55.6%). CONCLUSION: We present a state-of-the-art event extraction system that relies on the strengths of structured prediction and model combination through stacking. Akin to results on other tasks, stacking outperforms intersection and union and leads to very strong results. The utility of model combination hinges on complementary views of the data, and we show that our sub-systems capture different graph properties of event structures. Finally, by removing low precision novel events, we show that performance from stacking can be further improved.
David McClosky, Sebastian Riedel 0001, Mihai Surdeanu, Andrew McCallum, Christopher D. Manning
BMC Bioinform.1
2011 Event Extraction as Dependency Parsing
David McClosky, Mihai Surdeanu, Christopher D. Manning
ACL1
2010 Automatic Domain Adaptation for Parsing
David McClosky, Eugene Charniak, Mark Johnson 0001
HLT-NAACL1
2009 Improving Unsupervised Dependency Parsing with Richer Contexts and Smoothing
William P. Headden III, Mark Johnson 0001, David McClosky
HLT-NAACL3
2008 Evaluating Unsupervised Part-of-Speech Tagging for Grammar Induction
William P. Headden III, David McClosky, Eugene Charniak
COLING2
2008 When is Self-Training Effective for Parsing?
David McClosky, Eugene Charniak, Mark Johnson 0001
COLING1
2006 A Look at Parsing and Its Applications
Matthew Lease, Eugene Charniak, Mark Johnson 0001, David McClosky
AAAI4
2006 Reranking and Self-Training for Parser Adaptation
abstract
Statistical parsers trained and tested on the Penn Wall Street Journal (WSJ) treebank have shown vast improvements over the last 10 years. Much of this improvement, however, is based upon an ever-increasing number of features to be trained on (typically) the WSJ treebank data. This has led to concern that such parsers may be too finely tuned to this corpus at the expense of portability to other genres. Such worries have merit. The standard "Charniak parser" checks in at a labeled precision-recall f-measure of 89.7% on the Penn WSJ test set, but only 82.9% on the test set from the Brown treebank corpus.This paper should allay these fears. In particular, we show that the reranking parser described in Charniak and Johnson (2005) improves performance of the parser on Brown to 85.2%. Furthermore, use of the self-training techniques described in (McClosky et al., 2006) raise this to 87.8% (an error reduction of 28%) again without any use of labeled Brown data. This is remarkable since training the parser and reranker on labeled Brown data achieves only 88.4%.
David McClosky, Eugene Charniak, Mark Johnson 0001
ACL1
2006 Effective Self-Training for Parsing
David McClosky, Eugene Charniak, Mark Johnson 0001
HLT-NAACL1