James R. Curran

dblp:96/3617 · also James Richard Curran · DBLP profile ↗
← Back
49ranked-venue papers
7as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 7 first-authorApplied, interdisciplinary, general and emerging computing · 4Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
24 papers
Information extraction and text analysis · 94% Knowledge representation and reasoning · 1% Question answering and dialogue systems · 1%
Software engineering, system software, and programming languages
2 papers
Programming languages and type systems · 72% Compilers and program optimization · 28%

Topics — the 30 heaviest of 39, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.882015
Identifying Cascading Errors using Constraints in Dependency Parsing · ACL (1) 2015
Parser Showdown at the Wall Street Corral: An Empirical Investigation of Error Types in Parser Output · EMNLP-CoNLL 2012
Dependency Hashing for n-best CCG Parsing · ACL (1) 2012
Natural language and speech › Information extraction and text analysis
named entity recognition
0.732019
NNE: A Dataset for Nested Named Entity Recognition in English Newswire · ACL (1) 2019
Analysing recall loss in named entity slot filling · EMNLP 2014
Learning multilingual named entity recognition from Wikipedia · Artif. Intell. 2013
Natural language and speech › Information extraction and text analysis › entity typing
fine-grained entity typing
0.412019
NNE: A Dataset for Nested Named Entity Recognition in English Newswire · ACL (1) 2019
Natural language and speech › Information extraction and text analysis › named entity recognition
nested named entity recognition
0.412019
NNE: A Dataset for Nested Named Entity Recognition in English Newswire · ACL (1) 2019
Natural language and speech › Information extraction and text analysis › syntactic parsing › grammar-based parsing
combinatory categorial grammar parsing
0.452012
Dependency Hashing for n-best CCG Parsing · ACL (1) 2012
Formalism-Independent Parser Evaluation with CCG and DepBank · ACL 2007
Multi-Tagging for Lexicalized-Grammar Parsing · ACL 2006
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
0.212015
Identifying Cascading Errors using Constraints in Dependency Parsing · ACL (1) 2015
Natural language and speech › Information extraction and text analysis
slot filling
0.212014
Analysing recall loss in named entity slot filling · EMNLP 2014
Natural language and speech › Information extraction and text analysis › data annotation › corpus annotation
syntactic annotation
0.222010
Rebanking CCGbank for Improved NP Interpretation · ACL 2010
Adding Noun Phrase Structure to the Penn Treebank · ACL 2007
Natural language and speech › Information extraction and text analysis › named entity recognition
cross-lingual named entity recognition
0.212013
Learning multilingual named entity recognition from Wikipedia · Artif. Intell. 2013
Natural language and speech › Information extraction and text analysis
entity linking
0.212013
Evaluating Entity Linking with Wikipedia · Artif. Intell. 2013
Natural language and speech › Information extraction and text analysis › relation extraction
quotation attribution
0.212013
Automatically Detecting and Attributing Indirect Quotations · EMNLP 2013
Natural language and speech › Information extraction and text analysis
sequence labeling
0.112012
A Sequence Labelling Approach to Quote Attribution · EMNLP-CoNLL 2012
Natural language and speech › Information extraction and text analysis › data annotation
corpus annotation
0.112019
NNE: A Dataset for Nested Named Entity Recognition in English Newswire · ACL (1) 2019
Natural language and speech › Information extraction and text analysis › syntactic parsing
parsing efficiency
0.112010
Faster Parsing by Supertagger Adaptation · ACL 2010
Compilers and program optimization
parsing
0.112010
Faster Parsing by Supertagger Adaptation · ACL 2010
Natural language and speech › Information extraction and text analysis › lexical semantics
semantic drift
0.112009
Reducing Semantic Drift with Bagging and Distributional Similarity · ACL/IJCNLP 2009
Programming languages and type systems › grammar formalisms
combinatory categorial grammar
0.112009
Fully Lexicalising CCGbank with Hat Categories · EMNLP 2009
Programming languages and type systems
grammar formalisms
0.112009
Fully Lexicalising CCGbank with Hat Categories · EMNLP 2009
Natural language and speech › Information extraction and text analysis › distributional semantics
distributional similarity
0.122009
Scaling Distributional Similarity to Large Corpora · ACL 2006
Reducing Semantic Drift with Bagging and Distributional Similarity · ACL/IJCNLP 2009
Natural language and speech › Information extraction and text analysis
lexical semantics
0.122005
Supersense Tagging of Unknown Nouns Using Semantic Similarity · ACL 2005
Scaling Context Space · ACL 2002
Natural language and speech › Information extraction and text analysis › syntactic parsing
parser evaluation
0.112007
Formalism-Independent Parser Evaluation with CCG and DepBank · ACL 2007
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › exponential family
maximum entropy models
0.112006
Question classification with log-linear models · SIGIR 2006
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging
0.112006
Multi-Tagging for Lexicalized-Grammar Parsing · ACL 2006
Natural language and speech › Question answering and dialogue systems › question understanding
question classification
0.112006
Question classification with log-linear models · SIGIR 2006
Natural language and speech › Information extraction and text analysis › text similarity
semantic similarity
0.112005
Supersense Tagging of Unknown Nouns Using Semantic Similarity · ACL 2005
Natural language and speech › Information extraction and text analysis › lexical semantics
supersense tagging
0.112005
Supersense Tagging of Unknown Nouns Using Semantic Similarity · ACL 2005
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.012013
Evaluating Entity Linking with Wikipedia · Artif. Intell. 2013
Computer vision › Segmentation and scene understanding
object extraction
0.012004
Object-Extraction and Question-Parsing using CCG · EMNLP 2004
Natural language and speech › Information extraction and text analysis › semantic parsing
question parsing
0.012004
Object-Extraction and Question-Parsing using CCG · EMNLP 2004
Natural language and speech › Information extraction and text analysis › syntactic parsing
n-best parsing
0.012012
Dependency Hashing for n-best CCG Parsing · ACL (1) 2012

Methods — techniques the papers use, named apart from their topics

nested annotation · 0.4domain adaptation · 0.2constraints · 0.2combinatory categorial grammar · 0.2coreference resolution · 0.2sequence labeling · 0.2multilingual learning · 0.2empirical comparison · 0.1dependency hashing · 0.1supertagger adaptation · 0.1corpus reannotation · 0.1statistical weighting · 0.1
YearPublicationVenuePosition
2019 NNE: A Dataset for Nested Named Entity Recognition in English Newswire
abstract
Named entity recognition (NER) is widely used in natural language processing applications and downstream tasks.However, most NER tools target flat annotation from popular datasets, eschewing the semantic information available in nested entity mentions.We describe NNE-a fine-grained, nested named entity dataset over the full Wall Street Journal portion of the Penn Treebank (PTB).Our annotation comprises 279,795 mentions of 114 entity types with up to 6 layers of nesting.We hope the public release of this large dataset for English newswire will encourage development of new techniques for nested NER.
Nicky Ringland, Xiang Dai 0001, Ben Hachey, Sarvnaz Karimi, Cécile Paris, James R. Curran
ACL (1)6
2019 Article Segmentation in Digitised Newspapers with a 2D Markov Model
abstract
Document analysis and recognition is increasingly used to digitise collections of historical books, newspapers and other periodicals. In the digital humanities, it is often the goal to apply information retrieval (IR) and natural language processing (NLP) techniques to help researchers analyse and navigate these digitised archives. The lack of article segmentation is impairing many IR and NLP systems, which assume text is split into ordered, error-free documents. We define a document analysis and image processing task for segmenting digitised newspapers into articles and other content, e.g. adverts, and we automatically create a dataset of 11602 articles. Using this dataset, we develop and evaluate an innovative 2D Markov model that encodes reading order and substantially outperforms the current state-of-the-art, reaching similar accuracy to human annotators.
Andrew Naoum, Joel Nothman, James R. Curran
ICDAR3
2018 A Data-Driven Method for Helping Teachers Improve Feedback in Computer Programming Automated Tutors
Jessica McBroom, Kalina Yacef, Irena Koprinska, James R. Curran
AIED (1)4
2015 Identifying Cascading Errors using Constraints in Dependency Parsing
abstract
Dominick Ng, James R. Curran. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Dominick Ng, James R. Curran
ACL (1)2
2014 docrep: A lightweight and efficient document representation framework
Tim Dawborn, James R. Curran
COLING2
2014 Limited memory incremental coreference resolution
Kellie Webster, James R. Curran
COLING2
2014 Analysing recall loss in named entity slot filling
abstract
State-of-the-art fact extraction is heavily constrained by recall, as demonstrated by recent performance in TAC Slot Filling. We isolate this recall loss for NE slots by systematically analysing each stage of the slot filling pipeline as a filter over correct answers. Recall is critical as candidates never generated can never be recovered, whereas precision can always be increased in downstream processing. We provide precise, empirical confirma-tion of previously hypothesised sources of recall loss in slot filling. While NE type constraints substantially reduce the search space with only a minor recall penalty, we find that 10 % to 39 % of slot fills will be entirely ignored by most systems. One in six correct answers are lost if coreference is not used, but this can be mostly retained by simple name matching rules.
Glen Pink, Joel Nothman, James R. Curran
EMNLP3
2013 Automatically Detecting and Attributing Indirect Quotations
abstract
Direct quotations are used for opinion mining and information extraction as they have an easy to extract span and they can be attributed to a speaker with high accuracy.However, simply focusing on direct quotations ignores around half of all reported speech, which is in the form of indirect or mixed speech.This work presents the first large-scale experiments in indirect and mixed quotation extraction and attribution.We propose two methods of extracting all quote types from news articles and evaluate them on two large annotated corpora, one of which is a contribution of this work.We further show that direct quotation attribution methods can be successfully applied to indirect and mixed quotation attribution.* *These authors contributed equally to this work.by quotation marks, which makes them easy to extract.However, annotated resources suggest that direct quotations represent only a limited portion of all quotations, i.e., around 30% in the Penn Attribution Relation Corpus (PARC), which covers Wall Street Journal articles, and 52% in the Sydney Morning Herald Corpus (SMHC), with the remainder being indirect (Ex.1c) or mixed (Ex.1b)quotations.Retrieving only direct quotations can miss key content that can change the interpretation of the quotation (Ex.1b) and will entirely miss indirect quotations.
Silvia Pareti, Timothy O'Keefe, Ioannis Konstas, James R. Curran, Irena Koprinska
EMNLP4
2013 Evaluating Entity Linking with Wikipedia
Ben Hachey, Will Radford, Joel Nothman, Matthew Honnibal, James R. Curran
Artif. Intell.5
2013 Learning multilingual named entity recognition from Wikipedia
Joel Nothman, Nicky Ringland, Will Radford, Tara Murphy, James R. Curran
Artif. Intell.5
2012 Dependency Hashing for n-best CCG Parsing
Dominick Ng, James R. Curran
ACL (1)2
2012 Improving Combinatory Categorial Grammar Parse Reranking with Dependency Grammar Features
Sunghwan Mac Kim, Dominick Ng, James R. Curran
COLING4
2012 Parser Showdown at the Wall Street Corral: An Empirical Investigation of Error Types in Parser Output
Jonathan K. Kummerfeld, David Hall 0006, James R. Curran, Daniel Klein 0001
EMNLP-CoNLL3
2012 A Sequence Labelling Approach to Quote Attribution
Timothy O'Keefe, Silvia Pareti, James R. Curran, Irena Koprinska, Matthew Honnibal
EMNLP-CoNLL3
2012 The Challenges of Parsing Chinese with Combinatory Categorial Grammar
Daniel Tse, James R. Curran
HLT-NAACL2
2011 Graph-Based Named Entity Linking with Wikipedia
Ben Hachey, Will Radford, James R. Curran
WISE3
2011 Parsing Noun Phrases in the Penn Treebank
abstract
Noun phrases (nps) are a crucial part of natural language, and can have a very complex structure. However, this np structure is largely ignored by the statistical parsing field, as the most widely used corpus is not annotated with it. This lack of gold-standard data has restricted previous efforts to parse nps, making it impossible to perform the supervised experiments that have achieved high performance in so many Natural Language Processing (nlp) tasks. We comprehensively solve this problem by manually annotating np structure for the entire Wall Street Journal section of the Penn Treebank. The inter-annotator agreement scores that we attain dispel the belief that the task is too difficult, and demonstrate that consistent np annotation is possible. Our gold-standard np data is now available for use in all parsers. We experiment with this new data, applying the Collins (2003) parsing model, and find that its recovery of np structure is significantly worse than its overall performance. The parser's F-score is up to 5.69% lower than a baseline that uses deterministic rules. Through much experimentation, we determine that this result is primarily caused by a lack of lexical information. To solve this problem we construct a wide-coverage, large-scale np Bracketing system. With our Penn Treebank data set, which is orders of magnitude larger than those used previously, we build a supervised model that achieves excellent results. Our model performs at 93.8% F-score on the simple task that most previous work has undertaken, and extends to bracket longer, more complex nps that are rarely dealt with in the literature. We attain 89.14% F-score on this much more difficult task. Finally, we implement a post-processing module that brackets nps identified by the Bikel (2004) parser. Our np Bracketing model includes a wide variety of features that provide the lexical information that was missing during the parser experiments, and as a result, we outperform the parser's F-score by 9.04%. These experiments demonstrate the utility of the corpus, and show that many nlp applications can now make use of np structure.
David Vadas, James R. Curran
Comput. Linguistics2
2010 Rebanking CCGbank for Improved NP Interpretation
Matthew Honnibal, James R. Curran, Johan Bos
ACL2
2010 Faster Parsing by Supertagger Adaptation
Jonathan K. Kummerfeld, Jessika Roesner, Tim Dawborn, James Haggerty, James R. Curran, Stephen Clark
ACL5
2010 Chinese CCGbank: extracting CCG derivations from the Penn Chinese Treebank
Daniel Tse, James R. Curran
COLING2
2010 Data Mining for Generating Hints in a Python Tutor
Anna Katrina Dominguez, Kalina Yacef, James R. Curran
EDM3
2010 Data Mining to Generate Individualised Feedback
Anna Katrina Dominguez, Kalina Yacef, James R. Curran
Intelligent Tutoring Systems (2)3
2009 Reducing Semantic Drift with Bagging and Distributional Similarity
Tara McIntosh, James R. Curran
ACL/IJCNLP2
2009 Analysing Wikipedia and Gold-Standard Corpora for NER Training
Joel Nothman, Tara Murphy, James R. Curran
EACL3
2009 Fully Lexicalising CCGbank with Hat Categories
Matthew Honnibal, James R. Curran
EMNLP2
2009 Challenges for automatically extracting molecular interactions from full-text articles
abstract
BACKGROUND: The increasing availability of full-text biomedical articles will allow more biomedical knowledge to be extracted automatically with greater reliability. However, most Information Retrieval (IR) and Extraction (IE) tools currently process only abstracts. The lack of corpora has limited the development of tools that are capable of exploiting the knowledge in full-text articles. As a result, there has been little investigation into the advantages of full-text document structure, and the challenges developers will face in processing full-text articles. RESULTS: We manually annotated passages from full-text articles that describe interactions summarised in a Molecular Interaction Map (MIM). Our corpus tracks the process of identifying facts to form the MIM summaries and captures any factual dependencies that must be resolved to extract the fact completely. For example, a fact in the results section may require a synonym defined in the introduction. The passages are also annotated with negated and coreference expressions that must be resolved.We describe the guidelines for identifying relevant passages and possible dependencies. The corpus includes 2162 sentences from 78 full-text articles. Our corpus analysis demonstrates the necessity of full-text processing; identifies the article sections where interactions are most commonly stated; and quantifies the proportion of interaction statements requiring coherent dependencies. Further, it allows us to report on the relative importance of identifying synonyms and resolving negated expressions. We also experiment with an oracle sentence retrieval system using the corpus as a gold-standard evaluation set. CONCLUSION: We introduce the MIM corpus, a unique resource that maps interaction facts in a MIM to annotated passages within full-text articles. It is an invaluable case study providing guidance to developers of biomedical IR and IE systems, and can be used as a gold-standard evaluation set for full-text IR tasks.
Tara McIntosh, James R. Curran
BMC Bioinform.2
2008 Parsing Noun Phrase Structure with CCG
David Vadas, James R. Curran
ACL2
2007 Formalism-Independent Parser Evaluation with CCG and DepBank
Stephen Clark, James R. Curran
ACL2
2007 Linguistically Motivated Large-Scale NLP with C&C and Boxer
James R. Curran, Stephen Clark, Johan Bos
ACL1
2007 Adding Noun Phrase Structure to the Penn Treebank
David Vadas, James R. Curran
ACL2
2007 Wide-Coverage Efficient Statistical Parsing with CCG and Log-Linear Models
abstract
This article describes a number of log-linear parsing models for an automatically extracted lexicalized grammar. The models are “full” parsing models in the sense that probabilities are defined for complete parses, rather than for independent events derived by decomposing the parse tree. Discriminative training is used to estimate the models, which requires incorrect parses for each sentence in the training data as well as the correct parse. The lexicalized grammar formalism used is Combinatory Categorial Grammar (CCG), and the grammar is automatically extracted from CCGbank, a CCG version of the Penn Treebank. The combination of discriminative training and an automatically extracted grammar leads to a significant memory requirement (up to 25 GB), which is satisfied using a parallel implementation of the BFGS optimization algorithm running on a Beowulf cluster. Dynamic programming over a packed chart, in combination with the parallel implementation, allows us to solve one of the largest-scale estimation problems in the statistical parsing literature in under three hours. A key component of the parsing system, for both training and testing, is a Maximum Entropy supertagger which assigns CCG lexical categories to words in a sentence. The supertagger makes the discriminative training feasible, and also leads to a highly efficient parser. Surprisingly, given CCG's “spurious ambiguity,” the parsing speeds are significantly higher than those reported for comparable parsers in the literature. We also extend the existing parsing techniques for CCG by developing a new model and efficient parsing algorithm which exploits all derivations, including CCG's nonstandard derivations. This model and parsing algorithm, when combined with normal-form constraints, give state-of-the-art accuracy for the recovery of predicate-argument dependencies from CCGbank. The parser is also evaluated on DepBank and compared against the RASP parser, outperforming RASP overall and on the majority of relation types. The evaluation on DepBank raises a number of issues regarding parser evaluation. This article provides a comprehensive blueprint for building a wide-coverage CCG parser. We demonstrate that both accurate and highly efficient parsing is possible with CCG.
Stephen Clark, James R. Curran
Comput. Linguistics2
2006 Multi-Tagging for Lexicalized-Grammar Parsing
abstract
With performance above 97% accuracy for newspaper text, part of speech (POS) tagging might be considered a solved problem. Previous studies have shown that allowing the parser to resolve POS tag ambiguity does not improve performance. However, for grammar formalisms which use more fine-grained grammatical categories, for example TAG and CCG, tagging accuracy is much lower. In fact, for these formalisms, premature ambiguity resolution makes parsing infeasible.We describe a multi-tagging approach which maintains a suitable level of lexical category ambiguity for accurate and efficient CCG parsing. We extend this multi-tagging approach to the POS level to overcome errors introduced by automatically assigned POS tags. Although POS tagging accuracy seems high, maintaining some POS tag ambiguity in the language processing pipeline results in more accurate CCG supertagging.
James R. Curran, Stephen Clark, David Vadas
ACL1
2006 Scaling Distributional Similarity to Large Corpora
abstract
Accurately representing synonymy using distributional similarity requires large volumes of data to reliably represent infrequent words. However, the naïve nearest-neighbour approach to comparing context vectors extracted from large corpora scales poorly (O(n2) in the vocabulary size).In this paper, we compare several existing approaches to approximating the nearest-neighbour search for distributional similarity. We investigate the trade-off between efficiency and accuracy, and find that SASH (Houle and Sakuma, 2005) provides the best balance.
James Gorman, James R. Curran
ACL2
2006 Web Text Corpus for Natural Language Processing
Vinci Liu, James R. Curran
EACL2
2006 Random Indexing using Statistical Weight Functions
James Gorman, James R. Curran
EMNLP2
2006 Building a search engine to drive problem-based learning
abstract
Search engines pervade the digital world, mediating most access to information instantaneously. We have found that students can build search engine components, and even entire search engines, in the context of problem-based learning in introductory and intermediate computer science courses. The courses cover a broad range of topics in algorithms, data structures, and web design, with a heavy emphasis on programming. Additionally, the internet is coupled with the syllabus at many places, from web design and HTML to graph algorithms and pattern matching. This connection enlivens the discussion of otherwise dry topics like searching, sorting, indexing and hashing. Moreover, the challenge of web-scale computing motivates the continuing students in their later study of formal topics like algorithmic complexity, while non-continuing students acquire transferable analytical skills. We report on the experience in search engine projects for driving problem-based learning in computer science courses, for both high school and university students. Our experience shows that such projects are effective in both introductory and intermediate courses, and readily encompass student groups with diverse programming abilities.
Steven Bird, James R. Curran
ITiCSE2
2006 Partial Training for a Lexicalized-Grammar Parser
Stephen Clark, James R. Curran
HLT-NAACL2
2006 Question classification with log-linear models
abstract
Question classification has become a crucial step in modern question answering systems. Previous work has demonstrated the effectiveness of statistical machine learning approaches to this problem. This paper presents a new approach to building a question classifier using log-linear models. Evidence from a rich and diverse set of syntactic and semantic features is evaluated, as well as approaches which exploit the hierarchical structure of the question classes.
Phil Blunsom, Krystle Kocik, James R. Curran
SIGIR3
2005 Supersense Tagging of Unknown Nouns Using Semantic Similarity
abstract
The limited coverage of lexical-semantic resources is a significant problem for NLP systems which can be alleviated by automatically classifying the unknown words. Supersense tagging assigns unknown nouns one of 26 broad semantic categories used by lexicographers to organise their manual insertion into WORDNET. Ciaramita and Johnson (2003) present a tagger which uses synonym set glosses as annotated training examples. We describe an unsupervised approach, based on vector-space similarity, which does not require annotated examples but significantly outperforms their tagger. We also demonstrate the use of an extremely large shallow-parsed corpus for calculating vector-space semantic similarity.
James R. Curran
ACL1
2004 Parsing the WSJ Using CCG and Log-Linear Models
abstract
This paper describes and evaluates log-linear parsing models for Combinatory Categorial Grammar (CCG). A parallel implementation of the L-BFGS optimisation algorithm is described, which runs on a Beowulf cluster allowing the complete Penn Treebank to be used for estimation. We also develop a new efficient parsing algorithm for CCG which maximises expected recall of dependencies. We compare models which use all CCG derivations, including non-standard derivations, with normal-form models. The performances of the two models are comparable and the results are competitive with existing wide-coverage CCG parsers.
Stephen Clark, James R. Curran
ACL2
2004 Wide-Coverage Semantic Representations from a CCG Parser
Johan Bos, Stephen Clark, Mark Steedman, James R. Curran, Julia Hockenmaier
COLING4
2004 The Importance of Supertagging for Wide-Coverage CCG Parsing
Stephen Clark, James R. Curran
COLING2
2004 Object-Extraction and Question-Parsing using CCG
Stephen Clark, Mark Steedman, James R. Curran
EMNLP3
2003 Bootstrapping POS-taggers using unlabelled data
Stephen Clark, James R. Curran, Miles Osborne
CoNLL2
2003 Language Independent NER using a Maximum Entropy Tagger
James R. Curran, Stephen Clark
CoNLL1
2003 Investigating GIS and Smoothing for Maximum Entropy Taggers
James R. Curran, Stephen Clark
EACL1
2003 Log-Linear Models for Wide-Coverage CCG Parsing
Stephen Clark, James R. Curran
EMNLP2
2002 Scaling Context Space
abstract
Context is used in many NLP systems as an indicator of a term's syntactic and semantic function. The accuracy of the system is dependent on the quality and quantity of contextual information available to describe each term. However, the quantity variable is no longer fixed by limited corpus resources. Given fixed training time and computational resources, it makes sense for systems to invest time in extracting high quality contextual information from a fixed corpus. However, with an effectively limitless quantity of text available, extraction rate and representation size need to be considered. We use thesaurus extraction with a range of context extracting tools to demonstrate the interaction between context quantity, time and size on a corpus of 300 million words.
James R. Curran, Marc Moens
ACL1
2002 A Very Very Large Corpus Doesn't Always Yield Reliable Estimates
James R. Curran, Miles Osborne
CoNLL1