Mark Steedman

dblp:s/MarkSteedman · DBLP profile ↗
← Back
83ranked-venue papers
10as first author
15since 2021 · last 2025
0000-0003-2509-0797ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 80 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Empirical Study on Data Attributes Insufficiency of Evaluation Benchmarks for LLMs
abstract
Previous benchmarks for evaluating large language models (LLMs) have primarily emphasized quantitative metrics, such as data volume. However, this focus may neglect key qualitative data attributes that can significantly impact the final rankings of LLMs, resulting in unreliable leaderboards. In this paper, we investigate whether current LLM benchmarks adequately consider these data attributes. We specifically examine three attributes: diversity, redundancy, and difficulty. To explore these attributes, we propose a framework with three separate modules, each designed to assess one of the attributes. Using a method that progressively incorporates these attributes, we analyze their influence on the benchmark. Our experimental results reveal a meaningful correlation between LLM rankings on the revised benchmark and the original benchmark when these attributes are accounted for. These findings indicate that existing benchmarks often fail to meet all three criteria, highlighting a lack of consideration for multifaceted data attributes in current evaluation datasets.
Chuang Liu 0009, Renren Jin, Mark Steedman, Deyi Xiong
COLING6
2025 MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
abstract
The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are capable of handling hundreds of images with interleaved text tokens in a single forward pass. In this work, we introduce MMLongBench, the first benchmark covering a diverse set of long-context vision-language tasks, to evaluate LCVLMs effectively and thoroughly. MMLongBench is composed of 13,331 examples spanning five different categories of downstream tasks, such as Visual RAG and Many-Shot ICL. It also provides broad coverage of image types, including various natural and synthetic images. To assess the robustness of the models to different input lengths, all examples are delivered at five standardized input lengths (8K-128K tokens) via a cross-modal tokenization scheme that combines vision patches and text tokens. Through a thorough benchmarking of 46 closed-source and open-source LCVLMs, we provide a comprehensive analysis of the current models' vision-language long-context ability. Our results show that: i) performance on a single task is a weak proxy for overall long-context capability; ii) both closed-source and open-source models face challenges in long-context vision-language tasks, indicating substantial room for future improvement; iii) models with stronger reasoning ability tend to exhibit better long-context performance. By offering wide task coverage, various image types, and rigorous length control, MMLongBench provides the missing foundation for diagnosing and advancing the next generation of LCVLMs.
Zhaowei Wang 0003, Wenhao Yu 0002, Xiyu Ren, Yu Zhao 0043, Rohit Saxena, Ginny Y. Wong, Simon See, Pasquale Minervini, Yangqiu Song, Mark Steedman
NeurIPS12
2025 Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
abstract
Abstract Recent machine translation (MT) metrics calibrate their effectiveness by correlating with human judgment. However, these results are often obtained by averaging predictions across large test sets without any insights into the strengths and weaknesses of these metrics across different error types. Challenge sets are used to probe specific dimensions of metric behavior but there are very few such datasets and they either focus on a limited number of phenomena or a limited number of language pairs. We introduce ACES, a contrastive challenge set spanning 146 language pairs, aimed at discovering whether metrics can identify 68 translation accuracy errors. These phenomena range from basic alterations at the word/character level to more intricate errors based on discourse and real-world knowledge. We conducted a large-scale study by benchmarking ACES on 47 metrics submitted to the WMT 2022 and WMT 2023 metrics shared tasks. We also measure their sensitivity to a range of linguistic phenomena. We further investigate claims that large language models (LLMs) are effective as MT evaluators, addressing the limitations of previous studies by using a dataset that covers a range of linguistic phenomena and language pairs and includes both low- and medium-resource languages. Our results demonstrate that different metric families struggle with different phenomena and that LLM-based methods are unreliable. We expose a number of major flaws with existing methods: Most metrics ignore the source sentence; metrics tend to prefer surface level overlap; and over-reliance on language-agnostic representations leads to confusion when the target language is similar to the source language. To further encourage detailed evaluation beyond singular scores, we expand ACES to include error span annotations, denoted as SPAN-ACES, and we use this dataset to evaluate span-based error metrics, showing that these metrics also need considerable improvement. Based on our observations, we provide a set of recommendations for building better MT metrics, including focusing on error labels instead of scores, ensembling, designing metrics to explicitly focus on the source sentence, focusing on semantic content rather than relying on the lexical overlap, and choosing the right pre-trained model for obtaining representations.
Nikita Moghe, Arnisa Fazla, Chantal Amrhein, Tom Kocmi, Mark Steedman, Alexandra Birch, Rico Sennrich, Liane Guillou
Comput. Linguistics5
2025 A language-agnostic model of child language acquisition
abstract
This work reimplements a recent semantic bootstrapping child language acquisition (CLA) model, which was originally designed for English, and trains it to learn a new language: Hebrew. The model learns from pairs of utterances and logical forms as meaning representations, and acquires both syntax and word meanings simultaneously. The results show that the model mostly transfers to Hebrew, but that a number of factors, including the richer morphology in Hebrew, makes the learning slower and less robust. This suggests that a clear direction for future work is to enable the model to leverage the similarities between different word forms.
Louis Mahon, Omri Abend, Uri Berger, Katherine Demuth, Mark Johnson 0001, Mark Steedman
Comput. Speech Lang.6
2024 Human Temporal Inferences Go Beyond Aspectual Class
abstract
Past work in NLP has proposed the task of classifying English verb phrases into situation aspect categories, assuming that these categories play an important role in tasks requiring temporal reasoning. We investigate this assumption by gathering crowd-sourced judgements about aspectual entailments from non-expert, native English participants. The results suggest that aspectual class alone is not sufficient to explain the response patterns of the participants. We propose that looking at scenarios which can feasibly accompany an action description contributes towards a better explanation of the participants' answers. A further experiment using GPT-3.5 shows that its outputs follow different patterns than human answers, suggesting that such conceivable scenarios cannot be fully accounted for in the language alone. We release our dataset to support further research.
Katarzyna Prus, Mark Steedman, Adam Lopez
EACL (1)2
2024 A Usage-centric Take on Intent Understanding in E-Commerce
abstract
Identifying and understanding user intents is a pivotal task for E-Commerce.Despite its essential role in product recommendation and business user profiling analysis, intent understanding has not been consistently defined or accurately benchmarked.In this paper, we focus on predicative user intents as "how a customer uses a product", and pose intent understanding as a natural language reasoning task, independent of product ontologies.We identify two weaknesses of FolkScope, the SOTA E-Commerce Intent Knowledge Graph: categoryrigidity and property-ambiguity.They limit its ability to strongly align user intents with products having the most desirable property, and to recommend useful products across diverse categories.Following these observations, we introduce a Product Recovery Benchmark featuring a novel evaluation framework and an example dataset.We further validate the above FolkScope weaknesses on this benchmark.Our code and dataset are available at https://github.com/stayones/Usgae-Centric- Intent-Understanding.
Wendi Zhou, Pavlos Vougiouklis, Mark Steedman, Jeff Z. Pan
EMNLP4
2023 Extrinsic Evaluation of Machine Translation Metrics
abstract
Automatic machine translation (MT) metrics are widely used to distinguish the quality of machine translation systems across large test sets (i.e., system-level evaluation).However, it is unclear if automatic metrics can reliably distinguish good translations from bad at the sentence level (i.e., segment-level evaluation).We investigate how useful MT metrics are at detecting segment-level quality by correlating metrics with the translation utility for downstream tasks.We evaluate the segment-level performance of widespread MT metrics (chrF, COMET, BERTScore, etc.) on three downstream cross-lingual tasks (dialogue state tracking, question answering, and semantic parsing).For each task, we have access to a monolingual task-specific model and a translation model.We calculate the correlation between the metric's ability to predict a good/bad translation with the success/failure on the final task for machine-translated test sentences.Our experiments demonstrate that all metrics exhibit negligible correlation with the extrinsic evaluation of downstream outcomes.We also find that the scores provided by neural metrics are not interpretable, in large part due to having undefined ranges.We synthesise our analysis into recommendations for future MT metrics to produce labels rather than scores for more informative interaction between machine translation and multilingual language understanding.
Nikita Moghe, Tom Sherborne, Mark Steedman, Alexandra Birch
ACL (1)3
2023 Smoothing Entailment Graphs with Language Models
abstract
Nick McKenna, Tianyi Li, Mark Johnson, Mark Steedman. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Nick McKenna, Mark Johnson 0001, Mark Steedman
IJCNLP (1)4
2023 Parsing dialog turns with prosodic features in English
abstract
Parsing spoken dialogue presents challenges that parsing text does not, including a lack of clear sentence boundaries.We know from previous work that prosody helps in parsing single sentences [1], but we want to show the effect of prosody on parsing speech that isn't segmented into sentences.In experiments on the English Switchboard corpus, we find prosody helps our model both with parsing and with accurately identifying sentence boundaries.However, we find that the bestperforming parser is not necessarily the parser that produces the best sentence segmentation performance.We suggest that the best parses instead come from modelling sentence boundaries jointly with other syntactic boundaries.
Elizabeth Nielsen, Mark Steedman, Sharon Goldwater
INTERSPEECH2
2022 Sentence-Incremental Neural Coreference Resolution
abstract
We propose a sentence-incremental neural coreference resolution system which incrementally builds clusters after marking mention boundaries in a shift-reduce method.The system is aimed at bridging two recent approaches at coreference resolution: (1) state-of-the-art non-incremental models that incur quadratic complexity in document length with high computational cost, and (2) memory networkbased models which operate incrementally but do not generalize beyond pronouns.For comparison, we simulate an incremental setting by constraining non-incremental systems to form partial coreference chains before observing new sentences.In this setting, our system outperforms comparable state-of-the-art methods by 2 F1 on OntoNotes and 6.8 F1 on the CODI-CRAC 2021 corpus.In a conventional coreference setup, our system achieves 76.3 F1 on OntoNotes and 45.5 F1 on CODI-CRAC 2021, which is comparable to state-of-the-art baselines.We also analyze variations of our system and show that the degree of incrementality in the encoder has a surprisingly large effect on the resulting performance.1 1 Code is available at: https://github.com/mgrenander/sentence-incremental-coref
Matt Grenander, Shay B. Cohen, Mark Steedman
EMNLP3
2022 Erratum for "Formal Basis of a Language Universal"
abstract
In the paper “Formal Basis of a Language Universal” by Miloš Stanojević and Mark Steedman in Computational Linguistics 47:1 (https://doi.org/10.1162/coli_a_00394), there is an error in example (12) on page 17.The two occurrences of the notation ∖W should appear as |W. The paper has been updated so that the paragraph reads:In the full theory, these rules are generalized to “second level” cases, in which the secondary function is of the form (Y|Z)|W such as the following “forward crossing” instance, in which — matches either / or ∖ in both input and output:(12) The Forward Crossing Level Two Composition Rule X/×Y (Y∖Z)|W ⇒B× (X∖Z)|W (B×2)
Milos Stanojevic, Mark Steedman
Comput. Linguistics2
2021 Prosodic segmentation for parsing spoken dialogue
abstract
Elizabeth Nielsen, Mark Steedman, Sharon Goldwater. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Elizabeth Nielsen, Mark Steedman, Sharon Goldwater
ACL/IJCNLP (1)2
2021 Multivalent Entailment Graphs for Question Answering
abstract
Drawing inferences between open-domain natural language predicates is a necessity for true language understanding.There has been much progress in unsupervised learning of entailment graphs for this purpose.We make three contributions: (1) we reinterpret the Distributional Inclusion Hypothesis to model entailment between predicates of different valencies, like DEFEAT(Biden, Trump) WIN(Biden);(2) we actualize this theory by learning unsupervised Multivalent Entailment Graphs of open-domain predicates; and (3) we demonstrate the capabilities of these graphs on a novel question answering task.We show that directional entailment is more helpful for inference than non-directional similarity on questions of fine-grained semantics.We also show that drawing on evidence across valencies answers more questions than by using only the same valency evidence.
Nick McKenna, Liane Guillou, Mohammad Javad Hosseini, Sander Bijl de Vroe, Mark Johnson 0001, Mark Steedman
EMNLP (1)6
2021 Cross-lingual Intermediate Fine-tuning improves Dialogue State Tracking
abstract
Recent progress in task-oriented neural dialogue systems is largely focused on a handful of languages, as annotation of training data is tedious and expensive.Machine translation has been used to make systems multilingual, but this can introduce a pipeline of errors.Another promising solution is using cross-lingual transfer learning through pretrained multilingual models.Existing methods train multilingual models with additional codemixed task data or refine the cross-lingual representations through parallel ontologies.In this work, we enhance the transfer learning process by intermediate fine-tuning of pretrained multilingual models, where the multilingual models are fine-tuned with different but related data and/or tasks.Specifically, we use parallel and conversational movie subtitles datasets to design cross-lingual intermediate tasks suitable for downstream dialogue tasks.We use only 200K lines of parallel data for intermediate fine-tuning which is already available for 1782 language pairs.We test our approach on the cross-lingual dialogue state tracking task for the parallel Mul-tiWoZ (English→Chinese, Chinese→English) and Multilingual WoZ (English→German, English→Italian) datasets.We achieve impressive improvements (> 20% on joint goal accuracy) on the parallel MultiWoZ dataset and the Multilingual WoZ dataset over the vanilla baseline with only 10% of the target language task data and zero-shot setup respectively.
Nikita Moghe, Mark Steedman, Alexandra Birch
EMNLP (1)2
2021 Formal Basis of a Language Universal
abstract
Abstract Steedman (2020) proposes as a formal universal of natural language grammar that grammatical permutations of the kind that have given rise to transformational rules are limited to a class known to mathematicians and computer scientists as the “separable” permutations. This class of permutations is exactly the class that can be expressed in combinatory categorial grammars (CCGs). The excluded non-separable permutations do in fact seem to be absent in a number of studies of crosslinguistic variation in word order in nominal and verbal constructions. The number of permutations that are separable grows in the number n of lexical elements in the construction as the Large Schröder Number Sn−1. Because that number grows much more slowly than the n! number of all permutations, this generalization is also of considerable practical interest for computational applications such as parsing and machine translation. The present article examines the mathematical and computational origins of this restriction, and the reason it is exactly captured in CCG without the imposition of any further constraints.
Milos Stanojevic, Mark Steedman
Comput. Linguistics2
2020 Max-Margin Incremental CCG Parsing
abstract
Incremental syntactic parsing has been an active research area both for cognitive scientists trying to model human sentence processing and for NLP researchers attempting to combine incremental parsing with language modelling for ASR and MT.Most effort has been directed at designing the right transition mechanism, but less has been done to answer the question of what a probabilistic model for those transition parsers should look like.A very incremental transition mechanism of a recently proposed CCG parser when trained in straightforward locally normalised discriminative fashion produces very bad results on English CCGbank.We identify three biases as the causes of this problem: label bias, exposure bias and imbalanced probabilities bias.While known techniques for tackling these biases improve results, they still do not make the parser state of the art.Instead, we tackle all of these three biases at the same time using an improved version of beam search optimisation that minimises all beam search violations instead of minimising only the biggest violation.The new incremental parser gives better results than all previously published incremental CCG parsers, and outperforms even some widely used non-incremental CCG parsers.
Milos Stanojevic, Mark Steedman
ACL2
2020 Aspectuality Across Genre: A Distributional Semantics Approach
abstract
The interpretation of the lexical aspect of verbs in English plays a crucial role for recognizing textual entailment and learning discourse-level inferences.We show that two elementary dimensions of aspectual class, states vs. events, and telic vs. atelic events, can be modelled effectively with distributional semantics.We find that a verb's local context is most indicative of its aspectual class, and demonstrate that closed class words tend to be stronger discriminating contexts than content words.Our approach outperforms previous work on three datasets.Lastly, we contribute a dataset of human-human conversations annotated with lexical aspect and present experiments that show the correlation of telicity with genre and discourse goals.
Thomas Kober 0001, Malihe Alikhani, Matthew Stone, Mark Steedman
COLING4
2020 The role of context in neural pitch accent detection in English
abstract
Prosody is a rich information source in natural language, serving as a marker for phenomena such as contrast.In order to make this information available to downstream tasks, we need a way to detect prosodic events in speech.We propose a new model for pitch accent detection, inspired by the work of Stehwien et al. (2018), who presented a CNN-based model for this task.Our model makes greater use of context by using full utterances as input and adding an LSTM layer.We find that these innovations lead to an improvement from 87.5 percent to 88.7 percent accuracy on pitch accent detection on American English speech in the Boston University Radio News Corpus, a state-of-the-art result.We also find that a simple baseline that just predicts a pitch accent on every content word yields 82.2 percent accuracy, and we suggest that this is the appropriate baseline for this task.Finally, we conduct ablation tests that show pitch is the most important acoustic feature for this task and this corpus.
Elizabeth Nielsen, Mark Steedman, Sharon Goldwater
EMNLP (1)2
2019 Duality of Link Prediction and Entailment Graph Induction
abstract
Link prediction and entailment graph induction are often treated as different problems.In this paper, we show that these two problems are actually complementary.We train a link prediction model on a knowledge graph of assertions extracted from raw text.We propose an entailment score that exploits the new facts discovered by the link prediction model, and then form entailment graphs between relations.We further use the learned entailments to predict improved link prediction scores.Our results show that the two tasks can benefit from each other.The new entailment score outperforms prior state-of-the-art results on a standard entialment dataset and the new link prediction scores show improvements over the raw link prediction scores.
Mohammad Javad Hosseini, Shay B. Cohen, Mark Johnson 0001, Mark Steedman
ACL (1)4
2019 Wide-Coverage Neural A* Parsing for Minimalist Grammars
abstract
Minimalist Grammars (Stabler, 1997) are a computationally oriented, and rigorous formalisation of many aspects of Chomsky's (1995) Minimalist Program.This paper presents the first ever application of this formalism to the task of realistic wide-coverage parsing.The parser uses a linguistically expressive yet highly constrained grammar, together with an adaptation of the A* search algorithm currently used in CCG parsing (Lewis and Steedman, 2014;Lewis et al., 2016), with supertag probabilities provided by a bi-LSTM neural network supertagger trained on MGbank, a corpus of MG derivation trees.We report on some promising initial experimental results for overall dependency recovery as well as on the recovery of certain unbounded long distance dependencies.Finally, although like other MG parsers, ours has a high order polynomial worst case time complexity, we show that in practice its expected time complexity is O(n 3 ).The parser is publicly available.
John Torr, Milos Stanojevic, Mark Steedman, Shay B. Cohen
ACL (1)3
2018 Character-Level Models versus Morphology in Semantic Role Labeling
abstract
Character-level models have become a popular approach specially for their accessibility and ability to handle unseen data.However, little is known on their ability to reveal the underlying morphological structure of a word, which is a crucial skill for high-level semantic analysis tasks, such as semantic role labeling (SRL).In this work, we train various types of SRL models that use word, character and morphology level information and analyze how performance of characters compare to words and morphology for several languages.We conduct an in-depth error analysis for each morphological typology and analyze the strengths and limitations of character-level models that relate to out-of-domain data, training data size, long range dependencies and model complexity.Our exhaustive analyses shed light on important characteristics of character-level models and their semantic capability.
Gözde Gül Sahin, Mark Steedman
ACL (1)2
2018 Data Augmentation via Dependency Tree Morphing for Low-Resource Languages
abstract
Neural NLP systems achieve high scores in the presence of sizable training dataset.Lack of such datasets leads to poor system performances in the case low-resource languages.We present two simple text augmentation techniques using dependency trees, inspired from image processing.We "crop" sentences by removing dependency links, and we "rotate" sentences by moving the tree fragments around the root.We apply these techniques to augment the training sets of low-resource languages in Universal Dependencies project.We implement a character-level sequence tagging model and evaluate the augmented datasets on part-of-speech tagging task.We show that crop and rotate provides improvements over the models trained with non-augmented data for majority of the languages, especially for languages with rich case marking systems.
Gözde Gül Sahin, Mark Steedman
EMNLP2
2018 The Lost Combinator
abstract
Let me begin by thanking the Association for Computational Linguistics and its Executive Committee for conferring on me the great honor of their Lifetime Achievement Award for 2018, which of course I share with all the wonderful students and colleagues that have made many essential contributions to this work over many years.At the heart of the work that I have been pursuing over my research lifetime so far, whether in parsing and sentence processing, spoken language understanding, semantics, or even in musical understanding by machine, there lies a theory of natural language grammar that brings parsing, compositional semantics, statistical modeling, and logical inference into the closest possible relation. This theory of grammar is combinatory, in the sense that its operations are type-dependent and restricted to strictly string-adjacent phonologically or graphologically-realized inputs, and categorial, in the sense that those operands pair a syntactic type with a type-transparent semantic representation or logical form.I’d like to use this opportunity to briefly address three questions that revolve around the theory of grammar, both combinatory and otherwise. The first question concerns the way that Combinatory Categorial Grammar (CCG) was developed with a number of colleagues, over a number of stages and in slightly different forms. The second is an essentially evolutionary question of why natural language grammar should take a combinatory form. The third question is that of what the future holds for CCG and other structural theories of grammar in computational linguistics and NLP in the age of deep learning.I have called this talk “The Lost Combinator” in homage to the Victorian era poem “The Lost Chord,” in the hope of suggesting that the theoretical development of CCG has always been empirical, rather than axiomatic, in search of the simplest explanation of the facts of language, rather than for confirmation of linguistic received opinion, however intuitively salient.In the late 1960s (when I was a psychology undergraduate at the University of Sussex under Stuart Sutherland, and then started as a graduate student in artificial intelligence at Edinburgh under Christopher Longuet-Higgins), a broad community of theoretical linguists, psychologists, and computational linguists saw themselves as all working on the same problem, under the definition provided by the “transformational” theory of grammar proposed by Chomsky (1957, 1965), using theories of psycholinguistic processing, language acquisition, and language evolution proposed by Lashley (1951), Miller, Galanter, and Pribram (1960), Miller (1967), and Lenneberg (1967), theories of natural language semantics proposed by Carnap (1956), Montague (1970), and Lewis (1970), and computational models of parsing such as those proposed by Thorne, Bratley, and Dewar (1968) and Woods (1970). (I myself was so convinced that this program would succeed that I believed it was time to apply the same methods to other cognitive faculties, taking as my research project for Ph.D. their application to the interpretation of music by machine, following the lead of Max Clowes [1971] in machine vision.)Almost immediately, this consensus fell apart. First, Chomsky himself was among the first (1965) to recognize that transformational rules, though descriptively revealing, were so expressive as to have little explanatory force, and required many apparently arbitrary constraints (Ross 1967). Second, psychologists realized that psycholinguistic measures of processing difficulty of sentences bore almost no relation to their transformational derivational complexity (Marslen-Wilson 1973; Fodor, Bever, and Garrett 1974). Finally, computational linguists attempting to implement transformational grammars as parsers realized that they were spending all their time implementing even more constraints on rules, in order to limit search arising from overgeneration (Friedman 1971; Gross 1978). (Meanwhile, I realized that the problem had not in fact been solved, and returned to natural language processing, thanks to a postdoc at Sussex with Philip Johnson-Laird.)This disillusion wasn’t just a case of internal academic squabbling. There were also a couple of influential reports commissioned by the U.S. and UK governments that ended funding for machine translation (MT) and artificial intelligence (AI) (Pierce et al. 1966; Lighthill 1973). As a result of the second of these reports, which determined that AI was never going to work, PhDs in artificial intelligence like my classmate Geoff Hinton and myself spent ten years or so after graduation in psychology departments (in my case, at the Universities of Sussex and Warwick), until yet another report said AI was working after all and that Britain and the U.S. were falling behind Japan in this vital area. As a result, I could get hired again in computer science, first briefly back at Edinburgh, and then at the University of Pennsylvania (I learned a lesson from this odyssey that I have tried to remember whenever I have been appointed to a committee to report on anything, which is that while reports very rarely do any good, they can very easily do a great deal of harm.)Meanwhile, as a result of these conflicts, the scientific study of language fragmented. The linguists swiftly abjured any responsibility for their grammars (“Competence”) bearing any relation to processing (“Performance”). Because the psychologists could hardly abandon Performance, they in turn became agnostic about grammar, retreating to context-free surface grammar (which they tended to refer to as “parsing strategies”), or a touchingly optimistic belief in its emergence from neural models. Meanwhile, the computational linguists (whose machines were growing exponentially in size and speed from the 16K byte core of the machine that supported the whole group when I started my graduate studies, on to levels that would soon permit parsing the entire contents of the then embrionic Web) similarly found that very little of what the linguists and psychologists cared about was usable at scale, and that none of it significantly improved overall performance over very much simpler context-free or even finite-state methods that the linguists had shown to be incomplete. The reason of course was Zipf’s law, which means that the events with respect to which the low-level methods are incomplete are off in the long tail.It also became apparent to a few computationalists working on speech, MT, and information retrieval that the real problem was not grammar but ambiguity and its resolution by world-knowledge, and that the solution lay in probabilistic models (Bar-Hillel 1960/1964; Spärck Jones 1964/1986; Wilks 1975; Jelinek and Lafferty 1991) (although it was not immediately apparent how to combine statistical models with grammar-based systems without making obviously false independence assumptions).Nevertheless, as any red-blooded psychologist had always insisted, the divorce between competence and performance that everyone else had accepted did not make any sense. The grammar and the processor had to have evolved in lock-step, as a package deal, for what could be the evolutionary selective advantage of a grammar that you cannot process, or a parser without a grammar?It seemed equally obvious that surface syntax and the underlying semantic or conceptual representation must also be closely related, since the only reasonable basis for child language acquisition that has ever been on offer is that the child attaches language-specific grammar to a universal conceptual relation or “language of mind” (Miller 1967; Bowerman 1973; Wexler and Culicover 1980). It seemed to follow that radically new theories of grammar were needed.Theoretical linguists agree that the central problem for the theory of grammar is discontinuity or non-adjacent dependency between predicates and their arguments:Chomsky described discontinuity in terms of movement, which was known to be formally very unconstrained. By contrast, the ATN parser used in the LUNAR project (Woods, Kaplan, and Nash-Webber 1972) reduced all discontinuity to local operations on registers (Thorne, Bratley, and Dewar 1968; Bobrow and Fraser 1969; Woods 1970).In particular, unbounded wh-dependencies like the above were handled by: (a) putting a pointer into a * or HOLD register as soon as the “which” was encountered without regard to where it would end up; and (b) retrieving the pointer from HOLD when the verb needing an object “had” was encountered without regard to where it had started out. (It also included an ingenious mechanism for coordination called SYSCONJ, which one finds even now being reinvented on an almost yearly basis—cf. Woods [2010].) A * register was also used for wh-constructions within a systemic grammar framework by Winograd (1972, pages 52–53) in his inspiring conversational program SHRDLU.However, it was unclear how to generalize the HOLD register to handle the multiple long-range dependencies, including crossing dependencies, that are found in many other languages. In particular, if the HOLD register were assumed to be a stack, then the ATN becomes a two-stack machine (since we are already implicitly using one stack as a PDA to parse the context-free core grammar).On the computational side at least, the reaction to this impass took two distinct forms. Both reactions took the form of trying to reduce the two major operators of the transformation theory, substitution of immediate constituents, or what is nowadays called “Merge,” and “Move,” or displacement of non-immediate constituents, to one. On the one hand, Lexical Functional Grammar (Bresnan and Kaplan 1982) and Head-driven Phrase Structure Grammar (Pollard and Sag 1994) followed Kay (1979) in making unification the basis of movement and merger. Because unification can pass information across unbounded structures, this can be thought of as reducing Merge to Move.On the other hand, Generalized Phrase Structure Grammar (Gazdar 1981), Tree Adjoining Grammar (TAG; Joshi and Levy 1982), and Combinatory Categorial Grammar (CCG, Ades and Steedman, 1982) sought to reduce Move to various forms of local merger. In particular, the latter authors suggested that the same stack could be used to capture both long-range dependency and recursion in CCG.1Natural language grammar exhibits discontinuity because semantically language is an applicative system. Applicative systems (such as programming languages) support the twin notions of: (a) Application of a function/concept to an argument/entity; and (b) Abstraction, or the definition of a new function/concept in terms of existing ones.Language is in that sense inherently computational. It seems to follow that linguistics is (or should be) inherently computational as well. (Of course, it does not follow that computationalists have nothing to learn from linguistics.)There are two ways of modeling abstraction in applicative systems: Taking abstraction itself as a primitive operation (λ-calculus, LISP):(2)a.fatherEsau⇒Isaacb.grandfather=λx.father(fatherx)c.grandfatherEsau⇒Abrahamor Defining abstraction in terms of a collection of operators on strictly adjacent terms aka Combinators, such as function composition (Combinatory Calculus, MIRANDA).(3)b′.grandfather=BfatherfatherThe latter does the work of the λ-calculus without using any variables.Despite the resemblance of the “traces” (or copies) and “operators” (or complementizer positions) of the transformational theory to the λ-operators and variables of applicative systems of the first kind, natural language actually seems to be a system of the second, combinatory kind. The evidence stems from the fact that natural language deals with all sorts of fragments that linguists do not normally think of as semantically typable constituents, without the use of any phonologically realized equivalent of variables, such as pronouns:(4)a.Give[Anna books]?and[Manny records]?b.(Mother to child): There’s adoggie![Youlike]?#the doggie.c.Food that you must[washVP/NP[before eating](VP∖VP)/NP]?.d.ik denk dat ik1Henk2Cecilia3[zag1leren2zingen3]?These fragments are diagnostic of a Combinatory Calculus based on Bn, T, and the “duplicator” Sn, plus application (Steedman 1987; Szabolcsi 1989; Steedman and Baldridge all dependencies, such as and so logical form. syntactic are operators over phonologically realized and their logical forms. are restricted by a Combinatory which in they cannot the already in the language-specific but must be with and project the such language-specific information is in the The combinatory like composition are and such as and are to be over the as if they were as in the of and long-range are by of a such as with an adjacent with by of function combinatory composition of the syntactic shown with composition of logical forms in the to the logical forms shown as for the capture the in syntactic we pass over we also based on the capture like (whose syntactic is similarly we also based on composition saw to the of transformational theory to of adjacent to combinatory the form of the transformational theory has also proposed that Move should be as an form of or Merge though without any basis for the other than internal Merge as the it would be this of application and abstraction in a or all as in the to this to of such as As we saw in like and are using such to crossing CCG slightly than context-free CCG is not as expressive as In particular, we can only capture that are what is called where is to the of the by and (Steedman for the of the form and it is obvious by that we cannot recognize the following to be for the of this form for the these these of the of are The two are among the of this by is the of the is the of the number of ways of two of three by the about one in a the order that so were to be this would to about one in number of much more in than the number of all for around of the are There are obvious for the problem of in machine translation and neural semantic parsing, to which we I to and at first as a under least, that was what Joshi and then as a of where much of the development of CCG was out. students in a of and Joshi 1987; and that the of Ades and Steedman was by that both CCG and were equivalent to Grammar (Gazdar a new of the by the and of by these fell within the of what Joshi called which proposed as a for what could as a theory of natural that they are and and limit on crossing the is much much than the including the multiple and even the of so it seems to and CCG as with to the its CCG was assumed to be as a grammar for parsing, because of the derivational ambiguity by and the combinatory rules, these also under and as any grammar with the same as CCG the same of in the parser it is there in the In this is just another in the of derivational ambiguity that all natural language and can be handled by the same statistical models as other particular, the dependency models by and are and Steedman and CCG is also to parsing with which can be using and long and Steedman and is now used in those that for between semantic and syntactic processing, such as machine translation and and machine and parsing and Steedman 1987; and et al. et al. et al. and semantic parser and et al. et al. of my work with the same has returned to their application in musical and and Steedman and shown that CCG grammars of the same and parsing models of the same statistical are required there as It is only in the of their compositional semantics that music and language very than on this of work in like to by two more The first is an evolutionary should natural language be a combinatory in the first The second is a question about the future development of CCG and other grammar-based theories to be to NLP in the age of deep and neural take these questions in order in the two language like a combinatory applicative system because T, and evolved in to support of there was any language (Steedman like and are for of to to make can form the that you to is to that they and other can with like are to make with arbitrary of including and including other is yet to be to be to do the the and problem, using like to in order to that are of to of and so such as of and can apply to a and much work that the and other have to be there in the already for the to be to with was to that from the even if had the problem of can be as the problem of search for a of in a or of possible such search has the same as parser there are both and for the The latter a more in evolutionary as a mechanism that is to both semantic interpretation and parsing, rather than the evolution of like the the for linguistic as as the operators for competence grammar, the two to as what was to above as an evolutionary hope to have convinced you that CCG grammars are both and as as semantic parsers that are to parse the like other CCG parsers are by the of parsing models based on only a of no how we using and the they are and so as we have and off in the long As a parsers are to on performance overall by using models et al. more about the of grammars and parsers than about the of deep in the of semantic parser for arbitrary such as and Steedman I would that CCG and other grammar-based parsers have already been by of deep neural and and the question of whether models for in are because we have to the universal semantic that the child to CCG for natural and that we to be using both in semantic parser and in Because the language is any of linguistic logical like the universal language of it be more with and to semantic parsers for by deep neural force, rather than by CCG semantic parser is it possible that the problem of parsing could be by neural by stack et al. et al. actually learn as has been seems that semantic parsers and neural machine translation to have difficulty with long-range because the evidence for their is so both and as a case, verb or a complementizer I is actually by which that if you are an language with like and then you not in be to in like or like and you be to both and or at the time of a translation system no of learned these syntactic from a we get an sentence to translation means is the that the the a is the that the to the if we with we (which back into the a is the that the holds the a we with using a the translation again an that is is the that they said had the is the they said they the contrast, CCG parsers do rather on and Steedman Steedman, and which is a that to be rather determined by the parsing methods are similarly when with long-range even when the sentence is in this think that the the and that it is et think that the the and it is in at least, these could be learned by grammar-based semantic parser from using the methods of et al. and et al. make it that there be a for in like where long-range like deep and are to The future in parsing for such lies with systems using neural for and grammars for problem in NLP the fact that natural language understanding inference as as semantics, and we have no of the representation question is the almost it many it is almost equally to the information in a form that is not immediately with the form of the sentences like the following a different a rather than a an an a a and at a representation language that is we are using CCG parsers to the for between in order to of between over of the same using over then an it and it under such as et al. of in the then that can be to a relation and Steedman can be across from multiple and Steedman can then the semantics for relation with the and the entire using this now both and semantic an with the as and the as questions the in this we parse questions into the same semantics which is now the language of the the we use the and the and the following that in the then the the of is in the in the this to work, we to be in their in both and then be to like in of a semantic in the language of the to learn between semantic and the language of the this project is the the function of a of the semantic in semantic like those of and and Steedman and while the form a similarly of the of Carnap and Fodor, Fodor, and Garrett semantic are essentially but with the advantage that they can be with logical operators such as and for the of semantic underlying natural language semantics in to another different to semantics that to use reduced of to using operations such as and and in of compositional It is an question whether can be with to of a and et al. It is that of be to the of the of the long like and work in do they work in In particular, can they learn all the syntactic in the long like and crossing in a way that support semantic they are not actually but are a finite-state or a then by on as for natural language processing, we are in of of the computational linguistic project of also computational of language and if we that like is a real and that learn their first language by of the sentences of their language the of the universal language of we the of what that universal semantic language not get an to that question we can above using and and such as as for the language of to use machine for what it is such variables and their for use in a natural language work was supported in by a Award and a a University of Edinburgh and my and the and all my and students over many
Mark Steedman
Comput. Linguistics1
2018 Learning Typed Entailment Graphs with Global Soft Constraints
abstract
This paper presents a new method for learning typed entailment graphs from text. We extract predicate-argument structures from multiple-source news corpora, and compute local distributional similarity scores to learn entailments between predicates with typed arguments (e.g., person contracted disease). Previous work has used transitivity constraints to improve local decisions, but these constraints are intractable on large graphs. We instead propose a scalable method that learns globally consistent similarity scores based on new soft constraints that consider both the structures across typed entailment graphs and inside each graph. Learning takes only a few hours to run over 100K predicates and our results show large improvements over local similarity scores on two entailment data sets. We further show improvements over paraphrases and entailments from the Paraphrase Database, and prior state-of-the-art entailment graphs. We show that the entailment graphs improve performance in a downstream task.
Mohammad Javad Hosseini, Nathanael Chambers, Siva Reddy, Xavier R. Holt, Shay B. Cohen, Mark Johnson 0001, Mark Steedman
Trans. Assoc. Comput. Linguistics7
2017 Universal Semantic Parsing
abstract
Universal Dependencies (UD) offer a uniform cross-lingual syntactic representation, with the aim of advancing multilingual applications.Recent work shows that semantic parsing can be accomplished by transforming syntactic dependencies to logical forms.However, this work is limited to English, and cannot process dependency graphs, which allow handling complex phenomena such as control.In this work, we introduce UDEPLAMBDA, a semantic interface for UD, which maps natural language to logical forms in an almost language-independent fashion and can process dependency graphs.We perform experiments on question answering against Freebase and provide German and Spanish translations of the WebQuestions and GraphQuestions datasets to facilitate multilingual evaluation.Results show that UDEPLAMBDA outperforms strong baselines across languages and datasets.For English, it achieves a 4.9 F 1 point improvement over the state-of-the-art on Graph-Questions.ENTITY ⇒ λx.word(x a ); e.g.Oscar ⇒ λx.Oscar(x a ) EVENT ⇒ λx.word(x e ); e.g. won ⇒ λx.won(x e ) FUNCTIONAL ⇒ λx.TRUE; e.g. an ⇒ λx.TRUE COPY ⇒ λ f gx.∃y.f (x) ∧ g(y) ∧ rel(x, y) e.g.nsubj, dobj, nmod, advmod INVERT ⇒ λ f gx.∃y.f (x) ∧ g(y) ∧ rel i (y, x) e.g.amod, acl MERGE ⇒ λ f gx.f (x) ∧ g(x) e.g.compound, appos, amod, acl HEAD ⇒ λ f gx.f (x) e.g.case, punct, aux, mark .
Siva Reddy, Oscar Täckström, Slav Petrov, Mark Steedman, Mirella Lapata
EMNLP4
2016 Evaluating Induced CCG Parsers on Grounded Semantic Parsing
abstract
We compare the effectiveness of four different syntactic CCG parsers for a semantic slotfilling task to explore how much syntactic supervision is required for downstream semantic analysis.This extrinsic, task-based evaluation also provides a unique window into the semantics captured (or missed) by unsupervised grammar induction systems.
Yonatan Bisk, Siva Reddy, John Blitzer, Julia Hockenmaier, Mark Steedman
EMNLP5
2016 Shift-Reduce CCG Parsing using Neural Network Models
abstract
Bharat Ram Ambati, Tejaswini Deoskar, Mark Steedman. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Bharat Ram Ambati, Tejaswini Deoskar, Mark Steedman
HLT-NAACL3
2016 Assessing Relative Sentence Complexity using an Incremental CCG Parser
abstract
Given a pair of sentences, we present computational models to assess if one sentence is simpler to read than the other.While existing models explored the usage of phrase structure features using a non-incremental parser, experimental evidence suggests that the human language processor works incrementally.We empirically evaluate if syntactic features from an incremental CCG parser are more useful than features from a non-incremental phrase structure parser.Our evaluation on Simple and Standard Wikipedia sentence pairs suggests that incremental CCG features are indeed more useful than phrase structure features achieving 0.44 points gain in performance.Incremental CCG parser also gives significant improvements in speed (12 times faster) in comparison to the phrase structure parser.Furthermore, with the addition of psycholinguistic features, we achieve the strongest result to date reported on this task.
Bharat Ram Ambati, Siva Reddy, Mark Steedman
HLT-NAACL3
2016 Transforming Dependency Structures to Logical Forms for Semantic Parsing
abstract
The strongly typed syntax of grammar formalisms such as CCG, TAG, LFG and HPSG offers a synchronous framework for deriving syntactic structures and semantic logical forms. In contrast—partly due to the lack of a strong type system—dependency structures are easy to annotate and have become a widely used form of syntactic analysis for many languages. However, the lack of a type system makes a formal mechanism for deriving logical forms from dependency structures challenging. We address this by introducing a robust system based on the lambda calculus for deriving neo-Davidsonian logical forms from dependency trees. These logical forms are then used for semantic parsing of natural language to Freebase. Experiments on the Free917 and Web-Questions datasets show that our representation is superior to the original dependency trees and that it outperforms a CCG-based representation on this task. Compared to prior work, we obtain the strongest result to date on Free917 and competitive results on WebQuestions.
Siva Reddy, Oscar Täckström, Michael Collins 0001, Tom Kwiatkowski, Dipanjan Das 0001, Mark Steedman, Mirella Lapata
Trans. Assoc. Comput. Linguistics6
2015 Orthogonality of Syntax and Semantics within Distributional Spaces
abstract
Jeff Mitchell, Mark Steedman. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Jeff Mitchell 0001, Mark Steedman
ACL (1)2
2015 A Computationally Efficient Algorithm for Learning Topical Collocation Models
abstract
Zhendong Zhao, Lan Du, Benjamin Börschinger, John K Pate, Massimiliano Ciaramita, Mark Steedman, Mark Johnson. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Zhendong Zhao, Lan Du 0002, Benjamin Börschinger, John K. Pate, Massimiliano Ciaramita, Mark Steedman, Mark Johnson 0001
ACL (1)6
2015 Lexical Event Ordering with an Edge-Factored Model
abstract
Extensive lexical knowledge is necessary for temporal analysis and planning tasks.We address in this paper a lexical setting that allows for the straightforward incorporation of rich features and structural constraints.We explore a lexical event ordering task, namely determining the likely temporal order of events based solely on the identity of their predicates and arguments.We propose an "edgefactored" model for the task that decomposes over the edges of the event graph.We learn it using the structured perceptron.As lexical tasks require large amounts of text, we do not attempt manual annotation and instead use the textual order of events in a domain where this order is aligned with their temporal order, namely cooking recipes.
Omri Abend, Shay B. Cohen, Mark Steedman
HLT-NAACL3
2015 An Incremental Algorithm for Transition-based CCG Parsing
abstract
Bharat Ram Ambati, Tejaswini Deoskar, Mark Johnson, Mark Steedman. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Bharat Ram Ambati, Tejaswini Deoskar, Mark Johnson 0001, Mark Steedman
HLT-NAACL4
2014 Lexical Inference over Multi-Word Predicates: A Distributional Approach
abstract
Representing predicates in terms of their argument distribution is common practice in NLP.Multi-word predicates (MWPs) in this context are often either disregarded or considered as fixed expressions.The latter treatment is unsatisfactory in two ways: (1) identifying MWPs is notoriously difficult, (2) MWPs show varying degrees of compositionality and could benefit from taking into account the identity of their component parts.We propose a novel approach that integrates the distributional representation of multiple sub-sets of the MWP's words.We assume a latent distribution over sub-sets of the MWP, and estimate it relative to a downstream prediction task.Focusing on the supervised identification of lexical inference relations, we compare against state-of-the-art baselines that consider a single sub-set of an MWP, obtaining substantial improvements.To our knowledge, this is the first work to address lexical relations between MWPs of varying degrees of compositionality within distributional semantics.
Omri Abend, Shay B. Cohen, Mark Steedman
ACL (1)3
2014 Improving Dependency Parsers using Combinatory Categorial Grammar
abstract
Subcategorization information is a useful feature in dependency parsing. In this paper, we explore a method of incorporating this information via Combinatory Categorial Grammar (CCG) categories from a supertagger. We experiment with two popular dependency parsers (Malt and MST) for two languages: English and Hindi. For both languages, CCG categories improve the overall accuracy of both parsers by around 0.3-0.5% in all experiments. For both parsers, we see larger improvements specifically on dependencies at which they are known to be weak: long distance dependencies for Malt, and verbal arguments for MST. The result is particularly interesting in the case of the fast greedy parser (Malt), since improving its accuracy without significantly compromising speed is relevant for large scale applications such as parsing the web.
Bharat Ram Ambati, Tejaswini Deoskar, Mark Steedman
EACL3
2014 Generalizing a Strongly Lexicalized Parser using Unlabeled Data
abstract
Statistical parsers trained on labeled data suffer from sparsity, both grammatical and lexical.For parsers based on strongly lexicalized grammar formalisms (such as CCG, which has complex lexical categories but simple combinatory rules), the problem of sparsity can be isolated to the lexicon.In this paper, we show that semi-supervised Viterbi-EM can be used to extend the lexicon of a generative CCG parser.By learning complex lexical entries for low-frequency and unseen words from unlabeled data, we obtain improvements over our supervised model for both indomain (WSJ) and out-of-domain (questions and Wikipedia) data.Our learnt lexicons when used with a discriminative parser such as C&C also significantly improve its performance on unseen words.
Tejaswini Deoskar, Christos Christodoulopoulos 0001, Alexandra Birch, Mark Steedman
EACL4
2014 A Generative Model for User Simulation in a Spatial Navigation Domain
abstract
We propose the use of a generative model to simulate user behaviour in a novel taskoriented dialog domain, where user goals are spatial routes across artificial landscapes.We show how to derive an efficient feature-based representation of spatial goals, admitting exact inference and generalising to new routes.The use of a generative model allows us to capture a range of plausible behaviour given the same underlying goal.We evaluate intrinsically using held-out probability and perplexity, and find a substantial reduction in uncertainty brought by our spatial representation.We evaluate extrinsically in a human judgement task and find that our model's behaviour does not differ significantly from the behaviour of real users.
Aciel Eshky, Ben Allison, Subramanian Ramamoorthy, Mark Steedman
EACL4
2014 A* CCG Parsing with a Supertag-factored Model
abstract
We introduce a new CCG parsing model which is factored on lexical category assignments.Parsing is then simply a deterministic search for the most probable category sequence that supports a CCG derivation.The parser is extremely simple, with a tiny feature set, no POS tagger, and no statistical model of the derivation or dependencies.Formulating the model in this way allows a highly effective heuristic for A * parsing, which makes parsing extremely fast.Compared to the standard C&C CCG parser, our model is more accurate out-of-domain, is four times faster, has higher coverage, and is greatly simplified.We also show that using our parser improves the performance of a state-ofthe-art question answering system. 1
Mike Lewis, Mark Steedman
EMNLP2
2014 Extracting common sense knowledge from text for robot planning
abstract
Autonomous robots often require domain knowledge to act intelligently in their environment. This is particularly true for robots that use automated planning techniques, which require symbolic representations of the operating environment and the robot's capabilities. However, the task of specifying domain knowledge by hand is tedious and prone to error. As a result, we aim to automate the process of acquiring general common sense knowledge of objects, relations, and actions, by extracting such information from large amounts of natural language text, written by humans for human readers. We present two methods for knowledge acquisition, requiring only limited human input, which focus on the inference of spatial relations from text. Although our approach is applicable to a range of domains and information, we only consider one type of knowledge here, namely object locations in a kitchen environment. As a proof of concept, we test our approach using an automated planner and show how the addition of common sense knowledge can improve the quality of the generated plans.
Peter Kaiser 0001, Mike Lewis, Ronald P. A. Petrick, Tamim Asfour, Mark Steedman
ICRA5
2014 Robust Semantics for Semantic Parsing
Mark Steedman
PACLIC1
2014 Improved CCG Parsing with Semi-supervised Supertagging
abstract
Current supervised parsers are limited by the size of their labelled training data, making improving them with unlabelled data an important goal. We show how a state-of-the-art CCG parser can be enhanced, by predicting lexical categories using unsupervised vector-space embeddings of words. The use of word embeddings enables our model to better generalize from the labelled data, and allows us to accurately assign lexical categories without depending on a POS-tagger. Our approach leads to substantial improvements in dependency parsing results over the standard supervised CCG parser when evaluated on Wall Street Journal (0.8%), Wikipedia (1.8%) and biomedical (3.4%) text. We compare the performance of two recently proposed approaches for classification using a wide variety of word embeddings. We also give a detailed error analysis demonstrating where using embeddings outperforms traditional feature sets, and showing how including POS features can decrease accuracy.
Mike Lewis, Mark Steedman
Trans. Assoc. Comput. Linguistics2
2014 Large-scale Semantic Parsing without Question-Answer Pairs
abstract
In this paper we introduce a novel semantic parsing approach to query Freebase in natural language without requiring manual annotations or question-answer pairs. Our key insight is to represent natural language via semantic graphs whose topology shares many commonalities with Freebase. Given this representation, we conceptualize semantic parsing as a graph matching problem. Our model converts sentences to semantic graphs using CCG and subsequently grounds them to Freebase guided by denotations as a form of weak supervision. Evaluation experiments on a subset of the Free917 and WebQuestions benchmark datasets show our semantic parser improves over the state of the art.
Siva Reddy, Mirella Lapata, Mark Steedman
Trans. Assoc. Comput. Linguistics3
2013 Unsupervised Induction of Cross-Lingual Semantic Relations
abstract
Creating a language-independent meaning representation would benefit many crosslingual NLP tasks.We introduce the first unsupervised approach to this problem, learning clusters of semantically equivalent English and French relations between referring expressions, based on their named-entity arguments in large monolingual corpora.The clusters can be used as language-independent semantic relations, by mapping clustered expressions in different languages onto the same relation.Our approach needs no parallel text for training, but outperforms a baseline that uses machine translation on a cross-lingual question answering task.We also show how to use the semantics to improve the accuracy of machine translation, by using it in a simple reranker.
Mike Lewis, Mark Steedman
EMNLP2
2013 Combined Distributional and Logical Semantics
abstract
We introduce a new approach to semantics which combines the benefits of distributional and formal logical semantics. Distributional models have been successful in modelling the meanings of content words, but logical semantics is necessary to adequately represent many function words. We follow formal semantics in mapping language to logical representations, but differ in that the relational constants used are induced by offline distributional clustering at the level of predicate-argument structure. Our clustering algorithm is highly scalable, allowing us to run on corpora the size of Gigaword. Different senses of a word are disambiguated based on their induced types. We outperform a variety of existing approaches on a wide-coverage question answering task, and demonstrate the ability to make complex multi-sentence inferences involving quantifiers on the FraCaS suite.
Mike Lewis, Mark Steedman
Trans. Assoc. Comput. Linguistics2
2012 A Probabilistic Model of Syntactic and Semantic Acquisition from Child-Directed Utterances and their Meanings
Tom Kwiatkowski, Sharon Goldwater, Luke Zettlemoyer, Mark Steedman
EACL4
2012 Generative Goal-Driven User Simulation for Dialog Management
Aciel Eshky, Ben Allison, Mark Steedman
EMNLP-CoNLL3
2012 Learning STRIPS Operators from Noisy and Incomplete Observations
Kira Mourão, Luke Zettlemoyer, Ronald P. A. Petrick, Mark Steedman
UAI4
2011 A Bayesian Mixture Model for PoS Induction Using Multiple Features
Christos Christodoulopoulos 0001, Sharon Goldwater, Mark Steedman
EMNLP3
2011 Lexical Generalization in CCG Grammar Induction for Semantic Parsing
Tom Kwiatkowski, Luke Zettlemoyer, Sharon Goldwater, Mark Steedman
EMNLP4
2011 Semi-supervised CCG Lexicon Extension
Emily Thomforde, Mark Steedman
EMNLP2
2011 Grammar Induction from Text Using Small Syntactic Prototypes
Prachya Boonkwan, Mark Steedman
IJCNLP2
2010 Learning action effects in partially observable domains
abstract
We investigate the problem of learning action effects in partially observable STRIPS planning domains. Our approach is based on a voted kernel perceptron learning model, where action and state information is encoded in a compact vector representation as input to the learning mechanism, and resulting state changes are produced as output. Our approach relies on deictic features that assume an attentional mechanism that reduces the size of the representation. We evaluate our approach on a number of partially observable planning domains, and show that it can quickly learn the dynamics of such domains, with low average error rates. We show that our approach handles noisy domains, conditional effects, and that it scales independently of the number of objects in a domain.
Kira Mourão, Ronald P. A. Petrick, Mark Steedman
ECAI3
2010 Two Decades of Unsupervised POS Induction: How Far Have We Come?
Christos Christodoulopoulos 0001, Sharon Goldwater, Mark Steedman
EMNLP3
2010 Inducing Probabilistic CCG Grammars from Logical Form with Higher-Order Unification
Tom Kwiatkowski, Luke Zettlemoyer, Sharon Goldwater, Mark Steedman
EMNLP4
2010 A Multi-Dimensional Analysis of Japanese Benefactives: The Case of the Yaru-Construction
Akira Ohtani, Mark Steedman
PACLIC2
2009 Unbounded Dependency Recovery for Parser Evaluation
Laura Rimell, Stephen Clark, Mark Steedman
EMNLP3
2009 Note on Japanese Epistemic Verb Constructions: A Surface-Compositional Analysis
Akira Ohtani, Mark Steedman
PACLIC2
2008 On Japanese Desiderative Constructions
Akira Ohtani, Mark Steedman
PACLIC2
2008 The Grammar of Scope
Mark Steedman
WoLLIC1
2008 On Becoming a Discipline
abstract
The title of this column, Last Words, reminds me of an occasion in 2005, when I had the privilege of attending the award ceremony for the prestigious Benjamin Franklin Medal, given annually to a few scientists who have made outstanding lifetime contributions to science.This time, a computational linguist, Aravind Joshi, was among them, so several past, present, and future presidents and officers of the ACL joined the Great and the Good at the ceremony at the Franklin Institute in Philadelphia.The eight medal recipients were each represented by a short video presentation, which mostly consisted of voice-over by a narrator, interspersed with sound-bites from the recipients about their life and work, in the last of which they had clearly been asked to deliver as their last words a brief take-home message.I couldn't help noticing that the warmest applause was reserved for the physicist, a distinguished pioneer of string theory.I was initially puzzled by the enthusiasm on the part of a mostly lay audience for such theoretical work, which for all its elegance and beauty, could not (as far as I could see) be expected to have nearly as much impact on their everyday lives as that of some of the other recipients, who that year included not only Aravind, but another computer scientist whose impact on information processing will be obvious to the members of ACL, Andrew Viterbi.But then I recalled that the physicist's take-home message had had nothing to do with string theory.This admirable man's last words to us had been the following:Everything is made of particles.So physics is very important.
Mark Steedman
Comput. Linguistics1
2007 On Natural Language Processing and Plan Recognition
Christopher W. Geib, Mark Steedman
IJCAI2
2007 Case, Coordination, and Information Structure in Japanese
Ahkira Otani, Mark Steedman
PACLIC2
2007 CCGbank: A Corpus of CCG Derivations and Dependency Structures Extracted from the Penn Treebank
abstract
This article presents an algorithm for translating the Penn Treebank into a corpus of Combinatory Categorial Grammar (CCG) derivations augmented with local and long-range word-word dependencies. The resulting corpus, CCGbank, includes 99.4% of the sentences in the Penn Treebank. It is available from the Linguistic Data Consortium, and has been used to train wide-coverage statistical parsers that obtain state-of-the-art rates of dependency recovery. In order to obtain linguistically adequate CCG analyses, and to eliminate noise and inconsistencies in the original annotation, an extensive analysis of the constructions and annotations in the Penn Treebank was called for, and a substantial number of changes to the Treebank were necessary. We discuss the implications of our findings for the extraction of other linguistically expressive grammars from the Treebank, and for the design of future treebanks.
Julia Hockenmaier, Mark Steedman
Comput. Linguistics2
2004 Wide-Coverage Semantic Representations from a CCG Parser
Johan Bos, Stephen Clark, Mark Steedman, James R. Curran, Julia Hockenmaier
COLING3
2004 Object-Extraction and Question-Parsing using CCG
Stephen Clark, Mark Steedman, James R. Curran
EMNLP2
2004 An Annotation Scheme for Information Status in Dialogue
Malvina Nissim, Shipra Dingare, Jean Carletta, Mark Steedman
LREC4
2003 Bootstrapping statistical parsers from small datasets
Mark Steedman, Anoop Sarkar, Miles Osborne, Rebecca Hwa, Stephen Clark, Julia Hockenmaier, Paul Ruhlen, Jeremiah Crim
EACL1
2003 Example Selection for Bootstrapping Statistical Parsers
Mark Steedman, Rebecca Hwa, Stephen Clark, Miles Osborne, Anoop Sarkar, Julia Hockenmaier, Paul Ruhlen, Jeremiah Crim
HLT-NAACL1
2002 Building Deep Dependency Structures using a Wide-Coverage CCG Parser
abstract
This paper describes a wide-coverage statistical parser that uses Combinatory Categorial Grammar (CCG) to derive dependency structures. The parser differs from most existing wide-coverage treebank parsers in capturing the long-range dependencies inherent in constructions such as coordination, extraction, raising and control, as well as the standard local predicate-argument dependencies. A set of dependency structures used for training and testing the parser is obtained from a treebank of CCG normal-form derivations, which have been derived (semi-) automatically from the Penn Treebank. The parser correctly recovers over 80% of labelled dependencies, and around 90% of unlabelled dependencies.
Stephen Clark, Julia Hockenmaier, Mark Steedman
ACL3
2002 Generative Models for Statistical Parsing with Combinatory Categorial Grammar
abstract
This paper compares a number of generative probability models for a wide-coverage Combinatory Categorial Grammar (CCG) parser. These models are trained and tested on a corpus obtained by translating the Penn Treebank trees into CCG normal-form derivations. According to an evaluation of unlabeled word-word dependencies, our best model achieves a performance of 89.9%, comparable to the figures given by Collins (1999) for a linguistically less expressive grammar. In contrast to Gildea (2001), we find a significant improvement from modeling word-word dependencies.
Julia Hockenmaier, Mark Steedman
ACL2
2002 Acquiring Compact Lexicalized Grammars from a Cleaner Treebank
Julia Hockenmaier, Mark Steedman
LREC2
1999 Alternating Quantifier Scope in CCG
abstract
The paper shows that movement or equivalent computational structure-changing operations of any kind at the level of logical form can be dispensed with entirely in capturing quantifier scope ambiguity. It offers a new semantics whereby the effects of quantifier scope alternation can be obtained by an entirely monotonic derivation, without type-changing rules. The paper follows Fodor (1982), Fodor and Sag (1982), and Park (1995, 1996) in viewing many apparent scope ambiguities as arising from referential categories rather than true generalized quantifiers.
Mark Steedman
ACL1
1995 Dynamic Semantics for Tense and Aspect
Mark Steedman
IJCAI1
1994 Animated conversation: rule-based generation of facial expression, gesture & spoken intonation for multiple conversational agents
abstract
We describe an implemented system which automatically generates and animates conversations between multiple human-like agents with appropriate and synchronized speech, intonation, facial expressions, and hand gestures. Conversation is created by a dialogue planner that produces the text as well as the intonation of the utterances. The speaker/listener relationship, the text, and the intonation in turn drive facial expressions, lip motions, eye gaze, head motion, and arm gestures generators. Coordinated arm, wrist, and hand motions are invoked to create semantically meaningful gestures. Throughout we will use examples from an actual synthesized, fully animated conversation.
Justine Cassell, Catherine Pelachaud, Norman I. Badler, Mark Steedman, Brett Achorn, Tripp Becket, Brett Douville, Scott Prevost, Matthew Stone
SIGGRAPH4
1994 Specifying intonation from context for speech synthesis
Scott Prevost, Mark Steedman
Speech Communication2
1993 Generating Contextually Appropriate Intonation
Scott Prevost, Mark Steedman
EACL2
1993 Using context to specify intonation in speech synthesis
abstract
A generator based on Combinatory Categorial Grammar using a simple and domain-independent discourse model can be used to direct synthesis of intonation contours for responses to data-base queries, conveying distinctions of contrast and emphasis determined by the discourse model and the state of the knowledge-base. 1 INTRODUCTION One source of unnaturalness in the output of many text-tospeech systems stems from the involvement of algorithmically generated default intonation contours, applied under minimal control from syntax and semantics. The intelligibility of the speech produced by these systems is a tribute to both the resilience of human language understanding and the ingenuity of the algorithms. It has often been noted, however, that the results frequently sound unnatural when taken in context, and may occasionally mislead the hearer. It is for this reason that a number of discourse-model-based speech generation systems have been proposed, in which intonation contour is determi...
Scott Prevost, Mark Steedman
EUROSPEECH2
1991 Type-Raising and Directionality in Combinatory Grammar
abstract
The form of rules in combinatory categorial grammars (CCG) is constrained by three principles, called "adjacency", "consistency" and "inheritance". These principles have been claimed elsewhere to constrain the combinatory rules of composition and type raising in such a way as to make certain linguistic universals concerning word order under coordination follow immediately. The present paper shows that the three principles have a natural expression in a unification-based interpretation of CCG in which directional information is an attribute of the arguments of functions grounded in string position. The universals can thereby be derived as consequences of elementary assumptions. Some desirable results for grammars and parsers follow, concerning type-raising rules.
Mark Steedman
ACL1
1990 Structure and Intonation in Spoken Language Undestanding
abstract
The structure imposed upon spoken sentences by intonation seems frequently to be orthogonal to their traditional surface-syntactic structure. However, the notion of "intonational structure" as formulated by Pierrehumbert, Selkirk, and others, can be subsumed under a rather different notion of syntactic surface structure that emerges from a theory of grammar based on a "Combinatory" extension to Categorial Grammar. Interpretations of constituents at this level are in turn directly related to "information structure", or discourse-related notions of "theme", "rheme", "focus" and "presupposition". Some simplifications appear to follow for the problem of integrating syntax and other high-level modules in spoken language systems.
Mark Steedman
ACL1
1990 Narrated Animation: A Case for Generation
Norman I. Badler, Mark Steedman, Bonnie L. Webber
INLG2
1988 Temporal Ontology and Temporal Reference
Marc Moens, Mark Steedman
Comput. Linguistics2
1987 Temporal Ontology in Natural Language
abstract
A semantics of linguistic categories like tense, aspect, and certain temporal adverbials, and a theory of their use in defining the temporal relations of events, both require a more complex structure on the domain underlying the meaning representations than is commonly assumed. The paper proposes an ontology based on such notions as causation and consequence, rather than on purely temporal primitives. We claim that any manageable logic or other formal system for natural language temporal descriptions will have to embody such an ontology, as will any usable temporal database for knowledge about events which is to be interrogated using natural language.
Marc Moens, Mark Steedman
ACL2
1987 A Lazy way to Chart-Parse with Categorial Grammars
abstract
There has recently been a revival of interest in Categorial Grammars (CG) among computational linguists. The various versions noted below which extend pure CG by including operations such as functional composition have been claimed to offer simple and uniform accounts of a wide range of natural language (NL) constructions involving bounded and unbounded "movement" and coordination "reduction" in a number of languages. Such grammars have obvious advantages for computational applications, provided that they can be parsed efficiently. However, many of the proposed extensions engender proliferating semantically equivalent surface syntactic analyses. These "spurious analyses" have been claimed to compromise their efficient parseability.The present paper describes a simple parsing algorithm for our own "combinatory" extension of CG. This algorithm offers a uniform treatment for "spurious" syntactic ambiguities and the "genuine" structural ambiguities which any processor must cope with, by exploiting the associativity of functional composition and the procedural neutrality of the combinatory rules of grammar in a bottom-up, left-to-right parser which delivers all semantically distinct analyses via a novel unification-based extension of chart-parsing.
Remo Pareschi, Mark Steedman
ACL2