Nathan Schneider 0001

dblp:31/9014 · DBLP profile ↗
← Back
51ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0002-5994-671XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 7 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractions
abstract
Francesca Padovani, Xiulin Yang, Bastian Bunzeck, Jaap Jumelet, Yevgen Matusevych, Nathan Schneider, Arianna Bisazza. Proceedings of the 30th Conference on Computational Natural Language Learning. 2026.
Francesca Padovani, Xiulin Yang, Bastian Bunzeck, Jaap Jumelet, Yevgen Matusevych, Nathan Schneider 0001, Arianna Bisazza
CoNLL6
2026 Sense and Sensitivity: "Reasoning" Models are More Robust, but can Diverge from Human Consensus in a Legal Interpretation Task
abstract
Can LLMs make metalinguistic judgments?While LLM embeddings are often regarded as high-quality semantic representations, it is not clear that prompting an LLM is a useful way to obtain metalinguistic insights (e.g., whether a DIY gun kit is a "firearm").While some prior work has suggested LLM prompting can simulate surveys with human participants, computational studies in the domain of legal interpretation have found that LLMs are unreliable for metalinguistic judgments due to prompt sensitivity.However, these studies did not directly compare humans and LLMs on identical tasks, nor did they test so-called "reasoning" models.The current study addresses these gaps by directly comparing the robustness of human and LLM judgments (with and without reasoning) in an English-language legal interpretation task.Our results show that LLMs were more sensitive to irrelevant prompt features compared to human participants.Enabling reasoning improved the stability of LLM responses.However, even reasoning model outputs had only moderate correlations with human judgments, and all models sometimes output interpretations that no humans reached in response to the same prompt.We conclude that while reasoning decreases prompt sensitivity, LLMs are still poor proxies for human metalinguistic judgments.
Dawson Petersen, Abhishek Purushothama, Nathan Schneider 0001
CoNLL3
2026 Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
abstract
Grasping the semantics of rare constructions (form-meaning pairings) has been shown to be a challenging problem that has currently only been solved by the largest LLMs.It remains an open question if open-source models have robust constructional understanding, and if so, what learning dynamics underlie the acquisition of this knowledge.Focusing on a set of rare PAIRED-FOCUS constructions in English (e.g."let alone", "much less"), we construct a novel dataset to test their meanings using both scalar adjectival semantics and general world knowledge.Testing a wide range of models differing in parameter count, architecture, and pretraining dataset size, we find that several modestly sized models are sensitive to both the forms and the meanings of PAIRED-FOCUS constructions, though models trained on human-scale data fail at all meaning evaluations.Turning to training dynamics for a set of open-checkpoint models, we find that PAIRED-FOCUS understanding emerges later in training than PAIRED-FOCUS syntactic knowledge, and that learning of PAIRED-FOCUS semantics is correlated with gains in some domains of world knowledge.Overall, our empirical results support the conclusion that modestly sized open-source models can grasp the rare PAIRED-FOCUS constructions, and demonstrate a connection between knowledge of PAIRED-FOCUS constructions and other meaning domains.
Wesley Scivetti, Ethan Wilcox, Nathan Schneider 0001, Kanishka Misra, Leonie Weissweiler
CoNLL3
2025 Scope Ambiguity Resolution of Negated Connectives in English Corpora
Micaela Wells, Brandon Waldon, Nathan Schneider 0001
CogSci3
2025 Multilingual Supervision Improves Semantic Disambiguation of Adpositions
abstract
Adpositions display a remarkable amount of ambiguity and flexibility in their meanings, and are used in different ways across languages. We conduct a systematic corpus-based cross-linguistic investigation into the lexical semantics of adpositions, utilizing SNACS (Schneider et al., 2018), an annotation framework with data available in several languages. Our investigation encompasses 5 of these languages: Chinese, English, Gujarati, Hindi, and Japanese. We find substantial distributional differences in adposition semantics, even in comparable corpora. We further train classifiers to disambiguate adpositions in each of our languages. Despite the cross-linguistic differences in adpositional usage, sharing annotated data across languages boosts overall disambiguation performance, leading to the highest published scores on this task for all 5 languages.
Wesley Scivetti, Lauren Levine, Nathan Schneider 0001
COLING3
2025 Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning
abstract
Humans have a remarkable ability to acquire and understand grammatical phenomena that are seen rarely, if ever, during childhood.Recent evidence suggests that language models with human-scale pretraining data may possess a similar ability by generalizing from frequent to rare constructions.However, it remains an open question how widespread this generalization ability is, and to what extent this knowledge extends to meanings of rare constructions, as opposed to just their forms.We fill this gap by testing human-scale transformer language models on their knowledge of both the form and meaning of the (rare and quirky) English LET-ALONE construction.To evaluate our LMs we construct a bespoke synthetic benchmark that targets syntactic and semantic properties of the construction.We find that human-scale LMs are sensitive to form, even when related constructions are filtered from the dataset.However, human-scale LMs do not make correct generalizations about LET-ALONE's meaning.These results point to an asymmetry in the current architectures' sample efficiency between language form and meaning, something which is not present in human language learners.1
Wesley Scivetti, Tatsuya Aoyama, Ethan Wilcox, Nathan Schneider 0001
EMNLP4
2024 J-SNACS: Adposition and Case Supersenses for Japanese Joshi
abstract
Many languages use adpositions (prepositions or postpositions) to mark a variety of semantic relations, with different languages exhibiting both commonalities and idiosyncrasies in the relations grouped under the same lexeme. We present the first Japanese extension of the SNACS framework (Schneider et al., 2018), which has served as the basis for annotating adpositions in corpora from several languages. After establishing which of the set of particles (joshi) in Japanese qualify as case markers and adpositions as defined in SNACS, we annotate 10 chapters (≈10k tokens) of the Japanese translation of Le Petit Prince (The Little Prince), achieving high inter-annotator agreement. We find that, while a majority of the particles and their uses are captured by the existing and extended SNACS annotation guidelines from the previous work, some unique cases were observed. We also conduct experiments investigating the cross-lingual similarity of adposition and case marker supersenses, showing that the language-agnostic SNACS framework captures similarities not clearly observed in multilingual embedding space.
Tatsuya Aoyama, Chihiro Taguchi, Nathan Schneider 0001
LREC/COLING3
2024 CuRIAM: Corpus Re Interpretation and Metalanguage in U.S. Supreme Court Opinions
abstract
Most judicial decisions involve the interpretation of legal texts. As such, judicial opinions use language as the medium to comment on or draw attention to other language (for example, through definitions and hypotheticals about the meaning of a term from a statute). Language used this way is called metalanguage. Focusing on the U.S. Supreme Court, we view metalanguage as reflective of justices’ interpretive processes, bearing on current debates and theories about textualism in law and political science. As a step towards large-scale metalinguistic analysis with NLP, we identify 9 categories prominent in metalinguistic discussions, including key terms, definitions, and different kinds of sources. We annotate these concepts in a corpus of U.S. Supreme Court opinions. Our analysis of the corpus reveals high interannotator agreement, frequent use of quotes and sources, and several notable frequency differences between majority, concurring, and dissenting opinions. We observe fewer instances than expected of several legal interpretive categories. We discuss some of the challenges in developing the annotation schema and applying it and provide recommendations for how this corpus can be used for broader analyses.
Michael Kranzlein, Nathan Schneider 0001, Kevin Tobia
LREC/COLING2
2024 UCxn: Typologically-Informed Annotation of Constructions Atop Universal Dependencies
abstract
The Universal Dependencies (UD) project has created an invaluable collection of treebanks with contributions in over 140 languages. However, the UD annotations do not tell the full story. Grammatical constructions that convey meaning through a particular combination of several morphosyntactic elements—for example, interrogative sentences with special markers and/or word orders—are not labeled holistically. We argue for (i) augmenting UD annotations with a ‘UCxn’ annotation layer for such meaning-bearing grammatical constructions, and (ii) approaching this in a typologically informed way so that morphosyntactic strategies can be compared across languages. As a case study, we consider five construction families in ten languages, identifying instances of each construction in UD treebanks through the use of morphosyntactic patterns. In addition to findings regarding these particular constructions, our study yields important insights on methodology for describing and identifying constructions in language-general and language-particular ways, and lays the foundation for future constructional enrichment of UD treebanks.
Leonie Weissweiler, Nina Böbel, Kirian Guiller, Santiago Herrera, Wesley Scivetti, Arthur Lorenzi Almeida, Nurit Melnik, Archna Bhatia, Hinrich Schütze, Lori S. Levin, Amir Zeldes, Joakim Nivre, William Croft 0001, Nathan Schneider 0001
LREC/COLING14
2024 Lost in Translationese? Reducing Translation Effect Using Abstract Meaning Representation
abstract
Translated texts bear several hallmarks distinct from texts originating in the language.Though individual translated texts are often fluent and preserve meaning, at a large scale, translated texts have statistical tendencies which distinguish them from text originally written in the language ("translationese") and can affect model performance.We frame the novel task of translationese reduction and hypothesize that Abstract Meaning Representation (AMR), a graph-based semantic representation which abstracts away from the surface form, can be used as an interlingua to reduce the amount of translationese in translated texts.By parsing English translations into an AMR and then generating text from that AMR, the result more closely resembles originally English text across three quantitative macro-level measures, without severely compromising fluency or adequacy.We compare our AMR-based approach against three other techniques based on machine translation or paraphrase generation.This work makes strides towards reducing translationese in text and highlights the utility of AMR as an interlingua.
Shira Wein, Nathan Schneider 0001
EACL (1)2
2024 Modeling Nonnative Sentence Processing with L2 Language Models
abstract
We study LMs pretrained sequentially on two languages ("L2LMs") for modeling nonnative sentence processing.In particular, we pretrain GPT2 on 6 different first languages (L1s), followed by English as the second language (L2).We examine the effect of the choice of pretraining L1 on the model's ability to predict human reading times, evaluating on English readers from a range of L1 backgrounds.Experimental results show that, while all of the LMs' word surprisals improve prediction of L2 reading times, especially for human L1s distant from English, there is no reliable effect of the choice of L2LM's L1.We also evaluate the learning trajectory of a monolingual English LM: for predicting L2 as opposed to L1 reading, it peaks much earlier and immediately falls off, possibly mirroring the difference in proficiency between the native and nonnative populations.Lastly, we provide examples of L2LMs' surprisals, which could potentially generate hypotheses about human L2 reading.
Tatsuya Aoyama, Nathan Schneider 0001
EMNLP2
2024 Assessing the Cross-linguistic Utility of Abstract Meaning Representation
abstract
Abstract Semantic representations capture the meaning of a text. Abstract Meaning Representation (AMR), a type of semantic representation, focuses on predicate-argument structure and abstracts away from surface form. Though AMR was developed initially for English, it has now been adapted to a multitude of languages in the form of non-English annotation schemas, cross-lingual text-to-AMR parsing, and AMR-to-(non-English) text generation. We advance prior work on cross-lingual AMR by thoroughly investigating the amount, types, and causes of differences that appear in AMRs of different languages. Further, we compare how AMR captures meaning in cross-lingual pairs versus strings, and show that AMR graphs are able to draw out fine-grained differences between parallel sentences. We explore three primary research questions: (1) What are the types and causes of differences in parallel AMRs? (2) How can we measure the amount of difference between AMR pairs in different languages? (3) Given that AMR structure is affected by language and exhibits cross-lingual differences, how do cross-lingual AMR pairs compare to string-based representations of cross-lingual sentence pairs? We find that the source language itself does have a measurable impact on AMR structure, and that translation divergences and annotator choices also lead to differences in cross-lingual AMR pairs. We explore the implications of this finding throughout our study, concluding that, although AMR is useful to capture meaning across languages, evaluations need to take into account source language influences if they are to paint an accurate picture of system output, and meaning generally.
Shira Wein, Nathan Schneider 0001
Comput. Linguistics2
2023 ELQA: A Corpus of Metalinguistic Questions and Answers about English
abstract
We present ELQA, a corpus of questions and answers in and about the English language.Collected from two online forums, the >70k questions (from English learners and others) cover wide-ranging topics including grammar, meaning, fluency, and etymology.The answers include descriptions of general properties of English vocabulary and grammar as well as explanations about specific (correct and incorrect) usage examples.Unlike most NLP datasets, this corpus is metalinguistic-it consists of language about language.As such, it can facilitate investigations of the metalinguistic capabilities of NLU models, as well as educational applications in the language learning domain.To study this, we define a free-form question answering task on our dataset and conduct evaluations on multiple LLMs (Large Language Models) to analyze their capacity to generate metalinguistic answers.
Shabnam Behzad, Keisuke Sakaguchi, Nathan Schneider 0001, Amir Zeldes
ACL (1)3
2023 Syntactic Inductive Bias in Transformer Language Models: Especially Helpful for Low-Resource Languages?
abstract
A line of work on Transformer-based language models such as BERT has attempted to use syntactic inductive bias to enhance the pretraining process, on the theory that building syntactic structure into the training process should reduce the amount of data needed for training.But such methods are often tested for highresource languages such as English.In this work, we investigate whether these methods can compensate for data sparseness in lowresource languages, hypothesizing that they ought to be more effective for low-resource languages.We experiment with five low-resource languages: Uyghur, Wolof, Maltese, Coptic, and Ancient Greek.We find that these syntactic inductive bias methods produce uneven results in low-resource settings, and provide surprisingly little benefit in most cases.
Luke Gessler, Nathan Schneider 0001
CoNLL2
2022 Accounting for Language Effect in the Evaluation of Cross-lingual AMR Parsers
abstract
Cross-lingual Abstract Meaning Representation (AMR) parsers are currently evaluated in comparison to gold English AMRs, despite parsing a language other than English, due to the lack of multilingual AMR evaluation metrics. This evaluation practice is problematic because of the established effect of source language on AMR structure. In this work, we present three multilingual adaptations of monolingual AMR evaluation metrics and compare the performance of these metrics to sentence-level human judgments. We then use our most highly correlated metric to evaluate the output of state-of-the-art cross-lingual AMR parsers, finding that Smatch may still be a useful metric in comparison to gold English AMRs, while our multilingual adaptation of S2match (XS2match) is best for comparison with gold in-language AMRs.
Shira Wein, Nathan Schneider 0001
COLING2
2022 MASALA: Modelling and Analysing the Semantics of Adpositions in Linguistic Annotation of Hindi
abstract
We present a completed, publicly available corpus of annotated semantic relations of adpositions and case markers in Hindi. We used the multilingual SNACS annotation scheme, which has been applied to a variety of typologically diverse languages. Building on past work examining linguistic problems in SNACS annotation, we use language models to attempt automatic labelling of SNACS supersenses in Hindi and achieve results competitive with past work on English. We look towards upstream applications in semantic role labelling and extension to related languages such as Gujarati.
Aryaman Arora, Nitin Venkateswaran, Nathan Schneider 0001
LREC3
2022 Xposition: An Online Multilingual Database of Adpositional Semantics
abstract
We present Xposition, an online platform for documenting adpositional semantics across languages in terms of supersenses (Schneider et al., 2018). More than just a lexical database, Xposition houses annotation guidelines, structured lexicographic documentation, and annotated corpora. Guidelines and documentation are stored as wiki pages for ease of editing, and described elements (supersenses, adpositions, etc.) are hyperlinked for ease of browsing. We describe how the platform structures information; its current contents across several languages; and aspects of the design of the web application that supports it, with special attention to how it supports datasets and standards that evolve over time.
Luke Gessler, Nathan Schneider 0001, Joseph C. Ledford, Austin Blodgett
LREC2
2022 DocAMR: Multi-Sentence AMR Representation and Evaluation
abstract
Tahira Naseem, Austin Blodgett, Sadhana Kumaravel, Tim O’Gorman, Young-Suk Lee, Jeffrey Flanigan, Ramón Astudillo, Radu Florian, Salim Roukos, Nathan Schneider. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Tahira Naseem, Austin Blodgett, Sadhana Kumaravel, Tim O'Gorman, Young-Suk Lee 0001, Jeffrey Flanigan, Ramón Fernandez Astudillo, Radu Florian, Salim Roukos, Nathan Schneider 0001
NAACL-HLT10
2022 Linguistic Frameworks Go Toe-to-Toe at Neuro-Symbolic Language Modeling
abstract
We examine the extent to which, in principle, different syntactic and semantic graph representations can complement and improve neural language modeling.Specifically, by conditioning on a subgraph encapsulating the locally relevant sentence history, can a model make better next-word predictions than a pretrained sequential language model alone?With an ensemble setup consisting of GPT-2 and ground-truth graphs from one of 7 different formalisms, we find that the graph information indeed improves perplexity and other metrics.Moreover, this architecture provides a new way to compare different frameworks of linguistic representation.In our oracle graph setup, training and evaluating on English WSJ, semantic constituency structures prove most useful to language modeling performance-outpacing syntactic constituency structures as well as syntactic and semantic dependency structures.
Jakob Prange, Nathan Schneider 0001, Lingpeng Kong
NAACL-HLT2
2021 Probabilistic, Structure-Aware Algorithms for Improved Variety, Accuracy, and Coverage of AMR Alignments
abstract
Austin Blodgett, Nathan Schneider. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Austin Blodgett, Nathan Schneider 0001
ACL/IJCNLP (1)2
2021 Putting Words in BERT's Mouth: Navigating Contextualized Vector Spaces with Pseudowords
abstract
We present a method for exploring regions around individual points in a contextualized vector space (particularly, BERT space), as a way to investigate how these regions correspond to word senses.By inducing a contextualized "pseudoword" as a stand-in for a static embedding in the input layer, and then performing masked prediction of a word in the sentence, we are able to investigate the geometry of the BERT-space in a controlled manner around individual instances.Using our method on a set of carefully constructed sentences targeting ambiguous English words, we find substantial regularity in the contextualized space, with regions that correspond to distinct word senses; but between these regions there are occasionally "sense voids"-regions that do not correspond to any intelligible sense. 1Learn pseudoword in place of that is customized to reconstruct .
Taelin Karidi, Yichu Zhou, Nathan Schneider 0001, Omri Abend, Vivek Srikumar
EMNLP (1)3
2021 Supertagging the Long Tail with Tree-Structured Decoding of Complex Categories
abstract
Abstract Although current CCG supertaggers achieve high accuracy on the standard WSJ test set, few systems make use of the categories’ internal structure that will drive the syntactic derivation during parsing. The tagset is traditionally truncated, discarding the many rare and complex category types in the long tail. However, supertags are themselves trees. Rather than give up on rare tags, we investigate constructive models that account for their internal structure, including novel methods for tree-structured prediction. Our best tagger is capable of recovering a sizeable fraction of the long-tail supertags and even generates CCG categories that have never been seen in training, while approximating the prior state of the art in overall tag accuracy with fewer parameters. We further investigate how well different approaches generalize to out-of-domain evaluation sets.
Jakob Prange, Nathan Schneider 0001, Vivek Srikumar
Trans. Assoc. Comput. Linguistics2
2020 Supervised Grapheme-to-Phoneme Conversion of Orthographic Schwas in Hindi and Punjabi
abstract
Hindi grapheme-to-phoneme (G2P) conversion is mostly trivial, with one exception: whether a schwa represented in the orthography is pronounced or unpronounced (deleted).Previous work has attempted to predict schwa deletion in a rule-based fashion using prosodic or phonetic analysis.We present the first statistical schwa deletion classifier for Hindi, which relies solely on the orthography as the input and outperforms previous approaches.We trained our model on a newly-compiled pronunciation lexicon extracted from various online dictionaries.Our best Hindi model achieves state of the art performance, and also achieves good performance on a closely related language, Punjabi, without modification.
Aryaman Arora, Luke Gessler, Nathan Schneider 0001
ACL3
2020 (Re)construing Meaning in NLP
abstract
Human speakers have an extensive toolkit of ways to express themselves.In this paper, we engage with an idea largely absent from discussions of meaning in natural language understanding-namely, that the way something is expressed reflects different ways of conceptualizing or construing the information being conveyed.We first define this phenomenon more precisely, drawing on considerable prior work in theoretical cognitive semantics and psycholinguistics.We then survey some dimensions of construed meaning and show how insights from construal could inform theoretical and practical work in NLP.
Sean Trott, Tiago Timponi Torrent, Nancy Chang, Nathan Schneider 0001
ACL4
2020 Comparison by Conversion: Reverse-Engineering UCCA from Syntax and Lexical Semantics
abstract
Building robust natural language understanding systems will require a clear characterization of whether and how various linguistic meaning representations complement each other.To perform a systematic comparative analysis, we evaluate the mapping between meaning representations from different frameworks using two complementary methods: (i) a rule-based converter, and (ii) a supervised delexicalized parser that parses to one framework using only information from the other as features.We apply these methods to convert the STREUSLE corpus (with syntactic and lexical semantic annotations) to UCCA (a graph-structured full-sentence meaning representation).Both methods yield surprisingly accurate target representations, close to fully supervised UCCA parser quality-indicating that UCCA annotations are partially redundant with STREUSLE annotations.Despite this substantial convergence between frameworks, we find several important areas of divergence.
Daniel Hershcovich, Nathan Schneider 0001, Dotan Dvir, Jakob Prange, Miryam de Lhoneux, Omri Abend
COLING2
2020 A Human Evaluation of AMR-to-English Generation Systems
abstract
Most current state-of-the art systems for generating English text from Abstract Meaning Representation (AMR) have been evaluated only using automated metrics, such as BLEU, which are known to be problematic for natural language generation.In this work, we present the results of a new human evaluation which collects fluency and adequacy scores, as well as categorization of error types, for several recent AMR generation systems.We discuss the relative quality of these systems and how our results compare to those of automatic metrics, finding that while the metrics are mostly successful in ranking systems overall, collecting human judgments allows for more nuanced comparisons.We also analyze common errors made by these systems.
Emma Manning, Shira Wein, Nathan Schneider 0001
COLING3
2020 A Corpus of Adpositional Supersenses for Mandarin Chinese
abstract
Adpositions are frequent markers of semantic relations, but they are highly ambiguous and vary significantly from language to language. Moreover, there is a dearth of annotated corpora for investigating the cross-linguistic variation of adposition semantics, or for building multilingual disambiguation systems. This paper presents a corpus in which all adpositions have been semantically annotated in Mandarin Chinese; to the best of our knowledge, this is the first Chinese corpus to be broadly annotated with adposition semantics. Our approach adapts a framework that defined a general set of supersenses according to ostensibly language-independent semantic criteria, though its development focused primarily on English prepositions (Schneider et al., 2018). We find that the supersense categories are well-suited to Chinese adpositions despite syntactic differences from English. On a Mandarin translation of The Little Prince, we achieve high inter-annotator agreement and analyze semantic correspondences of adposition tokens in bitext.
Siyao Peng, Yang Liu 0213, Yilun Zhu 0001, Austin Blodgett, Yushi Zhao, Nathan Schneider 0001
LREC6
2019 Made for Each Other: Broad-Coverage Semantic Structures Meet Preposition Supersenses
abstract
Universal Conceptual Cognitive Annotation (UCCA; Abend and Rappoport, 2013) is a typologically-informed, broad-coverage semantic annotation scheme that describes coarse-grained predicate-argument structure but currently lacks semantic roles. We argue that lexicon-free annotation of the semantic roles marked by prepositions, as formulated by Schneider et al. (2018), is complementary and suitable for integration within UCCA. We show empirically for English that the schemes, though annotated independently, are compatible and can be combined in a single semantic graph. A comparison of several approaches to parsing the integrated representation lays the groundwork for future research on this task.
Jakob Prange, Nathan Schneider 0001, Omri Abend
CoNLL2
2018 Comprehensive Supersense Disambiguation of English Prepositions and Possessives
abstract
Nathan Schneider, Jena D. Hwang, Vivek Srikumar, Jakob Prange, Austin Blodgett, Sarah R. Moeller, Aviram Stern, Adi Bitan, Omri Abend. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Nathan Schneider 0001, Jena D. Hwang, Vivek Srikumar, Jakob Prange, Austin Blodgett, Sarah R. Moeller, Aviram Stern, Adi Bitan, Omri Abend
ACL (1)1
2018 Discourse Coherence: Concurrent Explicit and Implicit Relations
abstract
Theories of discourse coherence posit relations between discourse segments as a key feature of coherent text.Our prior work suggests that multiple discourse relations can be simultaneously operative between two segments for reasons not predicted by the literature.Here we test how this joint presence can lead participants to endorse seemingly divergent conjunctions (e.g., but and so) to express the link they see between two segments.These apparent divergences are not symptomatic of participant naïveté or bias, but arise reliably from the concurrent availability of multiple relations between segments -some available through explicit signals and some via inference.We believe that these new results can both inform future progress in theoretical work on discourse coherence and lead to higher levels of performance in discourse parsing.
Hannah Rohde, Alexander Johnson, Nathan Schneider 0001, Bonnie L. Webber
ACL (1)3
2018 Detecting and Using Buzz from Newspapers to Understand Patterns of Movement
abstract
Meaningful leading indicators of mass movement are difficult to discover given the dearth of available data about involuntary movement. As a first step, we propose analyzing whether we can use the changing dynamics of newspaper content as one possible indirect indicator of such displacement. Specifically, we explore whether news media buzz correlates with patterns of migration in Iraq. We consider different methods for detecting buzz and empirically evaluate them on a corpus of 1.4 million articles.
Julia Hocket, Yaguang Liu, Yifang Wei, Lisa Singh, Nathan Schneider 0001
IEEE BigData5
2018 Semantic Supersenses for English Possessives
Austin Blodgett, Nathan Schneider 0001
LREC2
2018 Abstract Meaning Representation of Constructions: The More We Include, the Better the Representation
Claire Bonial, Bianca Badarau, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Tim O'Gorman, Martha Palmer, Nathan Schneider 0001
LREC8
2018 Parsing Tweets into Universal Dependencies
abstract
Yijia Liu, Yi Zhu, Wanxiang Che, Bing Qin, Nathan Schneider, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Wanxiang Che, Bing Qin 0001, Nathan Schneider 0001, Noah A. Smith
NAACL-HLT5
2018 A Structured Syntax-Semantics Interface for English-AMR Alignment
abstract
Ida Szubert, Adam Lopez, Nathan Schneider. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Ida Szubert, Adam Lopez, Nathan Schneider 0001
NAACL-HLT3
2016 Inconsistency Detection in Semantic Annotation
Nora Hollenstein, Nathan Schneider 0001, Bonnie L. Webber
LREC2
2015 Big Data Small Data, In Domain Out-of Domain, Known Word Unknown Word: The Impact of Word Representations on Sequence Labelling Tasks
abstract
Word embeddings -distributed word representations that can be learned from unlabelled data -have been shown to have high utility in many natural language processing applications.In this paper, we perform an extrinsic evaluation of four popular word embedding methods in the context of four sequence labelling tasks: part-of-speech tagging, syntactic chunking, named entity recognition, and multiword expression identification.A particular focus of the paper is analysing the effects of task-based updating of word representations.We show that when using word embeddings as features, as few as several hundred training instances are sufficient to achieve competitive results, and that word embeddings lead to improvements over out-of-vocabulary words and also out of domain.Perhaps more surprisingly, our results indicate there is little difference between the different word embedding methods, and that simple Brown clusters are often competitive with word embeddings across all tasks we consider.
Lizhen Qu, Gabriela Ferraro, Liyuan Zhou, Weiwei Hou, Nathan Schneider 0001, Timothy Baldwin
CoNLL5
2015 Getting the Roles Right: Using FrameNet in NLP
abstract
Collin F. Baker, Nathan Schneider, Miriam R. L. Petruck, Michael Ellsworth. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Tutorial Abstracts. 2015.
Collin F. Baker, Nathan Schneider 0001, Miriam R. L. Petruck, Michael Ellsworth
HLT-NAACL2
2015 The Logic of AMR: Practical, Unified, Graph-Based Sentence Semantics for NLP
abstract
Meaning Representation formalism is rapidly emerging as an important practical form of structured sentence semantics which, thanks to the availability of largescale annotated corpora, has potential as a convergence point for NLP research.This tutorial unmasks the design philosophy, data creation process, and existing algorithms for AMR semantics.It is intended for anyone interested in working with AMR data, including parsing text into AMRs, generating text from AMRs, and applying AMRs to tasks such as machine translation and summarization.
Nathan Schneider 0001, Jeffrey Flanigan, Tim O'Gorman
HLT-NAACL1
2015 A Corpus and Model Integrating Multiword Expressions and Supersenses
abstract
This paper introduces a task of identifying and semantically classifying lexical expressions in running text.We investigate the online reviews genre, adding semantic supersense annotations to a 55,000 word English corpus that was previously annotated for multiword expressions.The noun and verb supersenses apply to full lexical expressions, whether single-or multiword.We then present a sequence tagging model that jointly infers lexical expressions and their supersenses.Results show that even with our relatively small training corpus in a noisy domain, the joint task can be performed to attain 70% class labeling F 1 .
Nathan Schneider 0001, Noah A. Smith
HLT-NAACL1
2014 Automatic Classification of Communicative Functions of Definiteness
Archna Bhatia, Chu-Cheng Lin, Nathan Schneider 0001, Yulia Tsvetkov, Fatima Talib Al-Raisi, Laleh Roostapour, Jordan Bender, Abhimanu Kumar, Lori S. Levin, Mandy Simons, Chris Dyer
COLING3
2014 A Dependency Parser for Tweets
abstract
We describe a new dependency parser for English tweets, TWEEBOPARSER. The parser builds on several contributions: new syntactic annotations for a corpus of tweets (TWEEBANK), with conventions informed by the domain; adaptations to a statistical parsing algorithm; and a new approach to exploiting out-of-domain Penn Treebank data. Our experiments show that the parser achieves over 80% unlabeled attachment accuracy on our new, high-quality test set and measure the benefit of our contributions. Our dataset and parser can be found at http://www.ark.cs.cmu.edu/TweetNLP.
Lingpeng Kong, Nathan Schneider 0001, Swabha Swayamdipta, Archna Bhatia, Chris Dyer, Noah A. Smith
EMNLP2
2014 Comprehensive Annotation of Multiword Expressions in a Social Web Corpus
Nathan Schneider 0001, Spencer Onuffer, Nora Kazour, Emily Danchik, Michael T. Mordowanec, Henrietta Conrad, Noah A. Smith
LREC1
2014 Augmenting English Adjective Senses with Supersenses
Yulia Tsvetkov, Nathan Schneider 0001, Dirk Hovy, Archna Bhatia, Manaal Faruqui, Chris Dyer
LREC2
2014 Frame-Semantic Parsing
abstract
Frame semantics is a linguistic theory that has been instantiated for English in the FrameNet lexicon. We solve the problem of frame-semantic parsing using a two-stage statistical model that takes lexical targets (i.e., content words and phrases) in their sentential contexts and predicts frame-semantic structures. Given a target in context, the first stage disambiguates it to a semantic frame. This model uses latent variables and semi-supervised learning to improve frame disambiguation for targets unseen at training time. The second stage finds the target's locally expressed semantic arguments. At inference time, a fast exact dual decomposition algorithm collectively predicts all the arguments of a frame at once in order to respect declaratively stated linguistic constraints, resulting in qualitatively better structures than naïve local predictors. Both components are feature-based and discriminatively trained on a small set of annotated frame-semantic parses. On the SemEval 2007 benchmark data set, the approach, along with a heuristic identifier of frame-evoking targets, outperforms the prior state of the art by significant margins. Additionally, we present experiments on the much larger FrameNet 1.5 data set. We have released our frame-semantic parser as open-source software.
Dipanjan Das 0001, Desai Chen, André F. T. Martins, Nathan Schneider 0001, Noah A. Smith
Comput. Linguistics4
2014 Discriminative Lexical Semantic Segmentation with Gaps: Running the MWE Gamut
abstract
We present a novel representation, evaluation measure, and supervised models for the task of identifying the multiword expressions (MWEs) in a sentence, resulting in a lexical semantic segmentation. Our approach generalizes a standard chunking representation to encode MWEs containing gaps, thereby enabling efficient sequence tagging algorithms for feature-rich discriminative models. Experiments on a new dataset of English web text offer the first linguistically-driven evaluation of MWE identification with truly heterogeneous expression types. Our statistical sequence model greatly outperforms a lookup-based segmentation procedure, achieving nearly 60% F1 for MWE identification.
Nathan Schneider 0001, Emily Danchik, Chris Dyer, Noah A. Smith
Trans. Assoc. Comput. Linguistics1
2013 Improved Part-of-Speech Tagging for Online Conversational Text with Word Clusters
Olutobi Owoputi, Brendan T. O'Connor 0001, Chris Dyer, Kevin Gimpel, Nathan Schneider 0001, Noah A. Smith
HLT-NAACL5
2013 Supersense Tagging for Arabic: the MT-in-the-Middle Attack
Nathan Schneider 0001, Behrang Mohit, Chris Dyer, Kemal Oflazer, Noah A. Smith
HLT-NAACL1
2013 Design Patterns in Fluid Construction Grammar Luc Steels (editor) Universitat Pompeu Fabra and Sony Computer Science Laboratory, Paris Amsterdam: John Benjamins Publishing Company (Constructional Approaches to Language series, edited by Mirjam Fried and Jan-Ola Östman, volume 11), 2012, xi+332 pp; hardbound, ISBN 978-90-272-0433-2, €99.00, $149.00
abstract
In computational modeling of natural language phenomena, there are at least three modes of research. The currently dominant statistical paradigm typically prioritizes instance coverage: Data-driven methods seek to use as much information observed in data as possible in order to generalize linguistic analyses to unseen instances. A second approach prioritizes detailed description of grammatical phenomena, that is, forming and defending theories with a focus on a small number of instances. A third approach might be called integrative: Rather than addressing phenomena in isolation, different approaches are brought together to address multiple challenges in a unified framework, and the behavior of the system is demonstrated with a small number of instances. Design Patterns in Fluid Construction Grammar (DPFCG) exemplifies the third approach, introducing a linguistic formalism called Fluid Construction Grammar (FCG) that addresses parsing, production, and learning in a single computational framework.The book emphasizes grammar-engineering, following broad-coverage descriptive paradigms that can be traced back to Generalized Phrase Structure Grammar (GPSG) Gazdar et al. (1985), Lexical Functional Grammar (LFG) Bresnan (2000), Head-Driven Phrase-Structure Grammar (HPSG) Sag and Wasow (1999), and Combinatory Categorial Grammar (CCG) Steedman (1996). In all of these cases, a formal meta-framework allows computational linguists to formalize their hypotheses and intuitions about a language's grammatical behavior and then explore how these representational choices affect the processing of natural language utterances. Many of the aforementioned approaches have engendered large-scale platforms that can be used and reused to provide formal description of grammars for different languages, such as Par-Gram for LFG Butt et al. (2002) and the LinGO Grammar Matrix for HPSG Bender, Flickinger, and Oepen (2002).FCG offers a similar grammar engineering framework that follows the principles of Construction Grammar (CxG) Goldberg (2003; Hoffmann and Trousdale (2013). CxG treats constructions as the basic units of grammatical organization in language. The constructions are viewed as learned associations between form (e.g., sounds, morphemes, syntactic phrases) and function (semantics, pragmatics, discourse meaning, etc.). CxG does not impose a strict separation between lexicon and grammar—indeed, it is perhaps best known as treating semi-productive idioms like “the X-er, the Y-er” and “X let alone Y” on equal footing with lexemes and “core” syntactic patterns Fillmore, Kay, and O'Connor (1988; Kay and Fillmore (1999). FCG, like other CxG formalisms—namely, Embodied Construction Grammar Bergen and Chang (2005; Feldman, Dodge, and Bryant (2009) and Sign-Based Construction Grammar Boas and Sag (2012)—is unification-based.1 The studies in this book describe constructions and how they can be combined in order to model natural language interpretation or generation as feature structure unification in a general search procedure.The book has five parts, covering the groundwork, basic linguistic applications, processing matters, advanced case studies, and, finally, features of FCG that make it fluid and robust. Each chapter identifies general strategies (design patterns) that might merit reuse in new FCG grammars, or perhaps in other computational frameworks.Part I: Introduction lays the groundwork for the rest of the book. “Introducing Fluid Construction Grammar” (by Luc Steels) presents the aims of the FCG formalism. FCG was designed as a framework for describing linguistic units (constructions—their form and meaning), with an emphasis on language variation and evolution (“fluidity”). The constructionist approach to language is described and the argument for applying it to study language variation and change is defended. Psychological validity is explicitly ruled out as a modeling goal (“The emphasis is on getting working systems, and this is difficult enough”; page 4). The architects of FCG set out to include both sides of the processing coin, however—parsing (interpretation) and production (generation). The concept of search in processing is emphasized, though some of the explanations of processing steps are too abstract for the reader to comprehend at this point. A further desideratum—robustness to noisy input containing disfluencies, fragments, and errors—is given as motivating a constructionist approach.The next chapter, “A First Encounter with Fluid Construction Grammar” (by Steels), describes the mechanisms of FCG in detail. In FCG, a working analysis hypothesized in processing is known as transient structure; the transduction of form to meaning (and vice versa) selects a sequence of constructions that apply to the transient structure to gradually expand it until reaching a final analysis. Identifying constructions that may apply to a transient structure presents a non-trivial search problem, also addressed by the architects of FCG. The sheer number of technical details make this chapter somewhat overwhelming. Most of the chapter is devoted to the low-level feature structures and the operations manipulating them. Templates—a practical means of avoiding boilerplate code when defining constructions—are then introduced, and do most of the heavy lifting in the rest of the book.In Part II: Grammatical Structures, we begin to see how constructions are defined in practice. “A Design Pattern for Phrasal Constructions” (by Steels) illustrates how constructions are used to describe the combination of multiple units into higher-level, typed phrases. Skeletal constructions compose with smaller units to form hierarchical structures (essentially similar to the Immediate Constituents Analysis of (Bloomfield 1933) and follow-up work in structuralist linguistics [Harris 1946]), and a range of additional constructions impose form (e.g., ordering) constraints and add new meaning to the newly created phrases. This chapter is of a tutorial nature, illustrating the step-by-step application of four kinds of noun phrase constructions to expand transient structures in processing. Over twenty templates are introduced in this chapter; they encapsulate design patterns dealing with hierarchical structure, agreement, and feature percolation. An aspect of phrasal constructions that is not yet dealt with is the complex linking of semantic arguments and the morphosyntactic categorizations of the composed elements.“A Design Pattern for Argument Structure Constructions” (by Remi van Trijp) then builds on the formal machinery presented in the previous chapter to explicitly address the complex mappings between semantic arguments (agent, patient, etc.) and syntactic arguments (subject, object, etc.). This mapping is a complex matter due to language-specific conceptualization of semantic arguments and different means of morphosyntactic realization used by different languages. In FCG, each lexical item introduces its linking potential in terms of the different types of semantic and syntactic arguments that it may take, with no particular mapping between them. Each argument structure construction imposes a partial mapping between the syntactic and semantic arguments to yield a particular argument structure instantiation, one of the multiple alternatives that may be available for a single lexical item. This account stands in sharp contrast to the lexicalist view of argument-structure (the view taken in LFG, HPSG, and CCG) whereby each lexical entry dictates all the necessary linking information. The construction-based approach is defended for its ability to deal with unknown words2 and constructional coercion3 Goldberg (1995). The argument structure design pattern allows FCG to crudely recover a partial specification of the form-meaning mapping of these elements, which is important for robust processing (see subsequent discussion).Part III: Managing Processing addresses how FCG transduces between a linguistic string and a meaning representation, where the two directions (parsing and production) share a common declarative representation of linguistic knowledge (the grammar). This entails assembling an analysis incrementally on the basis of the grammar, the input, and any partial analysis that has already been created. With FCG (and unification grammars more broadly) this search is nontrivial, and streamlining search (i.e., minimizing nondeterminism and avoiding dead ends) is a key motivator of many of the grammar design patterns suggested in the book.“Search in Linguistic Processing” (by Joris Bleys, Kevin Stadler, and Joachim De Beule) deals mainly with the problem of choosing which of multiple compatible constructions to apply next. Whereas the default heuristic search in FCG is a greedy, depth-first search (which can backtrack if the user-defined end-goal has not yet been achieved) the FCG framework allows for a guided search through scores that reflect the relative tendency of a construction to apply next. The authors suggest that such scoring can be informed by general principles, for instance: (i) specific constructions are preferred to more general ones, and (ii) previously co-applied constructions are preferred. Choosing appropriate constructions to apply early on dramatically reduces the time needed for processing the utterance.“Organizing Constructions in Networks” (by Pieter Wellens) takes this idea to the next level, and proposes to organize the different constructions in networks of conditional dependencies. A conditional dependency links two constructions where one provides necessary information for the application of the other. These dependency networks can be updated whenever an input is processed so that the system learns to search more efficiently when the same constructions are encountered in the future. Using these networks to guide the search thus significantly reduces the search for compatible constructions. An empirical effort to quantify this effect indeed shows a sharp reduction in search time; unlike the held-out experimental paradigm accepted in statistical NLP, however, the parsed/produced sentence is assumed to have been seen already by the system.Part IV: Case Studies addresses three challenging linguistic phenomena in FCG. “Feature Matrices and Agreement” (by van Trijp) on German case offers a new unification-based solution to the problem of feature indeterminacy. For instance, in the sentence Er findet und hilft Frauen ‘He finds and helps women’, the first verb requires an accusative object, whereas the second requires a dative object; the coordination is allowed only because Frauen can be either accusative or dative. Kindern ‘children’, which can only be dative, is not licensed here. Encoding case in a single feature on the Frauen construction wouldn't work because the feature would have to unify with contradictory values (from the verbs' inflectional features). Instead, case restrictions specified lexically for a noun or verb can be expressed with a distinctive feature matrix, with each matrix slot holding a variable or the value + or -. Unification then does the right thing—allowing Frauen and forbidding Kindern—without resorting to type hierarchies or disjunctive features.“Construction Sets and Unmarked Forms” (by Katrien Beuls) on Hungarian verbal agreement models a phenomenon whereby morphosyntactic, semantic, and phonological factors affect the choice between poly- and mono-personal agreement—that is, the decision whether a Hungarian transitive verb should agree with its object or just with its subject. The case and definiteness of the object and the person hierarchy relationship between subject and object determine which kind of agreement obtains, and phonological constraints determine its form. To make the different levels of structure interact properly, constructions are grouped into sets (lexical, morphological, etc.) and those sets are considered in a fixed order during processing. Construction sets also allow for efficient handling of unmarked forms (null affixes)—they are considered only after the overt affixes have had the opportunity to apply, thereby functioning as defaults.“Syntactic Indeterminacy and Semantic Ambiguity” (by Michael Spranger and Martin Loetzsch) on German spatial phrases models the German spatial terms for front, back, left, and right. To model spatial language in situated interaction with robots, two problems must be overcome. The first is syntactic indeterminacy: Any of these spatial relations may be realized as an adjective, an adverb, or a preposition. The second is semantic ambiguity, specifically when the perspective (e.g., whose ‘left’?) is implicit. Both are forms of underspecification which could cause early splits in the search space if handled naïvely. Much in the spirit of the argument structure constructions (see above), the solutions (which are too technical to explain here) involve (a) disjunctive representations of potential values of a feature, and (b) deferring decisions until a more opportune stage.Part V: Robustness and Fluidity (by Steels and van Trijp) surveys the different features of the system that ensure robustness in the face of variation, disfluencies, and noise. Natural language is fluid and open-ended. There is variation between speakers, there are disfluencies and speech errors, and noise may corrupt the speech signal. All of these may jeopardize the interpretability of the signal, but human listeners are adept at processing such input. In the spirit of usage-based grammar Tomasello (2003), FCG emphasizes the attainment of a communicative goal, rather than ensuring grammaticality of parsed/produced utterances. This is accomplished with a diagnostic-repair process that runs in parallel to parsing/production. Diagnostics can test for unknown words, unfamiliar meanings, missing constructions, and so on. Diagnostic tests are implemented by reversing the direction of the transduction process: a speaker may assume the hearer's point of view to analyze what she has produced in order to see whether communicative success has been attained. Likewise, a hearer may produce a phrase according to his own interpretation of the speaker's form, and check for a match. If a test fails, repair strategies such as proposing new constructions, relaxing the matching process for construction application, and coercing constructions to adapt to novel language use are considered.Fluidity and robustness are the hallmarks of FCG, and the computational framework has been used in experiments that assume embedded communication in robotic agents. This research program is developed at length by Steels (2012b).Discussion. Like the legacy of the GPSG book Gazdar et al. (1985), this book's main merit is not necessarily in its technical details or computational choices, but in demonstrating the feasibility of implementing the constructional approach in a full-fledged computational framework. We suggest that the CxG perspective presents a formidable challenge to the computational linguistics/natural language processing community. It posits a different notion of modularity than is observed by most NLP systems: Rather than treat different levels of linguistic structure independently, CxG recognizes that multiple formal components (phonological, lexical, morphological, syntactic) may be tied by convention to a specific meaning or function. Systematically describing these “cross-cutting” constructions and their processing, especially in a way that scales to large data encompassing both form and meaning and accommodates both parsing and generation, would in our view make for a more comprehensive account of language processing than our field is able to offer today. Thus, we hope this book will be provocative even outside of the grammar engineering community.This book is not without its weaknesses. In parts the writing is quite technical and terse, which can be daunting for readers new to FCG. Contextualization with respect to other strands of computational linguistics and AI research is, for the most part, lacking, though a second FCG book Steels (2012a) picks up some of the slack on this front.4DPFCG does not address the feasibility of learning constructions directly from data,5 nor does it discuss the expressive power of the formalism in relation to learnability results (such as that of Gold [1967]). As admitted by the authors, much more work would be needed to build life-size grammars. Still, we hope that readers of DPFCG will appreciate the authors' vision for a model of linguistic form and function that is at once formal, computational, fluid, and robust.
Nathan Schneider 0001, Reut Tsarfaty
Comput. Linguistics1
2012 Recall-Oriented Learning of Named Entities in Arabic Wikipedia
Behrang Mohit, Nathan Schneider 0001, Rishav Bhowmick, Kemal Oflazer, Noah A. Smith
EACL2
2010 Probabilistic Frame-Semantic Parsing
Dipanjan Das 0001, Nathan Schneider 0001, Desai Chen, Noah A. Smith
HLT-NAACL2